Calculus for ML
Formulas are drawn by KaTeX, which loads from the internet. You seem to be offline, so formulas show as plain text for now.
A beginner-friendly, interactive guide

Calculus for Machine Learning

Learn the maths of change from zero. Every idea starts with a plain-English picture, then a worked example, then the formal definition, and then something you can drag, slide, rotate and play with. Everything points toward one goal: understanding how a machine-learning model learns.

What is calculus, in one sentence? It is the maths of change: how fast something is changing right now (the derivative), and how little changes add up to a total (the integral).

Why care? A model learns by repeating one small move: "if I nudge this number a tiny bit, does the error go up or down, and by how much?" That question is a derivative. Training a neural network is that question asked millions of times. If you understand it, gradient descent and backpropagation stop being magic.

Companion guide. This guide goes hand in hand with the Linear Algebra guide. You can follow both side by side: whenever a vector or matrix appears, we remind you what it is, and link to the exact place that explains it.

How every topic is taught

Each concept follows the same six steps, in this order. The order matters: understanding comes before formulas.

1 · Intuition

A picture or everyday story. No symbols yet. If you only read this part, you will already get the idea.

2 · Example

A small problem with real numbers, solved slowly, step by step. You can redo it with pen and paper.

3 · Definition

Now the precise statement and the notation. Because you have seen the idea, the symbols just give it a name.

4 · Play

An interactive visual. Drag points, move sliders and watch the numbers change. In 3D you can also rotate the whole scene. Each one tells you what to try.

5 · Why · Where · How

Every concept answers three questions in plain English: Why do we need it? Where is it used? How is it used? So you always know why the idea is worth your time.

6 · Check

Short questions to test yourself, a recap of the key points, a NumPy code block to type out, and practice problems with worked answers.

Try your first interactive

The boxes marked Interactive react to you. Drag the dot along the curve and watch the straight line that just touches it.

Drag the dot left and right. The straight line that just touches the curve is the tangent. Its steepness is the derivative: positive when the curve climbs, negative when it falls, and exactly zero at the top of a hill or the bottom of a valley. That zero is where a model's error is lowest, and where training wants to go.

The roadmap

Fourteen chapters, in an order where each one builds on the last. Click any card to jump there.

What matters most for ML. Spend your effort on chapters 2.3 to 2.12: derivatives, gradients, the Jacobian, the chain rule, backpropagation, the Hessian and Taylor expansions. Limits (2.2) only need to be understood well enough to understand a derivative, and the vector-calculus tools in 2.13 (divergence, curl, line and surface integrals) are for awareness.

How to read the maths symbols

You will meet these symbols again and again. Do not memorise them now. Come back to this table whenever one looks strange.

SymbolSay it asMeaning
$f(x)$"f of x"A rule that turns the number $x$ into another number.
$\lim_{x \to a} f(x)$"the limit of f as x approaches a"The value $f(x)$ gets closer and closer to as $x$ gets close to $a$.
$f'(x)$, $\dfrac{df}{dx}$"f prime", "d f d x"The derivative: how fast $f$ changes as $x$ changes.
$\dfrac{\partial f}{\partial x}$"partial f partial x"How fast $f$ changes when only $x$ changes and everything else is held still.
$\nabla f$"del f", "the gradient of f"The list of all partial derivatives. It points uphill.
$J$"the Jacobian"The matrix of all partial derivatives of a vector-valued function.
$H$, $\nabla^2 f$"the Hessian"The matrix of all second derivatives (curvature).
$\displaystyle\int_a^b f(x)\,dx$"the integral of f from a to b"The area under the curve of $f$ between $a$ and $b$.
$\Delta x$, $dx$"delta x", "d x"A small change in $x$ ($dx$ means "infinitely small").
$e$, $\ln x$"e", "natural log of x"$e \approx 2.718$, and $\ln$ is the inverse of $e^x$.
$\mathbf{x}^\top$, $A^\top$"transpose"Flip a vector or matrix (rows ↔ columns).
$\approx$"approximately equal"Close, but not exactly equal.

Tips for studying

  • Play first, read second. If a paragraph confuses you, go to the interactive below it and move things around. Then re-read.
  • Do the examples by hand once. It feels slow, but nothing builds intuition faster.
  • Always ask "what does it mean for the slope?" Almost every derivative idea is a statement about steepness. Keep that picture in your head.
  • Don't rush. One chapter a week is a good pace. Come back to earlier chapters whenever you need to.
  • Use the tools. Press / to search, use the sidebar to jump around, switch dark mode with the button at the top, and mark chapters complete as you go.

If a section feels too hard, it is almost always because a word from an earlier section is fuzzy. Go back, find that word in the sidebar, and reread it. Nobody gets calculus in a single pass. Seeing it twice is normal.

Chapter 2.1

Functions & Mathematical Foundations

Calculus studies how functions change, so before we study change we must be comfortable with functions themselves. This chapter builds your toolbox: the function families you will differentiate in the next chapters, and the three special ones that power neural networks: sigmoid, softmax and ReLU. We start from the simplest picture: a function is a machine.

  • See a function as a machine, a table, a formula and a graph, and know its domain and range
  • Chain functions together (composition) and undo them (inverse functions)
  • Know the families: linear, polynomial, exponential, logarithmic, trigonometric
  • Understand sigmoid and softmax deeply: their shape, saturation and why every classifier uses them
  • Understand piecewise functions such as ReLU, leaky ReLU, absolute value, step, clip and Huber
  • Get a first look at functions of many inputs, which is what a loss function really is

You only need school algebra here. Where an idea from the Linear Algebra guide helps, we say so and link to it, for example vectors and matrices. Every symbol is explained the first time it appears. If you feel rusty, that is fine: go slowly and play with each picture.

What is a function? core

Picture a vending machine. You press one button (the input) and one specific snack drops out (the output). Press the same button tomorrow and you get the same snack. That is all a function is: a rule that turns each input into exactly one output, the same way every time.

The square-root key on a calculator is a function: type 9, press the key, and 3 comes out. Type 16 and 4 comes out.

We can describe one rule in four ways, and they are all the same function: in words ("double it, then add 3"), as a table of inputs and outputs, as a formula, and as a graph, a picture of all the (input, output) pairs. Calculus will mostly work with the graph, because a graph lets us see how fast the output changes.

A taxi fare. A taxi charges 3 dollars to start and 2 dollars for every kilometre. If you ride $x$ kilometres, the fare is "double $x$, then add 3". Call this rule $f$ and write $f(x) = 2x + 3$.

  1. Ride 0 km: $f(0) = 2\cdot 0 + 3 = 3$.
  2. Ride 1 km: $f(1) = 2\cdot 1 + 3 = 5$.
  3. Ride 5 km: $f(5) = 2\cdot 5 + 3 = 13$.
  4. The rule also accepts numbers a taxi would never see: $f(-2) = 2\cdot(-2) + 3 = -4 + 3 = -1$.

Data version. A model that predicts a house price from its area is a function: area goes in, one predicted price comes out.

A function $f$ is a rule that assigns to every allowed input $x$ exactly one output, written $f(x)$ and read "f of x". We write $y = f(x)$.

  • $x$ is the input (the independent variable). $y$ is the output (the dependent variable, because it depends on $x$).
  • "$f : \mathbb{R} \to \mathbb{R}$" says: $f$ takes a real number and gives a real number. ($\mathbb{R}$ is the set of all real numbers, the whole number line.)
  • The graph of $f$ is the set of all points $(x,\, f(x))$ drawn in the plane.
  • Vertical line test. A curve is the graph of a function exactly when no vertical line touches it more than once. (Two touches would mean one input with two outputs.)
  • The letter does not matter: $f(t) = 2t + 3$ is the same function as $f(x) = 2x + 3$.
Why do we need it?

We need a precise way to say "this quantity depends on that one". Calculus asks how fast the output changes when the input changes, and that question only makes sense once we have a function.

Where is it used?

Everywhere in machine learning: a trained model is a function from features to a prediction, a loss is a function from weights to an error number, and an activation (sigmoid, ReLU) is a function applied to every neuron.

How is it used?

Write the rule, plug in an input, read off the output. To understand a function, plot it. To use it for learning, ask the calculus question: "if I nudge the input a tiny bit, how much does the output move?"

Pick a rule, then drag the dot along the curve (or use the keyboard-friendly slider). Watch the input go into the machine and the output come out. Try the square-root rule and drag $x$ to a negative number: the machine refuses, because that input is not allowed (more about this in the next section).

Choose a curve and drag the blue dot along the x-axis to move the vertical line. For a real function the line always touches the curve at most once. Switch to the circle: the line can touch it twice, so a circle is not the graph of a function. Which x gives exactly one touch on the circle?

$f(x)$ is not "f times x". It means "the output of $f$ when the input is $x$". And $f(x+1)$ is the output at the input $x+1$, which is usually not $f(x) + 1$. For $f(x)=x^2$: $f(3+1) = 16$ but $f(3)+1 = 10$.

One input, one output. Many inputs may share an output ($x^2$ gives 4 for both 2 and $-2$). That is fine. What is forbidden is one input with two outputs.

Quick check: is $y = \pm\sqrt{x}$ (both signs) a function of $x$?

No. At $x = 4$ it gives both $2$ and $-2$, so a vertical line at $x=4$ touches it twice. The expression $\sqrt{x}$ on its own (the positive root only) is a function.

Domain and range core

Every machine has rules about what it accepts. A toaster accepts bread, not soup. The calculator's square-root key refuses $-4$, and the "divide by" key refuses $0$ and shows an error.

So a function comes with two natural questions:

  • What may I put in? The set of allowed inputs is the domain.
  • What can come out? The set of outputs that really happen is the range.

On a graph, the domain is the part of the x-axis the curve lives above or below. The range is the part of the y-axis the curve reaches.

FunctionDomain (allowed $x$)Range (outputs $y$)Why
$x^2$all real numbers$y \ge 0$squares are never negative
$\sqrt{x}$$x \ge 0$$y \ge 0$no real root of a negative number
$1/x$$x \ne 0$$y \ne 0$cannot divide by zero; and $1/x$ is never 0
$\ln x$$x > 0$all real numberslogarithm needs a positive input

Finding a domain, step by step. For $f(x) = \dfrac{1}{x-3}$: the only danger is dividing by zero. The bottom is zero when $x - 3 = 0$, that is $x = 3$. So the domain is "every $x$ except 3". For $f(x)=\sqrt{x-1}$: we need $x - 1 \ge 0$, so $x \ge 1$.

The domain of $f$ is the set of all inputs $x$ for which $f(x)$ is defined. The range of $f$ is the set of all outputs $f(x)$ that actually occur.

Interval notation. $[a, b]$ means $a \le x \le b$ (square bracket: the end is included). $(a, b)$ means $a < x < b$ (round bracket: the end is left out). $(0, \infty)$ means "every number above 0"; infinity is a direction, not a number, so it always gets a round bracket.

Three danger signs when you look for a domain: (1) dividing by zero, (2) an even root (like $\sqrt{\ }$) of a negative number, (3) a logarithm of zero or of a negative number.

Why do we need it?

Before using a function we must know where it works. Feeding it a forbidden input gives an error or garbage, and knowing the range tells us what values to expect back.

Where is it used?

A probability output must lie in the range $(0,1)$, which is why classifiers end with sigmoid or softmax. The loss $-\ln p$ needs $p>0$. A model trained on ages 20 to 60 has no reliable domain at age 90.

How is it used?

Check the three danger signs before you trust a formula. In code, clip inputs of logs and divisions (for example np.log(p + 1e-12)) so a stray zero cannot crash training.

Pick a function. The green band on the x-axis is its domain; the orange band on the y-axis is its range. Hollow rings mark ends that are left out. Drag the blue dot along the x-axis: inside the domain a dashed path goes up to the curve and across to the y-axis; outside it, the machine reports an error. Notice how the range is only the part of the y-axis the curve actually reaches.

Domain is about inputs, range is about outputs. Do not mix them up: for $\sqrt{x}$ both are "numbers 0 or more", but they are different sets (one on the x-axis, one on the y-axis).

Never take a log or a divide of an exact zero in code. A probability that rounds to $0$ makes $\ln p = -\infty$ and the whole loss becomes "inf" or "nan".

Quick check: what is the domain of $f(x) = \dfrac{\sqrt{x}}{x-4}$?

Two danger signs. The root needs $x \ge 0$. The division needs $x \ne 4$. Together: all $x \ge 0$ except $x = 4$.

Composition: machines in a row core

Put two machines on an assembly line. The first machine's output drops straight into the second machine. First wash the potato, then chop it. The pair of machines acts like one bigger machine that hides the middle step.

That is composition. Order matters: chopping first and washing second is a different (and messier) recipe.

Let $g(x) = x + 2$ (runs first) and $f(x) = x^2$ (runs second). Start with $x = 3$.

  1. First machine: $g(3) = 3 + 2 = 5$.
  2. Second machine takes that 5: $f(5) = 5^2 = 25$.
  3. So "$f$ after $g$" turns 3 into 25.

Now swap the order. $f(3) = 9$, then $g(9) = 9 + 2 = 11$. We got 11, not 25: order matters.

The formula. Replace $x$ inside $f$ by the whole of $g(x)$: $f(g(x)) = (x+2)^2 = x^2 + 4x + 4$. Check at $x=3$: $9 + 12 + 4 = 25$ ✓.

A one-neuron example. Let $g(x) = 2x - 1$ (a linear step) and $\sigma$ be the sigmoid (you meet it later in this chapter). At $x = 1$: $g(1) = 1$ and $\sigma(1) \approx 0.731$. A neuron is exactly "linear step, then squash".

The composition of $f$ and $g$ is the function $$(f\circ g)(x) = f\big(g(x)\big).$$ Read "$f$ after $g$": $g$ is the inner function (acts first), $f$ is the outer function (acts second).

  • Domain: $x$ must be allowed by $g$, and $g(x)$ must be allowed by $f$. Example: $f=\sqrt{\ }$, $g(x)=x-1$ gives $\sqrt{x-1}$, whose domain is $x \ge 1$.
  • Order matters: in general $f\circ g \ne g\circ f$.
  • Grouping does not matter: $(f\circ g)\circ h = f\circ(g\circ h)$, so we may write $f\circ g\circ h$.
  • A deep neural network with $L$ layers is a composition $f_L \circ \dots \circ f_2 \circ f_1$.
Why do we need it?

Complicated behaviour is built by chaining simple steps. Composition lets us name and study each step separately, and later it tells us how nudges travel through a chain.

Where is it used?

Every neural network layer-stack, an image pipeline (resize, normalise, model), a loss that wraps a model, $\mathrm{loss}(\mathrm{model}(x))$, and a model that wraps a feature map, such as $\sigma(\mathbf{w}\cdot\mathbf{x}+b)$.

How is it used?

Work from the inside out: compute the inner output first, then feed it to the outer function. To get a formula, replace the letter $x$ in the outer formula by the whole inner formula (in brackets!).

Looking ahead. If a nudge to the input moves the inner output by some factor, and a nudge to that output moves the final result by another factor, the total effect is the product of the two factors. This is the chain rule (Chapter 2.8), and backpropagation (Chapter 2.9) is that rule applied through a whole network.

Choose an inner function $g$ (blue) and an outer function $f$. The green curve is $f(g(x))$. Drag the $x$ slider and follow the number through both machines. The dashed orange curve is the other order, $g(f(x))$: compare the two. Try $f=\sqrt{x}$ with $g = x - 1$ and slide $x$ below 1: the composite is only defined where $g(x)$ is allowed by $f$.

Quick check: if $f(x)=3x$ and $g(x)=x-2$, what are $f(g(4))$ and $g(f(4))$?

$g(4) = 2$, then $f(2) = 6$. And $f(4) = 12$, then $g(12) = 10$. So $f(g(4)) = 6$ but $g(f(4)) = 10$.

Inverse functions: the undo button core

Some machines have an undo machine. A function turns Celsius into Fahrenheit; its inverse turns Fahrenheit back into Celsius. Put a number through both and you are back where you started.

On a graph, an inverse swaps the roles of input and output. If $(2, 7)$ is on the graph of $f$, then $(7, 2)$ is on the graph of the inverse. Swapping $x$ and $y$ is a mirror image across the diagonal line $y = x$.

An undo button only works if the machine never maps two different inputs to the same output. Otherwise, looking at the output, you could not tell which input it came from.

Find the inverse of $f(x) = 2x + 3$. Write $y = 2x + 3$ and solve for $x$:

  1. Start: $y = 2x + 3$.
  2. Subtract 3: $y - 3 = 2x$.
  3. Divide by 2: $x = \dfrac{y - 3}{2}$.
  4. Now rename the input of the new function to $x$: $f^{-1}(x) = \dfrac{x-3}{2}$.

Check: $f(5) = 13$ and $f^{-1}(13) = (13-3)/2 = 5$ ✓. Temperature check: $F = 1.8C + 32$ gives $100^\circ\text{C} \to 212^\circ\text{F}$, and the inverse $C = (F - 32)/1.8$ gives $(212-32)/1.8 = 100$ ✓.

A function with no inverse: $f(x)=x^2$ sends both 2 and $-2$ to 4. Given the output 4, which input was it? If we restrict the domain to $x \ge 0$, the problem disappears and the inverse is $\sqrt{x}$.

A function $f$ is one-to-one if different inputs always give different outputs. (Horizontal line test: no horizontal line touches the graph twice.) Then it has an inverse function $f^{-1}$ that undoes it: $$f^{-1}\big(f(x)\big) = x \qquad\text{and}\qquad f\big(f^{-1}(y)\big) = y.$$

  • The domain of $f^{-1}$ is the range of $f$, and the range of $f^{-1}$ is the domain of $f$.
  • The graph of $f^{-1}$ is the graph of $f$ reflected in the line $y=x$.
  • Careful: $f^{-1}$ is not $1/f$. The small $-1$ means "inverse function", not "to the power $-1$".
  • How to find it: write $y = f(x)$, solve for $x$, then swap the names.
Why do we need it?

Often we know the result and want the cause: given the output, which input made it? An inverse answers that. It is also how we build new functions: $\ln$ and $\sqrt{\ }$ are inverses of other functions.

Where is it used?

The logit (inverse of the sigmoid) in logistic regression, $\ln$ as the inverse of $\exp$, undoing data normalisation to get real units back, and sampling random numbers by inverting a cumulative distribution.

How is it used?

First check the function is one-to-one (restrict its domain if not). Then solve $y=f(x)$ for $x$. To verify, compute $f^{-1}(f(x))$ and confirm you get $x$ back.

The blue curve is $f$, the dashed line is $y = x$, and the orange curve is the inverse (the blue curve mirrored). Drag the blue dot: its orange mirror swaps the two coordinates. Move the horizontal test line (purple): a function with an inverse is touched at most once. Pick x² (all x) to see it fail, then x² (x ≥ 0 only) to see the repair.

$f^{-1}(x) \ne \dfrac{1}{f(x)}$. For $f(x)=2x+3$: $f^{-1}(x) = (x-3)/2$, but $1/f(x) = 1/(2x+3)$. They are different functions.

Restricting the domain is a legitimate fix. $\sqrt{x}$ is the inverse of "$x^2$ on $x\ge0$". The function $x^2$ on all numbers has no inverse.

Quick check: find the inverse of $f(x) = 5x - 10$.

$y = 5x - 10 \Rightarrow y + 10 = 5x \Rightarrow x = (y+10)/5$. So $f^{-1}(x) = (x+10)/5 = \tfrac{x}{5} + 2$. Check: $f(3) = 5$, and $f^{-1}(5) = 15/5 = 3$ ✓.

Linear functions: a straight line core

Walk up a straight ramp. Every metre you move forward raises you by the same amount, no matter where you are on the ramp. That constant steepness is the whole idea of a linear function.

Two numbers describe a ramp: how steep it is (the slope), and how high it starts at $x=0$ (the intercept). The taxi fare from the start of the chapter is one: 2 dollars per km, plus 3 dollars to start.

Take $f(x) = 2x + 1$.

$x$0123
$f(x)$1357

Each step of 1 to the right adds exactly 2 to the output. That "2" is the slope.

From two points to a line. A line passes through $(1, 3)$ and $(3, 7)$.

  1. Slope = rise over run $= \dfrac{7 - 3}{3 - 1} = \dfrac{4}{2} = 2$.
  2. Intercept: use $y = mx + b$ at $(1,3)$: $3 = 2\cdot 1 + b$, so $b = 1$.
  3. The line is $y = 2x + 1$.

A linear function has the form $$f(x) = m\,x + b.$$

  • $m$ is the slope: $m = \dfrac{\text{rise}}{\text{run}} = \dfrac{\Delta y}{\Delta x} = \dfrac{y_2 - y_1}{x_2 - x_1}$. ($\Delta$, "delta", means "change in".) Positive $m$: the line climbs. Negative: it falls. Zero: flat.
  • $b$ is the intercept: $f(0) = b$, where the line crosses the y-axis.
  • The key property: $f(x+1) - f(x) = \big(m(x+1) + b\big) - (mx + b) = m$. The change per unit step is the same everywhere.
  • A vertical line ($x = 3$) is not a function of $x$ (it fails the vertical line test).

A word on names. In linear algebra, "linear" is stricter: it means $f(cx) = c\,f(x)$ and $f(x+y) = f(x)+f(y)$, which forces $b = 0$. A line with $b\ne0$ is then called affine. Machine learning is relaxed about this and calls $wx + b$ "linear". With many inputs it is $\mathbf{w}\cdot\mathbf{x} + b$, a dot product plus a constant.

Why do we need it?

A line is the simplest possible "cause and effect": one more unit of input always gives $m$ more output. It is the first model to try, and it is what every smooth curve looks like when you zoom in far enough (Chapter 2.12).

Where is it used?

Linear regression, the weighted sum $\mathbf{w}\cdot\mathbf{x}+b$ inside every neuron, linearly decaying learning-rate schedules, and tangent lines (the best straight-line copy of a curve at one point).

How is it used?

Get the slope from two points (rise over run), get the intercept from the value at $x=0$, then predict by plugging in. To learn a line from data, choose the $m$ and $b$ that make the errors smallest.

Looking ahead. The slope of a line is its rate of change, and it is the same everywhere. For a curve the steepness changes from point to point. Finding the steepness at one single point is exactly what the derivative does (Chapter 2.3). The derivative of the line $mx+b$ is just $m$.

Move the sliders for $m$ and $b$. Increase $m$: the line steepens; make it negative: the line falls. Change $b$: the whole line slides up and down without changing its tilt. Then move the probe $x$ and check that going 1 step to the right always raises the line by exactly $m$ (the orange triangle).

Drag the two points. The readout computes the slope as rise over run and then the intercept. Make the points line up vertically: the "line" is no longer a function of $x$.

Quick check: a line goes through $(0, 5)$ and $(2, 1)$. What is $f(10)$?

Intercept $b = 5$ (the value at $x=0$). Slope $= (1-5)/(2-0) = -2$. So $f(x) = -2x + 5$ and $f(10) = -20 + 5 = -15$.

Polynomial functions: sums of powers core

A polynomial is built from building blocks: $1$, $x$, $x^2$, $x^3$, … Each block is multiplied by a number (a coefficient) and the pieces are added up. With only $1$ and $x$ you get a line. Add $x^2$ and the line can bend once, into a parabola (a bowl). Add $x^3$ and it can bend twice, and so on.

So more powers means more flexibility: a higher-degree polynomial can wiggle more. That is both its power and its danger.

Let $p(x) = x^2 - 4x + 3$.

  1. Factor it: $x^2 - 4x + 3 = (x-1)(x-3)$ (check: $(x-1)(x-3) = x^2 - 3x - x + 3$ ✓).
  2. It is zero when a factor is zero, so the roots are $x = 1$ and $x = 3$. Check: $p(1) = 1 - 4 + 3 = 0$ ✓ and $p(3) = 9 - 12 + 3 = 0$ ✓.
  3. At $x = 0$: $p(0) = 3$ (the constant term).
  4. The bottom of the bowl is halfway between the roots, at $x = 2$: $p(2) = 4 - 8 + 3 = -1$.

A cubic: $x^3 - x = x(x-1)(x+1)$ has three roots, $-1, 0, 1$, and bends twice.

A polynomial of degree $n$ is $$p(x) = a_n x^n + a_{n-1}x^{n-1} + \dots + a_2 x^2 + a_1 x + a_0, \qquad a_n \ne 0.$$ The numbers $a_0,\dots,a_n$ are the coefficients; the degree is the highest power present. A root is an $x$ with $p(x) = 0$.

  • Degree 1 is a line, degree 2 a parabola, degree 3 a cubic.
  • A degree-$n$ polynomial has at most $n$ real roots and at most $n-1$ turning points (peaks and valleys).
  • End behaviour is decided by the leading term $a_n x^n$ alone, when $|x|$ is huge: even degree means both ends go the same way (both up if $a_n>0$); odd degree means the ends go opposite ways.
  • If you know the roots $r_1,\dots,r_n$, then $p(x) = a_n (x - r_1)\cdots(x - r_n)$.
  • The domain is all real numbers. Polynomials are smooth: no gaps, jumps or corners.
Why do we need it?

Polynomials are the easiest flexible curves: only multiplication and addition, so they are cheap to compute and (you will see) easy to differentiate. They can also imitate other smooth functions near a point.

Where is it used?

Polynomial regression and polynomial features, the quadratic shape of the mean-squared-error loss in the weights, Taylor series (Chapter 2.11), splines in curve fitting, and polynomial kernels in SVMs.

How is it used?

Choose a degree, then let training find the coefficients. Too low a degree underfits (cannot bend enough); too high overfits (it wiggles through the noise). The degree is a "complexity knob".

The polynomial is $a(x-r_1)(x-r_2)\cdots$ and the roots are the dots on the x-axis. Drag a root and watch the curve cross the axis exactly there. Change the degree (more roots), then flip the sign of $a$ with the slider. Count the bends: there are always one fewer than the degree (when the roots are different). Look at the two far ends: which way do they go for even and odd degree?

Eight noisy data points (drag them up and down) come from a smooth curve (dashed grey). The blue polynomial is the best fit of the chosen degree. Start at degree 1: too stiff (underfits). Raise the degree: the fit improves, and at degree 7 it passes through every point, so the training error is almost 0. But look between the points and compare the test error (distance from the true curve): it usually gets worse again. That is overfitting.

"Higher degree" is not "better". More flexibility always lowers the error on the points you trained on, but it can raise the error on new points. We measure fit on data the model has not seen.

Roots can be complex. $x^2+1$ has no real root (its bowl never touches the axis). The count "at most $n$" refers to real roots.

Quick check: how many turning points can a degree-4 polynomial have at most, and how do its ends behave if $a_4 > 0$?

At most $4 - 1 = 3$ turning points. Even degree with positive leading coefficient: both ends go up.

Exponential functions and the number $e$ core

In a linear function you add the same amount every step. In an exponential function you multiply by the same factor every step. A bacterium that splits in two every hour: 1, 2, 4, 8, 16, … It grows slowly at first, then explosively, because the bigger it is, the faster it grows.

Run the film backwards for decay: a cup of coffee's caffeine halves every few hours: 1, 1/2, 1/4, 1/8, … It shrinks toward zero but never reaches it.

One base is special: the number $e \approx 2.71828$. We meet it properly in Chapter 2.2 (where $e$ appears as a limit) and Chapter 2.3 (where its steepness is worked out), but the short story is: $e^x$ is the exponential whose steepness at any point equals its height at that point. That makes it the most natural exponential, and calculus loves it.

$f(x) = 2^x$:

$x$$-2$$-1$012310
$2^x$$\tfrac14$$\tfrac12$12481024

Decay: $(1/2)^x$ gives $1, \tfrac12, \tfrac14, \dots$ for $x = 0, 1, 2, \dots$. And $e^0 = 1$, $e^1 \approx 2.718$, $e^2 \approx 7.389$, $e^{-1} \approx 0.368$.

Exponent rules (always true, for $b>0$):

  • $b^{x+y} = b^x\, b^y$   (example: $2^{3+2} = 2^5 = 32 = 8\cdot 4$ ✓)
  • $b^{0} = 1$ and $b^{-x} = 1/b^x$   (example: $2^{-3} = 1/8$)
  • $(b^x)^y = b^{xy}$   (example: $(2^3)^2 = 8^2 = 64 = 2^6$ ✓)

An exponential function with base $b>0$, $b \ne 1$, is $f(x) = b^x$.

  • Domain: all real numbers. Range: $y > 0$ (always positive, never zero).
  • It always passes through $(0, 1)$, because $b^0 = 1$.
  • If $b>1$ it grows (increasing); if $0\lt b<1$ it decays (decreasing).
  • The natural exponential uses $b = e \approx 2.71828\ldots$, written $e^x$ or $\exp(x)$.
  • Every exponential is secretly an $e$-exponential: since $b = e^{\ln b}$ (the number $\ln b$ is defined in the next section), $b^x = (e^{\ln b})^x = e^{x\ln b}$.
Why do we need it?

Many things grow or shrink by a fixed percentage per step rather than a fixed amount. Exponentials describe these, and $e^x$ is always positive, which is exactly what we need to turn any score into a valid weight or probability.

Where is it used?

Inside sigmoid and softmax, learning-rate decay schedules, exponential moving averages (momentum, Adam), weight decay, the Gaussian bell curve $e^{-x^2/2}$, compound interest and epidemic growth.

How is it used?

Pick a base (or a rate $k$ in $e^{kx}$). For a growth base, the doubling time is $\ln 2/\ln b$; for a decay base, the half-life is $\ln 2/(-\ln b)$. Remember $e^{a+b}=e^a e^b$ to simplify.

Slide the base $b$ below 1 (decay) and above 1 (growth); the curve always passes through $(0,1)$. Drag the dot and read the orange tangent's slope. The readout divides the slope by the height: for every base this ratio is the same number everywhere, namely $\ln b$ (the logarithm, explained in the next section). Press Set b = e: now the ratio is exactly 1, so slope = height.

$x^2$ is not $2^x$. In $x^2$ the variable is the base (a polynomial); in $2^x$ the variable is the exponent (an exponential). The second one eventually beats every polynomial: $2^{10}=1024$ but $10^2 = 100$, and the gap only grows.

An exponential never reaches 0 or goes negative: $b^x > 0$ for every $x$.

Quick check: simplify $e^{3}\cdot e^{-3}$ and $(e^{2})^{3}$.

$e^3 e^{-3} = e^{3-3} = e^0 = 1$. And $(e^2)^3 = e^{2\cdot3} = e^6$.

Logarithms: the inverse of exponentials core

A logarithm answers one question: "what power?" To what power must 2 be raised to get 8? Three, because $2\cdot2\cdot2=8$. We write $\log_2 8 = 3$.

So the logarithm is the undo button of the exponential (the inverse function from the last section). Exponential: power in, number out. Logarithm: number in, power out.

It also turns multiplication into addition. $1000 \times 100 = 100\,000$: the zeros add up, $3 + 2 = 5$. A logarithm just counts those zeros. That is why earthquakes (Richter), loudness (decibels) and acidity (pH) are measured on log scales: huge ranges become small, friendly numbers.

  • $\log_{10} 1000 = 3$, because $10^3 = 1000$.
  • $\log_2 32 = 5$, because $2^5 = 32$.
  • $\log_2 \tfrac18 = -3$, because $2^{-3} = \tfrac18$.
  • $\ln e^2 = 2$ (here $\ln$ means $\log_e$, the "natural log").
  • $\log_b 1 = 0$ for any base (since $b^0 = 1$) and $\log_b b = 1$.

Why not $\log 0$ or $\log(-5)$? Because $b^y$ is always positive, no power $y$ can give $0$ or a negative number. So logs only accept positive inputs.

Multiplication becomes addition: $\log_2(8 \cdot 4) = \log_2 32 = 5$, and $\log_2 8 + \log_2 4 = 3 + 2 = 5$ ✓.

For a base $b>0$, $b\ne1$, and $x > 0$: $$y = \log_b x \quad\Longleftrightarrow\quad b^{y} = x.$$ $\ln x = \log_e x$ is the natural logarithm; $\log_{10}$ is the common log; $\log_2$ counts bits. Domain: $x>0$. Range: all real numbers. Its graph passes through $(1, 0)$, has the y-axis as a vertical asymptote, and is the mirror image of $b^x$ in the line $y=x$. As inverses: $$\log_b(b^x) = x, \qquad b^{\log_b x} = x.$$

The rules, derived. Let $p = \log_b x$ and $q = \log_b y$, so $x = b^p$ and $y = b^q$.

  • Product: $xy = b^p b^q = b^{p+q}$, so $\log_b(xy) = p + q = \log_b x + \log_b y$.
  • Quotient: $x/y = b^p/b^q = b^{p-q}$, so $\log_b(x/y) = \log_b x - \log_b y$.
  • Power: $x^k = (b^p)^k = b^{pk}$, so $\log_b(x^k) = k\log_b x$.
  • Change of base: let $y = \log_b x$, so $b^y = x$. Take $\ln$ of both sides: $y \ln b = \ln x$, hence $\log_b x = \dfrac{\ln x}{\ln b}$.
Why do we need it?

Logs turn products into sums and powers into multiplications, and they squeeze enormous ranges into small ones. Sums are far easier to handle than products, both on paper and in a computer.

Where is it used?

The log-likelihood and the cross-entropy loss ($-\ln p$), entropy in bits, log-softmax, log-scaled plot axes for losses and learning rates, and feature transforms like $\log(\text{price})$.

How is it used?

To avoid tiny numbers, take logs of probabilities and add instead of multiplying. To undo an exponential, apply $\ln$. To change base, divide by $\ln b$.

The blue curve is $b^x$, the orange curve is $\log_b x$, and they mirror each other in the dashed line $y=x$. Drag the blue dot: the statement "$b^x = y$" on the blue curve becomes "$\log_b y = x$" on the orange curve. Change $b$, and press Set b = e. Notice the orange curve never touches the y-axis and passes through $(1,0)$.

Pick a base and choose two numbers $A$ and $B$. Compare each left side with its right side: they always agree (up to rounding). Try $A=8$, $B=4$ with base 2 to get whole numbers, then try ugly numbers and see that the rules still hold.

A model gives the right answer's probability $p$ for each of $n$ independent examples. The chance of all of them is the product $p^n$. Raise $n$: the product (left, ordinary scale) collapses to zero and the computer finally gives up (underflow), while the sum of logs (right) is just a calm straight line.

$\ln$ vs $\log$. In mathematics and most ML papers, "$\log$" with no base means the natural log $\ln$. In school (and on some calculators) it means $\log_{10}$. Check which one a formula uses; NumPy's np.log is the natural log.

Logs do not split sums: $\ln(A+B) \ne \ln A + \ln B$. The rules only work for products, quotients and powers.

Quick check: simplify $\ln(e^5) + \log_2 8 - \log_{10}(0.01)$.

$\ln(e^5) = 5$, $\log_2 8 = 3$, and $\log_{10}0.01 = \log_{10}10^{-2} = -2$. Total: $5 + 3 - (-2) = 10$.

Trigonometric functions: circles and waves

Watch a point travel round a circle of radius 1 at a steady speed. Look at it from the side (its height) and the height goes up and down smoothly, over and over. Plot that height against the angle travelled and you get a wave. That wave is the sine function. The point's horizontal position gives the cosine, the same wave shifted by a quarter turn.

Angles are measured in radians: the length of the arc you walk along the unit circle. A full circle has length $2\pi \approx 6.283$, so $360^\circ = 2\pi$ radians, $180^\circ = \pi$, $90^\circ=\pi/2$. Calculus (and every programming language) uses radians.

The point on the unit circle at angle $\theta$ is $(\cos\theta,\ \sin\theta)$. Here are the angles you should recognise:

Angle$0$$30^\circ = \pi/6$$45^\circ = \pi/4$$60^\circ = \pi/3$$90^\circ = \pi/2$$180^\circ = \pi$
$\cos\theta$1$\sqrt3/2 \approx 0.866$$\sqrt2/2 \approx 0.707$$1/2$0$-1$
$\sin\theta$0$1/2$$\sqrt2/2 \approx 0.707$$\sqrt3/2 \approx 0.866$10

Converting: radians $=$ degrees $\times \pi/180$. For $30^\circ$: $30\times\pi/180 = \pi/6 \approx 0.524$.

For any angle $\theta$ (in radians, counter-clockwise from the positive x-axis), the point on the unit circle is $(\cos\theta, \sin\theta)$. Also $\tan\theta = \dfrac{\sin\theta}{\cos\theta}$.

  • Domain of $\sin$, $\cos$: all real numbers. Range: $[-1, 1]$.
  • Periodic: $\sin(\theta + 2\pi) = \sin\theta$ and the same for cosine (period $2\pi$). $\tan$ has period $\pi$ and blows up where $\cos\theta = 0$ (at $\pm\pi/2, \pm3\pi/2,\dots$).
  • Pythagoras on the circle: the point $(\cos\theta,\sin\theta)$ is at distance 1 from the origin, so $\cos^2\theta + \sin^2\theta = 1$ always.
  • $\sin(-\theta) = -\sin\theta$ (odd) and $\cos(-\theta) = \cos\theta$ (even); $\cos\theta = \sin(\theta + \pi/2)$.
  • The wave family: $y = A\sin(\omega x + \varphi) + D$ has amplitude $A$ (height of a peak above the middle), angular frequency $\omega$ (period $2\pi/\omega$), phase $\varphi$ (slides the wave left by $\varphi/\omega$) and offset $D$ (moves the middle line).
Why do we need it?

Many things repeat or rotate: seasons, sound, an angle of a robot arm. Trig functions are the exact mathematics of repeating and of rotating, and they connect angles to lengths.

Where is it used?

Positional encodings in Transformers (sines and cosines of many frequencies), Fourier features and audio signals, rotation matrices, seasonality in time series, cosine similarity (the $\cos\theta$ in the dot product), and cosine learning-rate schedules.

How is it used?

Give angles in radians (NumPy's np.sin expects them). Choose amplitude, frequency and phase to match a repeating pattern. Remember the bounded range $[-1,1]$: sin and cos never blow up.

Drag the blue dot round the circle (or use the slider and the quick-angle buttons). The orange height is $\sin\theta$ and the blue horizontal distance is $\cos\theta$. On the right the same two numbers are drawn as waves against the angle. Find the angles where $\sin\theta = 0$, where $\cos\theta=0$, and where $\sin\theta=\cos\theta$.

The dashed grey curve is plain $\sin x$. Change $A$ (taller), $\omega$ (more waves, a shorter period), $\varphi$ (slides left or right) and $D$ (moves the middle line). The green arrow measures one period. Check the readout: period $=2\pi/\omega$. A Transformer uses many sines like this with different $\omega$ so every position gets a unique pattern.

Degrees vs radians. np.sin(90) is not 1: NumPy reads 90 as 90 radians. Use np.sin(np.pi/2) or np.deg2rad(90).

$\sin^2\theta$ means $(\sin\theta)^2$, not $\sin(\theta^2)$.

Quick check: what are $\sin(\pi)$, $\cos(\pi)$ and the period of $\sin(3x)$?

At $\theta=\pi$ (halfway round, the point $(-1,0)$): $\sin\pi = 0$, $\cos\pi = -1$. The period of $\sin(3x)$ is $2\pi/3\approx 2.094$.

The sigmoid function: a smooth on/off switch core

Imagine a dimmer switch instead of a light switch. Turn the knob far to the left: the light is off (0). Far to the right: fully on (1). In the middle: half bright. The sigmoid is that dimmer. It takes any number (a "score", positive or negative, huge or tiny) and squashes it into a value between 0 and 1.

That is why it is perfect for probabilities. A big positive score means "very likely yes" (close to 1). A big negative score means "very likely no" (close to 0). A score of 0 means "50-50".

The graph is an "S" with two flat tails. In the flat tails the output barely changes even when the input changes a lot. This is called saturation, and it matters a great deal for learning (see the callout below).

The formula is $\sigma(x) = \dfrac{1}{1 + e^{-x}}$. Compute $\sigma(1)$ by hand:

  1. $e^{-1} \approx 0.3679$.
  2. $1 + 0.3679 = 1.3679$.
  3. $\sigma(1) = 1 / 1.3679 \approx 0.7311$.
$x$$-5$$-2$0125
$\sigma(x)$0.00670.11920.50.73110.88080.9933

A spam filter. It computes a score of 2 for an email. $\sigma(2)\approx 0.88$: "88% chance it is spam". A score of $-5$ gives $0.0067$: almost certainly not spam. Going from 2 to 5 only moves the probability from 0.88 to 0.993: that flattening is saturation.

The sigmoid (or logistic function) is $$\sigma(x) = \frac{1}{1 + e^{-x}} = \frac{e^{x}}{1 + e^{x}}.$$ (The two forms agree: multiply top and bottom of the first by $e^x$.)

Properties, each derived:

  • Range $(0,1)$. Since $e^{-x}>0$, the bottom $1+e^{-x}$ is bigger than 1, so $0<\sigma(x)<1$. It gets as close as you like to 0 and 1 but never touches them.
  • $\sigma(0) = \tfrac12$. Because $e^0 = 1$ gives $1/(1+1)$.
  • Symmetry: $\sigma(-x) = 1 - \sigma(x)$. Indeed $1 - \sigma(x) = 1 - \dfrac{1}{1+e^{-x}} = \dfrac{e^{-x}}{1+e^{-x}} = \dfrac{1}{e^{x}+1} = \sigma(-x)$ (multiply top and bottom by $e^x$ in the last step). The curve is point-symmetric about $(0, \tfrac12)$.
  • Always increasing, since $e^{-x}$ falls as $x$ grows.
  • Inverse = the logit. Solve $p = \dfrac{1}{1+e^{-z}}$ for $z$: $1 + e^{-z} = \dfrac1p$, so $e^{-z} = \dfrac{1-p}{p}$, so $-z = \ln\dfrac{1-p}{p}$, so $$z = \ln\frac{p}{1-p} = \operatorname{logit}(p).$$ Here $\frac{p}{1-p}$ is the odds and $z$ is the log-odds.
  • Saturation numbers. $\sigma(x)\ge 0.95$ once $x \ge \ln 19 \approx 2.94$, and $\sigma(x) \ge 0.99$ once $x\ge\ln 99\approx 4.60$. (Solve $\sigma(x) = 0.95$: $e^{-x} = 0.05/0.95 = 1/19$.)
  • Slope preview. The steepness of the curve is $\sigma'(x) = \sigma(x)\,(1-\sigma(x))$, at most $\tfrac14$ (at $x=0$) and nearly $0$ in the tails. We prove this in Chapter 2.3.
  • Shape knobs: $\sigma(k(x - x_0))$ is the same S, centred at $x_0$ and steeper when $k$ is larger. As $k\to\infty$ it becomes an abrupt step.
  • Cousin: $\tanh x = 2\sigma(2x) - 1$ has the same S-shape but ranges over $(-1, 1)$.
Why do we need it?

A model's raw score can be any number, but a probability must lie between 0 and 1. Sigmoid converts one into the other smoothly, so we can still use calculus on it (a hard cut-off at 0 or 1 has no useful slope).

Where is it used?

Logistic regression and every binary classifier's output layer, the gates of LSTM and GRU networks, "attention gates", the Swish/SiLU activation $x\,\sigma(x)$, and mapping any score to a "probability of yes".

How is it used?

Compute a score $z = \mathbf{w}\cdot\mathbf{x} + b$, then the probability $p=\sigma(z)$. Predict "yes" when $p>0.5$, which is the same as $z>0$. Train with the loss slope; the slope of $\sigma$ is the part that shrinks when $|z|$ is large.

Drag the blue dot along the x-axis. In the red shaded tails the output is within 5% of 0 or 1: saturated, and the curve is almost flat there. Raise $k$ to make the switch sharper; change $x_0$ to slide the midpoint. Look at the "slope here" number: biggest in the middle ($0.25k$), almost zero in the tails.

Slide the probability $p$. The odds are $p/(1-p)$ ("9 to 1" for $p=0.9$) and the log-odds is $\ln$ of the odds. At $p=0.5$ the log-odds is 0. The sigmoid takes the log-odds back to $p$: the round trip always returns the same $p$.

Saturation causes slow learning. Training works by nudging the weights in the direction that lowers the error, and the size of the nudge depends on the slope of every function in the chain. In the flat tails of the sigmoid the slope is almost 0 (for example $\sigma'(5) \approx 0.0066$, versus $0.25$ at the centre), so the signal that flows back through a saturated sigmoid is tiny. In deep networks these small factors multiply and the signal vanishes: the vanishing gradient problem (Chapter 2.9). That is one reason hidden layers now often use ReLU (below), while sigmoid is kept for outputs and gates.

$\sigma$ outputs are probabilities only if the score is meaningful; a sigmoid of any number is between 0 and 1, even when the model is wrong.

Quick check: if $\sigma(z) = 0.2$, what is $\sigma(-z)$, and what is $z$ roughly?

By symmetry $\sigma(-z) = 1 - 0.2 = 0.8$. And $z = \operatorname{logit}(0.2) = \ln(0.2/0.8) = \ln 0.25 \approx -1.386$.

The softmax function: scores to probabilities core

Sigmoid answers a yes/no question. Softmax answers "which one of many?". A network looking at a photo produces one raw score (called a logit) per class: say cat 2.0, dog 1.0, bird 0.1. We want a list of probabilities: all positive, adding up to 1, with a bigger score getting a bigger share.

The recipe has two steps. (1) Make every score positive with $e^{\text{score}}$. (2) Divide each by the total, so you get a "share of the pie". The name says it all: it is a soft version of "take the max". The winner gets the biggest slice, but the others still get something, instead of everything going to the winner.

A temperature knob $T$ controls how soft: a small $T$ makes it almost winner-takes-all, a large $T$ makes the slices almost equal.

Scores $z = [2.0,\ 1.0,\ 0.1]$.

  1. Exponentiate: $e^{2.0} \approx 7.389$, $e^{1.0} \approx 2.718$, $e^{0.1} \approx 1.105$.
  2. Add them: $7.389 + 2.718 + 1.105 = 11.212$.
  3. Divide each by the total: $\dfrac{7.389}{11.212} \approx 0.659$, $\dfrac{2.718}{11.212}\approx 0.242$, $\dfrac{1.105}{11.212}\approx 0.099$.
  4. Check: $0.659 + 0.242 + 0.099 = 1.000$ ✓.

Adding the same number to every score changes nothing: $[102, 101, 100.1]$ gives the same three probabilities, because the common factor $e^{100}$ cancels from top and bottom.

Temperatureprobabilities for $[2.0, 1.0, 0.1]$behaviour
$T = 0.5$0.864, 0.117, 0.019sharper: the winner dominates
$T = 1$0.659, 0.242, 0.099ordinary softmax
$T = 5$0.400, 0.327, 0.273flatter: nearly equal shares

For scores $\mathbf{z} = [z_1, \dots, z_n]$, the softmax is the vector with entries $$\operatorname{softmax}(\mathbf{z})_i = \frac{e^{z_i}}{\sum_{j=1}^{n} e^{z_j}}.$$ With temperature: $\operatorname{softmax}(\mathbf{z}/T)$.

  • Valid probabilities. Each $e^{z_i}>0$ and is smaller than the total, so each output is in $(0,1)$. Adding them: $\sum_i p_i = \dfrac{\sum_i e^{z_i}}{\sum_j e^{z_j}} = 1$.
  • Order is kept. $e^x$ is increasing, so a bigger score always gets a bigger probability.
  • Shift invariance. $\dfrac{e^{z_i + c}}{\sum_j e^{z_j + c}} = \dfrac{e^c\,e^{z_i}}{e^c\sum_j e^{z_j}} = \dfrac{e^{z_i}}{\sum_j e^{z_j}}$. Only the differences between scores matter. (Multiplying all scores by a number does change the answer: that is exactly what temperature does.)
  • Two classes = sigmoid. $p_1 = \dfrac{e^{z_1}}{e^{z_1}+e^{z_2}}$. Divide top and bottom by $e^{z_1}$: $p_1 = \dfrac{1}{1+e^{-(z_1-z_2)}} = \sigma(z_1 - z_2)$. Softmax generalises the sigmoid.
  • Temperature limits. $T\to 0$: all the weight goes to the largest score (the "hard max", argmax). $T\to\infty$: every class gets $1/n$.
  • Numerically safe version. Subtract the largest score first, $z_i - \max_j z_j$ (allowed by shift invariance). This keeps every exponent $\le 0$ so $e^{(\cdot)}$ can never overflow. In logs: $\ln p_i = z_i - \ln\sum_j e^{z_j}$ (the "log-sum-exp").
  • Slope preview. $\dfrac{\partial p_i}{\partial z_j} = p_i(\delta_{ij} - p_j)$, a table of slopes that you derive in Chapter 2.5 and use in Chapter 2.14. ($\delta_{ij}$ is 1 when $i=j$ and 0 otherwise.)
Why do we need it?

A classifier with many classes must output a proper probability for each, positive and adding to 1. Softmax does this while staying smooth, so calculus can tell the network how to improve each score.

Where is it used?

The output layer of image classifiers and language models (a probability for each of tens of thousands of next words), attention weights in Transformers, policies in reinforcement learning, and the cross-entropy loss. Temperature controls the randomness of text sampling.

How is it used?

Compute logits, subtract their maximum, exponentiate, divide by the sum. Pick the biggest probability (or sample from them). Lower $T$ for confident, repeatable outputs; raise $T$ for more varied ones. For training, take $\ln$ of the right class's probability.

Move the four logit sliders and watch the bars on the left (logits) turn into the bars on the right (probabilities that add to 1). Lower the temperature $T$: the biggest bar swallows the rest. Raise $T$: all bars level out. Then tick Add 5 to every logit: the left bars all jump, the right bars do not move at all (shift invariance).

The three logits are $[c,\ c-1,\ c-2]$, so the true answer never changes (only differences matter). Slide $c$ up. The naive recipe computes $e^{c}$ directly and breaks once $c$ passes about 709 (the number is too big for a computer). The stable recipe, subtracting the largest logit first, never breaks.

Move $z_1$ and $z_2$. The blue dot (softmax probability of class 1) always lies on the sigmoid curve drawn against the difference $z_1 - z_2$. Make $z_1 = z_2$: a 50-50 tie.

Softmax outputs are not guaranteed to be calibrated. They always look like probabilities, even when the model is guessing. A softmax of $[10, 0, 0]$ says 99.99% for class 1, however wrong the model might be.

Softmax only sees differences. If you add the same number to every logit, nothing changes. That is a feature, but it also means the absolute size of a logit alone carries no meaning.

Quick check: what is softmax of $[0, 0, 0, 0]$, and of $[5, 5]$?

Equal scores give equal shares: $[0.25, 0.25, 0.25, 0.25]$ and $[0.5, 0.5]$. Each $e^0=1$ (or each $e^5$), the sum is $4$ (or $2e^5$), and the shares are $1/4$ (or $1/2$).

Piecewise functions: ReLU, leaky ReLU, absolute value, step, clip core

Some rules change depending on where the input is. A phone plan: "10 dollars flat up to 5 GB, then 2 dollars for every extra GB". Income tax works the same way, with different rates for different brackets. A piecewise function is several simple rules glued together, each used on its own stretch of inputs.

The most famous one in machine learning is ReLU ("rectified linear unit"): if the input is negative, output 0; otherwise pass it through unchanged. It is a flat line that bends into a ramp at 0. Simple, cheap, and it powers almost every modern deep network.

The phone plan. Let $x$ be the gigabytes used. $$\text{cost}(x) = \begin{cases} 10 & \text{if } x \le 5 \\ 10 + 2(x-5) & \text{if } x > 5 \end{cases}$$ For $x = 3$: first rule, cost $=10$. For $x = 8$: second rule, $10 + 2\cdot 3 = 16$. At $x=5$ both rules agree ($10$), so the pieces meet with no jump.

Absolute value. $|x| = x$ if $x \ge 0$, and $-x$ if $x<0$. So $|{-3}| = 3$ and $|4| = 4$.

The ML family (all piecewise):

NameRuleShape
ReLU$\max(0, x)$: $0$ for $x<0$, $x$ for $x\ge0$flat, then a 45° ramp
Leaky ReLU$x$ for $x\ge0$, $\alpha x$ for $x<0$ (small $\alpha$, e.g. 0.01)a gentle slope on the left
Absolute value$|x|$a V (used in the L1 loss)
Step$0$ for $x<0$, $1$ for $x\ge0$a jump
Clip$\min(\max(x, -c), c)$: stays in $[-c, c]$flat, ramp, flat
Huber loss$\tfrac12x^2$ if $|x|\le\delta$, else $\delta(|x| - \tfrac\delta2)$a bowl with straight sides

A piecewise function is defined by cases: $$f(x) = \begin{cases} f_1(x) & \text{if } x \text{ is in region 1} \\ f_2(x) & \text{if } x \text{ is in region 2} \\ \ \vdots \end{cases}$$ The regions must not overlap, and together they must cover the domain, so every $x$ has exactly one output.

  • The places where one piece hands over to the next are the break points. If the pieces meet there, the graph is connected. If the graph has a corner there it is a kink (ReLU at 0, $|x|$ at 0). If the pieces do not meet, there is a jump (the step function at 0).
  • Max and min give compact forms: $\text{ReLU}(x)=\max(0,x)$, $|x| = \max(x,-x)$, $\text{clip}=\min(\max(x,-c),c)$.
  • Slope preview. On each straight piece the slope is constant: ReLU has slope $0$ for $x<0$ and $1$ for $x>0$. At the kink there is no single slope. How calculus deals with kinks is in Chapter 2.3; frameworks simply pick a convention (slope 0 or 1 at exactly 0).

Why ReLU is so popular: (1) it is the cheapest possible non-linear function; (2) for $x>0$ it never saturates: its slope is exactly 1, so learning signals pass through unchanged (compare the flat tails of the sigmoid); (3) many ReLU "hinges" added together can trace any smooth curve, as closely as you like, as a chain of straight pieces (see the second widget below); (4) it outputs exact zeros, which makes layers sparse. Its weakness: for $x<0$ the slope is 0, so a unit that gets stuck on the negative side stops learning (a "dead ReLU"). Leaky ReLU keeps a small slope $\alpha$ there, so the signal never fully dies.

Why do we need it?

Real rules often differ by region, and a network built only from straight (linear) layers can only ever draw straight things. Bending the line, even once, is what lets a network model curves.

Where is it used?

ReLU in most CNNs and the hidden layers of many networks; leaky ReLU in GANs; Huber loss in robust regression and in the DQN reinforcement-learning loss; absolute value in the L1 loss; clipping of gradients and of probability ratios (PPO); the step function in the original perceptron.

How is it used?

Apply the function to every number coming out of a linear layer: np.maximum(0, z). Choose by region: write each piece, test the condition, use the matching piece. For losses, pick the one whose shape penalises errors the way you want.

Pick a function and drag the blue dot along the x-axis. The readout tells you which piece is used, the output, and the slope of that piece. Cross $x=0$ in ReLU: the slope changes from 0 to 1 (a kink). Use the knob: for leaky ReLU it is $\alpha$ (try 0.3 to see the left slope clearly), for clip it is $c$, for Huber it is $\delta$. Switch to Step to see a real jump.

The thick black curve is $b + c_1\,\text{ReLU}(x+2) + c_2\,\text{ReLU}(x) + c_3\,\text{ReLU}(x-2)$. Each coloured thin curve is one hinge, bending at its own knot (the dots at $-2$, $0$, $2$). Press Bump: three hinges build a triangle-shaped bump. Change the sliders and read the slope of each stretch: every hinge adds to the slope after its knot. With more hinges, a network can trace any curve.

Check the break points. If two pieces do not agree at the joint, the function jumps (like the step function) and cannot be trained with slopes. If they agree but with different slopes, you get a kink (like ReLU), which is usually fine in practice.

"Leaky" does not mean "smooth". Leaky ReLU still has a corner at 0; it only fixes the zero slope on the left.

Quick check: evaluate leaky ReLU with $\alpha=0.1$ at $x=-4$, and ReLU at $x=-4$ and $x=2.5$.

Leaky ReLU at $-4$: negative input, so $\alpha x = 0.1\cdot(-4) = -0.4$. ReLU at $-4$ is $\max(0,-4) = 0$. ReLU at $2.5$ is $2.5$.

Functions of many inputs: a first look

So far the machine had one slot for input. Real models have many. A house price depends on area and number of bedrooms and age. A neural network's loss depends on millions of weights.

With two inputs we can still draw the picture: the two inputs $x$ and $y$ are positions on a floor, and the output is the height of a surface above that spot. A loss function is a landscape of hills and valleys. Training means walking downhill. The rest of this guide teaches you how to find "downhill" using calculus.

A machine can also have several outputs, such as a layer that turns 3 numbers into 4. Both ideas are previewed here and handled properly in Chapter 2.4 and Chapter 2.5.

Let $f(x, y) = \dfrac{x^2 + y^2}{4}$ (a bowl).

  1. $f(0, 0) = 0$: the bottom of the bowl.
  2. $f(2, 0) = 4/4 = 1$.
  3. $f(2, 2) = (4 + 4)/4 = 2$.
  4. $f(-2, 2) = (4 + 4)/4 = 2$ as well: the bowl is symmetric.

A function with two outputs: $\mathbf{F}(x) = (2x,\ x^2)$ turns the input 3 into the vector $(6, 9)$.

A function of $n$ inputs is written $f(x_1, \dots, x_n)$, or $f(\mathbf{x})$ with the inputs gathered into a vector $\mathbf{x}\in\mathbb{R}^n$. It maps $\mathbb{R}^n \to \mathbb{R}$. Its graph is a surface in $n+1$ dimensions; for $n=2$ it is a surface over the $xy$-plane.

A function with $m$ outputs, $\mathbf{F}:\mathbb{R}^n\to\mathbb{R}^m$, is a vector-valued function (like one layer of a network). A slice (hold $y$ fixed and let only $x$ move) turns the surface back into an ordinary one-input curve. Slopes of slices are the partial derivatives of Chapter 2.4.

Why do we need it?

Models depend on many numbers at once, so we need functions of many inputs. Seeing the surface gives you the right picture for loss landscapes, even though real ones have millions of directions.

Where is it used?

The loss of any model as a function of its weights, the prediction as a function of many features, a neural-network layer ($\mathbb{R}^n \to \mathbb{R}^m$), and optimisation landscapes in gradient descent.

How is it used?

Think of the inputs as a point on a floor and the output as a height. To study one input at a time, take a slice. To improve a model, move the inputs in the direction that lowers the height.

Rotate by dragging the background. Drag the blue dot over the floor: it stays on the surface and its height is $f(x,y)$. Turn on Show slices: the orange curve holds $y$ fixed and the green curve holds $x$ fixed, each is an ordinary one-input function. Try the Top view, then Front, and compare the bowl with the saddle: on a saddle you go up in one direction and down in the other.

Quick check: for $f(x,y) = x^2 - y^2$, what are $f(2,0)$ and $f(0,2)$?

$f(2,0) = 4 - 0 = 4$ and $f(0,2) = 0 - 4 = -4$. Same distance from the origin, opposite heights: this is the saddle (up along one axis, down along the other).

Recap, cheat sheet and practice

  • A function gives exactly one output for each allowed input. Its domain is the allowed inputs; its range is the outputs that occur.
  • Composition $f(g(x))$ chains machines (order matters); an inverse $f^{-1}$ undoes $f$ and is its mirror image in $y=x$, and exists only for one-to-one functions.
  • Lines $mx+b$ have a constant slope; polynomials add powers (degree = flexibility, and too much flexibility overfits).
  • Exponentials $b^x$ multiply by a fixed factor per step; $e^x$ is the natural one. Logarithms are their inverse and turn products into sums (log-likelihood, no underflow).
  • Trig functions are the unit circle seen as waves: use radians, period $2\pi/\omega$.
  • Sigmoid squashes any score into $(0,1)$, with saturated flat tails. Softmax turns a score list into probabilities that add to 1 (shift-invariant, temperature-controlled). Two-class softmax is a sigmoid.
  • Piecewise functions (ReLU, leaky ReLU, $|x|$, step, clip, Huber) use a different rule per region; glued ReLU hinges can draw any curve.
  • Looking ahead: the slope of each function is what calculus will compute, and a loss is a function of many inputs.

Cheat sheet

FunctionFormulaDomain → rangeRemember
Line$mx + b$$\mathbb{R}\to\mathbb{R}$slope = rise / run
Polynomial$a_nx^n+\dots+a_0$$\mathbb{R}\to$ dependsat most $n$ roots, $n-1$ bends
Exponential$b^x$, $e^x$$\mathbb{R}\to(0,\infty)$$b^{x+y}=b^xb^y$; passes $(0,1)$
Logarithm$\ln x$, $\log_b x=\frac{\ln x}{\ln b}$$(0,\infty)\to\mathbb{R}$$\ln(xy)=\ln x+\ln y$
Sine / cosine$A\sin(\omega x+\varphi)+D$$\mathbb{R}\to[D-A,\ D+A]$ (that is $[-1,1]$ for plain $\sin$, $\cos$)radians; period $2\pi/\omega$; $\sin^2+\cos^2=1$
Sigmoid$\dfrac{1}{1+e^{-x}}$$\mathbb{R}\to(0,1)$$\sigma(-x)=1-\sigma(x)$; inverse is logit; max slope $\tfrac14$
Softmax$\dfrac{e^{z_i}}{\sum_j e^{z_j}}$$\mathbb{R}^n\to$ probabilitiessums to 1; subtract max; temperature $z/T$
ReLU / leaky$\max(0,x)$ / $\max(\alpha x, x)$$\mathbb{R}\to[0,\infty)$ / $\mathbb{R}$kink at 0; slope 0 or 1 (or $\alpha$)
Absolute value, clip$|x|$, $\min(\max(x,-c),c)$$\mathbb{R}\to[0,\infty)$ / $[-c,c]$corners, not jumps
Code it · NumPy

import numpy as np

# 1. A function is a rule: write it as code
def f(x):
    return 2 * x + 3
print(f(5), f(-2))                       # 13 -1

# 2. Composition and the inverse
g = lambda x: x + 2
sq = lambda x: x ** 2
print(sq(g(3)), g(sq(3)))                # 25 11   (order matters)
f_inv = lambda y: (y - 3) / 2
print(f_inv(f(5)))                       # 5.0

# 3. Exponential and logarithm are inverses
x = np.array([0.5, 1.0, 2.0])
print(np.exp(x))                         # [1.64872127 2.71828183 7.3890561 ]
print(np.log(np.exp(x)))                 # [0.5 1.  2. ]
print(np.log2(32), np.log10(1000))       # 5.0 3.0

# 4. Trig functions use radians
print(np.sin(np.pi / 6))                 # 0.49999999999999994 (that is 1/2, up to rounding)
print(np.sin(np.deg2rad(90)))            # 1.0

# 5. Sigmoid, logit and tanh
def sigmoid(z):
    return 1 / (1 + np.exp(-z))
def logit(p):
    return np.log(p / (1 - p))
z = np.array([-5.0, 0.0, 1.0, 5.0])
print(sigmoid(z))                        # [0.00669285 0.5        0.73105858 0.99330715]
print(logit(sigmoid(1.0)))               # 1.0 (up to rounding)
print(np.tanh(1.0), 2 * sigmoid(2.0) - 1)  # 0.7615941559557649 0.7615941559557646

# 6. Softmax (stable version: subtract the max first)
def softmax(z, T=1.0):
    z = np.asarray(z, dtype=float) / T
    e = np.exp(z - z.max())
    return e / e.sum()
print(softmax([2.0, 1.0, 0.1]))          # [0.65900114 0.24243297 0.09856589]
print(softmax([2.0, 1.0, 0.1], T=0.5))   # [0.86377712 0.11689952 0.01932336]
print(softmax([1002.0, 1001.0, 1000.1])) # [0.65900114 0.24243297 0.09856589]  (same: shift invariance)

# 7. Piecewise functions
relu = lambda x: np.maximum(0, x)
leaky = lambda x, a=0.1: np.where(x >= 0, x, a * x)
clip = lambda x, c=1.5: np.clip(x, -c, c)
v = np.array([-4.0, -1.0, 0.0, 2.5])
print(relu(v))                           # [0.  0.  0.  2.5]
print(leaky(v))                          # [-0.4 -0.1  0.   2.5]
print(clip(v))                           # [-1.5 -1.   0.   1.5]

# 8. Why we add logs: 400 probabilities of 0.1
print(0.1 ** 400)                        # 0.0   (underflow!)
print(400 * np.log(0.1))                 # -921.0340371976182
Test yourself

1. What is the domain of $f(x) = \sqrt{5 - x}$?

The root needs $5 - x \ge 0$. Adding $x$ to both sides gives $x \le 5$. At $x=5$ the output is $\sqrt0 = 0$, which is fine, so 5 is included.

2. Let $f(x) = 2x$ and $g(x) = x^2$. What is $f(g(3))$?

Inner function first: $g(3) = 9$. Then $f(9) = 18$. (The other order, $g(f(3)) = g(6) = 36$, is different.)

3. You are told $\sigma(3) \approx 0.953$. What is $\sigma(-3)$?

The sigmoid is symmetric: $\sigma(-x) = 1 - \sigma(x) = 1 - 0.953 = 0.047$. A sigmoid output is never negative.

4. What happens to the softmax output if you add 10 to every logit?

The factor $e^{10}$ appears in every numerator and in the denominator, so it cancels. Softmax only depends on the differences between logits.

5. For positive $x$ and $y$, which statement is always true?

Logs turn products into sums. The correct power rule is $\ln(x^2) = 2\ln x$, and the correct quotient rule is $\ln(x/y) = \ln x - \ln y$.

6. What is a "dead" ReLU unit?

ReLU gives 0 for negative inputs and also has slope 0 there, so no learning signal gets through. Leaky ReLU keeps a small slope $\alpha$ on the negative side to avoid this.

Practice problems

A. Find the inverse of $f(x) = \dfrac{2x+1}{3}$ and check it at $x=2$.

Write $y = \dfrac{2x+1}{3}$. Multiply by 3: $3y = 2x + 1$. Subtract 1: $3y - 1 = 2x$. Divide by 2: $x = \dfrac{3y-1}{2}$. So $f^{-1}(x) = \dfrac{3x-1}{2}$. Check: $f(2) = 5/3$, and $f^{-1}(5/3) = (5 - 1)/2 = 2$ ✓.

B. Factor $p(x) = x^3 - 4x$, list its roots, say how the ends behave, and compute $p(1)$.

$x^3 - 4x = x(x^2 - 4) = x(x-2)(x+2)$. The roots are $-2$, $0$ and $2$ (three roots, so two bends). Odd degree with positive leading coefficient: the left end goes down and the right end goes up. $p(1) = 1 - 4 = -3$ (negative, as expected between the roots $0$ and $2$).

C. Solve $3e^{2x} = 21$ for $x$.

Divide by 3: $e^{2x} = 7$. Take $\ln$ of both sides: $2x = \ln 7 \approx 1.9459$. So $x = \tfrac12\ln 7 \approx 0.973$. Check: $3e^{1.946} = 3 \times 7.00 = 21$ ✓.

D. A model's score for an email is $z = \ln 4$. What probability does the sigmoid give? What is the logit of that probability?

$\sigma(\ln 4) = \dfrac{1}{1 + e^{-\ln 4}} = \dfrac{1}{1 + \tfrac14} = \dfrac{1}{5/4} = 0.8$. The logit goes back: $\ln\dfrac{0.8}{0.2} = \ln 4 \approx 1.386$ ✓ (the odds are 4 to 1).

E. Compute the softmax of the logits $[0,\ \ln 2,\ \ln 3]$.

$e^0 = 1$, $e^{\ln 2} = 2$, $e^{\ln 3} = 3$. The sum is $6$. So the probabilities are $[\tfrac16, \tfrac26, \tfrac36] = [0.167, 0.333, 0.5]$, and they add to 1 ✓.

F. The Huber loss with $\delta = 1$ is $\tfrac12 x^2$ if $|x|\le1$, else $|x| - \tfrac12$. Compute it at $x = 0.5$ and $x = 3$, and compare with $\tfrac12x^2$ at $x=3$. Why is Huber popular with outliers?

At $x = 0.5$: $|x|\le 1$, so $\tfrac12(0.25) = 0.125$. At $x=3$: $|x|>1$, so $3 - 0.5 = 2.5$. The plain squared loss at 3 would be $\tfrac12\cdot9 = 4.5$. For a big error (an outlier) Huber grows only in a straight line, so one wild point cannot dominate the loss, while near zero it keeps the smooth bowl shape.

Chapter 2.2

Limits & Continuity

A limit asks a gentle question: as the input creeps closer and closer to a spot, what value is the output heading towards? That is all we need. In the next chapter, the derivative will turn out to be one single limit. So this chapter is deliberately short on formalism and long on pictures: we want you to feel what "getting closer and closer" means.

  • Understand a limit as "where the output is heading", and read it from a table and a zoomed graph
  • Use one-sided and two-sided limits, and know when a limit does not exist
  • Understand infinite limits (vertical asymptotes) and limits at infinity (horizontal asymptotes)
  • Define continuity and recognise the three kinds of discontinuity: removable, jump, infinite
  • Know and derive the key limit identities: $\frac{\sin x}{x}\to1$, $(1+\frac1n)^n\to e$, $\frac{e^x-1}{x}\to 1$, $\frac{\ln(1+x)}{x}\to 1$
  • See why the derivative is a limit, so you are ready for Chapter 2.3

This chapter uses the function toolbox from Chapter 2.1 (domain, exponentials, logs, sine, sigmoid). We only need limits well enough to understand derivatives. So the proofs are light, and the one formal definition ($\varepsilon$-$\delta$) is a clearly marked optional glimpse.

The intuition of a limit: sneaking up on a value core

Imagine walking towards a door. You take a step, then half the remaining distance, then half again. You never quite touch the door in this game, but it is perfectly clear where you are heading: the door. A limit is that "where you are heading".

Now the important twist. A limit does not care what happens at the door. It only cares about the values on the way there. The function may be missing a value at that spot, or have a wrong one, and the limit does not change.

Why would anyone need this? Think of a car's speedometer. "Speed at exactly 3:00:00" would be distance divided by zero time, which is $0/0$ and meaningless. But the average speed over the last second, then the last tenth of a second, then the last thousandth, settles down to a clear number. That number is the limit, and it is the speed. This is exactly how the derivative is defined.

Let $f(x) = \dfrac{x^2 - 1}{x - 1}$. At $x = 1$ it is $\frac{0}{0}$, so $f(1)$ does not exist. But what is $f$ heading to as $x$ gets close to 1?

$x$0.90.990.9991.0011.011.1
$f(x)$1.91.991.999?2.0012.012.1

Check one entry: $f(0.9) = \dfrac{0.81 - 1}{0.9 - 1} = \dfrac{-0.19}{-0.1} = 1.9$ ✓. From both sides, the output heads to 2. We write $\displaystyle\lim_{x\to1}\frac{x^2-1}{x-1} = 2$, even though $f(1)$ itself does not exist.

Another one: $\dfrac{\sin x}{x}$ is also $\frac00$ at $x=0$. At $x = 0.1$ it equals $0.99833$, at $x=0.01$ it equals $0.99998$. It is heading to 1.

We write $$\lim_{x \to a} f(x) = L$$ and say "the limit of $f(x)$ as $x$ approaches $a$ is $L$" when $f(x)$ gets as close to $L$ as we like by taking $x$ close enough to $a$ (but $x\ne a$).

  • "$x\to a$" means $x$ gets closer and closer to $a$, without ever being equal to it. That is why $0/0$ spots are fine.
  • The value $f(a)$ plays no role. It may be missing, or different from $L$.
  • If $f(x)$ does not head to a single number, we say the limit does not exist (DNE).
  • Two ways to find a limit: numerically (a table of values nearer and nearer) and graphically (zoom in on the point). Later we also use algebra.
Why do we need it?

Instant quantities, such as the speed at one instant or the slope at one point, come out as $0/0$ when computed directly. Limits give a clean meaning to "what the ratio becomes as the gap shrinks to nothing".

Where is it used?

The derivative is a limit (Chapter 2.3), and so is an integral (Chapter 2.13). Limits also describe training that "converges" (loss approaching its minimum), infinite series (Taylor, Chapter 2.11) and a function's behaviour at extreme inputs, as in a saturating sigmoid.

How is it used?

Build a table of $f(x)$ for $x$ closer and closer to $a$ from both sides, or zoom the graph in. If the values settle, that number is the limit. If an algebraic trick (cancelling) is available, use it to be sure.

Pick a function. Move the zoom slider up: the distance $h$ between the probe points and $a$ shrinks by ten each step, and the picture zooms in. The blue and orange dots close in on the purple line, which is the limit $L$. For the first three functions the exact spot is a hole (the hollow ring): the function has no value there, yet the limit is perfectly clear. In the last one the function is fine at $a$ and the limit equals its value.

A limit is not a value of the function. $\lim_{x\to a}f(x)$ and $f(a)$ can be different, or one can exist without the other. We care about the neighbourhood, not the point itself.

"Closer and closer" is not "reaches". $x\to a$ never equals $a$. The sequence $0.9, 0.99, 0.999, \dots$ approaches 1 but is never 1, and 1 is still the limit.

A table can fool you. If a function wiggles very fast, a few sample points may look settled while the true limit does not exist. Pair tables with graphs and algebra (we meet such a case, $\sin(1/x)$, in the next section).

Quick check: if $g(x) = x + 1$ for $x \ne 3$ and $g(3) = 100$, what is $\lim_{x\to3} g(x)$?

The values near 3 (like $2.99 \to 3.99$ and $3.01\to4.01$) head to $4$. The odd value $g(3)=100$ is ignored. The limit is $4$.

One-sided limits: left and right core

Picture a river with a broken bridge. You can walk up to the gap from the left bank or from the right bank. Where you are standing as you reach the edge may be different on each side: the two banks may be at different heights.

A function can do the same. Approaching $a$ from the left (only using $x\lt a$) may head to one value, and approaching from the right (only $x>a$) may head to another. We give each side its own limit.

Let $f(x) = x$ for $x < 1$, and $f(x) = x^2 + 1$ for $x \ge 1$.

  • From the left: $f(0.9) = 0.9$, $f(0.99) = 0.99$, $f(0.999)=0.999$: heading to $\mathbf{1}$.
  • From the right: $f(1.1) = 1.21+1 = 2.21$, $f(1.01) = 1.0201 + 1 = 2.0201$, $f(1.001) \approx 2.002$: heading to $\mathbf{2}$.

Different banks, different heights: the function jumps at $x=1$. (And $f(1) = 2$ is just the value the function actually takes there.)

More cases. $\sqrt{x}$ near 0 only exists on the right ($x \ge 0$): $\lim_{x\to0^+}\sqrt x = 0$, and the left side makes no sense. The step function (0 for $x<0$, 1 for $x\ge0$) has left limit 0 and right limit 1. ReLU near 0 has left limit 0 and right limit 0 (no jump, only a corner).

The left-hand limit $\displaystyle\lim_{x\to a^-} f(x) = L^-$ uses only $x\lt a$ ("$a^-$" means "from below $a$"). The right-hand limit $\displaystyle\lim_{x\to a^+} f(x) = L^+$ uses only $x>a$.

  • Each one can be a number, $+\infty$, $-\infty$, or fail to exist.
  • They can differ (a jump), or agree.
  • A one-sided limit may not even make sense if the function does not exist on that side of $a$ (like $\sqrt x$ at $0^-$).
  • Some functions have neither: $\sin(1/x)$ near 0 oscillates between $-1$ and $1$ faster and faster, and never settles.
Why do we need it?

Functions can behave differently on the two sides of a point: jumps, switches, corners. Looking at each side separately tells us exactly what kind of change happens there.

Where is it used?

The step activation and threshold decisions, piecewise losses like Huber, and ReLU's corner. In Chapter 2.3, ReLU's "slope from the left" is 0 and "slope from the right" is 1: one-sided limits of slopes.

How is it used?

Evaluate the function at points just left and just right of $a$ (for example $a\pm0.001$) and compare. Different results mean a jump; both heading to $\pm\infty$ means a blow-up.

Pick a function and shrink the distance $h$ with the slider. The blue dot walks in from the left and the orange dot from the right. For the jump they end up at different heights. Try $|x|/x$, $\sqrt{x}$ and $\sin(1/x)$ too: for $\sin(1/x)$ the dots never settle, however small $h$ gets.

Quick check: for the step function ($0$ for $x<0$, $1$ for $x\ge0$), what are $\lim_{x\to0^-}$ and $\lim_{x\to0^+}$?

From the left the values are all $0$, so the left-hand limit is $0$. From the right the values are all $1$, so the right-hand limit is $1$. They differ: a jump.

Two-sided limits, and how to compute them core

The ordinary limit $\lim_{x\to a}f(x)$ allows approach from both sides at once. It is only meaningful if the function agrees with itself no matter which road you take. Two walkers, one from each bank, must arrive at the same spot. If they arrive at different heights, we say the limit does not exist.

To compute a limit you have three tools: a table (numbers), a graph (picture), and algebra. The algebra trick for the common $0/0$ trouble is: simplify first, so the troublemaker disappears, then substitute.

Example 1: factor and cancel. $\displaystyle\lim_{x\to1}\frac{x^2-1}{x-1}$.

  1. Substitute $x=1$: $\frac{0}{0}$. Stuck, so simplify first.
  2. Factor the top: $x^2-1 = (x-1)(x+1)$.
  3. Cancel $(x-1)$. This is allowed because $x\to1$ means $x\ne1$, so $x-1\ne0$. The expression equals $x+1$ for every $x\ne1$.
  4. Now substitute: $1+1 = \mathbf{2}$ ✓ (matches the table in the first section).

Example 2: the conjugate trick. $\displaystyle\lim_{x\to0}\frac{\sqrt{x+4}-2}{x}$.

  1. Substitute $x = 0$: $\frac{2-2}{0} = \frac00$. Stuck.
  2. Multiply top and bottom by the conjugate $\sqrt{x+4}+2$ (this multiplies by 1): $\dfrac{(\sqrt{x+4}-2)(\sqrt{x+4}+2)}{x(\sqrt{x+4}+2)}$.
  3. The top is $(x+4) - 4 = x$, using $(A-B)(A+B) = A^2-B^2$.
  4. Cancel $x$: we get $\dfrac{1}{\sqrt{x+4}+2}$.
  5. Substitute $x=0$: $\dfrac{1}{2+2} = \mathbf{\tfrac14}$.

Example 3: a limit that does not exist. $\dfrac{|x|}{x}$ at 0: from the right it is $+1$, from the left $-1$. The two sides disagree, so there is no limit.

Two-sided limit rule. $\displaystyle\lim_{x\to a}f(x) = L$ exactly when $$\lim_{x\to a^-}f(x) = L \quad\text{and}\quad \lim_{x\to a^+}f(x) = L.$$ For a limit that is a number, both one-sided limits must exist, be finite, and be equal. (If both sides blow up the same way we write $\pm\infty$, see the next sections.)

Limit laws (if $\lim f = L$ and $\lim g = M$ both exist):

  • Sum / difference: $\lim (f \pm g) = L \pm M$. Constant multiple: $\lim (c f) = cL$.
  • Product: $\lim (f g) = LM$. Quotient: $\lim (f/g) = L/M$, provided $M\ne0$.
  • Power: $\lim (f^n) = L^n$. Compositions of nice functions can be done "from the inside out".
  • Direct substitution: for polynomials, exponentials, sine, cosine and (where defined) logs and roots, just plug in $a$: $\lim_{x\to a} f(x) = f(a)$. Trouble ($0/0$) means: simplify first.
Why do we need it?

We need rules that let us compute limits exactly, rather than guess from a table. The $0/0$ case is the important one: it is exactly the form that a derivative takes.

Where is it used?

Deriving every derivative rule in Chapter 2.3 (each one is a limit simplified by algebra), proving that loss functions are well behaved, and checking numerical code where "$0/0$" would crash (for example sin(x)/x at $x=0$).

How is it used?

1. Try direct substitution. 2. If you get $0/0$, factor, cancel or multiply by a conjugate. 3. Substitute again. 4. Cross-check with a table or graph, and test both sides if the function might jump.

Pick a limit. The readout shows what direct substitution gives, the algebra that fixes it, and a numeric check from both sides ($h = 0.1,\ 0.01,\ 0.001$). Try the ones that fail: for $|x|/x$ the two sides disagree, and for $1/x^2$ both sides blow up.

"$0/0$" is not an answer; it is a signal. It means "I need to simplify". The limit of a $0/0$ expression can be any number (we got 2, then $\tfrac14$, then 6) or may not exist.

Cancel only common factors, never common terms. In $\frac{x^2-1}{x-1}$ you may cancel the factor $(x-1)$, but you may not "cancel the $x$'s" in $\frac{x+3}{x}$.

Quick check: find $\displaystyle\lim_{x\to2}\frac{x^2-4}{x-2}$.

Substituting gives $0/0$. Factor: $x^2-4 = (x-2)(x+2)$. Cancel $(x-2)$: left with $x+2$. At $x=2$: $4$. The limit is $4$.

Infinite limits: when the output blows up

Divide 1 by a smaller and smaller positive number: $1/0.1 = 10$, $1/0.01 = 100$, $1/0.001 = 1000$. The answer grows without any ceiling. Pick any huge number you like; if $x$ is close enough to 0, $1/x$ is bigger than it.

We describe this by saying the limit is infinity. Infinity is not a number you can reach; it is a short way to say "grows without bound". On a graph, the curve shoots up (or down) along a vertical line it never touches: a vertical asymptote.

  • $\dfrac1{x^2}$ near 0: $x = \pm0.1\to100$, $x=\pm0.01\to10\,000$. Both sides go to $+\infty$: $\displaystyle\lim_{x\to0}\frac{1}{x^2} = +\infty$.
  • $\dfrac1{x}$ near 0: from the right $0.01\to 100$ (up); from the left $-0.01\to-100$ (down). So $\lim_{x\to0^+}\frac1x = +\infty$ and $\lim_{x\to0^-}\frac1x = -\infty$. The two-sided limit does not exist, since the sides disagree.
  • In ML: the cross-entropy loss for a correct class with predicted probability $p$ is $-\ln p$. As $p\to0^+$ it blows up: $-\ln 0.01 = 4.6$, $-\ln 10^{-6} = 13.8$, and it keeps growing without bound. A model that is confidently wrong is punished without limit.

$\displaystyle\lim_{x\to a} f(x) = +\infty$ means: $f(x)$ becomes larger than any chosen number, as soon as $x$ is close enough to $a$. Similarly for $-\infty$ (more negative than any number). One-sided versions: $x\to a^-$, $x\to a^+$.

  • Strictly speaking the limit "does not exist" (it is not a real number), but writing $\pm\infty$ tells us how it fails.
  • If $\lim_{x\to a^\pm}f(x)=\pm\infty$ (on at least one side), the line $x=a$ is a vertical asymptote of the graph.
  • Typical sources: a fraction whose bottom $\to0$ while the top does not, e.g. $\dfrac{1}{x-a}$, $\dfrac{1}{(x-a)^2}$, and $\ln x \to -\infty$ as $x\to0^+$.
  • Sign rule: $\dfrac{\text{positive}}{\text{tiny positive}} \to +\infty$, $\dfrac{\text{positive}}{\text{tiny negative}} \to -\infty$.
Why do we need it?

We need to describe functions that explode at some point, and to recognise danger spots where a formula must not be used (division by a vanishing number).

Where is it used?

The cross-entropy loss $-\ln p$ (infinite at $p=0$), numerical instability when dividing by tiny values (for example normalising by a standard deviation that is almost 0), and the logit $\ln\frac{p}{1-p}$, which blows up at $p=0$ and $p=1$.

How is it used?

Find where the denominator (or the log's input) becomes 0, then check each side's sign. In code, add a tiny constant ($10^{-8}$) or clip probabilities away from 0 and 1, so the blow-up never actually happens.

Pick a function and move the asymptote position $a$. Reduce the distance $h$ with the slider: the blue (left) and orange (right) dots shoot off the top or bottom of the picture. Compare $1/(x-a)$ (the two sides disagree: up on the right, down on the left) with $1/(x-a)^2$ (both sides go up) and $\ln|x-a|$ (both go down).

$\infty$ is not a number. You cannot do $\infty - \infty$ or $\frac{\infty}{\infty}$ as ordinary arithmetic. We only use $\infty$ to describe a trend.

Mind the sides. $\frac{1}{x}$ and $\frac{1}{x^2}$ both have a vertical asymptote at 0, but they behave differently on the left side.

Quick check: what are $\lim_{x\to2^+}\frac{1}{x-2}$ and $\lim_{x\to2^-}\frac{1}{x-2}$?

From the right, $x-2$ is a tiny positive number, so $1/(x-2)\to+\infty$. From the left, $x-2$ is a tiny negative number, so $1/(x-2)\to-\infty$.

Limits at infinity: where does the curve level off? core

Now change the question. Instead of asking what happens near a point, ask: what happens far away, as $x$ gets enormous? Does the output settle down to a steady value, grow forever, or keep wiggling?

Think of a cup of coffee cooling in a room. After a long time it settles at room temperature, and that steady value is the limit as time goes to infinity. On a graph this is a horizontal asymptote: a flat line the curve gets closer and closer to.

The sigmoid is the perfect example: as the score gets very large, the output creeps up to 1 and never passes it. That is exactly the "saturation" from Chapter 2.1.

  • $\displaystyle\lim_{x\to\infty}\frac{1}{x} = 0$ (since $1/10 = 0.1$, $1/1000 = 0.001$, …).
  • $\displaystyle\lim_{x\to\infty}\sigma(x) = 1$ and $\displaystyle\lim_{x\to-\infty}\sigma(x) = 0$. Also $e^{-x}\to0$ as $x\to\infty$.
  • A rational function. $\displaystyle\lim_{x\to\infty}\frac{3x^2+1}{x^2+2}$. Divide the top and bottom by the highest power, $x^2$:
    1. $\dfrac{3x^2+1}{x^2+2} = \dfrac{3 + 1/x^2}{1 + 2/x^2}$.
    2. As $x\to\infty$, $1/x^2\to0$ and $2/x^2\to0$.
    3. So the limit is $\dfrac{3+0}{1+0} = \mathbf{3}$.
    Numeric check: at $x = 10$ it is $301/102 \approx 2.951$; at $x=100$ it is $30001/10002 \approx 2.9995$ ✓.

$\displaystyle\lim_{x\to\infty} f(x) = L$ means $f(x)$ gets as close to $L$ as we like once $x$ is large enough. Then the line $y = L$ is a horizontal asymptote. Likewise for $x\to-\infty$. (The limit can also be $\pm\infty$ if the function grows forever.)

  • Rational functions $\dfrac{\text{degree } m}{\text{degree } n}$: if $m\lt n$, the limit is $0$; if $m=n$, it is the ratio of the leading coefficients; if $m>n$, it is $\pm\infty$. (Divide by the highest power of $x$ in the bottom to see why.)
  • Exponential beats polynomial: $\displaystyle\lim_{x\to\infty}\frac{x^k}{e^x} = 0$ for every $k$. (Numerically: $20^2/e^{20}\approx 8\times10^{-7}$.) And logarithms grow slower than any power: $\ln x/x\to0$.
  • Sine and cosine have no limit at infinity (they keep oscillating), but $\dfrac{\sin x}{x}\to 0$ because the top stays between $-1$ and $1$ while the bottom grows.
Why do we need it?

We want to know what a function does for extreme inputs and in the long run: does it saturate, settle, or explode? This predicts how a model behaves for very large scores or after very many training steps.

Where is it used?

Saturation of sigmoid, tanh and softmax for large logits; learning-rate decay to 0 and weight decay $e^{-\lambda t}$; a loss that levels off at its minimum; and why $e^{-x}$ weights in kernels vanish for far-away points.

How is it used?

Divide top and bottom by the highest power of $x$ to settle rational functions. Compare growth rates (log, polynomial, exponential) for the rest. Then read the horizontal asymptote off the result.

Left: the function over a normal window. Right: the same function viewed from far away, at $x = 10^k$ (the horizontal axis is $k$). Slide $k$ up: the dot follows the curve into the distance and the table shows $f(x)$ approaching the purple line $L$. Try the rational function, the sigmoid, and $x^2/(x+1)$ (which has no level: it grows forever).

A curve can cross its horizontal asymptote. $\frac{\sin x}{x}$ crosses the line $y=0$ again and again, yet its limit is 0. The asymptote describes the long-run trend, not a barrier.

Computers are not limits. Give a computer $x = 10^{400}$ and it overflows. Limits describe the exact mathematics; code must still avoid overflow (this is why softmax subtracts the maximum).

Quick check: find $\displaystyle\lim_{x\to\infty}\frac{5x+2}{2x-7}$.

Same degree on top and bottom. Divide by $x$: $\dfrac{5 + 2/x}{2 - 7/x} \to \dfrac{5}{2}$. The limit is $2.5$ (the ratio of the leading coefficients).

Continuity: draw it without lifting the pencil core

A function is continuous if you can draw its graph without lifting your pencil: no holes, no jumps, no sudden vertical escapes. In everyday terms: a small change in the input only causes a small change in the output. Nothing "snaps".

A bike's speed is continuous: it cannot jump from 10 to 30 km/h in zero time. A light switch, on the other hand, is not.

Check three functions at $x=1$:

  • $f(x)=x^2$: $f(1)=1$, and the limit as $x\to1$ is $1$. They match: continuous.
  • $f(x)=\dfrac{x^2-1}{x-1}$: $f(1)$ does not exist (hole). Not continuous at 1. (If we define $f(1)=2$, the hole is filled and it becomes continuous.)
  • Step function at $0$: the left limit is 0, the right limit is 1. No single limit: not continuous.

The intermediate value idea. A continuous function cannot skip values. Let $p(x)=x^3-x-1$. Then $p(1)=-1$ and $p(2)=5$: it goes from negative to positive, so it must cross 0 somewhere in between. Halving the interval: $p(1.5) = 0.875>0$, so the root is in $(1, 1.5)$; $p(1.25)\approx-0.297<0$, so it is in $(1.25, 1.5)$; $p(1.375)\approx0.225>0$, so it is in $(1.25, 1.375)$. This "bisection" method homes in on the root $\approx1.3247$.

A function $f$ is continuous at $x=a$ if all three hold:

  1. $f(a)$ is defined,
  2. $\displaystyle\lim_{x\to a}f(x)$ exists,
  3. $\displaystyle\lim_{x\to a}f(x) = f(a)$.

In one line: $\displaystyle\lim_{x\to a}f(x)=f(a)$ ("the limit equals the value"). $f$ is continuous on an interval if it is continuous at every point of it.

  • Continuous everywhere: polynomials, $e^x$, $\sin x$, $\cos x$, $|x|$, ReLU, sigmoid, tanh.
  • Continuous where defined: $\ln x$ on $x>0$, $\sqrt x$ on $x\ge0$, and rational functions $\frac{p}{q}$ wherever $q\ne0$.
  • Building blocks: sums, products, quotients (bottom $\ne0$) and compositions of continuous functions are continuous. So a network made only of continuous layers and activations is continuous.
  • Intermediate value theorem: if $f$ is continuous on $[a,b]$ and $f(a)$, $f(b)$ have opposite signs, then $f(x) = 0$ for some $x$ between.
  • Preview: to have a derivative at a point, a function must be continuous there (Chapter 2.3). The reverse is false: ReLU and $|x|$ are continuous but have a corner at 0.
Why do we need it?

Continuity is the promise that "nearby inputs give nearby outputs". Without it, a tiny nudge could cause a huge change, and we could neither predict nor train by small steps.

Where is it used?

Gradient-based training needs a loss that varies smoothly with the weights; stable predictions (small input noise should not flip the output wildly); root-finding by bisection; and the guarantee that composing layers keeps a network continuous.

How is it used?

Test the three conditions at suspicious points (where a denominator is 0, where pieces join). Between those points, trust the building-block rules. Prefer continuous activations and losses when you plan to differentiate.

Pick a function and drag the blue dot along the x-axis. At every $a$ the readout tests the three conditions. Move $a$ across the special point $x=1$ for each function and see which condition fails. Find the function where condition 3 fails even though the limit exists (the misplaced dot), and the one where all three pass again because the hole was filled.

Continuous does not mean smooth. $|x|$ and ReLU are continuous (no pencil lift) but have a sharp corner at 0. Continuity is the first level of "well-behaved"; having a derivative is the next.

Check continuity at a point or on an interval. $\frac1x$ is continuous on $x>0$ and on $x<0$, but it is not continuous on any interval containing 0, because 0 is not in its domain.

Quick check: is $f(x) = \dfrac{x}{x-2}$ continuous at $x=2$? at $x=5$?

At $x=2$ the function is not defined (division by 0), so condition 1 fails: not continuous there (it blows up: an infinite discontinuity). At $x=5$ it is a quotient of continuous functions with a non-zero bottom, so it is continuous: $f(5) = 5/3$.

Discontinuities: removable, jump and infinite core

When the pencil must lift, it can lift in three different ways:

  • Removable (a hole). One point is missing or misplaced, but the curve on both sides lines up. Just put the dot where it belongs and the problem is fixed.
  • Jump. The two banks of the river sit at different heights. No dot can repair it.
  • Infinite. The curve shoots off to $\pm\infty$ along a vertical asymptote.
TypeExample at the problem spotOne-sided limitsFixable?
Removable$\dfrac{x^2-1}{x-1}$ at $x=1$both $= 2$ (equal)yes: define $f(1)=2$
Jumpstep function at $0$; a tiered price that leaps from 10 to 20 dollars at 5 GB$0$ and $1$ (different)no
Infinite$\dfrac1x$ at $x=0$$-\infty$ and $+\infty$no

A fourth, rarer kind is oscillating: $\sin(1/x)$ at 0 swings between $-1$ and $1$ infinitely often, so no one-sided limit exists at all.

In ML. Accuracy is a step-like function of the weights: it only changes when a prediction flips, so it is made of flat pieces with jumps. Its slope is 0 almost everywhere, which tells gradient descent nothing. That is why we train on a smooth, continuous surrogate (cross-entropy or MSE) instead of accuracy itself.

Let $f$ be not continuous at $a$. Look at the one-sided limits $L^-$ and $L^+$:

  • Removable: $L^- = L^+ = L$ is a finite number, but $f(a)$ is undefined or $\ne L$. Fix: redefine $f(a)=L$.
  • Jump: $L^-$ and $L^+$ are both finite but $L^-\ne L^+$. The jump size is $L^+-L^-$.
  • Infinite: at least one of $L^-$, $L^+$ is $\pm\infty$; the line $x=a$ is a vertical asymptote.

Jump and infinite discontinuities are called non-removable. A function is continuous at $a$ exactly when none of these (nor the oscillating kind) happens.

Why do we need it?

Knowing which kind of break a function has tells you what to do: patch a removable hole, treat a jump with care, or avoid an infinite spot completely.

Where is it used?

Spotting the problem of threshold and step activations, analysing piecewise losses, finding where a formula such as $\frac{\sin x}{x}$ or $\frac{p}{1-p}$ needs a special case in code, and deciding which loss can be differentiated.

How is it used?

Compute the two one-sided limits and $f(a)$. Equal limits with a wrong or missing value: removable, and you can patch it in code with an if for that single point. Different finite limits: jump. Infinite: asymptote.

Pick a type and set where it happens ($a$) and how big it is ($s$). The classifier measures the two one-sided limits and $f(a)$, and names the discontinuity. Set $s = 0$ and watch the break vanish. For the hole, tick no value at a, and notice you can repair it; for the jump, there is nothing to patch.

"Removable" means the limit exists. That is the only requirement. A hole does not stop the function from having a clear limit, so it does no harm to calculus; you only need to patch the single value.

A jump is not an asymptote. In a jump, both one-sided limits are ordinary finite numbers; the function stays bounded.

Quick check: classify the discontinuity of $f(x) = \dfrac{x^2-9}{x-3}$ at $x=3$, and of $g(x) = \dfrac{1}{x-3}$ at $x=3$.

$f$: $(x-3)(x+3)/(x-3) = x+3$, so both sides head to $6$ while $f(3)$ is undefined: removable (define $f(3) = 6$). $g$: one side $\to+\infty$, the other $\to-\infty$: infinite.

Important limit identities (and how to derive them) core

A handful of limits turn up again and again, because they are exactly what the derivative of $\sin x$, $e^x$ and $\ln x$ is made of. Each is a $0/0$ puzzle with a clean answer:

  • $\dfrac{\sin x}{x}\to 1$: for a tiny angle, the sine is almost the angle itself (a tiny arc is almost straight).
  • $\left(1+\dfrac1n\right)^n \to e$: compound interest. Splitting a year into more and more steps gives a bigger and bigger payout, but it levels off at $e\approx2.71828$.
  • $\dfrac{e^x-1}{x}\to1$ and $\dfrac{\ln(1+x)}{x}\to1$: near 0, $e^x$ is almost $1+x$ and $\ln(1+x)$ is almost $x$.

You do not need to memorise them. We derive them, step by step, from a picture or from $e$ itself.

Compound interest. Put 1 dollar in a bank that pays 100% per year. If interest is paid once, you have $(1+1)^1 = 2$. If it is paid twice (50% each half-year), you have $(1+\tfrac12)^2 = 2.25$. Four times: $(1+\tfrac14)^4 \approx 2.4414$. Monthly: $(1+\tfrac1{12})^{12}\approx 2.6130$. Daily: $\approx2.7146$. A million times a year: $\approx 2.718280$. It never passes $e = 2.718281828\ldots$

A tiny angle. $\sin(0.1) = 0.0998334$, so $\sin(0.1)/0.1 = 0.99833$. $\sin(0.01)/0.01 = 0.99998$.

Others at $x=0.01$: $\frac{e^{0.01}-1}{0.01} = 1.00502$, $\frac{\ln 1.01}{0.01} = 0.99503$, $\frac{1-\cos0.01}{0.01^2} = 0.499996$.

The identities (all as $x\to0$ unless stated): $$\lim\frac{\sin x}{x}=1,\quad \lim_{n\to\infty}\Big(1+\frac1n\Big)^n=e,\quad \lim\,(1+x)^{1/x}=e,$$ $$\lim\frac{e^x-1}{x}=1,\quad \lim\frac{\ln(1+x)}{x}=1,\quad \lim\frac{a^x-1}{x}=\ln a,\quad \lim\frac{1-\cos x}{x^2}=\frac12.$$ And at infinity: $x^k/e^x\to0$ and $\ln x/x\to0$ ("exponential beats polynomial beats logarithm").

Derivations.

  • $e$ as a limit. This is the definition of $e$: $e=\lim_{n\to\infty}(1+\frac1n)^n$. Put $x = 1/n$ (so $x\to0^+$ when $n\to\infty$) to get $(1+x)^{1/x}\to e$. The table above shows the values creeping up to 2.71828.
  • $\ln(1+x)/x\to1$. Rewrite with the power rule of logs: $\dfrac{\ln(1+x)}{x} = \tfrac1x\ln(1+x) = \ln\big((1+x)^{1/x}\big)$. As $x\to0$ the inside heads to $e$ (above we showed this from the right, $x=1/n>0$; from the left it also holds, and we use that without proof), and $\ln$ is continuous, so the whole thing heads to $\ln e = 1$.
  • $(e^x-1)/x\to1$. Let $u = e^x - 1$. As $x\to0$, $u\to0$, and $x = \ln(1+u)$. Then $\dfrac{e^x-1}{x} = \dfrac{u}{\ln(1+u)} = \dfrac{1}{\ln(1+u)/u}\to\dfrac11 = 1$ (using the previous identity).
  • $(a^x-1)/x\to\ln a$. Write $a^x = e^{x\ln a}$ and let $t = x\ln a\to0$: $\dfrac{e^{t}-1}{x} = \ln a\cdot\dfrac{e^t-1}{t}\to\ln a\cdot1$.
  • $\sin x/x\to1$ (squeeze). On the unit circle, for $0\lt x<\frac\pi2$, compare three areas: the small triangle ($\tfrac12\sin x$) is inside the circular sector ($\tfrac12 x$), which is inside the big triangle ($\tfrac12\tan x$). So $\sin x\lt x<\tan x$. Divide by $\sin x$: $1<\dfrac{x}{\sin x}<\dfrac1{\cos x}$, and flipping: $\cos x<\dfrac{\sin x}{x}<1$. As $x\to0^+$, $\cos x\to1$, so $\frac{\sin x}{x}$ is squeezed between two things heading to 1. Since $\frac{\sin x}{x}$ has the same value at $-x$, the left side agrees.
  • $(1-\cos x)/x^2\to\frac12$. Multiply top and bottom by $1+\cos x$: $\dfrac{1-\cos^2x}{x^2(1+\cos x)} = \dfrac{\sin^2 x}{x^2(1+\cos x)} = \Big(\dfrac{\sin x}{x}\Big)^2\dfrac{1}{1+\cos x}\to 1\cdot\dfrac12$.

The squeeze theorem used above: if $g(x)\le f(x)\le h(x)$ near $a$ and both $g$ and $h$ head to $L$, then $f$ heads to $L$ too.

Why do we need it?

These limits are the raw ingredients of derivatives: $(\sin x)'=\cos x$ comes from $\frac{\sin x}{x}\to1$, $(e^x)'=e^x$ from $\frac{e^x-1}{x}\to1$, and $(\ln x)'=\frac1x$ from $\frac{\ln(1+x)}{x}\to1$ (all in Chapter 2.3).

Where is it used?

The number $e$ in softmax and sigmoid, compound growth and continuous decay ($e^{-\lambda t}$), small-angle approximations in physics and robotics, and numerically stable code: np.expm1 and np.log1p exist precisely because $e^x-1$ and $\ln(1+x)$ lose accuracy near 0.

How is it used?

Spot the pattern (a $0/0$ with $\sin$, $e^x-1$ or $\ln(1+x)$), rewrite so it matches a known identity (sometimes by a substitution like $u=3x$), and read off the answer. Example: $\frac{\sin 3x}{x} = 3\cdot\frac{\sin 3x}{3x}\to3$.

Invest 1 dollar at 100% yearly interest, paid $n$ times a year. Move the slider through the standard schedules (yearly, half-yearly, …, a million times). The orange curve is $(1+1/n)^n$ plotted against $\log_{10}n$: it rises quickly, then flattens just under the purple line at $e$. Ask: how far below $e$ is it at $n = 12$? At $n=10^6$?

Slide the angle $x$ (in radians) towards 0. Compare three lengths: the blue vertical segment $\sin x$, the orange arc $x$, and the green vertical segment $\tan x$. They always satisfy $\sin x\lt x<\tan x$, and as $x$ shrinks the three nearly coincide. That forces $\frac{\sin x}{x}$ to be squeezed between $\cos x$ and 1, both heading to 1.

Pick an identity. Slide $k$ so that $x=\pm10^{-k}$ closes in on 0 from both sides. The table shows the two sides agreeing with the target value (the purple line). Notice that the left and right values are different at first but converge to the same number. At $x=0$ itself every one of these is $\frac00$, so the graph has a hole there.

Computers lose digits near 0. $\frac{1-\cos x}{x^2}$ with $x=10^{-8}$ prints 0.0 in floating point, because $\cos x$ rounds to exactly 1 and the subtraction wipes out the answer. The identities say the true value is $\frac12$. Use the stable forms (np.expm1, np.log1p, $2\sin^2(x/2)$).

$(1+\frac1n)^n$ is not "1 to a power". The base tends to 1 and the exponent tends to infinity at the same time, a tug of war. Neither side wins: the result is $e$. And at absurd sizes ($n\approx10^{16}$) the computer rounds $1+\frac1n$ to exactly 1 and returns 1.

Quick check: find $\displaystyle\lim_{x\to0}\frac{\sin 3x}{x}$ and $\displaystyle\lim_{x\to0}\frac{e^{2x}-1}{x}$.

First: $\dfrac{\sin3x}{x} = 3\cdot\dfrac{\sin 3x}{3x}$. With $u=3x\to0$ the fraction $\to1$, so the limit is $3$. Second: $\dfrac{e^{2x}-1}{x} = 2\cdot\dfrac{e^{2x}-1}{2x}\to2\cdot1 = 2$.

Optional: the $\varepsilon$–$\delta$ glimpse

This section is optional. You can skip it and still follow the rest of the guide. It shows how mathematicians make "gets as close as we like" precise.

Think of a game between a challenger and you. The challenger names a tolerance on the output, $\varepsilon$ (the Greek letter epsilon): "I want $f(x)$ to be within $\varepsilon$ of $L$". You must answer with a tolerance on the input, $\delta$ (delta): "Then keep $x$ within $\delta$ of $a$, and I guarantee it."

If you can answer every challenge, however tiny the $\varepsilon$, then $L$ really is the limit.

Let $f(x)=2x+1$, $a=1$, $L=3$. The challenger says $\varepsilon = 0.1$.

  1. We need $|f(x) - 3| < 0.1$, that is $|2x+1-3| = |2x-2| = 2|x-1| < 0.1$.
  2. Divide by 2: $|x - 1| < 0.05$.
  3. So $\delta = 0.05$ works. In general, for any $\varepsilon$ choose $\delta = \varepsilon/2$: then $|x-1|<\delta$ gives $|f(x)-3| = 2|x-1| < 2\delta = \varepsilon$ ✓.

$\displaystyle\lim_{x\to a}f(x)=L$ means: for every $\varepsilon>0$ there is a $\delta>0$ such that $$0<|x-a|<\delta \ \Longrightarrow\ |f(x)-L|<\varepsilon.$$ In words: all inputs within $\delta$ of $a$ (except $a$ itself) give outputs within $\varepsilon$ of $L$. On a graph: the curve over the vertical strip $a\pm\delta$ stays inside the horizontal band $L\pm\varepsilon$.

Why do we need it?

It replaces the fuzzy "gets close" with a checkable promise. Every limit law and every theorem about limits is proved from this definition.

Where is it used?

Mostly in proofs. Its spirit appears in practice as tolerances: stop training when the loss is within $\varepsilon$ of its minimum, or compare floats with np.isclose (tolerance), and in convergence guarantees for optimisers.

How is it used?

To prove a limit, take an arbitrary $\varepsilon$, solve $|f(x)-L|<\varepsilon$ for $|x-a|$, and read off a $\delta$ (often a simple multiple of $\varepsilon$). You will rarely need to do this yourself.

The purple band is the output tolerance $L\pm\varepsilon$ chosen by the challenger. Slide $\delta$ to set the blue strip $a\pm\delta$. The piece of curve over the strip is green when it stays inside the band (you win) and red when it escapes. Press Find a δ to let the computer pick the biggest one that works, then shrink $\varepsilon$ and see $\delta$ shrink with it.

Quick check: for $f(x) = 5x$ at $a=0$ with $L=0$, which $\delta$ works for a given $\varepsilon$?

$|f(x)-0| = 5|x| < \varepsilon$ exactly when $|x| < \varepsilon/5$. So $\delta = \varepsilon/5$ works.

Where this leads: the derivative is a limit core

How steep is a curve at one single point? Steepness needs two points (rise over run), but we have only one. The trick is to take a second point very close by and draw the straight line through both, called a secant. Slide the second point closer and closer to the first. The secant line turns into the line that just touches the curve: the tangent. Its steepness is the slope of the curve at that point.

This is the speedometer story again: the average speed over a shorter and shorter time is the limit that gives the speed at one instant.

Take $f(x)=x^2$ at $a=1$. The secant through $(1, 1)$ and $(1+h,\ (1+h)^2)$ has slope $$\frac{f(1+h)-f(1)}{h} = \frac{(1+h)^2-1}{h}.$$

  1. At $h=0$ this is $\frac00$: stuck, as in every limit problem of this chapter.
  2. Expand the top: $(1+h)^2 - 1 = 1 + 2h + h^2 - 1 = 2h + h^2$.
  3. Cancel an $h$ (allowed: $h\ne0$ while $h\to0$): $\dfrac{2h+h^2}{h} = 2 + h$.
  4. Let $h\to0$: the slope heads to $\mathbf{2}$.

Numbers: $h=1$ gives $3$, $h=0.1$ gives $2.1$, $h=0.01$ gives $2.01$, $h=0.001$ gives $2.001$. The tangent slope at $x=1$ is 2.

The derivative of $f$ at $a$ is the limit of the secant slopes: $$f'(a)=\lim_{h\to0}\frac{f(a+h)-f(a)}{h}.$$ Every tool of this chapter will be used: the $\frac00$ trouble is a removable-hole type of limit, the algebra tricks (expand, factor, cancel) get it out, and identities such as $\frac{\sin h}{h}\to1$ give the derivatives of $\sin$, $e^x$ and $\ln x$. If the limit exists, $f$ is differentiable at $a$ (needing continuity there). We develop all of this in Chapter 2.3.

Why do we need it?

Machine learning asks "if I nudge this weight, how does the loss change?". That is a slope at a point, and a slope at a single point is only meaningful as a limit.

Where is it used?

Gradient descent, backpropagation, sensitivity analysis and every optimiser in the rest of this guide, since the gradient is a list of such limits. Even the finite-difference check of a gradient in code is this formula with a small, fixed $h$.

How is it used?

Write the secant slope $\frac{f(a+h)-f(a)}{h}$, simplify the algebra, and let $h\to0$. In code, use a small $h$ like $10^{-5}$ with the centred version $\frac{f(a+h)-f(a-h)}{2h}$ to check derivatives numerically.

The blue line is the secant through the dark point $P$ and the orange point $Q$ at distance $h$ to its right (or left). Slide $k$ up: $h=10^{-k}$ shrinks, and the blue secant swings into the dashed orange tangent. Compare the secant slope in the table with the tangent slope. Switch to "from the left": the same limit appears from the other side.

Quick check: for $f(x)=x^2$ at $a=3$, simplify the secant slope $\frac{(3+h)^2-9}{h}$ and take $h\to0$.

$(3+h)^2 - 9 = 9 + 6h + h^2 - 9 = 6h + h^2$. Divide by $h$: $6 + h$. As $h\to0$ this heads to $6$. So the slope of $x^2$ at $x=3$ is $6$ (compare: at $x=1$ it was $2$; the slope is $2a$).

Recap, cheat sheet and practice

  • A limit is where $f(x)$ is heading as $x\to a$ (without ever using $x=a$). The value $f(a)$ plays no role.
  • One-sided limits use only $x\lt a$ or only $x>a$. The two-sided limit is a number exactly when both one-sided limits exist, are finite and agree.
  • For a $\frac00$ puzzle: factor and cancel, or multiply by the conjugate, then substitute.
  • Infinite limits mean a vertical asymptote; limits at infinity mean a horizontal asymptote (divide by the highest power; exponential beats polynomial beats log).
  • Continuous at $a$: $f(a)$ defined, limit exists, limit $=f(a)$. A discontinuity is removable (hole), a jump, or infinite.
  • Key identities (all derived): $\frac{\sin x}{x}\to1$, $(1+\frac1n)^n\to e$, $\frac{e^x-1}{x}\to1$, $\frac{\ln(1+x)}{x}\to1$, $\frac{1-\cos x}{x^2}\to\frac12$.
  • The derivative is the limit of secant slopes, $f'(a)=\lim_{h\to0}\frac{f(a+h)-f(a)}{h}$: the road into Chapter 2.3.

Cheat sheet

IdeaNotation / testPicture or trick
Limit$\lim_{x\to a}f(x)=L$output heads to $L$; $f(a)$ ignored
One-sided$\lim_{x\to a^-}$, $\lim_{x\to a^+}$approach from the left / from the right
Two-sided exists$L^-=L^+$ (finite)both walkers meet at the same height
$0/0$ troublefactor, cancel, conjugatesimplify, then substitute
Vertical asymptote$\lim f=\pm\infty$ at $a$curve shoots up or down along $x=a$
Horizontal asymptote$\lim_{x\to\pm\infty}f=L$curve levels off at $y=L$
Continuity$\lim_{x\to a}f(x)=f(a)$no pencil lift
Removable / jump / infiniteequal limits, wrong value / unequal finite / $\pm\infty$hole / step / asymptote
Identities$\frac{\sin x}{x},\ \frac{e^x-1}{x},\ \frac{\ln(1+x)}{x}\to1$; $(1+\frac1n)^n\to e$tiny-angle, tiny-growth approximations
Derivative preview$f'(a)=\lim_{h\to0}\frac{f(a+h)-f(a)}{h}$secant swings into tangent
Code it · NumPy

import numpy as np

# 1. Numeric limits: look at f(a + h) for smaller and smaller h
f = lambda x: (x**2 - 1) / (x - 1)
for h in [0.1, 0.01, 0.001, 1e-6]:
    print(h, round(f(1 + h), 5), round(f(1 - h), 5))
# 0.1 2.1 1.9
# 0.01 2.01 1.99
# 0.001 2.001 1.999
# 1e-06 2.0 2.0          (both sides head to 2)

# 2. The sin(x)/x limit (sin(0)/0 is 0/0, so we stay close to 0 but not at 0)
x = np.array([0.5, 0.1, 0.01, 0.001])
print(np.sin(x) / x)     # [0.95885108 0.99833417 0.99998333 0.99999983]  -> 1

# 3. (1 + 1/n)^n heads to e
for n in [1, 10, 100, 10_000, 1_000_000]:
    print(n, round((1 + 1 / n) ** n, 6))
# 1 2.0
# 10 2.593742
# 100 2.704814
# 10000 2.718146
# 1000000 2.71828
print(round(np.e, 6))    # 2.718282

# 4. Stable versions of the identities near 0 (avoid rounding loss)
h = 1e-8
print(round(np.expm1(h) / h, 9))      # 1.000000005  (e^h - 1)/h -> 1
print(round(np.log1p(h) / h, 9))      # 0.999999995  ln(1+h)/h -> 1
print((1 - np.cos(h)) / h**2)         # 0.0  naive formula: all digits lost!
print(2 * np.sin(h / 2) ** 2 / h**2)  # 0.5  stable form of (1 - cos h)/h^2 -> 1/2

# 5. Infinite limits and limits at infinity
print(1 / 0.001, 1 / 0.000001)        # 1000.0 1000000.0   (1/x blows up from the right)
big = np.array([10.0, 100.0, 1000.0])
print((3 * big**2 + 1) / (big**2 + 2))   # [2.95098039 2.9995001  2.999995  ]  -> 3

# 6. Secant slopes of x^2 at a = 1 approach the slope 2
for h in [1, 0.1, 0.01, 0.001]:
    print(h, round(((1 + h) ** 2 - 1) / h, 5))
# 1 3.0
# 0.1 2.1
# 0.01 2.01
# 0.001 2.001
Test yourself

1. What is $\displaystyle\lim_{x\to3}\frac{x^2-9}{x-3}$?

Direct substitution gives $0/0$. Factor: $x^2-9 = (x-3)(x+3)$, cancel $(x-3)$, and substitute $x=3$ into $x+3$: $6$.

2. At a point $a$, a function has left-hand limit 2 and right-hand limit 5. Which statement is true?

A two-sided limit needs both one-sided limits to be equal. Different finite values mean a jump discontinuity, and the function never settles on one value.

3. What is $\displaystyle\lim_{x\to\infty}\frac{4x^2+x}{2x^2-3}$?

Same degree on top and bottom: divide by $x^2$ to get $\dfrac{4+1/x}{2-3/x^2}\to\dfrac42 = 2$, the ratio of the leading coefficients.

4. Which of these has a removable discontinuity at $x=2$?

$\frac{x^2-4}{x-2} = x+2$ for $x\ne2$: both sides head to 4, and only the value at 2 is missing (a hole you can fill). $\frac1{x-2}$ blows up, the step jumps, and $|x-2|$ is continuous (it has a corner, not a break).

5. What is $\displaystyle\lim_{x\to0}\frac{\sin x}{x}$?

By the squeeze: $\cos x<\frac{\sin x}{x}<1$ near 0 and $\cos x\to1$, so the ratio is squeezed to 1. (It is $\frac00$ at $x=0$ itself, but the limit only looks nearby.)

6. $f$ is continuous at $a$ when…

Continuity needs all three: $f(a)$ defined, the limit exists, and they are equal. Corners are allowed (ReLU and $|x|$ are continuous at 0).

Practice problems

A. Find $\displaystyle\lim_{x\to2}\frac{x^2-5x+6}{x-2}$.

Substitution gives $\frac00$. Factor the top: $x^2-5x+6=(x-2)(x-3)$. Cancel $(x-2)$: left with $x-3$. At $x=2$: $2-3=-1$. The limit is $-1$.

B. Find $\displaystyle\lim_{x\to0}\frac{\sqrt{1+x}-1}{x}$.

It is $\frac00$. Multiply top and bottom by the conjugate $\sqrt{1+x}+1$: the top becomes $(1+x)-1 = x$. So the expression is $\dfrac{x}{x(\sqrt{1+x}+1)} = \dfrac{1}{\sqrt{1+x}+1}$. At $x=0$: $\dfrac{1}{1+1}=\dfrac12$. (Numeric check at $x=0.01$: $0.49876$.)

C. Find $\displaystyle\lim_{x\to\infty}\frac{2x^3-x}{5x^3+4}$ and $\displaystyle\lim_{x\to\infty}\frac{x+1}{x^2+1}$.

First: same degree. Divide by $x^3$: $\dfrac{2-1/x^2}{5+4/x^3}\to\dfrac25$. Second: top degree 1 is lower than bottom degree 2. Divide by $x^2$: $\dfrac{1/x+1/x^2}{1+1/x^2}\to\dfrac01=0$.

D. For $f(x)=\dfrac{|x-3|}{x-3}$, find both one-sided limits at $x=3$ and name the discontinuity.

For $x>3$, $|x-3|=x-3$, so $f=1$: right-hand limit $1$. For $x<3$, $|x-3|=-(x-3)$, so $f=-1$: left-hand limit $-1$. They are different finite numbers: a jump of size $1-(-1)=2$. (Also $f(3)$ is not defined, and no choice of $f(3)$ can repair a jump.)

E. Show that $\displaystyle\lim_{x\to0}\frac{\sin 3x}{x}=3$.

Write $\dfrac{\sin3x}{x} = 3\cdot\dfrac{\sin 3x}{3x}$. Let $u=3x$; as $x\to0$, $u\to0$, and $\frac{\sin u}{u}\to1$. So the limit is $3\cdot1=3$. (Numeric check: at $x=0.01$, $\sin(0.03)/0.01 = 2.99955$.)

F. Compute the slope of the secant for $f(x)=x^2$ between $x=3$ and $x=3+h$, and take $h\to0$. Check it with $h = 0.01$.

Slope $=\dfrac{(3+h)^2-9}{h}=\dfrac{6h+h^2}{h}=6+h\to6$. With $h=0.01$ the secant slope is $6.01$, already close to the limit $6$. This limit is the derivative of $x^2$ at $x=3$, the topic of the next chapter.

Chapter 2.3

Differentiation of Univariate Functions

This is the heart of calculus. A derivative answers one simple question: if I nudge the input a tiny bit, how much does the output move? Learn to see it as a slope, learn to compute it from scratch, and learn the handful of rules that let you do it quickly. Then watch it steer a model downhill.

  • See the derivative as the slope of the tangent line and as a rate of change (speed, sensitivity)
  • Compute a derivative from first principles: secant line to tangent line, by hand, for $x^2$, $x^3$, $1/x$ and $\sqrt{x}$
  • Use and derive the power, product, quotient and chain rules, and the derivatives of $e^x$, $a^x$, $\ln x$, $\sin x$, $\cos x$, $\tan x$
  • Know the derivatives of the sigmoid and ReLU, and why a flat sigmoid starves learning
  • Meet the second and third derivatives ($f''$, $f'''$)
  • Connect it all to ML: loss functions, gradient descent, optimisation, sensitivity, learning curves

What is a derivative? core

Think of a road over hills. On a flat road the steepness is zero. Going uphill the steepness is positive, and going downhill it is negative. A sign at each spot could tell you "the road is climbing at 8% here".

A derivative is that sign, for a graph. At every point of a curve it tells you how steep the curve is right there. A straight line has the same steepness everywhere. A curve changes its steepness from point to point, so the derivative is itself a function: feed in $x$, get back the steepness at $x$.

Another way to say it: nudge $x$ a tiny bit to the right. How much does the output move up or down, for each unit of nudge?

Take $f(x) = x^2$. Remember "slope = rise ÷ run". We will nudge $x$ by $0.001$ and measure.

  1. At $x = 1$: $f(1) = 1$ and $f(1.001) = 1.002001$. The rise is $0.002001$. Divide by the run $0.001$: slope $\approx 2.001$.
  2. At $x = 2$: $f(2) = 4$ and $f(2.001) = 4.004001$. The rise is $0.004001$, so slope $\approx 4.001$.
  3. At $x = 0$: $f(0.001) = 0.000001$, so slope $\approx 0.001$, almost flat (the bottom of the bowl).
  4. At $x = -1$ the curve is going down: slope $\approx -2$.

The pattern is: slope $= 2x$. That rule, "$x$ goes in, $2x$ comes out", is the derivative of $x^2$.

The derivative of $f$ at $x$ is the slope of the curve at $x$:

$$f'(x) = \lim_{h \to 0} \frac{f(x+h) - f(x)}{h}$$

Here $h$ is a tiny nudge in $x$ and $f(x+h)-f(x)$ is how much $f$ moves. ("$\lim_{h\to 0}$" means: see what the ratio gets closer and closer to as $h$ shrinks; you met this in Chapter 2.2.) We prove it in the section on first principles below.

Names for the same thing: $f'(x)$ ("f prime of x"), $\dfrac{df}{dx}$ ("d f d x"), $\dfrac{d}{dx}f(x)$, or $y'$ if $y = f(x)$. A function with a derivative at a point is called differentiable there.

  • $f'(x) > 0$: the curve is rising at $x$.
  • $f'(x) < 0$: the curve is falling.
  • $f'(x) = 0$: the curve is flat (a hilltop, a valley bottom, or a shelf).
Why do we need it?

We need a number that says "which way is up, and how steeply" at one exact point. Without it we could only compare two far-apart points, never describe the curve right where we stand.

Where is it used?

Training every model: the derivative of the loss tells gradient descent which way to move each weight. Also physics (speed), economics (marginal cost) and sensitivity analysis.

How is it used?

Compute $f'(x)$ at your current point. Its sign says "which way is uphill" and its size says "how steep". To go downhill, step against it.

Drag the dot along the top curve (or along the green curve below, either works). The orange line is the tangent. Watch the green dot below: its height is the slope of the tangent. Find where the slope is zero. Then pick $|x|$ and drag to $x=0$: there is a sharp corner, and no single tangent.

Not every curve has a derivative everywhere. A sharp corner (like the tip of $|x|$ or the kink of ReLU) has no single tangent. A jump has none either. We will meet ReLU's kink again below.

The derivative is a function, and also a number. $f'(x)$ is a whole new function of $x$. $f'(2)$ is one number: the slope at $x=2$.

Quick check: for $f(x)=x^2$, what is the slope at $x=3$, and is the curve rising or falling there?

Using the pattern $f'(x)=2x$: slope $=6$. It is positive, so the curve is rising.

Geometric interpretation: zoom in until it is straight core

Stand in a field. The Earth is round, but the ground looks flat. Zoom in far enough on any smooth curve and the same thing happens: the curve looks like a straight line.

That straight line is the tangent line. It is the best straight-line copy of the curve near one point. Its slope is the derivative. This one picture is the idea behind almost everything in this chapter, and behind linearization later.

Take $f(x)=x^2$ at the point $a=1$, where $f(1)=1$ and the slope is $f'(1)=2$.

  1. The line through $(1, 1)$ with slope $2$ is $y = 1 + 2(x-1) = 2x - 1$.
  2. At $x = 1.1$: the curve gives $1.1^2 = 1.21$. The line gives $2(1.1)-1 = 1.2$. Gap: $0.01$.
  3. At $x = 1.01$: the curve gives $1.0201$, the line gives $1.02$. Gap: $0.0001$.

Come ten times closer and the gap becomes a hundred times smaller. That is "looks straight when you zoom in".

The tangent line to $y=f(x)$ at $x=a$ is the line through the point $(a, f(a))$ with slope $f'(a)$:

$$y = f(a) + f'(a)\,(x - a).$$

Near $a$ the curve and the tangent line almost agree: $f(x) \approx f(a) + f'(a)(x-a)$. This is the linear approximation.

Why do we need it?

Curves are hard to work with and straight lines are easy. If a curve looks like a line close up, we can answer "what happens if I move a little?" with simple arithmetic.

Where is it used?

Gradient descent (it assumes the loss is locally a slope), Newton's method, error estimates, and the Taylor series and linearization chapters ahead.

How is it used?

Compute $f(a)$ and $f'(a)$ once. Then predict nearby values with $f(a)+f'(a)(x-a)$ instead of evaluating $f$ again.

Slide zoom to the right. At first the blue curve and the orange tangent clearly separate. Keep zooming: they merge. The red bars at the window edges are the gap between curve and tangent; watch the number in the readout shrink. Move a to try other points, and try other functions.

Quick check: the tangent line to $f(x)=x^2$ at $a=3$?

$f(3)=9$ and $f'(3)=6$, so $y = 9 + 6(x-3) = 6x - 9$.

Rate of change: speed versus distance core

A car has two dials. The odometer says how far you have gone. The speedometer says how fast you are going right now. Speed is the rate of change of distance: how quickly distance is growing.

You can also compute a speed over a whole trip: "120 km in 2 hours is 60 km/h". That is the average rate. The speedometer shows the instantaneous rate: the average over a trip so short that it is just one moment. That instantaneous rate is the derivative.

The same idea works for any pair of quantities: dollars per item, loss per training step, temperature per hour.

A car's distance is $s(t) = t^2$ metres after $t$ seconds.

  1. Average over $t=1$ to $t=3$: $s(3)-s(1) = 9-1 = 8$ metres in $2$ seconds: $8/2 = 4$ m/s.
  2. Squeeze the trip around $t = 2$: from $2$ to $2.1$: $(4.41-4)/0.1 = 4.1$ m/s. From $2$ to $2.01$: $(4.0401 - 4)/0.01 = 4.01$ m/s.
  3. The averages approach $4$. So the speed at the moment $t=2$ is $4$ m/s, which matches $s'(t)=2t = 4$.

Units: the derivative has units "(output units) per (input unit)": metres per second here.

The average rate of change of $f$ from $x$ to $x+h$ is $\dfrac{f(x+h)-f(x)}{h}$ (rise over run: the slope of a secant line, which is a straight line through two points of the curve).

The instantaneous rate of change at $x$ is its limit as $h\to 0$, which is $f'(x) = \dfrac{df}{dx}$. Read $\dfrac{df}{dx}$ as "a tiny change in $f$ divided by the tiny change in $x$ that caused it".

Why do we need it?

Averages hide what happens at one moment. A trip can average 60 km/h while sometimes standing still. To describe "now", we need the rate over an infinitely short time.

Where is it used?

Physics (velocity, acceleration), finance (marginal cost), epidemic growth rates, and ML: the rate at which the loss falls per training step or per unit change of a weight.

How is it used?

Ask "output change per unit input change". Write it as $df/dx$ and keep the units: they are a quick check that you built the right ratio.

The top graph is distance $s(t)=3t^2 - t^3/3$ (metres). Drag the dot to choose the time $t$. The purple line joins the dot to a second point $\Delta t$ seconds later: its slope is the average speed. Slide $\Delta t$ down towards 0: the purple line turns into the orange tangent and the average speed becomes the speedometer reading (the green graph below).

Quick check: a model's loss falls from 5.0 to 3.8 over 6 training steps. What is the average rate of change, with units?

$(3.8-5.0)/6 = -0.2$ loss units per step. Negative means the loss is falling.

Derivative from first principles: secant to tangent core

We want the slope at one point. But slope needs two points (rise over run). The trick: pick a second point a small distance $h$ away and draw the straight line through both. That line is a secant. Its slope is easy: rise ÷ run.

Now slide the second point closer and closer to the first. The secant line swings and settles into the tangent. The secant's slope settles down to one number. That number is the derivative.

Why not just set $h=0$? Then rise and run are both $0$, and $0/0$ means nothing. So we do algebra first, to cancel the $h$, and then let $h$ go to $0$. That is exactly what a limit is for.

Worked example: $f(x)=x^2$. Follow every step.

  1. Write the secant slope: $\dfrac{f(x+h)-f(x)}{h} = \dfrac{(x+h)^2 - x^2}{h}$.
  2. Expand the square: $(x+h)^2 = x^2 + 2xh + h^2$.
  3. Subtract $x^2$: the top becomes $2xh + h^2$.
  4. Both terms on top contain $h$, so factor it out and cancel: $\dfrac{2xh + h^2}{h} = \dfrac{h(2x + h)}{h} = 2x + h$.
  5. Now let $h \to 0$. The leftover $h$ vanishes: the slope is $2x$.

So $\dfrac{d}{dx}x^2 = 2x$, which matches the numbers in the first section. See the trend with real numbers at $x=1$: with $h = 1$ the secant slope is $3$; with $h = 0.1$ it is $2.1$; with $h=0.01$ it is $2.01$. Each one equals $2x+h = 2+h$.

The derivative is the limit of the secant slope as the gap shrinks to zero:

$$f'(x) = \lim_{h\to 0} \frac{f(x+h)-f(x)}{h}.$$

The recipe, every time: (1) write the ratio, (2) expand and simplify until the $h$ on the bottom cancels, (3) let $h \to 0$. The more careful name for this is the difference quotient.

Why do we need it?

It is the definition: everything else (every rule) is a shortcut that was proven from it. If you can do it from first principles, you never have to trust a formula blindly.

Where is it used?

To prove the derivative rules, and in practice as the "nudge test" (finite differences) that checks gradients in a neural network: nudge a weight by $h$, see how the loss moves.

How is it used?

Write $\frac{f(x+h)-f(x)}{h}$, simplify algebraically until $h$ cancels, then set $h=0$. In code: use a small $h$ such as $10^{-5}$ and compare with your formula.

Pick a function. Drag h from large to small (or press Shrink h). The purple secant line swings into the orange tangent. In the table, each row is the secant slope for one value of $h$: watch the numbers close in on the exact derivative.

Worked example: $f(x)=x^3$.

  1. Ratio: $\dfrac{(x+h)^3 - x^3}{h}$.
  2. Expand the cube: $(x+h)^3 = x^3 + 3x^2h + 3xh^2 + h^3$.
  3. Subtract $x^3$: the top is $3x^2h + 3xh^2 + h^3$.
  4. Divide every term by $h$: $3x^2 + 3xh + h^2$.
  5. Let $h\to 0$: the terms with $h$ vanish, leaving $3x^2$.

So $\dfrac{d}{dx}x^3 = 3x^2$. Check at $x=1$: the secant slopes are $3+3h+h^2$, which is $3.31$ for $h=0.1$ and gets closer to $3$.

Worked example: $f(x)=1/x$.

  1. Ratio: $\dfrac{\frac{1}{x+h} - \frac{1}{x}}{h}$.
  2. Put the top over a common denominator: $\dfrac{1}{x+h} - \dfrac1x = \dfrac{x - (x+h)}{x(x+h)} = \dfrac{-h}{x(x+h)}$.
  3. Divide by $h$: $\dfrac{-h}{x(x+h)}\cdot\dfrac1h = \dfrac{-1}{x(x+h)}$.
  4. Let $h\to0$: $\dfrac{-1}{x\cdot x} = -\dfrac{1}{x^2}$.

So $\dfrac{d}{dx}\dfrac1x = -\dfrac1{x^2}$. At $x=2$ the slope is $-\tfrac14$. Check with $h=0.1$: $(1/2.1 - 0.5)/0.1 \approx -0.238$, close to $-0.25$. It is always negative: $1/x$ always falls as $x$ grows (for $x>0$).

Worked example: $f(x)=\sqrt{x}$. A square root on top calls for a trick: multiply top and bottom by the conjugate $\sqrt{x+h}+\sqrt{x}$. (This uses $(A-B)(A+B) = A^2-B^2$.)

  1. Ratio: $\dfrac{\sqrt{x+h}-\sqrt{x}}{h}$.
  2. Multiply top and bottom by $\sqrt{x+h}+\sqrt{x}$. The top becomes $(x+h) - x = h$.
  3. So the ratio is $\dfrac{h}{h\,(\sqrt{x+h}+\sqrt{x})}$. Cancel the $h$: $\dfrac{1}{\sqrt{x+h}+\sqrt{x}}$.
  4. Let $h\to0$: $\dfrac{1}{\sqrt{x}+\sqrt{x}} = \dfrac{1}{2\sqrt{x}}$.

So $\dfrac{d}{dx}\sqrt{x} = \dfrac{1}{2\sqrt x}$. At $x=4$ the slope is $\tfrac14$. Check with $h=0.1$: $(\sqrt{4.1}-2)/0.1 \approx 0.2485$. The slope is huge near $x=0$ (the curve starts out vertical) and flattens as $x$ grows.

You cannot skip the algebra. Plugging $h=0$ into $\frac{f(x+h)-f(x)}{h}$ gives $\frac00$. Simplify first, then let $h \to 0$.

Computers use a small, not zero, $h$. In code, $h\approx10^{-5}$ works well. Much smaller (like $10^{-15}$) gets ruined by rounding errors. Using the symmetric version $\frac{f(x+h)-f(x-h)}{2h}$ (a "central difference") is more accurate, and is what this guide's "nudge checks" use.

Quick check: use first principles to find the derivative of $f(x)=3x+5$.

$\dfrac{(3(x+h)+5) - (3x+5)}{h} = \dfrac{3h}{h} = 3$. It is $3$ for every $h$, and so $f'(x)=3$: a straight line's slope never changes.

Warm-up rules: constants, multiples and sums

Three tiny facts make big formulas easy.

  • A constant never changes, so its slope is $0$. Lifting a whole graph up by 5 does not change how steep it is anywhere.
  • Stretching a graph taller by a factor $c$ stretches every slope by $c$.
  • Adding two graphs adds their heights, and so adds their slopes.

Together: you can differentiate a long expression one piece at a time.

Differentiate $f(x) = 3x^2 + 5x - 7$. We know $\frac{d}{dx}x^2 = 2x$, $\frac{d}{dx}x = 1$ and $\frac{d}{dx}7 = 0$.

  1. The $3x^2$ piece: $3 \cdot 2x = 6x$.
  2. The $5x$ piece: $5\cdot 1 = 5$.
  3. The $-7$ piece: $0$.
  4. Add: $f'(x) = 6x + 5$.

Check at $x=2$: $f'(2)=17$. Nudge: $f(2)=15$, $f(2.001)=15.017003$, slope $\approx 17.003$ ✓.

For a constant $c$ and differentiable $f$, $g$:

$$\frac{d}{dx}c = 0, \qquad \frac{d}{dx}\big[c\,f(x)\big] = c\,f'(x), \qquad \frac{d}{dx}\big[f(x)\pm g(x)\big] = f'(x) \pm g'(x).$$

Why the sum rule is true (from first principles): the secant slope of $f+g$ is

$$\frac{[f(x+h)+g(x+h)] - [f(x)+g(x)]}{h} = \frac{f(x+h)-f(x)}{h} + \frac{g(x+h)-g(x)}{h},$$

and as $h\to0$ each piece goes to its own derivative. The constant-multiple rule works the same way: $c$ comes straight out of the ratio. The word for "these two rules together" is linearity of the derivative.

Why do we need it?

Real functions are sums of simple pieces. Linearity lets us differentiate a messy sum by handling each term alone, and drop any constant.

Where is it used?

Every loss that is a sum over data points (the derivative of a sum of losses is the sum of derivatives), regularisation terms like $\lambda w^2$ added to a loss, and polynomial models.

How is it used?

Split the expression at the plus and minus signs. Differentiate each term. Pull out constant factors. Terms with no $x$ in them vanish.

Change $c$ first: the blue curve slides up and down but the green slope graph does not move at all (constants vanish). Change $b$: the green line shifts up or down by $b$. Change $a$: the green line tilts, since $f'(x)=2ax+b$. Drag the slider for $x$ and check that the orange tangent's slope equals the green height.

Quick check: differentiate $f(x) = 4x^2 - 3x + 9$.

$4\cdot2x - 3\cdot 1 + 0 = 8x - 3$.

The power rule core

Look at what first principles gave us: $x^2 \to 2x$ and $x^3 \to 3x^2$. The recipe is a pattern:

Bring the power down in front, then lower the power by one.

So $x^4 \to 4x^3$, $x^{10}\to 10x^9$. The pattern even works for the strange powers we did by hand: $1/x = x^{-1}$ gave $-x^{-2}$, and $\sqrt x = x^{1/2}$ gave $\tfrac12 x^{-1/2}$.

  • $\dfrac{d}{dx}x^5 = 5x^4$.
  • $\dfrac{d}{dx}x = \dfrac{d}{dx}x^1 = 1\cdot x^0 = 1$.
  • $\dfrac{d}{dx}\dfrac{1}{x^2} = \dfrac{d}{dx}x^{-2} = -2x^{-3} = -\dfrac{2}{x^3}$.
  • $\dfrac{d}{dx}\sqrt[3]{x} = \dfrac{d}{dx}x^{1/3} = \tfrac13 x^{-2/3}$.
  • $\dfrac{d}{dx}\big(4x^3 - 2x^2 + 5x - 9\big) = 12x^2 - 4x + 5$. At $x=2$ this is $48 - 8 + 5 = 45$.

For any real number $n$:

$$\frac{d}{dx}x^n = n\,x^{n-1}.$$

Where it comes from. Expand $(x+h)^n$ with the binomial pattern. For small powers:

$$(x+h)^2 = x^2 + 2xh + h^2,\quad (x+h)^3 = x^3 + 3x^2h + 3xh^2 + h^3,\quad (x+h)^4 = x^4 + 4x^3h + 6x^2h^2 + 4xh^3 + h^4.$$

In every case the expansion is $x^n + n\,x^{n-1}h + (\text{terms with } h^2 \text{ or higher})$. Subtract $x^n$ and divide by $h$: you get $n\,x^{n-1} + (\text{terms that still contain } h)$. Let $h\to0$ and only $n\,x^{n-1}$ is left. For every whole number $n$ the binomial pattern always gives this, so the rule is proved for whole numbers. For $n=-1$ and $n=\tfrac12$ we proved it by hand above. For other fractions and negative powers it also holds, and we accept that without a full proof here.

Why do we need it?

Polynomials and power laws are everywhere, and differentiating $x^n$ from first principles every time would be slow. One line replaces a page of algebra.

Where is it used?

Squared-error loss $(y-\hat y)^2$, weight decay $w^2$, polynomial regression, scaling laws (loss $\propto$ size$^{-0.05}$), and the Taylor series later on.

How is it used?

Rewrite roots and fractions as powers ($\sqrt x = x^{1/2}$, $1/x^3 = x^{-3}$), multiply by the power, subtract one from the power.

Slide the power $n$ (try 2, 3, 0.5, 0, −1) and the point $x$. The readout compares the formula $n\,x^{n-1}$ with a nudge test that knows nothing about the formula. Notice: $n=1$ gives a flat green line at height 1, $n=0$ gives 0, and negative $n$ gives negative slopes.

The power rule is for "variable to a fixed power". $x^3$ is a power function. But $3^x$ (fixed base, variable exponent) is an exponential and follows a different rule (below). Do not write $\frac{d}{dx}3^x = x\,3^{x-1}$.

Quick check: differentiate $f(x)=\dfrac{5}{x^3} + 2\sqrt{x}$.

Rewrite: $5x^{-3} + 2x^{1/2}$. Then $5(-3)x^{-4} + 2\cdot\tfrac12 x^{-1/2} = -\dfrac{15}{x^4} + \dfrac{1}{\sqrt x}$.

The product rule core

Picture a rectangle whose width is $u(x)$ and whose height is $v(x)$. Its area is $u \cdot v$. Now nudge $x$ a little. Both sides grow a little. How does the area change?

  • A thin strip is added along the right edge: its length is the height $v$, its thickness is how much the width grew.
  • A thin strip is added along the top edge: its length is the width $u$, its thickness is how much the height grew.
  • A tiny square appears in the corner. It is (tiny) × (tiny), so it is far smaller than the strips and disappears when the nudge shrinks.

So the area's growth is (how fast the width grows) × (height) plus (width) × (how fast the height grows).

Let $f(x) = x^2(3x+1)$, so $u = x^2$ and $v = 3x+1$.

  1. Derivatives of the parts: $u' = 2x$, $v' = 3$.
  2. Product rule: $f' = u'v + uv' = 2x(3x+1) + x^2\cdot 3 = 6x^2 + 2x + 3x^2 = 9x^2 + 2x$.
  3. Check by multiplying out first: $f = 3x^3 + x^2$, so $f' = 9x^2 + 2x$ ✓ (same answer).
  4. At $x = 2$: $f'(2) = 36 + 4 = 40$. Nudge check: $f(2) = 28$, $f(2.001) = 28.040019$, slope $\approx 40.02$ ✓.

Not every product can be multiplied out (try $x\,e^x$ or $x^2\sin x$). That is where this rule earns its keep.

If $f(x) = u(x)\,v(x)$, then

$$f'(x) = u'(x)\,v(x) + u(x)\,v'(x).$$

Derivation from the limit. Write the secant slope and do one clever step: add and subtract $u(x+h)\,v(x)$ on top.

$$\begin{aligned} \frac{u(x+h)v(x+h) - u(x)v(x)}{h} &= \frac{u(x+h)v(x+h) - u(x+h)v(x) + u(x+h)v(x) - u(x)v(x)}{h} \\[4pt] &= u(x+h)\,\frac{v(x+h)-v(x)}{h} + v(x)\,\frac{u(x+h)-u(x)}{h}. \end{aligned}$$

Now let $h\to0$. The fraction $\frac{v(x+h)-v(x)}{h}\to v'(x)$ and $\frac{u(x+h)-u(x)}{h}\to u'(x)$. Also $u(x+h)\to u(x)$ (a smooth curve does not jump). This leaves $u(x)v'(x) + v(x)u'(x)$. ∎

In words: "derivative of the first times the second, plus the first times the derivative of the second."

Why do we need it?

Many functions multiply two changing things, and the answer is not "multiply the two derivatives". We need the correct recipe for how a product moves when both factors move.

Where is it used?

Gating in LSTMs and attention (one signal times another), the loss term $x\ln x$ in entropy, "weight times activation" terms in backprop, and physics problems where mass and speed both change (momentum $m\,v$).

How is it used?

Name the two factors $u$ and $v$. Differentiate each separately. Combine as $u'v + uv'$. Simplify at the end.

The blue rectangle has width $u=3$ and height $v=2$. Choose how fast each side grows ($u'$ and $v'$), then nudge $x$ by $dx$. The orange strip is $u'v\,dx$, the green strip is $u\,v'\,dx$, and the red corner is $u'v'dx^2$. Slide $dx$ towards zero: the red corner nearly vanishes, and $\Delta A / dx$ approaches $u'v + uv'$.

Pick a product. Slide $x$. The readout lists $u$, $v$, $u'$, $v'$ and builds $u'v+uv'$ step by step. It then compares with a nudge test on the whole product. Try to find a spot where the rule fails (it never does).

The derivative of a product is NOT the product of the derivatives. Test with $x\cdot x = x^2$: the true derivative is $2x$, but $u'v' = 1\cdot1 = 1$. The extra pieces ($u'v$ and $uv'$) are the two strips in the picture.

Quick check: differentiate $f(x)=x\,e^x$ at $x=1$.

$u=x$, $v=e^x$. $f' = 1\cdot e^x + x\,e^x = (1+x)e^x$. At $x=1$: $2e \approx 5.437$.

The quotient rule core

A fraction $\dfrac{u}{v}$ changes in two ways. If the top $u$ grows, the fraction grows. If the bottom $v$ grows, the fraction shrinks (a bigger bucket divides the same water into smaller shares). So the rule has a plus and a minus, and a squared bottom to account for dividing.

We do not need new limit work. A fraction is just a product in disguise, so the product rule will give us the answer.

Let $f(x) = \dfrac{x}{x^2+1}$, so $u = x$ and $v = x^2+1$.

  1. $u' = 1$, $v' = 2x$.
  2. Top of the rule: $u'v - uv' = 1\cdot(x^2+1) - x\cdot 2x = x^2 + 1 - 2x^2 = 1 - x^2$.
  3. Bottom of the rule: $v^2 = (x^2+1)^2$.
  4. So $f'(x) = \dfrac{1 - x^2}{(x^2+1)^2}$.

At $x=1$: $f'(1)=0$, so the graph is flat there (its peak: $f(1)=\tfrac12$). At $x=2$: $f'(2) = \dfrac{1-4}{25} = -0.12$. Nudge check: $f(2) = 0.4$, $f(2.001)\approx 0.39988$, slope $\approx-0.12$ ✓.

$$\frac{d}{dx}\left[\frac{u}{v}\right] = \frac{u'v - u\,v'}{v^2}, \qquad v \ne 0.$$

Derivation from the product rule. Call the fraction $q = u/v$. Then $u = q\,v$. Differentiate both sides with the product rule:

$$u' = q'\,v + q\,v'.$$

Solve for $q'$: $\;q' = \dfrac{u' - q\,v'}{v}$. Now put back $q = u/v$:

$$q' = \frac{u' - \frac{u}{v}v'}{v} = \frac{\frac{u'v - uv'}{v}}{v} = \frac{u'v - uv'}{v^2}. \;∎$$

Memory line: "bottom × derivative of top, minus top × derivative of bottom, all over bottom squared". A special case worth knowing: $\dfrac{d}{dx}\dfrac1v = -\dfrac{v'}{v^2}$ (take $u=1$, so $u'=0$), which matches our earlier $\frac{d}{dx}\frac1x = -\frac1{x^2}$.

Why do we need it?

Ratios appear whenever one quantity is divided by another, and dividing by something that also changes needs its own correction (the minus term).

Where is it used?

Sigmoid $1/(1+e^{-x})$, softmax (an exponential divided by a sum), normalising a vector by its length, $\tan x = \sin x/\cos x$, and rates such as "loss per sample".

How is it used?

Name top $u$ and bottom $v$. Compute $u'$ and $v'$. Form $u'v - uv'$ (order matters!) and divide by $v^2$. Often it is simplest to leave the bottom squared.

Pick a fraction and slide $x$. Watch for flat spots: $x/(x^2+1)$ has $f'=0$ at $x=1$ (where $1-x^2=0$), a peak. For $\sin x / x$ the slope is negative from the start until about $x=4.49$ (the first valley of the wave), and positive after that.

The order in the top is not optional. It is $u'v - uv'$, not $uv' - u'v$. A swap flips the sign of your answer. A quick sanity check: for $1/x$, the slope should be negative, and the rule with $u=1$ gives $-v'/v^2 = -1/x^2$ ✓.

Often you can skip the rule. $\dfrac{x^3+x}{x} = x^2+1$ is easier by simplifying first.

Quick check: differentiate $f(x) = \dfrac{x+1}{x-1}$.

$u=x+1$, $v=x-1$, $u'=v'=1$. $f' = \dfrac{1\cdot(x-1) - (x+1)\cdot1}{(x-1)^2} = \dfrac{-2}{(x-1)^2}$. It is negative everywhere (the function falls on each side of $x=1$).

The chain rule core

Imagine two machines in a row. Machine $g$ takes $x$ and makes $u$. Machine $f$ takes $u$ and makes $y$. So $x \to u \to y$. This is a composition, written $y = f(g(x))$ (you met it in Chapter 2.1).

Think of three gears in a line. Turn the first gear (the input $x$). The second gear turns, say, 3 times as fast. The third gear turns 2 times as fast as the second. So the third gear turns $3 \times 2 = 6$ times as fast as the first. Rates multiply along the chain.

That is the whole chain rule: "how fast does $y$ change when $x$ changes?" equals "how fast $u$ changes with $x$" times "how fast $y$ changes with $u$".

Let $y = (3x+1)^2$. Break it into two machines: inner $u = 3x+1$, outer $y = u^2$.

  1. Inner rate: $\dfrac{du}{dx} = 3$.
  2. Outer rate: $\dfrac{dy}{du} = 2u$.
  3. Multiply: $\dfrac{dy}{dx} = 2u \cdot 3 = 6u$.
  4. Put $u$ back: $\dfrac{dy}{dx} = 6(3x+1)$.

At $x=1$: $u = 4$, so the slope is $6\cdot4 = 24$. Nudge check: $y(1) = 16$, $y(1.001) = 16.024009$, so the slope $\approx 24.009$ ✓.

A harder one: $y = \sin(x^2)$. Inner $u = x^2$ (rate $2x$), outer $\sin u$ (rate $\cos u$). So $\dfrac{dy}{dx} = \cos(x^2)\cdot 2x$. At $x = 1$ this is $2\cos 1 \approx 1.0806$.

Three links: $y = e^{\sin(x^2)}$. Let $u = x^2$, $v = \sin u$, $y = e^v$. Then $\dfrac{dy}{dx} = \dfrac{dy}{dv}\cdot\dfrac{dv}{du}\cdot\dfrac{du}{dx} = e^v\cdot\cos u\cdot 2x$. At $x=1$: $e^{\sin 1}\cos 1 \cdot 2 \approx 2.507$. Just keep multiplying rates, one per link.

If $y = f(u)$ and $u = g(x)$, then

$$\frac{dy}{dx} = \frac{dy}{du}\cdot\frac{du}{dx}, \qquad\text{or}\qquad \big(f(g(x))\big)' = f'(g(x))\cdot g'(x).$$

Why it is true. Nudge $x$ by $\Delta x$. That nudges $u$ by $\Delta u$, which nudges $y$ by $\Delta y$. Always

$$\frac{\Delta y}{\Delta x} = \frac{\Delta y}{\Delta u}\cdot\frac{\Delta u}{\Delta x}$$

(the $\Delta u$ on top and bottom cancel, like ordinary fractions). Now let $\Delta x \to 0$. Then $\Delta u\to0$ too, and the three ratios become $\frac{dy}{dx}$, $\frac{dy}{du}$ and $\frac{du}{dx}$. ∎ (When $g$ happens to leave $u$ unchanged, $\Delta u = 0$, a more careful proof is needed, but the result is the same.)

Recipe: (1) name the inner function $u$; (2) differentiate the outer function with respect to $u$; (3) differentiate the inner function; (4) multiply; (5) replace $u$. Chains can be longer: just multiply one rate per link.

This is the single-variable chain rule. When there are many inputs, it becomes a sum of products over paths (Chapter 2.8), and applying it layer after layer is backpropagation (Chapter 2.9).

Why do we need it?

Almost every function in ML is built by feeding one function into another. We need a way to get the slope of the whole machine from the slopes of its parts.

Where is it used?

Backpropagation (a neural network is a long chain of layers), the derivative of $(y-\hat y)^2$ with respect to a weight, the sigmoid's derivative, $\ln$ of a probability, and every automatic-differentiation library.

How is it used?

Split the formula into layers from the inside out. Write the slope of each layer. Multiply all the slopes together. Each layer only needs to know its own local slope.

Pick a machine and an $x$. Turn the x-crank slider. The second gear turns $g'(x)$ times as much, and the third turns $f'(u)$ times as much as the second, so overall $g'(x)\cdot f'(u)$ times as much as the first. A negative rate means that gear turns the opposite way. The readout checks this with a tiny nudge.

Choose a composite function. The readout names the inner and outer parts and multiplies their slopes. Check that it matches the nudge test everywhere, including where the slope is 0.

Do not forget the inner derivative. The most common mistake: $\frac{d}{dx}\sin(x^2) = \cos(x^2)$. It is missing the factor $2x$. The outer rule alone only handles the outer machine.

Product rule or chain rule? If two pieces are multiplied, use the product rule. If one piece is inside the other, use the chain rule. Many problems need both.

Quick check: differentiate $y = (x^2+1)^3$.

Inner $u = x^2+1$, $u' = 2x$. Outer $u^3$, derivative $3u^2$. So $y' = 3(x^2+1)^2\cdot 2x = 6x(x^2+1)^2$. At $x=1$: $6\cdot4 = 24$.

Exponential derivatives: why $e^x$ is special core

Money in a bank with interest, bacteria in a dish, a rumour on social media: in all of them the more you have, the faster it grows. The growth rate is proportional to the current size. That is the nature of an exponential function.

So the slope of $a^x$ should be a fixed multiple of its height. For some special base the multiple is exactly 1: the slope equals the height at every point. That base is the number $e \approx 2.71828$. The curve $e^x$ is its own derivative.

How steep is $a^x$ at $x=0$? Use the nudge $\dfrac{a^h - 1}{h}$ (because $a^0=1$):

base $a$$h=0.1$$h=0.01$$h=0.001$settles near
$2$0.71770.69560.69340.6931
$e$1.05171.00501.00051
$3$1.16121.10471.09921.0986

Base 2 is a little too gentle (slope 0.693), base 3 a little too steep (1.099). Somewhere between them, at $e$, the slope at $0$ is exactly $1$.

Derivation. Start from first principles for $a^x$:

$$\frac{a^{x+h} - a^x}{h} = \frac{a^x\,a^h - a^x}{h} = a^x\cdot\frac{a^h - 1}{h}.$$

The fraction $\frac{a^h-1}{h}$ does not depend on $x$. As $h\to0$ it goes to a constant $k(a)$ (the slope of $a^x$ at $x=0$, the numbers in the table). So $\dfrac{d}{dx}a^x = k(a)\,a^x$: the slope is always a fixed multiple of the height. The number $e$ is defined as the base for which $k(e)=1$. Hence

$$\frac{d}{dx}e^x = e^x.$$

Other bases. Write $a = e^{\ln a}$, so $a^x = e^{x\ln a}$. The chain rule (inner $u = x\ln a$, with rate $\ln a$) gives

$$\frac{d}{dx}a^x = (\ln a)\,a^x, \qquad\text{so } k(a) = \ln a \;(\text{matches: } \ln 2 = 0.693,\ \ln 3 = 1.099).$$

With a chain: $\dfrac{d}{dx}e^{g(x)} = g'(x)\,e^{g(x)}$. For example $\dfrac{d}{dx}e^{2x} = 2e^{2x}$, $\dfrac{d}{dx}e^{-x} = -e^{-x}$, $\dfrac{d}{dx}e^{-x^2} = -2x\,e^{-x^2}$.

Why do we need it?

Growth and decay, probabilities and softmax all use exponentials. The special fact "$e^x$ is its own slope" makes their derivatives clean instead of messy.

Where is it used?

Softmax and sigmoid (both are built from $e^x$), exponential learning-rate decay, the Gaussian bell curve $e^{-x^2/2}$, radioactive decay and compound interest.

How is it used?

For $e^{g(x)}$: copy the function and multiply by $g'(x)$. For $a^x$: copy it and multiply by $\ln a$.

Slide the base $a$. The blue curve is $a^x$ and the green curve is its slope. If green is below blue, the base is below $e$; if above, the base is above $e$. Press a = e: the two curves lie exactly on top of each other. Drag $x$: the ratio slope ÷ height never changes.

Pick an exponential. See "copy the function, multiply by the inner slope (or by ln a)", then check it against the nudge test.

$e^x$ is not a power function. The variable is in the exponent. Never write $\frac{d}{dx}e^x = x\,e^{x-1}$. And $\frac{d}{dx}a^x$ is not just $a^x$ unless $a=e$: it has the extra factor $\ln a$.

Quick check: differentiate $f(x) = 5e^{3x}$, and find the slope of $2^x$ at $x=3$.

$f' = 5\cdot3e^{3x} = 15e^{3x}$. For $2^x$: $(\ln 2)\,2^3 = 8\ln2 \approx 5.545$.

Logarithmic derivatives: $\ln x$ and friends core

$\ln x$ answers: "to what power must I raise $e$ to get $x$?" It undoes $e^x$: it is the inverse function (Chapter 2.1). The graph of an inverse is the original graph flipped across the diagonal line $y=x$.

Flip a hill and its steepness flips too: "rise over run" becomes "run over rise". So the slope of the inverse is one over the slope of the original. Where $e^x$ is very steep, $\ln x$ is very flat, and the other way around.

That already hints at the answer: $\ln x$ is steep for small $x$ and gets flatter and flatter. It grows, but slowly.

  • At $x = 2$: the slope of $\ln x$ is $\tfrac12$. Nudge check: $(\ln 2.001 - \ln 2)/0.001 \approx 0.49988$ ✓.
  • At $x = 0.1$: slope $= 10$ (very steep, just after the curve starts).
  • At $x = 10$: slope $= 0.1$ (nearly flat).
  • Picture check: the point $(1, e)$ on $e^x$ has slope $e$. Its mirror image $(e, 1)$ on $\ln x$ has slope $1/e$, which is $1/x$ at $x=e$ ✓.
$$\frac{d}{dx}\ln x = \frac{1}{x}\quad (x>0), \qquad \frac{d}{dx}\log_a x = \frac{1}{x\ln a}.$$

Derivation. Let $y = \ln x$. By definition $e^y = x$. Differentiate both sides with respect to $x$. On the left use the chain rule ($y$ depends on $x$); on the right, $\frac{d}{dx}x = 1$:

$$e^y\cdot\frac{dy}{dx} = 1 \;\Longrightarrow\; \frac{dy}{dx} = \frac{1}{e^y} = \frac1x.\;∎$$

Other bases. $\log_a x = \dfrac{\ln x}{\ln a}$ (change of base), and $\ln a$ is just a constant, so $\dfrac{d}{dx}\log_a x = \dfrac{1}{x\ln a}$. For example the slope of $\log_2 x$ at $x=4$ is $\dfrac{1}{4\ln2}\approx0.3607$.

With a chain: $\dfrac{d}{dx}\ln g(x) = \dfrac{g'(x)}{g(x)}$. For example $\dfrac{d}{dx}\ln(x^2+1) = \dfrac{2x}{x^2+1}$, which is $1$ at $x=1$. The quantity $g'/g$ is the relative rate of change (change as a fraction of the current size). It is one reason the log of a likelihood, a product of many probabilities, is so convenient to differentiate (you will meet this in Chapter 2.14).

Why do we need it?

Logs turn products into sums and tame huge or tiny numbers, so ML uses them constantly. We need their slopes to train anything that contains a log.

Where is it used?

Cross-entropy loss $-\ln p$ and log-likelihood (the loss $-\ln p$ has slope $-1/p$: huge when $p$ is tiny, which punishes confident mistakes), softmax, entropy, and log-scaled features.

How is it used?

For $\ln g(x)$: put $g'(x)$ on top of $g(x)$. For $\log_a$: divide by $\ln a$ too. Remember $g(x)$ must be positive.

Slide $a$. The blue point $(a, e^a)$ sits on $e^x$ and its mirror image (orange) $(e^a, a)$ sits on $\ln x$. The dashed purple line between them is bisected by the diagonal. Compare the two slopes in the readout: they are reciprocals, and the orange slope is always $1/x$.

Pick a log function. Slide $x$ and compare the rule with the nudge test. Notice how the slope of $\ln x$ (green) falls towards 0 as $x$ grows, and how $x\ln x$ has slope 0 at $x = 1/e \approx 0.37$ (its lowest point).

Only positive numbers have a log. $\ln x$ and its slope $1/x$ exist only for $x>0$. And $\frac{d}{dx}\ln x = \frac1x$ is not $\ln$ of anything: it is a plain fraction.

Quick check: differentiate $f(x) = \ln(5x)$. What do you notice?

Chain rule: $\dfrac{5}{5x} = \dfrac1x$. The same as $\ln x$. (Indeed $\ln 5x = \ln 5 + \ln x$, and the constant $\ln 5$ has slope $0$.)

Trigonometric derivatives: sin, cos, tan core

Walk anticlockwise round a circle of radius 1 at speed 1. Your height is $\sin x$ and your sideways position is $\cos x$, where $x$ is how far round you have gone (an angle in radians).

At each moment you move along the tangent, at speed 1. How fast is your height changing? Only the upward part of your motion counts. At the right-hand side of the circle you are moving straight up (height changes fastest), and at the very top you are moving sideways (height momentarily not changing).

That upward part of the motion is exactly $\cos x$. So the slope of $\sin x$ is $\cos x$. By the same reasoning your sideways position changes at rate $-\sin x$: the slope of $\cos x$ is $-\sin x$ (negative while you are in the top half, moving left).

  • At $x = 0$: $\sin$ rises steeply, slope $\cos 0 = 1$.
  • At $x = \pi/2$ (the top of the wave): slope $\cos(\pi/2) = 0$. Flat.
  • At $x = \pi$: slope $\cos\pi = -1$, falling steeply through zero.
  • Small-angle check: $\sin(0.001)/0.001 = 0.9999998 \approx 1$ ✓.

The slope graph of $\sin$ is the cosine wave, shifted a quarter-turn. And the slope graph of cosine is the sine wave flipped upside down.

$$\frac{d}{dx}\sin x = \cos x,\qquad \frac{d}{dx}\cos x = -\sin x,\qquad \frac{d}{dx}\tan x = \frac{1}{\cos^2 x} = 1 + \tan^2 x.$$

Derivation of $\sin'$. Use the angle-addition formula $\sin(x+h) = \sin x\cos h + \cos x\sin h$:

$$\frac{\sin(x+h)-\sin x}{h} = \sin x\cdot\frac{\cos h - 1}{h} + \cos x\cdot\frac{\sin h}{h}.$$

Two limits do all the work. From Chapter 2.2 we know $\frac{\sin h}{h}\to1$ and $\frac{1-\cos h}{h^2}\to\frac12$. The second one gives $\frac{\cos h-1}{h} = -h\cdot\frac{1-\cos h}{h^2}\to 0\cdot\frac12 = 0$. Numerically, the two ratios in the formula above: at $h=0.1$ they are $0.99833$ and $-0.04996$; at $h=0.01$ they are $0.99998$ and $-0.00500$; at $h=0.001$ they are $0.9999998$ and $-0.0005$. So the sum goes to $\sin x\cdot 0 + \cos x\cdot 1 = \cos x$. ∎

Derivation of $\cos'$. Note $\cos x = \sin(\tfrac\pi2 - x)$. Chain rule: outer $\sin$ gives $\cos(\tfrac\pi2 - x)$, inner $\tfrac\pi2 - x$ has slope $-1$. So $(\cos x)' = -\cos(\tfrac\pi2 - x) = -\sin x$. ∎

Derivation of $\tan'$. $\tan x = \dfrac{\sin x}{\cos x}$. Quotient rule with $u=\sin x$, $v=\cos x$:

$$\frac{\cos x\cdot\cos x - \sin x\cdot(-\sin x)}{\cos^2 x} = \frac{\cos^2x + \sin^2x}{\cos^2x} = \frac{1}{\cos^2 x}.\;∎$$

(using $\cos^2x+\sin^2x = 1$). At $x=\pi/4$ this is $1/(\tfrac{1}{\sqrt2})^2 = 2$.

Why do we need it?

Anything that repeats (waves, rotations, seasons) is described by sin and cos, and we want to know how fast those quantities change.

Where is it used?

Positional encodings in Transformers (sines and cosines of different frequencies), signal processing and Fourier features, cosine learning-rate schedules, and rotation matrices.

How is it used?

Use $(\sin)'=\cos$ and $(\cos)'=-\sin$ together with the chain rule: $\frac{d}{dx}\sin(3x) = 3\cos(3x)$.

Drag the dot round the circle (or use the slider). The orange arrow is the horizontal part of your motion (it equals $-\sin x$) and the green arrow is the vertical part ($\cos x$). On the right the tangent to the sine wave has exactly that slope. Find where the green arrow vanishes: the sine wave is at its top or bottom.

Pick a trig function and slide $x$ (in radians). Compare the rule with the nudge test. Look at $\tan x$: its slope $1 + \tan^2 x$ is never below 1, and explodes near $x=\pm\pi/2$.

Radians only! These formulas assume $x$ is in radians. In degrees there is a stray factor: $\frac{d}{dx}\sin(x^\circ) = \frac{\pi}{180}\cos(x^\circ)$. Always do calculus in radians.

Signs. Only $\cos$ has a minus sign in its derivative: $(\sin)'=\cos$ but $(\cos)'=-\sin$. After four derivatives you are back where you started.

Quick check: differentiate $f(x) = x\cos x$.

Product rule: $1\cdot\cos x + x\cdot(-\sin x) = \cos x - x\sin x$.

The rules at a glance: a derivative table and picker

You now own a small toolbox. There are only four combining rules (sum, product, quotient, chain) and a short list of basic derivatives. Any formula built from these pieces can be differentiated by splitting it up and applying them one at a time.

You do not need to memorise the table. Every line in it was derived above, and you can re-derive it. The table is just a place to look things up.

Differentiate $f(x) = x^2 e^{3x}\sin x$? Break it into the pieces it is made of.

  1. It is a product of three things. Group as $u = x^2e^{3x}$ and $v=\sin x$. Then $f' = u'v + uv'$.
  2. For $u = x^2\cdot e^{3x}$, the product rule again: $u' = 2x\,e^{3x} + x^2\cdot3e^{3x}$ (chain rule for $e^{3x}$).
  3. And $v' = \cos x$.
  4. So $f'(x) = \big(2x + 3x^2\big)e^{3x}\sin x + x^2e^{3x}\cos x$.

Each step used only a rule from this chapter.

Combining rules (with $c$ a constant):

RuleFormula
Constant multiple$(c\,f)' = c\,f'$
Sum / difference$(f\pm g)' = f'\pm g'$
Product$(uv)' = u'v + uv'$
Quotient$(u/v)' = \dfrac{u'v - uv'}{v^2}$
Chain$\big(f(g(x))\big)' = f'(g(x))\,g'(x)$

Basic derivatives:

$f(x)$$f'(x)$$f(x)$$f'(x)$
$c$$0$$\ln x$$1/x$
$x^n$$n\,x^{n-1}$$\log_a x$$\dfrac{1}{x\ln a}$
$e^x$$e^x$$\sin x$$\cos x$
$a^x$$(\ln a)\,a^x$$\cos x$$-\sin x$
$1/x$$-1/x^2$$\tan x$$1/\cos^2 x$
$\sqrt x$$\dfrac{1}{2\sqrt x}$$\sigma(x)$$\sigma(x)(1-\sigma(x))$
$\tanh x$$1-\tanh^2x$$\max(0,x)$$0$ or $1$ (kink at $0$)
Why do we need it?

A short, trusted list means you can differentiate nearly any formula you meet without going back to limits each time.

Where is it used?

Deriving gradients by hand for a new loss function, checking what automatic differentiation (autograd) in PyTorch and JAX should return, and exam-style or interview-style derivations.

How is it used?

Split the formula into sums, products, fractions and nested layers. Look up each basic piece, then combine. Always confirm with a nudge test on a number.

Choose a function from the menu and slide $x$. Left: the function and its tangent. Right: the derivative graph with the current slope as a dot. The readout gives the formula for $f'$ and checks it with a nudge test. Pick ReLU or $|x|$ and slide through 0 to meet a kink.

Quick check: differentiate $f(x) = \dfrac{\ln x}{x}$.

Quotient rule, $u=\ln x$, $v=x$: $f' = \dfrac{\frac1x\cdot x - \ln x\cdot 1}{x^2} = \dfrac{1-\ln x}{x^2}$. It is zero at $x=e$ (the peak of $\ln x/x$).

Derivatives of sigmoid and ReLU core

Neural networks put a little bend into every neuron using an activation function. Two of the most common:

  • Sigmoid $\sigma(x)$ squashes any number into the range $0$ to $1$ with a smooth S shape. It is steepest in the middle and almost flat in both tails.
  • ReLU is $\max(0,x)$: flat at $0$ for negative inputs, then a straight line of slope $1$. It has a sharp kink at $0$.

The slope of the activation decides how much of the learning signal gets through. A flat part blocks it: no slope, no learning. A sigmoid in its flat tail is "saturated", and a ReLU on its negative side is "off".

Sigmoid values and slopes, using $\sigma'(x)=\sigma(x)\big(1-\sigma(x)\big)$ (derived below):

$x$$\sigma(x)$$\sigma'(x)$meaning
$0$0.50000.2500the steepest point
$2$0.88080.1050still learning
$5$0.99330.0066almost flat
$-5$0.00670.0066almost flat (same, by symmetry)
$10$0.999950.000045saturated: signal nearly dead

Stacking makes it worse. Even in the best case the slope is $0.25$. Through $5$ sigmoid layers the signal is multiplied by at most $0.25^5 \approx 0.001$, and through $10$ layers by $0.25^{10}\approx 9.5\times10^{-7}$ (counting only the sigmoid slopes; the weights multiply the signal too). This is the vanishing gradient problem.

ReLU: slope $0$ for $x\lt 0$, slope $1$ for $x>0$. Passing through 10 active ReLUs multiplies the signal by $1^{10}=1$ (again counting only the activation slopes): nothing is lost to the activations.

Sigmoid. $\sigma(x) = \dfrac{1}{1+e^{-x}} = (1+e^{-x})^{-1}$. Derivation with the chain rule (inner $u = 1+e^{-x}$, outer $u^{-1}$):

$$\sigma'(x) = -(1+e^{-x})^{-2}\cdot(-e^{-x}) = \frac{e^{-x}}{(1+e^{-x})^2}.$$

Split the fraction into two pieces:

$$\sigma'(x) = \frac{1}{1+e^{-x}}\cdot\frac{e^{-x}}{1+e^{-x}} = \sigma(x)\cdot\big(1-\sigma(x)\big),$$

because $1-\sigma = \dfrac{(1+e^{-x}) - 1}{1+e^{-x}} = \dfrac{e^{-x}}{1+e^{-x}}$. ∎ The slope never exceeds $\tfrac14$ (at $x=0$, where $\sigma=\tfrac12$ and $\tfrac12\cdot\tfrac12 = \tfrac14$). A close cousin: $\tanh'(x) = 1-\tanh^2(x)$, with maximum slope $1$.

ReLU. $\text{ReLU}(x) = \max(0,x)$ has $\text{ReLU}'(x) = 0$ for $x\lt 0$ and $1$ for $x>0$. At $x=0$ the two one-sided slopes ($0$ from the left, $1$ from the right) disagree, so the derivative does not exist at exactly $0$. Libraries just pick a value (usually $0$). Leaky ReLU uses a small slope $\alpha$ (like $0.01$) for $x\lt 0$, so the signal never fully dies.

Why do we need it?

The activation's slope is a factor in every gradient flowing back through the network. We must know where it is large (learning passes) and where it is near 0 (learning stalls).

Where is it used?

Backpropagation through every neural network layer, logistic regression (sigmoid), LSTM and GRU gates, and choosing activations such as ReLU, Leaky ReLU, GELU or tanh to avoid vanishing gradients.

How is it used?

During training, the framework multiplies the incoming gradient by $\sigma'(z)=\sigma(z)(1-\sigma(z))$ (re-using the already computed $\sigma(z)$), or by 0/1 for ReLU. You watch for saturated or "dead" units.

Drag the dot along the S-curve, or the green dot on the slope graph. In the middle the slope reaches its maximum $0.25$. Slide out to $\pm5$ and beyond: the slope collapses to nearly 0, and the readout warns that the unit is saturated.

Each row shows how much of the learning signal survives after that many layers (bar length is on a log scale, the number is exact). With sigmoid at $z=0$ it shrinks by 4 times per layer. Push $z$ to 3 or more and it dies much faster. Switch to ReLU or tanh and compare. Set ReLU's $z$ below 0 to see a "dead" unit.

Drag the dot along the ReLU curve and slide it to exactly $x=0$ (it snaps to tenths). Everywhere else there is a clean tangent. At $0$ the two dashed lines are the slopes from the left and from the right, and they disagree. Switch to Leaky ReLU and change $\alpha$: the left slope is no longer 0.

Large inputs, not small ones, kill a sigmoid. The slope is largest near $x=0$. If a layer's inputs are huge (say $\pm 10$), the sigmoid is saturated whatever the weights do. This is one reason inputs are normalised.

"The derivative does not exist" is not a disaster. ReLU's kink is a single point. In practice you almost never land exactly on it, and picking either one-sided slope works.

Quick check: what is $\sigma'(3)$, given $\sigma(3)\approx0.9526$?

$\sigma'(3) = 0.9526\times(1-0.9526) = 0.9526\times0.0474 \approx 0.0452$. About 18% of the best-case $0.25$.

Higher-order derivatives: $f''$ and $f'''$

A derivative is itself a function, so we can take its derivative. Back to the car: distance $\to$ (derivative) speed $\to$ (derivative) acceleration. Acceleration is "how fast the speed is changing", the second derivative of distance.

On a graph the second derivative tells you how the curve bends:

  • $f''>0$: the slope is increasing, so the curve bends upward like a bowl or smile.
  • $f''\lt 0$: the slope is decreasing, so the curve bends downward like a frown or hilltop.
  • $f''=0$: momentarily straight (often where the bend switches).

This is a short preview. Curvature and its multi-variable version, the Hessian, get their full treatment in Chapter 2.10.

Take $f(x) = x^4 - 2x^3$.

  1. $f'(x) = 4x^3 - 6x^2$.
  2. $f''(x) = 12x^2 - 12x$ (differentiate $f'$).
  3. $f'''(x) = 24x - 12$.
  4. $f^{(4)}(x) = 24$, and every derivative after that is $0$.

At $x=1$: $f=-1$, $f'=-2$ (falling), $f''=0$ (the bend switches here), $f'''=12$. For $\sin x$ the derivatives cycle: $\sin,\ \cos,\ -\sin,\ -\cos,\ \sin,\dots$ Every derivative of $e^x$ is $e^x$. For the car $s=3t^2-t^3/3$: speed $v = 6t - t^2$, acceleration $a = 6-2t$.

The second derivative is the derivative of the derivative:

$$f''(x) = \frac{d}{dx}\big[f'(x)\big] = \frac{d^2f}{dx^2}.$$

The third is $f'''(x) = \dfrac{d^3f}{dx^3}$, and the $n$-th is written $f^{(n)}(x)$ or $\dfrac{d^nf}{dx^n}$. Read $\dfrac{d^2f}{dx^2}$ as "d two f, d x squared". Meaning: $f'$ is the slope, $f''$ is the rate at which the slope changes (the bend), $f'''$ is the rate at which the bend changes.

Why do we need it?

The slope says which way is downhill but not how quickly the ground bends. Knowing the bend tells us whether we are in a valley or on a hilltop, and how big a step is safe.

Where is it used?

The second-derivative test for minima, Newton's method, curvature of loss surfaces (the Hessian), choosing learning rates, Taylor approximations, and acceleration in physics.

How is it used?

Differentiate once, then differentiate the result. Look at the sign of $f''$: positive means a valley-shaped bend, negative means a hill-shaped bend.

Slide $x$. The top graph is $f$, the middle is $f'$ (its slope), the bottom is $f''$ (the slope of the middle graph). Where $f''>0$ the top curve bends like a smile and the middle graph is rising. Find the points where $f''$ crosses zero: the bend switches there.

Quick check: find $f''(x)$ for $f(x)=x^3-3x$ and say where the curve bends upward.

$f' = 3x^2 - 3$, $f'' = 6x$. Positive for $x>0$: bends upward there. Negative for $x\lt 0$: bends downward.

Machine learning link 1: a loss is a function of a parameter

A model has knobs called parameters (or weights). For each setting of the knobs, the model makes predictions, and a loss gives one number saying how wrong they are: big = bad, small = good.

Fix the data and turn one knob $w$. The loss changes with $w$, so the loss is just a function $L(w)$, drawn as a curve. Learning means finding the $w$ at the bottom of that curve. And the derivative $L'(w)$ at your current spot tells you which way the bottom is, and how steep the ground is.

Three data points $(x,y) = (1,2),\ (2,4),\ (3,5)$ and a model with one knob: $\hat y = w\,x$. The loss is the mean squared error $L(w) = \tfrac13\sum (w x_i - y_i)^2$.

  1. Expand: $\tfrac13\big[(w-2)^2 + (2w-4)^2 + (3w-5)^2\big] = \tfrac13\big[14w^2 - 50w + 45\big]$.
  2. Differentiate with the power rule: $L'(w) = \tfrac13(28w - 50)$.
  3. Check the sign: $L(0) = 15$, $L'(0) = -16.67$. Negative slope: turning $w$ up lowers the loss. Also $L(1)=3$, $L'(1)=-7.33$; $L(2) = 0.33$, $L'(2) = 2$ (positive: now too big).
  4. The slope is zero at $28w = 50$, so $w^* = 25/14 \approx 1.786$: the bottom of the curve.

Same answer by the chain rule, for any data: each term is $(w x_i - y_i)^2$. The outer power gives $2(w x_i - y_i)$ and the inner part has slope $x_i$, so $L'(w) = \tfrac{2}{n}\sum x_i(w x_i - y_i)$. Check: $\tfrac23[(w-2) + 2(2w-4) + 3(3w-5)] = \tfrac23(14w - 25) = \tfrac13(28w-50)$ ✓.

A loss function $L(w)$ measures how badly a model with parameter $w$ fits the data. Training means finding $w$ that makes $L$ small. The derivative

$$L'(w) = \frac{dL}{dw}$$

is the sensitivity of the loss to the parameter: how much $L$ changes per unit change in $w$. For mean squared error with a one-weight line $\hat y = wx$: $L(w)=\frac1n\sum_i (w x_i - y_i)^2$ and $L'(w)=\frac2n\sum_i x_i\,(w x_i - y_i)$. (Many weights at once: Chapter 2.4.)

Why do we need it?

"Learn from data" has to become a precise question: which parameter value is best? A loss turns that into "find the lowest point of a curve", and calculus knows how to look for it.

Where is it used?

Every supervised model: linear and logistic regression, neural networks, and language models (cross-entropy loss). Mean squared error is used for regression, cross-entropy for classification.

How is it used?

Write the loss as a function of the parameters, compute its derivative, and use the derivative to move the parameters downhill (next sections).

Left: three data points (drag them up and down) and the line $\hat y = wx$. Red bars are the errors. Right: the loss $L(w)$ for every $w$. Slide $w$ to move along the loss curve: the orange tangent shows $L'(w)$. Find the bottom (green dot). Then drag a data point and watch the whole loss curve and the best $w^*$ shift.

The loss is a function of the parameters, not of the data. The data is fixed while we train. We ask how the loss changes as $w$ changes. (The data points in the widget are movable only so you can see how the curve depends on them.)

Quick check: with the loss $L(w) = (w-4)^2$, what is $L'(1)$ and which way should $w$ move?

$L'(w) = 2(w-4)$, so $L'(1) = -6$. Negative slope means increasing $w$ lowers the loss: move $w$ up, towards 4.

Machine learning link 2: gradient descent in one dimension core

You are on a foggy hillside and want the valley floor. You cannot see it, but you can feel the slope under your feet. If the ground tilts up to your right, step left. If it tilts up to your left, step right. Always step against the slope.

How big a step? Steep ground: take a bigger step (you are far from the bottom). Almost flat: take a tiny step (you are close). So step size = learning rate × slope. The slope shrinks as you near the bottom, so the steps shrink automatically.

The learning rate is how bold you are. Too timid and you crawl. Too bold and you leap clear over the valley and land higher than before.

Minimise $L(w) = (w-3)^2$, with derivative $L'(w) = 2(w-3)$. Start at $w_0=0$ with learning rate $\eta = 0.1$ ($\eta$ is the Greek letter "eta").

  1. Slope at $0$: $L'(0) = 2(0-3) = -6$. Step: $-\eta L' = -0.1\cdot(-6) = +0.6$. New $w_1 = 0.6$.
  2. Slope at $0.6$: $2(0.6-3) = -4.8$. Step $= +0.48$. $w_2 = 1.08$.
  3. Slope at $1.08$: $2(1.08-3) = -3.84$. Step $= +0.384$. $w_3 = 1.464$.
  4. The steps shrink ($0.6, 0.48, 0.384, \dots$) and $w$ creeps towards $3$.

The distance to the bottom shrinks by a factor $(1 - 2\eta) = 0.8$ every step: $3 \to 2.4 \to 1.92 \to 1.536$. In general: for $0\lt \eta\lt 0.5$ the approach is smooth; at $\eta=0.5$ you land on $3$ in one step; for $0.5\lt \eta\lt 1$ you zig-zag across the bottom but still shrink in; at $\eta=1$ you bounce between two points forever; for $\eta>1$ the distance grows by $|1-2\eta|>1$ each step and you diverge.

Gradient descent (1D) repeats the update

$$w_{\text{new}} = w_{\text{old}} - \eta\,L'(w_{\text{old}}).$$

$\eta>0$ is the learning rate (a number you choose). The minus sign sends you downhill. It stops (in practice) when $L'$ is almost $0$, or after a set number of steps. This is the core move of training; with many weights, $L'$ becomes the gradient vector (Chapter 2.4), but the idea is identical.

Why do we need it?

Most losses cannot be solved by algebra. But we can always compute the slope at the current point. Following the slope downhill finds a low point without solving anything.

Where is it used?

Training linear and logistic regression, neural networks and deep learning (as SGD, Adam and friends), matrix factorisation, and almost any model fitted by minimising a loss.

How is it used?

Pick a start and a learning rate. Repeat: compute the slope, subtract $\eta\times$slope from the parameter. Watch the loss curve: too slow means raise $\eta$, exploding means lower it.

Drag the start point (orange dot) anywhere on the curve. Press Step a few times, or Run 25 steps. Try the four presets for the learning rate $\eta$: too small crawls, just right glides in, zig-zag bounces across the valley but still converges, and diverge climbs out of the valley. The small graph underneath is the learning curve (loss against step number).

The learning rate $\eta$ must be tuned. The "safe" range depends on how curved the loss is. For $L = (w-3)^2$ it is $\eta\lt 1$. For a sharper bowl like $L=10(w-3)^2$ the limit is $\eta\lt 0.1$. The general rule for a parabola-shaped loss: if its second derivative (the bend, $L''$) is a constant, every step multiplies the distance to the bottom by $1-\eta L''$, so descent converges only when $\eta\lt 2/L''$. Here $L''=2$ and $L''=20$. Loss curves that shoot up are the classic sign of a learning rate that is too high.

Gradient descent finds a flat spot, not necessarily the best one. On a curve with several valleys it can end in a shallow one. See the next section.

Quick check: one step of gradient descent on $L(w) = w^2$ from $w=4$ with $\eta=0.25$.

$L'(w) = 2w = 8$. New $w = 4 - 0.25\cdot 8 = 2$. The loss fell from $16$ to $4$.

Machine learning link 3: optimisation, where the derivative is zero core

Walk along a hilly path. At the very top of a hill you stop going up and start going down. At the bottom of a valley you stop going down and start going up. At both turning points the path is, for an instant, flat. So the slope is zero at every hilltop and every valley bottom.

That gives a strategy for finding the best value of anything: find where the derivative is zero. Then tell hill from valley by how the slope changes as you walk through:

  • slope goes from $+$ to $-$ (up, then down): a hilltop (maximum);
  • slope goes from $-$ to $+$ (down, then up): a valley (minimum);
  • slope touches 0 but keeps its sign: a flat shelf, neither.

Find the turning points of $f(x) = x^3 - 3x$.

  1. $f'(x) = 3x^2 - 3 = 3(x-1)(x+1)$. It is zero at $x=-1$ and $x=1$.
  2. Test a point in each region. $x=-2$: $f' = 9>0$ (rising). $x=0$: $f' = -3\lt 0$ (falling). $x=2$: $f' = 9>0$ (rising).
  3. At $x=-1$ the slope goes $+\to-$: a local maximum, $f(-1) = 2$.
  4. At $x=1$ the slope goes $-\to+$: a local minimum, $f(1) = -2$.
  5. Second-derivative check: $f''(x)=6x$. $f''(-1)=-6\lt 0$ (bends down: hill ✓). $f''(1)=6>0$ (bends up: valley ✓).

A shelf: $f(x)=x^3$ has $f'(0)=0$, but $f'=3x^2$ is positive on both sides, so $x=0$ is neither a max nor a min.

A critical point (or stationary point) is an $x$ where $f'(x)=0$. At a critical point:

  • First-derivative test: look at the sign of $f'$ just left and right. $+\to-$ is a local max; $-\to+$ is a local min; no sign change means neither.
  • Second-derivative test: if $f''>0$ it is a local min, if $f''\lt 0$ a local max, and if $f''=0$ the test is inconclusive.

A local minimum is the lowest point in its own neighbourhood. The global minimum is the lowest point of the whole function. Local minima need not be global.

Why do we need it?

Training a model is an optimisation problem: find the parameters with the lowest loss. "Slope equals zero" is the signpost that says "you may have arrived".

Where is it used?

Closed-form solutions (the normal equation of linear regression comes from setting the loss slope to 0), stopping rules for gradient descent ("stop when the gradient is tiny"), and understanding local minima and saddle points in neural-network loss landscapes.

How is it used?

Compute $f'$, solve $f'=0$ if you can, and classify each solution by the sign change or by $f''$. If you cannot solve it, run gradient descent until $f'\approx0$.

Drag the dot along the curve (or along the slope graph). The slope graph underneath crosses zero at every turning point. Press Mark critical points to label them. Then, on two valleys, press Roll downhill from here from several starting spots: gradient descent ends in different valleys. One valley is the global minimum, the other only a local one.

$f'=0$ is necessary, not sufficient. It could be a max, a min, or a shelf. Always classify.

Local is not global. Gradient descent on a bumpy loss may settle in a shallow valley. Neural-network losses have many flat spots; the surprise of modern deep learning is that simple gradient descent works well anyway.

Quick check: find and classify the critical point of $f(x) = x^2 - 6x + 5$.

$f' = 2x - 6 = 0$ gives $x=3$. $f''=2>0$, so it is a minimum, with $f(3) = 9 - 18 + 5 = -4$.

Machine learning link 4: sensitivity, $\Delta y \approx f'(x)\,\Delta x$ core

A lamp is wired to a dial. Nudge the dial a tiny bit and watch the brightness. If a small nudge changes the brightness a lot, the lamp is very sensitive to the dial. If hardly anything happens, it is insensitive. The ratio (brightness change) ÷ (dial change) is the derivative.

So the derivative lets us predict the effect of a small nudge without recomputing anything: new output $\approx$ old output + slope × nudge. It works because the curve looks straight up close.

$f(x) = x^2$ at $x = 3$, where $f'(3) = 6$.

  1. Nudge by $\Delta x = 0.1$. Prediction: $\Delta y \approx 6\times0.1 = 0.6$.
  2. Truth: $f(3.1) - f(3) = 9.61 - 9 = 0.61$. Error: $0.01$.
  3. Nudge by $\Delta x = 0.01$. Prediction $0.06$. Truth: $9.0601 - 9 = 0.0601$. Error: $0.0001$.

A nudge ten times smaller gives an error a hundred times smaller (the error is about $(\Delta x)^2$). ML version: loss $L(w) = (w-3)^2$ at $w = 5$ has $L'(5) = 4$. Raising $w$ by $0.1$ should raise the loss by about $0.4$; the truth is $0.41$.

For a small change $\Delta x$ in the input:

$$\Delta y = f(x+\Delta x) - f(x) \;\approx\; f'(x)\,\Delta x.$$

$|f'(x)|$ is the sensitivity of the output to the input at $x$. Large $|f'|$: sensitive. $f'=0$: insensitive to first order. The error of this estimate shrinks like $(\Delta x)^2$ (more on this with Taylor series and linearization). The same formula covers error propagation: an uncertainty $\pm\Delta x$ in the input becomes about $\pm|f'(x)|\Delta x$ in the output.

Why do we need it?

We constantly ask "what if this number were a bit different?" The derivative answers it instantly, with one multiplication instead of re-running the model.

Where is it used?

Finding which weights or features matter most (large $|\partial L/\partial w|$), gradient descent (it moves a weight in proportion to its sensitivity), saliency maps, adversarial examples (tiny nudges that flip a prediction), and uncertainty estimates.

How is it used?

Compute $f'(x)$ once. Then $\Delta y\approx f'(x)\Delta x$ for any small $\Delta x$. Compare $|f'|$ across inputs to rank them by sensitivity.

The picture always zooms in on the nudge. The green bar is the true change $\Delta y$. The orange bar is the prediction $f'(x)\Delta x$ (up the tangent). The red piece is the error. Shrink $\Delta x$ and watch the error shrink much faster than the nudge. Try sigmoid at $x=5$: tiny sensitivity. Try $\sin x$ at $x=0$ (slope 1, very sensitive) and then at $x=1.5$ (near the top of the wave, slope about 0.07, almost insensitive).

Only for small nudges. The prediction is the tangent line, and the tangent leaves the curve as you move away. For a big nudge the error is big too. Zoom out in the widget (use $\Delta x=1$) to see this.

Zero slope does not mean "no effect". At a minimum $f'=0$, but a big enough nudge still changes the output, as the curvature (the second derivative) kicks in.

Quick check: $f'(2)=5$. Estimate $f(2.02)-f(2)$.

$\Delta y\approx f'(2)\cdot\Delta x = 5\times0.02 = 0.1$.

Machine learning link 5: learning curves and their slope

While a model trains, we write down its loss after every step. Plot loss against step number and you get a learning curve. It is a function of time, so it has a slope too: the slope of the learning curve is how fast the model is improving, right now.

  • Steeply falling: fast progress.
  • Gently falling: slow progress, getting close to the bottom.
  • Slope near zero: a plateau. The model has converged, or is stuck.
  • Slope positive: the loss is rising. Something is wrong (learning rate too high), or it is the validation loss turning up (overfitting).

Take the gradient-descent run from before ($L=(w-3)^2$, $\eta=0.1$, start at $0$). The loss after each step is $9,\ 5.76,\ 3.6864,\ 2.3593,\ 1.5099,\dots$ (each is $0.64\times$ the last).

  1. Slope from step 0 to 1: $5.76 - 9 = -3.24$ per step.
  2. Slope from step 1 to 2: $3.6864 - 5.76 = -2.0736$.
  3. Slope from step 2 to 3: $2.3593 - 3.6864 = -1.3271$.

The slope is negative (improving) but its size shrinks: fast progress early, slow progress later. A learning curve that flattens out is the visual signal for convergence.

A learning curve plots the loss $L(t)$ against the training step $t$. Its slope is the rate of improvement:

$$\frac{dL}{dt}\approx\frac{L(t+k)-L(t-k)}{2k}\quad(\text{average over nearby steps}).$$

The training loss is measured on data the model learns from; the validation loss on held-out data. Early stopping stops training at the step where the validation curve stops falling, which is where its slope reaches $0$ (the bottom of a valley in $t$!).

Why do we need it?

You cannot see inside a training run. The learning curve and its slope are your instruments: they say whether training is working, finished, stuck or broken.

Where is it used?

Every training run (TensorBoard and Weights & Biases plot them), early stopping, learning-rate schedules and warm-up, and debugging exploding or stalled training.

How is it used?

Watch the curve. Falling quickly: carry on. Flat: stop, or change the learning rate. Rising or jagged: lower the learning rate. Validation turning up while training keeps falling: overfitting, stop early.

These curves are synthetic: made-up shapes with a little random noise (generated from a fixed random seed, so you always see the same curves), not the result of a real training run. Pick a scenario. Drag the blue dot along the training curve: the orange line is its local slope, and the readout translates the slope into a verdict. Look at overfitting: the training curve keeps falling, but the validation curve (green) has a slope of about 0 at the purple line and then rises. That is the moment to stop.

Noise hides the slope. Real curves are jagged, so one step's difference is unreliable. Look at the trend over many steps (smooth the curve, or average over a window as the widget does).

A flat curve has two very different meanings: "finished" or "stuck". Only the loss value tells which: low and flat is converged; high and flat is stuck, so change something.

Quick check: the loss goes 4.0, 3.0, 2.4, 2.1, 2.0 over four steps. Is the slope getting steeper or flatter?

Slopes per step: $-1.0,\ -0.6,\ -0.3,\ -0.1$. They are all negative but shrinking: flatter, so training is slowing down towards a plateau.

Recap, cheat sheet and practice

  • The derivative $f'(x)=\lim_{h\to0}\frac{f(x+h)-f(x)}{h}$ is the slope of the tangent line, the instantaneous rate of change, and the sensitivity of the output to a nudge: $\Delta y\approx f'(x)\Delta x$.
  • First principles: secant slope, simplify until $h$ cancels, let $h\to0$. This gave $x^2\to2x$, $x^3\to3x^2$, $1/x\to-1/x^2$, $\sqrt x\to1/(2\sqrt x)$.
  • Rules: power $nx^{n-1}$; product $u'v+uv'$; quotient $(u'v-uv')/v^2$; chain $f'(g(x))\,g'(x)$ ("rates multiply along the chain").
  • Special functions: $(e^x)'=e^x$, $(a^x)'=(\ln a)a^x$, $(\ln x)'=1/x$, $(\sin)'=\cos$, $(\cos)'=-\sin$, $(\tan)'=1/\cos^2$.
  • ML: $\sigma'=\sigma(1-\sigma)\le0.25$ (vanishing gradients); ReLU has slope 0 or 1 with a kink at 0; a loss is a function $L(w)$ of a parameter; gradient descent $w\leftarrow w-\eta L'(w)$; critical points have $f'=0$ (classify by sign change or $f''$); a learning curve's slope says how fast you are improving.
  • $f''$ is the slope of the slope (the bend): $f''>0$ bends up (valley), $f''\lt 0$ bends down (hill).

Cheat sheet

IdeaFormulaPicture / meaning
Derivative$\lim_{h\to0}\frac{f(x+h)-f(x)}{h}$slope of the tangent
Tangent line at $a$$y=f(a)+f'(a)(x-a)$best straight-line copy
Nudge estimate$\Delta y\approx f'(x)\Delta x$sensitivity
Sum / multiple$(cf+g)'=cf'+g'$differentiate piece by piece
Power$(x^n)'=nx^{n-1}$bring down, lower by one
Product$(uv)'=u'v+uv'$two strips of area
Quotient$(u/v)'=\frac{u'v-uv'}{v^2}$product rule in disguise
Chain$\frac{dy}{dx}=\frac{dy}{du}\frac{du}{dx}$gears: rates multiply
Exponential$(e^{g})'=g'e^{g}$, $(a^x)'=\ln a\cdot a^x$slope proportional to height
Logarithm$(\ln g)'=g'/g$, $(\log_a x)'=\frac1{x\ln a}$mirror of $e^x$: reciprocal slope
Trig$(\sin)'=\cos,\ (\cos)'=-\sin,\ (\tan)'=1/\cos^2$radians only
Sigmoid$\sigma'=\sigma(1-\sigma)$, max $0.25$flat tails: vanishing gradient
ReLU$0$ for $x\lt 0$, $1$ for $x>0$kink at $0$
Higher order$f''=(f')'$bend: $>0$ smile, $\lt 0$ frown
Gradient descent$w\leftarrow w-\eta L'(w)$step against the slope
Critical point$f'(x)=0$hill, valley or shelf
Code it · NumPy

import numpy as np

def num_deriv(f, x, h=1e-5):
    """Central-difference 'nudge test' for the derivative of f at x."""
    return (f(x + h) - f(x - h)) / (2 * h)

# 1. Check hand-derived formulas against the nudge test
print(round(num_deriv(lambda x: x**2, 3.0), 4))                       # 6.0      (2x at x=3)
f  = lambda x: x**2 * (3*x + 1)                                       # product rule example
print(round(num_deriv(f, 2.0), 4), 9*2.0**2 + 2*2.0)                  # 40.0 40.0
g  = lambda x: np.sin(x**2)                                           # chain rule example
print(round(num_deriv(g, 1.0), 4), round(2*1.0*np.cos(1.0**2), 4))    # 1.0806 1.0806

# 2. Sigmoid and its derivative sigma(1 - sigma)
sigmoid = lambda z: 1 / (1 + np.exp(-z))
z = np.array([-5.0, 0.0, 5.0])
s = sigmoid(z)
print(np.round(s * (1 - s), 5))                                       # [0.00665 0.25    0.00665]
print(0.25 ** 10)                                                     # 9.5367431640625e-07  (10 sigmoid layers, best case)

# 3. A loss as a function of one weight (three data points)
X = np.array([1.0, 2.0, 3.0]); Y = np.array([2.0, 4.0, 5.0])
L  = lambda w: np.mean((w * X - Y) ** 2)
dL = lambda w: 2 * np.mean(X * (w * X - Y))                           # chain rule by hand
print(round(L(0.0), 4), round(dL(0.0), 4), round(num_deriv(L, 0.0), 4))   # 15.0 -16.6667 -16.6667
print(np.sum(X * Y) / np.sum(X * X), 25 / 14)                         # 1.7857142857142858 1.7857142857142858  (where L' = 0)

# 4. Gradient descent on L(w) = (w - 3)^2 with three learning rates
for eta in (0.02, 0.1, 1.1):
    w = 0.0
    for k in range(25):
        w = w - eta * 2 * (w - 3)          # w_new = w_old - eta * L'(w_old)
    print(eta, round(w, 4))                # 0.02 -> 1.9188 (too slow), 0.1 -> 2.9887 (close to 3), 1.1 -> 289.1886 (it diverges)
Test yourself

1. What is the derivative of $f(x)=x^3$ at $x=2$?

Power rule: $f'(x)=3x^2$, so $f'(2)=3\cdot4=12$. (8 is $f(2)$, not the slope.)

2. Which is $\dfrac{d}{dx}\big[x^2e^x\big]$?

Product rule with $u=x^2$, $v=e^x$: $u'v+uv' = 2x\,e^x + x^2e^x = (2x+x^2)e^x$. The derivative of a product is not the product of the derivatives.

3. What is $\dfrac{d}{dx}\cos(3x)$?

Chain rule: the outer function $\cos u$ has derivative $-\sin u$, and the inner $u=3x$ has derivative $3$. Multiply: $-3\sin(3x)$.

4. What is $\dfrac{d}{dx}\ln(x^2)$ for $x>0$?

Chain rule: $g'/g = 2x/x^2 = 2/x$. Or note $\ln(x^2)=2\ln x$, whose derivative is $2\cdot\frac1x$.

5. One step of gradient descent on $L(w)=w^2$ from $w=5$ with learning rate $\eta=0.3$ gives a new $w$ of…

$L'(5)=2\cdot5=10$, so $w_{\text{new}} = 5-0.3\cdot10 = 2$. (8 would come from moving with the slope instead of against it.)

6. Why can a deep stack of sigmoid layers have vanishing gradients?

By the chain rule, rates multiply along the chain. $\sigma'(z)=\sigma(1-\sigma)$ is at most $0.25$, so ten layers pass at most $0.25^{10}\approx10^{-6}$ of the signal.

Practice problems

A. Use first principles to differentiate $f(x)=x^2+3x$.

$\dfrac{(x+h)^2+3(x+h) - x^2 - 3x}{h} = \dfrac{2xh + h^2 + 3h}{h} = 2x + h + 3$. Let $h\to0$: $f'(x)=2x+3$. (At $x=1$ the slope is $5$.)

B. Differentiate $f(x)=(2x^3-x)(x^2+5)$ and evaluate at $x=1$.

$u=2x^3-x$, $v=x^2+5$, $u'=6x^2-1$, $v'=2x$. $f'=(6x^2-1)(x^2+5)+(2x^3-x)(2x) = 6x^4+29x^2-5 + 4x^4-2x^2 = 10x^4+27x^2-5$. At $x=1$: $10+27-5=32$. (Check: $f=2x^5+9x^3-5x$, so $f'=10x^4+27x^2-5$ ✓.)

C. Use the quotient rule on $f(x)=\dfrac{e^x}{1+e^x}$ and compare with the sigmoid.

$u=e^x$, $v=1+e^x$, $u'=v'=e^x$. $f' = \dfrac{e^x(1+e^x)-e^x\cdot e^x}{(1+e^x)^2} = \dfrac{e^x}{(1+e^x)^2}$. This is the sigmoid ($\frac{e^x}{1+e^x}=\frac1{1+e^{-x}}$) and its derivative $\sigma(1-\sigma)$ again: $\sigma\cdot(1-\sigma)=\frac{e^x}{1+e^x}\cdot\frac{1}{1+e^x}$ ✓.

D. Differentiate the "softplus" function $f(x)=\ln(1+e^x)$. What famous function do you get?

Chain rule: inner $u=1+e^x$ (derivative $e^x$), outer $\ln u$ (derivative $1/u$). $f'(x)=\dfrac{e^x}{1+e^x}=\sigma(x)$, the sigmoid. So softplus is a smooth version of ReLU whose slope is the sigmoid.

E. Find and classify the critical points of $f(x)=x^3-6x^2+9x+1$.

$f'=3x^2-12x+9=3(x-1)(x-3)$, zero at $x=1$ and $x=3$. $f''=6x-12$: $f''(1)=-6\lt 0$ so $x=1$ is a local maximum, $f(1)=5$. $f''(3)=6>0$ so $x=3$ is a local minimum, $f(3)=1$.

F. (i) Do two gradient-descent steps on $L(w)=(w-2)^2$ from $w_0=6$ with $\eta=0.25$. (ii) For $L(w)=w^3-w$ at $w=2$, estimate the change in loss if $w$ rises by $0.05$.

(i) $L'=2(w-2)$. $L'(6)=8$, so $w_1=6-0.25\cdot8=4$. $L'(4)=4$, so $w_2=4-0.25\cdot4=3$. The distance to $2$ halves each step ($4\to2\to1$), because $1-2\eta=0.5$.

(ii) $L'(w)=3w^2-1$, so $L'(2)=11$. Estimate: $\Delta L\approx11\times0.05=0.55$. Truth: $2.05^3-2.05-6=0.5651$. The error is about $0.015$, from the curvature.

Chapter 2.4

Partial Differentiation & Gradients

Real models have many inputs, not one. This chapter teaches you to find the slope of a landscape: how steep it is east, how steep it is north, and which way is straight uphill. That one arrow, the gradient, is the compass of machine learning.

  • See a function of two (or more) inputs as a surface or a coloured height map
  • Find a partial derivative: the slope when only one input moves and the others stand still
  • Pack all the partial derivatives into the gradient, and say what it means
  • Compute the slope in any direction (the directional derivative) with a dot product
  • Know why the gradient points to steepest ascent and its negative to steepest descent
  • Read level sets and contour maps, and see why the gradient is perpendicular to them

Functions of several inputs: surfaces and height maps core

So far a function took one number in and gave one number out. Real problems have many inputs. A house price depends on its area and its age and its distance from the city. A model's error depends on all of its weights.

The easiest way to picture two inputs is a landscape. Stand on a map. Your position is a pair of numbers: how far east ($x$) and how far north ($y$). At every position the ground has a height. The rule "position in, height out" is a function of two inputs, $f(x, y)$.

Draw the height above every spot and you get a surface. Or look straight down from a helicopter and colour each spot by its height: that is a height map. Both pictures show the same function.

In machine learning the "position" is the set of weights and the "height" is the loss (the error). Training means walking downhill on that landscape.

Let $f(x, y) = x^2 + y^2$. To evaluate it, put in a position and compute the height.

  1. At $(0, 0)$: $f = 0^2 + 0^2 = 0$. (the lowest point)
  2. At $(1, 2)$: $f = 1^2 + 2^2 = 1 + 4 = 5$.
  3. At $(-2, 0)$: $f = (-2)^2 + 0^2 = 4$.
  4. At $(3, -1)$: $f = 9 + 1 = 10$.

Every point at the same distance from the centre has the same height, so this surface is a round bowl. The function $x^2 - y^2$ gives a different shape, a saddle (up in one direction, down in the other), and $xy$ gives another saddle turned by 45°.

A function of several variables takes a vector of inputs and returns one number:

$$f:\ \mathbb{R}^n \to \mathbb{R}, \qquad \mathbf{x} = \begin{bmatrix} x_1 \\ \vdots \\ x_n \end{bmatrix} \ \longmapsto\ f(\mathbf{x}).$$

For two inputs we write $f(x, y)$. Its graph is the set of points $(x, y, f(x, y))$ in 3D: a surface whose height above the floor point $(x, y)$ is the output. (A vector is just a list of numbers, so "many inputs" is "one vector input".)

With $n$ inputs the graph needs $n+1$ dimensions. We cannot draw that, but every idea in this chapter works for any $n$. We draw $n = 2$ and trust it for $n = 1\,000\,000$.

Why do we need it?

One input is rarely enough. To describe anything that depends on several things at once, such as a price, a temperature or an error, we need a rule that accepts several numbers and returns one.

Where is it used?

Every loss function: it takes all the weights of a model (thousands to billions of numbers) and returns one number, the error. Also cost surfaces in optimisation, potential energy in physics, and heat maps of any score.

How is it used?

Think "position in, height out". To see a function, plot its surface or its height map. To train a model, you move the position (the weights) so that the height (the loss) goes down.

Rotate the scene by dragging the background. Drag the purple dot to move the position $(x, y)$ on the floor: the dot rides on the surface and the dashed line to the floor shows where you stand. Use the function menu to switch between a bowl, a saddle, rolling hills and a loss with two valleys. Press Top to look straight down.

This is the surface seen from straight above. The colour is the height: for the bowl and the valleys, darker blue means higher. For the saddle, hills and $xy$, blue means above zero and orange means below zero. Drag the dot. Turn on contour lines to draw lines that join points of equal height.

The height is the output, not a third input. A function of two inputs $f(x, y)$ has the input pair $(x, y)$ on the floor, and the output $z = f(x, y)$ is the height. Do not confuse the vertical axis with an input.

Pictures need care. A surface hides what is behind it. That is why people also use height maps and contour lines (you will meet them below).

Quick check: for $f(x, y) = x^2 - y^2$, what are $f(2, 1)$ and $f(1, 2)$?

$f(2, 1) = 4 - 1 = 3$ and $f(1, 2) = 1 - 4 = -3$. The order of the inputs matters: $x$ goes first, $y$ second. That is exactly why this surface is a saddle: it is up along the $x$-direction and down along the $y$-direction.

Partial derivatives core

You stand on a hillside. In Chapter 2.3 you learned that the derivative is a slope. On a surface there is no single slope: it depends on which way you face.

  • Walk due east (only $x$ changes, $y$ stays fixed). You feel one slope. That is $\partial f/\partial x$.
  • Walk due north (only $y$ changes, $x$ stays fixed). You feel a different slope. That is $\partial f/\partial y$.

Picture it as slicing. Cut the surface with a vertical wall along the line $y = $ constant. The edge of the cut is an ordinary curve (a function of $x$ alone). Its slope is the partial derivative $\partial f/\partial x$.

So the recipe is: freeze every input except one, then take an ordinary derivative. The frozen ones are treated like plain constants.

Let $f(x, y) = x^2 y + 2y$. Find both partial derivatives, then evaluate them at $(3, 2)$.

$\partial f/\partial x$: freeze $y$ (treat it as a number, like 5).

  1. $x^2 y$ is "$x^2$ times a constant $y$". Its derivative in $x$ is $2x\cdot y = 2xy$.
  2. $2y$ has no $x$ in it, so it is a constant: its derivative is $0$.
  3. So $\dfrac{\partial f}{\partial x} = 2xy$. At $(3, 2)$: $2\cdot3\cdot2 = 12$.

$\partial f/\partial y$: freeze $x$.

  1. $x^2 y$ is "a constant $x^2$ times $y$". Its derivative in $y$ is $x^2$.
  2. $2y$ has derivative $2$.
  3. So $\dfrac{\partial f}{\partial y} = x^2 + 2$. At $(3, 2)$: $9 + 2 = 11$.

Check by nudging. $f(3, 2) = 9\cdot2 + 4 = 22$. Nudge only $x$ by $0.001$: $f(3.001, 2) = 9.006001\cdot2 + 4 = 22.012002$. The change is $0.012002$, and $0.012002/0.001 = 12.002 \approx 12$ ✓. Nudge only $y$ by $0.001$: $f(3, 2.001) = 18.009 + 4.002 = 22.011$, a change of $0.011$, so the slope is $11$ ✓.

What this means. At $(3, 2)$, a small step in $x$ changes $f$ about 12 times as much as the step; a small step in $y$ changes it about 11 times as much. If we nudge $x$ by $0.01$ and $y$ by $0.02$, the two effects simply add: $12\cdot0.01 + 11\cdot0.02 = 0.12 + 0.22 = 0.34$. The true change is $f(3.01, 2.02) - f(3, 2) = 22.341402 - 22 = 0.341402$ (very close to $0.34$) ✓.

The partial derivative of $f$ with respect to $x$ at the point $(x, y)$ is

$$\frac{\partial f}{\partial x}(x, y) = \lim_{h \to 0} \frac{f(x + h,\, y) - f(x,\, y)}{h},$$

and with respect to $y$:

$$\frac{\partial f}{\partial y}(x, y) = \lim_{h \to 0} \frac{f(x,\, y + h) - f(x,\, y)}{h}.$$

It is the ordinary derivative with the other inputs held fixed. Notation: $\dfrac{\partial f}{\partial x}$, $f_x$ and $\partial_x f$ all mean the same. The curly $\partial$ ("partial d") warns you that other inputs exist; the straight $d$ is for one-input functions.

For $n$ inputs there are $n$ partial derivatives $\dfrac{\partial f}{\partial x_1}, \dots, \dfrac{\partial f}{\partial x_n}$: one slope for each input, each found by freezing all the others. All the usual rules (power, product, chain, $e^x$, $\sin$ …) still work for the one variable that moves.

Why do we need it?

To fix a model we must know what each single knob does. "If I turn only this weight up a little, does the error rise or fall, and how fast?" Partial derivatives answer that for each knob separately.

Where is it used?

Every training step of linear regression, logistic regression and neural networks needs the partial derivative of the loss with respect to each weight. Also sensitivity analysis (which input matters most) and thermodynamics.

How is it used?

Pick one input. Treat all the others as constants. Differentiate with the normal rules. Repeat for each input. Check one of them by nudging the input by 0.001 and dividing the change in the output by 0.001.

Pick move only x: a vertical wall (teal) cuts the surface along $y = $ constant, and the orange curve is the cut edge. The pink line touches the cut at the dot; its slope is $\partial f/\partial x$. Drag the purple dot and watch the slope. Switch to move only y for $\partial f/\partial y$. Rotate and press Side to see the cut as an ordinary 2D curve.

The function is $f(x, y) = x^2 y + 2y$ at the point $(x_0, y_0)$ you choose. Move $\Delta x$ and $\Delta y$. The bars show the change in $f$ from nudging only $x$, only $y$, and both. For small nudges, "both" is just the sum of the two single effects, and it matches $f_x\,\Delta x + f_y\,\Delta y$. Make the nudges large and notice the match start to slip.

"Frozen" does not mean "zero". When you differentiate in $x$, $y$ is not deleted. It stays in the answer as a constant factor (as in $2xy$).

A partial derivative is a slope in one special direction only. Knowing $f_x$ and $f_y$ does not yet tell you the slope when you walk north-east. (The directional derivative below fixes that.)

You cannot cancel the $\partial$'s like fractions in a casual way. $\dfrac{\partial f}{\partial x}$ is one symbol meaning "slope in the $x$ direction".

Quick check: for $f(x, y) = x^2 y + 2y$, what is $\partial f/\partial y$ at $(1, 5)$?

$\partial f/\partial y = x^2 + 2$. It does not depend on $y$ at all. At $x = 1$ it is $1 + 2 = 3$ (the $y = 5$ is not used).

The gradient core

A hillside gives you two useful numbers at your feet: how steep it is going east, and how steep it is going north. Keeping two numbers separate is clumsy. Better: stack them into one arrow and draw it on the map at your position.

That arrow is the gradient. Think of a weather map with a little wind arrow at every town. The gradient is a "slope arrow" at every position: it points up the hill, and a longer arrow means a steeper hill.

One arrow is also easier to use. To go downhill you simply walk the opposite way.

Example 1. $f(x, y) = x^2 + y^2$. The two partial derivatives are $f_x = 2x$ and $f_y = 2y$, so

$$\nabla f(x, y) = \begin{bmatrix} 2x \\ 2y \end{bmatrix}.$$
  1. At $(1, 2)$: $\nabla f = [2, 4]^\top$. It points up and to the right, away from the centre of the bowl.
  2. At $(-3, 0)$: $\nabla f = [-6, 0]^\top$. It points left, again away from the centre, and it is longer (the bowl is steeper far out).
  3. At $(0, 0)$: $\nabla f = [0, 0]^\top$. The bottom of the bowl is flat.

Example 2. From the last section, $f = x^2y + 2y$ has $f_x = 2xy$ and $f_y = x^2 + 2$. At $(3, 2)$ the gradient is $[12, 11]^\top$.

Example 3 (three inputs). $f(x, y, z) = xyz$. Then $f_x = yz$, $f_y = xz$, $f_z = xy$. At $(1, 2, 3)$: $\nabla f = [6, 3, 2]^\top$. Nothing changes except that the list is longer.

The gradient of $f:\mathbb{R}^n\to\mathbb{R}$ is the column vector of all its partial derivatives:

$$\nabla f(\mathbf{x}) = \begin{bmatrix} \dfrac{\partial f}{\partial x_1} \\ \vdots \\ \dfrac{\partial f}{\partial x_n} \end{bmatrix}.$$

The symbol $\nabla$ is read "nabla" or "del". Two things to remember:

  • $\nabla f(\mathbf{x})$ has the same shape as the input $\mathbf{x}$ ($n$ numbers). That is why an update like $\mathbf{x} - \eta\,\nabla f(\mathbf{x})$ makes sense: you can subtract two vectors of the same size.
  • The gradient is a function: it gives one vector for every point. (A vector for every point is called a vector field; see Chapter 2.13.) In this guide the gradient is always a column. Its transpose, a row, is the Jacobian of a scalar function (Chapter 2.5).
Why do we need it?

A model can have a million knobs. We need one object that says, for all of them at once, which way to turn each one to raise the error fastest. The gradient is that object.

Where is it used?

Gradient descent, stochastic gradient descent, Adam and every other optimiser; backpropagation computes the gradient of the loss with respect to all the weights. Also edge detection in images (the gradient of brightness).

How is it used?

Compute every partial derivative, stack them into a column, and evaluate at your current position. Then step against it to lower the loss: new position = old position minus (learning rate) times gradient.

Drag the purple point. The orange pieces are the two partial derivatives: how far to go in $x$ and how far in $y$ (the east-slope and the north-slope). Together they build the green gradient arrow, which points uphill. Notice that the arrow gets long where the colours change fast. Turn on arrows everywhere to see the gradient as a field.

The gradient lives on the floor, not on the surface. It is a vector in the input space (it has one entry per input). It tells you which way to walk, not which way the surface tilts in 3D. (The 3D "tilt" arrow is a different object, drawn in the next section.)

Gradient versus derivative. For one input, $f'(x)$ is a single number. For $n$ inputs, $\nabla f$ is a list of $n$ numbers. When $n = 1$ the two agree.

Quick check: find $\nabla f$ for $f(x, y) = 3x + 5y$ and for $f(x, y) = xy$.

For $3x + 5y$: $f_x = 3$, $f_y = 5$, so $\nabla f = [3, 5]^\top$ at every point (a flat slope that never changes: a tilted plane). For $xy$: $f_x = y$ and $f_y = x$, so $\nabla f = [y, x]^\top$. Notice the swap.

What the gradient tells you (gradient interpretation) core

The gradient at a point tells you three things. Read them like a compass:

  1. Direction: the arrow points straight uphill, the way the ground rises fastest.
  2. Length: a long arrow means a steep hill, a short arrow a gentle one.
  3. Zero: if the arrow has length zero, the ground is flat here. You are at the top of a hill, the bottom of a valley, or the middle of a saddle.

Zoom in far enough on any smooth surface and it looks like a tilted flat sheet (the tangent plane). The gradient is simply the direction in which that sheet tilts upward.

The bowl $f = x^2 + y^2$ has $\nabla f = [2x, 2y]^\top$: two times the position vector. So the gradient always points away from the centre (uphill), and gets longer the farther out you go (steeper). At the centre it is zero: the bottom.

The saddle $f = x^2 - y^2$ has $\nabla f = [2x, -2y]^\top$. At $(1, 1)$: $[2, -2]^\top$, which points right and down: uphill means "more $x$, less $y$". At $(0, 0)$ it is zero, yet the origin is not a minimum or maximum.

Using the gradient to predict. For the bowl at $(1, 2)$: $f = 5$ and $\nabla f = [2, 4]^\top$. Take a small step $\mathbf{h} = [0.1, 0.1]^\top$.

  1. Prediction: $f \approx 5 + \nabla f\cdot\mathbf{h} = 5 + 2\cdot0.1 + 4\cdot0.1 = 5.6$.
  2. Truth: $f(1.1, 2.1) = 1.21 + 4.41 = 5.62$.
  3. The error is $0.02$, tiny compared with the step. The smaller the step, the better the prediction.

If $f$ is smooth near $\mathbf{a}$, then for a small step $\mathbf{h}$

$$f(\mathbf{a} + \mathbf{h}) \approx f(\mathbf{a}) + \nabla f(\mathbf{a})\cdot\mathbf{h}.$$

The right side is the tangent plane (or linear approximation) at $\mathbf{a}$: $z = f(\mathbf{a}) + \nabla f(\mathbf{a})\cdot(\mathbf{x} - \mathbf{a})$. The gradient is what tilts it. This gives the three facts:

  • $\nabla f(\mathbf{a})$ points in the direction of fastest increase.
  • $\|\nabla f(\mathbf{a})\|$ is the steepness (the largest slope in any direction).
  • $\nabla f(\mathbf{a}) = \mathbf{0}$ marks a critical point (also stationary point): a minimum, a maximum or a saddle. Telling them apart needs second derivatives (Chapter 2.10).

(The proof that it points to fastest increase comes two sections later, with the directional derivative.)

Why do we need it?

To decide what to do next on a landscape you cannot see all at once. Local information (the tilt under your feet) is enough to choose a good direction.

Where is it used?

Optimisers use the direction to move and the length to judge progress; training stops when the gradient is nearly zero. Linearisation (Chapter 2.12) and Newton's method build on the tangent plane.

How is it used?

Compute $\nabla f$ at your point. Walk along it to go up, against it to go down. If its length is near zero you have reached a flat spot, so check whether it is a minimum, a maximum or a saddle.

Drag the purple dot over the surface. The teal sheet is the tangent plane: it just touches the surface and tilts exactly like it. The green arrow on the floor is the gradient $\nabla f$ (which way to walk). The orange arrow points up the steepest slope of the surface. Press Jump to a flat spot: the sheet lies flat and the arrow vanishes. Rotate to look from the side.

Local only. The gradient describes the ground right under your feet. Far away the surface may bend, so the straight-uphill direction can change (the prediction above only works for small steps).

Zero gradient does not mean minimum. Peaks, valleys and saddles all have a zero gradient. In a loss landscape we hope for a valley, but we can get stuck at a saddle.

Quick check: $f(x, y) = x^2 + 3y^2$ at $(2, -1)$. Which way is uphill, and how steep is it?

$\nabla f = [2x, 6y]^\top = [4, -6]^\top$. Uphill is right and down (more $x$, less $y$). Steepness $= \sqrt{16 + 36} = \sqrt{52} \approx 7.21$.

Directional derivatives core

The east-slope and the north-slope are just two special directions. What if you walk north-east, or in any direction you like? The slope you feel in a chosen direction is the directional derivative.

Here is the key idea: the gradient already contains the answer for every direction. To get the slope along a direction, take the dot product of the gradient with that direction. (A dot product multiplies matching entries and adds them up. It is big when two arrows point the same way.)

Why a dot product? If the direction mostly agrees with the uphill arrow, you climb fast. If it is at a right angle to it, you walk along a level path and climb nothing. If it points the opposite way, you go downhill.

Take $f = x^2y + 2y$ at $(3, 2)$, where $\nabla f = [12, 11]^\top$. Walk in the direction $\mathbf{u} = [0.6, 0.8]^\top$ (a unit vector: $0.36 + 0.64 = 1$).

  1. Dot product: $D_{\mathbf{u}} f = 12\cdot0.6 + 11\cdot0.8 = 7.2 + 8.8 = 16$.
  2. So walking 1 unit that way raises $f$ by about 16 (for a tiny step, 16 times the step).

Where does the dot product come from? Walk along the line $(3 + 0.6t,\ 2 + 0.8t)$. The height is $g(t) = (3 + 0.6t)^2(2 + 0.8t) + 2(2 + 0.8t)$, an ordinary function of $t$. Its slope at $t = 0$ is, by the product rule,

$$g'(0) = \underbrace{2\cdot3\cdot0.6}_{3.6}\cdot 2 + 9\cdot0.8 + 2\cdot0.8 = 7.2 + 7.2 + 1.6 = 16 \ ✓.$$

Compare: $12\cdot0.6 = 7.2$ is the $x$-part and $11\cdot0.8 = 8.8$ the $y$-part. The same $16$.

Special cases. $\mathbf{u} = [1, 0]^\top$ gives $12\cdot1 + 11\cdot0 = 12 = f_x$. $\mathbf{u} = [0, 1]^\top$ gives $f_y = 11$. The partial derivatives are directional derivatives along the axes.

A direction that is not a unit vector. To walk toward $[1, 1]$, first make it length 1: $\mathbf{u} = [1, 1]/\sqrt2 \approx [0.707, 0.707]$. Then $D_{\mathbf{u}}f = (12 + 11)/\sqrt2 = 23/1.4142 \approx 16.26$.

The directional derivative of $f$ at $\mathbf{a}$ in the direction of a unit vector $\mathbf{u}$ is the slope along that line:

$$D_{\mathbf{u}} f(\mathbf{a}) = \lim_{h\to0}\frac{f(\mathbf{a} + h\mathbf{u}) - f(\mathbf{a})}{h} = \nabla f(\mathbf{a})\cdot\mathbf{u}.$$

Why the formula holds. Put $g(t) = f(\mathbf{a} + t\mathbf{u}) = f(a_1 + tu_1,\ a_2 + tu_2)$. Each input changes at its own rate ($u_1$ and $u_2$ per unit of $t$), and each change moves $f$ by its partial derivative times that rate. So $g'(0) = f_x u_1 + f_y u_2 = \nabla f\cdot\mathbf{u}$. (This is the chain rule for several variables; you will study it properly in Chapter 2.8.)

Because $\mathbf{u}$ must have length 1, the answer is "slope per unit of distance walked". Without that rule, a longer vector would give a bigger number just from taking a longer step.

Why do we need it?

Real moves are rarely along an axis. We need the rate of change for any direction of travel, and a way to find the best one.

Where is it used?

Line search in optimisation (how the loss changes along the update direction), sensitivity of a model to a change in a chosen mix of features, and gradient checking (nudge the weights in a random direction and compare).

How is it used?

Normalise your direction to length 1, compute the gradient, and take their dot product. The sign tells you uphill (positive), downhill (negative) or level (zero).

Left: drag the purple point and turn the angle dial. The green arrow is your direction $\mathbf{u}$ and the orange arrow is the gradient. Right: the green dot shows the slope $D_{\mathbf{u}}f$ in the direction $\theta$ (distance from the centre = slope, scaled so that the largest is 1). It always traces a circle, and the dot sits at the far end exactly when $\mathbf{u}$ lines up with the gradient. At right angles the slope is 0.

Use a unit vector. $\mathbf{u}$ must have length exactly 1. If you are given the direction $[3, 4]^\top$, divide by its length $5$ first: $\mathbf{u} = [0.6, 0.8]^\top$.

A directional derivative is a number, not a vector. It is one slope. The gradient is the vector that holds the slopes for all directions at once.

Quick check: $\nabla f = [3, 4]^\top$ at a point. Find the slope in the direction $[0, 1]^\top$ and in the direction $[4, -3]^\top$.

For $[0, 1]^\top$ (already unit): $3\cdot0 + 4\cdot1 = 4$. For $[4, -3]^\top$, its length is $5$, so $\mathbf{u} = [0.8, -0.6]^\top$ and the slope is $3\cdot0.8 - 4\cdot0.6 = 2.4 - 2.4 = 0$. That direction is perpendicular to the gradient: a level path.

Steepest ascent and steepest descent core

You are on a foggy hill and can only feel the ground at your feet. You want to climb as fast as possible. Which way do you step? Straight along the gradient. Every other direction wastes some of your step going sideways.

Want to go down as fast as possible, like a ball rolling to the valley? Step the exact opposite way: against the gradient.

And if you step at a right angle to the gradient, you walk along a level path: you neither climb nor descend.

This one fact is the whole idea behind gradient descent, the method that trains almost every model: look at the gradient, take a small step the other way, repeat.

Take $f = x^2y + 2y$ at $(3, 2)$ with $\nabla f = [12, 11]^\top$.

  1. Steepness: $\|\nabla f\| = \sqrt{12^2 + 11^2} = \sqrt{265} \approx 16.28$.
  2. Steepest ascent direction: divide by the length, $\mathbf{u}_\text{up} = [12, 11]/16.28 \approx [0.737, 0.676]$. Slope there: $+16.28$.
  3. Steepest descent direction: $\mathbf{u}_\text{down} \approx [-0.737, -0.676]$. Slope there: $-16.28$.
  4. A level direction, perpendicular to the gradient: $[-11, 12]/16.28 \approx [-0.676, 0.737]$. Slope: $\bigl(12\cdot(-11) + 11\cdot12\bigr)/16.28 = 0$.
  5. Our earlier direction $[0.6, 0.8]$ gave slope 16, which is $16/16.28 \approx 98\%$ of the best. It is only about $10.6^\circ$ away from the gradient.

One gradient-descent step on the bowl $f = x^2 + y^2$ from $(3, 4)$: $\nabla f = [6, 8]^\top$ and $f = 25$. With step size $\eta = 0.1$, the new point is $[3, 4]^\top - 0.1\,[6, 8]^\top = [2.4, 3.2]^\top$, where $f = 5.76 + 10.24 = 16$. The loss dropped from $25$ to $16$.

Why the gradient is the steepest direction. For a unit vector $\mathbf{u}$ at angle $\theta$ to $\nabla f$, the dot product rule $\mathbf{a}\cdot\mathbf{b} = \|\mathbf{a}\|\|\mathbf{b}\|\cos\theta$ (see the dot product) gives

$$D_{\mathbf{u}}f = \nabla f\cdot\mathbf{u} = \|\nabla f\|\,\underbrace{\|\mathbf{u}\|}_{=1}\cos\theta = \|\nabla f\|\cos\theta.$$

The cosine is at most $1$ (at $\theta = 0$, same direction), at least $-1$ (at $\theta = 180^\circ$) and $0$ at $90^\circ$. Therefore:

  • Steepest ascent: $\mathbf{u} = \dfrac{\nabla f}{\|\nabla f\|}$, with the largest possible slope $+\|\nabla f\|$.
  • Steepest descent: $\mathbf{u} = -\dfrac{\nabla f}{\|\nabla f\|}$, with slope $-\|\nabla f\|$.
  • No change: any $\mathbf{u}\perp\nabla f$ has slope $0$.

This gives the gradient descent update, with a small positive number $\eta$ called the learning rate (step size):

$$\mathbf{x}_{\text{new}} = \mathbf{x} - \eta\,\nabla f(\mathbf{x}).$$

If $\nabla f = \mathbf{0}$ there is no direction to prefer, and the method stops.

Why do we need it?

Models have too many weights to try combinations at random. The gradient gives, for free, the single best direction to lower the error, so training becomes "repeat one small downhill step".

Where is it used?

Training linear and logistic regression, neural networks, matrix factorisation and embeddings. Gradient ascent is the same idea upside down: maximising a reward or a likelihood.

How is it used?

Start at some weights. Compute the gradient of the loss. Subtract learning rate times gradient. Repeat until the gradient is almost zero or the loss stops improving. If the loss gets worse or explodes, shrink the learning rate.

Drag the purple start point. The purple dots show the steps $\mathbf{x} \leftarrow \mathbf{x} - \eta\nabla f$ (they cross the contour lines at right angles). Try the two-valley loss: start on the left and you reach the left valley, start on the right and you reach the right one. Raise $\eta$ until the path zig-zags. With a start far from the valleys (try the corner $(1.9, 1.9)$ on the two-valley loss) a big $\eta$ makes it blow up. Switch to ascent to walk uphill instead.

Here the height is a loss (error) that depends on two weights $x$ and $y$. Drag the purple start dot on the surface. The orange path is gradient descent: each dot is one step $-\eta\nabla f$. Look at it from Top (it crosses the contours at right angles) and from Side (the loss goes down at every step). Start on each side of the middle ridge to land in different valleys.

Steepest locally, not the shortest way to the bottom. The gradient is the best direction for a tiny step. On a long, narrow valley it points mostly toward the valley wall, so the path zig-zags. And a step that is too big overshoots.

You can reach different minima. On a loss with several valleys, where you end up depends on where you start.

Descent needs the minus sign. $+\eta\nabla f$ climbs; $-\eta\nabla f$ descends.

Quick check: at a point $\nabla f = [3, 4]^\top$. What is the steepest descent direction and its slope?

$\|\nabla f\| = 5$. Direction $= -[3, 4]/5 = [-0.6, -0.8]$, slope $= -5$.

Level sets core

On a hiking map, a thin line labelled "100 m" joins every place that is exactly 100 metres above the sea. Walk along that line and you never go up or down. This is a contour line, and mathematicians call it a level set: all the points where the function has one chosen value.

Picture slicing the surface with a perfectly horizontal sheet at height $c$. Where the sheet cuts the surface you get a curve. Drop that curve straight down to the floor and you get the level set.

Do this for several heights and you get a set of nested curves: a map of the whole surface drawn on flat paper.

The bowl $f = x^2 + y^2$. The level set at height $c$ is $x^2 + y^2 = c$:

  • $c = 0$: only the single point $(0, 0)$.
  • $c = 1$: a circle of radius $1$. $c = 4$: radius $2$. $c = 9$: radius $3$. (In general radius $\sqrt{c}$.)
  • $c = -2$: nothing at all, since $x^2 + y^2$ is never negative. The level set is empty.

The saddle $f = x^2 - y^2$: $c = 1$ gives $x^2 - y^2 = 1$, two curved branches opening left and right (a hyperbola). $c = -1$ gives two branches opening up and down. $c = 0$ gives $x^2 = y^2$, i.e. the two crossing lines $y = x$ and $y = -x$.

The product $f = xy$: $c = 2$ gives $y = 2/x$, a hyperbola in the first and third quadrants.

The level set (or level curve when $n=2$) of $f$ at level $c$ is

$$L_c = \{\,\mathbf{x}\in\mathbb{R}^n \;:\; f(\mathbf{x}) = c\,\}.$$

For two inputs, $L_c$ is usually a curve in the floor plane. For three inputs it is a surface (a "level surface", for example a sphere for $f = x^2 + y^2 + z^2$). For $n$ inputs it usually has $n-1$ dimensions. Different levels never cross each other, because one point has only one height. (A level set can be a single point, empty, or have a crossing such as at the saddle level $c = 0$, but two different levels never meet.)

Why do we need it?

A surface is hard to draw and hides what is behind it. Level sets turn a 3D landscape into a flat map you can read, and they are the right tool for a function of three or more inputs, which cannot be drawn as a surface at all.

Where is it used?

Contour plots of loss functions in every optimisation tutorial, weather maps (lines of equal pressure), decision boundaries of classifiers (the level set where the score is 0.5), and constrained optimisation (Lagrange multipliers, regularisation balls).

How is it used?

Choose a few levels, find the curve where $f$ equals each, and draw them. Closely spaced curves mean steep ground; closed loops mean a peak or a valley. Equal-loss curves tell you which weights are "equally good".

Drag the purple dot over the surface. The teal sheet is a flat cut at the dot's height $c = f(x, y)$. The orange curve is the level set $L_c$: where the sheet meets the surface. Its shadow on the floor is the contour line. Turn on show several levels to see the full map. Press Top to look down at the contour map.

A level set lives on the floor (in the input space), not up on the surface. The curve on the surface is its "lifted copy"; the level set itself is the shadow on the floor.

Level sets are not the graph. The graph is the surface (inputs plus output). A level set is only the inputs that give one particular output.

Quick check: describe the level set $x^2 + 4y^2 = 4$ of $f = x^2 + 4y^2$ at $c = 4$.

Divide by 4: $\dfrac{x^2}{4} + y^2 = 1$. It is an ellipse that crosses the $x$-axis at $\pm2$ and the $y$-axis at $\pm1$ (wider than tall, because $y$ has the bigger coefficient, so $y$ gets "expensive" faster).

Contour interpretation: reading the map core

A contour map is a flat picture of a landscape. Four rules let you read it, just like a hiker:

  • Lines close together = steep. The height changes a lot over a short distance.
  • Lines far apart = gentle (nearly flat).
  • Closed loops circle a peak or a valley. Look at the labels to tell which.
  • The uphill direction is always at a right angle to the contour line. The fastest way up (or down) crosses the lines head-on. Walking along a line is level.

The last rule is the key link to this chapter: the gradient is perpendicular to the level sets.

Perpendicular, with numbers. $f = x^2 + y^2$ has level set $x^2 + y^2 = 5$ through the point $(1, 2)$ (a circle of radius $\sqrt5$).

  1. The gradient at $(1, 2)$ is $[2, 4]^\top$.
  2. The circle's tangent direction at $(1, 2)$ is perpendicular to the radius $[1, 2]$, for example $\mathbf{t} = [-2, 1]$.
  3. Dot product: $\nabla f\cdot\mathbf{t} = 2\cdot(-2) + 4\cdot1 = -4 + 4 = 0$ ✓.

Close lines mean steep. Draw contours of $f = x^2 + y^2$ at heights $c = 1, 2, 3, 4$. They are circles of radius $\sqrt c$: $1,\ 1.414,\ 1.732,\ 2$. The gaps between them are $0.414,\ 0.318,\ 0.268$: they get smaller as you go out. The gradient length is $2r$, so it grows outward: the bowl gets steeper. In fact, the gap between contours that differ in height by $\Delta c$ is about $\Delta c/\|\nabla f\|$. At $r \approx 1.2$: $1/(2\cdot1.2) \approx 0.42$ ✓.

Theorem. At any point where $\nabla f \ne \mathbf{0}$, the gradient is perpendicular to the level set through that point.

Why. Walk along the level curve with a path $\mathbf{r}(t)$. The height never changes: $f(\mathbf{r}(t)) = c$ for all $t$. Differentiate both sides with respect to $t$. The right side is a constant, so its derivative is $0$. The left side, by the chain rule for several variables (Chapter 2.8), is $\nabla f\cdot\mathbf{r}'(t)$. So

$$\nabla f(\mathbf{r}(t))\cdot\mathbf{r}'(t) = 0.$$

The velocity $\mathbf{r}'(t)$ is the tangent to the level curve, and its dot product with the gradient is zero: they are perpendicular. $\blacksquare$

Summary of the contour dictionary. (1) $\nabla f\perp$ contours. (2) $\nabla f$ points toward the higher-labelled contour. (3) The spacing of contours is about $\Delta c/\|\nabla f\|$: dense means steep. (4) A closed loop surrounds an extreme point (a peak or a valley); contour lines that cross at a point mark a saddle.

Why do we need it?

The map picture lets you judge a loss landscape at a glance: where it is steep, where it is flat, where the valleys are, and which way an optimiser will move, without computing anything.

Where is it used?

Every picture of gradient descent in a textbook, the explanation of why plain gradient descent zig-zags in narrow valleys (which motivates momentum and Adam), and the geometry of regularised regression and constrained optimisation.

How is it used?

To read a map: find the closest and farthest line spacing (steep versus flat), read the labels to see which side is uphill, then draw the gradient perpendicular to the lines at the point you care about. To plan a step, move across the lines, not along them.

Drag the purple point. The thick orange curve is the contour (level set) through the point. The dashed orange line is its tangent, and the green arrow is the gradient. The small square shows the right angle between them. Move toward places where the thin contours bunch up: the gradient gets longer (steeper). Change the function and check that it stays true everywhere, except at flat spots.

The map shows contour lines, each labelled with its height. Answer the question about points A and B by pressing a button. Then check the numbers in the readout. Press New map for another landscape. Hint: close lines mean steep, and the label tells you the height.

Equal height steps. "Close lines mean steep" is only true when the lines are drawn at equal steps in height (every 100 m, say). Always check the labels.

Perpendicular needs equal axes. If a plot stretches one axis, the gradient will no longer look perpendicular to the contours (it still is, in the true coordinates).

Steep along the arrow is not steep everywhere. A long, thin valley is steep across and nearly flat along. The gradient mostly points across, which is why gradient descent zig-zags there.

Quick check: on a map the 10 m line and the 20 m line are 5 m apart at one place and 50 m apart at another. Where is it steeper, and roughly how steep is it?

Steeper at the first place. The slope is about (height change) / (horizontal gap) $= 10/5 = 2$ there, and $10/50 = 0.2$ at the second place. The gradient is 10 times larger at the first place.

Recap, cheat sheet and practice

  • A function of several inputs $f(x, y)$ is a landscape: position in, height out. Its picture is a surface, or a height map seen from above.
  • A partial derivative $\partial f/\partial x$ is the slope when only $x$ moves: freeze the other inputs, then differentiate as usual. For small nudges, the effects of $x$ and $y$ simply add.
  • The gradient $\nabla f$ is the column of all partial derivatives. It points uphill, its length is the steepness, and it is zero at flat spots (peaks, valleys, saddles).
  • The directional derivative is the slope along a unit direction $\mathbf{u}$: $D_{\mathbf{u}}f = \nabla f\cdot\mathbf{u} = \|\nabla f\|\cos\theta$.
  • So the steepest ascent is along $\nabla f$ (slope $\|\nabla f\|$), the steepest descent is along $-\nabla f$, and directions perpendicular to $\nabla f$ stay level. Gradient descent: $\mathbf{x}\leftarrow\mathbf{x} - \eta\nabla f(\mathbf{x})$.
  • A level set $\{f = c\}$ is a contour line. The gradient is perpendicular to the level sets. Close lines mean steep ground; closed loops circle peaks or valleys.

Cheat sheet

IdeaFormulaPicture
Partial derivative$\dfrac{\partial f}{\partial x} = \lim_{h\to0}\dfrac{f(x+h, y) - f(x, y)}{h}$slope of the cut $y = $ const
Gradient (column)$\nabla f = \bigl[\tfrac{\partial f}{\partial x_1}, \dots, \tfrac{\partial f}{\partial x_n}\bigr]^\top$arrow pointing uphill
Linear approximation$f(\mathbf{a}+\mathbf{h}) \approx f(\mathbf{a}) + \nabla f\cdot\mathbf{h}$tangent plane
Directional derivative$D_{\mathbf{u}}f = \nabla f\cdot\mathbf{u}$ ($\|\mathbf{u}\| = 1$)slope along $\mathbf{u}$
Steepest ascent / descent$\pm\nabla f/\|\nabla f\|$, slope $\pm\|\nabla f\|$straight up / down the hill
No change$\mathbf{u}\perp\nabla f$walk along the contour
Gradient descent step$\mathbf{x}\leftarrow\mathbf{x} - \eta\nabla f(\mathbf{x})$small step downhill
Level set$L_c = \{\mathbf{x} : f(\mathbf{x}) = c\}$contour line
Critical point$\nabla f = \mathbf{0}$flat: min, max or saddle
Code it · NumPy

import numpy as np

def f(p):                      # f(x, y) = x^2 * y + 2*y
    x, y = p
    return x**2 * y + 2*y

def grad(p):                   # analytic gradient: [2xy, x^2 + 2]
    x, y = p
    return np.array([2*x*y, x**2 + 2])

def num_grad(f, p, h=1e-6):    # nudge ONE input at a time (central difference)
    g = np.zeros_like(p, dtype=float)
    for i in range(len(p)):
        e = np.zeros_like(g); e[i] = h
        g[i] = (f(p + e) - f(p - e)) / (2*h)
    return g

p = np.array([3.0, 2.0])
print(f(p))                    # 22.0
print(grad(p))                 # [12. 11.]
print(num_grad(f, p))          # [12. 11.]   (agrees with the formula to about 6 decimals)

u = np.array([0.6, 0.8])       # a unit direction (0.36 + 0.64 = 1)
print(grad(p) @ u)             # 16.0   directional derivative = gradient . u
h = 1e-6
print((f(p + h*u) - f(p)) / h) # about 16.0000036   the same thing, found by nudging

g = grad(p)
print(np.linalg.norm(g))       # 16.2788...   the steepest possible slope
print(g / np.linalg.norm(g))   # [0.7372 0.6757]   steepest-ascent direction

# gradient descent on the bowl x^2 + y^2 (gradient = 2x, 2y), step size 0.1
q = np.array([3.0, 4.0]); eta = 0.1
for k in range(3):
    q = q - eta * 2*q
    print(k + 1, q, q @ q)     # 1 [2.4 3.2] 16.0   then   2 [1.92 2.56] 10.24   then   3 [1.536 2.048] 6.5536

# the gradient is perpendicular to the contour: circle x^2 + y^2 = 5 at (1, 2)
gc = np.array([2*1, 2*2]); t = np.array([-2, 1])
print(gc @ t)                  # 0
Test yourself

1. For $f(x, y) = x^3 y^2$, what is $\partial f/\partial x$ at $(1, 2)$?

Freeze $y$: $\partial f/\partial x = 3x^2y^2 = 3\cdot1\cdot4 = 12$. (The value 4 is $\partial f/\partial y = 2x^3y$ at this point.)

2. What is $\nabla f$ at $(2, 5)$ for $f(x, y) = x^2 + 3y$?

$f_x = 2x = 4$ and $f_y = 3$ (a constant). The gradient is a column of the two partial derivatives, $[4, 3]^\top$. The value $19$ is $f(2, 5)$, which is a height, not a gradient.

3. $\nabla f = [2, -1]^\top$ at a point. What is the slope in the direction $\mathbf{u} = [0.6, 0.8]^\top$?

$D_{\mathbf{u}}f = 2\cdot0.6 + (-1)\cdot0.8 = 1.2 - 0.8 = 0.4$. $\mathbf{u}$ is a unit vector, so no rescaling is needed.

4. If $\nabla f = [0, -2]^\top$, which direction is steepest descent?

Steepest descent is $-\nabla f/\|\nabla f\| = -[0, -2]/2 = [0, 1]$. It must be a unit vector (so not $[0,2]$), and $[1, 0]$ is perpendicular to the gradient, a level direction.

5. On a contour map drawn at equal height steps, where the lines are packed very close together, the ground is…

The gap between lines is about (height step) / $\|\nabla f\|$. A small gap means a large gradient, which means steep ground. Height is read from the labels, not from the spacing.

6. For $f = x^2 + y^2$, one gradient-descent step from $(3, 4)$ with $\eta = 0.5$ lands at…

$\nabla f = [6, 8]^\top$, so the new point is $[3, 4] - 0.5\cdot[6, 8] = [0, 0]$, the bottom of the bowl. ($(2.4, 3.2)$ is the answer for $\eta = 0.1$; $(6, 8)$ would come from adding instead of subtracting.)

Practice problems

A. Find $\nabla f$ for $f(x, y) = x^2y^3$ and evaluate it at $(2, 1)$.

$f_x = 2xy^3$ (freeze $y$) and $f_y = 3x^2y^2$ (freeze $x$). At $(2, 1)$: $f_x = 2\cdot2\cdot1 = 4$ and $f_y = 3\cdot4\cdot1 = 12$. So $\nabla f(2, 1) = [4, 12]^\top$.

B. Find $\nabla f$ for $f(x, y) = e^{xy}$ at $(0, 2)$.

Freeze $y$ and use the chain rule: $f_x = y\,e^{xy}$. Freeze $x$: $f_y = x\,e^{xy}$. At $(0, 2)$: $e^0 = 1$, so $f_x = 2\cdot1 = 2$ and $f_y = 0\cdot1 = 0$. The gradient is $[2, 0]^\top$.

C. Find the directional derivative of $f = x^2 + xy$ at $(1, 2)$ toward the vector $[3, 4]$.

$\nabla f = [2x + y,\ x]^\top = [4, 1]^\top$ at $(1, 2)$. Make the direction a unit vector: $\|[3, 4]\| = 5$, so $\mathbf{u} = [0.6, 0.8]$. Then $D_{\mathbf{u}}f = 4\cdot0.6 + 1\cdot0.8 = 2.4 + 0.8 = 3.2$.

D. For $f = xy$ at $(2, 3)$, find the direction of fastest increase and the largest slope. What is the direction of fastest decrease?

$\nabla f = [y, x]^\top = [3, 2]^\top$. Its length is $\sqrt{9 + 4} = \sqrt{13} \approx 3.606$. Fastest increase: $[3, 2]/\sqrt{13} \approx [0.832, 0.555]$, with slope $\approx 3.606$. Fastest decrease: $[-0.832, -0.555]$, with slope $-3.606$.

E. Show that the gradient of $f(x, y) = x - 2y$ is perpendicular to its level sets.

$\nabla f = [1, -2]^\top$ everywhere. The level set $x - 2y = c$ is a straight line $y = (x - c)/2$ with direction vector $\mathbf{t} = [2, 1]$ (go 2 across, 1 up). Then $\nabla f\cdot\mathbf{t} = 1\cdot2 + (-2)\cdot1 = 0$. So the gradient is perpendicular to every level line. (Contours of a tilted plane are parallel straight lines, evenly spaced: the slope is the same everywhere.)

F. Do two steps of gradient descent on $f = x^2 + 4y^2$ from $(2, 1)$ with $\eta = 0.1$. Give the loss after each.

$\nabla f = [2x, 8y]^\top$. At $(2, 1)$: $f = 4 + 4 = 8$ and $\nabla f = [4, 8]$. Step 1: $[2, 1] - 0.1[4, 8] = [1.6, 0.2]$, with $f = 2.56 + 0.16 = 2.72$. Then $\nabla f = [3.2, 1.6]$. Step 2: $[1.6, 0.2] - 0.1[3.2, 1.6] = [1.28, 0.04]$, with $f = 1.6384 + 0.0064 = 1.6448$. The loss went $8 \to 2.72 \to 1.6448$. Notice that $y$ shrinks much faster than $x$: the gradient is bigger in the steep $y$-direction. That is the zig-zag behaviour of narrow valleys.

Chapter 2.5

Gradients of Vector-Valued Functions

A neural-network layer takes a vector in and gives a vector out. How does the output react when you nudge the input? The answer is a table of slopes called the Jacobian matrix. It is the single most important object you will carry into backpropagation.

  • Tell scalar-valued and vector-valued functions apart, and meet curves, maps and neural layers
  • Build the Jacobian matrix: shape $m\times n$, entry $(i, j) = \partial F_i/\partial x_j$
  • Know the shape of every derivative: vector by vector, scalar by vector (the gradient), vector by scalar (velocity)
  • See the Jacobian as the local linear map, and its determinant as the local area scale
  • Compute Jacobians of a linear map, polar coordinates, a dense layer and softmax
  • Apply the chain rule with Jacobians (a product of matrices) and check the shapes

Why this chapter matters. Everything in a neural network is a chain of vector functions (layer after layer). The Jacobian tells you how each layer passes a nudge along, and multiplying Jacobians is exactly what backpropagation does (Chapter 2.9). Take your time here.

Scalar-valued functions: many numbers in, one number out core

A scalar is just a single plain number. A scalar-valued function is a function whose answer is one number, no matter how many numbers go in.

Examples: the temperature at a spot on a map (position in, one temperature out); the price of a house from its size, age and distance to the city; and, most important for us, the loss of a model, which turns all of the weights into one number: "how wrong am I?".

You already know the derivative of such a function: it is the gradient from Chapter 2.4, one slope for each input. This chapter extends the idea to functions whose answer is a list.

$f(x_1, x_2) = x_1^2 + x_2^2$ takes a vector in $\mathbb{R}^2$ and returns a number.

  1. Input $\mathbf{x} = [3, 4]^\top$ (two numbers).
  2. Output $f(\mathbf{x}) = 9 + 16 = 25$ (one number).
  3. Derivative: $\nabla f = [2x_1, 2x_2]^\top = [6, 8]^\top$ (two numbers, one per input).

Shapes: in $n = 2$ numbers, out $m = 1$ number, derivative $n = 2$ numbers.

A scalar-valued function of $n$ variables is a function

$$f:\ \mathbb{R}^n\to\mathbb{R}, \qquad \mathbf{x} = [x_1, \dots, x_n]^\top \mapsto f(\mathbf{x}) \in \mathbb{R}.$$

Its derivative is the gradient $\nabla f(\mathbf{x})\in\mathbb{R}^n$: a column vector of the $n$ partial derivatives $\partial f/\partial x_j$. In this chapter we will see it is the special case $m = 1$ of the Jacobian.

Why do we need it?

To measure "how good or bad" with one number, we need a function that squeezes a lot of information down to a single score. Only a single score can be minimised.

Where is it used?

Every loss function: mean squared error, cross-entropy, a regularised loss. Also a model's single output (a house price), a likelihood, and a reward in reinforcement learning.

How is it used?

Compute the score for the current weights, then compute its gradient (one slope per weight) and nudge the weights against it. The output being a single number is what makes "lower is better" well-defined.

Drag the purple point (the input $\mathbf{x}$ has two numbers). The colour shows the single output $f(\mathbf{x})$. The readout shows the derivative as a column (the gradient, 2 × 1) and as a row (its transpose, 1 × 2, which is the Jacobian of a scalar function). The numbers are the same; only the layout differs.

"Scalar" refers to the output. A function of a million variables is still scalar-valued if it returns one number. The input can be a vector; the output is what we call scalar or vector.

Quick check: is $f(x, y, z) = xy + z$ scalar-valued? What are the sizes of its input, output and gradient?

Yes. The input is 3 numbers, the output is 1 number, and the gradient $[y, x, 1]^\top$ has 3 numbers.

Vector-valued functions: a list in, a list out core

Now let the answer be a list. A GPS maps the time $t$ to your position: two numbers (latitude and longitude). A paint mixer maps three dials to the three colours (red, green, blue) it produces. A neural-network layer maps its input vector to an output vector.

The easy way to think about it: a vector-valued function is several scalar-valued functions stacked in a column, all sharing the same inputs. Each output number has its own "recipe" that uses all the inputs.

Let $\mathbf{F}:\mathbb{R}^2\to\mathbb{R}^3$ be

$$\mathbf{F}(x_1, x_2) = \begin{bmatrix} F_1 \\ F_2 \\ F_3 \end{bmatrix} = \begin{bmatrix} x_1x_2 \\ x_1 + x_2 \\ x_1^2 \end{bmatrix}.$$
  1. Input $[2, 3]^\top$ (2 numbers).
  2. $F_1 = 2\cdot3 = 6$, $F_2 = 2 + 3 = 5$, $F_3 = 2^2 = 4$.
  3. Output $\mathbf{F}(2, 3) = [6, 5, 4]^\top$ (3 numbers).

Notice that $F_3$ does not use $x_2$ at all. We will see that fact show up as a zero in the Jacobian.

A vector-valued function is a function $\mathbf{F}:\mathbb{R}^n\to\mathbb{R}^m$,

$$\mathbf{F}(\mathbf{x}) = \begin{bmatrix} F_1(\mathbf{x}) \\ F_2(\mathbf{x}) \\ \vdots \\ F_m(\mathbf{x}) \end{bmatrix},$$

where each component function $F_i:\mathbb{R}^n\to\mathbb{R}$ is scalar-valued. The case $m = 1$ is the scalar-valued case from the last section. Capital bold $\mathbf{F}$ reminds us that the output is a vector. (A linear map $\mathbf{x}\mapsto A\mathbf{x}$ is the simplest example.)

Why do we need it?

Most things we model have several outputs, and most models are built from stages that each turn a vector into another vector. We need to be able to talk about, and differentiate, such functions.

Where is it used?

Every neural-network layer, a softmax that outputs one probability per class, word embeddings, the motion of a robot arm (joint angles in, hand position out), and coordinate changes (polar to Cartesian).

How is it used?

Evaluate it component by component. To differentiate it, differentiate each component with respect to each input. That table of derivatives is the Jacobian (next sections).

The function is $\mathbf{F}(x_1, x_2) = [x_1x_2,\ x_1 + x_2,\ x_1^2]^\top$. Move $x_1$ and watch all three bars react. Then move only $x_2$ and notice that the third bar does not move at all: $F_3$ does not depend on $x_2$. Which input changes which output is exactly what the Jacobian records.

Count the numbers. In $\mathbf{F}:\mathbb{R}^n\to\mathbb{R}^m$, $n$ is the number of inputs and $m$ the number of outputs. The two numbers do not have to match. Getting these two straight now saves a lot of confusion with matrix shapes later.

Quick check: $\mathbf{F}(x, y) = [x + y,\ x - y]^\top$. What is $\mathbf{F}(5, 2)$, and which $\mathbb{R}^n\to\mathbb{R}^m$ is it?

$\mathbf{F}(5, 2) = [7, 3]^\top$. It takes 2 numbers and returns 2 numbers, so $\mathbb{R}^2\to\mathbb{R}^2$.

Vector functions you will meet: curves, maps and layers

Vector-valued functions come in three everyday flavours. Each answers a different question:

  • A curve $\mathbf{r}:\mathbb{R}\to\mathbb{R}^m$. One number in (time $t$), a position out. Question: where am I at time $t$? The picture is a path.
  • A map of the plane $\mathbf{F}:\mathbb{R}^2\to\mathbb{R}^2$. A point in, a point out. Question: where does this point land? The picture is the plane being stretched, rotated or bent, like a rubber sheet.
  • A neural layer $\mathbf{x}\mapsto\sigma(W\mathbf{x} + \mathbf{b})$. A vector in, a vector out. Question: what features does the layer compute?
  1. Curve. $\mathbf{r}(t) = [\cos t,\ \sin t]^\top$. At $t = 0$: $[1, 0]$. At $t = \pi/2$: $[0, 1]$. As $t$ grows it goes round the unit circle.
  2. Map. Polar to Cartesian, $\mathbf{F}(r, \theta) = [r\cos\theta,\ r\sin\theta]^\top$. The input $(2, \pi/6)$ ("2 steps at 30°") lands at $[2\cdot0.866,\ 2\cdot0.5] = [1.732,\ 1]$.
  3. Layer. $W = \begin{bmatrix} 1 & -1 \\ 2 & 1 \end{bmatrix}$, $\mathbf{b} = [0.5, -1]^\top$, ReLU activation. For $\mathbf{x} = [1, 2]^\top$: $W\mathbf{x} = [1 - 2,\ 2 + 2] = [-1, 4]$, add $\mathbf{b}$ to get $\mathbf{z} = [-0.5, 3]$, then ReLU (replace negatives by 0): $\mathbf{a} = [0, 3]$.

All three are vector-valued functions $\mathbb{R}^n\to\mathbb{R}^m$, with different sizes:

Kind$n$ (inputs)$m$ (outputs)Typical name
Curve12 or 3$\mathbf{r}(t)$
Map of the plane22$\mathbf{F}(x, y)$
Neural layerany $n$any $m$$\mathbf{a} = \sigma(W\mathbf{x} + \mathbf{b})$

A vector function is simply any function whose output is a vector. We can compose them (feed the output of one into the next), and a neural network is exactly a long composition of layers.

Why do we need it?

These three shapes cover almost everything: following something through time, changing coordinates, and stacking layers of a network. Seeing the common pattern lets one tool (the Jacobian) handle all of them.

Where is it used?

Curves: trajectories in physics and robotics, training paths of the weights. Maps: polar and spherical coordinates, normalising flows and image warping. Layers: every neural network.

How is it used?

Name the sizes $n$ and $m$ first. Then evaluate component by component. Later, differentiate the same way, and the Jacobian will have shape $m\times n$.

Drag the purple point on the input plane (left). The green dot on the output plane (right) shows where it lands. The blue and orange grid lines show how the whole grid is bent. The thick blue line is "first input fixed", the thick orange line is "second input fixed". Try each map: polar turns a rectangle into a disc, the linear map slants the grid evenly, the wavy one bends it, and the squaring map wraps the plane twice around the origin (two different inputs, such as $(1, 0)$ and $(-1, 0)$, land on the same output).

A function is the whole machine, not one output. A map of the plane has two output numbers at each point; to draw it we need two pictures (input plane and output plane), or a deformed grid as above.

Quick check: what are $n$ and $m$ for a layer that takes a 784-pixel image and returns 10 class scores?

$n = 784$ inputs and $m = 10$ outputs, so the layer is a function $\mathbb{R}^{784}\to\mathbb{R}^{10}$.

The Jacobian matrix core

A vector function with $n$ inputs and $m$ outputs has $m\times n$ slopes. For each output and each input you can ask: "if I nudge this input, how fast does this output move?" Write all those slopes in a table. That table is the Jacobian matrix.

There is a second, geometric way to see it. Zoom in on a point of a bent map (like the polar grid) until it looks straight. Zoomed in, the map looks like a linear map, a matrix that stretches and turns the tiny neighbourhood. That matrix is the Jacobian. It is the "best straight-line copy" of the function, just as the derivative was the slope of the best straight line in one variable.

How to read the table. Each row belongs to one output (it is the gradient of that output, written sideways). Each column belongs to one input (it says how the whole output vector reacts when that one input is nudged).

Example 1. $\mathbf{F}(x, y) = [x^2y,\ 5x + y^2]^\top$, with $n = 2$ inputs and $m = 2$ outputs, so the Jacobian is $2\times2$.

  1. Row 1 (output $F_1 = x^2y$): $\partial F_1/\partial x = 2xy$, $\partial F_1/\partial y = x^2$.
  2. Row 2 (output $F_2 = 5x + y^2$): $\partial F_2/\partial x = 5$, $\partial F_2/\partial y = 2y$.
  3. So $J = \begin{bmatrix} 2xy & x^2 \\ 5 & 2y \end{bmatrix}$. At $(1, 2)$: $J = \begin{bmatrix} 4 & 1 \\ 5 & 4 \end{bmatrix}$.

Check by nudging. $\mathbf{F}(1, 2) = [2, 9]^\top$. Nudge the input by $\mathbf{h} = [0.01, -0.02]^\top$. The Jacobian predicts the change $J\mathbf{h} = [4\cdot0.01 + 1\cdot(-0.02),\ 5\cdot0.01 + 4\cdot(-0.02)] = [0.02, -0.03]$, so $\mathbf{F}\approx[2.02, 8.97]$. The true value is $\mathbf{F}(1.01, 1.98) = [1.0201\cdot1.98,\ 5.05 + 3.9204] = [2.019798,\ 8.9704]$ ✓.

Example 2: polar to Cartesian. $\mathbf{F}(r, \theta) = [r\cos\theta,\ r\sin\theta]^\top$. Differentiate each component:

$$J = \begin{bmatrix} \partial x/\partial r & \partial x/\partial\theta \\ \partial y/\partial r & \partial y/\partial\theta \end{bmatrix} = \begin{bmatrix} \cos\theta & -r\sin\theta \\ \sin\theta & r\cos\theta \end{bmatrix}, \qquad \det J = r\cos^2\theta + r\sin^2\theta = r.$$

At $(r, \theta) = (2, \pi/6)$: $J = \begin{bmatrix} 0.866 & -1 \\ 0.5 & 1.732 \end{bmatrix}$ and $\det J = 0.866\cdot1.732 + 1\cdot0.5 = 1.5 + 0.5 = 2 = r$ ✓.

What the determinant says. A tiny rectangle of size $\Delta r\times\Delta\theta$ in the $(r, \theta)$ plane lands on a patch of area about $r\,\Delta r\,\Delta\theta$ in the $(x, y)$ plane. Far from the origin (large $r$), the same angle step sweeps out a longer arc, so the area is stretched by $r$. This is the famous "extra $r$" that appears when you integrate in polar coordinates: the area scale of any change of coordinates is $|\det J|$.

For $\mathbf{F}:\mathbb{R}^n\to\mathbb{R}^m$, the Jacobian matrix at $\mathbf{x}$ is the $m\times n$ matrix whose $(i, j)$ entry is the partial derivative of output $i$ with respect to input $j$:

$$J_{\mathbf{F}}(\mathbf{x}) = \frac{\partial\mathbf{F}}{\partial\mathbf{x}} = \begin{bmatrix} \dfrac{\partial F_1}{\partial x_1} & \cdots & \dfrac{\partial F_1}{\partial x_n} \\ \vdots & \ddots & \vdots \\ \dfrac{\partial F_m}{\partial x_1} & \cdots & \dfrac{\partial F_m}{\partial x_n} \end{bmatrix} \in\mathbb{R}^{m\times n}.$$
  • Layout convention (used everywhere in this guide). One row per output, one column per input: shape $m\times n$, entry $(i, j) = \partial F_i/\partial x_j$. This is called numerator layout. Some books use the transpose ($n\times m$, "denominator layout"); always check which one a book uses. We use the same convention as the matrix-calculus chapter of the Linear Algebra guide.
  • Row $i$ is the gradient of $F_i$, written as a row: $(\nabla F_i)^\top$. Column $j$ is $\partial\mathbf{F}/\partial x_j$, the response of the whole output to nudging input $j$.
  • Best linear copy: $\mathbf{F}(\mathbf{x} + \mathbf{h}) \approx \mathbf{F}(\mathbf{x}) + J_{\mathbf{F}}(\mathbf{x})\,\mathbf{h}$ for small $\mathbf{h}$. (This is multiplication of an $m\times n$ matrix by an $n$-vector, giving an $m$-vector, so the shapes fit.)
  • When $m = n$ the Jacobian is square, and $\det J$ is the local area (or volume) scale factor. A negative determinant means the map flips orientation (like a mirror). $\det J = 0$ means the map squashes a patch onto a line or a point.
Why do we need it?

With many outputs and many inputs, one slope is not enough. We need the whole table of "which input moves which output, and how fast" in order to pass a small change through a function.

Where is it used?

Backpropagation multiplies the Jacobians of the layers. Also change of variables in integrals (the determinant), Newton's method for systems of equations, the extended Kalman filter, robot arms (joint speeds to hand speed), and normalising flows.

How is it used?

Differentiate every output component with respect to every input, and arrange the numbers with outputs as rows and inputs as columns. Evaluate at your point. Multiply it by a small input change to predict the output change.

Drag the purple point on the input plane. The small square (side $h$) becomes the curved orange patch on the output plane. The two arrows are the columns of the Jacobian times $h$ (blue: first input nudged, green: second input nudged); they form the straight dashed parallelogram that approximates the patch. Shrink $h$ and the match becomes almost perfect. Compare the true area ratio with $\det J$. On the squaring map, drag to the centre: $\det J \to 0$ and the patch collapses.

The map $\mathbf{F}(\theta, \varphi) = 2[\sin\varphi\cos\theta,\ \sin\varphi\sin\theta,\ \cos\varphi]$ takes 2 numbers (longitude $\theta$, angle from the top $\varphi$) to a point on a sphere in 3D, so $n = 2$, $m = 3$ and $J$ is $3\times2$. Drag the purple dot over the sphere (rotate to see the back). The blue arrow is column 1 (move $\theta$) and the green arrow is column 2 (move $\varphi$). The two columns span the teal tangent plane. Drag to the top of the sphere: the blue column shrinks to zero (all longitudes meet at the pole).

Rows are outputs, columns are inputs. A common slip is to transpose it. A quick test: the Jacobian multiplies an input-sized vector to give an output-sized vector, so it must have $n$ columns and $m$ rows.

The Jacobian depends on the point. Unless the function is linear, the matrix changes as you move (see how the arrows change as you drag).

It is a local picture. The parallelogram only matches the patch for small $h$. Far away the true map bends.

Quick check: find the Jacobian of $\mathbf{F}(x, y) = [x + y,\ xy]^\top$ at $(2, 3)$.

$J = \begin{bmatrix} 1 & 1 \\ y & x \end{bmatrix}$. At $(2, 3)$: $J = \begin{bmatrix} 1 & 1 \\ 3 & 2 \end{bmatrix}$, and $\det J = 2 - 3 = -1$ (this map flips orientation near that point).

The derivative of a vector with respect to a vector core

"Derivative of a vector $\mathbf{F}$ with respect to a vector $\mathbf{x}$" is just another name for the Jacobian. The question it answers is: when the whole input vector is nudged a little, how does the whole output vector change?

Because both sides are vectors, the answer is a matrix. A good habit: always write down the shape first. Output size $m$, input size $n$, so the derivative is $m\times n$. If you can state the shape, you can often spot an error before you compute anything.

The simplest vector function is a linear one, $\mathbf{F}(\mathbf{x}) = A\mathbf{x}$. Its derivative is as nice as it gets: it is the matrix $A$ itself, everywhere. (The scalar version: the derivative of $a x$ is $a$.)

Let $A = \begin{bmatrix} 2 & 1 & 0 \\ -1 & 3 & 4 \end{bmatrix}$ ($2\times3$) and $\mathbf{F}(\mathbf{x}) = A\mathbf{x}$ with $\mathbf{x}\in\mathbb{R}^3$. So $n = 3$, $m = 2$, and the Jacobian will be $2\times3$.

  1. Write the outputs: $F_1 = 2x_1 + 1x_2 + 0x_3$ and $F_2 = -1x_1 + 3x_2 + 4x_3$.
  2. Differentiate: $\partial F_1/\partial x_1 = 2$, $\partial F_1/\partial x_2 = 1$, $\partial F_1/\partial x_3 = 0$. And $\partial F_2/\partial x_1 = -1$, $\partial F_2/\partial x_2 = 3$, $\partial F_2/\partial x_3 = 4$.
  3. Collect: $J = \begin{bmatrix} 2 & 1 & 0 \\ -1 & 3 & 4 \end{bmatrix} = A$.

Why in general. $F_i = \sum_j A_{ij}x_j$. Differentiating with respect to $x_j$, only the term with that $j$ survives: $\partial F_i/\partial x_j = A_{ij}$. That is exactly the $(i, j)$ entry of $A$. Adding a constant vector, $\mathbf{F}(\mathbf{x}) = A\mathbf{x} + \mathbf{b}$, changes nothing, because constants have zero derivative.

(This is why the matrix of a linear map from the Linear Algebra guide is its own derivative: a linear map already is its best linear copy.)

For $\mathbf{F}:\mathbb{R}^n\to\mathbb{R}^m$ the derivative of the vector $\mathbf{F}$ with respect to the vector $\mathbf{x}$ is the Jacobian,

$$\frac{\partial\mathbf{F}}{\partial\mathbf{x}} = J_{\mathbf{F}}\in\mathbb{R}^{m\times n}, \qquad \Bigl(\frac{\partial\mathbf{F}}{\partial\mathbf{x}}\Bigr)_{ij} = \frac{\partial F_i}{\partial x_j}.$$

The shapes of every derivative in this chapter (numerator layout):

FunctionNameDerivativeShape
$f:\mathbb{R}\to\mathbb{R}$ordinary function$f'(x)$$1\times1$
$f:\mathbb{R}^n\to\mathbb{R}$scalar-valued$(\nabla f)^\top$ (gradient, as a row)$1\times n$
$\mathbf{r}:\mathbb{R}\to\mathbb{R}^m$curve$\mathbf{r}'(t)$ (velocity)$m\times1$
$\mathbf{F}:\mathbb{R}^n\to\mathbb{R}^m$vector-valued$J_{\mathbf{F}}$$m\times n$

For a linear map, $\mathbf{F}(\mathbf{x}) = A\mathbf{x} + \mathbf{b}$, $J_{\mathbf{F}} = A$ at every point.

Why do we need it?

A layer of a network is a vector function. To train it we must know how its output vector responds to a change in its input vector (and weights). The Jacobian is that response, in matrix form.

Where is it used?

Linear layers (Jacobian equals the weight matrix), activation layers (diagonal Jacobian), softmax, normalisation layers, and the sensitivity of any model output to its input features.

How is it used?

State $n$ and $m$ and write the shape $m\times n$. Fill each entry $(i, j)$ with $\partial F_i/\partial x_j$. Check that the Jacobian times an input-sized vector gives an output-sized vector.

Set the number of inputs $n$ and outputs $m$. The grid is the Jacobian, with one cell per pair (output $i$, input $j$). Try $m = 1$ (a scalar function: the Jacobian becomes a single row, the transposed gradient), $n = 1$ (a curve: a single column, the velocity) and $m = n$ (a square matrix that has a determinant).

Edit the matrix $A$ (it is $2\times3$: three inputs, two outputs) and move the inputs. The computer estimates the Jacobian by nudging each input a tiny bit and watching the output (a numerical Jacobian). It equals $A$ every time, wherever $\mathbf{x}$ is, and also when a bias is added.

Do not confuse the Jacobian of $A\mathbf{x}$ with $A\mathbf{x}$'s derivative with respect to $A$. Here $A$ is fixed and $\mathbf{x}$ moves. Taking derivatives with respect to the matrix $A$ is a different (and bigger) object, treated in Chapter 2.6.

Shape mistakes are the number-one bug. If the Jacobian has the wrong shape, the chain rule's matrix products will not even be defined. Always write $m\times n$ first.

Quick check: $\mathbf{F}:\mathbb{R}^4\to\mathbb{R}^2$. What shape is its Jacobian, and how many slopes does it hold?

$2\times4$ (2 outputs = rows, 4 inputs = columns). It holds $2\cdot4 = 8$ slopes.

The derivative of a scalar with respect to a vector (the gradient) core

When the output is a single number, the table of slopes has only one row: one slope per input. That one row is the gradient, laid on its side.

Here is the one subtle point. There are two natural ways to write the same $n$ numbers:

  • As a column ($n\times1$): the gradient $\nabla f$. It lives in the same space as the input $\mathbf{x}$, so you can add it to or subtract it from $\mathbf{x}$. This is the form for gradient descent.
  • As a row ($1\times n$): the Jacobian $J_f = (\nabla f)^\top$. It is a matrix that multiplies on the left of a vector, which is what the chain rule needs.

Same numbers, different layout. In this guide: the gradient is a column, and the Jacobian of a scalar function is its transpose.

A tiny model predicts $\hat y = \mathbf{w}\cdot\mathbf{x}$ and its loss is $L(\mathbf{w}) = (\hat y - y)^2$. Take the data point $\mathbf{x} = [1, 2, -1]^\top$ with target $y = 3$, and the weights $\mathbf{w} = [1, 0, 2]^\top$.

  1. Prediction: $\hat y = 1\cdot1 + 0\cdot2 + 2\cdot(-1) = -1$.
  2. Error: $e = \hat y - y = -1 - 3 = -4$. Loss: $L = e^2 = 16$.
  3. Since $e = \sum_j w_jx_j - y$, we have $\partial e/\partial w_j = x_j$.
  4. Chain rule on $L = e^2$: $\partial L/\partial w_j = 2e\cdot x_j$. With $2e = -8$: $\partial L/\partial\mathbf{w} = -8\cdot[1, 2, -1] = [-8, -16, 8]$.

As a column: $\nabla L = [-8, -16, 8]^\top$ ($3\times1$). As a Jacobian: $J_L = [-8,\ -16,\ 8]$ ($1\times3$).

Check by nudging. Raise $w_1$ by $0.001$: $\hat y = -0.999$, $e = -3.999$, $L = 15.992001$. The change is $-0.007999$, and $-0.007999/0.001\approx-8$ ✓. (Gradient descent would now move $\mathbf{w}$ against $\nabla L$: $w_1$ goes up, $w_2$ goes up, $w_3$ goes down.)

For $f:\mathbb{R}^n\to\mathbb{R}$:

$$\nabla f(\mathbf{x}) = \begin{bmatrix} \partial f/\partial x_1 \\ \vdots \\ \partial f/\partial x_n \end{bmatrix}\ (n\times1), \qquad J_f(\mathbf{x}) = \frac{\partial f}{\partial\mathbf{x}} = \bigl[\partial f/\partial x_1\ \cdots\ \partial f/\partial x_n\bigr] = (\nabla f)^\top\ (1\times n).$$

So $\nabla f = J_f^\top$. The linear approximation can be written either way: $f(\mathbf{x}+\mathbf{h}) \approx f(\mathbf{x}) + J_f\,\mathbf{h} = f(\mathbf{x}) + \nabla f\cdot\mathbf{h}$ (a $1\times n$ times $n\times1$ is a number).

A useful special case: for $f(\mathbf{x}) = \mathbf{a}^\top\mathbf{x}$ ($=a_1x_1 + \dots + a_nx_n$) we get $\partial f/\partial x_j = a_j$, so $\nabla f = \mathbf{a}$ and $J_f = \mathbf{a}^\top$.

Why do we need it?

Training needs the slope of the loss with respect to every weight, collected in one object, so that the weights can be moved all at once. And we need it in a form that fits the chain rule.

Where is it used?

Every loss in machine learning: for linear and logistic regression, neural networks and SVMs, the gradient of the loss with respect to the weights is what gradient descent follows. In backpropagation the first thing computed is the row $\partial L/\partial\mathbf{a}$ at the output.

How is it used?

Compute the partial derivative for each weight and stack them as a column to get $\nabla L$; update $\mathbf{w}\leftarrow\mathbf{w} - \eta\nabla L$. When chaining, use the row version $(\nabla L)^\top$ on the left of Jacobian matrices.

The loss is $L(\mathbf{w}) = (\mathbf{w}\cdot\mathbf{x} - y)^2$ with $\mathbf{x} = [1, 2, -1]$ and $y = 3$. Move the three weight sliders. The gradient (a column of 3 slopes) tells you which weights to raise or lower. Press Take a gradient step a few times and watch the loss fall. Then compare with the numerical check.

"Gradient" is a column, "Jacobian of a scalar function" is a row. They hold the same numbers. If a formula multiplies a gradient on the left of a matrix, transpose it first. Books that use the other layout will show the gradient as a row; the maths is the same, only the bookkeeping differs.

Quick check: $f(\mathbf{x}) = 3x_1 - x_2 + 4x_3$. Write $\nabla f$ and $J_f$, with shapes.

$\nabla f = [3, -1, 4]^\top$ ($3\times1$) and $J_f = [3, -1, 4]$ ($1\times3$). It is $\mathbf{a}^\top\mathbf{x}$ with $\mathbf{a} = [3, -1, 4]^\top$.

The derivative of a vector with respect to a scalar (velocity) core

Flip the last case around: one input, several outputs. Time $t$ goes in, a position $\mathbf{r}(t)$ comes out. The derivative says how fast each coordinate changes, and stacked together it is the velocity.

The velocity is an arrow attached to the moving point. It points along the path (the direction of travel), and its length is the speed. The Jacobian has one column: $m\times1$.

  1. $\mathbf{r}(t) = [t,\ t^2]^\top$ (a parabola). Differentiate each entry: $\mathbf{r}'(t) = [1,\ 2t]^\top$. At $t = 1.5$: $\mathbf{r}' = [1, 3]$, so the point is moving right 1 and up 3 per second; the speed is $\sqrt{1 + 9}\approx3.16$.
  2. $\mathbf{r}(t) = [\cos t,\ \sin t]^\top$ (a circle). $\mathbf{r}'(t) = [-\sin t,\ \cos t]^\top$. At $t = \pi/2$: the point is at $[0, 1]$ (top) and $\mathbf{r}' = [-1, 0]$ (moving left). The speed is $\sqrt{\sin^2t + \cos^2t} = 1$ always. The velocity is perpendicular to the position (check: $[0, 1]\cdot[-1, 0] = 0$).
  3. A 3D helix $\mathbf{r}(t) = [2\cos t,\ 2\sin t,\ 0.4(t - 2\pi)]^\top$ has $\mathbf{r}'(t) = [-2\sin t,\ 2\cos t,\ 0.4]^\top$ and constant speed $\sqrt{4 + 0.16}\approx2.04$.

For a curve $\mathbf{r}:\mathbb{R}\to\mathbb{R}^m$ the derivative is taken entry by entry:

$$\mathbf{r}'(t) = \frac{d\mathbf{r}}{dt} = \lim_{h\to0}\frac{\mathbf{r}(t + h) - \mathbf{r}(t)}{h} = \begin{bmatrix} r_1'(t) \\ \vdots \\ r_m'(t) \end{bmatrix}\ (m\times1).$$

It is a column vector, tangent to the path. Its length $\|\mathbf{r}'(t)\|$ is the speed. It is the Jacobian when $n = 1$.

Why do we need it?

To describe how something moves or changes over time, we need the rate of change of every coordinate together, as one arrow.

Where is it used?

Physics and robotics (velocity of a moving object), the path of the weights during training (each step is a small piece of a curve in weight space), recurrent networks that evolve a hidden state over time, and neural ODEs.

How is it used?

Differentiate each component with respect to $t$ and stack the results. Then evaluate at a chosen $t$ to get the velocity arrow, and use its length for the speed. Over a short time $\Delta t$, the position changes by about $\mathbf{r}'(t)\,\Delta t$.

Drag the purple point along the curve (or move the slider for $t$). The orange arrow is the velocity $\mathbf{r}'(t)$: it always touches the curve. Watch its length (the speed) change: on the spiral it speeds up, on the circle it stays constant. The readout shows each component derivative.

The orange curve is the helix $\mathbf{r}(t)$. Drag the purple dot along it (it snaps to the curve), or use the slider. The green arrow is the velocity $\mathbf{r}'(t)$: three components, one for each axis. Press Top: from above the helix is a circle, and the arrow's horizontal part is the circle's velocity. Press Side: the vertical part is the constant climb 0.4.

Velocity is a vector, speed is a number. Velocity has a direction (along the path), speed is just its length.

Same function, different questions. $\mathbf{r}'(t)$ (derivative of a vector by a scalar) is a column, while the derivative of a scalar by a vector is a row in the numerator layout. Both are just Jacobians with the right shapes.

Quick check: $\mathbf{r}(t) = [3t,\ t^2,\ 5]^\top$. Find $\mathbf{r}'(2)$.

$\mathbf{r}'(t) = [3,\ 2t,\ 0]^\top$ (the constant 5 has derivative 0). At $t = 2$: $[3, 4, 0]^\top$, with speed $\sqrt{9 + 16} = 5$.

Jacobians inside a neural network: dense layer and softmax core

Crucial for neural networks. The two Jacobians in this section appear in every network you will ever train. In Chapter 2.9 (backpropagation) you will multiply them together, layer by layer. Spend time on these.

A dense layer does two things: first a linear step $\mathbf{z} = W\mathbf{x} + \mathbf{b}$ (mix the inputs), then an activation $\mathbf{a} = \sigma(\mathbf{z})$ applied to every entry separately (bend the result, for example ReLU or sigmoid).

Nudge an input $x_j$. It first reaches output $i$ through the weight $W_{ij}$ (the linear step), and then the activation squashes or passes that nudge by its own slope $\sigma'(z_i)$. So the Jacobian is just the weight matrix with each row scaled by the activation's slope.

Softmax turns a vector of scores (logits) into probabilities that add up to 1. Raise one score and its probability rises, but the others must fall to keep the total at 1. So the Jacobian has positive numbers on the diagonal and negative numbers elsewhere.

Dense layer with ReLU. $W = \begin{bmatrix} 1 & -1 \\ 2 & 1 \end{bmatrix}$, $\mathbf{b} = [0.5, -1]^\top$, $\mathbf{x} = [1, 2]^\top$ (the layer from the earlier example). We found $\mathbf{z} = [-0.5, 3]$ and $\mathbf{a} = [0, 3]$.

  1. ReLU's slope is $1$ where $z>0$ and $0$ where $z<0$. So $\sigma'(\mathbf{z}) = [0, 1]$.
  2. Scale row 1 of $W$ by $0$ and row 2 by $1$: $J = \operatorname{diag}(0, 1)\,W = \begin{bmatrix} 0 & 0 \\ 2 & 1 \end{bmatrix}$.
  3. Meaning: output 1 is "switched off" (its $z$ is negative), so nudging the input does nothing to it (row of zeros). Output 2 passes the nudge on, with weights $[2, 1]$.

Softmax. For $\mathbf{z} = [0, 0, 0]$ all three probabilities are $s_i = 1/3$. Then $J_{ii} = s_i(1 - s_i) = \tfrac13\cdot\tfrac23 = \tfrac29$ and $J_{ij} = -s_is_j = -\tfrac19$, so

$$J = \begin{bmatrix} 2/9 & -1/9 & -1/9 \\ -1/9 & 2/9 & -1/9 \\ -1/9 & -1/9 & 2/9 \end{bmatrix}, \quad\text{each row sums to } \tfrac29 - \tfrac19 - \tfrac19 = 0.$$

For $\mathbf{z} = [2, 1, 0]$ we get $\mathbf{s}\approx[0.665,\ 0.245,\ 0.090]$ and a Jacobian with first row $[0.223,\ -0.163,\ -0.060]$.

Dense layer. For $\mathbf{a} = \sigma(\mathbf{z})$ with $\mathbf{z} = W\mathbf{x} + \mathbf{b}$ ($W$ is $m\times n$):

$$a_i = \sigma(z_i),\quad z_i = \sum_j W_{ij}x_j + b_i \quad\Longrightarrow\quad \frac{\partial a_i}{\partial x_j} = \sigma'(z_i)\,W_{ij},$$ $$J_{\mathbf{a}}(\mathbf{x}) = \operatorname{diag}\bigl(\sigma'(\mathbf{z})\bigr)\,W \qquad (m\times m)(m\times n) = m\times n.$$

(The step $\partial a_i/\partial x_j = \sigma'(z_i)\,\partial z_i/\partial x_j$ is the chain rule from one variable; $\partial z_i/\partial x_j = W_{ij}$ is the linear-map Jacobian from before; the bias drops out.) Common slopes: sigmoid $\sigma' = \sigma(1 - \sigma)$ (at most $0.25$); $\tanh' = 1 - \tanh^2$; ReLU' $= 1$ for $z>0$, else $0$.

Softmax. $s_i = e^{z_i}/S$ with $S = \sum_k e^{z_k}$. Differentiate with the quotient rule:

  • $j = i$: $\dfrac{\partial s_i}{\partial z_i} = \dfrac{e^{z_i}S - e^{z_i}e^{z_i}}{S^2} = s_i - s_i^2 = s_i(1 - s_i)$.
  • $j \neq i$: $\dfrac{\partial s_i}{\partial z_j} = \dfrac{0\cdot S - e^{z_i}e^{z_j}}{S^2} = -s_is_j$.
$$J_{\mathbf{s}} = \operatorname{diag}(\mathbf{s}) - \mathbf{s}\mathbf{s}^\top.$$

It is symmetric, and every row and every column sums to zero. Each column sums to zero because $\sum_i s_i = 1$ never changes, so the rises and falls cancel. Each row sums to zero because adding the same number to all the logits does not change any $s_i$ at all. That second fact is why softmax can safely be computed after subtracting the largest logit (it avoids huge numbers like $e^{1000}$).

Why do we need it?

To train a network we must know how a change in a layer's input (or weights) changes its output, and how that change moves on to the next layer. These two Jacobians do that for the most common building blocks.

Where is it used?

Every fully connected layer in an MLP, the output layer of every classifier (softmax with cross-entropy), attention weights in Transformers (a softmax over scores), and the analysis of vanishing gradients (slopes of $\sigma'$ smaller than 1 shrink the signal).

How is it used?

Dense layer: compute $\mathbf{z}$, evaluate $\sigma'$ at each $z_i$, scale the rows of $W$. Softmax: compute $\mathbf{s}$ and form $\operatorname{diag}(\mathbf{s}) - \mathbf{s}\mathbf{s}^\top$. In practice libraries never build these matrices for big layers; they just apply the same rules to vectors.

The layer has 3 inputs and 2 outputs: $\mathbf{a} = \sigma(W\mathbf{x} + \mathbf{b})$. Edit $W$ and $\mathbf{b}$, move the inputs, and pick an activation. The Jacobian ($2\times3$) is $W$ with each row scaled by the slope $\sigma'(z_i)$. It matches the numerical Jacobian (nudge each input). With ReLU, a negative $z_i$ gives a whole row of zeros; with sigmoid or tanh, big $|z_i|$ makes the row tiny ("saturation", the cause of vanishing gradients).

Move the three scores (logits). The bars are the probabilities $s_i$. In the Jacobian below, positive entries are blue and negative entries are orange. Raise $z_1$ and look at column 1: $s_1$ goes up (positive, on the diagonal) while $s_2$ and $s_3$ go down (negative). Every row sums to $0$, and the matrix is symmetric. Set all three scores equal to see $\tfrac29$ and $-\tfrac19$.

Activations act entry by entry, so their Jacobian is diagonal. $a_i$ depends only on $z_i$. Softmax is different: every output depends on all inputs, so its Jacobian is a full matrix.

A zero slope kills the signal. ReLU with $z<0$ and sigmoid with large $|z|$ give (nearly) zero rows. Gradients passing through them become (nearly) zero, a major reason deep networks can be hard to train.

Softmax Jacobian is singular. Its rows sum to zero, so $J\mathbf{1} = \mathbf{0}$: it has no inverse.

Quick check: the sigmoid has $\sigma(0) = 0.5$. What is $\sigma'(0)$, and what does that say about its Jacobian?

$\sigma'(0) = 0.5\cdot0.5 = 0.25$, the largest slope the sigmoid ever has. So the sigmoid step alone shrinks every nudge by a factor of at least 4 (each row of $W$ is multiplied by at most $0.25$).

The chain rule using Jacobians core

Think of a pipeline: $\mathbf{x}\ \xrightarrow{\ \mathbf{F}\ }\ \mathbf{h}\ \xrightarrow{\ \mathbf{G}\ }\ \mathbf{y}$. Nudge the input a little. Function $\mathbf{F}$ turns that nudge into a nudge of $\mathbf{h}$, by multiplying with its Jacobian $J_{\mathbf{F}}$. Then $\mathbf{G}$ turns the nudge of $\mathbf{h}$ into a nudge of $\mathbf{y}$, by multiplying with $J_{\mathbf{G}}$.

So the total effect is two matrices applied one after the other: a matrix product. Like unit conversion (metres to feet, then feet to inches): the conversion factors multiply. In one variable the factors are numbers, $(g\circ f)' = g'(f(x))\,f'(x)$. With vectors the factors are matrices.

The order matters: $\mathbf{F}$ acts first, so $J_{\mathbf{F}}$ sits on the right (next to $\mathbf{x}$), and $J_{\mathbf{G}}$ on the left.

Let $\mathbf{F}:\mathbb{R}^2\to\mathbb{R}^2$, $\mathbf{F}(x, y) = [x^2y,\ 5x + y^2]^\top$ (from before) and $\mathbf{G}:\mathbb{R}^2\to\mathbb{R}^3$, $\mathbf{G}(u, v) = [u + v,\ uv,\ u^2]^\top$. Find the Jacobian of $\mathbf{G}\circ\mathbf{F}$ at $(1, 2)$.

  1. Shapes first: $\mathbf{G}\circ\mathbf{F}:\mathbb{R}^2\to\mathbb{R}^3$, so the answer is $3\times2$. And $J_{\mathbf{G}}$ is $3\times2$, $J_{\mathbf{F}}$ is $2\times2$: $(3\times2)(2\times2) = 3\times2$ ✓.
  2. Inner: $\mathbf{F}(1, 2) = [2, 9]$ and $J_{\mathbf{F}}(1, 2) = \begin{bmatrix} 4 & 1 \\ 5 & 4 \end{bmatrix}$.
  3. Outer, evaluated at the inner output $(u, v) = (2, 9)$: $J_{\mathbf{G}} = \begin{bmatrix} 1 & 1 \\ v & u \\ 2u & 0 \end{bmatrix} = \begin{bmatrix} 1 & 1 \\ 9 & 2 \\ 4 & 0 \end{bmatrix}$.
  4. Multiply: $\begin{bmatrix} 1 & 1 \\ 9 & 2 \\ 4 & 0 \end{bmatrix}\begin{bmatrix} 4 & 1 \\ 5 & 4 \end{bmatrix} = \begin{bmatrix} 4 + 5 & 1 + 4 \\ 36 + 10 & 9 + 8 \\ 16 + 0 & 4 + 0 \end{bmatrix} = \begin{bmatrix} 9 & 5 \\ 46 & 17 \\ 16 & 4 \end{bmatrix}$.

Check by nudging. With $\mathbf{h} = [0.01, -0.02]$ the prediction is $[9\cdot0.01 - 5\cdot0.02,\ 46\cdot0.01 - 17\cdot0.02,\ 16\cdot0.01 - 4\cdot0.02] = [-0.01,\ 0.12,\ 0.08]$. Directly: $\mathbf{G}(\mathbf{F}(1, 2)) = [11, 18, 4]$ and $\mathbf{G}(\mathbf{F}(1.01, 1.98))$ changes it by $[-0.0098,\ 0.1184,\ 0.0796]$ ✓.

If $\mathbf{F}:\mathbb{R}^n\to\mathbb{R}^m$ and $\mathbf{G}:\mathbb{R}^m\to\mathbb{R}^p$ are differentiable, then

$$J_{\mathbf{G}\circ\mathbf{F}}(\mathbf{x}) = J_{\mathbf{G}}\bigl(\mathbf{F}(\mathbf{x})\bigr)\; J_{\mathbf{F}}(\mathbf{x}) \qquad (p\times m)(m\times n) = p\times n.$$
  • Check shapes: the inner sizes ($m$ and $m$) must match; the outer sizes ($p$ and $n$) are the answer. If they don't match, you have the factors in the wrong order.
  • Why: near $\mathbf{x}$, $\mathbf{F}$ behaves like the linear map $\mathbf{h}\mapsto J_{\mathbf{F}}\mathbf{h}$, and near $\mathbf{F}(\mathbf{x})$, $\mathbf{G}$ behaves like $\mathbf{k}\mapsto J_{\mathbf{G}}\mathbf{k}$. Doing one then the other is the product of the two matrices (matrix product = composing linear maps).
  • One variable ($n = m = p = 1$): $g'(f(x))\cdot f'(x)$, the rule you know. Longer chains simply multiply more Jacobians: $J_3J_2J_1$.
  • Scalar output (a loss). If the last function is a scalar loss $L$ ($p = 1$), then $J_L$ is a row and $(\nabla_{\mathbf{x}}(L\circ\mathbf{F}))^\top = J_L\,J_{\mathbf{F}}$, or, transposing both sides, $\nabla_{\mathbf{x}} = J_{\mathbf{F}}^\top\,\nabla_{\mathbf{h}}L$. So going backwards, the gradient is multiplied by the transposed Jacobians. That is backpropagation (Chapter 2.9), and the full chain rule gets its own chapter (Chapter 2.8).
Why do we need it?

A network is a long chain of functions. We can only differentiate it by differentiating one piece at a time and then combining the pieces. The Jacobian chain rule is the rule for combining them.

Where is it used?

Backpropagation through every layer of every deep network, automatic differentiation libraries (PyTorch, JAX, TensorFlow), sensitivity analysis of multi-stage pipelines, and the change of variables in several steps.

How is it used?

Write each stage's Jacobian at the right point (the output of the stage before it). Check the shapes. Multiply, with the last stage on the left. In practice you multiply a gradient vector by one Jacobian at a time, never forming the giant product.

The pipeline is $\mathbf{x}\in\mathbb{R}^2\to\mathbf{F}\to\mathbf{h}\in\mathbb{R}^2\to\mathbf{G}\to\mathbf{y}\in\mathbb{R}^3$ (the functions from the example). Move $x_1, x_2$: the product $J_{\mathbf{G}}J_{\mathbf{F}}$ matches the numerical Jacobian of the whole pipeline. Now switch to wrong order: the shapes $(2\times2)(3\times2)$ do not fit, so the product does not exist. Shapes catch mistakes!

Evaluate each Jacobian at the right point. $J_{\mathbf{G}}$ is evaluated at $\mathbf{F}(\mathbf{x})$, not at $\mathbf{x}$. This is the most common slip.

Order matters. Matrix products do not commute. The function applied last has its Jacobian on the left.

Huge Jacobians are never formed in practice. A layer with 4096 inputs and 4096 outputs has a $4096\times4096$ Jacobian (16 million numbers). Backprop multiplies a vector by it instead, which is much cheaper (more in Chapter 2.9).

Quick check: $\mathbf{F}:\mathbb{R}^5\to\mathbb{R}^3$ and $\mathbf{G}:\mathbb{R}^3\to\mathbb{R}^7$. What is the shape of $J_{\mathbf{G}\circ\mathbf{F}}$, and of each factor?

$J_{\mathbf{F}}$ is $3\times5$, $J_{\mathbf{G}}$ is $7\times3$, and the product $J_{\mathbf{G}}J_{\mathbf{F}}$ is $(7\times3)(3\times5) = 7\times5$.

Recap, cheat sheet and practice

  • A scalar-valued function $\mathbb{R}^n\to\mathbb{R}$ returns one number; a vector-valued function $\mathbf{F}:\mathbb{R}^n\to\mathbb{R}^m$ returns $m$ numbers, which are $m$ scalar functions stacked. Curves, maps of the plane and neural layers are all vector functions.
  • The Jacobian $J_{\mathbf{F}}$ is the $m\times n$ matrix of all partial derivatives, with entry $(i, j) = \partial F_i/\partial x_j$: rows are outputs, columns are inputs. It is the best linear copy: $\mathbf{F}(\mathbf{x} + \mathbf{h})\approx\mathbf{F}(\mathbf{x}) + J\mathbf{h}$. For $m = n$, $\det J$ is the local area scale.
  • Special cases: scalar by vector gives the gradient (a column, $n\times1$) and its transpose is the Jacobian (a row, $1\times n$); vector by scalar gives the velocity $\mathbf{r}'(t)$ ($m\times1$); a linear map $A\mathbf{x} + \mathbf{b}$ has $J = A$.
  • Dense layer $\sigma(W\mathbf{x} + \mathbf{b})$: $J = \operatorname{diag}(\sigma'(\mathbf{z}))\,W$. Softmax: $J = \operatorname{diag}(\mathbf{s}) - \mathbf{s}\mathbf{s}^\top$ (symmetric, rows sum to 0). Polar: $\det J = r$.
  • Chain rule: $J_{\mathbf{G}\circ\mathbf{F}}(\mathbf{x}) = J_{\mathbf{G}}(\mathbf{F}(\mathbf{x}))\,J_{\mathbf{F}}(\mathbf{x})$, shapes $(p\times m)(m\times n) = p\times n$. This product of matrices is what backpropagation does (Chapter 2.9).

Cheat sheet

FunctionDerivativeShapeExample
$f:\mathbb{R}\to\mathbb{R}$$f'(x)$$1\times1$$x^2\mapsto2x$
$f:\mathbb{R}^n\to\mathbb{R}$$\nabla f$ (column), $J_f = (\nabla f)^\top$ (row)$n\times1$, $1\times n$loss $\to$ gradient
$\mathbf{r}:\mathbb{R}\to\mathbb{R}^m$$\mathbf{r}'(t)$ (velocity)$m\times1$$[\cos t, \sin t]\mapsto[-\sin t, \cos t]$
$\mathbf{F}:\mathbb{R}^n\to\mathbb{R}^m$$J_{\mathbf{F}}$, $(J)_{ij} = \partial F_i/\partial x_j$$m\times n$layer
$A\mathbf{x} + \mathbf{b}$$A$$m\times n$linear layer
$\sigma(W\mathbf{x} + \mathbf{b})$$\operatorname{diag}(\sigma'(\mathbf{z}))W$$m\times n$dense layer
softmax$(\mathbf{z})$$\operatorname{diag}(\mathbf{s}) - \mathbf{s}\mathbf{s}^\top$$k\times k$classifier output
polar $\to$ Cartesian$\begin{bmatrix}\cos\theta & -r\sin\theta \\ \sin\theta & r\cos\theta\end{bmatrix}$, $\det = r$$2\times2$area scale
$\mathbf{G}\circ\mathbf{F}$$J_{\mathbf{G}}(\mathbf{F}(\mathbf{x}))\,J_{\mathbf{F}}(\mathbf{x})$$p\times n$chain rule
Code it · NumPy

import numpy as np
np.set_printoptions(precision=4, suppress=True)

def num_jac(F, x, h=1e-6):
    """Numerical Jacobian: nudge one input at a time. Shape (m, n)."""
    x = np.asarray(x, dtype=float)
    cols = []
    for j in range(len(x)):
        e = np.zeros_like(x); e[j] = h
        cols.append((F(x + e) - F(x - e)) / (2*h))     # column j = dF/dx_j
    return np.stack(cols, axis=1)

# 1) a 2 -> 2 function and its exact Jacobian
F  = lambda x: np.array([x[0]**2 * x[1], 5*x[0] + x[1]**2])
JF = lambda x: np.array([[2*x[0]*x[1], x[0]**2],
                         [5.0,         2*x[1]]])
x = np.array([1.0, 2.0])
print(F(x))                       # [2. 9.]
print(JF(x))                      # [[4. 1.] [5. 4.]]
print(num_jac(F, x))              # [[4. 1.] [5. 4.]]   (matches)

# 2) polar -> Cartesian: the determinant is r
P  = lambda p: np.array([p[0]*np.cos(p[1]), p[0]*np.sin(p[1])])
p0 = np.array([2.0, np.pi/6])
Jp = num_jac(P, p0)
print(Jp)                         # [[ 0.866 -1.   ] [ 0.5    1.7321]]
print(np.linalg.det(Jp))          # about 2.0   (= r)

# 3) a linear map: the Jacobian is the matrix itself
A = np.array([[2., 1., 0.], [-1., 3., 4.]])
print(num_jac(lambda v: A @ v, np.array([1., 2., -1.])))   # [[ 2. 1. 0.] [-1. 3. 4.]]  = A

# 4) dense layer a = relu(W x + b):  J = diag(relu'(z)) @ W
W = np.array([[1., -1.], [2., 1.]]); b = np.array([0.5, -1.]); xx = np.array([1., 2.])
z = W @ xx + b
print(z)                          # [-0.5  3. ]
J_layer = np.diag((z > 0).astype(float)) @ W
print(J_layer)                    # [[0. 0.] [2. 1.]]   (neuron 1 is off, so its row is zero)
print(num_jac(lambda v: np.maximum(0, W @ v + b), xx))     # [[0. 0.] [2. 1.]]

# 5) softmax:  J = diag(s) - s s^T   (symmetric, every row sums to 0)
def softmax(z):
    e = np.exp(z - z.max()); return e / e.sum()
zs = np.array([2., 1., 0.]); s = softmax(zs)
Js = np.diag(s) - np.outer(s, s)
print(s)                          # [0.6652 0.2447 0.09  ]
print(Js)                         # [[ 0.2227 -0.1628 -0.0599] [-0.1628  0.1848 -0.022 ] [-0.0599 -0.022   0.0819]]
print(Js.sum(axis=1))             # [0. 0. 0.]   rows sum to zero
print(np.abs(Js - num_jac(softmax, zs)).max() < 1e-8)   # True

# 6) chain rule: J_(G o F) = J_G(F(x)) @ J_F(x)      shapes (3x2)(2x2) = 3x2
G  = lambda h: np.array([h[0] + h[1], h[0]*h[1], h[0]**2])
JG = lambda h: np.array([[1., 1.], [h[1], h[0]], [2*h[0], 0.]])
print(JG(F(x)) @ JF(x))                      # [[ 9.  5.] [46. 17.] [16.  4.]]
print(num_jac(lambda v: G(F(v)), x))         # the same numbers
Test yourself

1. $\mathbf{F}:\mathbb{R}^3\to\mathbb{R}^2$. What is the shape of its Jacobian?

Rows = outputs ($m = 2$), columns = inputs ($n = 3$): $m\times n = 2\times3$. Check: $J$ times a 3-vector must give a 2-vector.

2. The Jacobian of $\mathbf{F}(x, y) = [x + y,\ xy]^\top$ at $(2, 3)$ is…

Row 1 (for $x + y$): $[1, 1]$. Row 2 (for $xy$): $[\partial/\partial x, \partial/\partial y] = [y, x] = [3, 2]$. The third option is the transpose (rows and columns swapped).

3. $f:\mathbb{R}^4\to\mathbb{R}$. In this guide's convention, which statement is true?

The gradient is a column in the input space; the Jacobian of a scalar function has one row (one output), so it is the transpose.

4. $\mathbf{F}(\mathbf{x}) = A\mathbf{x} + \mathbf{b}$ with a fixed matrix $A$ and vector $\mathbf{b}$. Its Jacobian is…

$\partial F_i/\partial x_j = A_{ij}$, and the constant $\mathbf{b}$ has derivative 0. A linear (or affine) map is its own best linear copy.

5. $\mathbf{F}:\mathbb{R}^2\to\mathbb{R}^3$ and $\mathbf{G}:\mathbb{R}^3\to\mathbb{R}^4$. The Jacobian of $\mathbf{G}\circ\mathbf{F}$ is…

$J_{\mathbf{G}}$ is $4\times3$ and $J_{\mathbf{F}}$ is $3\times2$. The product $(4\times3)(3\times2)$ is $4\times2$, with the outer function's Jacobian on the left. $J_{\mathbf{F}}J_{\mathbf{G}}$ would be $(3\times2)(4\times3)$, which is not defined.

6. For polar-to-Cartesian coordinates, $\det J = r$. What does it tell you?

The determinant of a Jacobian is the local area scale. It is positive (orientation kept) for $r > 0$, and it is $0$ only at the origin $r = 0$, where the whole angle direction collapses to a point.

Practice problems

A. Find the Jacobian of $\mathbf{F}(x, y, z) = [xy,\ yz]^\top$ and evaluate it at $(1, 2, 3)$.

$n = 3$, $m = 2$, so $J$ is $2\times3$. Row 1 (for $xy$): $[y, x, 0]$. Row 2 (for $yz$): $[0, z, y]$. So $J = \begin{bmatrix} y & x & 0 \\ 0 & z & y \end{bmatrix}$, and at $(1, 2, 3)$: $J = \begin{bmatrix} 2 & 1 & 0 \\ 0 & 3 & 2 \end{bmatrix}$. The zeros appear where an output does not depend on an input ($xy$ has no $z$).

B. Find the Jacobian and determinant of $\mathbf{F}(x, y) = [e^x\cos y,\ e^x\sin y]^\top$ at $(0, \pi/2)$.

$J = \begin{bmatrix} e^x\cos y & -e^x\sin y \\ e^x\sin y & e^x\cos y \end{bmatrix}$ and $\det J = e^{2x}(\cos^2y + \sin^2y) = e^{2x}$. At $(0, \pi/2)$: $\cos = 0$, $\sin = 1$, $e^0 = 1$, so $J = \begin{bmatrix} 0 & -1 \\ 1 & 0 \end{bmatrix}$ (a rotation by $90^\circ$) and $\det J = 1$: areas near that point are unchanged.

C. For $f(\mathbf{x}) = \|\mathbf{x}\|^2 = x_1^2 + x_2^2 + x_3^2$, write $\nabla f$ and $J_f$ at $\mathbf{x} = [1, -2, 3]^\top$.

$\partial f/\partial x_j = 2x_j$, so $\nabla f = 2\mathbf{x} = [2, -4, 6]^\top$ ($3\times1$) and $J_f = [2, -4, 6]$ ($1\times3$).

D. A ReLU layer has $W = \begin{bmatrix} 2 & 0 \\ 1 & -1 \end{bmatrix}$ and pre-activations $\mathbf{z} = [0.3, -0.7]$. Find its Jacobian.

ReLU' is $1$ for $z>0$ and $0$ for $z<0$, so $\sigma'(\mathbf{z}) = [1, 0]$. $J = \operatorname{diag}(1, 0)\,W = \begin{bmatrix} 2 & 0 \\ 0 & 0 \end{bmatrix}$. The second output is switched off, so its row is zero.

E. $\mathbf{F}(\mathbf{x}) = [2x_1,\ x_1 + x_2]^\top$ and $L(u, v) = u^2 + v^2$. Use the chain rule to find $\nabla(L\circ\mathbf{F})$ at $(1, 1)$, and check it directly.

$\mathbf{F}(1, 1) = [2, 2]$. $J_L = [2u, 2v] = [4, 4]$ ($1\times2$). $J_{\mathbf{F}} = \begin{bmatrix} 2 & 0 \\ 1 & 1 \end{bmatrix}$. Product: $[4, 4]\begin{bmatrix} 2 & 0 \\ 1 & 1 \end{bmatrix} = [8 + 4,\ 0 + 4] = [12, 4]$, so $\nabla(L\circ\mathbf{F}) = [12, 4]^\top$. Directly, $L\circ\mathbf{F} = 4x_1^2 + (x_1 + x_2)^2$, so $\partial/\partial x_1 = 8x_1 + 2(x_1 + x_2) = 12$ and $\partial/\partial x_2 = 2(x_1 + x_2) = 4$ ✓.

F. For two classes with logits $\mathbf{z} = [0, 0]$, find the softmax Jacobian and compare with the sigmoid's slope.

$\mathbf{s} = [0.5, 0.5]$. $J = \operatorname{diag}(\mathbf{s}) - \mathbf{s}\mathbf{s}^\top = \begin{bmatrix} 0.5 - 0.25 & -0.25 \\ -0.25 & 0.5 - 0.25 \end{bmatrix} = \begin{bmatrix} 0.25 & -0.25 \\ -0.25 & 0.25 \end{bmatrix}$. Rows sum to $0$ ✓. The entry $0.25$ is exactly $\sigma'(0)$: with two classes, $s_1 = \sigma(z_1 - z_2)$, so softmax is the sigmoid in disguise.

Chapter 2.6

Matrix Calculus

Calculus when the input is a list of numbers, or a whole table of numbers. You will learn one skill: take any expression with vectors and matrices, find its derivative, and know exactly what shape the answer must have. We derive every result step by step, so nothing has to be memorised.

  • Name the kinds of derivative (scalar, vector, matrix) and read the shape of each answer
  • Use the notation: gradient $\nabla f$, Jacobian $J$, Hessian $H$, and know the two layout conventions
  • Use differentials ($d\mathbf{x}$, $dX$) and the trace trick to derive gradients without index gymnastics
  • Derive $\nabla(\mathbf{x}^\top\mathbf{x})$, $A\mathbf{x}$, $\mathbf{a}^\top\mathbf{x}$, $\mathbf{x}^\top A\mathbf{x}$, $\|A\mathbf{x}-\mathbf{b}\|^2$, $\log\mathbf{x}$ and $e^{\mathbf{x}}$, each two ways
  • Check any derivative numerically

This chapter builds on partial derivatives and gradients (2.4) and the Jacobian (2.5). It also uses matrices and the dot product from the Linear Algebra guide. That guide has a shorter overview: Matrix calculus in the Linear Algebra guide. Here we go deeper: we derive each identity from scratch, and we add differentials and the trace trick. The conventions are the same: the gradient is a column vector, and the Jacobian of a map from $\mathbb{R}^n$ to $\mathbb{R}^m$ is an $m\times n$ matrix.

The map of derivatives: what depends on what core

Picture a control panel. It has knobs (the inputs) and meters (the outputs). A derivative answers one question: "if I turn this knob a tiny bit, how much does that meter move?"

The knobs can be one number, a list of numbers (a vector), or a whole grid (a matrix). The meters can be the same. So there are nine combinations. You do not need to memorise nine different things. There is one rule behind all of them:

  • There is one slope for every (meter, knob) pair. So the number of entries in a derivative is (number of outputs) × (number of inputs).
  • The derivative is just those slopes, arranged in a sensible grid.
  1. One knob, one meter. $f(x)=x^2$. One slope: $f'(x)=2x$.
  2. Three knobs, one meter. A loss that depends on three weights. Three slopes, one per weight. We stack them into a list (a vector): the gradient.
  3. Three knobs, two meters. $2\times3=6$ slopes. We lay them in a table with one row per meter and one column per knob: the Jacobian.
  4. A matrix of knobs, one meter. A weight matrix $W$ with 6 entries and one loss. Six slopes, laid out in the same grid as $W$. This is the gradient with respect to a matrix.

Here is the whole map. A scalar is one number. $\mathbf{x}\in\mathbb{R}^n$ is a vector. $X$ is a $p\times q$ matrix.

Output ↓   Input →scalar $x$vector $\mathbf{x}$ ($n$ entries)matrix $X$ ($p\times q$)
scalar $f$$f'(x)$, a numbergradient $\nabla f$: $n\times1$ columngradient matrix $\nabla_X f$: $p\times q$
vector $\mathbf{y}$ ($m$ entries)$\mathbf{y}'(x)$: $m\times1$ column (a velocity)Jacobian $J$: $m\times n$a 3-index table $m\times p\times q$ (avoid)
matrix $Y$ ($r\times s$)$r\times s$ (derivative of each entry)3-index table (avoid)4-index table $r\times s\times p\times q$ (avoid)

Scalar derivatives (top left), vector derivatives (the middle cells) and matrix derivatives (the right column) are the three families this chapter walks through. The "avoid" cells hold a lot of numbers. We never write them out. Instead we use differentials (later in this chapter) to get the answer we need.

Why do we need it?

Models have many inputs (weights) and sometimes many outputs. We need a clear way to say "the sensitivity of everything to everything" without getting lost in indices.

Where is it used?

The gradient in gradient descent, the Jacobian of each layer in backpropagation, the gradient of a loss with respect to a weight matrix $W$ in a linear layer, and the Hessian in second-order optimisers.

How is it used?

Ask two questions: what are my inputs and outputs (number, vector, matrix)? Then count: entries = outputs × inputs. That tells you the shape of the answer before you compute a single slope.

Pick the type of the input and the type of the output. The picture shows the derivative: one block for every output entry, and each block has the shape of the input (one slope per knob). Look at how the entry count is always (outputs) × (inputs). Try vector → vector: each block is one row, so together they form the $m\times n$ Jacobian.

Quick check: a layer maps 4 inputs to 3 outputs. How many slopes are in its derivative, and what is its shape?

$3\times4=12$ slopes. It is a vector-to-vector map, so the derivative is the Jacobian, a $3\times4$ matrix (one row per output, one column per input).

Scalar derivatives: the "nudge" picture

You already know that $f'(x)$ is the slope of the curve. Here is the same idea in the form we will use all chapter: nudge and respond.

Move the input by a tiny amount, called $dx$. The output moves by a tiny amount, called $df$. The slope $f'(x)$ is the conversion rate between them: $df \approx f'(x)\,dx$. For matrices, we will keep exactly this form, only $dx$ will become a vector or a matrix.

Let $f(x)=x^2$ at $x=3$. Then $f(3)=9$ and $f'(3)=6$.

  1. Nudge by $dx=0.1$: the new value is $3.1^2=9.61$. The true change is $0.61$. The prediction $f'(3)\,dx = 6\times0.1=0.6$.
  2. Nudge by $dx=0.01$: true change $3.01^2-9=0.0601$. Prediction $0.06$.
  3. Nudge by $dx=0.001$: true change $0.006001$. Prediction $0.006$.

The prediction gets better as the nudge gets smaller. The leftover error here is exactly $dx^2$ (it is $0.01$, $0.0001$, $0.000001$). It shrinks much faster than the nudge itself.

The derivative is $f'(x)=\dfrac{df}{dx}=\displaystyle\lim_{dx\to0}\frac{f(x+dx)-f(x)}{dx}$. The differential form says the same thing:

$$df = f'(x)\,dx.$$

Think of $dx$ as "a nudge so small that squares of it ($dx^2$) can be ignored". The rules you know become rules about differentials:

  • Sum: $d(u+v)=du+dv$. Constant factor: $d(cu)=c\,du$.
  • Product: $d(uv)=u\,dv+v\,du$.
  • Chain: if $y=g(f)$ then $dy=g'(f)\,df$. For example $y=\sin(x^2)$: $dy=\cos(x^2)\cdot 2x\,dx$.
Why do we need it?

The "nudge and respond" form $df = f'(x)\,dx$ extends to vectors and matrices without changing shape. Slopes alone do not, but nudges do.

Where is it used?

Every first-order method: a gradient descent step is "choose $dx$ so that $df$ is negative". It is also the idea behind sensitivity analysis and error propagation.

How is it used?

To differentiate, write the differential of each piece with the rules above, then collect everything that multiplies $dx$. That coefficient is the derivative.

Drag the blue dot along the curve. Slide $dx$. The green bar is the true change $\Delta f$; the orange bar is the prediction $f'(x)\,dx$ from the tangent line; the red gap is the error. Make $dx$ ten times smaller and watch the error shrink about a hundred times (it behaves like $dx^2$).

$df$ and $dx$ are small nudges, not the final derivative. The derivative is the ratio $df/dx$. A tangent-line prediction is only trustworthy when $dx$ is small.

Quick check: $f(x)=x^3$ at $x=2$. Predict $df$ for $dx=0.01$, then compare with the truth.

$f'(2)=3\cdot2^2=12$, so the prediction is $12\times0.01=0.12$. The truth is $2.01^3-8=8.120601-8=0.120601$. The error ($0.000601$) is tiny.

Vector derivatives: gradient and velocity core

Now let the input be a vector: several knobs at once. Nudge knob 1 by $dx_1$ and knob 2 by $dx_2$. Each knob adds its own share to the change: (slope of knob 1) × $dx_1$ plus (slope of knob 2) × $dx_2$. That is a dot product:

$df \approx (\text{list of slopes})\cdot(\text{list of nudges})$.

The list of slopes is the gradient. The other direction also exists: one knob (say time $t$) and a vector of meters, like the position $(x(t),y(t))$ of a moving dot. Its derivative is the list of speeds: the velocity.

$f(x_1,x_2)=x_1^2+3x_1x_2$ at $(1,2)$. The value is $f=1+6=7$.

  1. $\partial f/\partial x_1 = 2x_1+3x_2 = 2+6=8$ and $\partial f/\partial x_2 = 3x_1=3$. So $\nabla f=[8,\,3]^\top$.
  2. Nudge by $d\mathbf{x}=[0.01,\,-0.02]^\top$. Prediction: $\nabla f^\top d\mathbf{x}=8(0.01)+3(-0.02)=0.08-0.06=0.02$.
  3. Truth: $f(1.01,\,1.98)=1.0201+3(1.01)(1.98)=1.0201+5.9994=7.0195$. So the true change is $0.0195$. ✓ Close to $0.02$.

Velocity. A point moves on a circle: $\mathbf{r}(t)=[\cos t,\,\sin t]^\top$. Differentiate each entry: $\mathbf{r}'(t)=[-\sin t,\,\cos t]^\top$. At $t=0$ this is $[0,1]^\top$: the dot starts at $(1,0)$ and moves straight up.

Gradient (scalar output, vector input $\mathbf{x}\in\mathbb{R}^n$). It is a column vector with the shape of $\mathbf{x}$:

$$\nabla f(\mathbf{x})=\begin{bmatrix}\partial f/\partial x_1\\ \vdots\\ \partial f/\partial x_n\end{bmatrix},\qquad df=\nabla f^\top d\mathbf{x}=\sum_{i=1}^n\frac{\partial f}{\partial x_i}\,dx_i.$$

Derivative of a vector with respect to a number (vector output $\mathbf{y}\in\mathbb{R}^m$, scalar input $t$): the column of ordinary derivatives,

$$\frac{d\mathbf{y}}{dt}=\begin{bmatrix}dy_1/dt\\ \vdots\\ dy_m/dt\end{bmatrix},\qquad d\mathbf{y}=\frac{d\mathbf{y}}{dt}\,dt.$$

The gradient lives in the same space as the input. That is why "move opposite to the gradient" makes sense: $\mathbf{x}-\eta\nabla f$ is a vector of the same shape as $\mathbf{x}$.

Why do we need it?

A model has thousands or millions of weights. One list of slopes, one per weight, tells us how to change all of them at once.

Where is it used?

Gradient descent, SGD and Adam for every trained model, gradient clipping, saliency maps (which pixels matter), and velocities in physics simulations and ODE solvers.

How is it used?

Compute the slope for each input, stack them in a column, and update $\mathbf{x}\leftarrow\mathbf{x}-\eta\nabla f$. To predict a small change use $df=\nabla f^\top d\mathbf{x}$.

The shading shows $f$ (darker blue = higher) and the thin lines are level curves. Drag the blue dot to move the point. Drag the orange dot to choose the nudge $d\mathbf{x}$. The readout compares the true change with $\nabla f^\top d\mathbf{x}$. Make the nudge small and they agree. Nudge along a level curve and $df\approx0$. Nudge along the gradient and $df$ is biggest.

Slide $t$. The dot is at $\mathbf{r}(t)=[\cos t,\ \sin 2t]^\top$ and the orange arrow is the derivative $\mathbf{r}'(t)=[-\sin t,\ 2\cos 2t]^\top$. The arrow always touches the path and points the way the dot is moving. Where is the dot momentarily slowest?

Some books write the gradient as a row. In this guide it is always a column, the same shape as $\mathbf{x}$. If a formula you find online looks "transposed", this is usually why (see the section on layout below).

Quick check: $f(\mathbf{x})=x_1x_2$ at $(3,5)$. What is $\nabla f$, and what does $d\mathbf{x}=[0.1,0]^\top$ do to $f$?

$\nabla f=[x_2,\,x_1]^\top=[5,3]^\top$. The change is about $\nabla f^\top d\mathbf{x}=5(0.1)+3(0)=0.5$. (True: $3.1\cdot5-15=0.5$.)

Derivatives of vector functions: the Jacobian core

A layer of a neural network takes a vector in and gives a vector out. Every output has its own gradient (its own list of slopes). Stack those lists as rows and you get a table: the Jacobian. Row $i$ answers "how does output $i$ respond to each input?". Column $j$ answers "what does turning input $j$ do to all the outputs?".

The key fact: for a small nudge, the output nudge is the Jacobian times the input nudge. The Jacobian is the best straight-line (linear) description of the function near the point. (Chapter 2.5 develops this picture. Here we focus on notation and shapes.)

$F(x_1,x_2,x_3)=(x_1x_2,\ \ x_2+x_3^2)$. Three inputs, two outputs, so $J$ is $2\times3$.

  1. Gradient of output 1: $[x_2,\ x_1,\ 0]$. Gradient of output 2: $[0,\ 1,\ 2x_3]$.
  2. Stack as rows: $J=\begin{bmatrix}x_2&x_1&0\\0&1&2x_3\end{bmatrix}$.
  3. At $\mathbf{x}=(1,2,3)$: $F=(2,\,11)$ and $J=\begin{bmatrix}2&1&0\\0&1&6\end{bmatrix}$.
  4. Nudge $d\mathbf{x}=[0.1,\,0,\,-0.1]^\top$. Then $J\,d\mathbf{x}=[0.2,\ -0.6]^\top$.
  5. Truth: $F(1.1,\,2,\,2.9)=(2.2,\ 10.41)$, so the true change is $[0.2,\ -0.59]^\top$. ✓

For $F:\mathbb{R}^n\to\mathbb{R}^m$ with outputs $F_1,\dots,F_m$, the Jacobian is the $m\times n$ matrix

$$J_F(\mathbf{x})=\begin{bmatrix}\partial F_1/\partial x_1&\cdots&\partial F_1/\partial x_n\\ \vdots&&\vdots\\ \partial F_m/\partial x_1&\cdots&\partial F_m/\partial x_n\end{bmatrix},\quad (J_F)_{ij}=\frac{\partial F_i}{\partial x_j},\quad d\mathbf{y}=J_F\,d\mathbf{x}.$$
  • Row $i$ is $(\nabla F_i)^\top$. When $m=1$, $J_f=(\nabla f)^\top$ (a single row).
  • A linear map $F(\mathbf{x})=A\mathbf{x}$ has $J=A$ everywhere.
  • Derivative of a vector function = a derivative of a vector with respect to a vector. This is why the Jacobian is also called "the derivative" of a vector function.
Why do we need it?

Layers output vectors. To pass "how sensitive is the loss?" backwards through a layer we need one object that says how every output reacts to every input.

Where is it used?

Backpropagation through every layer (Chapter 2.8 and 2.9), the Jacobian determinant in normalising flows, Gauss–Newton and Levenberg–Marquardt curve fitting, and robot-arm kinematics.

How is it used?

Compute (or let autograd compute) $J$, check that its shape is outputs × inputs, and use $d\mathbf{y}\approx J\,d\mathbf{x}$ to predict small changes.

Choose a function. The table shows the Jacobian from the formula and the Jacobian found by nudging each input up and down. They match. Edit $\mathbf{x}$ (and $A$ for the linear map). Look at the pattern of zeros: the elementwise functions give a diagonal matrix (each output depends on one input), softmax gives a full matrix (every output depends on every input).

Quick check: $F(\mathbf{x})=A\mathbf{x}+\mathbf{b}$ with $A$ a $3\times2$ matrix. What is $J$ and what is its shape?

$J=A$, shape $3\times2$. The constant $\mathbf{b}$ shifts the output but does not change how it reacts to a nudge. (This is the "affine" case: $d\mathbf{y}=A\,d\mathbf{x}$ exactly.)

Matrix derivatives: matrices as inputs and outputs

A layer of a neural network stores its weights in a matrix $W$. The loss $L$ is one number. Every single entry $W_{ij}$ is a knob, so there is one slope $\partial L/\partial W_{ij}$ per entry. We arrange those slopes in the same grid as $W$. The result is the gradient matrix: a "map" showing which weights the loss cares about most.

What if the output is also a matrix, like $Y=X^2$ or $Y=X^{-1}$? Then each of the output entries has its own gradient matrix. That is a four-index table. It is huge and awkward, so in practice we never write it. We use a differential instead (next sections).

  1. Sum of entries. $f(X)=\sum_{i,j}X_{ij}$. Every entry has slope 1, so $\nabla_Xf$ is a matrix of ones.
  2. A custom function. $f(X)=X_{11}^2+2X_{12}+3X_{21}X_{22}$. Slopes: $\partial f/\partial X_{11}=2X_{11}$, $\partial f/\partial X_{12}=2$, $\partial f/\partial X_{21}=3X_{22}$, $\partial f/\partial X_{22}=3X_{21}$. At $X=\begin{bmatrix}1&0\\2&1\end{bmatrix}$: $\nabla_Xf=\begin{bmatrix}2&2\\3&6\end{bmatrix}$.
  3. Sum of squares. $f(X)=\sum X_{ij}^2$. Each slope is $2X_{ij}$, so $\nabla_Xf=2X$.
  4. A matrix-valued function. $Y=X^2$ for a $2\times2$ matrix $X=\begin{bmatrix}a&b\\c&d\end{bmatrix}$ gives $Y=\begin{bmatrix}a^2+bc&ab+bd\\ac+cd&bc+d^2\end{bmatrix}$. Four output entries, each with a $2\times2$ gradient matrix: for example $\nabla_X Y_{11}=\begin{bmatrix}2a&c\\b&0\end{bmatrix}$. That is $4\times4=16$ numbers.

For a scalar function $f$ of a $p\times q$ matrix $X$, the gradient matrix $\nabla_Xf$ has the same shape as $X$:

$$(\nabla_Xf)_{ij}=\frac{\partial f}{\partial X_{ij}},\qquad df=\sum_{i,j}\frac{\partial f}{\partial X_{ij}}\,dX_{ij}=\operatorname{tr}\!\big((\nabla_Xf)^\top dX\big).$$

The last form uses the trace (explained below). It is the matrix version of $df=\nabla f^\top d\mathbf{x}$.

Derivative of a matrix function $Y(X)$: one gradient matrix per entry of $Y$. Two ways to tame it: (1) work with the differential $dY$, which has the same shape as $Y$ and is easy to compute; (2) vectorise: stack the entries of $X$ into one long vector, and then it is an ordinary Jacobian (awareness: this is where the Kronecker product appears).

Why do we need it?

Weights live in matrices. We need the slope of the loss for each weight, laid out so that the update $W\leftarrow W-\eta\,\nabla_WL$ makes sense entry by entry.

Where is it used?

Every linear layer, convolution kernel and attention projection matrix in a neural network; covariance estimation, matrix factorisation and PCA objectives.

How is it used?

Get the gradient matrix with the same shape as the weight matrix, then subtract a small multiple of it. In code, W.grad in PyTorch always has the shape of W.

Choose a function of the $3\times3$ matrix $X$. The table shows the gradient matrix from the formula and from the nudge-each-entry check. Edit the entries of $X$ and $A$. Notice that the answer is always $3\times3$, and that for $\operatorname{tr}(AX)$ the answer is $A^\top$ (not $A$): you will see why in the trace-trick section.

Quick check: $X$ is $4\times5$ and $f(X)=\sum X_{ij}^2$. What is the shape of $\nabla_Xf$, and what is it?

It is $4\times5$ (the shape of $X$), and it equals $2X$.

Notation: gradient, Jacobian, Hessian, and the layout question core

There are three names to keep straight, and they are three sizes of the same idea:

  • Gradient $\nabla f$: the slopes of one output with respect to all inputs.
  • Jacobian $J$: the slopes of all outputs with respect to all inputs (a stack of gradients).
  • Hessian $H$: the slope of the slope: how each slope changes as each input changes. It measures curvature.

There is also an annoying fact of life: different books arrange these numbers differently (rows or columns). This is the main reason why formulas from two sources can look "transposed" from each other. We will be honest about it and then pick one convention.

$f(x_1,x_2)=x_1^2+3x_1x_2$ again.

  1. Gradient: $\nabla f=\begin{bmatrix}2x_1+3x_2\\3x_1\end{bmatrix}$.
  2. Hessian: differentiate each entry of the gradient with respect to $x_1$ and $x_2$. The first entry gives $[2,\ 3]$. The second gives $[3,\ 0]$. So $H=\begin{bmatrix}2&3\\3&0\end{bmatrix}$.
  3. $H$ is symmetric: $\partial^2f/\partial x_1\partial x_2=\partial^2f/\partial x_2\partial x_1=3$. This always holds when the second derivatives are continuous.

Notation used in this guide.

  • Gradient of $f:\mathbb{R}^n\to\mathbb{R}$: $\nabla f$ (or $\nabla_{\mathbf{x}}f$), an $n\times1$ column.
  • Jacobian of $F:\mathbb{R}^n\to\mathbb{R}^m$: $J_F$ (also written $\partial F/\partial\mathbf{x}^\top$ or $DF$), an $m\times n$ matrix with $(J_F)_{ij}=\partial F_i/\partial x_j$.
  • Hessian of $f$: $H=\nabla^2f$, the $n\times n$ matrix with $H_{ij}=\dfrac{\partial^2f}{\partial x_i\,\partial x_j}$. It is the Jacobian of the gradient: $H=J_{\nabla f}$.
  • Differentials: $df=\nabla f^\top d\mathbf{x}$, $\ d\mathbf{y}=J\,d\mathbf{x}$, $\ df=\operatorname{tr}(G^\top dX)\Rightarrow G=\nabla_Xf$.

The layout question. There are two conventions for arranging "$\partial\mathbf{y}/\partial\mathbf{x}$":

Numerator layoutDenominator layout
Ruleshape = (size of $\mathbf{y}$) × (size of $\mathbf{x}$)shape = (size of $\mathbf{x}$) × (size of $\mathbf{y}$)
$\partial f/\partial\mathbf{x}$ for scalar $f$a row ($1\times n$)a column ($n\times1$)
Jacobian of $F:\mathbb{R}^n\to\mathbb{R}^m$$m\times n$$n\times m$ (the transpose)
Chain rule order$J_{f\circ g}=J_f\,J_g$ (natural order)order reversed

This guide uses a mixed convention, the most common one in machine learning: the Jacobian is $m\times n$ (numerator layout), but the gradient of a scalar is written as a column of the same shape as the input. They fit together through $\nabla f=J_f^\top$ when $m=1$. The chain rule for gradients then reads $\nabla(f\circ g)=J_g^\top\,\nabla f$. The gradient with respect to a matrix $X$ always has the shape of $X$. We avoid writing $\partial f/\partial\mathbf{x}$ for vectors, to avoid any doubt. The Linear Algebra guide uses the same rules.

Why do we need it?

Without agreed notation, a formula like "$A^\top(A\mathbf{x}-\mathbf{b})$ or $(A\mathbf{x}-\mathbf{b})^\top A$?" is a coin toss. Shapes decide it, but only if you know the convention.

Where is it used?

Optimisation papers and textbooks (each picks a layout), deep-learning libraries (autograd returns gradients with the shape of the parameter), and second-order methods that use the Hessian (Newton's method, L-BFGS, natural gradient).

How is it used?

State the convention once, then check the shape of every formula. If a source looks transposed from yours, check whether it uses the other layout, and transpose the answer.

Set the number of inputs $n$ and outputs $m$. Switch between our guide's convention, pure numerator layout and pure denominator layout. The blocks show the shape of each derivative in that layout. Notice that only the Jacobian's orientation and the gradient (row or column) change. The Hessian is the same in all of them, and so is the gradient matrix of a weight matrix in ours and in denominator layout.

When you read another book, look at its very first example. If its gradient of a scalar is a row, it uses numerator layout throughout. Do not mix formulas from two layouts without transposing. The Hessian of a smooth function is symmetric, so it never changes.

Quick check: in this guide, $F:\mathbb{R}^5\to\mathbb{R}^3$. What is the shape of $J_F$, and what is the shape of $J_F^\top$?

$J_F$ is $3\times5$ (outputs × inputs). $J_F^\top$ is $5\times3$: it is what multiplies a $3\times1$ output-side gradient to give a $5\times1$ input-side gradient in backpropagation.

Reading the shape of a derivative core

Shapes are your best friend in matrix calculus. Before computing anything, you can usually tell what the answer must look like. And after computing it, a wrong shape instantly tells you that you made a mistake. This is the cheapest bug-check there is.

Two simple rules cover almost everything:

  • Scalar output: the gradient has exactly the shape of the input. (Number in, number out. Vector in, vector. Matrix in, matrix.)
  • Everything else: output shape first, then input shape. A vector-to-vector map gives (output length) × (input length).
  1. $f(\mathbf{x})=\|A\mathbf{x}-\mathbf{b}\|^2$ with $\mathbf{x}\in\mathbb{R}^4$: scalar output, so the gradient has shape $4\times1$.
  2. Check the formula $2A^\top(A\mathbf{x}-\mathbf{b})$ with $A$ of size $3\times4$: $A^\top$ is $4\times3$ and $(A\mathbf{x}-\mathbf{b})$ is $3\times1$. The product is $(4\times3)(3\times1)=4\times1$. ✓ The shape fits.
  3. Try the wrong formula $2A(A\mathbf{x}-\mathbf{b})$: $(3\times4)(3\times1)$ cannot be multiplied. ✗ We caught the mistake without computing a thing.
  4. A layer $F:\mathbb{R}^4\to\mathbb{R}^3$: Jacobian is $3\times4$. A weight matrix $W$ ($3\times4$) and a scalar loss: $\nabla_WL$ is $3\times4$.

For input shape $S_{in}$ and output shape $S_{out}$, the derivative has shape:

CaseShape of the derivative
scalar output, any input$S_{in}$ (the gradient has the input's shape)
vector output ($m$), scalar input$m\times1$
vector output ($m$), vector input ($n$)$m\times n$ (Jacobian)
any other$S_{out}$ followed by $S_{in}$ (a higher-index table)

Counting check: the number of entries is always (size of the output) × (size of the input). A shape check of a formula means: write the shape of each factor and make sure neighbouring inner sizes match, as in matrix multiplication.

Why do we need it?

Matrix formulas are long and easy to garble. A shape check is a free proof-reader: wrong formulas often refuse to multiply.

Where is it used?

Every time you write a backward pass by hand, implement a loss in NumPy or PyTorch, or debug a "size mismatch" error. Gradients with the wrong shape are one of the most common bugs in deep-learning code.

How is it used?

Write the shape under each symbol, from right to left, and check that each product is allowed. The final shape must equal the shape the table above predicts.

Choose the type of the input and the output, then set their sizes. The calculator tells you the shape of the derivative and how many numbers it holds. Try scalar-valued $f$ of a $3\times4$ matrix (the answer is $3\times4$), then vector-to-vector with $n=4$, $m=3$ (the Jacobian, $3\times4$), then matrix-to-matrix (a big table).

Pick a formula and set the sizes $m$ and $n$ of $A$ (it is $m\times n$, and $\mathbf{x}$ has $n$ entries). The widget multiplies the shapes step by step and marks each product. Notice that the wrong formulas only work when $m=n$, which is exactly why a square test matrix can hide a bug. Always test with $m\ne n$.

A shape check is necessary, not sufficient. $A^\top(A\mathbf{x}-\mathbf{b})$ and $A^\top(\mathbf{b}-A\mathbf{x})$ both fit, but only one has the right sign. Use shapes to catch the wrong arrangement, then use a numerical check (later) for the rest.

Quick check: $W$ is $5\times8$, $\mathbf{x}$ has 8 entries, $L=\mathbf{c}^\top(W\mathbf{x})$ with $\mathbf{c}\in\mathbb{R}^5$. What shape is $\nabla_WL$?

The same as $W$: $5\times8$. (We will derive $\nabla_WL=\mathbf{c}\,\mathbf{x}^\top$, which is $(5\times1)(1\times8)=5\times8$. ✓)

Matrix differential notation: nudge the whole matrix core

Taking a derivative "with respect to a matrix" sounds scary. The trick is to stop asking for the derivative and ask for the nudge response instead. Write $dX$ for "a tiny nudge of the whole matrix": a matrix of the same shape as $X$ whose entries are tiny nudges $dX_{ij}$.

Then work out how everything else responds to that nudge, using the familiar product rule. At the very end, read the gradient off the response. No indices, no big tables.

One warning: matrices do not commute ($AB\neq BA$ in general). So the order of factors in a differential rule matters, even though in one variable it never did.

Take $X=\begin{bmatrix}1&2\\3&4\end{bmatrix}$ and $Y=\begin{bmatrix}0&1\\1&0\end{bmatrix}$. Nudge $X$ by $dX=\begin{bmatrix}0.1&0\\0&0\end{bmatrix}$ and $Y$ by $dY=\begin{bmatrix}0.1&0\\0&0\end{bmatrix}$.

  1. Original product: $XY=\begin{bmatrix}2&1\\4&3\end{bmatrix}$.
  2. New product: $(X+dX)(Y+dY)=\begin{bmatrix}1.1&2\\3&4\end{bmatrix}\begin{bmatrix}0.1&1\\1&0\end{bmatrix}=\begin{bmatrix}2.11&1.1\\4.3&3\end{bmatrix}$. True change: $\begin{bmatrix}0.11&0.1\\0.3&0\end{bmatrix}$.
  3. Rule: $dX\,Y=\begin{bmatrix}0&0.1\\0&0\end{bmatrix}$ and $X\,dY=\begin{bmatrix}0.1&0\\0.3&0\end{bmatrix}$. Sum: $\begin{bmatrix}0.1&0.1\\0.3&0\end{bmatrix}$.
  4. Compare: true change $0.11$ against predicted $0.1$ in the top-left. The leftover $0.01$ is $dX\,dY$, a product of two tiny nudges (second order), which the rule ignores.

Rules for differentials ($A$ is constant; $X,Y$ vary; $c$ is a number):

RuleWhy it holds
$d(A)=0$, $\ d(cX)=c\,dX$, $\ d(X+Y)=dX+dY$constants do not move; nudging is linear
$d(XY)=dX\,Y+X\,dY$entry $(i,j)$ is $\sum_kX_{ik}Y_{kj}$. Apply the ordinary product rule to each term: $\sum_k(dX_{ik}Y_{kj}+X_{ik}dY_{kj})$, which is entry $(i,j)$ of $dX\,Y+X\,dY$
$d(X^\top)=(dX)^\top$transposing just swaps positions of entries; nudges swap with them
$d\operatorname{tr}(X)=\operatorname{tr}(dX)$$\operatorname{tr}X=\sum_iX_{ii}$ is a sum, and nudging a sum nudges each term
$d(X^{-1})=-X^{-1}\,dX\,X^{-1}$from $XX^{-1}=I$: nudge both sides, $dX\,X^{-1}+X\,d(X^{-1})=0$. Solve for $d(X^{-1})$ by multiplying on the left by $X^{-1}$
$d(X\circ Y)=dX\circ Y+X\circ dY$, $\ d\,f(X)=f'(X)\circ dX$the same product and chain rule, applied entry by entry ($\circ$ = entrywise product)

Consequence: $d(X^2)=dX\,X+X\,dX$, which is not $2X\,dX$ unless $X$ and $dX$ commute. In one variable, $(1/x)'=-1/x^2$ is the scalar version of $d(X^{-1})$.

Why do we need it?

It lets us differentiate expressions made of matrix products, inverses and traces without ever writing out an entry. Long index sums disappear.

Where is it used?

Deriving the gradients of least squares, ridge regression, Gaussian log-likelihoods (with $\Sigma^{-1}$), PCA objectives, and the backward pass of a layer written in matrix form.

How is it used?

Write $d(\text{your expression})$ using the rules, working from the outside in. Keep every factor in its original order. Then use the trace trick (next section) to read the gradient off the result.

Choose a rule. The widget nudges $X$ and $Y$ by a tiny multiple $\varepsilon$ of fixed directions $D_X$, $D_Y$, then compares the true change divided by $\varepsilon$ with the rule's prediction. Make $\varepsilon$ smaller and the difference vanishes. Then switch on use the wrong rule: for $d(X^2)$ the "$2X\,dX$" version never matches, because $X$ and $dX$ do not commute.

  • Keep the order: $d(XY)=dX\,Y+X\,dY$, never $Y\,dX+X\,dY$.
  • $d(X^{-1})=-X^{-1}\,dX\,X^{-1}$ has the inverse on both sides.
  • $dX$ has the same shape as $X$. If a differential has a different shape from the thing it nudges, something is wrong.
Quick check: what is $d(AXB)$ for constant matrices $A$ and $B$?

$A$ and $B$ do not move, so the product rule gives $d(AXB)=A\,dX\,B$. (Check sizes: if $X$ is $p\times q$ then $A$ is $r\times p$, $B$ is $q\times s$, and the result is $r\times s$.)

The trace trick core

The trace of a square matrix is the sum of its diagonal entries. Why would calculus care? Because of one magic property: you can rotate matrices inside a trace: $\operatorname{tr}(AB)=\operatorname{tr}(BA)$. Even when $AB$ and $BA$ are completely different matrices (even different sizes), their traces are equal.

This is useful because a differential response like $df$ is a single number, and we want to bring the nudge $dX$ to the very end of the expression, where we can read off whatever stands in front. Rotating inside a trace does exactly that.

$A=\begin{bmatrix}1&2&0\\0&1&3\end{bmatrix}$ ($2\times3$) and $B=\begin{bmatrix}1&0\\2&1\\0&1\end{bmatrix}$ ($3\times2$).

  1. $AB=\begin{bmatrix}5&2\\2&4\end{bmatrix}$ ($2\times2$). Its trace is $5+4=9$.
  2. $BA=\begin{bmatrix}1&2&0\\2&5&3\\0&1&3\end{bmatrix}$ ($3\times3$). Its trace is $1+5+3=9$.
  3. Different matrices, different sizes, the same trace.

$\operatorname{tr}(A)=\sum_iA_{ii}$ for a square matrix $A$. Properties:

  • Linear: $\operatorname{tr}(A+B)=\operatorname{tr}A+\operatorname{tr}B$, $\ \operatorname{tr}(cA)=c\operatorname{tr}A$. Also $\operatorname{tr}(A^\top)=\operatorname{tr}(A)$.
  • Cyclic: $\operatorname{tr}(AB)=\operatorname{tr}(BA)$ whenever both products exist. For three: $\operatorname{tr}(ABC)=\operatorname{tr}(BCA)=\operatorname{tr}(CAB)$. But in general $\operatorname{tr}(ACB)\ne\operatorname{tr}(ABC)$: you may rotate, not swap.
  • A scalar is its own trace: a $1\times1$ matrix $a$ has $\operatorname{tr}(a)=a$. So any scalar expression can be wrapped in $\operatorname{tr}$ and then rotated.
  • Dot product of matrices: $\operatorname{tr}(A^\top B)=\sum_{i,j}A_{ij}B_{ij}$.

Proof of $\operatorname{tr}(AB)=\operatorname{tr}(BA)$: $\operatorname{tr}(AB)=\sum_i\sum_kA_{ik}B_{ki}=\sum_k\sum_iB_{ki}A_{ik}=\operatorname{tr}(BA)$. We only swapped the order of two sums and two numbers.

Why it gives a gradient. By the last property, $\operatorname{tr}(G^\top dX)=\sum_{i,j}G_{ij}\,dX_{ij}$. Compare with $df=\sum_{ij}\frac{\partial f}{\partial X_{ij}}dX_{ij}$. If $df=\operatorname{tr}(G^\top dX)$ for every nudge, then $G_{ij}=\partial f/\partial X_{ij}$, i.e. $G=\nabla_Xf$.

Why do we need it?

It turns a matrix expression into a form where $dX$ stands alone at the end, so the gradient can be read off. It replaces pages of index algebra by two moves: rotate, then read.

Where is it used?

Deriving gradients for linear and ridge regression, matrix factorisation, Gaussian log-likelihoods with covariance matrices, and the backward pass of linear layers. Also in statistics: for a random vector $\mathbf{x}$ with mean zero and covariance $\Sigma$, the cyclic trick gives $\mathbb{E}[\mathbf{x}^\top A\mathbf{x}]=\operatorname{tr}(A\Sigma)$.

How is it used?

Wrap the scalar in $\operatorname{tr}$. Use the cyclic property to move $dX$ to the right end: $\operatorname{tr}(M\,dX)$. Then $G^\top=M$, so the gradient is $G=M^\top$.

Edit the entries. In two matrices mode, $A$ is $2\times3$ and $B$ is $3\times2$: the products $AB$ ($2\times2$) and $BA$ ($3\times3$) differ in size, yet their traces agree (diagonal entries are highlighted). In three matrices mode, rotating $ABC\to BCA\to CAB$ never changes the trace, but swapping to $ACB$ usually does.

You may rotate the factors of a trace, not shuffle them. $\operatorname{tr}(ABC)=\operatorname{tr}(CAB)$ but usually $\ne\operatorname{tr}(ACB)$. Also, "scalar equals its trace" applies only to $1\times1$ things: $\mathbf{x}^\top A\mathbf{x}$ is a scalar, but $A\mathbf{x}$ is not.

Quick check: rewrite the scalar $\mathbf{a}^\top X\mathbf{b}$ so that $X$ is at the end, inside a trace.

$\mathbf{a}^\top X\mathbf{b}=\operatorname{tr}(\mathbf{a}^\top X\mathbf{b})=\operatorname{tr}(\mathbf{b}\,\mathbf{a}^\top X)$ (rotate $\mathbf{b}$ to the front). Here $X$ is at the right end.

The recipe: differential, rearrange, read off core

Now put the two tools together. Ask the function: "if I nudge the input by $dX$, how much does $f$ change?" The answer is always a number built from $dX$. Rewrite it in the form "(something) multiplying $dX$", and that something is the gradient. It is like reading the price list off a bill: the bill says "3 apples at $a$ each plus 2 pears at $p$ each", and the multipliers tell you the quantities.

$f(X)=\operatorname{tr}(AX)$ for a $2\times2$ matrix $A$. We want $\nabla_Xf$.

  1. Nudge: $df=\operatorname{tr}(A\,dX)$ ($A$ is constant, trace is linear).
  2. We need the form $\operatorname{tr}(G^\top dX)$. We already have $\operatorname{tr}(\underbrace{A}_{G^\top}\,dX)$.
  3. So $G^\top=A$, and therefore $\nabla_Xf=A^\top$.

Check with entries: $\operatorname{tr}(AX)=\sum_{i,j}A_{ij}X_{ji}$. The coefficient of $X_{ji}$ is $A_{ij}$, so the gradient at position $(j,i)$ is $A_{ij}$: that is $A^\top$. ✓

The recipe.

  1. Differential. Write $df$ using the rules (constants do not move; product rule in order).
  2. Rearrange. If $df$ is a scalar, wrap it in $\operatorname{tr}$ and use cyclic rotations (and transposes of scalars) so that $dX$ is at the right end, with nothing to its right.
  3. Read off. The three standard forms:
$$df=\mathbf{g}^\top d\mathbf{x}\ \Rightarrow\ \nabla f=\mathbf{g},\qquad d\mathbf{y}=M\,d\mathbf{x}\ \Rightarrow\ J=M,\qquad df=\operatorname{tr}(M\,dX)\ \Rightarrow\ \nabla_Xf=M^\top.$$

Why the answer is unique. Suppose $\operatorname{tr}(M^\top dX)=\operatorname{tr}(G^\top dX)$ for every $dX$. Choose $dX$ to have a 1 in position $(i,j)$ and zeros elsewhere. The left side is $M_{ij}$ and the right is $G_{ij}$. So $M=G$.

Why do we need it?

It is one procedure that works for every function built from products, transposes, traces and inverses. You derive instead of looking up.

Where is it used?

Deriving custom loss gradients, checking autograd results by hand, writing backward passes for new layers, and the derivations in papers on Gaussian processes, PCA and matrix completion.

How is it used?

Three lines on paper: write $df$, rotate to put $dX$ last, read off the multiplier (transposing for a matrix gradient). Then run a numerical check and compare shapes.

Pick a function and press Next to reveal one line at a time. Each line follows from the one above. Try to predict the next line before you press the button. Then check that the final shape matches the input.

  • The recipe needs $df$ to be a scalar before you wrap it in a trace. If your function is vector-valued, use $d\mathbf{y}=M\,d\mathbf{x}\Rightarrow J=M$ instead.
  • Do not forget the final transpose for matrix gradients: $\operatorname{tr}(M\,dX)$ gives $M^\top$, not $M$.
Quick check: use the recipe to find $\nabla_X\operatorname{tr}(X^\top B)$.

$df=\operatorname{tr}(dX^\top B)$. A trace equals the trace of its transpose, so $\operatorname{tr}(dX^\top B)=\operatorname{tr}(B^\top dX)$. Then $G^\top=B^\top$, so $\nabla_Xf=B$. (This is the "dot product of matrices" $\sum X_{ij}B_{ij}$, whose slope for $X_{ij}$ is clearly $B_{ij}$.)

Important derivative 1: $\nabla(\mathbf{x}^\top\mathbf{x})=\nabla\|\mathbf{x}\|^2=2\mathbf{x}$ core

In one variable, $x^2$ has slope $2x$. The vector version of "$x$ squared" is $\mathbf{x}^\top\mathbf{x}=x_1^2+x_2^2+\cdots+x_n^2$: the sum of the squares of the entries. It is also the squared length $\|\mathbf{x}\|^2$ (Pythagoras, from the L2 norm).

Each entry $x_k$ appears in exactly one term, $x_k^2$. So nudging $x_k$ changes the total by $2x_k\,dx_k$ and nothing else. The gradient is just "$2\times$ the vector itself". It points straight away from the origin, and it is longer the further you are from the origin.

$\mathbf{x}=[3,4]^\top$, so $f=\mathbf{x}^\top\mathbf{x}=9+16=25$.

  1. Prediction: $\nabla f=2\mathbf{x}=[6,8]^\top$.
  2. Nudge $x_1$ from $3$ to $3.01$: $f=3.01^2+4^2=9.0601+16=25.0601$. The change is $0.0601$. The prediction is $6\times0.01=0.06$ ✓.
  3. Nudge $x_2$ from $4$ to $4.01$: $f=9+16.0801=25.0801$. Change $0.0801$, prediction $8\times0.01=0.08$ ✓.
  4. Another case, in 3 dimensions: $\mathbf{x}=[1,-2,2]^\top$ has $f=1+4+4=9$ and $\nabla f=[2,-4,4]^\top$.
$$\nabla_{\mathbf{x}}\,(\mathbf{x}^\top\mathbf{x})=2\mathbf{x},\qquad \nabla_{\mathbf{x}}\,\|\mathbf{x}\|^2=2\mathbf{x},\qquad \nabla_{\mathbf{x}}\,\tfrac12\|\mathbf{x}\|^2=\mathbf{x}.$$

The two expressions are the same function, since $\|\mathbf{x}\|^2=\mathbf{x}^\top\mathbf{x}$. The gradient is radial: it is perpendicular to the circular level curves $\|\mathbf{x}\|=\text{const}$. Its length is $2\|\mathbf{x}\|$. A gradient-descent step with learning rate $\eta$ gives $\mathbf{x}-\eta\,2\mathbf{x}=(1-2\eta)\mathbf{x}$: it simply shrinks the vector. (The gradient of the length $\|\mathbf{x}\|$ itself, not squared, is $\mathbf{x}/\|\mathbf{x}\|$. We derive that in Chapter 2.7.)

Why do we need it?

The squared length is the simplest smooth measure of "how big" or "how far". Its gradient is the building block of every squared-error loss.

Where is it used?

L2 regularisation and weight decay (the penalty $\lambda\|\mathbf{w}\|^2$ adds $2\lambda\mathbf{w}$ to the gradient), squared distance in k-means, and mean squared error.

How is it used?

When a loss contains $\|\mathbf{w}\|^2$, add $2\mathbf{w}$ (or $\mathbf{w}$ for $\frac12\|\mathbf{w}\|^2$) to the gradient. Each update then shrinks the weights slightly: "weight decay".

Choose a method and press Next until the end. The first method works entry by entry; the second uses differentials. Both land on the same answer, and both end with a numerical check.

Drag the blue point. The circles are level curves of $\|\mathbf{x}\|^2$. The green arrow is $\nabla f=2\mathbf{x}$: it is always perpendicular to the circle through the point and twice as long as the position vector. Press Take a descent step a few times: each step multiplies $\mathbf{x}$ by $(1-2\eta)$, so the point walks straight to the origin. What happens if you make $\eta$ bigger than $1$?

$\nabla\|\mathbf{x}\|^2=2\mathbf{x}$, but $\nabla\|\mathbf{x}\|=\mathbf{x}/\|\mathbf{x}\|$ (a unit vector). Do not confuse the squared length with the length. Also, the gradient of $\|\mathbf{x}\|^2$ at $\mathbf{x}=\mathbf{0}$ is $\mathbf{0}$: the bottom of the bowl.

Quick check: what is $\nabla_{\mathbf{w}}\big(\tfrac\lambda2\|\mathbf{w}\|^2\big)$?

The constant $\tfrac\lambda2$ just comes along: $\tfrac\lambda2\cdot2\mathbf{w}=\lambda\mathbf{w}$. This is exactly the weight-decay term.

Important derivatives 2 and 3: the Jacobian of $A\mathbf{x}$, and $\nabla(\mathbf{a}^\top\mathbf{x})$ core

A linear function has the same slope everywhere. $f(x)=5x$ has slope $5$ at every point. The vector versions are:

  • $A\mathbf{x}$: a whole list of weighted sums, one per row of $A$. Its table of slopes is $A$.
  • $\mathbf{a}^\top\mathbf{x}$: one weighted sum. Its list of slopes is $\mathbf{a}$. (It is the single-row case of $A\mathbf{x}$.)

Because the slope never changes, a nudge $d\mathbf{x}$ gives an output change of exactly $A\,d\mathbf{x}$ at any point, with no leftover error.

$A=\begin{bmatrix}2&1\\0&3\\1&-1\end{bmatrix}$ ($3\times2$) and $\mathbf{x}=[1,2]^\top$.

  1. $F(\mathbf{x})=A\mathbf{x}=[2+2,\ 0+6,\ 1-2]^\top=[4,6,-1]^\top$.
  2. Outputs: $F_1=2x_1+x_2$, $F_2=3x_2$, $F_3=x_1-x_2$. Partial derivatives: $\partial F_1/\partial x_1=2$, $\partial F_1/\partial x_2=1$, and so on. They are exactly the entries of $A$. So $J=A$ ($3\times2$).
  3. Nudge $d\mathbf{x}=[0.5,-1]^\top$: $A\,d\mathbf{x}=[1-1,\ -3,\ 0.5+1]^\top=[0,-3,1.5]^\top$. And $F(\mathbf{x}+d\mathbf{x})-F(\mathbf{x})$: $A[1.5,1]^\top=[4,3,0.5]^\top$, minus $[4,6,-1]^\top$ gives $[0,-3,1.5]^\top$. Exactly equal ✓.
  4. Weighted sum. $\mathbf{a}=[3,-1,2]^\top$, $f=\mathbf{a}^\top\mathbf{x}=3x_1-x_2+2x_3$. The slopes are $3$, $-1$, $2$, which is $\mathbf{a}$. At $\mathbf{x}=[1,2,0]^\top$, $f=1$.
$$F(\mathbf{x})=A\mathbf{x}\ \Rightarrow\ J_F=A,\qquad f(\mathbf{x})=\mathbf{a}^\top\mathbf{x}\ \Rightarrow\ \nabla f=\mathbf{a}\ \ (J_f=\mathbf{a}^\top).$$

For $A$ of size $m\times n$: $J=A$ is $m\times n$ (outputs × inputs ✓). For the affine map $A\mathbf{x}+\mathbf{b}$ the Jacobian is still $A$. The Jacobian of the single row $\mathbf{a}^\top$ is the row $\mathbf{a}^\top$, and the gradient (a column) is its transpose $\mathbf{a}$. The one-variable cousin: $(ax)'=a$.

Why do we need it?

Every linear layer, every linear model and every weighted sum is of this form. Its derivative is the simplest of all, so it is the "base case" of backpropagation.

Where is it used?

Fully connected layers ($\mathbf{z}=W\mathbf{x}+\mathbf{b}$: the Jacobian with respect to $\mathbf{x}$ is $W$), linear and logistic regression, and the final score $\mathbf{w}^\top\mathbf{x}$ of a classifier.

How is it used?

To pass a gradient backwards through $\mathbf{z}=W\mathbf{x}+\mathbf{b}$, multiply by $W^\top$: $\nabla_{\mathbf{x}}L=W^\top\nabla_{\mathbf{z}}L$ (this is the Jacobian rule $J^\top\nabla$ from Chapter 2.8).

Four short derivations: pick one and step through it. Compare the "by entries" and "differential" versions of each. They finish with the same answer.

Left: the input plane. Drag the blue point $\mathbf{x}$ and the orange point $\mathbf{x}+d\mathbf{x}$. Right: the output plane, where the blue and orange points are $A\mathbf{x}$ and $A(\mathbf{x}+d\mathbf{x})$. The green arrow is $A\,d\mathbf{x}$. It always joins the two output points exactly, wherever you drag. Edit $A$ to see how it stretches the plane.

Shapes: the Jacobian of $A\mathbf{x}$ is $A$, but the gradient flowing back through $A\mathbf{x}$ uses $A^\top$. Do not mix them up: $J$ goes with a forward nudge ($d\mathbf{y}=J\,d\mathbf{x}$), $J^\top$ goes with a backward gradient.

Quick check: what is $\nabla_{\mathbf{x}}(\mathbf{c}^\top A\mathbf{x})$ for a constant vector $\mathbf{c}$ and matrix $A$?

Write $\mathbf{c}^\top A\mathbf{x}=(A^\top\mathbf{c})^\top\mathbf{x}$, a weighted sum with weights $\mathbf{a}=A^\top\mathbf{c}$. So the gradient is $A^\top\mathbf{c}$.

Important derivative 4: $\nabla(\mathbf{x}^\top A\mathbf{x})=(A+A^\top)\mathbf{x}$ core

In one variable, $ax^2$ has slope $2ax$. The matrix version is $\mathbf{x}^\top A\mathbf{x}=\sum_{i,j}A_{ij}x_ix_j$: a sum of terms of the form (number) × $x_i$ × $x_j$. This is called a quadratic form. The 3D picture is a bowl, a dome or a saddle (Chapter 1.12 of the Linear Algebra guide).

Why does the slope have two parts? Because each input $x_k$ appears twice: once as the left factor $x_i$ (when $i=k$) and once as the right factor $x_j$ (when $j=k$). Each appearance contributes a piece. If $A$ is symmetric the two pieces are equal, and we get $2A\mathbf{x}$, the exact analogue of $2ax$.

$A=\begin{bmatrix}2&1\\0&3\end{bmatrix}$ (not symmetric).

  1. Multiply out: $A\mathbf{x}=[2x_1+x_2,\ 3x_2]^\top$, so $f=x_1(2x_1+x_2)+x_2(3x_2)=2x_1^2+x_1x_2+3x_2^2$.
  2. Partial derivatives: $\partial f/\partial x_1=4x_1+x_2$ and $\partial f/\partial x_2=x_1+6x_2$.
  3. At $\mathbf{x}=[1,2]^\top$: $\nabla f=[6,\,13]^\top$ and $f=2+2+12=16$.
  4. Formula: $A+A^\top=\begin{bmatrix}4&1\\1&6\end{bmatrix}$, and $(A+A^\top)\mathbf{x}=[4+2,\ 1+12]^\top=[6,13]^\top$ ✓.
  5. The naive guess $2A\mathbf{x}=2[4,6]^\top=[8,12]^\top$ is wrong. The naive rule holds only when $A$ is symmetric.
$$\nabla_{\mathbf{x}}(\mathbf{x}^\top A\mathbf{x})=(A+A^\top)\mathbf{x}\ \ \Big(=2A\mathbf{x}\ \text{ if }A=A^\top\Big),\qquad \nabla^2(\mathbf{x}^\top A\mathbf{x})=A+A^\top.$$

A quadratic form only "sees" the symmetric part $\tfrac12(A+A^\top)$ of $A$: the antisymmetric part contributes nothing, because $\mathbf{x}^\top K\mathbf{x}=0$ whenever $K^\top=-K$. The Hessian (the matrix of second derivatives) is the constant $A+A^\top$. One-variable cousin: $(ax^2)'=2ax$, $(ax^2)''=2a$.

Why do we need it?

Quadratic forms measure "energy", "curvature" and "variance". Almost every smooth loss looks like a quadratic bowl near its minimum, so this is the gradient you meet most.

Where is it used?

Least squares and ridge regression (the loss expands into $\mathbf{x}^\top A^\top A\mathbf{x}$), Gaussian distributions (the exponent $-\tfrac12(\mathbf{x}-\boldsymbol\mu)^\top\Sigma^{-1}(\mathbf{x}-\boldsymbol\mu)$), PCA (maximise $\mathbf{w}^\top\Sigma\mathbf{w}$), and Newton's method.

How is it used?

Check whether $A$ is symmetric. If yes, use $2A\mathbf{x}$; if not, use $(A+A^\top)\mathbf{x}$. Setting the gradient to zero finds the bottom of the bowl.

Start with 2×2 by hand (the numbers from the example). Then step through the general derivation by indices, and finally the three-line derivation with differentials. Notice where each method gets the two parts $A\mathbf{x}$ and $A^\top\mathbf{x}$.

Rotate with the background, and use Top to look down. Drag the blue point (it slides on the surface). The green arrow on the floor is the gradient $\nabla f=(A+A^\top)\mathbf{x}$, the direction of steepest uphill; the orange arrow is the same direction lifted onto the surface. The purple patch is the tangent plane. Press the presets: a bowl (positive definite), a saddle (indefinite), a dome (negative definite), and not symmetric: then $2A\mathbf{x}$ and $(A+A^\top)\mathbf{x}$ disagree. Find a point with gradient $\mathbf{0}$.

  • $2A\mathbf{x}$ is only correct when $A=A^\top$. In general it is $(A+A^\top)\mathbf{x}$. Many bugs come from forgetting this when $A$ is not symmetric (for example a weight matrix).
  • $\mathbf{x}^\top A\mathbf{x}$ is a number. Its gradient is a vector. Do not drop the $\mathbf{x}$ from the answer.
Quick check: $A=\begin{bmatrix}3&1\\1&2\end{bmatrix}$ and $\mathbf{x}=[1,-1]^\top$. What is $\nabla(\mathbf{x}^\top A\mathbf{x})$?

$A$ is symmetric, so the gradient is $2A\mathbf{x}=2[3-1,\ 1-2]^\top=2[2,-1]^\top=[4,-2]^\top$. Check with $f=3x_1^2+2x_1x_2+2x_2^2$: $\partial f/\partial x_1=6x_1+2x_2=4$, $\partial f/\partial x_2=2x_1+4x_2=-2$ ✓.

Important derivative 5: $\nabla\|A\mathbf{x}-\mathbf{b}\|^2=2A^\top(A\mathbf{x}-\mathbf{b})$ core

This is the loss of least squares: $A\mathbf{x}$ is the model's prediction, $\mathbf{b}$ the targets, and $\mathbf{r}=A\mathbf{x}-\mathbf{b}$ the vector of errors (the residual). The loss is the squared length of the error vector.

Think of it as "(inside)² → 2 × inside × (slope of the inside)", the chain rule in vector form. The inside is $\mathbf{r}$. Its derivative with respect to $\mathbf{x}$ is $A$ (the Jacobian of $A\mathbf{x}$). The "2 × inside" part is $2\mathbf{r}$. The transpose $A^\top$ appears because $2\mathbf{r}$ lives in output space, and $A^\top$ carries it back to input space: it asks "which inputs caused these errors?".

$A=\begin{bmatrix}1&0\\0&2\\1&1\end{bmatrix}$, $\mathbf{b}=[2,1,4]^\top$, at $\mathbf{x}=[1,1]^\top$.

  1. Prediction: $A\mathbf{x}=[1,2,2]^\top$. Residual: $\mathbf{r}=A\mathbf{x}-\mathbf{b}=[-1,\ 1,\ -2]^\top$.
  2. Loss: $\|\mathbf{r}\|^2=1+1+4=6$.
  3. $A^\top\mathbf{r}=[1(-1)+0(1)+1(-2),\ \ 0(-1)+2(1)+1(-2)]^\top=[-3,\ 0]^\top$.
  4. Gradient: $\nabla f=2A^\top\mathbf{r}=[-6,\ 0]^\top$.
  5. Direct check: $f=(x_1-2)^2+(2x_2-1)^2+(x_1+x_2-4)^2$. $\partial f/\partial x_1=2(x_1-2)+2(x_1+x_2-4)=2(-1)+2(-2)=-6$. $\partial f/\partial x_2=4(2x_2-1)+2(x_1+x_2-4)=4+(-4)=0$ ✓.
  6. Going downhill (the negative gradient) increases $x_1$, because the first and third predictions are both too small.
$$f(\mathbf{x})=\|A\mathbf{x}-\mathbf{b}\|^2\ \Rightarrow\ \nabla f=2A^\top(A\mathbf{x}-\mathbf{b})=2A^\top A\,\mathbf{x}-2A^\top\mathbf{b},\qquad \nabla^2f=2A^\top A.$$

Setting $\nabla f=\mathbf{0}$ gives the normal equations $A^\top A\,\mathbf{x}=A^\top\mathbf{b}$ (Chapter 1.10 of the Linear Algebra guide). In the example, $A^\top A=\begin{bmatrix}2&1\\1&5\end{bmatrix}$, $A^\top\mathbf{b}=[6,6]^\top$, and the solution is $\mathbf{x}^\star=[8/3,\ 2/3]^\top$ with loss $1$. One-variable cousin: $\big((ax-b)^2\big)'=2a(ax-b)$.

Why do we need it?

It is the gradient of the most-used loss in all of data science: squared error of a linear model. It tells us which way to move the weights to make the predictions better.

Where is it used?

Linear regression, ridge regression (add $2\lambda\mathbf{x}$), the last layer of many networks, signal reconstruction, and the Gauss–Newton method for nonlinear least squares.

How is it used?

Compute the residual $\mathbf{r}=A\mathbf{x}-\mathbf{b}$, then the gradient $2A^\top\mathbf{r}$, then step $\mathbf{x}\leftarrow\mathbf{x}-\eta\,2A^\top\mathbf{r}$. Or solve $A^\top A\mathbf{x}=A^\top\mathbf{b}$ directly.

Step through each version: the short one with differentials, the one with entries (chain rule per term), and the one that multiplies out first and reuses the earlier results. The last one ends with the normal equations.

Colour shows the error $\|A\mathbf{x}-\mathbf{b}\|$ over the $(x_1,x_2)$ plane (darker = larger). Drag the blue point: the green arrow is $\nabla f=2A^\top(A\mathbf{x}-\mathbf{b})$ and it crosses the level curves at right angles. The orange path is gradient descent from the point. The purple ring is the exact answer from the normal equations. Use the step size slider: below $2$ the path settles (at $1$ the error along the steepest direction of the bowl disappears in a single step), and above $2$ it diverges. Edit $A$ or $\mathbf{b}$ too.

  • The factor $2$ is real. Many books define the loss as $\tfrac12\|A\mathbf{x}-\mathbf{b}\|^2$ so that the gradient is $A^\top(A\mathbf{x}-\mathbf{b})$ with no $2$. Check which one you have.
  • $A^\top$ (not $A$) multiplies the residual. The shape check catches this: $A(A\mathbf{x}-\mathbf{b})$ does not fit unless $A$ is square.
Quick check: what is the gradient of the ridge loss $\|A\mathbf{x}-\mathbf{b}\|^2+\lambda\|\mathbf{x}\|^2$?

Add the two gradients: $2A^\top(A\mathbf{x}-\mathbf{b})+2\lambda\mathbf{x}$. Setting it to zero gives $(A^\top A+\lambda I)\mathbf{x}=A^\top\mathbf{b}$.

Important derivatives 6 and 7: $\log\mathbf{x}$ and $e^{\mathbf{x}}$ (elementwise) core

In one variable: $(\ln x)'=1/x$ and $(e^x)'=e^x$. In ML we constantly apply such a function to every entry of a vector: $\log\mathbf{x}=[\ln x_1,\ \ln x_2,\dots]$. Each output depends on only the input in its own position. Turning knob $x_2$ moves meter $2$ and leaves all other meters alone.

So most slopes are zero. The only non-zero slopes sit on the diagonal: the Jacobian is a diagonal matrix. This is very good news: no need to build a big matrix, just multiply entry by entry.

  1. $\mathbf{x}=[1,2,4]^\top$. Then $\ln\mathbf{x}=[0,\ 0.693,\ 1.386]^\top$ and $J=\operatorname{diag}(1/x_i)=\operatorname{diag}(1,\ 0.5,\ 0.25)$.
  2. $e^{\mathbf{x}}$ for $\mathbf{x}=[0,1,2]^\top$: $[1,\ 2.718,\ 7.389]^\top$ and $J=\operatorname{diag}(1,\ 2.718,\ 7.389)$.
  3. Scalar loss. $f(\mathbf{x})=\sum_i\ln x_i$ (like a log-likelihood of independent pieces). Then $\partial f/\partial x_k=1/x_k$, so $\nabla f=[1/x_1,\dots,1/x_n]^\top$. At $[1,2,4]$: $[1,\ 0.5,\ 0.25]^\top$.
  4. Log-sum-exp $f(\mathbf{x})=\ln\sum_je^{x_j}$. Then $\partial f/\partial x_k=\dfrac{e^{x_k}}{\sum_je^{x_j}}$, which is entry $k$ of softmax. At $\mathbf{x}=[0,1,2]$: $\nabla f\approx[0.090,\ 0.245,\ 0.665]^\top$ (these add to 1).

Let $g$ be a scalar function applied to each entry: $\mathbf{y}=g(\mathbf{x})$, meaning $y_i=g(x_i)$. Write $\circ$ (or $\odot$) for the entrywise product. Then

$$J=\operatorname{diag}\big(g'(x_1),\dots,g'(x_n)\big),\qquad d\mathbf{y}=g'(\mathbf{x})\circ d\mathbf{x}.$$
  • $\log\mathbf{x}$: $J=\operatorname{diag}(1/\mathbf{x})$ (needs every $x_i>0$). $\quad e^{\mathbf{x}}$: $J=\operatorname{diag}(e^{\mathbf{x}})$.
  • Backward pass: a gradient $\mathbf{v}$ arriving at the output becomes $J^\top\mathbf{v}=g'(\mathbf{x})\circ\mathbf{v}$ at the input. No matrix needed: $n$ multiplications instead of $n^2$.
  • $\nabla\sum_i\ln x_i=1/\mathbf{x}$, $\quad\nabla\sum_ie^{x_i}=e^{\mathbf{x}}$, $\quad\nabla\ln\sum_je^{x_j}=\operatorname{softmax}(\mathbf{x})$.
Why do we need it?

Activation functions, log-likelihoods and probabilities all apply $\log$, $\exp$ or another scalar function entry by entry. Their derivatives must be cheap, and the diagonal structure makes them so.

Where is it used?

Every activation function (ReLU, sigmoid, tanh) in a neural network, the log in cross-entropy and log-likelihood, and the softmax (log-sum-exp) at the end of every classifier.

How is it used?

In backpropagation, multiply the incoming gradient entry by entry with $g'(\mathbf{x})$. Frameworks do exactly this and never build the diagonal matrix.

Step through each derivation. Notice the moment where "off-diagonal" slopes turn out to be zero, and how log-sum-exp produces the softmax.

Choose a function and edit $\mathbf{x}$ and the incoming gradient $\mathbf{v}$. The Jacobian is printed as a table: only the diagonal is non-zero. Check that $J^\top\mathbf{v}$ (matrix times vector) equals the cheap entrywise product $g'(\mathbf{x})\circ\mathbf{v}$. For $\log$, make an entry of $\mathbf{x}$ negative and see the message. When you pick $e^x$, the log-sum-exp gradient (softmax) is also checked.

  • $\log\mathbf{x}$ (a vector) is not a scalar. Its derivative is a matrix (diagonal), not "$1/\mathbf{x}$". Only for a scalar sum like $\sum\ln x_i$ is the gradient the vector $1/\mathbf{x}$.
  • $\log$ needs positive inputs. Real code works with $\log(\text{softmax})$ in one combined, stable function, never $\ln$ of a tiny number.
  • Do not confuse $e^{\mathbf{x}}$ (entrywise) with the matrix exponential $e^{A}$.
Quick check: what is $\nabla_{\mathbf{x}}\sum_i x_i\ln x_i$ (the negative entropy)?

For each entry, $\frac{d}{dx}(x\ln x)=\ln x+1$ (product rule). So $\nabla f=\ln\mathbf{x}+\mathbf{1}$.

Check any derivative numerically core

How do you know a derivative you derived by hand is right? Test it with brute force. Nudge each input up a tiny bit and down a tiny bit, and measure the slope from the two function values. If that matches your formula, you are almost surely right. It costs two function evaluations per input, so it is far too slow for training. But it is perfect for catching mistakes.

$f(\mathbf{x})=\mathbf{x}^\top\mathbf{x}$ at $\mathbf{x}=[3,4]^\top$, with $h=0.001$.

  1. Nudge $x_1$: $f(3.001,4)=9.006001+16=25.006001$ and $f(2.999,4)=8.994001+16=24.994001$.
  2. Slope: $(25.006001-24.994001)/(2\times0.001)=0.012/0.002=6$. The formula $2x_1=6$. ✓
  3. Same for $x_2$: $8$ ✓. For a quadratic like this, the two-sided difference is exact (apart from round-off).

The central finite difference for each input $x_k$ (with $\mathbf{e}_k$ the vector with a 1 in position $k$):

$$\frac{\partial f}{\partial x_k}\approx\frac{f(\mathbf{x}+h\,\mathbf{e}_k)-f(\mathbf{x}-h\,\mathbf{e}_k)}{2h}.$$

The error shrinks like $h^2$ as $h$ gets smaller, until round-off in the computer's arithmetic takes over. A good all-round value is $h\approx10^{-5}$. Compare with the formula using the relative error $\dfrac{\|\text{formula}-\text{numerical}\|}{\max(1,\|\text{formula}\|)}$. Values around $10^{-6}$ or smaller mean "correct". Values around $10^{-2}$ or bigger mean "there is a bug".

Why do we need it?

Hand-derived gradients are easy to get wrong by a transpose, a missing 2 or a sign. A numerical check catches all of these in seconds.

Where is it used?

"Gradient checking" when writing a custom loss or layer, unit tests in ML libraries (for example PyTorch's gradcheck), and verifying hand-written backpropagation.

How is it used?

Pick a random point, compute the analytic gradient and the finite-difference gradient, and compare. Always test at a random point, never at $\mathbf{x}=\mathbf{0}$, where many wrong formulas happen to give 0.

Pick one of the important derivatives from this chapter. Edit $A$, $\mathbf{x}$ and $\mathbf{b}$ (the formulas hold for any values). The formula and the nudge-each-entry gradient should agree to many digits. Slide the step $h$: if it is too large, the answer is rough; if it is far too small, round-off noise appears. Try the not symmetric $A$ with the xᵀAx choice, and look at the wrong guess $2A\mathbf{x}$ in the last line.

  • A numerical check is a test, not a training method: it needs $2n$ function evaluations for $n$ inputs.
  • Check at a random point and with non-square, non-symmetric matrices. Special values hide bugs.
  • At a kink (like $|x|$ at $0$ or ReLU at $0$) the numerical slope depends on $h$: that is not a bug in your formula (see Chapter 2.7 on subgradients).
Quick check: your formula and the numerical gradient differ by exactly a factor of 2 in every entry. What is the likely mistake?

A missing or extra factor $2$: for instance using $A^\top(A\mathbf{x}-\mathbf{b})$ for $\|A\mathbf{x}-\mathbf{b}\|^2$ (correct for $\tfrac12\|\cdot\|^2$), or the other way round.

Recap, cheat sheet and practice

  • A derivative has one slope per (output, input) pair. Entries = outputs × inputs. Scalar out: the gradient has the shape of the input. Vector to vector: the $m\times n$ Jacobian.
  • Conventions here: $\nabla f$ is a column; $J_F$ is $m\times n$; $H$ is symmetric $n\times n$; $\nabla_Xf$ has the shape of $X$. Other books may use the transposed layout.
  • Differentials replace derivatives: $d(XY)=dX\,Y+X\,dY$, $d(X^\top)=(dX)^\top$, $d(X^{-1})=-X^{-1}dX\,X^{-1}$, $d\operatorname{tr}X=\operatorname{tr}dX$. Keep the order of factors.
  • Trace trick: $\operatorname{tr}(AB)=\operatorname{tr}(BA)$, a scalar is its own trace, and $df=\operatorname{tr}(G^\top dX)\Rightarrow\nabla_Xf=G$.
  • Recipe: write $df$, rearrange so that $dX$ (or $d\mathbf{x}$) is at the right end, read off the coefficient, transpose if needed, then check the shape and a numerical gradient.
  • Results: $\nabla\mathbf{x}^\top\mathbf{x}=2\mathbf{x}$; $J_{A\mathbf{x}}=A$; $\nabla\mathbf{a}^\top\mathbf{x}=\mathbf{a}$; $\nabla\mathbf{x}^\top A\mathbf{x}=(A+A^\top)\mathbf{x}$; $\nabla\|A\mathbf{x}-\mathbf{b}\|^2=2A^\top(A\mathbf{x}-\mathbf{b})$; elementwise $\log$ and $\exp$ have diagonal Jacobians.

Cheat sheet

FunctionDerivativeHow we got it
$\mathbf{x}^\top\mathbf{x}=\|\mathbf{x}\|^2$$\nabla=2\mathbf{x}$$df=2\mathbf{x}^\top d\mathbf{x}$
$A\mathbf{x}$$J=A$$d\mathbf{y}=A\,d\mathbf{x}$
$\mathbf{a}^\top\mathbf{x}$$\nabla=\mathbf{a}$$df=\mathbf{a}^\top d\mathbf{x}$
$\mathbf{x}^\top A\mathbf{x}$$\nabla=(A+A^\top)\mathbf{x}$; $H=A+A^\top$product rule, then transpose the scalar term
$\|A\mathbf{x}-\mathbf{b}\|^2$$\nabla=2A^\top(A\mathbf{x}-\mathbf{b})$$df=2\mathbf{r}^\top A\,d\mathbf{x}$, $\mathbf{r}=A\mathbf{x}-\mathbf{b}$
$\log\mathbf{x}$, $e^{\mathbf{x}}$ (entrywise)$J=\operatorname{diag}(1/\mathbf{x})$, $\operatorname{diag}(e^{\mathbf{x}})$each output uses only its own input
$\ln\sum_je^{x_j}$$\nabla=\operatorname{softmax}(\mathbf{x})$chain rule on $\ln S$
$\operatorname{tr}(AX)$$\nabla_X=A^\top$$df=\operatorname{tr}(A\,dX)$
$\mathbf{a}^\top X\mathbf{b}$$\nabla_X=\mathbf{a}\mathbf{b}^\top$rotate $\mathbf{b}$ to the front of the trace
Code it · NumPy

import numpy as np

def num_grad(f, x, h=1e-5):
    """Central-difference gradient of a scalar function f at the vector x."""
    g = np.zeros_like(x)
    for k in range(x.size):
        e = np.zeros_like(x); e[k] = h
        g[k] = (f(x + e) - f(x - e)) / (2 * h)
    return g

# 1) x^T x  ->  2x
x = np.array([3.0, 4.0])
print(num_grad(lambda v: v @ v, x))             # [6. 8.]   same as 2*x

# 2) x^T A x  ->  (A + A^T) x   (A is NOT symmetric)
A = np.array([[2.0, 1.0], [0.0, 3.0]])
x = np.array([1.0, 2.0])
print((A + A.T) @ x)                            # [ 6. 13.]  the right formula
print(2 * A @ x)                                # [ 8. 12.]  the naive (wrong) formula
print(num_grad(lambda v: v @ A @ v, x))         # [ 6. 13.]  numbers agree with (A + A^T) x

# 3) ||Ax - b||^2  ->  2 A^T (Ax - b)
A = np.array([[1.0, 0.0], [0.0, 2.0], [1.0, 1.0]])
b = np.array([2.0, 1.0, 4.0])
x = np.array([1.0, 1.0])
r = A @ x - b
print(r, r @ r)                                 # [-1.  1. -2.] 6.0
print(2 * A.T @ r)                              # [-6.  0.]
print(num_grad(lambda v: np.sum((A @ v - b) ** 2), x))   # [-6.  0.]
print(np.linalg.solve(A.T @ A, A.T @ b))        # [2.66666667 0.66666667] = [8/3, 2/3], the normal equations

# 4) Jacobian of an elementwise function is diagonal
x = np.array([1.0, 2.0, 4.0])
print(np.diag(1 / x))                           # J of log(x): diag(1, 0.5, 0.25)
v = np.array([1.0, 1.0, 1.0])
print(np.diag(1 / x).T @ v, v / x)              # [1.   0.5  0.25] twice: the entrywise product is enough

# 5) The trace trick: tr(AB) = tr(BA) even when the sizes differ
A = np.array([[1, 2, 0], [0, 1, 3]]); B = np.array([[1, 0], [2, 1], [0, 1]])
print(np.trace(A @ B), np.trace(B @ A))         # 9 9

# 6) Gradient with respect to a matrix: f(X) = tr(AX) has gradient A^T (the shape of X)
A = np.array([[1.0, 2.0], [3.0, 4.0]]); X = np.array([[0.5, -1.0], [2.0, 1.0]])
G = np.zeros_like(X)
for i in range(2):
    for j in range(2):
        E = np.zeros_like(X); E[i, j] = 1e-5
        G[i, j] = (np.trace(A @ (X + E)) - np.trace(A @ (X - E))) / 2e-5
print(np.round(G, 6))                           # [[1. 3.] [2. 4.]]
print(A.T)                                      # [[1. 3.] [2. 4.]]  same
Test yourself

1. $F:\mathbb{R}^5\to\mathbb{R}^3$. In this guide, what is the shape of its Jacobian?

One row per output (3) and one column per input (5). The $5\times3$ shape is the denominator-layout convention, or $J^\top$.

2. What is $\nabla_{\mathbf{x}}(\mathbf{x}^\top A\mathbf{x})$ when $A$ is not symmetric?

The differential is $\mathbf{x}^\top(A+A^\top)d\mathbf{x}$. For symmetric $A$ it reduces to $2A\mathbf{x}$, but not in general.

3. $f(\mathbf{x})=\|A\mathbf{x}-\mathbf{b}\|^2$ with $A$ of size $3\times4$. Which gradient has a valid shape?

The gradient must have the shape of $\mathbf{x}$ ($4\times1$). $A^\top$ is $4\times3$ and the residual is $3\times1$, so the product is $4\times1$. The other options do not multiply, or have the wrong shape ($4\times4$).

4. Which of these is true for matrices $A$ ($2\times3$) and $B$ ($3\times2$)?

Both traces are sums of the same numbers, $\sum_{i,k}A_{ik}B_{ki}$. $AB$ is $2\times2$ and $BA$ is $3\times3$, and both are square, so both have traces.

5. What is the Jacobian of $\mathbf{y}=e^{\mathbf{x}}$ (entrywise) for $\mathbf{x}\in\mathbb{R}^3$?

Output $i$ depends only on $x_i$, so every off-diagonal slope is zero, and the diagonal entries are $e^{x_i}$ (different for each entry).

6. $d(X^{-1})$ equals…

Nudge both sides of $XX^{-1}=I$: $dX\,X^{-1}+X\,d(X^{-1})=0$. Multiply on the left by $X^{-1}$ to solve. The inverse appears on both sides because matrices do not commute.

Practice problems

A. Find $\nabla f$ for $f(\mathbf{x})=\mathbf{c}^\top\mathbf{x}+\mathbf{x}^\top\mathbf{x}$ at $\mathbf{x}=[1,2]^\top$ with $\mathbf{c}=[3,-1]^\top$, using differentials.

$df=\mathbf{c}^\top d\mathbf{x}+2\mathbf{x}^\top d\mathbf{x}=(\mathbf{c}+2\mathbf{x})^\top d\mathbf{x}$, so $\nabla f=\mathbf{c}+2\mathbf{x}=[3,-1]+[2,4]=[5,\,3]^\top$. Check: $f=3x_1-x_2+x_1^2+x_2^2$ gives $\partial_1=3+2x_1=5$, $\partial_2=-1+2x_2=3$ ✓.

B. For $A=\begin{bmatrix}1&2\\3&4\end{bmatrix}$ and $\mathbf{x}=[1,1]^\top$, compute $\nabla(\mathbf{x}^\top A\mathbf{x})$ and compare with $2A\mathbf{x}$.

$A+A^\top=\begin{bmatrix}2&5\\5&8\end{bmatrix}$, so $\nabla=[7,13]^\top$. And $2A\mathbf{x}=2[3,7]^\top=[6,14]^\top$, which is different (since $A$ is not symmetric). Check: $f=x_1^2+5x_1x_2+4x_2^2$; $\partial_1=2x_1+5x_2=7$, $\partial_2=5x_1+8x_2=13$ ✓.

C. Derive $\nabla_{\mathbf{w}}\|X\mathbf{w}-\mathbf{y}\|^2+\lambda\|\mathbf{w}\|^2$ and the solution of the ridge equations, for a data matrix $X$.

With $\mathbf{r}=X\mathbf{w}-\mathbf{y}$, $df=2\mathbf{r}^\top X\,d\mathbf{w}+2\lambda\mathbf{w}^\top d\mathbf{w}$. So $\nabla=2X^\top(X\mathbf{w}-\mathbf{y})+2\lambda\mathbf{w}$. Setting it to $\mathbf{0}$: $(X^\top X+\lambda I)\mathbf{w}=X^\top\mathbf{y}$. Shape check: both sides are $n\times1$.

D. Use the recipe to find $\nabla_X\operatorname{tr}(X^\top AX)$ for a square $A$, in terms of $X$ and $A$.

$df=\operatorname{tr}(dX^\top AX)+\operatorname{tr}(X^\top A\,dX)$. The first term equals its transpose $\operatorname{tr}(X^\top A^\top dX)$. So $df=\operatorname{tr}\big(X^\top(A^\top+A)\,dX\big)$, and $G^\top=X^\top(A+A^\top)$, giving $\nabla_X=(A+A^\top)X$. (Chapter 2.7 revisits this.)

E. $\mathbf{y}=\ln\mathbf{x}$ and $L=\mathbf{c}^\top\mathbf{y}$ for $\mathbf{x}\in\mathbb{R}^3$. Find $\nabla_{\mathbf{x}}L$.

$L=\sum_ic_i\ln x_i$, so $\partial L/\partial x_k=c_k/x_k$ and $\nabla L=\mathbf{c}/\mathbf{x}$ (entrywise). Via Jacobians: $J^\top\mathbf{c}=\operatorname{diag}(1/\mathbf{x})\,\mathbf{c}$, the same thing.

F. Shape check: $W$ is $4\times6$, $\mathbf{x}\in\mathbb{R}^6$, $\mathbf{g}=\nabla_{\mathbf{z}}L\in\mathbb{R}^4$ where $\mathbf{z}=W\mathbf{x}$. Show that $\nabla_WL=\mathbf{g}\mathbf{x}^\top$ and $\nabla_{\mathbf{x}}L=W^\top\mathbf{g}$.

$d\mathbf{z}=dW\,\mathbf{x}+W\,d\mathbf{x}$ and $dL=\mathbf{g}^\top d\mathbf{z}$. With $\mathbf{x}$ fixed: $dL=\mathbf{g}^\top dW\,\mathbf{x}=\operatorname{tr}(\mathbf{x}\mathbf{g}^\top dW)$, so $G^\top=\mathbf{x}\mathbf{g}^\top$, $G=\mathbf{g}\mathbf{x}^\top$: $(4\times1)(1\times6)=4\times6$ ✓. With $W$ fixed: $dL=\mathbf{g}^\top W\,d\mathbf{x}$, so $\nabla_{\mathbf{x}}L=W^\top\mathbf{g}$ ($6\times4$ times $4\times1$) ✓.

Chapter 2.7

Useful Gradient Identities

There are about a dozen gradient formulas that show up again and again in machine learning. This chapter gives you all of them, but with a twist: you are not asked to memorise a single one. Each comes with a derivation you can follow, and all of them come from one reusable recipe.

  • Learn the derivation recipe: write the differential, rearrange into $\operatorname{tr}(G^\top dX)$, read off $G$
  • Derive the gradients of linear, quadratic and transpose expressions, and of trace expressions $\operatorname{tr}(AX)$, $\operatorname{tr}(X^\top AX)$, $\operatorname{tr}(AXB)$
  • See what symmetry buys you, and the gradients of vector norms ($L_2$, $L_1$, $L_\infty$) and matrix norms (Frobenius; nuclear and spectral for awareness)
  • Meet $\log\det X$ and $d(X^{-1})$, and the chain-rule identities $\nabla_{\mathbf{x}}f(A\mathbf{x})=A^\top\nabla f$
  • Finish with a one-page identity table and a routine to check any gradient numerically

This chapter uses the tools of Chapter 2.6: differentials, the trace trick, the gradient/Jacobian conventions (gradient = column; Jacobian $m\times n$; $\nabla_Xf$ has the shape of $X$), and shape checks. The Linear Algebra guide has a shorter table: Matrix calculus in the Linear Algebra guide. Here, every row is derived, and we add the trace, norm, log-determinant and chain-rule families.

Derive, don't memorise: the recipe core

Imagine you had to remember the multiplication table up to $1000\times1000$. Impossible. Instead you learn how to multiply, and then any product is a few steps away. Gradient identities are the same. There are dozens of them, but they all come out of one short procedure.

The procedure asks the function a question: "if I nudge my input a tiny bit, how much do you change?" Whatever the function says, we tidy the answer until the nudge stands alone at the end. The coefficient in front of the nudge is the gradient.

A first run of the recipe on a function we have not seen yet: $f(W)=\tfrac12\|W\mathbf{x}-\mathbf{y}\|^2$, where $W$ is $m\times n$, $\mathbf{x}\in\mathbb{R}^n$ and $\mathbf{y}\in\mathbb{R}^m$ are fixed.

  1. Differential. Let $\mathbf{r}=W\mathbf{x}-\mathbf{y}$. Then $d\mathbf{r}=dW\,\mathbf{x}$, and $df=\mathbf{r}^\top d\mathbf{r}=\mathbf{r}^\top dW\,\mathbf{x}$.
  2. Rearrange. This is a scalar, so wrap it in a trace and rotate $\mathbf{x}$ to the front: $df=\operatorname{tr}(\mathbf{x}\,\mathbf{r}^\top dW)$.
  3. Read off. $G^\top=\mathbf{x}\mathbf{r}^\top$, so $\nabla_Wf=\mathbf{r}\,\mathbf{x}^\top=(W\mathbf{x}-\mathbf{y})\mathbf{x}^\top$, which is $m\times n$, the shape of $W$ ✓.

This is exactly the gradient for the weights of a linear layer with squared error. We never looked it up.

The recipe for a scalar function $f$ of a vector $\mathbf{x}$ or matrix $X$:

  1. Differential. Write $df$. Constants do not move. Use $d(XY)=dX\,Y+X\,dY$, $d(X^\top)=(dX)^\top$, $d(X^{-1})=-X^{-1}dX\,X^{-1}$, $d\operatorname{tr}=\operatorname{tr}\,d$. Keep factor order.
  2. Rearrange. Get every term into the form $\operatorname{tr}(M\,dX)$ (or $\mathbf{g}^\top d\mathbf{x}$), using the toolkit below.
  3. Read off. $df=\mathbf{g}^\top d\mathbf{x}\Rightarrow\nabla f=\mathbf{g}$. $\ df=\operatorname{tr}(M\,dX)\Rightarrow\nabla_Xf=M^\top$.
  4. Check. Shape of the answer = shape of the variable. Then a numerical check.

The rearrangement toolkit (only four moves):

MoveRuleUse it to…
wrap a scalar$s=\operatorname{tr}(s)$turn a number into something rotatable
rotate$\operatorname{tr}(PQ)=\operatorname{tr}(QP)$move $dX$ to the right end
transpose a scalar or trace$s=s^\top$, $\operatorname{tr}(P)=\operatorname{tr}(P^\top)$, $(PQ)^\top=Q^\top P^\top$flip a $dX^\top$ into a $dX$
collect$\operatorname{tr}(P\,dX)+\operatorname{tr}(Q\,dX)=\operatorname{tr}((P+Q)\,dX)$merge terms
Why do we need it?

Memorised tables fail the first time your loss is slightly different. A procedure that you can run on paper works on any expression made of products, transposes, traces and inverses.

Where is it used?

Writing the backward pass of a new layer, deriving the update of an EM or matrix-factorisation algorithm, checking the output of autograd, and reading the derivations in papers on Gaussian processes, PCA and variational inference.

How is it used?

Do the four steps on paper, then confirm with a finite-difference check (the last section of this chapter). When the two agree, you can trust the formula.

Pick an example and press Next to reveal one line at a time. Before you press, try to guess which of the four moves (wrap, rotate, transpose, collect) comes next. Both examples end with a shape check.

  • The recipe only works on scalar functions. For a vector-valued function, use $d\mathbf{y}=M\,d\mathbf{x}\Rightarrow J=M$.
  • Do not skip the final transpose ($\operatorname{tr}(M\,dX)\Rightarrow M^\top$). A shape check usually catches a forgotten transpose when $X$ is not square.
  • You may rotate a trace, but you may not swap two factors inside it.
Quick check: run the recipe on $f(\mathbf{x})=\mathbf{x}^\top B\mathbf{x}$ and say after which step the answer appears.

$df=(d\mathbf{x})^\top B\mathbf{x}+\mathbf{x}^\top B\,d\mathbf{x}$. Transpose the first term (a scalar): $\mathbf{x}^\top B^\top d\mathbf{x}$. Collect: $df=\mathbf{x}^\top(B+B^\top)\,d\mathbf{x}$. Read off: $\nabla f=(B+B^\top)\mathbf{x}$, the quadratic-form identity of Chapter 2.6.

Derivative of linear functions core

A linear function is the easiest thing to differentiate: its slope is the same everywhere, like a ramp with a constant incline. Whatever the input (a number, a vector, a matrix), the answer is just "the weights".

A linear function of a matrix $X$ is a weighted sum of its entries: $f(X)=\sum_{i,j}C_{ij}X_{ij}$. The slope for entry $X_{ij}$ is its weight $C_{ij}$. So the gradient is the weight matrix $C$ itself. Everything else in this family is a disguise of that one fact.

  1. Vector. $f(\mathbf{x})=3x_1-x_2+2x_3$: gradient $[3,-1,2]^\top=\mathbf{a}$ for $f=\mathbf{a}^\top\mathbf{x}$.
  2. Weighted sum of a matrix. $f(X)=1\cdot X_{11}+3\cdot X_{12}+2\cdot X_{21}+4\cdot X_{22}$. The gradient is $\begin{bmatrix}1&3\\2&4\end{bmatrix}=C$, and $f=\operatorname{tr}(C^\top X)$.
  3. A trace in disguise. $\operatorname{tr}(AX)$ with $A=\begin{bmatrix}1&2\\3&4\end{bmatrix}$ equals $A_{11}X_{11}+A_{12}X_{21}+A_{21}X_{12}+A_{22}X_{22}=X_{11}+2X_{21}+3X_{12}+4X_{22}$. The weights, placed in the grid of $X$, form $\begin{bmatrix}1&3\\2&4\end{bmatrix}=A^\top$ ✓.
  4. Bilinear. $f(X)=\mathbf{a}^\top X\mathbf{b}=\sum_{ij}a_iX_{ij}b_j$: the weight of $X_{ij}$ is $a_ib_j$, so the gradient is $\mathbf{a}\mathbf{b}^\top$.

With constant $\mathbf{a},\mathbf{b},\mathbf{c}$ and matrices $A,B,C$ of fitting shapes:

FunctionDerivativeReason
$\mathbf{a}^\top\mathbf{x}$, also $\mathbf{x}^\top\mathbf{a}$$\nabla=\mathbf{a}$$df=\mathbf{a}^\top d\mathbf{x}$
$A\mathbf{x}+\mathbf{c}$$J=A$$d\mathbf{y}=A\,d\mathbf{x}$ (the constant drops out)
$\operatorname{tr}(C^\top X)=\sum C_{ij}X_{ij}$$\nabla_X=C$weights of the entries
$\operatorname{tr}(AX)$$\nabla_X=A^\top$$df=\operatorname{tr}(A\,dX)$
$\mathbf{a}^\top X\mathbf{b}$$\nabla_X=\mathbf{a}\mathbf{b}^\top$$df=\operatorname{tr}(\mathbf{b}\mathbf{a}^\top dX)$
$\operatorname{tr}(X)$$\nabla_X=I$$df=\operatorname{tr}(I\,dX)$

For a linear function the gradient does not depend on where you stand: the Hessian is zero.

Why do we need it?

Linear pieces are in every model: a prediction $\mathbf{w}^\top\mathbf{x}+b$, a layer $W\mathbf{x}+\mathbf{b}$. They are the building blocks that all bigger gradients are assembled from.

Where is it used?

Linear and logistic regression scores, every dense layer of a neural network, the output projection in a Transformer, and the "linear term" of every quadratic model.

How is it used?

Spot the weights: whatever multiplies the variable (arranged in the shape of the variable) is the gradient. Use the table as a quick lookup, but be able to rebuild each row with the recipe.

Step through the two derivations. In the first, the gradient simply appears as the weights. In the second you see where the transpose in $\nabla\operatorname{tr}(AX)=A^\top$ comes from.

The plane shows $f(x,y)=a_1x+a_2y$ as shading, with the gradient drawn as an arrow on a grid of points. All arrows are identical: the slope does not depend on where you stand. Change $a_1$ and $a_2$ and drag the blue point. The green arrow is $\nabla f=\mathbf{a}$. The level lines are straight and perpendicular to it.

Pick an identity and edit the matrices and vectors. The formula and the brute-force gradient should always agree, whatever you type. Notice that $\nabla\operatorname{tr}(AX)$ is $A^\top$ (look for the transposed pattern), but $\nabla\operatorname{tr}(A^\top X)$ is $A$.

  • $\nabla_X\operatorname{tr}(AX)=A^\top$ but $\nabla_X\operatorname{tr}(A^\top X)=A$. The transpose depends on how the trace is written. When in doubt, count: the weight of $X_{ij}$.
  • Linear in $X$ does not mean the gradient is the same matrix for any shape: use the shape check ($\mathbf{a}\mathbf{b}^\top$ must have the shape of $X$).
Quick check: $X$ is $2\times3$, $\mathbf{a}\in\mathbb{R}^2$, $\mathbf{b}\in\mathbb{R}^3$. What is $\nabla_X(\mathbf{a}^\top X\mathbf{b})$ and what is its shape?

$\mathbf{a}\mathbf{b}^\top$, shape $(2\times1)(1\times3)=2\times3$, matching $X$. Entry $(i,j)$ is $a_ib_j$.

Derivative of quadratic functions core

In one variable a quadratic is $ax^2+bx+c$: a parabola. Its slope is $2ax+b$ and it has one flat spot, at $x=-b/(2a)$.

The vector version has the same three ingredients: a bowl (or dome, or saddle) $\tfrac12\mathbf{x}^\top A\mathbf{x}$, a ramp $\mathbf{b}^\top\mathbf{x}$ that tilts it, and a constant lift $c$. The ramp slides the flat spot away from the origin. Finding where the gradient is zero tells you where the bottom of the bowl is.

$A=\begin{bmatrix}2&0\\0&4\end{bmatrix}$, $\mathbf{b}=[-2,-4]^\top$, $c=0$:

  1. $f(\mathbf{x})=\tfrac12\mathbf{x}^\top A\mathbf{x}+\mathbf{b}^\top\mathbf{x}=x_1^2+2x_2^2-2x_1-4x_2$.
  2. Gradient (school way): $[2x_1-2,\ \ 4x_2-4]$. Formula: $A\mathbf{x}+\mathbf{b}=[2x_1-2,\ 4x_2-4]$ ✓.
  3. Set it to zero: $A\mathbf{x}=-\mathbf{b}$, so $\mathbf{x}^\star=[1,1]^\top$ and $f(\mathbf{x}^\star)=1+2-2-4=-3$.
  4. Completing the square: $f=\tfrac12(\mathbf{x}-\mathbf{x}^\star)^\top A(\mathbf{x}-\mathbf{x}^\star)-3=(x_1-1)^2+2(x_2-1)^2-3$. Expanding gives back $x_1^2+2x_2^2-2x_1-4x_2$ ✓.

For $f(\mathbf{x})=\tfrac12\mathbf{x}^\top A\mathbf{x}+\mathbf{b}^\top\mathbf{x}+c$:

$$\nabla f=\tfrac12(A+A^\top)\mathbf{x}+\mathbf{b}\ \ \Big(=A\mathbf{x}+\mathbf{b}\ \text{ if }A=A^\top\Big),\qquad \nabla^2f=\tfrac12(A+A^\top).$$

Without the $\tfrac12$: $\nabla(\mathbf{x}^\top A\mathbf{x}+\mathbf{b}^\top\mathbf{x})=(A+A^\top)\mathbf{x}+\mathbf{b}$. For a symmetric invertible $A$, the flat spot is $\mathbf{x}^\star=-A^{-1}\mathbf{b}$: a minimum if $A$ is positive definite (a bowl), a maximum if negative definite, a saddle if indefinite.

Shifted form (a squared distance measured with a weight matrix, as in the Gaussian exponent): $\nabla\big[(\mathbf{x}-\mathbf{c})^\top A(\mathbf{x}-\mathbf{c})\big]=(A+A^\top)(\mathbf{x}-\mathbf{c})$, because $d(\mathbf{x}-\mathbf{c})=d\mathbf{x}$.

Why do we need it?

Near its minimum almost every smooth loss looks like a quadratic. Knowing the quadratic's gradient and flat spot lets us understand, and speed up, optimisation.

Where is it used?

Least squares and ridge regression, the Gaussian log-density (Mahalanobis distance), Newton's method (it solves a quadratic model exactly in one step), trust-region methods and the analysis of why gradient descent zig-zags.

How is it used?

Read off $A$ and $\mathbf{b}$, compute $\nabla f=A\mathbf{x}+\mathbf{b}$, and solve $A\mathbf{x}=-\mathbf{b}$ for the flat spot. Never form $A^{-1}$ explicitly in code: use a linear solver.

Step through the derivation with differentials, then see the same function reduced to a sum of squares by completing the square, and finally the shifted version $(\mathbf{x}-\mathbf{c})^\top A(\mathbf{x}-\mathbf{c})$.

The contours are level curves of $f(\mathbf{x})=\tfrac12\mathbf{x}^\top A\mathbf{x}+\mathbf{b}^\top\mathbf{x}$. The small arrows are the gradient field: they point uphill. Drag the blue point. Press Jump to the flat spot to move it to the purple ring where $\nabla f=\mathbf{0}$. Edit $A$: make it positive definite (a bowl, the arrows point away from the centre), then indefinite (a saddle, the arrows cross). Change $\mathbf{b}$ and watch the flat spot slide.

  • Watch the $\tfrac12$. With $\tfrac12\mathbf{x}^\top A\mathbf{x}$ the gradient is $A\mathbf{x}$ (for symmetric $A$); without it, $2A\mathbf{x}$.
  • If $A$ is not symmetric, replace $A$ by $\tfrac12(A+A^\top)$ everywhere (the symmetric-matrix section below explains why this loses nothing).
  • $A\mathbf{x}=-\mathbf{b}$ has a unique solution only if $A$ is invertible. If $A$ is singular the surface is flat in some direction.
Quick check: $f(\mathbf{x})=\mathbf{x}^\top\mathbf{x}-6x_1$ in two variables. Where is its minimum?

Write $f=\tfrac12\mathbf{x}^\top(2I)\mathbf{x}+\mathbf{b}^\top\mathbf{x}$ with $\mathbf{b}=[-6,0]$. The gradient is $2\mathbf{x}+\mathbf{b}=[2x_1-6,\ 2x_2]$. It is zero at $\mathbf{x}^\star=[3,0]^\top$ and $f=9-18=-9$.

Derivatives involving a transpose

Transposing a table flips rows and columns. A nudge of every entry flips in the same way: nudging a matrix and then transposing it gives the same result as transposing first and then nudging. That is the rule $d(X^\top)=(dX)^\top$.

So a transpose inside the function usually just means a transposed pattern in the gradient. Two small tricks do all the work: a number equals its own transpose (so you can flip a whole term), and $(PQ)^\top=Q^\top P^\top$ (the order reverses).

$X$ is $2\times2$, $\mathbf{a}=[1,2]^\top$, $\mathbf{b}=[3,4]^\top$. Compare the two functions $\mathbf{a}^\top X\mathbf{b}$ and $\mathbf{a}^\top X^\top\mathbf{b}$.

  1. $\mathbf{a}^\top X\mathbf{b}=\sum_{ij}a_iX_{ij}b_j$: the weight of $X_{ij}$ is $a_ib_j$, so the gradient is $\mathbf{a}\mathbf{b}^\top=\begin{bmatrix}3&4\\6&8\end{bmatrix}$.
  2. $\mathbf{a}^\top X^\top\mathbf{b}=\sum_{ij}a_jX_{ij}b_i$ (since $(X^\top)_{ji}=X_{ij}$): the weight of $X_{ij}$ is $b_ia_j$, so the gradient is $\mathbf{b}\mathbf{a}^\top=\begin{bmatrix}3&6\\4&8\end{bmatrix}$.
  3. The two gradients are transposes of each other. Check with the number trick: $\mathbf{a}^\top X^\top\mathbf{b}$ is a number, so it equals its transpose $\mathbf{b}^\top X\mathbf{a}$, whose gradient is $\mathbf{b}\mathbf{a}^\top$ ✓.
  • $d(X^\top)=(dX)^\top$, and $\operatorname{tr}(P^\top)=\operatorname{tr}(P)$.
  • $\nabla_X(\mathbf{a}^\top X^\top\mathbf{b})=\mathbf{b}\mathbf{a}^\top$ (compare $\nabla_X(\mathbf{a}^\top X\mathbf{b})=\mathbf{a}\mathbf{b}^\top$).
  • $\nabla_X\operatorname{tr}(X^\top A)=A$, but $\nabla_X\operatorname{tr}(XA)=A^\top$. Also $\nabla_X\operatorname{tr}(AX^\top)=A$.
  • For vectors: $\nabla_{\mathbf{x}}(\mathbf{b}^\top A^\top\mathbf{x})=A\mathbf{b}$ (compare $\nabla_{\mathbf{x}}(\mathbf{b}^\top A\mathbf{x})=A^\top\mathbf{b}$). The Jacobian of $A^\top\mathbf{x}$ is $A^\top$.
  • Flipping the argument. If $g(X)=f(X^\top)$, then $\nabla g(X)=\big(\nabla f(X^\top)\big)^\top$: the gradient of a function of the transpose is the transpose of the gradient.
Why do we need it?

Weight matrices appear both as $W$ and as $W^\top$ (forward and backward passes, tied embeddings). We must know how a transpose changes the gradient, or the shapes will not match.

Where is it used?

Backpropagation ($W^\top$ carries gradients backwards), weight tying in language models (the same matrix used as $E$ and $E^\top$), covariance expressions $X^\top X$ and $XX^\top$, and PCA.

How is it used?

Bring any $dX^\top$ to $dX$ by transposing the whole (scalar) term, then read off as usual. Then check the shape. A wrongly transposed gradient usually has the wrong shape for non-square $X$.

Pick an example and step through it. Watch for the moment where a transpose is used to flip a term so that $dX$ (not $dX^\top$) appears.

Choose an identity, edit the entries, and compare the formula with the brute-force gradient. Look at the pairs: $\mathbf{a}^\top X\mathbf{b}$ against $\mathbf{a}^\top X^\top\mathbf{b}$, and $\operatorname{tr}(X^\top A)$ against $\operatorname{tr}(XA)$. In each pair the answers are transposes of one another.

  • Transposing a number changes nothing. Transposing a matrix changes its shape. Only use "a scalar equals its transpose" on genuinely $1\times1$ expressions.
  • $(PQ)^\top=Q^\top P^\top$: the order reverses. Forgetting this is the most common error in this section.
Quick check: $\nabla_X\,\mathbf{u}^\top X\mathbf{v}$ and $\nabla_X\,\mathbf{v}^\top X^\top\mathbf{u}$. Are the answers equal?

The first is $\mathbf{u}\mathbf{v}^\top$. The second is a number equal to its transpose $\mathbf{u}^\top X\mathbf{v}$, so it is the same function, and its gradient is also $\mathbf{u}\mathbf{v}^\top$. Yes, equal.

Trace identities: $\operatorname{tr}(AX)$, $\operatorname{tr}(X^\top AX)$, $\operatorname{tr}(AXB)$ core

A trace of a matrix expression is just a long weighted sum of the entries of $X$, so its gradient is "the weights, in the shape of $X$".

  • $\operatorname{tr}(AX)$ and $\operatorname{tr}(AXB)$ are linear in $X$ (a weighted sum).
  • $\operatorname{tr}(X^\top AX)$ is quadratic. A nice way to see it: the columns of $X$ are vectors $\mathbf{x}_1,\mathbf{x}_2,\dots$, and $\operatorname{tr}(X^\top AX)=\sum_k\mathbf{x}_k^\top A\,\mathbf{x}_k$: a quadratic form for each column. By the vector rule, each column contributes $(A+A^\top)\mathbf{x}_k$, so the whole gradient is $(A+A^\top)X$.

$A=\begin{bmatrix}1&2\\0&3\end{bmatrix}$ and $X=\begin{bmatrix}1&0\\1&1\end{bmatrix}$, with columns $\mathbf{x}_1=[1,1]^\top$ and $\mathbf{x}_2=[0,1]^\top$.

  1. $\mathbf{x}_1^\top A\mathbf{x}_1=1+2+0+3=6$ and $\mathbf{x}_2^\top A\mathbf{x}_2=A_{22}=3$. So $\operatorname{tr}(X^\top AX)=9$.
  2. $A+A^\top=\begin{bmatrix}2&2\\2&6\end{bmatrix}$. Column 1: $(A+A^\top)\mathbf{x}_1=[4,8]^\top$. Column 2: $(A+A^\top)\mathbf{x}_2=[2,6]^\top$.
  3. So $\nabla_X\operatorname{tr}(X^\top AX)=\begin{bmatrix}4&2\\8&6\end{bmatrix}=(A+A^\top)X$ ✓.
  4. $\operatorname{tr}(AXB)$. With $A=\begin{bmatrix}1&0\\1&1\end{bmatrix}$ and $B=\begin{bmatrix}2&1\\0&1\end{bmatrix}$ the gradient is $A^\top B^\top=\begin{bmatrix}1&1\\0&1\end{bmatrix}\begin{bmatrix}2&0\\1&1\end{bmatrix}=\begin{bmatrix}3&1\\1&1\end{bmatrix}$.
Function$\nabla_X$Derivation (the recipe)
$\operatorname{tr}(AX)$$A^\top$$df=\operatorname{tr}(A\,dX)$
$\operatorname{tr}(AXB)$$A^\top B^\top$$df=\operatorname{tr}(A\,dX\,B)=\operatorname{tr}(BA\,dX)$, so $G^\top=BA$
$\operatorname{tr}(AX^\top B)$$BA$rotate: $\operatorname{tr}(AX^\top B)=\operatorname{tr}(BA\,X^\top)=\sum_{ij}(BA)_{ij}X_{ij}$, so the weights are $BA$
$\operatorname{tr}(X^\top AX)$$(A+A^\top)X$ (or $2AX$ if $A=A^\top$)$df=\operatorname{tr}(dX^\top AX)+\operatorname{tr}(X^\top A\,dX)$, transpose the first
$\operatorname{tr}(XAX^\top)$$X(A+A^\top)$see the recipe widget

Shape check for all: the result has the shape of $X$. Rule of thumb: write the function as $\sum(\text{weight})\cdot X_{ij}$ and put the weights into the grid of $X$.

Why do we need it?

Many matrix objectives are naturally written as traces: a sum of quadratic forms over all columns, a sum of inner products, a covariance term. These identities give their gradients at once.

Where is it used?

PCA (maximise $\operatorname{tr}(W^\top\Sigma W)$), matrix factorisation and recommender systems, canonical correlation analysis, the Gaussian log-likelihood (the $\operatorname{tr}(\Sigma^{-1}S)$ term), and metric learning.

How is it used?

Write the objective as a trace, apply the matching row of the table (or run the recipe), and update $X\leftarrow X-\eta\,G$. Always check the shape of $G$ and compare with finite differences.

Step through each identity. Notice how the same moves (differential, rotate, transpose, collect) are used each time, only in different combinations.

Each column $\mathbf{x}_k$ of $X$ gives a quadratic form $\mathbf{x}_k^\top A\mathbf{x}_k$. Their sum is the trace. The gradient of column $k$ is $(A+A^\top)\mathbf{x}_k$, which is exactly column $k$ of the matrix gradient $(A+A^\top)X$. Edit $A$ and $X$ and check that the two sides always agree.

Pick an identity and edit the numbers: the formula and the brute-force gradient must agree. In $\operatorname{tr}(AXB)$ use matrices where $A$ and $B$ are not equal and not symmetric. Try to guess what happens to $\nabla\operatorname{tr}(X^\top AX)$ if you make $A$ symmetric.

  • $\nabla\operatorname{tr}(AXB)=A^\top B^\top$, not $AB$ and not $BA$. The order and both transposes matter. For non-square $A$, $B$ the shape check catches most slips.
  • You may rotate $\operatorname{tr}(AXB)=\operatorname{tr}(BAX)$, but not $\operatorname{tr}(ABX)$.
  • The "$(A+A^\top)X$" identity needs a square $A$. If $A$ is symmetric it reduces to $2AX$.
Quick check: $X$ is $3\times2$, $A$ is $3\times3$ and symmetric. What is $\nabla_X\operatorname{tr}(X^\top AX)$ and its shape?

$2AX$, shape $(3\times3)(3\times2)=3\times2$, the shape of $X$.

Symmetric-matrix identities core

A symmetric matrix is a table that looks the same when you flip it over its diagonal: $A=A^\top$. Many matrices in ML are symmetric: covariance matrices, Gram matrices $X^\top X$, Hessians, graph Laplacians.

Why does symmetry help? In $\mathbf{x}^\top A\mathbf{x}=\sum A_{ij}x_ix_j$ the product $x_ix_j$ is the same as $x_jx_i$. Only the sum $A_{ij}+A_{ji}$ matters. So any matrix $A$ can be replaced by its symmetric part $S=\tfrac12(A+A^\top)$ without changing the function at all. The leftover (the antisymmetric part) is invisible. After that replacement the gradient is the simple $2S\mathbf{x}$.

$A=\begin{bmatrix}1&3\\1&2\end{bmatrix}$ and $\mathbf{x}=[1,2]^\top$.

  1. Split: $S=\tfrac12(A+A^\top)=\begin{bmatrix}1&2\\2&2\end{bmatrix}$ and $K=\tfrac12(A-A^\top)=\begin{bmatrix}0&1\\-1&0\end{bmatrix}$. Check: $S+K=A$ ✓.
  2. $\mathbf{x}^\top A\mathbf{x}=1(1+6)+2(1+4)=17$. $\ \mathbf{x}^\top S\mathbf{x}=1(1+4)+2(2+4)=17$. $\ \mathbf{x}^\top K\mathbf{x}=1(2)+2(-1)=0$.
  3. So the antisymmetric part contributes nothing: $17=17+0$.
  4. Gradient: $(A+A^\top)\mathbf{x}=[10,12]^\top$ and $2S\mathbf{x}=2[5,6]^\top=[10,12]^\top$ ✓. But $2A\mathbf{x}=[14,10]^\top$ ✗.

Every square matrix splits uniquely as $A=S+K$ with $S=\tfrac12(A+A^\top)$ (symmetric) and $K=\tfrac12(A-A^\top)$ (antisymmetric: $K^\top=-K$).

Why $\mathbf{x}^\top K\mathbf{x}=0$: it is a number, so it equals its transpose: $\mathbf{x}^\top K\mathbf{x}=\mathbf{x}^\top K^\top\mathbf{x}=-\mathbf{x}^\top K\mathbf{x}$. A number equal to its own negative is $0$.

Identities for symmetric $A=A^\top$:

$$\nabla(\mathbf{x}^\top A\mathbf{x})=2A\mathbf{x},\quad \nabla^2(\mathbf{x}^\top A\mathbf{x})=2A,\quad \nabla_X\operatorname{tr}(X^\top AX)=2AX,\quad \nabla\|B\mathbf{x}\|^2=2B^\top B\mathbf{x}.$$

The last uses that $B^\top B$ is symmetric for any matrix $B$. For a general (non-symmetric) $A$: use $S$ instead of $A$, i.e. $\nabla(\mathbf{x}^\top A\mathbf{x})=2S\mathbf{x}$.

Why do we need it?

It removes the awkward "$A+A^\top$" and gives the same simple form as one-variable calculus. It also tells us that when a quadratic form is learned from data, only the symmetric part is identifiable.

Where is it used?

Covariance and precision matrices in Gaussians, PCA ($\mathbf{w}^\top\Sigma\mathbf{w}$), the Gram matrix $X^\top X$ in regression, kernel methods, and Hessians (always symmetric).

How is it used?

Check whether the middle matrix is symmetric. If it is, use $2A\mathbf{x}$. If it is not, symmetrise it first: $S=\tfrac12(A+A^\top)$, or use $(A+A^\top)\mathbf{x}$.

Step through the argument: split $A$, show the antisymmetric part vanishes, and arrive at the gradient.

The dark curves are level curves of $\mathbf{x}^\top A\mathbf{x}$. The orange curves are those of $\mathbf{x}^\top S\mathbf{x}$ with $S$ the symmetric part. They coincide, however lopsided $A$ is. Drag the point: the green arrow is the true gradient $(A+A^\top)\mathbf{x}=2S\mathbf{x}$ and the red arrow is the naive $2A\mathbf{x}$. They differ unless $A$ is symmetric. Press Symmetrise A and the red arrow lands on the green one.

  • Symmetry of $A$ is an assumption you must check, not something you can assume. $W$ matrices in neural networks, for example, are usually not symmetric.
  • Replacing $A$ by $S$ is only valid for the quadratic form $\mathbf{x}^\top A\mathbf{x}$ (same vector on both sides). It is not valid for $\mathbf{a}^\top A\mathbf{b}$ with two different vectors.
Quick check: $A=\begin{bmatrix}0&5\\-5&0\end{bmatrix}$ (antisymmetric). What is $\nabla(\mathbf{x}^\top A\mathbf{x})$?

$A+A^\top=\mathbf{0}$, so the gradient is $\mathbf{0}$ everywhere. Indeed $\mathbf{x}^\top A\mathbf{x}=5x_1x_2-5x_2x_1=0$ for every $\mathbf{x}$: the function is identically zero.

Vector norm derivatives: $L_2$, squared $L_2$, $L_1$, $L_\infty$ core

A norm measures size (Linear Algebra guide, Chapter 1.2). Its gradient answers: "which way should I move to make the size grow fastest?"

  • $L_2$ (length). Walking straight away from the origin lengthens the vector at rate $1$. Walking around a circle does not change the length at all. So the gradient is the unit vector pointing away from the origin: $\mathbf{x}/\|\mathbf{x}\|$.
  • Squared $L_2$. The squared length grows faster the further out you are: slope $2\times$ distance. The gradient is $2\mathbf{x}$.
  • $L_1$ (taxi size). Each entry adds $|x_i|$. Each one has slope $+1$ on the right of $0$ and $-1$ on the left. At exactly $0$ there is a sharp corner (a kink): no single slope fits.
  • $L_\infty$ (largest entry). Only the biggest entry counts, so the gradient is a single $\pm1$ in that position and $0$ elsewhere (with a kink when two entries tie).
  1. $L_2$: $\mathbf{x}=[3,4]^\top$. $\|\mathbf{x}\|=5$ and $\nabla\|\mathbf{x}\|=[0.6,\ 0.8]^\top$ (a unit vector). Check: nudging $x_1$ by $0.01$ gives $\|[3.01,4]\|=5.0060$, a change of $0.0060\approx0.6\times0.01$ ✓.
  2. Squared $L_2$: $\|\mathbf{x}\|^2=25$ and $\nabla=[6,8]^\top$ (length 10 = twice the distance).
  3. $L_1$: $\mathbf{x}=[3,-2,0]^\top$. $\|\mathbf{x}\|_1=5$. A subgradient is $[1,\ -1,\ s]$ for any $s\in[-1,1]$ (the entry at $0$ is a kink).
  4. $L_\infty$: $\mathbf{x}=[3,-4,1]^\top$. The largest absolute entry is $|-4|$, so $\|\mathbf{x}\|_\infty=4$ and the gradient is $[0,\ -1,\ 0]^\top$ (moving $x_2$ more negative increases the max).
$$\nabla\|\mathbf{x}\|_2=\frac{\mathbf{x}}{\|\mathbf{x}\|_2}\ (\mathbf{x}\ne\mathbf{0}),\qquad \nabla\|\mathbf{x}\|_2^2=2\mathbf{x},\qquad \nabla\|\mathbf{x}-\mathbf{c}\|_2=\frac{\mathbf{x}-\mathbf{c}}{\|\mathbf{x}-\mathbf{c}\|_2}.$$ $$\nabla\|\mathbf{x}\|_1=\operatorname{sign}(\mathbf{x})\ \text{(a subgradient)},\qquad \nabla\|\mathbf{x}\|_\infty=\operatorname{sign}(x_k)\,\mathbf{e}_k\ \text{ where }k=\arg\max_i|x_i|\ \text{(unique max)}.$$

At $\mathbf{x}=\mathbf{0}$ the $L_2$ norm has a cone-shaped tip, so its gradient is not defined there.

Subgradients (awareness). At a kink the ordinary derivative does not exist. For a convex function, a subgradient at $\mathbf{x}$ is any vector $\mathbf{g}$ such that the straight line (plane) $f(\mathbf{x})+\mathbf{g}^\top(\mathbf{y}-\mathbf{x})$ stays below the graph everywhere. The set of all of them is the subdifferential. For $|x|$ at $0$ it is the whole interval $[-1,1]$. In code one usually picks $0$: that is what np.sign(0) returns. For $L_\infty$ at a tie, any weighted average of the tied $\pm\mathbf{e}_i$ is a subgradient.

Other $L_p$ norms ($p\gt1$) are smooth away from $0$: $\nabla\|\mathbf{x}\|_p^p=p\,|\mathbf{x}|^{p-1}\circ\operatorname{sign}(\mathbf{x})$ (entrywise).

Why do we need it?

Penalties and constraints are written with norms. To train with them, or to normalise by them, we need their gradients, including the awkward non-smooth cases.

Where is it used?

L1 regularisation (Lasso) for sparse weights, L2 weight decay, gradient clipping by norm, normalisation layers (dividing by $\|\mathbf{x}\|$), cosine similarity, and adversarial attacks (FGSM steps along $\operatorname{sign}(\nabla)$, the best move inside an $L_\infty$ budget).

How is it used?

Add the gradient of the penalty to the loss gradient: $\lambda\operatorname{sign}(\mathbf{w})$ for L1, $2\lambda\mathbf{w}$ for squared L2. Frameworks use $0$ at the kink. Divide by $\|\mathbf{x}\|$ only after checking it is not (nearly) zero.

Step through each derivation. In the $L_1$ derivation, notice the exact moment where the ordinary derivative stops existing.

Choose a norm. The curves are its level sets (the "unit balls" at different sizes). Drag the blue point: the green arrow is the gradient. For the $L_2$ and $L_\infty$ norms its length is always $1$ (for $L_1$ it is $\sqrt2$ off the axes), but for $\frac12\|\mathbf{x}\|^2$ it shrinks as you approach the origin. Press Take a gradient step a few times with the same step size: the squared norm shrinks smoothly, but the others keep bouncing around $0$ because their slope never gets smaller. Put the point on an axis for $L_1$ to see a kink.

Drag the point along the graph (it snaps to steps of $0.25$ so that you can land exactly on $0$). For the smooth $x^2$ there is one tangent line. For $|x|$ and ReLU, at $x=0$ the single tangent is replaced by a fan of lines, all lying below the graph: these are the subgradients. Away from $0$ the fan collapses to the ordinary tangent.

Rotate with the background and use Top and Side. Drag the blue point on the surface (it slides along it). The green arrow on the floor is the gradient (steepest uphill direction). The $L_2$ norm is a cone with a sharp tip, the $L_1$ norm a pyramid with creases along the axes, $L_\infty$ a flat-sided pyramid with creases on the diagonals, and $\tfrac12\|\mathbf{x}\|^2$ a smooth bowl. Gradients are undefined on the creases.

  • $\nabla\|\mathbf{x}\|_2=\mathbf{x}/\|\mathbf{x}\|$ is not defined at $\mathbf{0}$. In code add a tiny constant, $\mathbf{x}/(\|\mathbf{x}\|+\varepsilon)$, or use $\|\mathbf{x}\|^2$ if you can.
  • $\operatorname{sign}(\mathbf{x})$ is a subgradient, not "the" derivative: do not use it in a finite-difference check at points where an entry is exactly $0$.
  • $L_1$ and $L_\infty$ gradients keep a constant size however close you are. Optimisers need shrinking steps (or a "proximal" step) to settle at the minimum.
Quick check: what is $\nabla\,\lambda\|\mathbf{w}\|_1$ at $\mathbf{w}=[2,-3,0,5]^\top$, using $\operatorname{sign}(0)=0$?

$\lambda\operatorname{sign}(\mathbf{w})=\lambda[1,-1,0,1]^\top$. Every non-zero weight is pushed towards $0$ by the same amount $\lambda$ per step, regardless of its size. (Squared L2 would push big weights harder.)

Matrix norm derivatives: Frobenius, and a glimpse of nuclear and spectral core

The Frobenius norm treats a matrix as one long vector: square every entry, add, take the square root. So everything we learned about the $L_2$ norm carries over, with "entries of $X$" in place of "entries of $\mathbf{x}$".

The two other matrix norms you will hear about are based on singular values (Chapter 1.13 of the Linear Algebra guide): the spectral norm is the largest singular value (the biggest stretch the matrix makes), and the nuclear norm is the sum of all singular values (a convex stand-in for the rank, just as $L_1$ is the stand-in for counting non-zero entries). They are the matrix versions of $\max$ and of $L_1$, and their gradients are built from the SVD.

  1. $X=\begin{bmatrix}1&2\\3&4\end{bmatrix}$. $\|X\|_F^2=1+4+9+16=30$. $\nabla\|X\|_F^2=2X=\begin{bmatrix}2&4\\6&8\end{bmatrix}$. And $\|X\|_F=\sqrt{30}\approx5.477$, with gradient $X/\|X\|_F$.
  2. Matrix least squares. $A=\begin{bmatrix}1&0\\0&2\end{bmatrix}$, $X=\begin{bmatrix}1&1\\1&1\end{bmatrix}$, $B=\begin{bmatrix}0&1\\1&0\end{bmatrix}$. Then $AX=\begin{bmatrix}1&1\\2&2\end{bmatrix}$ and the residual $R=AX-B=\begin{bmatrix}1&0\\1&2\end{bmatrix}$. Loss: $1+0+1+4=6$. Gradient: $2A^\top R=2\begin{bmatrix}1&0\\0&2\end{bmatrix}\begin{bmatrix}1&0\\1&2\end{bmatrix}=\begin{bmatrix}2&0\\4&8\end{bmatrix}$.
  3. Awareness. $X=\operatorname{diag}(3,1)$ has singular values $3,1$ and $U=V=I$. The spectral norm is $3$, with gradient $\mathbf{u}_1\mathbf{v}_1^\top=\begin{bmatrix}1&0\\0&0\end{bmatrix}$ (only the entry that sets the largest stretch matters). The nuclear norm is $3+1=4$ with gradient $UV^\top=I$. Check: nudging $X_{11}$ by $0.01$ changes $\|X\|_*=3.01+1$ by $0.01$ ✓ (slope $1$).
$$\|X\|_F^2=\sum_{i,j}X_{ij}^2=\operatorname{tr}(X^\top X),\qquad \nabla\|X\|_F^2=2X,\qquad \nabla\|X\|_F=\frac{X}{\|X\|_F}.$$ $$\nabla_X\|AX-B\|_F^2=2A^\top(AX-B),\qquad \nabla_X\|XA-B\|_F^2=2(XA-B)A^\top.$$

Awareness (SVD $X=U\Sigma V^\top$ with singular values $\sigma_1\ge\sigma_2\ge\cdots$). The standard first-order fact is $d\sigma_i=\mathbf{u}_i^\top dX\,\mathbf{v}_i$. From it:

$$\nabla\|X\|_2=\mathbf{u}_1\mathbf{v}_1^\top\ \ (\text{spectral norm }\sigma_1,\ \text{if }\sigma_1\gt\sigma_2),\qquad \nabla\|X\|_*=UV^\top=\sum_i\mathbf{u}_i\mathbf{v}_i^\top\ \ (\text{nuclear norm }\textstyle\sum\sigma_i,\ \text{if all }\sigma_i\gt0).$$

$UV^\top$ is the matrix version of $\operatorname{sign}(x)$: it keeps the directions and sets every stretch to 1. Where singular values tie or hit $0$, these norms have kinks, and the formulas give just one subgradient.

Why do we need it?

Weights of layers are matrices, and we often want to penalise or control their size. Each matrix norm controls something different, and training needs its gradient.

Where is it used?

Frobenius: weight decay on matrices, multi-output regression $\|XW-Y\|_F^2$, matrix factorisation. Nuclear: low-rank matrix completion and recommender systems. Spectral: spectral normalisation of GAN discriminators and Lipschitz control of networks.

How is it used?

Frobenius: add $2\lambda W$ to the gradient. For the other two, use a library SVD: spectral normalisation estimates $\mathbf{u}_1,\mathbf{v}_1$ with a few power-iteration steps and divides $W$ by $\sigma_1$.

Step through the derivations. The two matrix least-squares versions differ only in which side $A$ multiplies $X$. Compare the position of $A^\top$ in the answers. The last example is the (awareness) singular-value rule.

Pick an identity, edit the matrices, and compare. The two least-squares versions use the same $A$, $B$ and $X$, but $A$ sits on different sides, and the answers differ. (Entering $X=0$ in the $\|X\|_F$ test shows the "not defined at zero" message.)

Edit the $3\times3$ matrix. The widget computes the SVD and compares the formulas $\mathbf{u}_1\mathbf{v}_1^\top$ and $UV^\top$ with brute-force gradients of $\sigma_1$ and of $\sigma_1+\sigma_2+\sigma_3$. Try making two singular values equal (for example $X=I$) or a row zero: then the norm has a kink and the check says so.

  • $\nabla\|X\|_F^2=2X$, but $\nabla\|X\|_F=X/\|X\|_F$. As with vectors, do not mix the squared and un-squared versions.
  • $\|AX-B\|_F^2$ gives $2A^\top(AX-B)$; for $\|XA-B\|_F^2$ the $A^\top$ moves to the right: $2(XA-B)A^\top$. The side where $A$ multiplies $X$ decides where $A^\top$ goes.
  • The SVD formulas are only valid when the relevant singular values are distinct (spectral) or positive (nuclear). Otherwise they are subgradients.
Quick check: $\nabla_W\big(\|W\|_F^2\cdot\tfrac\lambda2\big)$ for a weight matrix $W$?

$\tfrac\lambda2\cdot2W=\lambda W$. This is weight decay applied to a whole matrix.

The log-determinant and the inverse (awareness) core

The determinant $\det X$ says by what factor the matrix $X$ scales volume (area in 2D; Chapter 1.7 of the Linear Algebra guide). Its logarithm turns products into sums, which is why $\log\det$ appears so often.

How does $\log\det X$ react to a nudge $dX$? In one variable: $d\ln x=dx/x=x^{-1}dx$, the nudge measured in units of $x$. The matrix version is the same: $d\log\det X=\operatorname{tr}(X^{-1}dX)$, the nudge measured in units of $X$. The inverse is the matrix version of "divide by".

$X=\begin{bmatrix}2&1\\1&3\end{bmatrix}$. $\det X=2\cdot3-1\cdot1=5$ and $X^{-1}=\tfrac15\begin{bmatrix}3&-1\\-1&2\end{bmatrix}=\begin{bmatrix}0.6&-0.2\\-0.2&0.4\end{bmatrix}$.

  1. $\log\det X=\ln(x_{11}x_{22}-x_{12}x_{21})$. By the chain rule, $\partial/\partial x_{11}=x_{22}/\det=3/5=0.6$; $\ \partial/\partial x_{12}=-x_{21}/\det=-0.2$; $\ \partial/\partial x_{21}=-x_{12}/\det=-0.2$; $\ \partial/\partial x_{22}=x_{11}/\det=0.4$.
  2. So the gradient is $\begin{bmatrix}0.6&-0.2\\-0.2&0.4\end{bmatrix}=X^{-\top}$ ✓ (here equal to $X^{-1}$ because $X$ is symmetric).
  3. The inverse. Nudge $x_{11}$ by $\varepsilon$: the exact inverse of $\begin{bmatrix}2+\varepsilon&1\\1&3\end{bmatrix}$ has top-left entry $\frac{3}{5+3\varepsilon}\approx0.6-0.36\,\varepsilon$. The formula $-X^{-1}\,dX\,X^{-1}$ with $dX=\varepsilon E_{11}$ gives $-\varepsilon\,(0.6)(0.6)=-0.36\,\varepsilon$ in that entry ✓.
$$d\det X=\det X\cdot\operatorname{tr}(X^{-1}dX),\qquad d\log|\det X|=\operatorname{tr}(X^{-1}dX),\qquad d(X^{-1})=-X^{-1}\,dX\,X^{-1}.$$ $$\nabla_X\log|\det X|=X^{-\top},\qquad \nabla_X\det X=\det X\cdot X^{-\top},\qquad \nabla_X\operatorname{tr}(AX^{-1})=-X^{-\top}A^\top X^{-\top}.$$

Here $X^{-\top}$ means $(X^{-1})^\top$. For symmetric $X$ (a covariance matrix) all the transposes disappear.

Where the first formula comes from (awareness). For a tiny nudge $\varepsilon E$: $\det(X+\varepsilon E)=\det X\cdot\det(I+\varepsilon M)$ with $M=X^{-1}E$. The determinant of $I+\varepsilon M$ is the product of $(1+\varepsilon\lambda_i)$ over the eigenvalues $\lambda_i$ of $M$. Multiplying out and dropping $\varepsilon^2$ gives $1+\varepsilon\sum\lambda_i=1+\varepsilon\operatorname{tr}M$. So $d\det X=\det X\operatorname{tr}(X^{-1}E)\,\varepsilon$.

Why do we need it?

Gaussian probability densities contain $\log\det\Sigma$ and $\Sigma^{-1}$. To fit them, or to train models that use them, we need their gradients.

Where is it used?

Maximum-likelihood estimation of a covariance, Gaussian processes (the $\log\det K$ term), graphical lasso, normalising flows ($\log|\det J|$ of the Jacobian, Chapter 2.5), and Bayesian model evidence.

How is it used?

Use $\nabla\log\det X=X^{-\top}$ and $d(X^{-1})=-X^{-1}dX\,X^{-1}$ as building blocks in the recipe. In code, never invert: use a Cholesky factorisation and solve. Compute $\log\det$ as twice the sum of logs of the Cholesky diagonal.

Step through each derivation. The last one is a payoff: it uses three identities to prove, in a few lines, that the best Gaussian covariance is the sample covariance.

Start from the matrix $X_0$ and move in the direction $E$: $X(t)=X_0+tE$. The curve is $\log|\det X(t)|$. At $t=0$ the orange tangent has slope $\operatorname{tr}(X_0^{-1}E)$. The readout compares it with the numerical slope of the curve. Try $E=I$ (the curve is a sum of logs), then $E=\begin{bmatrix}0&0\\0&1\end{bmatrix}$. The curve plunges to $-\infty$ where $X(t)$ becomes singular (determinant $0$): that is where $\log\det$ is not defined.

Pick an identity and edit $X$ (and $A$). The formula and the brute-force gradient should agree whenever $X$ is invertible. Make $X$ singular (for example, a zero row) and see the message. Note how the transposes appear: use a non-symmetric $X$.

  • $\nabla\log\det X=X^{-\top}$ needs $X$ invertible ($\det X\ne0$). For the plain $\log$ one also needs $\det X\gt0$; with $\log|\det X|$ the formula holds for any invertible $X$.
  • $d(X^{-1})=-X^{-1}dX\,X^{-1}$ has two inverses, one on each side. It is not $-X^{-2}dX$ (that would be right only if $X$ and $dX$ commuted).
  • The transposes matter for non-symmetric $X$. For a symmetric covariance matrix, they all vanish.
Quick check: $X=\operatorname{diag}(a,b)$. What is $\nabla_X\log\det X$ at the diagonal entries?

$\log\det X=\ln a+\ln b$. The slopes with respect to the diagonal entries are $1/a$ and $1/b$, which are the diagonal entries of $X^{-\top}=\operatorname{diag}(1/a,1/b)$ ✓.

Chain-rule identities: $\nabla f(A\mathbf{x})=A^\top\nabla f$ core

Most functions in machine learning are compositions: do something to $\mathbf{x}$ to get $\mathbf{u}$, then feed $\mathbf{u}$ into a scalar function $f$. A nudge $d\mathbf{x}$ makes $\mathbf{u}$ move by $d\mathbf{u}=J\,d\mathbf{x}$, and that makes $f$ move by $\nabla f^\top d\mathbf{u}$. Chain the two: $df=\nabla f^\top J\,d\mathbf{x}$.

Reading off, the gradient is $J^\top\nabla f$. The transpose is the heart of backpropagation: a gradient measured at the output, $\nabla f$, travels backwards through the inner function by multiplying with $J^\top$. If the inner function is a matrix, $\mathbf{u}=A\mathbf{x}$, then $J=A$ and the backward step is multiplication by $A^\top$.

  1. Linear inner function. $f(\mathbf{u})=\tfrac12\|\mathbf{u}\|^2$ and $A=\begin{bmatrix}1&2\\3&0\\0&1\end{bmatrix}$, $\mathbf{x}=[1,1]^\top$. Then $\mathbf{u}=A\mathbf{x}=[3,3,1]^\top$, $\nabla f(\mathbf{u})=\mathbf{u}$, and $\nabla_{\mathbf{x}}f=A^\top\mathbf{u}=[1\cdot3+3\cdot3,\ \ 2\cdot3+1\cdot1]^\top=[12,\ 7]^\top$. Direct check: $f=\tfrac12\big[(x_1+2x_2)^2+9x_1^2+x_2^2\big]$, so $\partial_1f=(x_1+2x_2)+9x_1=3+9=12$ and $\partial_2f=2(x_1+2x_2)+x_2=6+1=7$ ✓.
  2. Nonlinear inner function. $\mathbf{g}(\mathbf{x})=(x_1x_2,\ x_1+x_2^2)$ and $f(\mathbf{u})=u_1^2+u_2$, at $\mathbf{x}=(2,1)$. Then $\mathbf{u}=(2,3)$, $\nabla f=[2u_1,\,1]=[4,1]^\top$, and $J_{\mathbf{g}}=\begin{bmatrix}x_2&x_1\\1&2x_2\end{bmatrix}=\begin{bmatrix}1&2\\1&2\end{bmatrix}$. So $J_{\mathbf{g}}^\top\nabla f=\begin{bmatrix}1&1\\2&2\end{bmatrix}\begin{bmatrix}4\\1\end{bmatrix}=[5,\ 10]^\top$. Direct check: $f\circ\mathbf{g}=(x_1x_2)^2+x_1+x_2^2$ gives $\partial_1=2x_1x_2^2+1=5$ and $\partial_2=2x_1^2x_2+2x_2=10$ ✓. (The wrong order $J\nabla f=[6,6]$ would fail.)
  3. Scalar outer function. $f(\mathbf{x})=e^{-\frac12\mathbf{x}^\top\mathbf{x}}$ (a Gaussian bump). Outer $h(q)=e^{-q}$ with $q=\tfrac12\|\mathbf{x}\|^2$, $\nabla q=\mathbf{x}$: $\nabla f=-e^{-q}\mathbf{x}$.

Let $f$ be a scalar function, $A$ a matrix, $\mathbf{g}$ a vector function with Jacobian $J_{\mathbf{g}}$:

$$\nabla_{\mathbf{x}}\,f(A\mathbf{x}+\mathbf{b})=A^\top\,\nabla f(A\mathbf{x}+\mathbf{b}),\qquad \nabla_{\mathbf{x}}\,f(\mathbf{g}(\mathbf{x}))=J_{\mathbf{g}}(\mathbf{x})^\top\,\nabla f(\mathbf{g}(\mathbf{x})),$$ $$\nabla_W\,f(W\mathbf{x}+\mathbf{b})=\nabla f\;\mathbf{x}^\top,\qquad \nabla_{\mathbf{b}}\,f(W\mathbf{x}+\mathbf{b})=\nabla f,\qquad \nabla_{\mathbf{x}}\,h(q(\mathbf{x}))=h'(q)\,\nabla q.$$

Shapes: $A^\top$ is $n\times m$ and $\nabla f$ is $m\times1$, so the result is $n\times1$ (the shape of $\mathbf{x}$). $\nabla f\,\mathbf{x}^\top$ is $(m\times1)(1\times n)=m\times n$ (the shape of $W$). Awareness: the Hessian chain rule (Chapter 2.10) gives $\nabla^2_{\mathbf{x}}f(A\mathbf{x})=A^\top(\nabla^2f)\,A$.

Why do we need it?

Nearly every loss is a function of a function of the weights. The chain-rule identities turn "differentiate a big composition" into "multiply a few simple pieces".

Where is it used?

Logistic regression ($f$ applied to $X\mathbf{w}$), every layer of a neural network, backpropagation, Gaussian densities, and any loss written as $\ell(X\mathbf{w})$ on a data matrix $X$.

How is it used?

Differentiate the outer function with respect to its input, evaluated at the inner output. Then multiply by $J^\top$ of the inner function ($A^\top$ if it is linear). For a weight matrix use the outer product $\nabla f\,\mathbf{x}^\top$.

Step through each derivation. The last example, logistic regression, uses the identity at full strength: one line gives the gradient of the whole data-set loss.

Choose an outer function $f$ and an inner function $\mathbf{u}$. The widget computes $\mathbf{u}$, then $\nabla f(\mathbf{u})$, then the Jacobian $J$, and finally $J^\top\nabla f$. It compares with the brute-force gradient of the whole composition. The line wrong order shows $J\nabla f$: because $J$ is square here, its shape fits, but the numbers are wrong, so only the numerical check can catch it. The last choice differentiates with respect to the matrix $W$.

  • The transpose is essential: it is $J^\top\nabla f$, not $J\nabla f$. When $J$ is square the shapes still fit, so a shape check cannot save you. Use the numerical check.
  • $\nabla f$ must be evaluated at the inner output $\mathbf{u}=\mathbf{g}(\mathbf{x})$, not at $\mathbf{x}$.
  • For $\nabla_W f(W\mathbf{x})$ the answer is an outer product (a matrix with the shape of $W$), not a dot product.
Quick check: $\mathbf{x}\in\mathbb{R}^5$, $A$ is $3\times5$, $f$ maps $\mathbb{R}^3\to\mathbb{R}$. What is the shape of $\nabla_{\mathbf{x}}f(A\mathbf{x})$ and what is its formula?

$A^\top\nabla f(A\mathbf{x})$: $(5\times3)(3\times1)=5\times1$, the shape of $\mathbf{x}$.

The one-page identity table, and how to check any gradient numerically core

Here is everything in one place. Use it like a dictionary: look up a pattern, but also know that each line is a few steps from the recipe. And whether you looked it up or derived it, always run the check at the bottom of this section. A gradient formula is a claim, and a numerical test is a quick way to prove the claim wrong, or to believe it.

The test is very simple: nudge each input up and down by a tiny amount, and see if the output changes at the rate your formula says.

Test the claim $\nabla(\mathbf{x}^\top A\mathbf{x})=(A+A^\top)\mathbf{x}$ with $A=\begin{bmatrix}2&1\\0&3\end{bmatrix}$ at $\mathbf{x}=[1,2]^\top$, using $h=10^{-5}$.

  1. Formula: $(A+A^\top)\mathbf{x}=[6,13]^\top$.
  2. Nudge $x_1$: $f=2x_1^2+x_1x_2+3x_2^2$, so $f(1.00001,2)=16.0000600002$ and $f(0.99999,2)=15.9999400002$. The slope is $(16.0000600002-15.9999400002)/(2\times10^{-5})=6.0000$ ✓.
  3. Nudge $x_2$ the same way: $13.0000$ ✓. The relative error is around $10^{-10}$.
  4. The wrong guess $2A\mathbf{x}=[8,12]^\top$ differs from $[6,13]$ by $2$: the check catches it at once.

The identity table. $\mathbf{x}\in\mathbb{R}^n$; $A,B$ constant matrices; $X$ a matrix variable; $\nabla_X$ has the shape of $X$. Each "key step" is the recipe's central move.

FunctionGradientKey step
$\mathbf{a}^\top\mathbf{x}$$\mathbf{a}$$df=\mathbf{a}^\top d\mathbf{x}$
$A\mathbf{x}$ (Jacobian)$A$$d\mathbf{y}=A\,d\mathbf{x}$
$\|\mathbf{x}\|^2$$2\mathbf{x}$number equals its transpose
$\mathbf{x}^\top A\mathbf{x}$$(A+A^\top)\mathbf{x}$; $2A\mathbf{x}$ if symmetricproduct rule, flip one term
$\tfrac12\mathbf{x}^\top A\mathbf{x}+\mathbf{b}^\top\mathbf{x}$$\tfrac12(A+A^\top)\mathbf{x}+\mathbf{b}$sum of the two rows above
$\|A\mathbf{x}-\mathbf{b}\|^2$$2A^\top(A\mathbf{x}-\mathbf{b})$$d\mathbf{r}=A\,d\mathbf{x}$, $df=2\mathbf{r}^\top d\mathbf{r}$
$\|\mathbf{x}\|_2$ / $\|\mathbf{x}\|_1$ / $\|\mathbf{x}\|_\infty$$\mathbf{x}/\|\mathbf{x}\|$ / $\operatorname{sign}(\mathbf{x})$ / $\operatorname{sign}(x_k)\mathbf{e}_k$chain rule / kink: subgradient / only the max counts
$\sum\ln x_i$, $\ \sum e^{x_i}$, $\ \ln\sum e^{x_i}$$1/\mathbf{x}$, $\ e^{\mathbf{x}}$, $\ \operatorname{softmax}(\mathbf{x})$entrywise; chain rule on $\ln S$
$f(A\mathbf{x}+\mathbf{b})$ / $f(\mathbf{g}(\mathbf{x}))$$A^\top\nabla f$ / $J_{\mathbf{g}}^\top\nabla f$$df=\nabla f^\top J\,d\mathbf{x}$
$\operatorname{tr}(AX)$ / $\mathbf{a}^\top X\mathbf{b}$$A^\top$ / $\mathbf{a}\mathbf{b}^\top$$\operatorname{tr}(A\,dX)$ / rotate $\mathbf{b}$ to the front
$\operatorname{tr}(X^\top AX)$$(A+A^\top)X$transpose the first term
$\operatorname{tr}(AXB)$$A^\top B^\top$rotate: $\operatorname{tr}(BA\,dX)$
$\|X\|_F^2$ / $\|X\|_F$$2X$ / $X/\|X\|_F$$\operatorname{tr}(X^\top X)$
$\|AX-B\|_F^2$ / $\|XA-B\|_F^2$$2A^\top(AX-B)$ / $2(XA-B)A^\top$$dR=A\,dX$ / $dX\,A$, then rotate
$\|X\|_2$ / $\|X\|_*$ (awareness)$\mathbf{u}_1\mathbf{v}_1^\top$ / $UV^\top$$d\sigma_i=\mathbf{u}_i^\top dX\,\mathbf{v}_i$
$\log\lvert\det X\rvert$ / $\det X$$X^{-\top}$ / $\det X\,X^{-\top}$$d\det X=\det X\operatorname{tr}(X^{-1}dX)$
$\operatorname{tr}(AX^{-1})$$-X^{-\top}A^\top X^{-\top}$$d(X^{-1})=-X^{-1}dX\,X^{-1}$
$f(W\mathbf{x}+\mathbf{b})$ w.r.t. $W$$\nabla f\,\mathbf{x}^\top$$df=\operatorname{tr}(\mathbf{x}\nabla f^\top dW)$

How to check any gradient numerically.

  1. Shape first. The gradient must have the shape of the variable. (Free, instant.)
  2. Pick a random point with non-square, non-symmetric matrices. Avoid special points ($\mathbf{0}$, $I$, kinks like $x_i=0$ for $L_1$).
  3. Central differences: for each entry, $\dfrac{f(\ldots+h\,\mathbf{e}\ldots)-f(\ldots-h\,\mathbf{e}\ldots)}{2h}$ with $h\approx10^{-5}$ (and double-precision numbers). Forward differences are worse: error $\propto h$ instead of $h^2$.
  4. Compare with a relative error: $\dfrac{\max|\text{formula}-\text{numeric}|}{\max(1,\max|\text{formula}|)}$. Below $10^{-6}$: correct. Between $10^{-6}$ and $10^{-3}$: suspicious (try another $h$ or point). Above $10^{-3}$: a bug.
  5. Cheap version for huge models: pick a random direction $\mathbf{v}$ and compare $\nabla f^\top\mathbf{v}$ with the single slope $\frac{f(\mathbf{x}+h\mathbf{v})-f(\mathbf{x}-h\mathbf{v})}{2h}$. Two function evaluations instead of $2n$.

Library versions: scipy.optimize.check_grad, torch.autograd.gradcheck, jax.test_util.check_grads.

Why do we need it?

A wrong gradient does not crash your program: it just trains badly, silently. A ten-second numerical check turns "I think it is right" into "it is right".

Where is it used?

Writing custom layers and losses in PyTorch or JAX, unit tests in ML libraries, reproducing a paper's backward pass, and debugging models whose loss will not go down.

How is it used?

Run the five steps above on a small version of your problem (for example 3 inputs and 2 outputs). When the relative error is below $10^{-6}$ at a few random points, trust the formula and scale up.

Each row is an identity from the table. Press Test to draw fresh random (non-symmetric, non-square) matrices, compute the formula and the brute-force gradient, and report the relative error. Press it again for new random numbers. Use Test all to run every row. Every identity should come out below $10^{-6}$.

Each problem lists the right formula and several tempting wrong ones, all tested at one random point. A wrong formula either has the wrong shape (the shape check catches it) or the right shape but the wrong numbers (only the numerical check catches it). In the matrix problems all matrices are square, so shapes cannot help. Press New random point a few times: the right answer stays at about $10^{-10}$ and the wrong ones stay large.

The curves show the error of a finite-difference derivative as the step $h$ changes (both axes are logarithmic: $10^{-1}$ on the right, $10^{-12}$ on the left). Orange: forward difference, with error that shrinks like $h$. Blue: central difference, with error that shrinks like $h^2$. On the left both curves go up again: round-off noise. Move the slider to see the errors at one $h$. About $10^{-5}$ is the sweet spot for central differences. Pick another function and check that the shape of the picture stays the same.

  • Do not use $h$ smaller than about $10^{-8}$: round-off errors grow. Do not use a huge $h$ either: the slope estimate becomes a poor average.
  • In 32-bit floating point (the default in many deep-learning libraries) use bigger steps (about $10^{-3}$) or, better, run the check in 64-bit.
  • A mismatch at a point where an entry is exactly $0$ (for $L_1$), or where two values tie (for $\max$), is not a bug: the function has a kink there.
Quick check: formula and numerical gradient differ by a factor of exactly $2$ in every entry, and the shape is right. Name two likely causes.

(1) A missing or extra factor $2$, for example using $A^\top(A\mathbf{x}-\mathbf{b})$ for $\|A\mathbf{x}-\mathbf{b}\|^2$. (2) Using $A\mathbf{x}$ instead of $2A\mathbf{x}$ for $\mathbf{x}^\top A\mathbf{x}$ with a symmetric $A$, or $2A\mathbf{x}$ for $\tfrac12\mathbf{x}^\top A\mathbf{x}$. If the formula is off by exactly the same factor $2$ in every entry, look for the missing or extra $\tfrac12$ first.

Recap, cheat sheet and practice

  • The recipe replaces the table: (1) write $df$ with the differential rules, (2) rearrange with the four moves (wrap a scalar in $\operatorname{tr}$, rotate, transpose, collect) until $dX$ is at the right end, (3) read off ($\operatorname{tr}(M\,dX)\Rightarrow\nabla_X=M^\top$, $\ \mathbf{g}^\top d\mathbf{x}\Rightarrow\mathbf{g}$), (4) check shape and numbers.
  • Linear functions give "the weights": $\mathbf{a}$, $A$, $A^\top$ (for $\operatorname{tr}(AX)$), $\mathbf{a}\mathbf{b}^\top$. Quadratics give $(A+A^\top)\mathbf{x}$, and $A\mathbf{x}+\mathbf{b}$ for $\tfrac12\mathbf{x}^\top A\mathbf{x}+\mathbf{b}^\top\mathbf{x}$ with symmetric $A$.
  • Symmetry: $\mathbf{x}^\top A\mathbf{x}$ only sees $\tfrac12(A+A^\top)$, so $\nabla=2S\mathbf{x}$. $\mathbf{x}^\top K\mathbf{x}=0$ for antisymmetric $K$. $B^\top B$ is always symmetric.
  • Trace identities: $\nabla\operatorname{tr}(AXB)=A^\top B^\top$, $\ \nabla\operatorname{tr}(X^\top AX)=(A+A^\top)X$. Transposes flip the pattern of the gradient.
  • Norms: $\nabla\|\mathbf{x}\|_2=\mathbf{x}/\|\mathbf{x}\|$, $\nabla\|\mathbf{x}\|^2=2\mathbf{x}$, $\nabla\|\mathbf{x}\|_1=\operatorname{sign}(\mathbf{x})$ (subgradient at kinks), $\nabla\|X\|_F^2=2X$, $\nabla\|AX-B\|_F^2=2A^\top(AX-B)$. Spectral and nuclear: $\mathbf{u}_1\mathbf{v}_1^\top$ and $UV^\top$.
  • Log-det and inverse: $d\log\det X=\operatorname{tr}(X^{-1}dX)$, $\nabla=X^{-\top}$, $\ d(X^{-1})=-X^{-1}dX\,X^{-1}$.
  • Chain rule: $\nabla_{\mathbf{x}}f(A\mathbf{x})=A^\top\nabla f$, $\ \nabla f(\mathbf{g}(\mathbf{x}))=J_{\mathbf{g}}^\top\nabla f$, $\ \nabla_Wf(W\mathbf{x})=\nabla f\,\mathbf{x}^\top$.
  • Always test: shape first, then central differences at a random point ($h\approx10^{-5}$), relative error below $10^{-6}$.

Cheat sheet: the four moves of the recipe

SituationMoveExample
$dX$ is not at the right endrotate inside a trace$\operatorname{tr}(P\,dX\,Q)=\operatorname{tr}(QP\,dX)$
You have a number, not a tracewrap: $s=\operatorname{tr}(s)$$\mathbf{a}^\top dX\,\mathbf{b}=\operatorname{tr}(\mathbf{b}\mathbf{a}^\top dX)$
You see $dX^\top$transpose the whole term$\operatorname{tr}(dX^\top M)=\operatorname{tr}(M^\top dX)$
Two terms with $dX$ at the endcollect$\operatorname{tr}(P\,dX)+\operatorname{tr}(Q\,dX)=\operatorname{tr}((P+Q)\,dX)$
Final answertranspose the coefficient$df=\operatorname{tr}(M\,dX)\Rightarrow\nabla_Xf=M^\top$
Code it · NumPy

import numpy as np

rng = np.random.default_rng(0)

def num_grad(f, X, h=1e-5):
    """Central-difference gradient of a scalar function f at X (any shape)."""
    G = np.zeros_like(X)
    for idx in np.ndindex(*X.shape):
        E = np.zeros_like(X); E[idx] = h
        G[idx] = (f(X + E) - f(X - E)) / (2 * h)
    return G

def check(name, f, G, X):
    """Compare a formula G(X) with the numerical gradient. Prints True if they agree."""
    a, n = G(X), num_grad(f, X)
    rel = np.abs(a - n).max() / max(1.0, np.abs(a).max())
    print(f"{name:28s}", rel < 1e-6)

A = rng.normal(size=(3, 3)); B = rng.normal(size=(3, 3))     # NOT symmetric
b = rng.normal(size=3); x = rng.normal(size=3); X = rng.normal(size=(3, 3))
Ai = np.linalg.inv

check("x^T A x",        lambda v: v @ A @ v,                     lambda v: (A + A.T) @ v, x)             # True
check("||Ax - b||^2",   lambda v: np.sum((A @ v - b) ** 2),      lambda v: 2 * A.T @ (A @ v - b), x)     # True
check("||x||_2",        lambda v: np.linalg.norm(v),             lambda v: v / np.linalg.norm(v), x)     # True
check("||x||_1",        lambda v: np.abs(v).sum(),               lambda v: np.sign(v), x)                # True (no entry is 0)
check("tr(AX)",         lambda M: np.trace(A @ M),               lambda M: A.T, X)                       # True
check("tr(X^T A X)",    lambda M: np.trace(M.T @ A @ M),         lambda M: (A + A.T) @ M, X)             # True
check("tr(A X B)",      lambda M: np.trace(A @ M @ B),           lambda M: A.T @ B.T, X)                 # True
check("||AX - B||_F^2", lambda M: np.sum((A @ M - B) ** 2),      lambda M: 2 * A.T @ (A @ M - B), X)     # True
check("log|det X|",     lambda M: np.log(abs(np.linalg.det(M))), lambda M: Ai(M).T, X)                   # True
check("tr(A X^-1)",     lambda M: np.trace(A @ Ai(M)),           lambda M: -Ai(M).T @ A.T @ Ai(M).T, X)  # True

# a WRONG formula is caught: 2Ax is not the gradient of x^T A x when A is not symmetric
check("x^T A x  (wrong: 2Ax)", lambda v: v @ A @ v, lambda v: 2 * A @ v, x)                             # False

# chain rule: f(Wx) with f(u) = sum(log(cosh(u)))  ->  grad_x = W^T tanh(Wx), grad_W = tanh(Wx) x^T
W = rng.normal(size=(4, 3))
f = lambda u: np.sum(np.log(np.cosh(u)))
check("f(Wx) w.r.t. x",  lambda v: f(W @ v), lambda v: W.T @ np.tanh(W @ v), x)                          # True
check("f(Wx) w.r.t. W",  lambda M: f(M @ x), lambda M: np.outer(np.tanh(M @ x), x), W)                   # True

# a cheap check for big models: one random direction v instead of every entry
v = rng.normal(size=3)
g = (A + A.T) @ x                                             # formula for the gradient of x^T A x
h = 1e-5
slope = ((x + h * v) @ A @ (x + h * v) - (x - h * v) @ A @ (x - h * v)) / (2 * h)
print(abs(g @ v - slope) < 1e-6)                              # True
Test yourself

1. $f(X)=\operatorname{tr}(AXB)$ with $A$, $X$, $B$ all $3\times3$. Which gradient is correct?

$df=\operatorname{tr}(A\,dX\,B)=\operatorname{tr}(BA\,dX)$, so $G^\top=BA$ and $G=(BA)^\top=A^\top B^\top$. Because all matrices are square, shapes cannot reveal a mistake here: only the derivation or a numerical test can.

2. A subgradient of $\|\mathbf{x}\|_1$ at $\mathbf{x}=[2,-1,0]^\top$ is…

Entries with $x_i\ne0$ contribute $\operatorname{sign}(x_i)$. At the kink $x_3=0$ every slope in $[-1,1]$ gives a line below the graph. A convex function always has at least one subgradient.

3. For a non-symmetric invertible $X$, $\nabla_X\log|\det X|$ equals…

$d\log\det X=\operatorname{tr}(X^{-1}dX)$, so $G^\top=X^{-1}$ and $G=X^{-\top}$. For symmetric $X$ this equals $X^{-1}$, which is why the transpose is easy to forget.

4. $\mathbf{x}\in\mathbb{R}^5$, $A$ is $3\times5$, $f:\mathbb{R}^3\to\mathbb{R}$. The gradient $\nabla_{\mathbf{x}}f(A\mathbf{x})$ is…

$df=\nabla f^\top A\,d\mathbf{x}$, so the gradient is $A^\top\nabla f$, of shape $(5\times3)(3\times1)=5\times1$, the shape of $\mathbf{x}$.

5. $K$ is antisymmetric ($K^\top=-K$). Then $\mathbf{x}^\top K\mathbf{x}$ equals…

A number equals its own transpose: $\mathbf{x}^\top K\mathbf{x}=\mathbf{x}^\top K^\top\mathbf{x}=-\mathbf{x}^\top K\mathbf{x}$. So it is $0$, and its gradient $(K+K^\top)\mathbf{x}$ is $\mathbf{0}$ too.

6. Your formula and the central-difference gradient (with $h=10^{-5}$, at a random point) have a relative error of $0.3$. What do you conclude?

A correct formula agrees to about $10^{-6}$ or better with a sensible $h$. A relative error of $0.3$ is enormous. Check shapes, factors of 2 and transposes, then test again.

Practice problems

A. Find $\nabla_{\mathbf{x}}(\mathbf{a}^\top\mathbf{x})^2$ using the chain rule, and check it at $\mathbf{a}=[1,2]^\top$, $\mathbf{x}=[3,1]^\top$.

Outer $h(q)=q^2$ with $q=\mathbf{a}^\top\mathbf{x}$, $\nabla q=\mathbf{a}$. So $\nabla=2(\mathbf{a}^\top\mathbf{x})\,\mathbf{a}$. At the point: $\mathbf{a}^\top\mathbf{x}=5$, gradient $=10[1,2]^\top=[10,20]^\top$. Direct: $f=(x_1+2x_2)^2$, $\partial_1=2(x_1+2x_2)=10$, $\partial_2=4(x_1+2x_2)=20$ ✓.

B. Find $\nabla_X\|X-C\|_F^2$, and say what the gradient descent update with step $\eta=\tfrac12$ does.

With $R=X-C$, $dR=dX$, $df=2\operatorname{tr}(R^\top dX)$, so $\nabla=2(X-C)$. The update $X\leftarrow X-\tfrac12\cdot2(X-C)=C$: it lands on $C$ in one step. (Same as $A=I$ in $\|AX-B\|_F^2$.)

C. Compute $\nabla_\Sigma\log\det\Sigma$ at $\Sigma=\operatorname{diag}(2,5)$.

$\Sigma^{-\top}=\Sigma^{-1}=\operatorname{diag}(1/2,\,1/5)$. Check: $\log\det\Sigma=\ln(\Sigma_{11}\Sigma_{22}-\Sigma_{12}\Sigma_{21})$, and $\partial/\partial\Sigma_{11}=\Sigma_{22}/\det=5/10=0.5$, $\partial/\partial\Sigma_{22}=2/10=0.2$, and the off-diagonal slopes are $0$ at a diagonal matrix. ✓

D. Find $\nabla_{\mathbf{x}}\|A\mathbf{x}\|_2$ (assume $A\mathbf{x}\ne\mathbf{0}$).

Chain rule with $\mathbf{u}=A\mathbf{x}$ and outer $f(\mathbf{u})=\|\mathbf{u}\|$, $\nabla f=\mathbf{u}/\|\mathbf{u}\|$, $J=A$. So $\nabla=A^\top\dfrac{A\mathbf{x}}{\|A\mathbf{x}\|}=\dfrac{A^\top A\mathbf{x}}{\|A\mathbf{x}\|}$. Shape $n\times1$ ✓.

E. Use the recipe to find $\nabla_W\big[\operatorname{tr}(W^\top W)+\operatorname{tr}(WA)\big]$.

$df=2\operatorname{tr}(W^\top dW)+\operatorname{tr}(A\,dW)$ (using $\nabla\|W\|_F^2=2W$ and $\operatorname{tr}(A\,dW)$ for the second term). So $G^\top=2W^\top+A$ and $G=2W+A^\top$.

F. Derive $\nabla_X\operatorname{tr}(X^{-1})$.

$df=\operatorname{tr}\big(d(X^{-1})\big)=-\operatorname{tr}(X^{-1}dX\,X^{-1})$. Rotate the last $X^{-1}$ to the front: $-\operatorname{tr}(X^{-2}\,dX)$. So $G^\top=-X^{-2}$ and $G=-(X^{-2})^\top=-X^{-\top}X^{-\top}$.

Chapter 2.8

The Chain Rule

Almost every function in machine learning is a chain: a small step, then another, then another. The chain rule tells you how to get the slope of the whole chain from the slope of each link. It is the one idea behind backpropagation, so we will go slowly and build it up from gears to graphs.

  • Use the scalar chain rule and see why it works (slopes multiply)
  • Use the multivariate chain rule: add up the contribution of every path
  • Write the vector chain rule and the Jacobian chain rule, and check the shapes
  • Draw a computation as a computational graph with local derivatives
  • Run forward differentiation and reverse differentiation by hand, and compare their cost
  • Explain why a scalar loss with millions of parameters needs reverse mode (this leads straight to Chapter 2.9, backpropagation)

The scalar chain rule: slopes multiply core

Think of three gears in a row. Turn the handle once and gear A turns 2 times. Each turn of A makes gear B turn 3 times. So one turn of the handle makes B turn $2\times 3 = 6$ times. When steps are chained, the rates multiply.

A car works the same way. Petrol used per kilometre, times kilometres driven per hour, gives petrol used per hour. Each link only knows its own rate. The whole chain's rate is the product.

A derivative is just a rate: how fast the output moves per unit move of the input (see Chapter 2.3). A function that feeds into another function is a chain of two gears.

Let $y = (2x+1)^3$. Break it into two steps: first $u = 2x + 1$ (the inner step), then $y = u^3$ (the outer step). We want $dy/dx$ at $x = 1$.

  1. Inner rate: $\dfrac{du}{dx} = 2$.
  2. Outer rate: $\dfrac{dy}{du} = 3u^2$. At $x=1$ we have $u = 2\cdot1+1 = 3$, so $\dfrac{dy}{du} = 3\cdot 3^2 = 27$. (We use the inner value $u=3$, not $x=1$.)
  3. Multiply: $\dfrac{dy}{dx} = 27 \cdot 2 = 54$.

Numeric check (nudge $x$ by $0.01$). At $x=1.01$ we get $u = 3.02$ and $y = 3.02^3 = 27.543608$. The rise is $27.543608 - 27 = 0.543608$, and $0.543608 / 0.01 = 54.36 \approx 54$. ✓ (A smaller nudge gets even closer to 54.)

A chain of three links. The sigmoid is $\sigma(z) = (1 + e^{-z})^{-1}$. Write it as $a = -z$, then $b = 1 + e^{a}$, then $\sigma = b^{-1}$:

  1. $\dfrac{da}{dz} = -1$,   $\dfrac{db}{da} = e^{a} = e^{-z}$,   $\dfrac{d\sigma}{db} = -b^{-2}$.
  2. Multiply the three rates: $\dfrac{d\sigma}{dz} = (-b^{-2})\cdot e^{-z}\cdot(-1) = \dfrac{e^{-z}}{(1+e^{-z})^2}$.
  3. Now split this as $\dfrac{1}{1+e^{-z}}\cdot\dfrac{e^{-z}}{1+e^{-z}}$. The first factor is $\sigma$. The second is $1-\sigma$ (because $1 - \frac{1}{1+e^{-z}} = \frac{e^{-z}}{1+e^{-z}}$). So $\sigma'(z) = \sigma(z)\,(1-\sigma(z))$.
  4. Check at $z=0$: $\sigma = 0.5$, so $\sigma' = 0.5\cdot0.5 = 0.25$. A nudge: $\sigma(0.01) = 0.502500$, and $(0.502500 - 0.5)/0.01 = 0.25$ ✓.

If $y = f(u)$ and $u = g(x)$, so that $y = f(g(x))$, then

$$\frac{dy}{dx} \;=\; f'\big(g(x)\big)\cdot g'(x) \;=\; \frac{dy}{du}\cdot\frac{du}{dx}.$$

Read it as "(outer slope, measured at the inner value) times (inner slope)". For longer chains keep multiplying: if $y = f(v)$, $v = h(u)$, $u = g(x)$ then $\dfrac{dy}{dx} = \dfrac{dy}{dv}\dfrac{dv}{du}\dfrac{du}{dx}$.

The fraction look of $\frac{dy}{du}\cdot\frac{du}{dx}$ is a memory aid: the "$du$" seems to cancel. It is not real cancelling (these are not ordinary fractions), but the pattern is a good way to remember which pieces go together.

Why do we need it?

Real models are functions of functions. Without the chain rule we would have to expand the whole formula before differentiating. With it, we only differentiate one simple link at a time and multiply.

Where is it used?

Every activation inside a neuron (sigmoid, tanh or ReLU of a weighted sum), the derivative of log-loss, exp and log in softmax, and the whole of backpropagation in PyTorch, JAX and TensorFlow.

How is it used?

Name the inner part $u$. Differentiate the outer function with respect to $u$, and the inner function with respect to $x$. Put the inner value back into the outer slope, then multiply.

Pick an inner function $g$ and an outer function $f$, then move the slider for $x$. Left: $u = g(x)$ with its tangent. Middle: $y = f(u)$ with its tangent at the inner value $u$. Right: the composite $y = f(g(x))$ with its tangent. The slope on the right always equals the left slope times the middle slope. Try $g = x^2$ with $f = \sin u$.

  • Evaluate the outer slope at the inner value. In the example it is $3u^2$ with $u=3$, not $3x^2$ with $x=1$.
  • Do not forget the inner slope. The derivative of $\sin(x^2)$ is $\cos(x^2)\cdot 2x$, not just $\cos(x^2)$.
  • The chain rule multiplies. The sum rule adds. They are different situations: "one after another" multiplies, "side by side" adds.
Quick check: differentiate $y = e^{3x}$ and find the slope at $x = 0$.

Inner $u = 3x$, so $du/dx = 3$. Outer $y = e^u$, so $dy/du = e^u$. Then $dy/dx = 3e^{3x}$. At $x=0$ this is $3e^0 = 3$.

Why the chain rule works: a derivation from nudges

Nudge the input $x$ by a tiny amount. The inner function reacts and $u$ moves a little. The outer function sees that moved $u$ and reacts too, so $y$ moves. Two small reactions, one after the other.

Each reaction is "slope times the nudge" (that is what a slope means, when the nudge is small). So the second reaction is (outer slope) times (the inner reaction), which is (outer slope) times (inner slope) times (the first nudge).

Take $y = (2x+1)^3$ at $x=1$ and nudge by $h = 0.1$.

  1. $\Delta x = 0.1$. Then $u$ goes from $3$ to $2(1.1)+1 = 3.2$, so $\Delta u = 0.2$ and $\Delta u / \Delta x = 2$.
  2. $y$ goes from $27$ to $3.2^3 = 32.768$, so $\Delta y = 5.768$ and $\Delta y/\Delta u = 28.84$.
  3. Multiply: $\dfrac{\Delta y}{\Delta u}\cdot\dfrac{\Delta u}{\Delta x} = 28.84 \cdot 2 = 57.68$, and indeed $\dfrac{\Delta y}{\Delta x} = \dfrac{5.768}{0.1} = 57.68$. (The two ratios multiply to the third exactly, because $\Delta u$ cancels.)
  4. As $h$ shrinks the ratios settle at $2$, $27$ and $54$.

Derivation. For a nudge $\Delta x \neq 0$, let $\Delta u = g(x+\Delta x) - g(x)$ and $\Delta y = f(u+\Delta u) - f(u)$. When $\Delta u \neq 0$ we can write

$$\frac{\Delta y}{\Delta x} = \frac{\Delta y}{\Delta u}\cdot\frac{\Delta u}{\Delta x}.$$

Now let $\Delta x \to 0$. Then $\Delta u \to 0$ as well (a differentiable function does not jump), so $\Delta y/\Delta u \to f'(u)$ and $\Delta u/\Delta x \to g'(x)$. The product of the limits is $f'(u)\,g'(x)$. ∎

(If $\Delta u$ happens to be exactly $0$ the division is not allowed. Mathematicians handle that case with a small extra argument. The picture above is the heart of the proof.)

Why do we need it?

You asked to learn how to derive identities, not memorise them. This shows the chain rule is just "slope times nudge", applied twice, so you can rebuild it at any time.

Where is it used?

The same nudge argument explains gradient checking (nudge a weight, watch the loss) and the "local derivative times upstream gradient" step inside backpropagation.

How is it used?

Whenever you doubt a derivative, nudge the input by a small $h$, compute $(f(x+h)-f(x))/h$ and compare. Chain-rule results must match.

Choose a chain, move $x$ and shrink the nudge $h$. The orange line (secant) joins the two points $x$ and $x+h$ on the composite curve. As $h$ gets small it turns into the green tangent. In the readout the two ratios always multiply to the third, and as $h \to 0$ they approach the true slopes.

Quick check: why can the two nudge ratios be multiplied so easily?

Because $\Delta u$ appears once on top and once on the bottom: $\frac{\Delta y}{\Delta u}\cdot\frac{\Delta u}{\Delta x} = \frac{\Delta y}{\Delta x}$. It cancels exactly, as with ordinary fractions. The chain rule is what is left when the nudges become infinitely small.

The multivariate chain rule: add up every path core

A shop's profit depends on the price and on the number sold. Both of those depend on one dial: how much you spend on advertising. Turn the dial a little. The price reacts, and that changes the profit. The number sold reacts, and that also changes the profit.

There are two routes from the dial to the profit. Along each route, the rates multiply (that is the gears idea). To get the total effect of the dial, add the routes.

Draw the routes as arrows: dial $\to$ price $\to$ profit, and dial $\to$ number sold $\to$ profit. The rule is "multiply along a path, add over paths".

Let $f(u,v) = uv + v^2$, with $u = t^2$ and $v = 3t$. Find $df/dt$ at $t = 1$.

  1. Values at $t=1$: $u = 1$, $v = 3$.
  2. Rates of the inner steps: $\dfrac{du}{dt} = 2t = 2$ and $\dfrac{dv}{dt} = 3$.
  3. Partial derivatives of the outer function (Chapter 2.4): $\dfrac{\partial f}{\partial u} = v = 3$ and $\dfrac{\partial f}{\partial v} = u + 2v = 1 + 6 = 7$.
  4. Path through $u$: $\dfrac{\partial f}{\partial u}\cdot\dfrac{du}{dt} = 3\cdot2 = 6$. Path through $v$: $\dfrac{\partial f}{\partial v}\cdot\dfrac{dv}{dt} = 7\cdot3 = 21$.
  5. Add the paths: $\dfrac{df}{dt} = 6 + 21 = 27$.

Check by expanding. $f = t^2\cdot3t + 9t^2 = 3t^3 + 9t^2$, so $f' = 9t^2 + 18t = 27$ at $t=1$ ✓. Check by nudging. $f(1.01) = 3(1.030301) + 9(1.0201) = 12.271803$, so $(12.271803 - 12)/0.01 = 27.18 \approx 27$ ✓.

Let $f(u_1,\dots,u_n)$ and let every $u_i$ depend on $t$. Then

$$\frac{df}{dt} \;=\; \sum_{i=1}^{n}\frac{\partial f}{\partial u_i}\,\frac{du_i}{dt}.$$

If each $u_i$ depends on several inputs $x_1,\dots,x_k$, apply the rule once for each input (the other inputs stay frozen):

$$\frac{\partial f}{\partial x_j} \;=\; \sum_{i=1}^{n}\frac{\partial f}{\partial u_i}\,\frac{\partial u_i}{\partial x_j}.$$

Why it works. A tiny nudge $\Delta t$ moves each $u_i$ by about $\frac{du_i}{dt}\Delta t$. By the linear-approximation rule from Chapter 2.4 (a small step $\Delta\mathbf{u}$ changes $f$ by about $\nabla f\cdot\Delta\mathbf{u}$), $\Delta f \approx \sum_i \frac{\partial f}{\partial u_i}\Delta u_i = \Big(\sum_i \frac{\partial f}{\partial u_i}\frac{du_i}{dt}\Big)\Delta t$. Divide by $\Delta t$.

Same variable used twice. For $f = x\cdot x$ there are two paths from $x$ to $f$ (through the left factor and through the right factor). Each contributes $x$, and the sum is $2x$. This "fan-out" case is why gradients are added in backpropagation.

Why do we need it?

A value is often used in several places (a weight shared by several outputs, an input feeding several neurons). We must count its effect through every place, or the derivative will be too small.

Where is it used?

Hidden neurons that feed many neurons in the next layer, shared weights in convolution and recurrent networks, residual connections ($\mathbf{x} + f(\mathbf{x})$ uses $\mathbf{x}$ twice), and every node with fan-out in a computational graph.

How is it used?

List all paths from the input to the output. Multiply the local derivatives along each path. Add the path products. If it feels hard, draw the graph first.

The four (or six) numbers on the arrows are local derivatives. Change them with the sliders. Each path from $x$ to $f$ has a colour and its product is shown. The total $\partial f/\partial x$ is their sum. Switch to 3 routes and set one slope to $0$ to see that route drop out. Check the nudge line: moving $x$ by $0.1$ moves $f$ by exactly $0.1$ times the total.

  • Count every path. A forgotten path means a missing term. Fan-out (one value used twice) is the usual place this happens.
  • Multiply along a path, add between paths. Never add along a path and never multiply between paths.
  • $\partial f/\partial u$ inside the rule is a partial derivative (other inputs of $f$ held still), while $du/dt$ is an ordinary derivative.
Quick check: $f = \sin(x)\cdot x^2$. Use "two paths from $x$" to find $f'(x)$.

Write $f = u\cdot v$ with $u = \sin x$ and $v = x^2$. Then $\partial f/\partial u = v = x^2$, $\partial f/\partial v = u = \sin x$, $du/dx = \cos x$, $dv/dx = 2x$. Sum of paths: $f' = x^2\cos x + \sin x\cdot 2x$. This is exactly the product rule: the product rule is a chain rule over two paths.

The vector chain rule: gradient dotted with velocity core

Picture a hiker on a hilly map. Her position $\mathbf{u}(t)$ changes with time: she walks along a path. The height of the ground under her is $f(\mathbf{u})$. How fast is her height changing?

Two things decide it. The gradient $\nabla f$ points uphill and says how steep the hill is. Her velocity $\mathbf{u}'(t)$ says which way and how fast she walks. If she walks straight uphill she climbs fast. If she walks along a contour line she does not climb at all. The "how much do two arrows agree" measure is the dot product. So the climb rate is gradient · velocity.

Take the function $f(u_1,u_2) = u_1^2 + u_1u_2$ and the path $\mathbf{u}(t) = (t^2,\; 3t)$. At $t = 1$:

  1. Position: $\mathbf{u} = (1, 3)$.
  2. Gradient (as a column): $\nabla f = \begin{bmatrix} 2u_1 + u_2 \\ u_1 \end{bmatrix} = \begin{bmatrix} 5 \\ 1 \end{bmatrix}$.
  3. Velocity: $\mathbf{u}'(t) = (2t,\,3) = (2, 3)$.
  4. Dot product: $\dfrac{df}{dt} = 5\cdot2 + 1\cdot3 = 13$.

Check: $f(t) = t^4 + 3t^3$, so $f'(t) = 4t^3 + 9t^2 = 13$ at $t=1$ ✓. (This is the sum-over-paths rule with $n = 2$, written as a dot product.)

Many inputs. Now let $\mathbf{u} = g(x,y) = (xy,\; x+y)$ and $f(\mathbf{u}) = u_1u_2$, so $f = xy(x+y)$. At $(x,y) = (2,3)$: $\mathbf{u} = (6,5)$, $\nabla f = (u_2, u_1) = (5, 6)$. The matrix of inner slopes is $\begin{bmatrix} \partial u_1/\partial x & \partial u_1/\partial y \\ \partial u_2/\partial x & \partial u_2/\partial y\end{bmatrix} = \begin{bmatrix} y & x \\ 1 & 1\end{bmatrix} = \begin{bmatrix} 3 & 2 \\ 1 & 1\end{bmatrix}$. Then

$$\begin{bmatrix}\partial f/\partial x\\ \partial f/\partial y\end{bmatrix} = \begin{bmatrix} 3 & 1 \\ 2 & 1\end{bmatrix}\begin{bmatrix} 5 \\ 6\end{bmatrix} = \begin{bmatrix} 15 + 6 \\ 10 + 6\end{bmatrix} = \begin{bmatrix} 21 \\ 16\end{bmatrix}.$$

Check by expanding: $f = x^2y + xy^2$, so $f_x = 2xy + y^2 = 12 + 9 = 21$ ✓ and $f_y = x^2 + 2xy = 4 + 12 = 16$ ✓.

Along a path. If $\mathbf{u}:\mathbb{R}\to\mathbb{R}^n$ is a curve and $f:\mathbb{R}^n\to\mathbb{R}$, then

$$\frac{d}{dt}f\big(\mathbf{u}(t)\big) = \nabla f\big(\mathbf{u}(t)\big)^{\top}\,\mathbf{u}'(t) = \nabla f\cdot\mathbf{u}'.$$

From many inputs. If $\mathbf{u} = g(\mathbf{x})$ with $\mathbf{x}\in\mathbb{R}^k$, $\mathbf{u}\in\mathbb{R}^n$, let $J_g$ be the $n\times k$ matrix with entry $(i,j) = \partial u_i/\partial x_j$ (the Jacobian, Chapter 2.5). Then the gradient (a column) is

$$\nabla_{\mathbf{x}} f\big(g(\mathbf{x})\big) \;=\; J_g(\mathbf{x})^{\top}\,\nabla_{\mathbf{u}} f\big(g(\mathbf{x})\big).$$

Derivation. By the multivariate rule, $\dfrac{\partial f}{\partial x_j} = \sum_i \dfrac{\partial f}{\partial u_i}\dfrac{\partial u_i}{\partial x_j} = \sum_i (J_g)_{ij}\,(\nabla f)_i = \big(J_g^{\top}\nabla f\big)_j$.

Shape check. $J_g^\top$ is $k\times n$ and $\nabla f$ is $n\times1$, so the product is $k\times1$, the same shape as $\mathbf{x}$ ✓. The transpose appears because we are going from "gradient with respect to $\mathbf{u}$" back to "gradient with respect to $\mathbf{x}$".

Why do we need it?

It turns the sum over paths into one dot product (or one matrix times a vector), which is what a computer does fast. It also gives the meaning: how fast a loss changes along the direction that the weights move.

Where is it used?

Gradient descent: the weights follow a path $\mathbf{w}(t)$ and the loss changes at rate $\nabla L\cdot\mathbf{w}'$. Also the directional derivative, the backward step through a layer ($J^\top\nabla$), and sensitivity analysis.

How is it used?

Compute the gradient of the outer function at the inner value, compute the Jacobian (or the velocity) of the inner function, and multiply: velocity with a dot product, or the transposed Jacobian with the gradient.

The coloured map shows $f(u_1,u_2)=u_1^2+u_1u_2$; the thin lines are contours (equal height). The blue curve is the hiker's path. Move the slider for $t$. The purple arrow is the gradient (uphill), the orange arrow is the velocity (both drawn shorter than true size). Watch the number df/dt in the readout: it is never negative on this path, because $4t^3+9t^2 = t^2(4t+9)$. At $t=0$ the gradient is the zero vector (a flat point of the map), so the climb rate is exactly $0$. Compare the angle between the arrows with the sign of the dot product.

The sheet is the surface $z = xy/2$. The path on the floor is a circle of radius 2. Drag the blue dot (it stays on the circle) to choose where you are. The green curve is the path lifted onto the surface, and the orange arrow is its tangent: its steepness is $df/dt$. Use the Top and Front buttons, and rotate. Where is the climb rate zero? (Look for the lifted path's highest and lowest points.)

  • The gradient is a column vector. The dot product is $\nabla f^\top\mathbf{u}'$. If you carry the gradient as a row, the transpose sits somewhere else, so keep one convention.
  • When you go backwards from $\nabla_{\mathbf{u}}f$ to $\nabla_{\mathbf{x}}f$ you multiply by $J^\top$, not $J$. The shapes tell you: only $J^\top\nabla f$ fits.
Quick check: $f(u_1,u_2) = u_1u_2$ and the path $\mathbf{u}(t) = (t, t^2)$. Find $df/dt$ at $t=2$ with the dot-product rule, and check with $f(t)=t^3$.

$\nabla f = (u_2, u_1) = (4, 2)$ at $\mathbf{u}=(2,4)$. Velocity $= (1, 2t) = (1, 4)$. Dot product $= 4\cdot1 + 2\cdot4 = 12$. Check: $f(t) = t\cdot t^2 = t^3$ and $f'(2) = 3\cdot4 = 12$ ✓.

The Jacobian chain rule: matrices multiply core

Zoom in on a smooth function far enough and it looks like a matrix: a straight, linear map. That matrix is the Jacobian (Chapter 2.5). Now chain two such functions. Zoomed in, the first one is a matrix and the second one is a matrix. Doing one then the other is multiplying the matrices.

That is the scalar gears rule in bigger clothes. A "rate" has become a table of rates, and "multiply the rates" has become "multiply the tables".

Let $G:\mathbb{R}^2\to\mathbb{R}^2$ be $G(x_1,x_2) = (x_1x_2,\; x_1+x_2)$, and $F:\mathbb{R}^2\to\mathbb{R}^3$ be $F(u_1,u_2) = (u_1^2,\; u_1+u_2,\; u_1u_2)$. We want the Jacobian of $F\circ G$ at $\mathbf{x}=(2,3)$. (Convention: a Jacobian of a map $\mathbb{R}^n\to\mathbb{R}^m$ has $m$ rows and $n$ columns, entry $(i,j) = \partial F_i/\partial x_j$.)

  1. Inner value: $\mathbf{u} = G(2,3) = (6, 5)$.
  2. $J_G = \begin{bmatrix} x_2 & x_1 \\ 1 & 1\end{bmatrix} = \begin{bmatrix} 3 & 2 \\ 1 & 1\end{bmatrix}$ (shape $2\times2$).
  3. $J_F$ at $\mathbf{u}=(6,5)$: $\begin{bmatrix} 2u_1 & 0 \\ 1 & 1 \\ u_2 & u_1\end{bmatrix} = \begin{bmatrix} 12 & 0 \\ 1 & 1 \\ 5 & 6\end{bmatrix}$ (shape $3\times2$).
  4. Multiply, $(3\times2)(2\times2) = 3\times2$: $$J_{F\circ G} = \begin{bmatrix} 12 & 0 \\ 1 & 1 \\ 5 & 6\end{bmatrix}\begin{bmatrix} 3 & 2 \\ 1 & 1\end{bmatrix} = \begin{bmatrix} 36 & 24 \\ 4 & 3 \\ 21 & 16\end{bmatrix}.$$ For the first row: $12\cdot3 + 0\cdot1 = 36$ and $12\cdot2 + 0\cdot1 = 24$. Second row: $1\cdot3+1\cdot1 = 4$ and $1\cdot2+1\cdot1 = 3$. Third row: $5\cdot3+6\cdot1 = 21$ and $5\cdot2+6\cdot1 = 16$.

Check. The third output is $x_1x_2(x_1+x_2)$, whose gradient we found above is $(21, 16)$ ✓. The first output is $(x_1x_2)^2$, so $\partial/\partial x_1 = 2x_1x_2\cdot x_2 = 2\cdot6\cdot3 = 36$ ✓ and $\partial/\partial x_2 = 2x_1x_2\cdot x_1 = 24$ ✓.

If $G:\mathbb{R}^n\to\mathbb{R}^m$ and $F:\mathbb{R}^m\to\mathbb{R}^p$, then

$$\boxed{\;J_{F\circ G}(\mathbf{x}) \;=\; J_F\big(G(\mathbf{x})\big)\;J_G(\mathbf{x})\;}\qquad (p\times n)=(p\times m)(m\times n).$$
  • Order: the outer function's Jacobian is on the left. Matrix multiplication is not commutative.
  • Shape rule: the inner size $m$ (the number of values passed from $G$ to $F$) must match. If it does not, you have mixed up the order.
  • Long chains: $F_k\circ\cdots\circ F_1$ has Jacobian $J_k\cdots J_2J_1$, each $J_i$ evaluated at the value that reaches step $i$.
  • Special cases: $n=m=p=1$ is the scalar rule. $p=1$ (a scalar output) makes $J_F$ a single row, $\nabla f^\top$, which gives $\nabla_{\mathbf{x}}f = J_G^\top\nabla_{\mathbf{u}} f$ from the previous section.

Derivation. Near $\mathbf{x}$, $G(\mathbf{x}+\boldsymbol{\delta}) \approx G(\mathbf{x}) + J_G\boldsymbol{\delta}$. Near $\mathbf{u}=G(\mathbf{x})$, $F(\mathbf{u}+\boldsymbol{\epsilon}) \approx F(\mathbf{u}) + J_F\boldsymbol{\epsilon}$. Put $\boldsymbol{\epsilon} = J_G\boldsymbol{\delta}$: $F(G(\mathbf{x}+\boldsymbol{\delta})) \approx F(G(\mathbf{x})) + J_FJ_G\boldsymbol{\delta}$. The matrix in front of $\boldsymbol{\delta}$ is the Jacobian of the composite.

Why do we need it?

It gives the derivative of any composition of vector functions from the Jacobians of the small parts, and the shape rule catches mistakes before you compute anything.

Where is it used?

Layer-by-layer gradients in neural networks (each layer has a Jacobian), normalising flows, robot kinematics, Gauss–Newton and Levenberg–Marquardt curve fitting, and sensitivity analysis of a whole pipeline.

How is it used?

Evaluate each step's Jacobian at the value that reaches it, multiply them with the last step on the left, and check that neighbouring sizes match. In code you rarely build the full matrices; you multiply a vector through them.

$G$ has $n$ inputs and $m$ outputs. $F$ has $m$ inputs and $p$ outputs. The coloured grids are $J_G$ ($m\times n$), $J_F$ ($p\times m$) and the product ($p\times n$). The two bordered sides (size $m$) must match. Change the sizes, then press Swap the order: the product $J_G J_F$ is only possible if $n = p$.

Move $x_1$ and $x_2$. The product $J_F J_G$ (built from the two small Jacobians) is compared with the Jacobian of the whole function $F\circ G$, found by nudging each input. At $(2, 3)$ you should see the matrix from the worked example. The difference stays tiny everywhere.

  • The outer Jacobian goes on the left: $J_F J_G$, not $J_GJ_F$. Check the shapes.
  • Evaluate each Jacobian at the right point: $J_F$ at $G(\mathbf{x})$, not at $\mathbf{x}$.
  • The product of Jacobians is for the value direction (forward). Gradients of a scalar loss move the other way and use the transposes: $J_G^\top J_F^\top\nabla$. We use exactly this in reverse mode below.
Quick check: $G:\mathbb{R}^5\to\mathbb{R}^7$ and $F:\mathbb{R}^7\to\mathbb{R}^2$. What are the shapes of $J_G$, $J_F$ and $J_{F\circ G}$?

$J_G$ is $7\times5$, $J_F$ is $2\times7$, and $J_{F\circ G} = J_FJ_G$ is $(2\times7)(7\times5) = 2\times5$.

Computational graphs: nodes, edges and local derivatives core

Any formula can be cut into tiny steps where each step does one simple thing: add, multiply, take a sine. Write each result on a small box and draw an arrow from every box that was used to the box that it helped make. That picture is a computational graph.

  • A node holds one value (an input, an in-between result, or the final output).
  • An edge (arrow) says "this value was used to make that one".
  • On every edge sits a local derivative: how fast the box at the tip of the arrow changes per unit change of the box at its tail. It depends on that one tiny operation only.

Think of an assembly line. Each worker knows only their own job and how sensitive their output is to each part they receive. Nobody needs to know the whole factory.

Take $f(x,y) = xy\,(x+y)$ at $x=2$, $y=3$. Cut it into three steps:

$$v_1 = x\cdot y,\qquad v_2 = x + y,\qquad f = v_1\cdot v_2.$$
  1. Forward values: $v_1 = 6$, $v_2 = 5$, $f = 30$.
  2. Local derivatives of each step (one rule per operation): $\dfrac{\partial v_1}{\partial x} = y = 3$, $\dfrac{\partial v_1}{\partial y} = x = 2$, $\dfrac{\partial v_2}{\partial x} = 1$, $\dfrac{\partial v_2}{\partial y} = 1$, $\dfrac{\partial f}{\partial v_1} = v_2 = 5$, $\dfrac{\partial f}{\partial v_2} = v_1 = 6$.
  3. Paths from $x$ to $f$: $x\to v_1\to f$ with product $3\cdot5 = 15$, and $x\to v_2\to f$ with product $1\cdot6 = 6$. So $\dfrac{\partial f}{\partial x} = 15 + 6 = 21$.
  4. Paths from $y$ to $f$: $y\to v_1\to f$ gives $2\cdot5 = 10$, and $y\to v_2\to f$ gives $1\cdot6 = 6$. So $\dfrac{\partial f}{\partial y} = 16$.

Check: $f = x^2y + xy^2$, so $f_x = 2xy + y^2 = 12 + 9 = 21$ ✓ and $f_y = x^2 + 2xy = 4+12 = 16$ ✓.

A computational graph is a directed graph with no loops. Each node $v_i$ is computed by one elementary operation from the nodes $v_j$ that point to it. The local derivative on the edge $j\to i$ is $\dfrac{\partial v_i}{\partial v_j}$, computed from that single operation (treating its other inputs as constants).

Chain rule on a graph. The derivative of the output $f$ with respect to an input $x$ is the sum, over all paths from $x$ to $f$, of the product of the local derivatives along the path:

$$\frac{\partial f}{\partial x} \;=\; \sum_{\text{paths }x\to f}\ \prod_{\text{edges }j\to i\text{ on the path}}\frac{\partial v_i}{\partial v_j}.$$

This is exactly the multivariate rule ("multiply along, add between"), applied to a graph.

OperationLocal derivativesIn words
$v = a + b$$\partial v/\partial a = 1,\ \partial v/\partial b = 1$an add node passes the gradient to both inputs unchanged
$v = a\cdot b$$\partial v/\partial a = b,\ \partial v/\partial b = a$a multiply node passes back the other input
$v = e^{a}$$\partial v/\partial a = e^{a} = v$reuses its own output
$v = \ln a$$\partial v/\partial a = 1/a$
$v = \sin a$$\partial v/\partial a = \cos a$
$v = \sigma(a)$$\partial v/\partial a = v(1-v)$reuses its own output
$v = \max(0,a)$$1$ if $a>0$, else $0$ReLU lets the gradient through or blocks it
Why do we need it?

A big formula is hard to differentiate in one go. A graph cuts it into steps so simple that each local derivative is a one-line rule, and a computer can do the bookkeeping.

Where is it used?

PyTorch builds a graph as your code runs ("dynamic graph"), TensorFlow and JAX trace one, and compilers such as XLA optimise it. Every neural network, loss and optimiser step is a node in such a graph.

How is it used?

Run the computation and record each operation as a node with its inputs. Keep the values. Then the chain rule on the graph gives every derivative using only local rules.

Pick a function and press Next step. Each press computes one node from the nodes that feed it. After the last forward step, the next press writes the local derivative on every edge (the blue-green boxes show values). Change $x$ and $y$ with the sliders and watch all the numbers update. Try $x=2$, $y=3$ on the first function and find the $5$ and $6$ from the worked example.

Choose a function and an input. Every path from that input to the output is listed with the product of its local derivatives, and the sum of the products is $\partial f/\partial(\text{input})$. Use the slider to highlight one path in red on the graph. A node used twice (fan-out) always gives at least two paths.

  • A local derivative uses the values from the forward pass (for a multiply node it is the other input's value). So the graph must remember them.
  • One operation = one node. If you hide two operations in one node, its local derivative is no longer a one-line rule.
  • "Local" means local to one node. The final derivative needs all the local derivatives along the paths.
Quick check: in the graph of $f = x\cdot x$ (the node uses $x$ twice), what is $\partial f/\partial x$ by paths?

There are two edges from $x$ to the multiply node, one for each use. Each local derivative is the other factor, which is $x$. The two paths add: $x + x = 2x$ ✓.

Forward differentiation: carry the derivative along with the value core

When you compute $f(x,y)$ step by step you carry a number at each node: its value. Forward differentiation carries a second number next to it: how fast that node changes when you nudge one chosen input. Every node holds a pair (value, derivative), and both are updated by the same operation.

You choose one input to nudge (the seed) and press "go". The derivative information travels forward, in the same direction as the computation. It is like pushing a small ripple through the circuit and watching how big it is when it reaches the output.

Again $f = xy(x+y)$ at $(2,3)$. This time we nudge $x$: the seed is $\dot x = 1$ and $\dot y = 0$ (dot means "rate of change with respect to the nudged input"). Each node carries (value, rate).

  1. $x$: (2, 1)   $y$: (3, 0).
  2. $v_1 = xy$: value $6$. Rate: $\dot v_1 = y\,\dot x + x\,\dot y = 3\cdot1 + 2\cdot0 = 3$.
  3. $v_2 = x+y$: value $5$. Rate: $\dot v_2 = \dot x + \dot y = 1 + 0 = 1$.
  4. $f = v_1v_2$: value $30$. Rate: $\dot f = v_2\,\dot v_1 + v_1\,\dot v_2 = 5\cdot3 + 6\cdot1 = 21$.

So $\partial f/\partial x = 21$ after one pass. For $\partial f/\partial y$ we must start again with the seed $\dot x=0,\ \dot y=1$: then $\dot v_1 = 3\cdot0+2\cdot1 = 2$, $\dot v_2 = 1$, $\dot f = 5\cdot2 + 6\cdot1 = 16$. That is a second pass.

In forward mode each node $v_i$ carries its value and its rate $\dot v_i = \dfrac{\partial v_i}{\partial(\text{seed})}$. The rate is computed from the rates of the nodes that feed it, with the local derivatives:

$$\dot v_i \;=\; \sum_{j\to i}\frac{\partial v_i}{\partial v_j}\,\dot v_j.$$

One pass has about the cost of evaluating $f$ (a small constant factor more). More generally the seed can be any vector $\mathbf{s}$ (a nudge direction). One pass then returns the Jacobian–vector product $J\mathbf{s}$: the full derivative along that direction. The seeds $\mathbf{e}_1,\dots,\mathbf{e}_n$ give the columns of $J$, so the whole Jacobian costs $n$ passes, one per input.

Why do we need it?

It is the easiest form of automatic differentiation: no storage of the graph is needed, because derivatives are produced in the same order as the values.

Where is it used?

Functions with few inputs and many outputs, Jacobian-vector products (jvp in JAX), directional derivatives, sensitivity analysis in simulations, and Hessian-vector products (forward over reverse).

How is it used?

Pick the input to vary. Set its rate to 1 and all other inputs' rates to 0. Run the program once, updating (value, rate) pairs. Read the rate at the output.

Choose which input to nudge, then press Next step. Every node shows its value and its rate $d$ (orange). The text shows how $d$ comes from the local derivatives and the rates of the inputs. At the end you get the derivative with respect to one input. Switch the seed from $x$ to $y$ and notice that you must walk the whole graph again.

  • One forward pass gives the derivative with respect to one input direction only. A function with a million inputs needs a million passes for its full gradient.
  • The rate $\dot v$ is not a derivative with respect to $v$. It is how fast $v$ changes per unit change of the seed input.
Quick check: $f = x\cdot y$ at $(4, 5)$. Run forward mode with the seed $\dot x=1,\ \dot y=0$. What do you get?

$\dot f = y\,\dot x + x\,\dot y = 5\cdot1 + 4\cdot0 = 5$. That is $\partial f/\partial x = y = 5$ ✓.

Reverse differentiation: send the sensitivity backwards core

Now flip the question. Instead of asking "if I nudge this input, how does the output change?" ask "how much does the output care about each node?" Start at the output, where the answer is easy: the output changes by exactly 1 per unit change of itself. Then walk backwards. A node's importance is passed back to the nodes that fed it, multiplied by how sensitive the node is to each of them.

Think of blame in a team project that went badly. Start from the final result, ask each member how much of the blame passes to the people who supplied their work, and keep passing it back. At the end, every starting point knows how much it is responsible for the final error, all from one sweep.

$f = xy(x+y)$ at $(2,3)$. First the forward pass (values): $v_1 = 6$, $v_2 = 5$, $f = 30$. Write $\bar v$ for $\partial f/\partial v$ (the sensitivity of the output to that node).

  1. Start: $\bar f = 1$.
  2. $f = v_1v_2$: $\bar v_1 = \bar f\cdot v_2 = 5$ and $\bar v_2 = \bar f\cdot v_1 = 6$.
  3. $v_2 = x + y$: it sends $\bar v_2\cdot 1 = 6$ to $x$ and $6$ to $y$.
  4. $v_1 = xy$: it sends $\bar v_1\cdot y = 5\cdot3 = 15$ to $x$ and $\bar v_1\cdot x = 5\cdot2 = 10$ to $y$.
  5. $x$ receives from two places, so add: $\bar x = 6 + 15 = 21$. Likewise $\bar y = 6 + 10 = 16$.

One backward sweep gave both derivatives, $(21, 16)$. The same numbers as forward mode, but forward mode needed two passes.

In reverse mode, after a forward pass that stores all values, we compute for each node its adjoint $\bar v_j = \dfrac{\partial f}{\partial v_j}$ going backwards from $\bar f = 1$:

$$\bar v_j \;=\; \sum_{j\to i}\ \bar v_i\;\frac{\partial v_i}{\partial v_j}.$$

In words: (gradient arriving from each consumer) times (the local derivative of that consumer), summed over all consumers. More generally, starting from a vector $\mathbf{u}$ instead of $1$ gives the vector–Jacobian product $\mathbf{u}^\top J$. The seed $\mathbf{u}=\mathbf{e}_i$ gives row $i$ of $J$, so the full Jacobian costs $m$ passes, one per output. For a scalar output ($m=1$) one pass gives the whole gradient, whatever the number of inputs.

In matrix form, for the chain $J = J_3J_2J_1$, reverse mode computes $\bar{\mathbf{x}} = J_1^\top\big(J_2^\top(J_3^\top\,\mathbf{u})\big)$ from the output end (the transposes appear because the gradient is a column).

Why do we need it?

A model has one loss and millions of weights. Reverse mode gives the derivative of that one number with respect to every weight in a single sweep, which forward mode cannot do cheaply.

Where is it used?

Backpropagation in every deep-learning framework (loss.backward() in PyTorch, jax.grad, tf.GradientTape), and gradient-based fitting of any scalar objective: regression, logistic loss, variational inference.

How is it used?

Run the forward pass and keep every value. Set the output's adjoint to 1. Visit the nodes in reverse order: multiply the node's adjoint by each local derivative, and add it into the adjoint of each input.

Press Next step. First the forward pass fills in the values. Then the red $\partial$ numbers appear from the output backwards: each is $\partial f/\partial(\text{that node})$. The text shows every multiplication, and the word adds appears when a node receives gradient from two places. Both inputs get their derivative in one sweep, and the finite-difference check at the end agrees.

  • A node that is used in several places must add the gradients that come back from each use. Overwriting instead of adding is a classic bug.
  • The backward pass needs the forward values (for example the other factor of a multiply). Reverse mode therefore stores them. This costs memory.
  • One backward pass gives the gradient of one scalar output. Two outputs need two backward sweeps.
Quick check: for $f = x\cdot y$ at $(4,5)$, what are $\bar x$ and $\bar y$ after the backward pass?

$\bar f = 1$. The multiply node sends back the other factor: $\bar x = 1\cdot y = 5$ and $\bar y = 1\cdot x = 4$.

Forward versus reverse: counting the cost core

Picture the Jacobian $J$ of the whole function as a table with $m$ rows (outputs) and $n$ columns (inputs). The two modes fill the table in different shapes:

  • Forward mode fills a whole column per pass (one input nudged, the effect on every output).
  • Reverse mode fills a whole row per pass (one output, the sensitivity to every input).

A table with a million columns and one row is filled by one reverse pass, or by a million forward passes. A table with one column and a million rows is the opposite. A training loss has one row and one column per weight.

A tiny network has $n = 1{,}000{,}000$ parameters and one loss ($m=1$). Say one evaluation of the loss takes cost $C$. A forward or reverse pass costs a small multiple of $C$, about $3C$ (a rule of thumb: we use 3 here, and honest values are between about 2 and 4).

  1. Forward mode: $n$ passes, so about $1{,}000{,}000\times3C = 3{,}000{,}000\,C$.
  2. Reverse mode: $m = 1$ pass, so about $1\times3C = 3C$.
  3. Ratio: reverse is about one million times cheaper for the gradient.
  4. For comparison, finite differences (nudge each weight and re-evaluate) need $2n = 2{,}000{,}000$ evaluations, so about $2{,}000{,}000\,C$.

Reverse mode's price is memory: all forward values must be kept until the backward pass uses them. Forward mode needs almost no extra memory.

For $F:\mathbb{R}^n\to\mathbb{R}^m$ whose evaluation costs $C$:

Forward modeReverse mode
One pass computes$J\mathbf{s}$ (a column combination)$\mathbf{u}^\top J$ (a row combination)
Cost of one pass (rule of thumb)about $2$–$3\,C$about $2$–$4\,C$ (the "cheap gradient" principle)
Passes for the full Jacobian$n$ (one per input)$m$ (one per output)
Extra memorysmallall intermediate values
Best when$n \ll m$ (few inputs)$m \ll n$ (few outputs, e.g. a loss)

Because the Jacobian of a chain is a product $J_3J_2J_1$ and matrix multiplication can be grouped either way, the two modes are just two ways of bracketing the same product: forward mode computes $J_3(J_2(J_1\mathbf{s}))$, reverse mode computes $((\mathbf{u}^\top J_3)J_2)J_1$. A thin vector at one end keeps every intermediate product thin.

Why do we need it?

The cost of getting all the derivatives can differ by a factor of a million between the two orders. Choosing the right mode is what makes training large models practical at all.

Where is it used?

Reverse mode for training any model with a scalar loss. Forward mode for a few parameters with many outputs (physics simulations, sensitivity of a curve), and in "forward-over-reverse" Hessian-vector products for second-order methods.

How is it used?

Count inputs $n$ and outputs $m$. If $m$ is much smaller than $n$, use reverse mode (backward(), grad). If $n$ is much smaller, use forward mode (jvp). Budget memory for reverse mode.

The table is the Jacobian ($m$ rows, $n$ columns). Move passes done to the right. Forward mode (left, blue) fills one column per pass. Reverse mode (right, orange) fills one row per pass. Press Loss network ($n=12$, $m=1$): reverse is finished after 1 pass, forward needs 12. Press One dial, many outputs: the opposite.

Set the number of parameters $n$ and outputs $m$ (log scales). The bars compare the number of function evaluations of the three ways to get the full Jacobian: finite differences ($2n$), forward mode ($n$ passes) and reverse mode ($m$ passes), where one AD pass costs the chosen number of function evaluations. Press GPT-style: 1 billion parameters and see how long each would take if one evaluation takes a second.

  • "Reverse is better" is not a law. It is better when there are more inputs than outputs. With few inputs and many outputs forward mode wins.
  • The "about 3×" is a typical constant, not an exact number. It depends on the operations. The important thing is that it does not grow with the number of inputs.
  • Reverse mode keeps all intermediate values. For a very deep network this is the main memory cost of training. Tricks such as gradient checkpointing recompute some values to save memory.
Quick check: a function has 3 inputs and 500 outputs. Which mode needs fewer passes for the full Jacobian, and how many?

Forward mode needs one pass per input: 3. Reverse mode needs one per output: 500. So forward mode is cheaper here, with 3 passes.

Recap, cheat sheet and practice

  • Scalar chain rule: for $y=f(g(x))$, $\dfrac{dy}{dx}=f'(g(x))\,g'(x)$. The slopes of chained steps multiply (gears).
  • Why: a nudge $\Delta x$ causes $\Delta u\approx g'\Delta x$, which causes $\Delta y\approx f'\Delta u$. The $\Delta u$ cancels.
  • Multivariate: multiply along each path and add the paths: $\dfrac{df}{dt}=\sum_i\dfrac{\partial f}{\partial u_i}\dfrac{du_i}{dt}$.
  • Vector form: $\dfrac{d}{dt}f(\mathbf{u}(t))=\nabla f\cdot\mathbf{u}'$, and $\nabla_{\mathbf{x}}f=J_g^\top\nabla_{\mathbf{u}}f$.
  • Jacobian chain rule: $J_{F\circ G}=J_F\,J_G$ with shapes $(p\times m)(m\times n)$; the outer map goes on the left.
  • Computational graph: nodes are values, edges carry local derivatives, and $\partial f/\partial x$ is the sum over paths of the product of local derivatives.
  • Forward mode carries (value, rate) and costs one pass per input. Reverse mode sends adjoints backwards, adds at fan-out, and costs one pass per output (plus memory). A scalar loss with many parameters therefore needs reverse mode. That is backpropagation, the next chapter.

Cheat sheet

IdeaFormulaRemember
Scalar chain rule$\dfrac{dy}{dx}=\dfrac{dy}{du}\dfrac{du}{dx}$outer slope at the inner value
Paths$\dfrac{\partial f}{\partial x}=\sum_{\text{paths}}\prod(\text{local derivatives})$multiply along, add between
Along a curve$\dfrac{d}{dt}f(\mathbf{u})=\nabla f\cdot\mathbf{u}'$gradient · velocity
Gradient back$\nabla_{\mathbf{x}}f=J_g^\top\nabla_{\mathbf{u}}f$transpose when going backwards
Jacobian chain$J_{F\circ G}=J_FJ_G$$p\times n=(p\times m)(m\times n)$
Forward mode$\dot v_i=\sum_j\frac{\partial v_i}{\partial v_j}\dot v_j$$n$ passes, small memory
Reverse mode$\bar v_j=\sum_i\bar v_i\frac{\partial v_i}{\partial v_j}$$m$ passes, stores values
Local rulesadd: 1, 1; multiply: other input; ReLU: 1 or 0each node knows only itself
Code it · NumPy

import numpy as np

# 1) scalar chain rule: y = (2x + 1)^3 at x = 1
x = 1.0
u = 2 * x + 1                      # inner value
dy_dx = (3 * u**2) * 2             # (outer slope at u) * (inner slope)
f = lambda x: (2 * x + 1) ** 3
h = 1e-6
print(dy_dx, (f(x + h) - f(x - h)) / (2 * h))     # 54.0  54.00000000044258

# 2) multivariate chain rule: f(u, v) = u*v + v^2, u = t^2, v = 3t, at t = 1
t = 1.0
u, v = t**2, 3 * t
df_dt = (v) * (2 * t) + (u + 2 * v) * 3          # sum over the two paths
g = lambda t: (t**2) * (3 * t) + (3 * t) ** 2
print(df_dt, (g(t + h) - g(t - h)) / (2 * h))     # 27.0  26.999999998444935

# 3) Jacobian chain rule: J_{F o G} = J_F(G(x)) @ J_G(x)
G = lambda x: np.array([x[0] * x[1], x[0] + x[1]])
F = lambda u: np.array([u[0] ** 2, u[0] + u[1], u[0] * u[1]])

def num_jac(fn, x, h=1e-6):
    cols = []
    for j in range(len(x)):
        e = np.zeros(len(x)); e[j] = h
        cols.append((fn(x + e) - fn(x - e)) / (2 * h))
    return np.stack(cols, axis=1)

x = np.array([2.0, 3.0])
u = G(x)
JG = np.array([[x[1], x[0]], [1, 1]])
JF = np.array([[2 * u[0], 0], [1, 1], [u[1], u[0]]])
J = JF @ JG
print(J)                                           # [[36. 24.]  [ 4.  3.]  [21. 16.]]
print(np.allclose(J, num_jac(lambda z: F(G(z)), x), atol=1e-5))   # True

# 4) the same function as a tiny graph: f = (x*y) * (x+y)
def forward_mode(x, y, dx, dy):
    v1, d1 = x * y, y * dx + x * dy              # carry (value, derivative)
    v2, d2 = x + y, dx + dy
    return v1 * v2, v2 * d1 + v1 * d2

print(forward_mode(2.0, 3.0, 1.0, 0.0)[1])        # 21.0   (one pass per input)
print(forward_mode(2.0, 3.0, 0.0, 1.0)[1])        # 16.0

def reverse_mode(x, y):
    v1, v2 = x * y, x + y                        # forward pass: store values
    f_bar = 1.0                                  # backward pass
    v1_bar, v2_bar = f_bar * v2, f_bar * v1
    x_bar = v1_bar * y + v2_bar * 1.0            # x is used twice: add
    y_bar = v1_bar * x + v2_bar * 1.0
    return x_bar, y_bar

print(reverse_mode(2.0, 3.0))                    # (21.0, 16.0)  both in ONE pass

# 5) forward vs reverse as matrix products: J = J3 @ J2 @ J1 (n = 5 inputs, m = 2 outputs)
rng = np.random.default_rng(0)
J1, J2, J3 = rng.normal(size=(4, 5)), rng.normal(size=(4, 4)), rng.normal(size=(2, 4))
Jfull = J3 @ J2 @ J1
v = rng.normal(size=5)                           # forward mode: J @ v, right to left
jvp = J3 @ (J2 @ (J1 @ v))
u_ = rng.normal(size=2)                          # reverse mode: u @ J, left to right
vjp = ((u_ @ J3) @ J2) @ J1
print(np.allclose(jvp, Jfull @ v), np.allclose(vjp, u_ @ Jfull))   # True True
Test yourself

1. Let $y = (5x-2)^2$. What is $dy/dx$ at $x = 1$?

Inner $u = 5x-2 = 3$, $du/dx = 5$. Outer $dy/du = 2u = 6$. Product: $6\cdot5 = 30$. Forgetting the inner slope gives 6; using $2x$ instead of $2u$ gives 10.

2. $f(u,v)$ has $\partial f/\partial u=3$ and $\partial f/\partial v=4$ at the point of interest. If $u=2x$ and $v=x^2$, what is $df/dx$ at $x=1$?

$du/dx = 2$ and $dv/dx = 2x = 2$. Paths: $3\cdot2 + 4\cdot2 = 6 + 8 = 14$. Multiplying numbers instead of adding the paths (for example $3\cdot4\cdot2 = 24$) or forgetting a factor are the usual mistakes.

3. $G:\mathbb{R}^4\to\mathbb{R}^3$ and $F:\mathbb{R}^3\to\mathbb{R}^5$. What is the shape of the Jacobian of $F\circ G$?

$J_F$ is $5\times3$ and $J_G$ is $3\times4$, so $J_FJ_G$ is $5\times4$: five outputs by four inputs.

4. A model has 1,000,000 parameters and a single scalar loss. How many backward (reverse-mode) passes give the full gradient?

Reverse mode needs one pass per output. There is one output (the loss), so one pass gives the derivative with respect to all million parameters. The cost of that pass is a small multiple of one forward evaluation.

5. You need the full gradient of $f:\mathbb{R}^{100}\to\mathbb{R}$ using forward mode. How many passes?

Forward mode needs one pass per input direction, so 100 passes, one for each coordinate seed. Reverse mode would need just 1.

6. In the backward pass, a node's value was used in two places. What do you do with the two gradients that come back?

Each use is a separate path from the node to the output, and the multivariate chain rule adds the contributions of all paths. For $f=x\cdot x$ the two uses each give $x$, and the sum is $2x$.

Practice problems

A. Differentiate $y = \sin(3x^2)$.

Inner $u = 3x^2$, $du/dx = 6x$. Outer $y=\sin u$, $dy/du = \cos u$. So $dy/dx = \cos(3x^2)\cdot 6x$.

B. Differentiate the softplus function $y = \ln(1 + e^{x})$ and recognise the answer.

Inner $u = 1 + e^x$, $du/dx = e^x$. Outer $y=\ln u$, $dy/du = 1/u$. So $dy/dx = \dfrac{e^x}{1+e^x}$. Dividing top and bottom by $e^x$ gives $\dfrac{1}{1+e^{-x}} = \sigma(x)$: the derivative of softplus is the sigmoid. Check at $x=0$: slope $= 1/2$, and $(\ln(1+e^{0.01}) - \ln 2)/0.01 \approx 0.5012$ ✓.

C. $f(u,v)=u^2v$ with $u = x+y$ and $v = xy$. Find $\partial f/\partial x$ at $(x,y)=(1,2)$ by summing paths, then check by expanding.

Values: $u = 3$, $v = 2$. Outer partials: $f_u = 2uv = 12$, $f_v = u^2 = 9$. Inner partials with respect to $x$: $\partial u/\partial x = 1$, $\partial v/\partial x = y = 2$. Paths: $12\cdot1 + 9\cdot2 = 12 + 18 = 30$. Check: $f = (x+y)^2xy$, so $f_x = 2(x+y)xy + (x+y)^2y = 2\cdot3\cdot2 + 9\cdot2 = 12 + 18 = 30$ ✓.

D. $f(\mathbf{u}) = u_1^2 + u_2$ and $G(x_1,x_2) = (x_1+x_2,\ x_1x_2)$. Find the gradient of $f\circ G$ at $(1,2)$ with the Jacobian chain rule.

$\mathbf{u} = G(1,2) = (3, 2)$. $J_F = [\,2u_1,\ 1\,] = [6,\ 1]$ (a $1\times2$ row). $J_G = \begin{bmatrix}1 & 1\\ x_2 & x_1\end{bmatrix} = \begin{bmatrix}1&1\\2&1\end{bmatrix}$. Product: $[6\cdot1 + 1\cdot2,\ 6\cdot1+1\cdot1] = [8,\ 7]$, so the gradient is $(8,7)$. Check: $f\circ G = (x_1+x_2)^2 + x_1x_2$, so $\partial_1 = 2(x_1+x_2)+x_2 = 6+2 = 8$ ✓ and $\partial_2 = 6 + 1 = 7$ ✓.

E. A sensor model has 3 tunable inputs and produces 1000 readings. Which mode would you use to get its full Jacobian, and why?

Forward mode: it needs 3 passes (one per input), while reverse mode would need 1000 (one per output). Each forward pass also uses little memory.

F. Draw the graph of $f = x\,(x+y)$ and find $\partial f/\partial x$ at $(1,2)$ by listing the paths.

Nodes: $v = x+y = 3$, $f = x\cdot v = 3$. The input $x$ feeds $f$ directly (local derivative $= v = 3$) and also through $v$ (local derivatives $1$ and then $\partial f/\partial v = x = 1$). Paths: $3 + 1\cdot1 = 4$. Check: $f = x^2 + xy$, $f_x = 2x + y = 4$ ✓.

Chapter 2.9

Backpropagation & Automatic Differentiation

Backpropagation sounds mysterious. It is not. It is the chain rule from Chapter 2.8, applied to a computational graph, in the clever order (from the output backwards). By the end of this chapter you will be able to say why, and to show it with every number written out.

  • Compare the three ways to get a derivative: symbolic, numerical and automatic
  • Build a computational graph, run the forward pass and read off local derivatives
  • Run the backward pass: local derivative times upstream gradient, added at fan-out
  • Work a tiny neural network completely by hand (every number), and verify it numerically
  • Write backprop for dense layers, activations and losses, with shapes
  • Understand forward mode (dual numbers), reverse mode, and the cost in time and memory
  • See vanishing and exploding gradients, and check gradients like a professional
  • Explain why backpropagation is essentially repeated application of the chain rule

Three ways to get a derivative core

Suppose you want the steepness of a hill at the spot where you stand. There are three ways.

  • Symbolic: you have the formula of the hill, and you do algebra on it, like a student with a pen. The answer is another formula.
  • Numerical: you ignore the formula. You take one small step, measure how much you rose, and divide by the step. The answer is an approximation.
  • Automatic (AD): you watch the computer compute the height, one tiny operation at a time. For each tiny operation you know its exact slope (a one-line rule). The chain rule glues these slopes together as the program runs. The answer is exact, and no big formula is ever written.

Deep-learning libraries use the third way. It is as exact as algebra and almost as cheap as evaluating the function itself.

Differentiate $f(x) = x\sin x$ at $x = 1$ in all three ways.

  1. Symbolic. Product rule: $f'(x) = \sin x + x\cos x$. At $x=1$: $0.841471 + 0.540302 = 1.381773$.
  2. Numerical. Take $h = 0.001$. $f(1.001) = 0.842853$ and $f(1) = 0.841471$, so the slope is about $(0.842853 - 0.841471)/0.001 = 1.38189$. Close, but already wrong in the fourth decimal place: the error is about $1.2\times10^{-4}$.
  3. Automatic (forward). Carry (value, rate) pairs, with $x = (1,\,1)$. The step $\sin x$ gives $(0.841471,\ \cos 1\cdot1 = 0.540302)$. The product step gives value $1\cdot0.841471 = 0.841471$ and rate $1\cdot0.841471 + 1\cdot0.540302 = 1.381773$. Exactly the symbolic answer, with no formula written.

If we make the numerical step smaller, it first gets better and then worse: with $h=10^{-6}$ the error is $10^{-7}$, with $h = 10^{-12}$ it is back to $10^{-4}$, and with $h=10^{-15}$ it is $0.06$. Computers store numbers with about 16 digits, and subtracting two almost equal numbers throws digits away. The widget below shows this.

SymbolicNumerical (finite differences)Automatic (AD)
What you geta formula for $f'$an approximate numberan exact number (to rounding)
Accuracyexacttruncation error plus rounding error; even with the best $h$ the error is only about $10^{-8}$ (forward) to $10^{-10}$ (central)exact, about $10^{-16}$
Cost for a gradient of $n$ inputscan be huge ("expression swell"), then still evaluate it$n+1$ or $2n$ evaluations of $f$reverse mode: a small multiple (rule of thumb: 2–4) of one evaluation, for any $n$
Handles loops, if-statements, program code?no: needs a closed formulayes (treats $f$ as a black box)yes: it follows the actual run
Main usepen-and-paper maths, computer algebrachecking gradientstraining models

Finite differences. Forward: $f'(x)\approx\dfrac{f(x+h)-f(x)}{h}$ (error about $h$). Central: $f'(x)\approx\dfrac{f(x+h)-f(x-h)}{2h}$ (error about $h^2$, usually much better). Both are derived from the definition of the derivative in Chapter 2.3.

Automatic differentiation is neither of the other two. It applies the chain rule to the elementary operations of the program, so its results are exact. It does not manipulate formulas, and it does not take steps.

Why do we need it?

Training needs the gradient of a loss with millions of weights, at every step. We need a method that is exact, fast, and works on real program code. Only automatic differentiation checks all three boxes.

Where is it used?

PyTorch autograd, JAX grad, TensorFlow GradientTape, Stan, and Julia's Zygote use AD. Symbolic differentiation lives in SymPy and Mathematica. Finite differences are used in gradient checks and for black-box functions such as simulators.

How is it used?

You write the forward computation normally. The library records the operations and gives you the gradient (loss.backward()). You may confirm it once with a finite-difference check on a few weights.

The horizontal axis is $\log_{10}h$ (so $-6$ means $h=10^{-6}$). The vertical axis is $\log_{10}$ of the error. Blue is the forward difference, orange is the central difference, and the green dashed line is automatic differentiation. Slide $h$ and watch the dot. Both curves fall, reach a floor, then climb again at tiny $h$ (rounding error). Which $h$ is best for each? Change the function and check.

Pick a family of nested functions and raise the depth $n$. The first number is the size of the formula $f_n$ written out in full. The second is the size of its naive symbolic derivative (the chain/product rules applied with no tidying). The third is the number of steps AD needs when each partial result is computed once and reused. Look at how fast the formula sizes grow, while the AD step count grows slowly. The derivative formula is printed when it is short.

  • AD is not "numerical" and it is not "symbolic". It has no step size and it never writes the derivative formula. People often confuse these.
  • Finite differences are still the right tool for checking a gradient (see the gradient-checking section) and for functions you cannot see inside.
  • "Exact" means exact up to the usual 16-digit rounding of the computer.
Quick check: why does the numerical error first fall and then rise as $h$ shrinks?

Two errors compete. The truncation error (from using a straight line instead of the curve) shrinks as $h$ shrinks. The rounding error (from subtracting two nearly equal numbers, each stored with about 16 digits, then dividing by a tiny $h$) grows as $h$ shrinks. The best $h$ balances them.

Computational graphs and the forward pass core

A program that computes a loss is a recipe: a list of tiny steps, each using results of earlier steps. Draw every result as a box and every "was used by" as an arrow. This is the computational graph (you met it in Chapter 2.8).

The forward pass is simply running the recipe from the inputs to the output, filling in the number in each box. Nothing about derivatives yet. But there is one important extra rule: keep every number. The backward pass will need them.

A mini "neuron with a loss": one input $x$, one weight $w$ (we call it $y$ here to fit the widget), and a target $1$.

$$a = x\cdot y,\qquad s = \sigma(a),\qquad e = s - 1,\qquad L = e^2.$$

At $x = 2$, $y = 0.5$:

  1. $a = 2\cdot0.5 = 1$.
  2. $s = \sigma(1) = \dfrac{1}{1 + e^{-1}} = 0.7311$.
  3. $e = 0.7311 - 1 = -0.2689$.
  4. $L = (-0.2689)^2 = 0.0723$.

All four numbers are stored. A real network does exactly this, with millions of boxes: the layer outputs, called activations, are the stored numbers.

The forward pass evaluates the nodes of the graph in an order where every node comes after the nodes it depends on (a topological order). Each node applies its elementary operation to its inputs' values:

$$v_i = \phi_i\big(v_{j_1}, v_{j_2},\dots\big).$$

The result is the output (the loss) and a table of all intermediate values $v_i$. That table is the memory that backpropagation reads. In a network, the stored values are the pre-activations $\mathbf{z}$, activations $\mathbf{h}$, and the inputs of every layer.

Why do we need it?

Every local derivative is evaluated at the values the program actually had. Without the forward pass we would not know which numbers to plug in.

Where is it used?

Every prediction a model makes is a forward pass. In training, the forward pass computes the loss and records the activations; PyTorch's model(x) does this and builds the graph at the same time.

How is it used?

Evaluate the steps in order, store each result, and finish at the loss. Then start the backward pass. When you only want a prediction (inference), you can skip storing values and save memory.

Pick a function and press Next step to compute one node at a time (the arrows point forward). The last press shows the local derivatives, which we will use next. Choose (sigmoid(x·y) − 1)² with $x=2$ and $y=0.5$ to see the example above. Notice the order: no node is computed before the nodes it needs.

  • The forward pass is not "just" prediction: in training it also stores the intermediate values. That stored memory is large for deep networks.
  • The order must respect dependencies. In a graph with no loops such an order always exists.
Quick check: with $x=2$, $y=0.5$ in the example, which stored value will the sigmoid's local derivative use?

The sigmoid's local derivative is $s(1-s)$, which uses its own stored output $s = 0.7311$: $0.7311\times0.2689 = 0.1966$. This is why the forward values must be kept.

Local derivatives: one rule per operation core

A single node does only one small thing: add, multiply, take $e^a$. For that one thing, the slope is a one-line rule you already know. The node does not care how the rest of the network looks. It only answers: "if my input changes a tiny bit, how much does my output change?"

Backpropagation is built from a short list of local rules, one per kind of node, like a recipe book. A whole deep network is nothing but these few node types used over and over.

The node $s = \sigma(a)$ at $a = 1$ (stored output $s = 0.7311$):

  1. Rule: $\partial s/\partial a = s(1-s)$ (derived in Chapter 2.8).
  2. Number: $0.7311\times(1 - 0.7311) = 0.7311\times0.2689 = 0.1966$.
  3. Check by nudging: $\sigma(1.001) = 0.731255$ and $\sigma(1) = 0.731059$, so the slope is about $(0.731255 - 0.731059)/0.001 = 0.196 \approx 0.1966$ ✓.

The multiply node $v = a\cdot b$ at $a = 3$, $b = 4$: $\partial v/\partial a = b = 4$ and $\partial v/\partial b = a = 3$. Check: nudge $a$ to $3.01$: $v = 12.04$, rise $0.04$ over $0.01$ gives $4$ ✓.

The local derivative of a node $v = \phi(a_1,\dots,a_k)$ is the vector of its partial derivatives $\partial\phi/\partial a_j$ evaluated at the stored input values. The standard list:

NodeLocal derivative(s)
$a + b$$1$ and $1$
$a - b$$1$ and $-1$
$a\cdot b$$b$ (for $a$) and $a$ (for $b$)
$a / b$$1/b$ and $-a/b^2$
$a^k$ (constant $k$)$k\,a^{k-1}$
$e^a$$e^a$ (the stored output)
$\ln a$$1/a$
$\sin a$, $\cos a$$\cos a$, $-\sin a$
$\tanh a$$1 - \tanh^2 a$ (uses the stored output)
$\sigma(a)$$\sigma(a)(1-\sigma(a))$ (uses the stored output)
$\mathrm{ReLU}(a)=\max(0,a)$$1$ if $a > 0$, otherwise $0$

Notice that several rules reuse the node's own output: another reason to store the forward values.

Why do we need it?

If every node type has a tiny, tested derivative rule, a framework can differentiate any program built from them, even one invented tomorrow. Nobody has to differentiate the whole model by hand.

Where is it used?

Every op in PyTorch and JAX (add, matmul, relu, softmax, conv2d…) ships with its own backward rule. Writing a custom layer means writing its local derivative.

How is it used?

During the forward pass the node keeps what its rule needs. During the backward pass it multiplies the incoming gradient by its local derivative and hands the result to each input.

Choose a node type and set its input(s) with the sliders. The graph shows the node's curve in $a$ (other input fixed) and its tangent. The readout compares the local-derivative rule with a tiny-nudge measurement. Try ReLU on both sides of $0$, tanh far from $0$ (the slope vanishes), and divide with $b$ near $0$.

  • ReLU has a corner at $0$ where the derivative is not defined. Frameworks simply pick $0$ (or $1$) there. It rarely matters in practice.
  • A local derivative is evaluated at the stored input values, not at "x".
  • For a node with two inputs there are two local derivatives, one per incoming edge.
Quick check: what are the local derivatives of $v = a\cdot b$ at $a = -2$, $b = 5$?

$\partial v/\partial a = b = 5$ and $\partial v/\partial b = a = -2$.

The backward pass: upstream gradient × local derivative core

After the forward pass we know the loss. Now we ask the question training cares about: "if I nudge this number, how much does the loss move?" We answer it for every number in one sweep, starting at the loss and walking backwards.

At the loss, the answer is trivial: the loss changes by exactly 1 per unit change of itself. Now step back one node. The node before it affects the loss only through the node in front. So its answer is: (the answer in front of it) × (how much I affect the node in front). The first part is called the upstream gradient. The second is the local derivative. That multiplication is the gear rule from the chain rule. Repeat all the way back.

Continue the mini-neuron from before: $a = xy$, $s = \sigma(a)$, $e = s - 1$, $L = e^2$ at $x=2$, $y = 0.5$. Forward values: $a = 1$, $s = 0.7311$, $e = -0.2689$, $L = 0.0723$. Now backwards:

  1. $\dfrac{\partial L}{\partial L} = 1$.
  2. $L = e^2$, local derivative $2e = -0.5378$. So $\dfrac{\partial L}{\partial e} = 1\times(-0.5378) = -0.5378$.
  3. $e = s - 1$, local derivative $1$. So $\dfrac{\partial L}{\partial s} = -0.5378\times1 = -0.5378$.
  4. $s = \sigma(a)$, local derivative $s(1-s) = 0.1966$. So $\dfrac{\partial L}{\partial a} = -0.5378\times0.1966 = -0.1057$ (more digits: $-0.10575$).
  5. $a = x\cdot y$: local derivatives $\partial a/\partial x = y = 0.5$ and $\partial a/\partial y = x = 2$. So $\dfrac{\partial L}{\partial x} = -0.10575\times0.5 = -0.0529$ and $\dfrac{\partial L}{\partial y} = -0.10575\times2 = -0.2115$.

Check. Nudge $y$ from $0.5$ to $0.51$: $L$ changes from $0.0723$ to $0.0702$ (rise $-0.0021$), and $-0.0021/0.01 = -0.209 \approx -0.2115$ ✓. A central difference with tiny $h$ gives $-0.21151$, matching to five digits.

The backward pass (reverse-mode differentiation). Given the stored forward values:

  1. Set the gradient at the output (the loss) to $1$, and all other gradients to $0$.
  2. Visit the nodes in reverse topological order. For a node $v_i$ with gradient $\bar v_i = \partial L/\partial v_i$ and inputs $v_j$: $$\bar v_j \;\mathrel{+}=\; \bar v_i\cdot\frac{\partial v_i}{\partial v_j}\qquad\text{(for every input }j\text{ of node }i\text{).}$$
  3. When every node has been visited, $\bar v_j = \partial L/\partial v_j$ for every node, including every weight.

The sign "$+=$" matters: a node used in several places receives a contribution from each, and they add. This is the multivariate chain rule of Chapter 2.8. So the backward pass is just: local derivative × upstream gradient, node after node, and sum at fan-out. Each node needs only its own local rule and the number arriving from the front.

Why do we need it?

Gradient descent needs $\partial L/\partial w$ for every weight. The backward pass delivers all of them at once, at a cost close to one extra forward pass.

Where is it used?

This is what loss.backward() does in PyTorch, jax.grad in JAX and tape.gradient in TensorFlow, for networks from tiny MLPs to large language models.

How is it used?

Run the forward pass, call backward, and read the stored .grad of each parameter. The optimiser then uses it (for example $w \leftarrow w - \eta\,\partial L/\partial w$).

Press Next step: first the forward pass fills the values, then the red $\partial L$ numbers flow from the output back to the inputs. Every step is written as (upstream) × (local). Start with the first function and $x=2$, $y=0.5$ to reproduce the worked example. Then pick x·y·(x + y): $x$ and $y$ are used twice, so watch the word adds. Then try relu(x·y − 1) + x with $y=0.5$: the ReLU blocks the gradient on one path.

  • Gradients add where a value is used more than once. Overwriting is a bug that silently gives wrong gradients.
  • The gradient at the output is 1 only because we differentiate the loss with respect to itself. With several outputs you start from a chosen weighting vector.
  • Gradients must be reset between training steps (optimizer.zero_grad()), because PyTorch accumulates them with "+=" just as above.
Quick check: $L = (a\cdot b)^2$ with $a=1$, $b=3$. Do the backward pass for $\partial L/\partial a$.

Forward: $p = ab = 3$, $L = p^2 = 9$. Backward: $\bar L = 1$; $\bar p = 2p = 6$; $\bar a = \bar p\cdot b = 6\cdot3 = 18$ and $\bar b = \bar p\cdot a = 6$. Check: $L = a^2b^2$, so $\partial L/\partial a = 2ab^2 = 18$ ✓.

Backpropagation: a tiny network, every number by hand core

Time to do it on a real (very small) neural network: 2 inputs, 2 hidden neurons, 1 output. The hidden neurons use the sigmoid. The output is a plain weighted sum. The loss is half the squared error. The network has $4 + 2 + 2 + 1 = 9$ numbers to learn: two weights into each hidden neuron (4), a bias for each (2), two weights into the output (2) and an output bias (1).

We will compute the forward pass, then the backward pass, then verify every one of the 9 gradients with a finite difference. After that, you can truthfully say you have done backpropagation.

Input $\mathbf{x} = (1, 2)$, target $t = 1$. Weights: $W_1 = \begin{bmatrix} 0.1 & 0.2 \\ -0.3 & 0.4\end{bmatrix}$, $\mathbf{b}_1 = (0.1, -0.1)$, $W_2 = [\,0.5,\ -0.5\,]$, $b_2 = 0.2$. (Values are shown rounded to 4 decimals, but each step is computed from the unrounded numbers. So a product of two printed numbers can differ from the printed result in the last digit.)

Forward pass.

  1. Pre-activations $\mathbf{z} = W_1\mathbf{x}+\mathbf{b}_1$: $z_1 = 0.1\cdot1 + 0.2\cdot2 + 0.1 = 0.6$, $\;z_2 = -0.3\cdot1 + 0.4\cdot2 - 0.1 = 0.4$.
  2. Activations $\mathbf{h} = \sigma(\mathbf{z})$: $h_1 = \sigma(0.6) = 0.6457$, $\;h_2 = \sigma(0.4) = 0.5987$.
  3. Output $\hat y = W_2\mathbf{h} + b_2 = 0.5\cdot0.6457 - 0.5\cdot0.5987 + 0.2 = 0.3228 - 0.2993 + 0.2 = 0.2235$.
  4. Error $r = \hat y - t = 0.2235 - 1 = -0.7765$, and loss $L = \tfrac12 r^2 = \tfrac12\cdot0.6030 = 0.3015$.

Backward pass (upstream × local at every step).

  1. $\dfrac{\partial L}{\partial\hat y} = r = -0.7765$ (since $L=\tfrac12r^2$ and $r = \hat y - t$).
  2. Output layer: $\dfrac{\partial L}{\partial W_2} = r\cdot\mathbf{h}^\top = (-0.7765\cdot0.6457,\ -0.7765\cdot0.5987) = (-0.5014,\ -0.4649)$ and $\dfrac{\partial L}{\partial b_2} = r = -0.7765$.
  3. Back to the hidden activations: $\dfrac{\partial L}{\partial h_j} = r\cdot W_2[j]$, so $\dfrac{\partial L}{\partial h_1} = -0.7765\cdot0.5 = -0.3883$ and $\dfrac{\partial L}{\partial h_2} = -0.7765\cdot(-0.5) = 0.3883$.
  4. Through the sigmoids: $\sigma'(z_j) = h_j(1-h_j)$, so $\sigma'(z_1) = 0.6457\cdot0.3543 = 0.2288$ and $\sigma'(z_2) = 0.5987\cdot0.4013 = 0.2403$. Then $\delta_1 = \dfrac{\partial L}{\partial z_1} = -0.3883\cdot0.2288 = -0.0888$ and $\delta_2 = \dfrac{\partial L}{\partial z_2} = 0.3883\cdot0.2403 = 0.0933$.
  5. Hidden layer weights: $\dfrac{\partial L}{\partial W_1[j,i]} = \delta_j\,x_i$. So $\dfrac{\partial L}{\partial W_1} = \begin{bmatrix} -0.0888\cdot1 & -0.0888\cdot2 \\ 0.0933\cdot1 & 0.0933\cdot2\end{bmatrix} = \begin{bmatrix} -0.0888 & -0.1777 \\ 0.0933 & 0.1866\end{bmatrix}$, and $\dfrac{\partial L}{\partial\mathbf{b}_1} = (\delta_1,\delta_2) = (-0.0888,\ 0.0933)$.

Verify numerically. Take $W_1[1,1]$. Add and subtract $h = 10^{-6}$, recompute $L$ each time, and divide the difference by $2h$: the result is $-0.088827$, matching $-0.0888$ ✓. All nine gradients agree with their finite differences to about $10^{-10}$ or better (the widget does this check for you).

One gradient step with learning rate $\eta = 0.1$ ($w \leftarrow w - \eta\,\partial L/\partial w$) lowers the loss from $0.3015$ to $0.1959$. Repeating forward, backward, update is called training.

For the network $\mathbf{z} = W_1\mathbf{x}+\mathbf{b}_1,\ \mathbf{h}=\sigma(\mathbf{z}),\ \hat y = W_2\mathbf{h}+b_2,\ L = \tfrac12(\hat y - t)^2$, backpropagation computes

$$\begin{aligned} r &= \hat y - t, & \dfrac{\partial L}{\partial W_2} &= r\,\mathbf{h}^\top, & \dfrac{\partial L}{\partial b_2} &= r,\\ \dfrac{\partial L}{\partial \mathbf{h}} &= W_2^\top r, & \boldsymbol{\delta} = \dfrac{\partial L}{\partial \mathbf{z}} &= \dfrac{\partial L}{\partial \mathbf{h}}\odot\mathbf{h}\odot(1-\mathbf{h}), & &\\ \dfrac{\partial L}{\partial W_1} &= \boldsymbol{\delta}\,\mathbf{x}^\top, & \dfrac{\partial L}{\partial \mathbf{b}_1} &= \boldsymbol{\delta}. & & \end{aligned}$$

The symbol $\odot$ means entry-by-entry multiplication. Every line is "(upstream gradient) × (local derivative)". Nothing else is going on.

Why do we need it?

This is the whole of training in miniature: the forward pass gives the loss, the backward pass gives how every weight should change, and a small step downhill improves the model. Seeing every number removes the magic.

Where is it used?

The same pattern, scaled up, trains every MLP, CNN, RNN and Transformer. Real networks have more layers and other activations, but each layer repeats these same steps.

How is it used?

Do forward, then backward, then update the weights by $-\eta$ times each gradient. Verify new code with a finite-difference check on a few weights. The next widgets let you do exactly that.

Press Next step to walk through the forward pass (4 steps) and the backward pass (5 steps), with every multiplication written below. At the end the table compares all 9 backprop gradients with finite differences. Then press Take one gradient step a few times and watch the loss fall. You can edit any weight, the input or the target in the boxes; the default values are the worked example.

  • The factor $\sigma'(z) = h(1-h)$ is at most $0.25$. If it is tiny (a saturated sigmoid) the gradient behind it is tiny too. We return to this in the vanishing-gradients section.
  • $\partial L/\partial W_1$ needs the input $x$ and the upstream $\delta$. $\partial L/\partial W_2$ needs the stored activation $h$. That is why the forward pass stores them.
  • Rounding: the numbers here are shown with 4 decimals, so a hand multiplication may differ from the printed one in the last digit.
Quick check: in the example, which of the nine weights has the gradient with the biggest size, and why?

$b_2$ has $-0.7765$ (the error itself) and $W_2[1]$, $W_2[2]$ have $-0.5014$ and $-0.4649$. They sit right next to the output, so no small factor (like $\sigma' \le 0.25$) is multiplied in. The hidden-layer weights have gradients of at most about $0.19$, because they pass through $W_2$ and then $\sigma'$.

Computational graphs in neural networks: dense layer, activation, loss core

A real network is a graph made of only three kinds of building blocks, repeated:

  • Dense (linear) layer: $\mathbf{y} = W\mathbf{x} + \mathbf{b}$, a bundle of weighted sums.
  • Activation: an entry-by-entry function such as ReLU or sigmoid, $\mathbf{h} = \varphi(\mathbf{z})$.
  • Loss: one number that compares the output with the target.

Each block has a tiny backward rule. Backpropagation through the whole network is just these rules, applied in reverse order, layer by layer, instead of node by node. A "layer" is a group of nodes that we treat together to use fast matrix code.

A dense layer with $\mathbf{x} = (1, 2, -1)$, $W = \begin{bmatrix} 1 & 0 & 2 \\ -1 & 3 & 1\end{bmatrix}$ and $\mathbf{b} = (0, 1)$. The upstream gradient (the gradient arriving from the layers in front) is $\mathbf{g} = \partial L/\partial\mathbf{y} = (2, -1)$.

  1. Forward: $\mathbf{y} = W\mathbf{x}+\mathbf{b} = (1 + 0 - 2 + 0,\ -1 + 6 - 1 + 1) = (-1,\ 5)$.
  2. $\dfrac{\partial L}{\partial W} = \mathbf{g}\,\mathbf{x}^\top = \begin{bmatrix} 2\cdot1 & 2\cdot2 & 2\cdot(-1)\\ -1\cdot1 & -1\cdot2 & -1\cdot(-1)\end{bmatrix} = \begin{bmatrix} 2 & 4 & -2 \\ -1 & -2 & 1\end{bmatrix}$ (same shape as $W$).
  3. $\dfrac{\partial L}{\partial\mathbf{b}} = \mathbf{g} = (2, -1)$.
  4. $\dfrac{\partial L}{\partial\mathbf{x}} = W^\top\mathbf{g} = \begin{bmatrix}1 & -1\\0 & 3\\2 & 1\end{bmatrix}\begin{bmatrix}2\\-1\end{bmatrix} = (2+1,\ 0-3,\ 4-1) = (3,\ -3,\ 3)$.

Check with $L = 2y_1 - y_2$ (so that $\partial L/\partial\mathbf{y} = \mathbf{g}$): $\partial L/\partial x_1 = 2\cdot W_{11} - 1\cdot W_{21} = 2 + 1 = 3$ ✓.

Backward rules for the three blocks (all with a column-vector gradient convention):

Block (forward)Given upstream gradientBackward rule
Dense: $\mathbf{y} = W\mathbf{x} + \mathbf{b}$
$W$: $m\times n$
$\mathbf{g} = \partial L/\partial\mathbf{y}$ ($m\times1$)$\dfrac{\partial L}{\partial W} = \mathbf{g}\mathbf{x}^\top$ ($m\times n$),   $\dfrac{\partial L}{\partial\mathbf{b}} = \mathbf{g}$,   $\dfrac{\partial L}{\partial\mathbf{x}} = W^\top\mathbf{g}$ ($n\times1$)
Activation: $\mathbf{h} = \varphi(\mathbf{z})$$\mathbf{g} = \partial L/\partial\mathbf{h}$$\dfrac{\partial L}{\partial\mathbf{z}} = \mathbf{g}\odot\varphi'(\mathbf{z})$
Loss: $L = \tfrac12\|\hat{\mathbf{y}}-\mathbf{t}\|^2$(start of the backward pass)$\dfrac{\partial L}{\partial\hat{\mathbf{y}}} = \hat{\mathbf{y}} - \mathbf{t}$
Loss: softmax then cross-entropy(start)$\dfrac{\partial L}{\partial\mathbf{z}} = \mathbf{p} - \mathbf{y}_{\text{one-hot}}$ (you will meet this derivation in Chapter 2.14; the log-sum-exp part of it is in Chapter 2.7)

Derivation for the dense layer. The loss depends on $W_{ij}$ only through $y_i = \sum_j W_{ij}x_j + b_i$, and $\partial y_i/\partial W_{ij} = x_j$. So $\partial L/\partial W_{ij} = g_i\,x_j$, which is the outer product $\mathbf{g}\mathbf{x}^\top$. Also $x_j$ feeds all outputs, with $\partial y_i/\partial x_j = W_{ij}$, so by the multivariate rule $\partial L/\partial x_j = \sum_i g_iW_{ij} = (W^\top\mathbf{g})_j$.

Shape rule. Every gradient has the same shape as the thing it is the gradient of. $\partial L/\partial W$ is $m\times n$ like $W$; $\partial L/\partial\mathbf{x}$ is $n\times1$ like $\mathbf{x}$. With a batch of $B$ examples the weight gradient is the sum over the batch, $\sum_b \mathbf{g}_b\mathbf{x}_b^\top$.

Why do we need it?

Treating a whole layer as one block lets us use fast matrix multiplication on GPUs, for both the forward and the backward pass, with only a few rules to implement per layer type.

Where is it used?

The linear layers, feed-forward blocks and attention projections of Transformers, fully connected heads of CNNs, and every library's Linear module use these rules. Convolutions, normalisation and softmax have their own similar rules.

How is it used?

Walk the layers from the last to the first. At each one, use the upstream gradient to get the weight gradient (using the stored input of that layer), then the gradient for the layer before it (using the transposed weights).

The boxes hold $W$ (2×3), $\mathbf{b}$, the input $\mathbf{x}$ and the upstream gradient $\mathbf{g}$. The readout shows $\mathbf{y}$ and the three backward results. For the check we use $L = \mathbf{g}\cdot\mathbf{y}$ (so that its gradient with respect to $\mathbf{y}$ is exactly $\mathbf{g}$) and nudge every entry of $W$ and $\mathbf{x}$. Change any number and the check still passes. Notice that the weight-gradient entry $(i,j)$ is $g_i\,x_j$.

Set the width of each layer and the batch size (log sliders). The table lists, for each dense layer, the shape of $W$, the values the forward pass must store, and the shapes of the backward results. Look at the last lines: the backward pass costs about twice the forward multiplications, as a rule of thumb ($\partial W$ and $\partial\mathbf{x}$ each cost about as much as the forward product), and memory grows with the batch size because activations are stored for every example.

  • $\partial L/\partial W = \mathbf{g}\mathbf{x}^\top$ and $\partial L/\partial\mathbf{x} = W^\top\mathbf{g}$. Mixing up the transpose is the commonest hand-written-backprop bug. Let the shapes tell you.
  • The activation backward rule is an entry-wise product. It is not a matrix product, because each activation acts on one entry only.
  • With a mean loss over $B$ examples, remember the $1/B$ factor.
Quick check: a dense layer maps 784 inputs to 128 outputs. What is the shape of $\partial L/\partial W$, and of $\partial L/\partial\mathbf{x}$?

$W$ is $128\times784$, so $\partial L/\partial W$ is $128\times784$ (the same shape as $W$). $\partial L/\partial\mathbf{x} = W^\top\mathbf{g}$ is $(784\times128)(128\times1) = 784\times1$.

Why backpropagation is just the chain rule, used cleverly core

Here is the one sentence to remember:

Backpropagation is the chain rule applied again and again from the output backwards, saving the partial products so that every weight can reuse them.

The gradient of one weight is always a product of local derivatives along a path from that weight to the loss. Many weights share the front part of their paths. Backprop multiplies that shared part once and passes it back, instead of recomputing it for each weight. That reuse is the "propagation".

In the tiny network, take the weight $W_1[1,1]$. Follow its path to the loss: $W_1[1,1]\to z_1\to h_1\to\hat y\to L$. The chain rule multiplies the four local derivatives:

$$\frac{\partial L}{\partial W_1[1,1]} = \underbrace{\frac{\partial L}{\partial\hat y}}_{r}\;\underbrace{\frac{\partial\hat y}{\partial h_1}}_{W_2[1]}\;\underbrace{\frac{\partial h_1}{\partial z_1}}_{h_1(1-h_1)}\;\underbrace{\frac{\partial z_1}{\partial W_1[1,1]}}_{x_1} = (-0.7765)(0.5)(0.2288)(1) = -0.0888.$$

Now see what is shared. The first factor $r$ is used by all nine gradients. The product $r\cdot W_2[1]\cdot\sigma'(z_1) = \delta_1$ is used by $W_1[1,1]$, $W_1[1,2]$ and $b_1[1]$. Backprop computes $r$, then $\delta_1$, then multiplies by $x_1$ or $x_2$ at the very end. Nothing is multiplied twice.

The dictionary.

Chain-rule idea (Chapter 2.8)What backprop does
Slopes of chained steps multiplyEach backward step multiplies the incoming gradient by a local derivative
Sum over all pathsWhere a value is used twice, the returning gradients are added
Jacobian chain rule $J=J_k\cdots J_1$For a scalar loss, multiply from the left with transposes: $J_1^\top(\cdots(J_k^\top\mathbf{1}))$
Reverse mode: one pass per outputOne backward pass gives the gradient of the loss for every weight
Local derivatives need valuesThe forward pass stores activations

Algorithm. (1) Forward pass: compute and store all values. (2) Set $\partial L/\partial L = 1$. (3) For each node in reverse order: multiply its gradient by each local derivative and add to the inputs' gradients. (4) Read the gradients of the weights. This is exactly the chain rule over the graph, and nothing more.

Why do we need it?

Computing every weight's gradient by a separate full chain would repeat huge amounts of work. Reusing the shared factors makes the cost grow in step with the network size, not with its square.

Where is it used?

The reuse idea is why training networks with hundreds of layers is possible at all. It is also why frameworks store activations, and why tricks like gradient checkpointing and mixed precision are about memory (and speed) rather than about maths.

How is it used?

When you debug a gradient, pick one weight and write its chain of local derivatives. Compare each factor with what the framework stored. It must be the same product.

Pick a weight of the tiny network. Its gradient is shown as the product of local derivatives along its path to the loss, with the numbers filled in. The final product equals the value backprop found. Then look at the second part: for a chain of $L$ single-weight layers, computing each weight's gradient separately repeats work (about $L^2/2$ multiplications), while backprop reuses the running product ($L$ multiplications). Move $L$.

  • "Just the chain rule" does not mean "trivial". The clever part is the order (from the loss backwards) and reuse of partial products. Reverse order is what makes one pass enough for all weights.
  • Backprop is not a learning algorithm by itself. It only computes gradients. The optimiser (such as gradient descent) uses them to change the weights.
Quick check: say in one sentence why backpropagation is repeated application of the chain rule.

Each weight's gradient is a product of local derivatives along its path to the loss (chain rule), and backprop evaluates these products from the loss backwards, one node at a time, adding where paths merge and reusing the shared partial products for all weights behind them.

Forward-mode differentiation and dual numbers core

In Chapter 2.8 we carried a pair (value, rate) through the graph. There is a neat way to see that pair as one number. Imagine a number $x + \varepsilon$: $x$ plus an infinitely tiny nudge $\varepsilon$, so tiny that $\varepsilon^2$ is zero for all practical purposes.

Run your program on this nudged number. Every operation keeps two parts: the ordinary part and the nudge part. At the end, the nudge part is exactly the derivative. You never wrote a derivative rule for the whole program. The arithmetic of nudged numbers produced it.

Evaluate $f(x) = x^2 + 3x$ at $x = 2 + \varepsilon$, using $\varepsilon^2 = 0$:

$$(2+\varepsilon)^2 + 3(2+\varepsilon) = (4 + 4\varepsilon + \varepsilon^2) + (6 + 3\varepsilon) = 10 + 7\varepsilon.$$

The ordinary part $10$ is $f(2)$. The nudge part $7$ is $f'(2) = 2\cdot2+3$ ✓.

A product: $(1+\varepsilon)\cdot\sin(1+\varepsilon)$. The rule $\sin(a + b\varepsilon) = \sin a + b\cos a\,\varepsilon$ gives $\sin(1+\varepsilon) = 0.8415 + 0.5403\varepsilon$. Then $(1+\varepsilon)(0.8415 + 0.5403\varepsilon) = 0.8415 + (0.5403 + 0.8415)\varepsilon + 0.5403\varepsilon^2 = 0.8415 + 1.3818\,\varepsilon$. So $f(1) = 0.8415$ and $f'(1) = 1.3818$ for $f = x\sin x$, the same as in the first section.

A dual number is $a + b\varepsilon$ with real $a, b$ and the rule $\varepsilon^2 = 0$ (but $\varepsilon\neq0$). Its arithmetic:

$$\begin{aligned} (a+b\varepsilon) \pm (c+d\varepsilon) &= (a\pm c) + (b\pm d)\varepsilon\\ (a+b\varepsilon)(c+d\varepsilon) &= ac + (ad + bc)\varepsilon \qquad\text{(the product rule appears!)}\\ \frac{a+b\varepsilon}{c+d\varepsilon} &= \frac{a}{c} + \frac{bc - ad}{c^2}\varepsilon\qquad (c\neq0)\\ \varphi(a+b\varepsilon) &= \varphi(a) + \varphi'(a)\,b\,\varepsilon\qquad\text{(one-input functions)} \end{aligned}$$

Why it works. The last line is the first-order Taylor expansion $\varphi(a+\Delta) \approx \varphi(a) + \varphi'(a)\Delta$ with $\Delta = b\varepsilon$ and all higher terms dropped because $\varepsilon^2=0$. Because each operation is computed correctly this way, the chain rule is applied automatically, step by step: $f(x+\varepsilon) = f(x) + f'(x)\,\varepsilon$ for any program built from these operations.

Forward-mode AD is exactly "run the program on dual numbers". To differentiate with respect to input $x_i$, give that input the nudge part $1$ and every other input the nudge part $0$. One run gives one derivative (one column of the Jacobian); a function with $n$ inputs needs $n$ runs. The cost of one run is a small constant times a normal run, since every number is carried as two numbers.

Why do we need it?

It shows that differentiation can be done by ordinary arithmetic on a slightly richer kind of number. Operator overloading gives a working forward-mode AD in a few dozen lines, and it is exact.

Where is it used?

JAX jvp, Julia's ForwardDiff package, C++ libraries such as Eigen's autodiff, and quick sensitivity analysis of physics and finance models with few inputs. Forward mode also gives cheap Hessian-vector products when combined with reverse mode.

How is it used?

Replace the number type with a dual number type (a pair), seed the input of interest with nudge part 1, run the code, and read the nudge part of the result as the derivative.

Choose an operation and set $a = a_0 + a_1\varepsilon$ and $b = b_0 + b_1\varepsilon$. The result has an ordinary part and a nudge part, with the algebra written out. The last line checks the nudge part by actually nudging: it should match. Try $a_1 = 1$, $b_1 = 0$ with multiplication: the nudge part is $b_0$, the derivative of $a\cdot b_0$. Then give both a nudge to see the product rule.

Choose a function. The blue curve is $f(x)$ and the orange curve is $f'(x)$, and both are produced by running the same code on dual numbers $x+\varepsilon$ at many points (no derivative formula was typed). The thin dashed curve is the exact formula for $f'$: it lies right under the orange one. Slide $x$ and read both parts of the dual result.

  • $\varepsilon$ is not a small number you pick. It is a formal symbol with $\varepsilon^2 = 0$. That is why there is no truncation error and no step size to tune.
  • One run still gives only one directional derivative. For a million parameters you would need a million runs. That is why training uses reverse mode instead.
  • Dual numbers carry two numbers instead of one, so each operation costs about two to three times more than the plain number.
Quick check: compute $(3+\varepsilon)^3$ with $\varepsilon^2 = 0$ (and so $\varepsilon^3 = 0$). What are $f(3)$ and $f'(3)$ for $f=x^3$?

$(3+\varepsilon)^3 = 27 + 3\cdot9\,\varepsilon + 3\cdot3\,\varepsilon^2 + \varepsilon^3 = 27 + 27\varepsilon$. So $f(3) = 27$ and $f'(3) = 27 = 3\cdot3^2$ ✓.

Computational complexity: time and memory core

What does backpropagation cost? Two things matter: time and memory.

  • Time. The backward pass visits every node once, with about the same amount of arithmetic as the forward pass. For a dense layer it is exactly two matrix products ($\partial L/\partial W$ and $\partial L/\partial\mathbf{x}$), each as big as the forward one. So the backward pass costs about 2× the forward pass, and a whole training step about 3× one prediction. This is a rule of thumb. The key point is that you do not need one extra pass per weight.
  • Memory. The backward pass needs the forward values, so the forward pass must keep them: all the layer inputs and activations, for every example in the batch. For big networks this is the number that fills your GPU.

A plain network with $L = 100$ layers, each $1000$ wide, batch size $B = 128$.

  1. Parameters: $100\times1000\times1000 = 10^8$ numbers.
  2. Forward multiply-adds: $L\,B\,n^2 = 100\cdot128\cdot10^6 = 1.28\times10^{10}$. Backward: about twice that, $2.56\times10^{10}$. One training step: about $3.84\times10^{10}$.
  3. Stored activations: $L\,B\,n = 100\cdot128\cdot1000 = 1.28\times10^7$ numbers, which is $51$ MB in 32-bit floats.
  4. Gradient checkpointing: keep only every $k$-th layer's activations and recompute the others during the backward pass. With $k=\sqrt{L}=10$ you store about $2\sqrt{L} = 20$ layers' worth instead of 100 (memory about 5× smaller) and pay about one extra forward pass in time.

For convolutional networks and Transformers the stored activations are usually much larger than in this plain example, which is why memory, not arithmetic, is often the limit.

Reverse mode (backpropagation). Time: about $2$–$4$ times the cost of evaluating the function (a rule of thumb, the "cheap gradient principle"), for a scalar output, independent of the number of inputs. Memory: proportional to the number of intermediate values stored, that is, the size of the forward computation.

Forward mode. Time: about $2$–$3$ times the cost of the function per input. Memory: little more than the function itself.

Whole Jacobian of $F:\mathbb{R}^n\to\mathbb{R}^m$. Forward: $n$ passes. Reverse: $m$ passes. A scalar loss has $m=1$.

Finite differences. $2n$ function evaluations for a gradient (central differences), which for millions of weights is hopeless.

Why do we need it?

Cost decides what can be trained. Knowing that the backward pass costs about 2× the forward pass, and that memory scales with stored activations, tells you how big a model and batch fit on your hardware.

Where is it used?

Planning GPU memory for training large models, choosing the batch size, gradient checkpointing (torch.utils.checkpoint), mixed precision (16-bit activations halve the memory), and comparing AD modes in scientific computing.

How is it used?

Estimate parameters as the sum of weight sizes, activations as batch × layer widths, and time as about 3× the forward cost. If memory is too high, lower the batch size, use checkpointing, or use smaller precision.

The program computes $f(x_1,\dots,x_n)=\sum_i \sin(x_i)\,x_{i+1}$ (a scalar), built as a graph by our own engine. Increase $n$. Forward mode runs once per input, so its total work grows with $n$. Reverse mode runs once. Both give the same gradient (the largest difference is shown). Work is counted as node evaluations plus multiply-adds on edges.

Set the depth $L$, the width $n$ and the batch size $B$. The bars compare memory for the weights and for the stored activations, and the time for forward and backward (in multiply-adds). Switch on gradient checkpointing to store only about $2\sqrt{L}$ layers: memory drops a lot, time rises by about one forward pass. Press Wide, shallow and Deep, big batch to see which of the two memories dominates.

  • "About 2–3×" is a rule of thumb. Exact costs depend on the layer types, memory traffic and hardware.
  • Forward mode is cheaper in memory, but reverse mode wins whenever there are more inputs than outputs.
  • Inference (prediction only) does not need the stored activations, which is why running a model needs much less memory than training it.
Quick check: one forward pass of a network costs $F$ multiply-adds. Roughly what does one full training step (forward, backward) cost?

The backward pass costs about $2F$ (one product for $\partial L/\partial W$ and one for $\partial L/\partial\mathbf{x}$ per layer), so the whole step costs about $3F$.

Automatic differentiation in action: a mini autodiff playground

Everything so far fits in a few dozen lines of code: parse a formula, turn it into a graph, run the forward pass, run the backward pass. This is what PyTorch and JAX do, at a much larger scale and speed. Here is the small version, running in your browser. Type a formula in $x$ and $y$ and watch the graph.

Type x*y + sin(x) with $x = 1$, $y = 2$.

  1. Graph: $v_1 = x\cdot y$, $v_2 = \sin x$, $f = v_1 + v_2$. Values: $v_1 = 2$, $v_2 = 0.8415$, $f = 2.8415$.
  2. Backward: $\bar f = 1$. The add node sends $1$ to both. $\bar v_1 = 1$, $\bar v_2 = 1$.
  3. $x$ gets $\bar v_1\cdot y = 2$ from the product and $\bar v_2\cdot\cos x = 0.5403$ from the sine, and they add: $\partial f/\partial x = 2.5403$.
  4. $y$ gets $\bar v_1\cdot x = 1$.
  5. Check with the formula: $\partial f/\partial x = y + \cos x = 2 + 0.5403$ ✓.

What the playground does (a miniature automatic-differentiation system):

  1. Parse the text into a tree (a safe parser written for this page; nothing is executed as code).
  2. Build the graph: every operation becomes a node; repeated variables become one node with fan-out.
  3. Forward pass: compute and store every value.
  4. Backward pass: local derivative × upstream gradient, added at fan-out (reverse mode).
  5. Forward mode as a second opinion: two runs with seeds $\dot x=1$ and $\dot y=1$.
  6. Compare with finite differences and with the symbolic derivative (tidied).

Supported: + − * / ^, exp log sin cos tanh sigmoid relu sqrt, numbers, pi, e, and the names x and y. Write 2*x, not 2x.

Why do we need it?

Seeing a whole AD system in miniature removes the last bit of mystery: there is no magic in loss.backward(), only a graph, stored values and local rules.

Where is it used?

The same design underlies PyTorch autograd, JAX, TensorFlow, Autograd, Tinygrad and Karpathy-style "micrograd" teaching code. Writing your own tiny version is a common way to learn backpropagation.

How is it used?

Type formulas, change $x$ and $y$, and compare the three derivative columns. Test tricky cases: a variable used many times, relu at a negative input, log at a negative input (no value), deep nesting.

Pick an example or type your own formula in the box (press Enter or just type). The graph shows each value; switch the display to the backward gradients (red) or to forward-mode derivatives (orange). The table compares reverse mode, forward mode, finite differences and the symbolic derivative: they agree. Try x*x*x, sigmoid(x*y), relu(x - y), exp(-x^2).

  • Real frameworks do the same thing with tensors (whole arrays) as the values, so each node is a big matrix operation.
  • At corners (like relu at exactly 0) the derivative is not defined; the engine picks 0 and a finite difference may disagree there.
  • The symbolic column is for comparison only: it can be long, while AD never builds it.
Quick check: for x*x the graph has the node $x$ used twice. What does the backward pass do at $x$?

The multiply node sends back $x$ to each of its two inputs (the other factor each time). They are the same node, so the gradients add: $x + x = 2x$.

Vanishing and exploding gradients core

Backpropagation multiplies one local derivative per layer. Multiply many numbers that are all a bit smaller than 1 and the result goes to nearly zero: the gradient vanishes, and the early layers barely learn. Multiply many numbers a bit bigger than 1 and it blows up: the gradient explodes, and training becomes unstable.

Think of whispering a message down a line of 30 people, each repeating it a little more quietly (vanishing), or a little more loudly (exploding). The first person's message does not reach the end in a useful form.

The sigmoid's slope $\sigma'(z) = \sigma(1-\sigma)$ is at most $0.25$ (at $z=0$). A chain of 10 sigmoid layers with weights near 1 multiplies the gradient by at most $0.25^{10}$:

  1. $0.25^2 = 0.0625$, $\;0.25^5 = 0.000977$, $\;0.25^{10} = 0.00000095 \approx 10^{-6}$.
  2. So the gradient reaching layer 1 is about a millionth of the gradient at the last layer.
  3. ReLU's slope is exactly $1$ for positive inputs, so $1^{10} = 1$: the signal passes (but blocks to $0$ where the input is negative).
  4. If the per-layer factor is $1.5$ instead, $1.5^{30} \approx 190{,}000$: the gradient explodes.

Along a chain, the gradient reaching layer $\ell$ from the loss is a product of one factor per layer in between:

$$\frac{\partial L}{\partial\mathbf{h}_{\ell}} = W_{\ell+1}^\top D_{\ell+1}\,W_{\ell+2}^\top D_{\ell+2}\cdots W_L^\top D_L\,\frac{\partial L}{\partial\mathbf{h}_L},\qquad D_k=\mathrm{diag}\big(\varphi'(\mathbf{z}_k)\big).$$

Its size behaves roughly like $(\text{typical weight size}\times\text{typical activation slope})^{\text{number of layers}}$. If that base is below 1 the gradient vanishes; above 1 it explodes.

Remedies. ReLU-type activations (slope 1 where active), careful weight initialisation (Xavier or He: scale the weights so the base is near 1), normalisation layers, residual connections (a "+1" path), gradient clipping (cap the size of an exploding gradient), and LSTM/GRU gates in recurrent networks.

Why do we need it?

It explains why very deep networks were hard to train for years, why sigmoid hidden layers fell out of favour, and why tricks like ResNets and normalisation matter.

Where is it used?

Diagnosing a network that "does not learn" (layer-wise gradient norms in TensorBoard or Weights & Biases), choosing initialisation and activations, and gradient clipping when training RNNs and Transformers.

How is it used?

Watch the gradient size per layer. If early layers have gradients many orders of magnitude smaller than late ones, change the activation or initialisation. If the loss suddenly becomes NaN or huge, clip gradients and lower the learning rate.

A random network with 10 neurons per layer. The curves show the size of the gradient reaching a layer, compared with the gradient at the last layer, against how many layers back from the loss you are (log scale: $-6$ means a millionth). The gradients here are computed by real backpropagation. Raise the depth: the sigmoid curve plunges. Press He scale ($\sqrt2$) so ReLU stays steady, then set the weight scale to $3$ and watch the gradients explode.

  • The numbers depend on the random weights, but the trend does not: a per-layer factor away from 1 compounds exponentially with depth.
  • ReLU can still lose signal: a neuron that is negative for every input has zero gradient forever ("dead ReLU").
  • Vanishing and exploding gradients are properties of the product, not of backprop itself; any method that multiplies many Jacobians has them.
Quick check: a 20-layer chain where each layer multiplies the gradient by 0.8. What reaches the first layer?

$0.8^{20} \approx 0.0115$: about 1% of the gradient at the last layer. With a factor of $0.5$ it would be $0.5^{20}\approx10^{-6}$.

Gradient checking: trust, but verify core

Backprop code is easy to get subtly wrong: a missing transpose, a forgotten factor, an overwritten gradient. The loss still goes down a little, so nothing crashes. The test that catches these bugs uses the oldest definition of a derivative: nudge the weight a little in each direction, and see how the loss changes. If it matches what backprop said, the code is right.

It is slow (two extra forward passes per weight), so you do it on a tiny model, a handful of weights, once, while developing. Then you switch it off.

Tiny network from before, weight $W_2[1] = 0.5$. Backprop said $\partial L/\partial W_2[1] = -0.5014$.

  1. Nudge: $W_2[1] = 0.5 + 10^{-5}$ gives $L(+) = 0.301483\ldots$ and $W_2[1] = 0.5 - 10^{-5}$ gives $L(-) = 0.301493\ldots$ (computed with the full forward pass each time).
  2. Central difference: $(L(+) - L(-))/(2\cdot10^{-5}) = -0.50136$.
  3. Relative error: $\dfrac{|a - n|}{|a| + |n|}$ with $a = -0.501362$ and $n = -0.501362$ (they agree to about 12 digits) is below $10^{-12}$. That is tiny, so the gradient passes.
  4. If we forget the sigmoid slope in the hidden layer, $\partial L/\partial W_1[1,1]$ comes out as $-0.3883$ instead of $-0.0888$, a relative error near $0.6$: the check fails and flags the bug.

Gradient check. For each weight $w_i$ compare the backprop value $a_i$ with the numerical value $n_i = \dfrac{L(w_i+h) - L(w_i-h)}{2h}$, using a small $h$ (about $10^{-5}$ to $10^{-6}$ in 64-bit arithmetic), via the relative error

$$\text{rel}_i = \frac{|a_i - n_i|}{|a_i| + |n_i|}.$$

Typical reading: below $10^{-7}$ excellent; around $10^{-5}$ fine for a non-smooth model; above $10^{-3}$ probably a bug. Practical rules: use 64-bit floats, check a few random weights, avoid points where the model has a corner (like ReLU at 0), and turn off dropout and other randomness while checking.

Why do we need it?

A wrong gradient does not crash the program; it silently trains a worse model. A gradient check is the cheapest way to find out whether your backward code matches the true derivative of your forward code.

Where is it used?

Whenever you write a custom layer or loss: torch.autograd.gradcheck, JAX's check_grads, the classic Stanford CS231n and Andrew Ng exercises, and unit tests of numerical libraries.

How is it used?

Pick a few weights. Compute the backprop gradient and the central difference for each. Print the relative error and require it to be tiny. Fix the bug and test again.

The table checks all 9 gradients of the tiny network against central differences. Choose a bug to inject into the backward pass and watch the relative errors jump from about $10^{-10}$ to large values. Then slide the step size $h$: if $h$ is large, truncation error appears; if it is tiny (below $10^{-9}$), rounding error appears. There is a sweet spot.

  • A passing check on a few weights is strong evidence, not proof. Check different weights and a few random inputs.
  • Do not use gradient checking during real training: it needs $2n$ forward passes.
  • If the check fails only for some weights, look at the layers they belong to (and at corners such as ReLU at $0$).
Quick check: why use the central difference $\frac{L(w+h)-L(w-h)}{2h}$ rather than $\frac{L(w+h)-L(w)}{h}$?

Its error shrinks like $h^2$ instead of $h$, so at the same $h$ it is much more accurate, and the relative error of a correct gradient is small enough to separate it from a bug.

What the gradient does: a loss surface in 3D

All of this work has one purpose: to know which way to move the weights so the loss goes down. Imagine the loss as a landscape: the two weights are the east and north positions, and the loss is the height. Backpropagation tells you the slope of the ground under your feet. Gradient descent takes a step downhill, and repeats.

One surprise: far out on a flat plateau the slope is almost zero, so the arrow is tiny and descent crawls. That is the vanishing gradient again, seen from above.

One sigmoid neuron: prediction $\hat y = \sigma(w x + b)$ with input $x = 2$, target $t = 1$ and loss $L = (\hat y - 1)^2$. At $(w, b) = (0, 0)$:

  1. Forward: $a = 2\cdot0 + 0 = 0$, $\hat y = \sigma(0) = 0.5$, $L = 0.25$.
  2. Backward: $\partial L/\partial\hat y = 2(\hat y - 1) = -1$; $\sigma'(0) = 0.25$; so $\partial L/\partial a = -0.25$.
  3. $\partial a/\partial w = x = 2$ and $\partial a/\partial b = 1$, so $\partial L/\partial w = -0.5$ and $\partial L/\partial b = -0.25$.
  4. Gradient descent with $\eta = 1$ moves to $(w,b) = (0.5, 0.25)$, where $a = 1.25$, $\hat y = 0.7773$ and $L = 0.0496$. The loss fell from $0.25$ to $0.0496$.

The gradient $\nabla L = (\partial L/\partial w,\ \partial L/\partial b)$ points in the direction of steepest increase. Gradient descent moves the opposite way:

$$(w, b) \leftarrow (w, b) - \eta\,\nabla L(w, b).$$

Backpropagation is how $\nabla L$ is computed for a network with millions of weights; the widget below uses our autodiff engine to get it for this neuron.

Why do we need it?

The gradient is the compass of training. Seeing it on a surface connects the abstract numbers from the backward pass to the actual downhill move the weights make.

Where is it used?

Every optimiser (SGD, momentum, Adam) takes the backprop gradient and decides how to step. Plateaus, valleys and saddles of real loss surfaces explain slow training.

How is it used?

Compute the gradient with backprop, step against it with a learning rate $\eta$, and check that the loss falls. Too large a step overshoots; too small crawls.

The sheet is the loss $L(w,b)$ of the one-neuron model. Drag the blue dot on the floor to choose $(w,b)$ (hold Shift to move it up, though it snaps back to the floor). The orange arrow on the floor is the gradient (uphill) from backprop. Press Descend to follow gradient descent from the dot (green path). Drag to the far left, where the loss is about 1: the arrow is almost zero and descent crawls. Rotate and use the Top and Front buttons.

Quick check: at $(w,b)=(0,0)$ the gradient is $(-0.5,-0.25)$. Which way does gradient descent move $w$ and $b$ with $\eta = 1$?

$w \leftarrow 0 - 1\cdot(-0.5) = 0.5$ and $b \leftarrow 0 - 1\cdot(-0.25) = 0.25$. Both increase, because the loss falls when $w$ and $b$ increase (the neuron needs a larger output to reach the target $1$).

Recap, cheat sheet and practice

  • Three ways to differentiate: symbolic (a formula, may swell), numerical (approximate, step-size trouble, good for checking) and automatic (exact, follows the program, cheap).
  • Computational graph: one elementary operation per node. The forward pass computes and stores every value. Each node has a local derivative that is a one-line rule.
  • Backward pass: set $\partial L/\partial L = 1$, then in reverse order send (upstream gradient) × (local derivative) to each input, adding where a value is used twice. This is reverse-mode differentiation.
  • Backpropagation = the chain rule applied from the output backwards, with the shared partial products saved and reused. Layer rules: dense $\partial W = \mathbf{g}\mathbf{x}^\top$, $\partial\mathbf{x} = W^\top\mathbf{g}$; activation $\mathbf{g}\odot\varphi'(\mathbf{z})$.
  • Forward mode runs the program on dual numbers $a+b\varepsilon$ ($\varepsilon^2=0$) and costs one pass per input; reverse mode costs one pass per output, so a scalar loss needs only one backward pass.
  • Cost (rule of thumb): backward $\approx 2$–$3\times$ forward in time; memory = the stored activations (checkpointing trades time for memory).
  • Trouble and checks: a product of many slopes below (above) 1 makes gradients vanish (explode); ReLU, good initialisation, residual links, normalisation and clipping help. Gradient checking with central differences catches backward-pass bugs.

Cheat sheet

IdeaFormula / ruleRemember
Backward step$\bar v_j \mathrel{+}= \bar v_i\,\dfrac{\partial v_i}{\partial v_j}$upstream × local, add at fan-out
Add nodepasses the gradient to both inputslocal derivatives 1, 1
Multiply nodesends back the other input$\partial(ab)/\partial a = b$
Sigmoid / tanh / ReLU$\sigma(1-\sigma)$ / $1-\tanh^2$ / $1$ or $0$reuse the stored output
Dense layer$\partial W=\mathbf{g}\mathbf{x}^\top,\ \partial\mathbf{b}=\mathbf{g},\ \partial\mathbf{x}=W^\top\mathbf{g}$gradient has the shape of the thing
Activation layer$\partial\mathbf{z}=\mathbf{g}\odot\varphi'(\mathbf{z})$entry-wise product
Dual numbers$f(a+\varepsilon)=f(a)+f'(a)\varepsilon$nudge part = derivative
Forward vs reverse$n$ passes vs $m$ passesloss: $m=1$, use reverse
Training step cost$\approx 3\times$ one forward pass (rule of thumb)memory = stored activations
Gradient check$\dfrac{|a-n|}{|a|+|n|}$ with $n=\dfrac{L(w+h)-L(w-h)}{2h}$tiny model, 64-bit, $h\approx10^{-5}$
Code it · Python (a micro-autograd, about 40 lines)

import math

class Value:
    """A number that remembers how it was made (a node of the computational graph)."""
    def __init__(self, data, parents=(), local=()):
        self.data = data
        self.grad = 0.0
        self.parents = parents      # nodes this one was computed from
        self.local = local          # local derivatives d(self)/d(parent)

    def __add__(self, o):
        o = o if isinstance(o, Value) else Value(o)
        return Value(self.data + o.data, (self, o), (1.0, 1.0))
    def __mul__(self, o):
        o = o if isinstance(o, Value) else Value(o)
        return Value(self.data * o.data, (self, o), (o.data, self.data))
    def __sub__(self, o): return self + (o * -1.0)
    def __pow__(self, k): return Value(self.data ** k, (self,), (k * self.data ** (k - 1),))
    def sigmoid(self):
        s = 1 / (1 + math.exp(-self.data))
        return Value(s, (self,), (s * (1 - s),))
    def relu(self):
        return Value(max(0.0, self.data), (self,), (1.0 if self.data > 0 else 0.0,))

    def backward(self):
        order, seen = [], set()
        def visit(v):                       # topological order: parents before children
            if id(v) not in seen:
                seen.add(id(v))
                for p in v.parents: visit(p)
                order.append(v)
        visit(self)
        self.grad = 1.0                     # dL/dL = 1
        for v in reversed(order):           # walk backwards
            for p, loc in zip(v.parents, v.local):
                p.grad += v.grad * loc      # upstream * local, ADDED at fan-out

# --- the tiny 2-2-1 network of this chapter, built from Value objects ---
def loss(params):
    W1a, W1b, W1c, W1d, b1a, b1b, W2a, W2b, b2 = params
    x1, x2, t = 1.0, 2.0, 1.0
    z1 = W1a * x1 + W1b * x2 + b1a
    z2 = W1c * x1 + W1d * x2 + b1b
    y = W2a * z1.sigmoid() + W2b * z2.sigmoid() + b2
    return (y - t) ** 2 * 0.5

vals = [0.1, 0.2, -0.3, 0.4, 0.1, -0.1, 0.5, -0.5, 0.2]
params = [Value(v) for v in vals]
L = loss(params)
L.backward()
print(round(L.data, 4))                              # 0.3015
print([round(p.grad, 4) for p in params])
# [-0.0888, -0.1777, 0.0933, 0.1866, -0.0888, 0.0933, -0.5014, -0.4649, -0.7765]

# --- gradient check: central finite differences ---
def numeric_grad(vals, i, h=1e-6):
    up = [Value(v + (h if j == i else 0)) for j, v in enumerate(vals)]
    dn = [Value(v - (h if j == i else 0)) for j, v in enumerate(vals)]
    return (loss(up).data - loss(dn).data) / (2 * h)

err = max(abs(p.grad - numeric_grad(vals, i)) for i, p in enumerate(params))
print(err < 1e-8)                                    # True: backprop matches finite differences

# --- fan-out: f = x*x uses x twice, so the two gradients ADD ---
x = Value(3.0)
f = x * x
f.backward()
print(f.data, x.grad)                                # 9.0 6.0
Test yourself

1. Which method gives an exact derivative (up to rounding), works on ordinary program code with loops and if-statements, and costs only a small multiple of one function evaluation for a scalar loss?

Symbolic needs a closed formula and can swell. Finite differences are approximate and need $2n$ evaluations. Reverse-mode AD is exact and costs about 2–4 evaluations for any $n$.

2. In the backward pass, a multiply node $v = a\cdot b$ receives upstream gradient $g$. What does it send to input $a$?

The local derivative $\partial v/\partial a$ is $b$. Upstream times local: $g\cdot b$. (And $g\cdot a$ goes to $b$.)

3. A value is used in two places. The gradients coming back along the two edges are $3$ and $4$. What is the gradient of that value?

Each use is a separate path to the loss, and the multivariate chain rule adds paths: $3 + 4 = 7$.

4. Roughly how does the cost of the backward pass compare with the forward pass for a typical network?

For each dense layer the backward pass does two matrix products (for $\partial W$ and for $\partial\mathbf{x}$), each as big as the forward product, so about 2× the forward cost (the whole step is about 3×). It does not grow with the number of weights beyond the network size.

5. With dual numbers ($\varepsilon^2 = 0$), what is $(2+\varepsilon)(5+3\varepsilon)$?

$2\cdot5 + (2\cdot3 + 1\cdot5)\varepsilon + 3\varepsilon^2 = 10 + 11\varepsilon$. The nudge part $2\cdot3 + 1\cdot5$ is the product rule.

6. The slope of the sigmoid is at most $0.25$. Through 12 sigmoid layers (weights of size about 1), about how small can the gradient become at the first layer relative to the last?

One factor of at most $0.25$ per layer multiplies: $0.25^{12} \approx 6\times10^{-8}$. That is a vanishing gradient.

Practice problems

A. A neuron has $p = w\cdot x$, $e = p - t$, $L = e^2$ with $w = 2$, $x = 3$, $t = 5$. Run the forward and backward passes and give $\partial L/\partial w$ and $\partial L/\partial x$.

Forward: $p = 6$, $e = 1$, $L = 1$. Backward: $\bar L = 1$; $\bar e = 2e = 2$; $\bar p = \bar e\cdot1 = 2$; $\bar w = \bar p\cdot x = 6$ and $\bar x = \bar p\cdot w = 4$. Check: $L = (wx-t)^2$ gives $\partial L/\partial w = 2(wx-t)x = 2\cdot1\cdot3 = 6$ ✓ and $\partial L/\partial x = 2\cdot1\cdot2 = 4$ ✓.

B. Use dual numbers to find $f(2)$ and $f'(2)$ for $f(x) = x^3 - 2x$.

$(2+\varepsilon)^3 = 8 + 12\varepsilon$ (higher powers vanish) and $2(2+\varepsilon) = 4 + 2\varepsilon$. So $f(2+\varepsilon) = (8-4) + (12-2)\varepsilon = 4 + 10\varepsilon$. Thus $f(2) = 4$ and $f'(2) = 10 = 3\cdot4 - 2$ ✓.

C. A deep chain of 8 tanh layers has an average factor of $0.7$ per layer. What fraction of the gradient reaches the first layer? What if the factor were $1.3$?

$0.7^8 \approx 0.058$: under 6% of the gradient survives (vanishing, mildly). With $1.3$: $1.3^8 \approx 8.2$: the gradient is amplified 8-fold (exploding, mildly). Over 40 layers the same factors give $0.7^{40}\approx6\times10^{-7}$ and $1.3^{40}\approx36{,}000$.

D. A dense layer has $W = \begin{bmatrix}2 & 0\\1 & -1\end{bmatrix}$, input $\mathbf{x} = (1, 3)$ and upstream gradient $\mathbf{g} = (1, 2)$. Find $\partial L/\partial W$ and $\partial L/\partial\mathbf{x}$.

$\partial L/\partial W = \mathbf{g}\mathbf{x}^\top = \begin{bmatrix}1\cdot1 & 1\cdot3\\ 2\cdot1 & 2\cdot3\end{bmatrix} = \begin{bmatrix}1 & 3\\2 & 6\end{bmatrix}$. $\partial L/\partial\mathbf{x} = W^\top\mathbf{g} = \begin{bmatrix}2 & 1\\0 & -1\end{bmatrix}\begin{bmatrix}1\\2\end{bmatrix} = (2+2,\ 0-2) = (4, -2)$. Shapes: $2\times2$ and $2\times1$ ✓.

E. A model has $10^7$ weights and one loss. Compare the number of forward-pass-equivalents needed for the full gradient by central finite differences and by backpropagation (take a backward pass as about 3 forward passes).

Finite differences: $2\times10^7$ evaluations. Backprop: about 3. The ratio is about $7\times10^6$: backprop is millions of times cheaper, and exact.

F. A gradient check for one weight gives backprop $= 0.2500$ and central difference $= 0.5000$. Compute the relative error and say what it suggests.

$\text{rel} = \dfrac{|0.25 - 0.5|}{0.25 + 0.5} = \dfrac{0.25}{0.75} = 0.33$. That is far above $10^{-3}$, so the backward code is wrong. The numerical value is exactly twice the backprop value, which suggests a missing factor of 2 (for example the derivative of $r^2$ was taken as $r$).

Chapter 2.10

Higher-Order Derivatives

The derivative tells you how steep the ground is. The second derivative tells you how the steepness itself is changing: is the ground curving up like a bowl, curving down like a hill, or twisting like a mountain pass? That one extra piece of information decides whether a flat spot is a valley, and how big a step training can safely take.

  • Read the second derivative as concavity (smile or frown) and as acceleration; meet the third and higher derivatives
  • Build the Hessian matrix of second partial derivatives, and know when mixed partials are equal
  • Measure curvature in any direction with $\mathbf{u}^\top H\mathbf{u}$ and read the Hessian through its eigenvalues
  • Tell a bowl (positive definite), a hill (negative definite) and a saddle (indefinite) apart, and apply the second-derivative test
  • Use curvature to take a smarter step (Newton's method) and to choose a safe learning rate ($\eta \lt 2/\lambda_{\max}$)

The second derivative: concavity and acceleration core

Think about driving. Your position tells where you are. Your speed tells how fast the position is changing (that is the derivative, from Chapter 2.3). Your acceleration tells how fast the speed is changing: press the gas and it is positive, brake and it is negative.

Acceleration is "the derivative of the derivative". That is exactly what a second derivative is. Nothing new is needed: you just differentiate twice.

Now picture a curve as a road. The first derivative is how steep the road is right now. The second derivative says whether the road is getting steeper (curving upward like a smile) or flattening out (curving downward like a frown).

A curve. Let $f(x) = x^3$.

  1. First derivative (power rule): $f'(x) = 3x^2$.
  2. Differentiate again: $f''(x) = 6x$.
  3. At $x = 2$: $f''(2) = 12 \gt 0$, so the curve is a smile there (the slope is growing).
  4. At $x = -2$: $f''(-2) = -12 \lt 0$, a frown (the slope is shrinking).
  5. At $x = 0$: $f''(0) = 0$. The curve switches from frown to smile right here.

Numeric check by nudging $x$ by $0.01$: the slope at $2.01$ is $3(2.01)^2 = 12.1203$ and at $1.99$ it is $3(1.99)^2 = 11.8803$. The slope changes by $0.24$ over a step of $0.02$, so the rate of change of the slope is $0.24/0.02 = 12$ ✓.

A falling ball. A dropped ball has fallen $s(t) = 4.9\,t^2$ metres after $t$ seconds. Speed: $s'(t) = 9.8\,t$ metres per second. Acceleration: $s''(t) = 9.8$ metres per second per second, a constant (gravity).

The second derivative of $f$ is the derivative of its derivative:

$$f''(x) \;=\; \frac{d}{dx}\bigl(f'(x)\bigr) \;=\; \lim_{h\to 0}\frac{f'(x+h) - f'(x)}{h}, \qquad\text{also written}\quad \frac{d^2 f}{dx^2}.$$

The notation $\dfrac{d^2 f}{dx^2}$ reads "d two f over d x squared": apply $\frac{d}{dx}$ twice.

  • $f''(x) \gt 0$: the graph is concave up (a smile; it would "hold water"). The slope is increasing.
  • $f''(x) \lt 0$: the graph is concave down (a frown). The slope is decreasing.
  • A point where $f''$ changes sign is an inflection point (the bend switches direction).

How to estimate it from values only (derivation). The slope halfway between $x$ and $x+h$ is about $\frac{f(x+h)-f(x)}{h}$, and the slope halfway between $x-h$ and $x$ is about $\frac{f(x)-f(x-h)}{h}$. These two slopes are $h$ apart, so the rate of change of the slope is their difference divided by $h$:

$$f''(x) \approx \frac{1}{h}\left(\frac{f(x+h)-f(x)}{h} - \frac{f(x)-f(x-h)}{h}\right) = \frac{f(x+h) - 2f(x) + f(x-h)}{h^2}.$$

Check with $f = x^3$, $x = 2$, $h = 0.1$: $\frac{2.1^3 - 2\cdot 8 + 1.9^3}{0.01} = \frac{9.261 - 16 + 6.859}{0.01} = \frac{0.12}{0.01} = 12$ ✓.

Why do we need it?

The slope alone cannot tell a valley from a hilltop: both have slope zero. Whether the slope is growing or shrinking around a flat spot is what tells them apart. It also tells us whether we are speeding up or slowing down.

Where is it used?

Physics (acceleration), checking whether a loss function is convex ($f'' \ge 0$ everywhere), the second-derivative test for minima, Newton's method, and limits on the learning rate in gradient descent.

How is it used?

Differentiate twice and look at the sign: positive means smile (the bottom of a bowl is nearby), negative means frown. The size tells how quickly the slope changes. In code, use the three-point formula above or an autodiff library.

Drag the dot along the top curve. The dashed line follows you down through the slope graph $f'$ and the curvature graph $f''$. Blue pieces are smiles ($f'' \ge 0$, slope going up); orange pieces are frowns ($f'' \lt 0$, slope going down). Notice that where $f'$ has a peak or dip, $f''$ crosses zero, and the top curve changes colour. Pick another function from the menu.

Move the time slider (or press Play). The car's position is $s(t) = t^3 - 6t^2 + 9t + 1$. Blue arrow = velocity $s'$, orange arrow = acceleration $s''$. Watch three moments: at $t=1$ the car stops (speed 0) and turns around, even though the acceleration is still pulling backward. At $t=2$ the acceleration is 0. Whenever speed and acceleration have the same sign, the car is speeding up.

$f''=0$ does not always mean an inflection point. For $f(x)=x^4$ we get $f''(x)=12x^2$, which is $0$ at $x=0$, yet the curve is a smile on both sides. An inflection needs $f''$ to change sign.

"Concave up" is about the bend, not about being positive. A curve can be below the axis and still smile. The sign of $f''$ only describes how the slope changes.

Quick check: $f(x) = x^2 - 4x$. What is $f''$, and what does it say about the shape?

$f'(x) = 2x - 4$ and $f''(x) = 2$. It is positive everywhere, so $f$ is a smile (a parabola opening upward) at every point, with its lowest point where $f' = 0$, at $x = 2$.

The third derivative and higher-order derivatives

Nothing stops us from differentiating again. The derivative of the acceleration is called the jerk: it is the sudden jolt you feel when a driver stamps on the brake. A smooth ride has a small jerk. That is the third derivative.

And we can keep going: fourth, fifth, and so on. Each one tells how fast the previous one is changing. It is like a ladder: each rung is the slope of the rung below.

Some functions run out of rungs quickly (a polynomial eventually reaches 0). Some never change ($e^x$). Some repeat in a loop ($\sin x$).

Differentiate three functions again and again.

order$x^4$$e^x$$\sin x$
$f$$x^4$$e^x$$\sin x$
$f'$$4x^3$$e^x$$\cos x$
$f''$$12x^2$$e^x$$-\sin x$
$f'''$$24x$$e^x$$-\cos x$
$f^{(4)}$$24$$e^x$$\sin x$ (back to the start)
$f^{(5)}$$0$$e^x$$\cos x$

Notice: $x^4$ reaches a constant, then zero. $e^x$ is its own derivative forever. $\sin x$ repeats every 4 steps. The fourth derivative of $x^4$ is $24 = 4\cdot3\cdot2\cdot1 = 4!$ (read "four factorial": the product of all whole numbers from 4 down to 1). Factorials will matter in Chapter 2.11.

The $n$-th derivative is the derivative taken $n$ times:

$$f^{(n)}(x) = \frac{d^n f}{dx^n}, \qquad f^{(0)} = f,\quad f^{(1)} = f',\quad f^{(2)} = f'',\quad f^{(3)} = f''' ,\ \dots$$

The small number in brackets is the order, so that it is not mistaken for a power. Useful patterns:

  • A polynomial of degree $n$ has $f^{(n)} = n!\times(\text{leading coefficient})$, and every derivative after that is $0$.
  • $\dfrac{d^n}{dx^n}e^x = e^x$, and $\sin$/$\cos$ cycle with period 4.
  • For $\ln(1+x)$: $f' = \frac{1}{1+x}$, $f'' = -\frac{1}{(1+x)^2}$, $f''' = \frac{2}{(1+x)^3}$, and in general $f^{(n)} = \dfrac{(-1)^{n-1}(n-1)!}{(1+x)^n}$.

A function whose derivatives all exist is called smooth. $\mathrm{ReLU}(x) = \max(0,x)$ is not smooth: it has a kink at $0$. Smooth replacements such as softplus and GELU exist partly for this reason.

Why do we need it?

Each higher derivative adds one more detail about the shape of a function near a point: height, slope, bend, how the bend changes... Together they let us rebuild the function nearby. That is the idea behind Taylor series (next chapter).

Where is it used?

Taylor series (the coefficient of $x^n$ uses $f^{(n)}(0)/n!$), smoothness requirements of optimisers, jerk limits in robot and vehicle motion planning, and comparing ReLU (kinked) with smooth activations such as GELU or softplus.

How is it used?

Differentiate repeatedly and look for a pattern (it often repeats or ends). On a computer, symbolic tools (SymPy diff(f, x, n)) or autodiff do it exactly; finite differences get noisy very quickly as the order grows.

Pick a function and move the order slider from 0 up to 6. The faint grey curve is $f$ itself; the blue curve is the derivative of the chosen order. For $x^4$ it flattens to a constant and then to zero. For $\sin x$ the blue curve cycles. For $e^x$ it never changes. Which order of $\ln(1+x)$ blows up fastest near $x=-1$?

Do not use many finite differences in a row. Each nudge divides a tiny number by another tiny number, so rounding noise grows with the order. Fourth and fifth derivatives from raw values are usually garbage. Use symbolic maths or autodiff.

Quick check: what is the 10th derivative of $x^9$?

$0$. A polynomial of degree 9 has $f^{(9)} = 9!$ (a constant), so $f^{(10)} = 0$.

The Hessian and the Hessian matrix core

With one input there is one slope and one bend. With two inputs (a landscape with east-west $x$ and north-south $y$), the ground has more to say. Walk east: does the slope eastward grow or shrink? Walk north: does the slope northward change? And there is a cross effect: as you walk north, does the east-west slope change? (A twisted roof tile does this.)

Each of those questions is a second partial derivative. For two inputs there are $2\times2 = 4$ of them. We write them in a small table, and that table is the Hessian. It is the "second derivative" for functions of many variables.

In the gradient chapter you collected the first partial derivatives into a column. Now we collect the second ones into a square.

Let $f(x,y) = x^2 y + y^2$. Find the Hessian at the point $(1, 2)$.

  1. First partials: $f_x = \dfrac{\partial f}{\partial x} = 2xy$ and $f_y = \dfrac{\partial f}{\partial y} = x^2 + 2y$.
  2. Differentiate $f_x = 2xy$ with respect to $x$: $f_{xx} = 2y$. With respect to $y$: $f_{xy} = 2x$.
  3. Differentiate $f_y = x^2 + 2y$ with respect to $x$: $f_{yx} = 2x$. With respect to $y$: $f_{yy} = 2$.
  4. At $(1,2)$: $f_{xx} = 2\cdot2 = 4$, $f_{xy} = f_{yx} = 2\cdot1 = 2$, $f_{yy} = 2$.
$$H(1,2) = \begin{bmatrix} 4 & 2 \\ 2 & 2 \end{bmatrix}.$$

Notice that the two off-diagonal entries are equal. That is not an accident (next section).

For $f:\mathbb{R}^n \to \mathbb{R}$ the Hessian matrix is the $n\times n$ table of all second partial derivatives:

$$H(\mathbf{x}) = \nabla^2 f(\mathbf{x}) = \begin{bmatrix} \dfrac{\partial^2 f}{\partial x_1^2} & \dfrac{\partial^2 f}{\partial x_1\partial x_2} & \cdots \\[2mm] \dfrac{\partial^2 f}{\partial x_2\partial x_1} & \dfrac{\partial^2 f}{\partial x_2^2} & \cdots \\ \vdots & \vdots & \ddots \end{bmatrix}, \qquad H_{ij} = \frac{\partial^2 f}{\partial x_i\,\partial x_j}.$$
  • It is the Jacobian of the gradient. The gradient $\nabla f$ (a column vector, as always in this guide) is a function from $\mathbb{R}^n$ to $\mathbb{R}^n$. Its Jacobian (an $n\times n$ matrix; Chapter 2.5) has entry $(i,j) = \partial(\partial f/\partial x_i)/\partial x_j$, which is exactly $H_{ij}$.
  • Shape: $n\times n$. For $n = 1$ it is the $1\times1$ matrix $[f''(x)]$, so the Hessian really is the second derivative again.
  • It is symmetric for the smooth functions of ML ($H_{ij} = H_{ji}$, see next section), so it has real eigenvalues and perpendicular eigenvectors (Linear Algebra: eigenvalues).
  • The symbol $\nabla^2 f$ is "del squared f". (Some books use it for the Laplacian, the trace of $H$. Here it always means the Hessian.)
Why do we need it?

The gradient only knows the slope. To know whether the ground curves up or down, and in which directions, we need all the second derivatives. One number is not enough once there are many inputs, so we need a matrix.

Where is it used?

Newton's method and quasi-Newton optimisers (L-BFGS), the second-derivative test, the Laplace approximation in Bayesian models, Fisher information, second-order Taylor expansions, and analyses of why training is slow or unstable.

How is it used?

Differentiate the gradient again, entry by entry, and arrange the results in an $n\times n$ symmetric matrix. Then look at its eigenvalues. In code: jax.hessian(f)(x), torch.autograd.functional.hessian, or a finite-difference grid.

Drag the blue point (it snaps to a grid so the numbers stay tidy) and read the Hessian. First choose $x^2y + y^2$ and move to $(1, 2)$: you should get $\begin{bmatrix}4&2\\2&2\end{bmatrix}$ like the worked example. Then try other functions: for the bowl the Hessian never changes; for the egg box it changes from place to place. The orange arrow is the gradient. The last line checks the exact formulas against a finite-difference estimate.

The Hessian has $n^2$ entries. A network with a million weights would need a table of $10^{12}$ numbers (about 4 000 gigabytes, i.e. 4 terabytes, at 4 bytes each, and much more work to invert). That is why deep learning mostly avoids the full Hessian and uses first-order methods, or tricks such as Hessian-vector products that never build the whole matrix.

The Hessian is not the Jacobian of a vector function. It belongs to a scalar function $f:\mathbb{R}^n\to\mathbb{R}$, such as a loss. (A vector function has a Jacobian, an $m\times n$ matrix, and no single Hessian.)

Quick check: for $f(x,y)=x^2+3xy+2y^2$, what is $H$?

$f_x = 2x+3y$, $f_y = 3x+4y$. So $f_{xx} = 2$, $f_{xy} = 3$, $f_{yx} = 3$, $f_{yy} = 4$, giving $H = \begin{bmatrix}2&3\\3&4\end{bmatrix}$ at every point (the function is a quadratic, so its second derivatives are constants).

Mixed partial derivatives: does the order matter?

A mixed partial derivative takes one derivative with respect to $x$ and one with respect to $y$. There are two ways to do it: differentiate in $x$ first and then $y$, or in $y$ first and then $x$. Does it matter?

Think of the ground as a twisted sheet. The "twist" is one single fact about the sheet. You can measure it by asking "how does the east-west slope change as I go north?" or "how does the north-south slope change as I go east?". Both questions measure the same twist, so the answers agree.

There is a neat picture. Take four corners of a small square: $(x,y)$, $(x+h,y)$, $(x,y+k)$, $(x+h,y+k)$. The quantity $f(x+h,y+k) - f(x+h,y) - f(x,y+k) + f(x,y)$ is "(top edge difference) minus (bottom edge difference)" and also "(right edge difference) minus (left edge difference)". It is one number, however you group it. Divided by $hk$ and shrunk to a point, it becomes the mixed partial.

Let $f(x,y) = x^2y^3$.

  1. $x$ first: $f_x = 2xy^3$. Then $y$: $f_{xy} = \dfrac{\partial}{\partial y}(2xy^3) = 6xy^2$.
  2. $y$ first: $f_y = 3x^2y^2$. Then $x$: $f_{yx} = \dfrac{\partial}{\partial x}(3x^2y^2) = 6xy^2$.
  3. Same answer: $f_{xy} = f_{yx} = 6xy^2$ ✓.

(Notation warning: $f_{xy}$ usually means "$x$ first, then $y$", and $\frac{\partial^2 f}{\partial y\,\partial x}$ is written right to left. Because the two agree for smooth functions, you rarely need to worry.)

Equality of mixed partials (Schwarz / Clairaut). If the second partial derivatives $f_{xy}$ and $f_{yx}$ exist and are continuous around a point, then at that point

$$\frac{\partial^2 f}{\partial x\,\partial y} = \frac{\partial^2 f}{\partial y\,\partial x}.$$

For a function of $n$ variables this says $H_{ij} = H_{ji}$: the Hessian is a symmetric matrix.

The continuity condition cannot be dropped. The function $f(x,y) = \dfrac{xy(x^2-y^2)}{x^2+y^2}$ (with $f(0,0)=0$) is a famous counterexample: at the origin $f_{xy}(0,0) = -1$ but $f_{yx}(0,0) = +1$. Why: along the $y$-axis ($x=0$) its $x$-slope is $f_x(0,y) = -y$, so differentiating that in $y$ gives $-1$; along the $x$-axis its $y$-slope is $f_y(x,0) = x$, so differentiating in $x$ gives $+1$. Its second derivatives are not continuous at the origin.

Why do we need it?

It cuts the work nearly in half and guarantees a nice matrix. A symmetric Hessian has real eigenvalues and perpendicular eigenvectors, so "curvature along each axis" is a meaningful picture.

Where is it used?

Every optimiser that uses the Hessian (Newton, L-BFGS, natural gradient), covariance and Fisher matrices (also symmetric), and checking hand-derived gradients in code by confirming that the Hessian comes out symmetric.

How is it used?

Compute only the upper triangle and copy it. If a computed Hessian is not symmetric, there is a bug (or the function is not smooth, as with ReLU networks at a kink).

Drag the point (or press Go to the origin). The two numbers $A$ and $B$ are found purely by nudging $f$ (no formulas): $A$ = nudge $y$ of the $x$-slope, $B$ = nudge $x$ of the $y$-slope. For the smooth functions they always agree. Now choose the famous counterexample and go to the origin: $A = -1$ but $B = +1$. Move away from the origin and they agree again (as long as the nudge $k$ is smaller than your distance from the origin).

Equal mixed partials are a statement about smooth functions. Polynomials, $e^x$, $\log$, sigmoid, tanh, softmax and squared error are all smooth, so their Hessians are symmetric. Networks that use ReLU have kinks, so their second derivatives are only defined away from the kinks (and the ReLU itself adds no curvature there, because it is a straight line on each side).

Quick check: $f(x,y) = e^{xy}$. Find $f_{xy}$ and $f_{yx}$.

$f_x = y\,e^{xy}$. Then $f_{xy} = e^{xy} + y\cdot x\,e^{xy} = (1+xy)e^{xy}$ (product rule). The other order: $f_y = x\,e^{xy}$ and $f_{yx} = e^{xy} + xy\,e^{xy} = (1+xy)e^{xy}$. Equal ✓.

Curvature: how sharply does it bend?

Drive around a bend. A gentle bend is part of a very big circle; a sharp hairpin is part of a very small circle. Curvature is "how tight the circle is". We pick the one circle that hugs the curve best at your spot, the curvature circle (also called the osculating circle, from the Latin for "kissing"). Its radius $R$ says how gently the road bends: big $R$ means gentle, small $R$ means sharp.

A straight road has an infinitely big circle (radius $\infty$): curvature zero.

In machine learning "curvature" almost always means the second derivative. At a flat spot ($f' = 0$) the two ideas are the same number.

The parabola $f(x) = \tfrac14 x^2$ at its bottom, $x = 0$.

  1. $f'(x) = \tfrac12 x$, so $f'(0) = 0$ (flat).
  2. $f''(x) = \tfrac12$.
  3. The curvature is $\kappa = \dfrac{f''}{(1 + f'^2)^{3/2}} = \dfrac{0.5}{(1+0)^{3/2}} = 0.5$.
  4. The curvature circle has radius $R = 1/\kappa = 2$. Its centre is $2$ above the bottom, at $(0, 2)$.

Away from the bottom, the slope is not zero. At $x = 2$: $f' = 1$, $f'' = 0.5$, so $\kappa = \dfrac{0.5}{(1+1)^{3/2}} = \dfrac{0.5}{2.828} = 0.177$ and $R \approx 5.66$. The parabola flattens out, so the circle gets bigger even though $f''$ stayed the same: the formula divides out the tilt.

For a curve $y = f(x)$ the signed curvature and the radius of curvature are

$$\kappa(x) = \frac{f''(x)}{\bigl(1 + f'(x)^2\bigr)^{3/2}}, \qquad R(x) = \frac{1}{|\kappa(x)|}.$$

$\kappa \gt 0$ means the curve bends upward (a smile), $\kappa \lt 0$ downward. The centre of the circle lies on the concave side, at distance $R$ from the curve.

At a critical point ($f'=0$) the denominator is $1$, so $\kappa = f''$. This is why, in optimisation, we say "the curvature is $f''$". Away from flat spots the two differ by the tilt factor $(1+f'^2)^{3/2}$, but $f''$ still tells the sign and how fast the slope changes, which is what a step of gradient descent feels.

Why do we need it?

We want one number for "how sharply does this bend?". A sharp bend means a small safe step; a gentle bend means we can stride. Curvature gives a picture (a circle) for the second derivative.

Where is it used?

Road and rail design, computer graphics (smooth curves), robot path planning, and in ML the idea of sharp versus flat minima: a very sharp bowl is hard for gradient descent and may generalise worse.

How is it used?

Compute $f'$ and $f''$ at the point and plug into $\kappa$. For optimisation at a flat spot just read $f''$: large positive means a narrow, steep bowl; small positive means a wide, gentle bowl.

Drag the dot along the curve. The circle is the best-fitting circle at that spot; its radius is $1/|\kappa|$. Blue circle = smile (centre above the curve), orange = frown (centre below). Look for the places where the circle gets enormous (an inflection or a long flat stretch). On $\sin x$, find where the circle is smallest (the peaks and troughs).

Quick check: a sharp bowl has $f'' = 10$ at its bottom; a wide one has $f''=0.1$. Which circle is bigger?

At the bottom $f'=0$, so $\kappa = f''$. Sharp bowl: $R = 1/10 = 0.1$. Wide bowl: $R = 1/0.1 = 10$. The wide bowl has the much bigger circle.

What the Hessian tells you: curvature in every direction core

Stand on a hilly landscape and pick a compass direction. Walk a short way along it and watch your height: it traces a curve, a smile or a frown. The curvature of that curve depends on the direction you chose. In a mountain pass it smiles along the ridge (towards the peaks) and frowns along the path that crosses the pass.

The Hessian packs the curvature of all directions into one matrix. Hand it a direction $\mathbf{u}$, and it gives back the curvature $\mathbf{u}^\top H\mathbf{u}$. Among all directions there is a most curved one and a least curved one: they are the eigenvectors of $H$, and the amounts of curvature are the eigenvalues.

Let $H = \begin{bmatrix}1&2\\2&1\end{bmatrix}$. Walk in three directions (unit vectors $\mathbf{u}$):

  • East, $\mathbf{u} = [1, 0]$: $\mathbf{u}^\top H\mathbf{u} = 1\cdot1\cdot1 = 1$. (Just the top-left entry.)
  • Northeast, $\mathbf{u} = [1,1]/\sqrt2$: $\mathbf{u}^\top H\mathbf{u} = \tfrac12(1 + 2 + 2 + 1) = 3$. A strong smile.
  • Southeast, $\mathbf{u} = [1,-1]/\sqrt2$: $\tfrac12(1 - 2 - 2 + 1) = -1$. A frown.

The eigenvalues of $H$ are $3$ and $-1$, with eigenvectors along those last two diagonals. So this ground smiles most (3) along one diagonal, frowns most ($-1$) along the other, and everything else lies in between.

Curvature along a direction. Walk from $\mathbf{x}$ in the unit direction $\mathbf{u}$ and let $g(t) = f(\mathbf{x} + t\mathbf{u})$ be the height after $t$ steps. Derive its second derivative:

  1. The chain rule gives the slope: $g'(t) = \nabla f(\mathbf{x}+t\mathbf{u})^\top\mathbf{u} = \sum_i f_i(\mathbf{x}+t\mathbf{u})\,u_i$, where $f_i = \partial f/\partial x_i$.
  2. Differentiate again. Each $f_i(\mathbf{x}+t\mathbf{u})$ changes at rate $\sum_j f_{ij}\,u_j$, so $g''(t) = \sum_i\sum_j u_i\,f_{ij}\,u_j$.
  3. That double sum is exactly $\mathbf{u}^\top H\mathbf{u}$:
$$\frac{d^2}{dt^2} f(\mathbf{x}+t\mathbf{u})\Big|_{t=0} = \mathbf{u}^\top H(\mathbf{x})\,\mathbf{u}.$$

Extremes. Write $H = \sum_k \lambda_k \mathbf{v}_k\mathbf{v}_k^\top$ with orthonormal eigenvectors $\mathbf{v}_k$ (Quadratic Forms & Definiteness). Any unit $\mathbf{u} = \sum_k c_k\mathbf{v}_k$ has $\sum c_k^2 = 1$, so

$$\mathbf{u}^\top H\mathbf{u} = \sum_k \lambda_k c_k^2 = \text{a weighted average of the eigenvalues}.$$

A weighted average lies between the smallest and the largest value: $\lambda_{\min} \le \mathbf{u}^\top H\mathbf{u} \le \lambda_{\max}$, with the ends reached exactly along the eigenvectors. Negative curvature in some direction exists exactly when $\lambda_{\min} \lt 0$.

Why do we need it?

In many dimensions, "is the ground curved up or down?" has no single answer: it depends on the direction. We need a tool that answers it for every direction at once, and finds the worst one.

Where is it used?

Finding the stiffest and flattest directions of a loss surface (these control the safe learning rate and the slow directions), detecting negative curvature to escape saddles, and the Laplace approximation, whose width along each eigen-direction is $1/\sqrt{\lambda}$.

How is it used?

Compute eigenvalues with np.linalg.eigvalsh(H) (it is symmetric). The largest tells the sharpest bend, the smallest the gentlest. For a specific direction use u @ H @ u.

The surface is $f(x,y) = \tfrac12\,\mathbf{x}^\top H\mathbf{x}$, so its Hessian is exactly the matrix you set with the three sliders (heights are shrunk by one common factor if they would not fit). Press a preset: bowl, hill, saddle, trough. Drag the purple probe over the surface. The teal and pink curves are the slices along the two eigen-directions. Turn the direction angle and watch the dark slice: its curvature $\mathbf{u}^\top H\mathbf{u}$ always lies between the two eigenvalues. Press Top to see the directions on the floor and Front to see a parabola.

The Hessian depends on the point. For a quadratic like the surface above it is the same everywhere. For a general function $H(\mathbf{x})$ changes from place to place, so "bowl" or "saddle" is a local statement.

Use unit vectors. $\mathbf{u}^\top H\mathbf{u}$ scales with the square of the length of $\mathbf{u}$. Always normalise $\mathbf{u}$ first if you want to compare directions fairly.

Quick check: $H = \begin{bmatrix}5&0\\0&1\end{bmatrix}$. Which direction curves most, and what is the curvature along $[0.6, 0.8]$?

The eigenvalues are $5$ (along $x$) and $1$ (along $y$), so the $x$-direction curves most. For $\mathbf{u} = [0.6, 0.8]$: $\mathbf{u}^\top H\mathbf{u} = 5(0.36) + 1(0.64) = 1.8 + 0.64 = 2.44$, between $1$ and $5$ as promised.

Positive definite Hessian: the bowl and the local minimum core

Put a marble on a flat spot. If the ground curves up in every direction around it (like the inside of a bowl), the marble stays: pushed any way, it rolls back. That flat spot is the bottom of a bowl, a local minimum.

"Curves up in every direction" means: the curvature $\mathbf{u}^\top H\mathbf{u}$ is positive for every direction $\mathbf{u}$. A Hessian with that property is called positive definite. It is the matrix version of "$f'' \gt 0$".

Let $f(x,y) = x^2 + xy + 2y^2$. Its gradient is $\nabla f = [2x + y,\; x + 4y]$, which is $[0,0]$ only at the origin. Its Hessian is $H = \begin{bmatrix}2&1\\1&4\end{bmatrix}$ everywhere.

  1. Trace $= 2 + 4 = 6$ and determinant $= 2\cdot4 - 1\cdot1 = 7$.
  2. The eigenvalues solve $\lambda^2 - 6\lambda + 7 = 0$, so $\lambda = 3 \pm \sqrt2 \approx 4.41$ and $1.59$. Both are positive: positive definite.
  3. Check by completing the square: $f = \left(x + \tfrac y2\right)^2 + \tfrac74 y^2$. A sum of squares, so $f \ge 0$, and $f = 0$ only at the origin. The origin is the lowest point ✓.

A symmetric matrix $H$ is positive definite (written $H \succ 0$) when $\mathbf{u}^\top H\mathbf{u} \gt 0$ for every $\mathbf{u}\neq\mathbf{0}$. Equivalently: all its eigenvalues are positive (Quadratic Forms & Definiteness).

Local minimum rule. If $\nabla f(\mathbf{c}) = \mathbf{0}$ and $H(\mathbf{c})$ is positive definite, then $\mathbf{c}$ is a (strict) local minimum. Reason, using the second-order Taylor expansion you will meet in Chapter 2.11: near $\mathbf{c}$,

$$f(\mathbf{c}+\boldsymbol\delta) \approx f(\mathbf{c}) + \underbrace{\nabla f(\mathbf{c})^\top\boldsymbol\delta}_{=\,0} + \tfrac12\,\boldsymbol\delta^\top H\,\boldsymbol\delta \;\gt\; f(\mathbf{c}).$$

If $H(\mathbf{x})$ is positive semidefinite (eigenvalues $\ge 0$) at every point, $f$ is convex: a single bowl whose local minimum is the global minimum.

Why do we need it?

It is the guarantee that a flat spot is a true bottom, not a hilltop or a pass. It also tells us optimisation is "easy": a bowl has no traps.

Where is it used?

Least-squares and ridge regression (Hessian $2X^\top X/n$ is positive semidefinite, positive definite with independent features), logistic regression (convex), convex optimisation, and the requirement behind Newton's method that the step goes downhill.

How is it used?

At a point with zero gradient, test the eigenvalues of $H$ (np.linalg.eigvalsh), or try a Cholesky factorisation: it works exactly when the matrix is positive definite. All positive means a local minimum.

Rotate the scene (drag the background), then press Top and Front. Both curvatures are positive, so every slice through the bottom is a smile. Drag the purple ball anywhere and press Roll downhill: it always ends at the bottom, where the gradient is zero. Press Climb uphill to see ascent leave the bowl. Make $\lambda_1$ large and $\lambda_2$ small: the bowl becomes a long narrow valley and the ball drifts along it slowly.

"Local" means local. A positive definite Hessian at one flat spot says nothing about other bowls elsewhere. A function can have many local minima (the wells in the classifier below).

A zero eigenvalue breaks the guarantee. If an eigenvalue is $0$ the bowl is a flat-bottomed trough in that direction and the test cannot decide. (We return to this in the second-derivative test.)

Quick check: is $H=\begin{bmatrix}2&3\\3&1\end{bmatrix}$ positive definite?

No. The determinant is $2\cdot1 - 3\cdot3 = -7 \lt 0$, so the eigenvalues have opposite signs (their product is the determinant). Along some direction the ground curves down. Positive diagonal entries are not enough.

Negative definite Hessian: the hill and the local maximum

Turn the bowl upside down and you get a hill. A marble balanced on the very top slides off if you nudge it in any direction. The top is a local maximum, and there the ground curves down in every direction: $\mathbf{u}^\top H\mathbf{u} \lt 0$ for all $\mathbf{u}$. We call that Hessian negative definite.

A hill and a bowl are the same shape flipped. So maximising $f$ is the same job as minimising $-f$.

Let $f(x,y) = -x^2 - xy - y^2$. Then $\nabla f = [-2x - y,\; -x - 2y]$, zero only at the origin, and $H = \begin{bmatrix}-2&-1\\-1&-2\end{bmatrix}$.

  1. Trace $= -4$, determinant $= 4 - 1 = 3$.
  2. $\lambda^2 + 4\lambda + 3 = 0$, so $\lambda = -1$ and $\lambda = -3$. Both negative: negative definite.
  3. Flip the sign: $-f = x^2 + xy + y^2$ has $H = \begin{bmatrix}2&1\\1&2\end{bmatrix}$ with eigenvalues $3$ and $1$, a bowl. Flipping $f$ flipped every eigenvalue ✓.

A symmetric matrix $H$ is negative definite ($H \prec 0$) when $\mathbf{u}^\top H\mathbf{u} \lt 0$ for every $\mathbf{u}\ne\mathbf{0}$, equivalently all eigenvalues are negative.

Local maximum rule. If $\nabla f(\mathbf{c}) = \mathbf{0}$ and $H(\mathbf{c}) \prec 0$, then $\mathbf{c}$ is a strict local maximum.

Link to minimising: $H_{-f} = -H_f$, so the eigenvalues of $-f$ are the negatives of those of $f$. Negative definite for $f$ means positive definite for $-f$.

Why do we need it?

Some problems are about going up: make a probability as large as possible. We need to recognise a true peak just as we recognise a true valley.

Where is it used?

Maximum-likelihood estimation (the log-likelihood should be a hill at the answer), the Laplace approximation (a Gaussian bump matched to the top of the log-posterior), expectation-maximisation, and reinforcement-learning objectives that are maximised.

How is it used?

Check the eigenvalues at the zero-gradient point: all negative means a local maximum. In practice we usually flip the sign and minimise the negative log-likelihood, which then has a positive definite Hessian.

Rotate the hill and look from the Front: every slice is a frown. Drag the ball to the top (the ring marks the peak), then drop it a little to one side and press Roll downhill: it slides off. Press Climb uphill (gradient ascent) and it goes to the top: ascent on a hill is the mirror image of descent in a bowl. Notice that all the eigenvalues in the readout are negative.

Gradient descent never finds a maximum on purpose. It walks downhill, so a hilltop is an unstable resting place. Only an exact start at the top (or exact arithmetic with zero gradient) leaves it there.

Quick check: $f$ has $H = \begin{bmatrix}-4&0\\0&-1\end{bmatrix}$ at a critical point. What are the eigenvalues of the Hessian of $-f$?

Flip the signs: $4$ and $1$. Both positive, so $-f$ has a bowl there and $f$ has a hill.

Saddle points: flat, but neither a valley nor a hill core

Picture a horse saddle, or a mountain pass between two peaks. Walk along the ridge line of the pass (towards one peak) and you go up on both sides: here the ground smiles. Cross the pass sideways (down into the valleys either side) and it frowns: the pass itself is the highest point of that path. At the centre the ground is level (the gradient is zero), yet it is not a valley bottom and not a hilltop. That spot is a saddle point.

A marble placed exactly there balances for a moment. Nudge it along the smile direction and it rolls back; nudge it along the frown direction and it rolls away.

$f(x,y) = x^2 - y^2$.

  1. $\nabla f = [2x, -2y] = [0,0]$ only at the origin.
  2. $H = \begin{bmatrix}2&0\\0&-2\end{bmatrix}$, eigenvalues $2$ and $-2$: mixed signs.
  3. Along the $x$-axis: $f(x,0) = x^2$, a smile (the origin looks like a minimum). Along the $y$-axis: $f(0,y) = -y^2$, a frown (the origin looks like a maximum).

A tilted example: $H = \begin{bmatrix}2&4\\4&2\end{bmatrix}$ has trace $4$ and determinant $4-16 = -12 \lt 0$, so eigenvalues $6$ and $-2$: also a saddle.

A critical point ($\nabla f = \mathbf{0}$) where the Hessian is indefinite (it has at least one positive and at least one negative eigenvalue) is a saddle point. Along an eigenvector with $\lambda \gt 0$ the function smiles; along one with $\lambda \lt 0$ it frowns.

Why saddles dominate in high dimensions. A critical point of a function of $n$ variables has $n$ curvatures (eigenvalues). It is a minimum only if all $n$ are positive. If each sign were an independent fair coin flip, the chance would be $2^{-n}$: for $n = 20$ that is about one in a million. Real Hessians are not coin flips, but the same effect holds: the more directions there are, the likelier it is that at least one bends down. Neural-network losses have millions of directions, so we expect saddles to far outnumber true minima. Random-matrix arguments and experiments on small networks suggest that critical points at high loss are mostly saddles, while local minima tend to sit at low loss. This is a useful picture, not a theorem about every network.

What gradient descent does near a saddle. The gradient is tiny near the saddle, so progress nearly stops (a plateau), then the small component along the frown direction grows by a factor $(1 + \eta|\lambda|)$ per step and the iterate slides off. Noise (as in SGD) and momentum help it escape.

Why do we need it?

If we only checked "the gradient is zero", we might stop at a saddle and call it a solution. Knowing about saddles explains why training sometimes stalls on a plateau and then suddenly moves on.

Where is it used?

Understanding training curves of deep networks, saddle-escaping optimisers (noise injection, momentum, Adam), the geometry of loss landscapes, and game theory and GANs (a Nash equilibrium is a saddle of the joint objective).

How is it used?

At a flat spot compute the Hessian's eigenvalues. Mixed signs mean saddle, and the eigenvector with a negative eigenvalue is a direction that goes downhill: a clean way to escape.

Rotate to look along the teal and pink slices: one smiles, one frowns. The ball starts almost on the "smile" axis. Press Roll downhill: it first slides toward the saddle point (the gradient is tiny there), then peels away along the frown direction. Try starting exactly on the smile axis, or turn the axes with $\varphi$. Press Climb uphill to see ascent do the mirror image.

Each random symmetric matrix below stands for the Hessian at one critical point in $n$ dimensions. The bars count how many of its $n$ eigenvalues are negative. Increase $n$. At $n=1$ half are minima. By $n=8$ the "all positive" bar (a true minimum) has almost vanished: nearly everything is a saddle. Press New sample to re-draw.

This is a picture, not a proof about your network. Real Hessians are not random matrices, and trained networks usually reach points with a few slightly negative or near-zero eigenvalues. The lesson is the counting: in many dimensions "all curvatures positive" is a demanding condition.

Zero gradient does not mean "done". Check the Hessian's eigenvalues, or at least watch whether the loss is still falling slowly.

Quick check: a critical point has $H$ with eigenvalues $5, 2, 0.1, -0.01$. What is it?

A saddle point: one eigenvalue is negative (tiny, but negative), so there is a (very gently) downhill direction. Along that direction progress will be slow, which looks like a plateau in training.

The second-derivative test (1D and many dimensions) core

Finding a flat spot (gradient zero) is step one. Step two is to ask: what kind of flat spot? Look at the curvature there. Smile: valley bottom. Frown: hilltop. Mixed: saddle. Zero: the test cannot tell, so look harder.

In one dimension there is one curvature, $f''$. In many dimensions there are $n$ curvatures, the eigenvalues of $H$, and we look at all of them.

One variable. $f(x) = \tfrac14x^4 - x^2$.

  1. $f'(x) = x^3 - 2x = x(x^2 - 2)$, which is $0$ at $x = 0$ and $x = \pm\sqrt2$.
  2. $f''(x) = 3x^2 - 2$.
  3. $f''(0) = -2 \lt 0$: a local maximum. $f''(\pm\sqrt2) = 6 - 2 = 4 \gt 0$: two local minima.

Many variables. $f(x,y) = (x^2-1)^2 + y^2$ has $\nabla f = [4x(x^2-1),\, 2y]$ and $H = \begin{bmatrix}12x^2-4&0\\0&2\end{bmatrix}$. Critical points: $(\pm1, 0)$ with $H = \mathrm{diag}(8, 2)$ (both positive: two minima) and $(0,0)$ with $H = \mathrm{diag}(-4, 2)$ (mixed: a saddle between the wells).

1D test. Suppose $f'(c) = 0$.

  • $f''(c) \gt 0$ $\Rightarrow$ local minimum.   $f''(c) \lt 0$ $\Rightarrow$ local maximum.
  • $f''(c) = 0$ $\Rightarrow$ inconclusive: $x^4$ (a minimum), $-x^4$ (a maximum) and $x^3$ (neither) all have $f'(0) = f''(0) = 0$. Look at higher derivatives: the first non-zero one decides (even order: min or max by its sign; odd order: not an extremum).

$n$-dimensional test. Suppose $\nabla f(\mathbf{c}) = \mathbf{0}$ and look at the eigenvalues of the symmetric $H(\mathbf{c})$:

  • All $\gt 0$ (positive definite) $\Rightarrow$ local minimum.
  • All $\lt 0$ (negative definite) $\Rightarrow$ local maximum.
  • Some $\gt 0$ and some $\lt 0$ (indefinite) $\Rightarrow$ saddle point.
  • Some $= 0$ and none of both signs $\Rightarrow$ inconclusive.

Shortcut for two variables. With $D = f_{xx}f_{yy} - f_{xy}^2 = \det H$ (the product of the eigenvalues): $D \gt 0$ and $f_{xx} \gt 0$ means minimum; $D \gt 0$ and $f_{xx} \lt 0$ means maximum; $D \lt 0$ means saddle; $D = 0$ is inconclusive.

Why do we need it?

"The gradient is zero" cannot tell a valley from a hilltop or a pass. The Hessian can. It is the standard way to certify that a point is really a minimum.

Where is it used?

Optimality conditions in optimisation theory, checking an analytic solution (for example that the least-squares solution is a minimum), analysing loss landscapes, and deciding whether an optimiser has stalled at a saddle.

How is it used?

1) Solve $\nabla f = 0$ for candidate points. 2) Compute $H$ there. 3) Get the eigenvalues (eigvalsh) and read off the signs. Treat tiny eigenvalues with care: numerically they may be zero.

Drag the dot. When $|f'|$ is almost zero the readout applies the test. Press Next flat spot to jump from one critical point to the next on the double-well curve: two minima and a maximum in between. Then choose $x^3$, $x^4$ and $-x^4$: all three have $f'(0) = f''(0) = 0$, so the test is silent, yet their shapes are different.

Colours show the height of $f$; the thin lines are level curves. Drag the blue point. The two short green/red lines are the Hessian's eigen-directions at the point (green = curves up, red = curves down). Press Next critical point to land on a flat spot and read the verdict: two wells have two minima and a saddle between them; the egg box has peaks, pits and passes; the monkey saddle is inconclusive. Away from flat spots, the test says "not a critical point".

"Inconclusive" does not mean "saddle". It means the second derivatives cannot decide. Example: $f = x^2y + y^2$ has $H(0,0) = \mathrm{diag}(0, 2)$, but along the curve $y = -x^2/2$ it equals $-x^4/4 \lt 0$ while along the $y$-axis it is positive, so the origin is actually a (degenerate) saddle.

It finds local, not global, minima. The two wells are both minima. Here they are equally deep, so both are global minima; in general only the lowest well is the global one.

Numerical zero. On a computer an eigenvalue like $10^{-12}$ is really zero. Use a tolerance.

Quick check: at a critical point $f_{xx}=3$, $f_{yy}=2$, $f_{xy}=2$. What is it?

$D = 3\cdot2 - 2^2 = 2 \gt 0$ and $f_{xx} = 3 \gt 0$: a local minimum. (Eigenvalues $(5\pm\sqrt{17})/2 \approx 4.56$ and $0.44$, both positive ✓.)

Newton's method: use the curvature to take a smarter step core

Gradient descent looks only at the slope: "downhill is that way, so take a small step". It has no idea how far the bottom is. Newton's method also looks at the curvature.

Here is the trick. Near where you stand, replace the curve by the parabola that matches its height, slope and curvature. A parabola has an obvious bottom, and you can jump straight to it. Then stand there, build a new parabola, and jump again.

If the ground really is a bowl-shaped parabola, you land on the bottom in one jump. Where the slope is gentle and the curvature is small, the parabola's bottom is far away, so Newton takes a big step; where the curvature is large, it takes a cautious one. The step size adjusts itself.

Minimise $f(x) = e^x - 2x$. Its minimum is where $f' = e^x - 2 = 0$, i.e. $x = \ln 2 \approx 0.6931$. Newton's update is $x \leftarrow x - f'(x)/f''(x)$ with $f'' = e^x$. Start at $x_0 = 0$.

  1. $x_0 = 0$: $f' = 1 - 2 = -1$, $f'' = 1$. Step $= -(-1)/1 = +1$, so $x_1 = 1$.
  2. $x_1 = 1$: $f' = e - 2 = 0.7183$, $f'' = e = 2.7183$. Step $= -0.2642$, so $x_2 = 0.7358$.
  3. $x_2 = 0.7358$: $f' = 0.0877$, $f'' = 2.0877$. Step $= -0.0420$, so $x_3 = 0.6940$.
  4. $x_3 = 0.6940$: next $x_4 = 0.69315$. The true answer is $0.693147\ldots$ Four steps give five correct digits.

Plain gradient descent with a safe step $\eta = 0.1$ needs about 32 steps to get within $0.001$ of the answer.

Derivation (1D). Near $x$, approximate $f$ by its second-order expansion (Chapter 2.11), a parabola in the step $\delta$:

$$m(\delta) = f(x) + f'(x)\,\delta + \tfrac12 f''(x)\,\delta^2.$$

Its bottom is where its slope is zero: $m'(\delta) = f'(x) + f''(x)\,\delta = 0$, so $\delta = -\dfrac{f'(x)}{f''(x)}$. Therefore

$$x_{\text{new}} = x - \frac{f'(x)}{f''(x)}.$$

In many dimensions the model is $m(\boldsymbol\delta) = f + \nabla f^\top\boldsymbol\delta + \tfrac12\boldsymbol\delta^\top H\boldsymbol\delta$. Setting its gradient (which is $\nabla f + H\boldsymbol\delta$ for symmetric $H$) to zero gives $H\boldsymbol\delta = -\nabla f$, so

$$\mathbf{x}_{\text{new}} = \mathbf{x} - H(\mathbf{x})^{-1}\,\nabla f(\mathbf{x}).$$

(In code we solve $H\boldsymbol\delta = -\nabla f$ rather than invert $H$.) Gradient descent is the same recipe with the curvature replaced by a fixed guess: $H \to \frac{1}{\eta} I$ gives $\boldsymbol\delta = -\eta\nabla f$.

Another view: Newton's method is the tangent-line root-finder applied to $f'$: it finds where $f' = 0$.

Why do we need it?

Gradient descent needs many small steps, and a bad step size makes it crawl or blow up. Using curvature chooses the step length (and direction) for us, and near a minimum converges extremely fast.

Where is it used?

Logistic regression solvers (iteratively reweighted least squares is Newton's method), L-BFGS and other quasi-Newton optimisers that approximate the Hessian, natural-gradient and K-FAC ideas in deep learning, and interior-point solvers.

How is it used?

At each iterate compute the gradient and Hessian, solve $H\boldsymbol\delta = -\nabla f$, step to $\mathbf{x}+\boldsymbol\delta$ (often with a line search or damping), and repeat until the gradient is tiny. Cost: about $n^3$ operations and $n^2$ memory per step.

The blue parabola matches the curve's height, slope and curvature at your point. Press Newton step: you jump to the parabola's bottom (the green dot). Press Gradient step to take a plain step of size $\eta\,|f'|$ instead. On $e^x - 2x$ Newton homes in within 4 steps. Now try ln cosh x starting at $x \approx 1.2$ or further: where the curve is almost flat the parabola's bottom is far away and Newton overshoots. On $x^3/3 - x$ start on the left side where it frowns: Newton jumps to the maximum.

The bowl is long and narrow (curvature 1 one way, 10 the other). Drag the blue start point, then press Step both several times. Orange is Newton: on a quadratic it lands exactly on the bottom in one step. Blue is gradient descent: it must keep $\eta$ small for the steep direction, so it crawls along the long one. At $\eta = 0.1$ the steep direction is cured in one step and the path then creeps along the flat direction; push $\eta$ above $0.1$ and it starts to zig-zag across the steep direction (above $0.2$ it explodes). Then pick saddle: Newton still jumps to the flat point, which is a saddle, not a minimum.

Newton heads for a flat spot of its parabola model, not necessarily a minimum. Where the curvature is negative it happily jumps to a maximum or a saddle (see the widgets). Practical versions fix the Hessian to be positive definite (add $\lambda I$), use a line search or trust region, or only take Newton steps near a minimum.

It is costly. Forming $H$ takes $n^2$ memory and solving takes $\sim n^3$ time. For a million weights that is impossible, which is why deep learning uses first-order methods (SGD, Adam) or cheap Hessian approximations (L-BFGS).

It can overshoot far from the answer. If the curvature is much smaller than it is near the bottom, the parabola's bottom is far away (see $\ln\cosh x$).

Quick check: $f(x) = 3x^2 - 12x + 5$. Take one Newton step from $x = 10$.

$f'(x) = 6x - 12 = 48$ and $f''(x) = 6$. Step $= -48/6 = -8$, so $x_{\text{new}} = 2$. That is the exact minimum ($f'(2) = 0$): on a quadratic Newton needs one step.

Why curvature controls the best learning rate core

Roll a marble down a steep, narrow gully by moving it in fixed hops. If each hop is too big, you jump clean across the gully to the other wall, which is even higher, then back, higher still: it blows up. A steep wall (high curvature) means you must hop small. In a flat, wide valley (low curvature) you can hop large.

So the biggest safe learning rate $\eta$ is set by the sharpest curvature. And the speed is set by the flattest direction: you must hop small to be safe in the steep direction, so you crawl along the flat one. That tension is what makes ill-conditioned problems hard.

One variable, $f(x) = \tfrac12\lambda x^2$ with curvature $\lambda = 5$. Gradient descent: $x \leftarrow x - \eta\,f'(x) = x - \eta\cdot5x = (1 - 5\eta)\,x$.

  • $\eta = 0.1$: factor $1 - 0.5 = 0.5$. Start at $x=2$: $2 \to 1 \to 0.5 \to 0.25$ (smooth shrinking).
  • $\eta = 0.2$: factor $1 - 1 = 0$. One step lands exactly on $0$. (And $\eta = 1/\lambda = 0.2$ is Newton's step.)
  • $\eta = 0.3$: factor $-0.5$. $2 \to -1 \to 0.5 \to -0.25$ (zig-zag but shrinking).
  • $\eta = 0.5$: factor $-1.5$. $2 \to -3 \to 4.5 \to -6.75$ (blows up).

The boundary is at factor $-1$, i.e. $\eta = 2/\lambda = 0.4$.

1D quadratic. For $f(x) = \tfrac12\lambda x^2$ ($\lambda \gt 0$), the update gives $x_{k+1} = (1 - \eta\lambda)\,x_k$, so $x_k = (1-\eta\lambda)^k x_0$. This shrinks to $0$ exactly when $|1 - \eta\lambda| \lt 1$, that is

$$0 \lt \eta \lt \frac{2}{\lambda}.$$

The best step is $\eta = 1/\lambda$ (one-step convergence). For $\eta\lambda \gt 1$ the iterates alternate sides; for $\eta\lambda \gt 2$ they grow.

Many dimensions. For $f(\mathbf{x}) = \tfrac12\mathbf{x}^\top H\mathbf{x}$, write $H = V\Lambda V^\top$ and use the eigen-directions as new axes (eigendecomposition). The problem splits into $n$ independent 1D problems with curvatures $\lambda_1,\dots,\lambda_n$. Every one must be stable, so

$$\boxed{\;0 \lt \eta \lt \frac{2}{\lambda_{\max}}\;}$$

The slowest direction shrinks by $|1 - \eta\lambda_{\min}|$ each step. Choosing $\eta$ so that the fastest and slowest directions have equal-size factors, $1 - \eta\lambda_{\min} = -(1 - \eta\lambda_{\max})$, gives $\eta^* = \dfrac{2}{\lambda_{\max} + \lambda_{\min}}$ and a best possible factor

$$\frac{\lambda_{\max} - \lambda_{\min}}{\lambda_{\max} + \lambda_{\min}} = \frac{\kappa - 1}{\kappa + 1}, \qquad \kappa = \frac{\lambda_{\max}}{\lambda_{\min}} \ (\text{the condition number}).$$

With $\lambda = (1, 10)$: $\eta \lt 0.2$, $\eta^* = 2/11 \approx 0.182$, factor $9/11 \approx 0.82$ per step. A bigger $\kappa$ pushes the factor towards 1: slow.

For a general smooth function the Hessian changes from place to place, and the same rule applies locally: the safe step shrinks where the loss is sharp. (Standard theory uses $\eta \le 1/L$, where $L$ bounds $\lambda_{\max}$.) Experiments with full-batch gradient descent on neural networks (reported as the "edge of stability") show the largest Hessian eigenvalue often rising during training until it hovers near $2/\eta$, while the loss keeps falling in a bumpy way. Treat this as an observed phenomenon, not a rule that always holds.

Why do we need it?

"Which learning rate should I use?" is the most common question in training. Curvature gives the principled answer: too big blows up, too small crawls, and the limit is $2/\lambda_{\max}$.

Where is it used?

Choosing and scheduling the learning rate, learning-rate warm-up (curvature can be large at the start), feature scaling (it makes the bowl rounder, reducing $\kappa$) and normalisation layers (often credited with similar effects), preconditioners, Adam and RMSprop (which rescale each direction), and convergence proofs.

How is it used?

If you can estimate $\lambda_{\max}$ (for example by a few power-iteration steps with Hessian-vector products), set $\eta$ below $2/\lambda_{\max}$. Otherwise try a few learning rates on a log scale and watch whether the loss goes down smoothly, zig-zags, or blows up.

The curve is $f = \frac12\lambda x^2$ and gradient descent starts at $x = 2$. Set the curvature $\lambda$, then slide $\eta$. Watch the multiplier $1-\eta\lambda$: positive and small, the steps shrink smoothly; negative, they bounce side to side; beyond $-1$ (that is $\eta \gt 2/\lambda$) they grow. Press Best step 1/λ to land on the bottom in one move. Double $\lambda$ and see how the safe $\eta$ halves. (If $\lambda$ is below $0.8$, the limit $2/\lambda$ is beyond the end of the $\eta$ slider.)

The bowl is $f = \frac12(\lambda_x x^2 + \lambda_y y^2)$: curvature $\lambda_x$ along $x$, $\lambda_y$ along $y$. Gradient descent starts at the blue ring. Set $\lambda_x = 1$, $\lambda_y = 8$. Slide $\eta$ up toward $2/\lambda_{\max} = 0.25$: the path zig-zags more and more across the steep $y$ direction; just beyond, it explodes. Press Best η: both directions shrink at the same rate, the fastest this method can go. Then make the two curvatures equal and see a perfectly round bowl.

The limit comes from the largest curvature, the speed from the smallest. Making the sharpest direction stable forces small steps, which slows the flat direction. That is why rescaling features (making the bowl round, $\kappa \approx 1$) speeds up training dramatically.

Real losses are not quadratic. Curvature changes along the path, so a rate that is safe in a flat region may blow up when you enter a sharp one. Warm-up and gradient clipping help.

Quick check: $f = \frac12(2x^2 + 5y^2)$. What is the largest safe $\eta$, and the best $\eta$?

$\lambda_{\max} = 5$, so $\eta \lt 2/5 = 0.4$. The best step is $2/(5+2) = 2/7 \approx 0.286$, with factor $(5-2)/(5+2) = 3/7 \approx 0.43$ per step.

Recap, cheat sheet and practice

  • The second derivative $f''$ is the slope of the slope. Positive: smile (concave up); negative: frown. Higher derivatives $f^{(n)}$ keep going; they are the raw material of Taylor series.
  • The Hessian $H_{ij} = \partial^2 f/\partial x_i\partial x_j$ is the $n\times n$ matrix of second partials. It is the Jacobian of the gradient, and symmetric for smooth functions (equal mixed partials).
  • The curvature along a unit direction $\mathbf{u}$ is $\mathbf{u}^\top H\mathbf{u}$, always between $\lambda_{\min}$ and $\lambda_{\max}$, reached along the eigenvectors.
  • At a critical point: all eigenvalues positive is a bowl (minimum), all negative a hill (maximum), mixed signs a saddle, a zero eigenvalue is inconclusive. In high dimensions saddles are expected to far outnumber minima.
  • Newton's method uses curvature: $\mathbf{x}\leftarrow\mathbf{x} - H^{-1}\nabla f$ (exact in one step on a quadratic, but costly and drawn to saddles).
  • On a quadratic, gradient descent is stable only if $\eta \lt 2/\lambda_{\max}$; speed is set by the condition number $\kappa = \lambda_{\max}/\lambda_{\min}$.

Cheat sheet

IdeaFormulaPicture
Second derivative$f'' = \dfrac{d}{dx}f' \approx \dfrac{f(x+h)-2f(x)+f(x-h)}{h^2}$smile ($+$) or frown ($-$)
Hessian$H_{ij} = \dfrac{\partial^2 f}{\partial x_i\partial x_j}$, $H = H^\top$table of curvatures
Curvature along $\mathbf{u}$$\mathbf{u}^\top H\mathbf{u}$, $\lambda_{\min}\le\cdot\le\lambda_{\max}$slice of the surface
Curvature circle$\kappa = \dfrac{f''}{(1+f'^2)^{3/2}}$, $R=1/|\kappa|$circle hugging the curve
Bowl / hill / saddleall $\lambda\gt0$ / all $\lambda\lt0$ / mixed signsmin / max / pass
2D test$D=f_{xx}f_{yy}-f_{xy}^2$; $D\gt0,f_{xx}\gt0$ min; $D\lt0$ saddlesign of $\det H$
Newton step$\boldsymbol\delta = -H^{-1}\nabla f$jump to the parabola's bottom
Safe learning rate$\eta \lt 2/\lambda_{\max}$; best $2/(\lambda_{\max}+\lambda_{\min})$hop smaller than the gully
Convergence factor$(\kappa-1)/(\kappa+1)$round bowl is fast
Code it · NumPy

import numpy as np

# f(x, y) = x^2 * y + y^2   (the worked example of this chapter)
def f(p):
    x, y = p
    return x**2 * y + y**2

def hess(p):                       # exact Hessian (from the formulas we derived)
    x, y = p
    return np.array([[2*y, 2*x],
                     [2*x, 2.0]])

def hessian_fd(f, p, h=1e-4):      # Hessian from function values only (central differences)
    p = np.asarray(p, float); n = p.size; H = np.zeros((n, n))
    for i in range(n):
        for j in range(n):
            ei = np.zeros(n); ej = np.zeros(n); ei[i] = h; ej[j] = h
            H[i, j] = (f(p+ei+ej) - f(p+ei-ej) - f(p-ei+ej) + f(p-ei-ej)) / (4*h*h)
    return H

p = np.array([1.0, 2.0])
print(hess(p))                     # [[4. 2.] [2. 2.]]
print(np.round(hessian_fd(f, p), 3))   # same numbers, found by nudging
print(np.linalg.eigvalsh(hess(p)))     # [0.7639 5.2361] = 3 -/+ sqrt(5): both positive

def classify(H, tol=1e-8):         # second-derivative test for a symmetric Hessian
    lam = np.linalg.eigvalsh(H)
    if np.all(lam > tol):  return "local minimum"
    if np.all(lam < -tol): return "local maximum"
    if lam.min() < -tol and lam.max() > tol: return "saddle point"
    return "inconclusive"

print(classify(np.diag([8., 2.])))     # local minimum   (two-wells function at (1, 0))
print(classify(np.diag([-4., 2.])))    # saddle point    (two-wells function at (0, 0))

H = np.array([[1., 2.], [2., 1.]])
u = np.array([1., 1.]) / np.sqrt(2)
print(u @ H @ u)                       # 3.0  curvature along u (about 3, up to rounding)
print(np.linalg.eigvalsh(H))           # [-1.  3.]  smallest and largest possible curvature

# Newton's method on g(x) = e^x - 2x   (minimum at ln 2 = 0.6931...)
x = 0.0
for k in range(4):
    x = x - (np.exp(x) - 2) / np.exp(x)    # x - g'(x) / g''(x)
    print(k + 1, x)                        # 1.0, 0.7358, 0.6940, 0.69315

# Learning rate vs curvature on f = 0.5 * (1*x^2 + 10*y^2): stable only if eta < 2/10 = 0.2
lam = np.array([1.0, 10.0])
for eta in (0.15, 0.19, 0.21):
    z = np.array([1.0, 1.0])
    for _ in range(100):
        z = z - eta * lam * z                  # gradient descent step (gradient = lam * z)
    print(eta, np.abs(z).max())                # tiny, tiny, huge
Test yourself

1. At a point where $f'(c) = 0$ and $f''(c) = -3$, what is $c$?

A negative second derivative means the curve is a frown, so a flat spot is a hilltop: a local maximum.

2. Why is the Hessian of a smooth function symmetric?

That is the theorem on equality of mixed partials (Schwarz/Clairaut). Being square does not make a matrix symmetric, and the Hessian need not be positive definite.

3. At a critical point the Hessian has eigenvalues $4$ and $-1$. This is a…

Mixed signs mean the ground curves up in one direction and down in another: a saddle.

4. On $f(x) = \tfrac12\cdot5\,x^2$ gradient descent is run. Which learning rate makes it diverge?

The limit is $2/\lambda = 2/5 = 0.4$. Only $0.45$ exceeds it (the multiplier is $1 - 0.45\cdot5 = -1.25$, magnitude above 1).

5. Newton's method is applied to $f(x) = 3x^2 - 12x + 5$ starting at $x = 10$. After one step, where is it?

$f' = 48$, $f'' = 6$, so the step is $-48/6 = -8$ and $x = 2$. For a quadratic one Newton step is exact.

6. Which statement about $\mathbf{u}^\top H\mathbf{u}$ for a unit vector $\mathbf{u}$ is true?

Writing $\mathbf{u}$ in the eigenvector basis turns $\mathbf{u}^\top H\mathbf{u}$ into a weighted average of the eigenvalues, so it cannot leave their range.

Practice problems

A. Find and classify the critical points of $f(x,y) = x^3 + y^3 - 3xy$.

$f_x = 3x^2 - 3y = 0$ and $f_y = 3y^2 - 3x = 0$. From the first, $y = x^2$; substituting, $x^4 = x$, so $x = 0$ or $x = 1$. Critical points: $(0,0)$ and $(1,1)$. The Hessian is $H = \begin{bmatrix}6x & -3\\ -3 & 6y\end{bmatrix}$.

At $(0,0)$: $H = \begin{bmatrix}0&-3\\-3&0\end{bmatrix}$, eigenvalues $\pm3$: saddle. At $(1,1)$: $H = \begin{bmatrix}6&-3\\-3&6\end{bmatrix}$, eigenvalues $3$ and $9$: local minimum with $f(1,1) = -1$.

B. For $H = \begin{bmatrix}2&1\\1&4\end{bmatrix}$ find the curvature along $\mathbf{u} = [0.6, 0.8]$ and check it lies in the allowed range.

$\mathbf{u}^\top H\mathbf{u} = 2(0.36) + 2\cdot1\cdot(0.6)(0.8) + 4(0.64) = 0.72 + 0.96 + 2.56 = 4.24$. The eigenvalues are $3\pm\sqrt2 \approx 1.59$ and $4.41$, and $1.59 \le 4.24 \le 4.41$ ✓.

C. Take two Newton steps for $f(x) = x^4$ from $x = 2$. What is the pattern?

$f' = 4x^3$, $f'' = 12x^2$, so $x - f'/f'' = x - x/3 = \tfrac23 x$. From $2$: $x_1 = 4/3 \approx 1.333$, $x_2 = 8/9 \approx 0.889$. The error shrinks by a factor $2/3$ each step (slow, because $f''(0) = 0$: the parabola model is poor at a flat-bottomed minimum).

D. $f = \frac12(2x^2 + 5y^2)$. Give the safe range for $\eta$, the best $\eta$ and the best per-step factor.

$\lambda = 2, 5$, so $\kappa = 2.5$. Safe: $\eta \lt 2/5 = 0.4$. Best: $\eta^* = 2/(2+5) = 2/7 \approx 0.286$. Factor: $(\kappa-1)/(\kappa+1) = 1.5/3.5 = 3/7 \approx 0.43$.

E. Show the second-derivative test is inconclusive for $f(x,y) = x^4 + y^2$ at the origin, then decide what the origin is.

$\nabla f = [4x^3, 2y] = \mathbf{0}$ at the origin and $H = \begin{bmatrix}12x^2&0\\0&2\end{bmatrix} = \mathrm{diag}(0, 2)$ there: one eigenvalue is $0$, so the test is inconclusive. But $f = x^4 + y^2 \ge 0$ and equals $0$ only at the origin, so it is a strict (global) minimum.

F. If each of the $n=10$ curvature signs at a critical point were an independent fair coin flip, what is the chance of a minimum? Why does this matter?

All 10 must be positive: $2^{-10} = 1/1024 \approx 0.1\%$. So with many directions, almost every critical point has some downhill direction: most are saddles.

Chapter 2.11

Taylor Series

Almost every function in machine learning is complicated. But close to a single point, any smooth function looks like a simple polynomial: a straight line, then a parabola, then something even closer. Taylor series is the recipe for building that polynomial. It is also the reason gradient descent and Newton's method work at all.

  • Build a polynomial that matches a function's value, slope, curvature and more at a point, and see why the coefficients are $f^{(n)}(a)/n!$
  • Use the first-order (tangent) and second-order (adds curvature) approximations, in one variable and in many
  • Expand $e^x$, $\sin x$, $\cos x$, $\ln(1+x)$, $\frac{1}{1-x}$ and the sigmoid, and use them for quick estimates
  • Know how big the error is (about $\delta^2$ for first order, $\delta^3$ for second) and where the approximation stops being good
  • Derive gradient descent from the first-order expansion and Newton's method from the second-order one

The Taylor expansion: a polynomial that copies a function core

Polynomials ($1$, $x$, $x^2$, $x^3$, ...) are the easiest functions there are: a computer only needs to add and multiply. Functions like $e^x$ or $\sin x$ are much harder. The big idea: near one point, we can copy a hard function with a polynomial.

How do we copy it? One feature at a time, from the most basic up:

  • Match the height at the point (a flat line at the right level).
  • Also match the slope (the line now tilts like the curve).
  • Also match the curvature (the line bends into a parabola).
  • Also match how the curvature changes, and so on. Each extra match makes the copy hug the curve for longer.

Each "feature" is one derivative from Chapter 2.10. The finished polynomial is the Taylor polynomial.

Copy $f(x) = e^x$ near $a = 0$ using $p(x) = c_0 + c_1x + c_2x^2 + c_3x^3 + \cdots$. Every derivative of $e^x$ is $e^x$, and $e^0 = 1$, so at $x = 0$ the height, slope, curvature, ... are all $1$.

  1. Match the height: $p(0) = c_0$ must be $1$, so $c_0 = 1$.
  2. Match the slope: $p'(x) = c_1 + 2c_2x + 3c_3x^2 + \cdots$, so $p'(0) = c_1 = 1$.
  3. Match the curvature: $p''(x) = 2c_2 + 6c_3x + \cdots$, so $p''(0) = 2c_2 = 1$ and $c_2 = \tfrac12$.
  4. Next: $p'''(0) = 6c_3 = 1$, so $c_3 = \tfrac16$. Next: $p^{(4)}(0) = 24c_4 = 1$, so $c_4 = \tfrac1{24}$.
$$e^x \approx 1 + x + \frac{x^2}{2} + \frac{x^3}{6} + \frac{x^4}{24} + \cdots$$

Numeric check at $x = 1$: $1 + 1 + 0.5 + 0.1667 + 0.0417 + 0.0083 = 2.7167$ after six terms, and $e = 2.7183$. Very close already.

The Taylor polynomial of degree $n$ of $f$ around the point $a$ is

$$P_n(x) = \sum_{k=0}^{n}\frac{f^{(k)}(a)}{k!}(x-a)^k = f(a) + f'(a)(x-a) + \frac{f''(a)}{2!}(x-a)^2 + \cdots + \frac{f^{(n)}(a)}{n!}(x-a)^n.$$

Where does $1/k!$ come from? Suppose $p(x) = c_0 + c_1(x-a) + c_2(x-a)^2 + \cdots$ and we want $p^{(k)}(a) = f^{(k)}(a)$ for every $k$. Differentiate the term $c_k(x-a)^k$ exactly $k$ times: the power rule gives $k\cdot(k-1)\cdots2\cdot1 = k!$, a constant. Lower powers have already vanished (their $k$-th derivative is $0$), and higher powers still contain a factor $(x-a)$ that is $0$ at $x=a$. So $p^{(k)}(a) = k!\,c_k$. Setting this equal to $f^{(k)}(a)$ gives

$$c_k = \frac{f^{(k)}(a)}{k!}.$$

Taking $n\to\infty$ gives the Taylor series. Around $a=0$ it is also called the Maclaurin series. Degree $0$ is a flat line at height $f(a)$; degree $1$ is the tangent line; degree $2$ is a parabola.

Why do we need it?

Hard functions are hard to compute and hard to reason about. A polynomial copy is easy to evaluate, differentiate and minimise, and it tells us how the function behaves nearby.

Where is it used?

How calculators and math libraries compute $\sin$, $\exp$ and $\log$; justification of gradient descent and Newton's method; Laplace approximations; second-order loss approximations in gradient boosting (XGBoost); and error and sensitivity estimates.

How is it used?

Pick the point $a$ you care about, compute $f$ and its first few derivatives there, divide the $k$-th derivative by $k!$, and add up the terms. Keep as many terms as the accuracy you need.

Choose a function (it starts as $e^x$). Add terms with the slider or the Next term button. The table shows each coefficient $c_k = f^{(k)}(0)/k!$. Watch the blue polynomial climb to cover more and more of the grey curve, always hugging it near $x = 0$ first. For $\sin x$ you will see the even-numbered terms are all zero. Try the sigmoid: its Taylor series has no $x^2$ term.

The order of the polynomial is not the number of non-zero terms. $\sin x$ has only odd powers, so its degree-4 polynomial is the same as its degree-3 one.

A Taylor polynomial is built at one point. Change $a$ and every coefficient changes. It copies the function near $a$, not everywhere (see "Local approximation" below).

Quick check: what is the coefficient of $x^3$ in the Taylor series of $\sin x$ at $0$?

$f'''(x) = -\cos x$, so $f'''(0) = -1$. Divide by $3! = 6$: the coefficient is $-\tfrac16$. So $\sin x \approx x - \tfrac{x^3}{6} + \cdots$.

First-order approximation: the tangent line core

Zoom in on any smooth curve. Zoom in more. Eventually the curve looks like a straight line, and that line is the tangent. So close to a point, the function behaves like its tangent line: "start at the current height, and add slope times how far you moved".

This is the simplest useful approximation, and it is what a gradient is for: the gradient is the slope, and the slope predicts what a small nudge does.

Estimate $e^{0.1}$ without a calculator, using $f(x) = e^x$ near $a = 0$. Here $f(0) = 1$ and $f'(0) = 1$.

  1. First-order formula: $f(a + h) \approx f(a) + f'(a)\,h$.
  2. With $a = 0$ and $h = 0.1$: $e^{0.1} \approx 1 + 1\cdot0.1 = 1.1$.
  3. True value: $e^{0.1} = 1.10517\ldots$ The error is $0.0052$.

Now double the step, $h = 0.2$: estimate $1.2$, true $1.2214$, error $0.0214$. Twice the step gave about four times the error ($0.0214/0.0052 \approx 4.1$). The error grows like $h^2$.

The first-order Taylor approximation (also called the linear approximation or tangent-line approximation) is

$$f(a + h) \approx f(a) + f'(a)\,h, \qquad\text{or}\qquad f(x) \approx f(a) + f'(a)(x - a).$$

Its graph is the tangent line at $a$. The error is the part we dropped, which is of size about $\tfrac12 f''(a)\,h^2$: proportional to the square of the step. In words: halve the step, quarter the error.

For $f:\mathbb{R}^n\to\mathbb{R}$ this becomes $f(\mathbf{x}+\boldsymbol\delta) \approx f(\mathbf{x}) + \nabla f(\mathbf{x})^\top\boldsymbol\delta$ (see the multivariable section below). Chapter 2.12 continues this idea as linearisation.

Why do we need it?

It turns a curved, hard question ("what will the loss be after this change?") into a simple one ("height plus slope times step"). That is enough to decide which way to move.

Where is it used?

Gradient descent (the step is chosen using it), sensitivity and error propagation, finite-difference gradient checks, linearising a network around a point (neural tangent kernel ideas), and the Gauss-Newton method.

How is it used?

Compute $f(a)$ and the slope $f'(a)$ (or the gradient). Predict the new value as $f(a)$ plus slope times step. Trust it only for small steps.

Drag the blue point $a$ along the curve. The orange line is the tangent. Slide the step $h$: the vertical red bar is the error. Halve $h$ and check that the error roughly quarters (the readout shows error$/h^2$, which stays near $\tfrac12 f''(a)$). In the right-hand panel, slide look closer and watch the curve flatten into its tangent line.

It is only good for small steps. The error is $\approx\tfrac12 f''h^2$. If the curvature $f''$ is large, "small" has to be very small. A straight line cannot follow a sharp bend.

The sign of the error tells the curvature. If the curve is a smile ($f''\gt0$) the tangent line lies below the curve; for a frown it lies above.

Quick check: estimate $\sqrt{4.1}$ using the tangent line of $\sqrt x$ at $a = 4$.

$f(4) = 2$ and $f'(x) = \dfrac{1}{2\sqrt x}$, so $f'(4) = \dfrac14$. Then $\sqrt{4.1} \approx 2 + \dfrac14(0.1) = 2.025$. The true value is $2.02485\ldots$ (error about $0.00015$).

Second-order approximation: adding the curvature core

The tangent line is straight, but the curve is not. So the line peels away from the curve, and it peels off to one side: above a frown, below a smile. We can fix this by letting the copy bend. Add a term that matches the second derivative, and the straight line becomes a parabola hugging the curve much longer.

That one extra term, $\tfrac12 f''(a)\,h^2$, is the curvature term from Chapter 2.10.

Estimate $e^{0.1}$ again, now adding the curvature term. At $a=0$: $f = f' = f'' = 1$.

  1. Formula: $f(a+h) \approx f(a) + f'(a)h + \tfrac12 f''(a)h^2$.
  2. $e^{0.1} \approx 1 + 0.1 + \tfrac12(0.01) = 1.105$.
  3. True value $1.105171$. Error $0.000171$, about 30 times smaller than the tangent line's $0.0052$.

Double the step to $h = 0.2$: estimate $1.22$, true $1.221403$, error $0.001403$. Twice the step gives about eight times the error ($0.001403/0.000171 \approx 8.2$): the error now grows like $h^3$.

The second-order Taylor approximation is

$$f(a + h) \approx f(a) + f'(a)\,h + \tfrac12 f''(a)\,h^2.$$

Its graph is a parabola that matches $f$ in height, slope and curvature at $a$. The leftover error is of size about $\tfrac16 f'''(a)h^3$: proportional to the cube of the step. Halve the step, divide the error by about eight.

If $f$ is itself a parabola (a quadratic), this approximation is exact, because all higher derivatives are zero.

Why do we need it?

A straight line cannot tell whether a flat spot is a valley or a hilltop. A parabola can: it has a bottom (or a top). Adding curvature also lets us take bigger, smarter steps.

Where is it used?

Newton's method and its relatives (L-BFGS), the Laplace approximation (a Gaussian fitted at a loss minimum), second-order terms in XGBoost, and analysing how big a learning rate can be.

How is it used?

Compute value, slope and curvature at the point and use the parabola as a stand-in for the function. To minimise, jump to the parabola's bottom; to analyse stability, read the curvature.

Drag the blue point. The orange line is order 1, the green parabola is order 2. Slide the step $h$ out from $0$: the line peels off first, the parabola stays with the curve for much longer. Halve $h$ and compare how the two errors shrink (about 4 times and about 8 times). Try the sigmoid: at $a=0$ the curvature is $0$, so the parabola equals the line there! Move away from $0$ to see them differ.

More terms are not always better far away. The extra terms improve the copy near $a$. Far from $a$ a high-degree polynomial can shoot off wildly. The next sections show where it works.

Quick check: use the second-order formula to estimate $\ln(1.1)$ at $a=0$ for $f=\ln(1+x)$.

$f(0) = 0$, $f'(0) = 1$, $f''(0) = -1$. So $\ln(1+h) \approx h - \tfrac12h^2$. With $h = 0.1$: $0.1 - 0.005 = 0.095$. The true value is $0.09531$.

The famous expansions, derived core

A handful of functions show up everywhere, so it pays to know their Taylor series around $0$ and, even better, to know how to rebuild them. Each one comes from the same two steps: (1) find the derivatives at $0$, (2) divide the $k$-th one by $k!$.

The patterns are simple: $e^x$ has all-equal derivatives; $\sin$ and $\cos$ cycle through $0, 1, 0, -1$; $\frac1{1-x}$ and $\ln(1+x)$ are geometric-looking.

$e^x$: every derivative at $0$ is $1$, so $c_k = \dfrac1{k!}$.

$\sin x$: $f, f', f'', f''', f^{(4)}, \dots$ at $0$ are $0, 1, 0, -1, 0, 1, \dots$ (from $\sin, \cos, -\sin, -\cos, \dots$). Only odd $k$ survive, with alternating signs, so $c_1 = 1$, $c_3 = -\tfrac1{3!}$, $c_5 = \tfrac1{5!}$.

$\cos x$: the derivatives at $0$ are $1, 0, -1, 0, 1, \dots$. Only even $k$: $c_0 = 1$, $c_2 = -\tfrac1{2!}$, $c_4 = \tfrac1{4!}$.

$\dfrac1{1-x}$: $f^{(k)}(x) = \dfrac{k!}{(1-x)^{k+1}}$ (the chain rule brings out one more factor each time). At $0$ it is $k!$, so $c_k = k!/k! = 1$.

$\ln(1+x)$: $f' = \dfrac1{1+x}$, $f'' = -\dfrac1{(1+x)^2}$, $f''' = \dfrac2{(1+x)^3}$, so $f^{(k)}(0) = (-1)^{k-1}(k-1)!$ for $k\ge1$ and $f(0)=0$. Then $c_k = \dfrac{(-1)^{k-1}(k-1)!}{k!} = \dfrac{(-1)^{k-1}}{k}$.

Sigmoid $\sigma(x) = \dfrac1{1+e^{-x}}$: use $\sigma' = \sigma(1-\sigma)$.

  1. $\sigma(0) = \tfrac12$, so $\sigma'(0) = \tfrac12\cdot\tfrac12 = \tfrac14$.
  2. Product rule on $\sigma' = \sigma - \sigma^2$: $\sigma'' = \sigma'(1 - 2\sigma)$. At $0$: $\tfrac14\cdot(1 - 1) = 0$.
  3. Again: $\sigma''' = \sigma''(1-2\sigma) - 2(\sigma')^2$. At $0$: $0 - 2\cdot\tfrac1{16} = -\tfrac18$.
  4. Coefficients: $c_0 = \tfrac12$, $c_1 = \tfrac14$, $c_2 = 0$, $c_3 = -\tfrac18/6 = -\tfrac1{48}$.
$$\begin{aligned} e^x &= 1 + x + \frac{x^2}{2!} + \frac{x^3}{3!} + \frac{x^4}{4!} + \cdots &&\text{(all } x)\\ \sin x &= x - \frac{x^3}{3!} + \frac{x^5}{5!} - \frac{x^7}{7!} + \cdots &&\text{(all } x)\\ \cos x &= 1 - \frac{x^2}{2!} + \frac{x^4}{4!} - \frac{x^6}{6!} + \cdots &&\text{(all } x)\\ \ln(1+x) &= x - \frac{x^2}{2} + \frac{x^3}{3} - \frac{x^4}{4} + \cdots &&(-1 \lt x \le 1)\\ \frac{1}{1-x} &= 1 + x + x^2 + x^3 + \cdots &&(|x| \lt 1)\\ \sigma(x) &= \frac12 + \frac{x}{4} - \frac{x^3}{48} + \frac{x^5}{480} - \cdots &&(|x| \lt \pi) \end{aligned}$$

The "valid for" notes are the radius of convergence (see the section on local approximation). Two more used all the time: $\sqrt{1+x} \approx 1 + \tfrac x2 - \tfrac{x^2}{8}$ and $\dfrac{1}{1+x}\approx 1 - x + x^2$. And a link between them (awareness, needs complex numbers): $\cos x$ and $i\sin x$ are the even and odd parts of $e^{ix}$.

Why do we need it?

These functions are the building blocks of activations, losses and probabilities. Their series give quick mental estimates and show how each behaves near zero, where inputs to a network often live.

Where is it used?

Sigmoid $\approx \frac12 + \frac x4$ for small inputs (so a unit fed small inputs behaves almost linearly), $\ln(1+x)\approx x$ (log1p), $e^x - 1\approx x$ (expm1), small-angle $\sin\theta\approx\theta$, softplus $\approx \ln2 + \frac x2 + \frac{x^2}{8}$, and the geometric series in discounted returns in reinforcement learning.

How is it used?

Recognise the pattern, keep the first one to three terms, and check the error is acceptable for your range of $x$. When you need a new series, do not memorise: derive the derivatives at $0$ and divide by $k!$.

Pick a function and slide the order $n$ from 0 to 10. Top: the grey curve is $f$, the blue curve is the degree-$n$ Taylor polynomial, the green band is where the error is smaller than your tolerance. Bottom: the error on a log scale (each step down is ten times smaller). Notice: each order widens the green band; the band is limited for $\ln(1+x)$ and $\frac{1}{1-x}$ even at high order. Drag the black dot to read the error at one point.

Know the "valid for" range. $\ln(1+x)$'s series only works for $-1 \lt x \le 1$ and $\frac{1}{1-x}$'s only for $|x|\lt1$, no matter how many terms you take. At $x=1.5$, adding terms makes the answer worse.

Sigmoid: $\sigma(0)=\tfrac12$, so the series starts at $\tfrac12$, not $0$. It is an odd function plus $\tfrac12$, so only odd powers appear after the constant.

Quick check: use the series to estimate $\sigma(0.4)$ with three terms, and compare with $0.59869$.

$\sigma(0.4) \approx \tfrac12 + \tfrac{0.4}{4} - \tfrac{0.4^3}{48} = 0.5 + 0.1 - \tfrac{0.064}{48} = 0.6 - 0.001333 = 0.598667$. The true value is $0.598688$, so the error is only about $0.00002$.

Multivariate Taylor expansion core

For a function of many inputs the picture is a surface. Zoom in on one spot and the surface looks like a flat tilted sheet (the tangent plane): that is the first-order approximation, built from the gradient. Look a little less closely and you see the sheet bend: it becomes a bowl, hill or saddle (a paraboloid): that is the second-order approximation, built from the Hessian.

The recipe is the same as in one variable: start from the current height, add the slope part, add the curvature part. Only now "slope" is a gradient (a vector) and "curvature" is a Hessian (a matrix), and the step $\boldsymbol\delta$ is a vector.

$f(x,y) = x^2y + y^2$ at $\mathbf{x} = (1,2)$, with a step $\boldsymbol\delta = (0.1,\,-0.1)$. From Chapter 2.10: $f_x = 2xy$, $f_y = x^2 + 2y$, $H = \begin{bmatrix}2y&2x\\2x&2\end{bmatrix}$.

  1. Height: $f(1,2) = 1\cdot2 + 4 = 6$.
  2. Gradient (a column): $\nabla f(1,2) = [2\cdot1\cdot2,\; 1 + 4] = [4, 5]$.
  3. Slope part: $\nabla f^\top\boldsymbol\delta = 4(0.1) + 5(-0.1) = -0.1$. So the first-order estimate is $6 - 0.1 = 5.9$.
  4. Hessian: $H(1,2) = \begin{bmatrix}4&2\\2&2\end{bmatrix}$. Then $\boldsymbol\delta^\top H\boldsymbol\delta = 4(0.01) + 2\cdot2\cdot(0.1)(-0.1) + 2(0.01) = 0.04 - 0.04 + 0.02 = 0.02$.
  5. Curvature part: $\tfrac12(0.02) = 0.01$. The second-order estimate is $5.9 + 0.01 = 5.91$.
  6. True value: $f(1.1,\,1.9) = 1.21\cdot1.9 + 3.61 = 5.909$. Errors: first order $0.009$, second order $0.001$.

For $f:\mathbb{R}^n\to\mathbb{R}$ (smooth), a point $\mathbf{x}$ and a small step $\boldsymbol\delta$:

$$\boxed{\,f(\mathbf{x}+\boldsymbol\delta) \;\approx\; f(\mathbf{x}) \;+\; \nabla f(\mathbf{x})^\top\boldsymbol\delta \;+\; \tfrac12\,\boldsymbol\delta^\top H(\mathbf{x})\,\boldsymbol\delta\,}$$

Each term is a single number: $\nabla f^\top\boldsymbol\delta$ is a row $(1\times n)$ times a column $(n\times1)$; $\boldsymbol\delta^\top H\boldsymbol\delta$ is $(1\times n)(n\times n)(n\times 1)$. Keeping only the first two terms gives the first-order (tangent plane) approximation.

Derivation from the 1D formula. Walk along the straight line from $\mathbf{x}$ to $\mathbf{x}+\boldsymbol\delta$ and let $g(t) = f(\mathbf{x} + t\boldsymbol\delta)$, so $g(0) = f(\mathbf{x})$ and $g(1) = f(\mathbf{x}+\boldsymbol\delta)$. Chapter 2.10 showed that $g'(0) = \nabla f(\mathbf{x})^\top\boldsymbol\delta$ and $g''(0) = \boldsymbol\delta^\top H(\mathbf{x})\boldsymbol\delta$ (curvature along a direction, without normalising). The one-variable second-order formula at $t = 1$ is $g(1)\approx g(0) + g'(0)\cdot 1 + \tfrac12 g''(0)\cdot1^2$, which is exactly the boxed formula ∎.

If $f$ is a quadratic $\tfrac12\mathbf{x}^\top A\mathbf{x} - \mathbf{b}^\top\mathbf{x} + c$ ($A$ symmetric) the formula is exact, with $H = A$. Higher orders involve third derivatives (a 3-index array); in ML we almost never go past second order.

Why do we need it?

A loss depends on millions of parameters at once. To predict what a small change of all of them does to the loss, we need a formula that works with vectors: gradient for the slope, Hessian for the bend.

Where is it used?

The derivation of gradient descent and Newton's method (below), trust-region optimisers, the Laplace approximation, natural gradients and K-FAC, Hessian-based pruning (Optimal Brain Surgeon), and sharp-versus-flat-minima analyses.

How is it used?

At the current parameters compute the loss $f$, the gradient $g$ and (if affordable) the Hessian $H$. The model $f + g^\top\delta + \tfrac12\delta^\top H\delta$ predicts the loss after any step $\delta$, and can be minimised over $\delta$.

The coloured sheet is $f$. The green patch is the first-order approximation (tangent plane) at the blue point $P$; the purple patch is the second-order one (paraboloid). Drag $P$ (blue) to move the base point and drag $Q$ (orange) to choose the step $\delta = Q - P$. The readout compares the true height at $Q$ with the plane and the paraboloid. Rotate and press Front to see how the plane peels off while the paraboloid hugs the surface. Move $Q$ close to $P$ and watch both errors collapse (the plane's like $\delta^2$, the paraboloid's like $\delta^3$).

Order matters in the quadratic term: $\boldsymbol\delta^\top H\boldsymbol\delta$ puts the step on both sides of the matrix; it is not $H\boldsymbol\delta$ (a vector) and not $H\boldsymbol\delta^2$.

The Hessian term is a curvature times the step squared. It is negligible for tiny steps, which is why first order suffices for small learning rates, but it dominates for large ones.

Quick check: for $f(x,y) = x^2 + y^2$ at $(1,1)$, what do first- and second-order approximations predict for $\boldsymbol\delta = (0.5, 0)$, and what is the truth?

$f = 2$, $\nabla f = [2, 2]$, $H = 2I$. First order: $2 + 2(0.5) = 3$. Second order: $+\tfrac12\cdot2\cdot0.25 = 0.25$, giving $3.25$. Truth: $f(1.5, 1) = 2.25 + 1 = 3.25$. Exactly equal, because $f$ is a quadratic.

Taylor approximation in practice

In real work you rarely sum a long series. You keep one, two or three terms and ask: is this good enough for the numbers I care about? Typical uses are quick estimates by hand ("$\sqrt{101}$ is about 10.05"), safer numerics for tiny inputs, estimating how an error in the input spreads to the output, and building simplified models of a loss.

The skill is to match the order to the size of the step: for a very small change one term is enough; for a bigger one add the curvature.

  1. Estimate by hand. $\sqrt{101} = 10\sqrt{1.01} \approx 10\left(1 + \tfrac{0.01}{2}\right) = 10.05$. True: $10.04988$.
  2. Error propagation. A circle's radius is measured as $r = 10 \pm 0.1$. Area $A = \pi r^2$, so $dA \approx A'(r)\,dr = 2\pi r\,dr = 2\pi\cdot10\cdot0.1 = 6.28$. (Exact change: $\pi(10.1^2 - 10^2) = 6.31$.)
  3. Sigmoid at small input. $\sigma(0.5) \approx \tfrac12 + \tfrac{0.5}{4} = 0.625$. True: $0.6225$.
  4. Why log1p exists. In floating point $1 + 10^{-17}$ rounds to exactly $1$, so $\ln(1+10^{-17})$ comes out $0$. But $\ln(1+x)\approx x$ gives the right answer $10^{-17}$. Libraries use such expansions internally (np.log1p, np.expm1).
  5. Softplus. $\ln(1+e^x) \approx \ln 2 + \tfrac x2 + \tfrac{x^2}{8}$ near $0$ (derivatives: $\sigma(0) = \tfrac12$ and $\sigma'(0) = \tfrac14$).

A practical recipe.

  1. Choose the base point $a$ where $f$ and its derivatives are easy (often $a=0$, or a nearby "nice" number).
  2. Write $f(a+h) \approx f(a) + f'(a)h + \tfrac12 f''(a)h^2$ (add more terms if needed).
  3. Estimate the size of the next term, $\tfrac16|f'''|\,|h|^3$, to know the error.
  4. Stop when that error is smaller than what you need.

Propagation of error. If an input is off by $\Delta x$, the output is off by about $|f'(x)|\,\Delta x$. For several inputs, $\Delta f \approx \nabla f^\top\Delta\mathbf{x}$.

Second-order loss models. Gradient-boosted trees (XGBoost) approximate each loss as $\ell(\hat y+\delta) \approx \ell + g\,\delta + \tfrac12 h\,\delta^2$ with first derivative $g$ and second derivative $h$ of the loss (here $h$ is not the step), and then set $\delta = -g/h$: Newton's step.

Why do we need it?

Exact calculations are often too expensive, too unstable, or impossible by hand. A short Taylor expansion gives a fast answer together with an estimate of how wrong it can be.

Where is it used?

Numerical libraries (log1p, expm1, stable softplus), measurement error and uncertainty estimates in science, delta-method statistics, XGBoost and LightGBM, and linearised models of neural networks.

How is it used?

Keep one term for very small changes, two terms for moderate ones. Compare the size of the last kept term with the first dropped one. For tiny inputs use log1p/expm1 instead of log(1+x) or exp(x)-1.

Pick a function and slide $x$ (the step away from $0$). The table shows the true value and the estimates from the constant, linear, quadratic and cubic polynomials, each with its error. For $x = 0.1$ the error shrinks by a factor of 10 or more with each extra order; at $x = 0.8$ it shrinks much more slowly. Try $\sqrt{1+x}$ at $x=0.01$ (the $\sqrt{101}$ example) and find where each approximation stops being good enough for 3 decimal places.

Always say how big the step is compared with the curvature. "$\sin x\approx x$" is excellent for $x = 0.1$ (error $0.00017$) and poor for $x = 1.5$ (it says $1.5$ but the truth is $0.997$).

Do not expand around a bad point. The expansion point must be close to where you evaluate, and the function must be smooth there. Around $x=0$ the series of $\sqrt x$ does not exist (infinite slope), which is why we wrote $\sqrt{101}$ as $10\sqrt{1.01}$.

Quick check: use $e^x \approx 1 + x + \tfrac{x^2}{2}$ to estimate $e^{-0.2}$.

$1 - 0.2 + 0.02 = 0.82$. True $e^{-0.2} = 0.81873$. The error is about $0.0013$.

The remainder: how big is the error? core

A Taylor polynomial of degree $n$ copies the first $n$ features of the function. The error is whatever feature number $n+1$ (and later) contributes, and for a small step that is dominated by the first term we dropped. That term has $h^{n+1}$ in it, and powers of a small number shrink fast:

  • Step $h = 0.1$: $h^2 = 0.01$, $h^3 = 0.001$, $h^4 = 0.0001$.
  • Halve the step, and an error of size $h^2$ falls by $4$, of size $h^3$ by $8$, of size $h^4$ by $16$.

On a graph that plots the logarithm of the error against the logarithm of the step, these are straight lines with slopes $2$, $3$, $4$: the slope tells the order.

$f(x) = e^x$ at $a=0$, step $h$. First order: $P_1 = 1+h$. Second order: $P_2 = 1 + h + h^2/2$.

$h$error of $P_1$$\text{error}/h^2$error of $P_2$$\text{error}/h^3$
$0.2$$0.021403$$0.535$$0.001403$$0.175$
$0.1$$0.005171$$0.517$$0.000171$$0.171$
$0.05$$0.001271$$0.508$$0.0000211$$0.169$

The ratios settle at $\tfrac12 f''(0) = 0.5$ and $\tfrac16 f'''(0) = 0.1667$: the next Taylor coefficients.

Remainder (Lagrange form). If $f$ has $n+1$ derivatives, then

$$f(a+h) = P_n(a+h) + R_n, \qquad R_n = \frac{f^{(n+1)}(\xi)}{(n+1)!}\,h^{n+1}$$

for some unknown number $\xi$ between $a$ and $a+h$. We do not know $\xi$, but if we can bound $|f^{(n+1)}|$ by a number $M$ on that interval, then

$$|R_n| \le \frac{M}{(n+1)!}\,|h|^{n+1}.$$
  • $n=1$ (tangent line): error $\approx \tfrac12 f''\,h^2$, so it is $O(h^2)$.
  • $n=2$ (parabola): error $\approx \tfrac16 f'''\,h^3$, so it is $O(h^3)$.
  • In general the error after degree $n$ is $O(h^{n+1})$.

Example of a bound. For $\sin 1$ with $P_5(1) = 1 - \tfrac16 + \tfrac1{120} = 0.841667$ (which is also the degree-6 polynomial, since the $x^6$ term is zero), $M = 1$ and $n = 6$: $|R| \le \tfrac1{7!} = 0.000198$. The true error is $0.000196$ ✓. (The Lagrange form is awareness: you rarely compute $\xi$, you just use the bound.)

Check of the Lagrange form for $e^{0.1}$ with $n=1$: the true error is $0.005171 = \tfrac12 e^{\xi}(0.1)^2$, which gives $e^{\xi} = 1.034$, so $\xi = 0.034$, inside $(0, 0.1)$ ✓.

Why do we need it?

An approximation without an error estimate is just a guess. The remainder tells us how far to trust it, and how much better a higher order or a smaller step would be.

Where is it used?

Proofs of convergence for gradient descent and Newton's method (they use exactly these remainder bounds), choosing finite-difference step sizes, designing numerical libraries, and deciding how small a learning rate must be for first-order reasoning to hold.

How is it used?

Look at the first dropped term to estimate the error, or bound the next derivative. In experiments, plot error against step on log-log axes: the slope reveals the order, and a wrong slope reveals a bug.

Each line is the error of a Taylor polynomial of order 1 (orange), 2 (green) and 3 (blue), plotted against the step $h$ with both axes on a log scale. The readout fits the slope of each line: about 2, 3, 4, so error $\propto h^{n+1}$. Drag the base point $a$ with the slider. (At a point where $f''(a) = 0$, such as $\sin$ at $0$, the order-1 line is steeper than expected, slope about 3: the dropped term vanished, so the next one decides. Likewise at such a point the order-3 line is steeper, since $f^{(4)}$ also vanishes there.) (With much smaller steps than shown here, the lines would eventually flatten near $10^{-16}$: that is rounding noise in the computer.)

"$O(h^2)$" hides a constant. It says how the error scales, not how big it is. A large $f''$ makes the constant large, so "small" must be smaller.

Rounding noise. In floating-point arithmetic the error cannot go below about $10^{-16}$ relative to the numbers involved. This is why very tiny finite-difference steps stop helping.

Quick check: a first-order estimate has error $0.04$ for $h = 0.2$. Roughly what error do you expect for $h = 0.1$?

First order means error $\propto h^2$. Half the step gives a quarter of the error: about $0.01$.

Local approximation: only good near the point

A Taylor polynomial is a local copy: it is perfect at the base point and good close to it, but far away it can be badly wrong. A map of your street is a fine approximation of the city near your house and useless on another continent.

For some functions ($e^x$, $\sin x$) adding terms keeps widening the good region without limit. For others the good region has a fixed width however many terms you add: the series simply stops working beyond a certain distance from the base point, the radius of convergence. Beyond it, adding more terms makes things worse.

The geometric series $\dfrac1{1-x} = 1 + x + x^2 + \cdots$. Add up the terms at two points:

  • $x = 0.5$: the partial sums are $1,\ 1.5,\ 1.75,\ 1.875,\ 1.9375,\ \dots \to 2$ and indeed $\frac1{1-0.5} = 2$ ✓. Converges.
  • $x = 1.5$: the partial sums are $1,\ 2.5,\ 4.75,\ 8.125,\ 13.19,\ \dots$ growing without bound, while the function value is $\frac1{1-1.5} = -2$. Diverges.

The danger point is $x=1$ where the function itself blows up. The series works only for $|x| \lt 1$: radius of convergence $R = 1$.

A Taylor series around $a$ has a radius of convergence $R$: it converges to $f$ for $|x - a| \lt R$ and diverges for $|x-a| \gt R$.

  • $e^x$, $\sin x$, $\cos x$: $R = \infty$ (work everywhere).
  • $\dfrac{1}{1-x}$ and $\ln(1+x)$ around $0$: $R = 1$ (they break at $x = 1$ and $x=-1$ respectively).
  • Sigmoid around $0$: $R = \pi$. This is surprising: $\sigma(4)$ is perfectly well-behaved, yet the series diverges at $x = 4$. The reason is hidden in the complex numbers: $\sigma$ blows up at $x = \pm i\pi$, at distance $\pi$ from $0$ (awareness).

Even inside $R$ you only get an accurate answer if you take enough terms, and the closer to the edge the more you need. Rule of thumb: trust a short Taylor polynomial only for steps much smaller than the distance to the nearest trouble spot.

Why do we need it?

It stops us from trusting an approximation outside the region where it was built. In optimisation this is the idea of a trust region: the model is only reliable for small steps.

Where is it used?

Trust-region methods and the reason learning rates must be small, proximal policy optimisation in reinforcement learning (it limits how far the policy moves because its approximation is only local), and understanding why a quadratic model of a loss fails far from the current weights.

How is it used?

Keep steps small relative to how quickly the curvature changes. If a step makes the actual loss disagree with the model's prediction, shrink the step (line search, smaller learning rate, trust region).

Pick a function and move $x$. Left: $f$ (grey), the degree-$N$ polynomial (blue), and the safe zone $|x| \lt R$ (green). Right: the partial sums $S_0, S_1, S_2, \dots$ at your chosen $x$ (dots) against the true value (dashed line). Inside the zone they settle onto the line; outside they run away. Try the geometric series at $x = 0.5$ and $x = 1.3$; the sigmoid at $x = 2$ and $x = 3.6$; and $e^x$ at $x = 3$ (always converges, but needs many terms).

A Taylor polynomial is not a good global model. Even a high-degree polynomial matches the function only near the base point and shoots off elsewhere. Never extrapolate far.

Convergence is not the same as accuracy at a given N. Inside the radius the series converges, but at $x$ near the edge a degree-3 polynomial may still be off by a lot.

Quick check: the series of $\ln(1+x)$ has radius $1$. Is it safe to use at $x = 2$?

No. $|2| \gt 1$, so the series diverges there: the partial sums $2,\ 0,\ 2.67,\ -1.33, \dots$ swing wider and wider, while the true value $\ln3 = 1.0986$ is perfectly fine. Use a base point closer to $x=2$ instead: around $a=1$ the radius is $2$, so $x=2$ is safely inside.

Why gradient descent works: derived from the first-order expansion core

You stand on a foggy hillside. You cannot see the valley. All you can feel is the slope under your feet, so you treat the ground near you as a flat, tilted sheet (the first-order Taylor model). On a tilted sheet, which small step lowers your height the most? Straight against the slope, the way water would run.

That is gradient descent. It is not a rule somebody made up: it is what the first-order Taylor approximation says is the best small move.

$f(x,y) = \tfrac12(x^2 + 4y^2)$ at $P = (2, 1)$, so $f(P) = \tfrac12(4+4) = 4$. The gradient is $\mathbf{g} = [x, 4y] = [2, 4]$, with $\|\mathbf{g}\|^2 = 4 + 16 = 20$. Take the step $\boldsymbol\delta = -\eta\,\mathbf{g}$ with $\eta = 0.1$, i.e. $\boldsymbol\delta = [-0.2, -0.4]$.

  1. First-order prediction of the change: $\mathbf{g}^\top\boldsymbol\delta = 2(-0.2) + 4(-0.4) = -2.0$, which equals $-\eta\|\mathbf{g}\|^2 = -0.1\cdot20$ ✓.
  2. Actual change: the new point is $(1.8, 0.6)$ and $f = \tfrac12(3.24 + 4\cdot0.36) = \tfrac12(4.68) = 2.34$. So the change is $2.34 - 4 = -1.66$.
  3. The difference, $+0.34$, is the curvature term: $\tfrac12\boldsymbol\delta^\top H\boldsymbol\delta = \tfrac12\left(1\cdot0.04 + 4\cdot0.16\right) = \tfrac12(0.68) = 0.34$ ✓. The loss did go down, but by a bit less than the flat-sheet model promised.

Derivation.

  1. Model. For a small step $\boldsymbol\delta$: $f(\mathbf{x}+\boldsymbol\delta) \approx f(\mathbf{x}) + \mathbf{g}^\top\boldsymbol\delta$, where $\mathbf{g} = \nabla f(\mathbf{x})$ (a column).
  2. Question. Among all steps of a fixed small length $\|\boldsymbol\delta\| = \varepsilon$, which makes $\mathbf{g}^\top\boldsymbol\delta$ as negative (as downhill) as possible?
  3. Answer by Cauchy–Schwarz (Linear Algebra): $\mathbf{g}^\top\boldsymbol\delta \ge -\|\mathbf{g}\|\,\|\boldsymbol\delta\|$, with equality exactly when $\boldsymbol\delta$ points opposite to $\mathbf{g}$. So the steepest descent direction is $\boldsymbol\delta = -\varepsilon\,\mathbf{g}/\|\mathbf{g}\|$.
  4. Take a step proportional to $-\mathbf{g}$: $\boldsymbol\delta = -\eta\,\mathbf{g}$ with learning rate $\eta \gt 0$. The model predicts a change of $\mathbf{g}^\top(-\eta\mathbf{g}) = -\eta\|\mathbf{g}\|^2 \le 0$: the loss goes down (unless the gradient is zero).
  5. Update rule: $\mathbf{x}_{\text{new}} = \mathbf{x} - \eta\,\nabla f(\mathbf{x})$.

Why $\eta$ must be small. The model dropped the curvature term. Adding it back (second-order, exact for quadratics) the true change is

$$\Delta f \approx -\eta\|\mathbf{g}\|^2 + \tfrac12\eta^2\,\mathbf{g}^\top H\mathbf{g}.$$

The first term is linear in $\eta$, the second is quadratic. For small $\eta$ the first wins and $f$ falls; for large $\eta$ the second wins and $f$ can go up. The break-even is $\eta = \dfrac{2\|\mathbf{g}\|^2}{\mathbf{g}^\top H\mathbf{g}}$, and since $\mathbf{g}^\top H\mathbf{g} \le \lambda_{\max}\|\mathbf{g}\|^2$, any $\eta \lt 2/\lambda_{\max}$ is safe: exactly the learning-rate limit of Chapter 2.10.

Why do we need it?

It tells us why "step against the gradient" is the right move, and exactly how it can fail (when the step is too big for the curvature). Knowing the reason lets you fix training problems instead of guessing.

Where is it used?

Every training loop: SGD, momentum, Adam all build on this step. It also underlies convergence proofs, learning-rate schedules and the step-size rules in line search.

How is it used?

Compute the gradient, subtract $\eta$ times it from the parameters. If the loss goes up, the first-order model failed: the step was too big compared with the curvature, so reduce $\eta$.

Press Next to walk through the derivation on a contour map (blue = low, light = high). Drag the start point $P$ if you like. In step 2 turn the direction dial: the dots show the predicted change $\mathbf{g}^\top\boldsymbol\delta$ for each direction; the best one is straight against the gradient. In the last step slide $\eta$ and compare the predicted change with the actual change: they agree for small $\eta$ and part ways for large $\eta$, until the loss goes up.

The gradient is the steepest direction only for tiny steps, and only in the usual (Euclidean) sense of distance. With big steps the curvature bends the path; with a different distance (or a different metric) the best direction changes. That is the idea behind natural gradient and, loosely, behind preconditioned methods such as Adam (for the ill-conditioned bowls they try to fix, see Chapter 2.10).

"Steepest" is not "fastest". In a long narrow valley the steepest direction points across the valley, not along it, so gradient descent zig-zags. Curvature (the next section) can fix that.

Quick check: $f(x) = x^2$ at $x = 3$. What does the first-order model predict for one gradient step with $\eta = 0.1$, and what happens?

$f' = 6$, so the step is $-0.6$ and the predicted change is $-\eta f'^2 = -0.1\cdot36 = -3.6$. Actual: $x = 2.4$, $f = 5.76$, change $5.76 - 9 = -3.24$. The loss fell, a little less than predicted (the missing curvature term is $+0.36$).

Why Newton's method works: derived from the second-order expansion core

Same foggy hillside, but now you can feel both the slope and the bend of the ground. Your model is no longer a flat sheet but a bowl (a parabola). A bowl has a lowest point you can compute directly. So instead of a cautious small step, jump straight to the bottom of the bowl, then look again.

Gradient descent and Newton's method are the same idea at two accuracy levels: flat-sheet model (first order) versus bowl model (second order).

Minimise $f(x) = e^x - 2x$ from $x = 0$. There $f = 1$, $f' = e^0 - 2 = -1$, $f'' = e^0 = 1$.

  1. Second-order model: $m(\delta) = 1 - \delta + \tfrac12\delta^2$.
  2. Its slope is $m'(\delta) = -1 + \delta$, which is zero at $\delta = 1$ (the bottom of the parabola).
  3. So the new point is $x = 0 + 1 = 1$. (The true minimum is $\ln 2 = 0.693$, so we overshot a little because the true curve bends up faster than the parabola. One more step from $x=1$ gives $0.736$.)

Derivation.

  1. Model. $m(\boldsymbol\delta) = f + \mathbf{g}^\top\boldsymbol\delta + \tfrac12\boldsymbol\delta^\top H\boldsymbol\delta$, the second-order Taylor expansion at the current point.
  2. Find its lowest point (requires $H$ positive definite, a bowl). Set its gradient with respect to $\boldsymbol\delta$ to zero. The gradient of $\mathbf{g}^\top\boldsymbol\delta$ is $\mathbf{g}$ and the gradient of $\tfrac12\boldsymbol\delta^\top H\boldsymbol\delta$ is $H\boldsymbol\delta$ (symmetric $H$; Chapter 2.6). So $\mathbf{g} + H\boldsymbol\delta = \mathbf{0}$.
  3. Solve: $\boldsymbol\delta = -H^{-1}\mathbf{g}$.
  4. Update: $\mathbf{x}_{\text{new}} = \mathbf{x} - H^{-1}\nabla f$. In 1D: $x_{\text{new}} = x - f'(x)/f''(x)$.

Gradient descent is Newton with a guessed curvature. Replace the true Hessian in the model by $\frac1\eta I$: $m_{\text{GD}}(\boldsymbol\delta) = f + \mathbf{g}^\top\boldsymbol\delta + \dfrac{1}{2\eta}\|\boldsymbol\delta\|^2$. Setting its gradient to zero, $\mathbf{g} + \boldsymbol\delta/\eta = \mathbf{0}$, gives $\boldsymbol\delta = -\eta\,\mathbf{g}$: the gradient step. So the learning rate $\eta$ is "one over an assumed curvature". If you guess a curvature smaller than the truth, the step is too long and it overshoots.

When it works well: if $f$ is a quadratic the model is exact, so one step lands on the minimum. Near a smooth minimum it converges very quickly (the number of correct digits roughly doubles each step). When it misbehaves: if $H$ is not positive definite the model's flat spot is not a minimum, and if the step is large the model is no longer trustworthy (see "Local approximation").

Why do we need it?

It explains where the Newton step comes from, so we can see its strengths (curvature-aware, scale-free) and weaknesses (needs a bowl, costs a matrix solve) rather than just trusting the formula.

Where is it used?

Logistic regression and generalised linear models (iteratively reweighted least squares), L-BFGS and trust-region optimisers, Gauss–Newton and Levenberg–Marquardt for least squares, and XGBoost's leaf values $-g/h$.

How is it used?

Build the quadratic model at the current point, solve $H\boldsymbol\delta = -\mathbf{g}$, and move. Safeguards: add $\lambda I$ to make $H$ positive definite, limit the step length (trust region) or do a line search.

Press Next to walk through the derivation. The blue parabola is the second-order model at your point (drag the point if you wish). Step 2 marks its bottom. In step 4 an orange parabola appears: the model gradient descent uses, where the true curvature is replaced by $1/\eta$. Slide $\eta$: when $1/\eta$ equals $f''$ the two parabolas coincide, and the gradient step is the Newton step. Press Jump & repeat to move to the new point. Pick the double-well and start near $x = 0.5$, where the curve frowns: the "bottom" of the model is really a top.

The model is local. Jumping to the bottom of a parabola that only matches the function near your point can overshoot badly. Newton's method is safest close to the answer.

The model's bottom must be a bottom. If $H$ has a negative eigenvalue the parabola curves down in that direction and the flat spot is a saddle or a peak; the step $-H^{-1}\mathbf{g}$ then walks to it (Chapter 2.10).

Quick check: $f(x) = x^2 - 6x$. Starting from $x=0$, where does one Newton step land, and why?

$f' = 2x - 6 = -6$ and $f'' = 2$. The step is $-(-6)/2 = 3$, so $x = 3$, which is the exact minimum ($f'(3) = 0$). The function is itself a parabola, so its second-order Taylor model is exact.

Recap, cheat sheet and practice

  • Near a point $a$, a smooth function is copied by the Taylor polynomial $\sum_k \frac{f^{(k)}(a)}{k!}(x-a)^k$. The $1/k!$ appears because differentiating $(x-a)^k$ exactly $k$ times gives $k!$.
  • First order: $f(a+h)\approx f(a)+f'(a)h$ (tangent line). Second order adds $\tfrac12f''(a)h^2$ (curvature). In many variables: $f(\mathbf{x}+\boldsymbol\delta)\approx f+\nabla f^\top\boldsymbol\delta+\tfrac12\boldsymbol\delta^\top H\boldsymbol\delta$.
  • Know the series of $e^x$, $\sin x$, $\cos x$, $\ln(1+x)$, $\frac1{1-x}$ and the sigmoid $\frac12+\frac x4-\frac{x^3}{48}$, and how to derive any of them.
  • The error after degree $n$ is about the first dropped term, $O(h^{n+1})$: $h^2$ for the tangent line, $h^3$ for the parabola. On a log–log plot the slope is the order.
  • The approximation is local: it is only trusted for steps small compared with the curvature, and series have a radius of convergence.
  • Gradient descent is the best step for the first-order model; Newton's method is the bottom of the second-order model. Gradient descent is Newton with the Hessian replaced by $\frac1\eta I$.

Cheat sheet

IdeaFormulaPicture
Taylor polynomial$P_n(x)=\sum_{k=0}^n \dfrac{f^{(k)}(a)}{k!}(x-a)^k$polynomial that copies $f$ near $a$
First order$f(a)+f'(a)h$; error $\sim\tfrac12f''h^2$tangent line
Second order$f(a)+f'(a)h+\tfrac12f''(a)h^2$; error $\sim\tfrac16f'''h^3$parabola
Multivariate$f+\nabla f^\top\boldsymbol\delta+\tfrac12\boldsymbol\delta^\top H\boldsymbol\delta$tangent plane, paraboloid
$e^x$, $\sin x$, $\cos x$$\sum\frac{x^k}{k!}$; $x-\frac{x^3}{3!}+\cdots$; $1-\frac{x^2}{2!}+\cdots$valid for all $x$
$\ln(1+x)$, $\frac1{1-x}$$x-\frac{x^2}2+\frac{x^3}3-\cdots$; $1+x+x^2+\cdots$radius $1$
Sigmoid$\frac12+\frac x4-\frac{x^3}{48}+\cdots$radius $\pi$
Remainder$R_n=\dfrac{f^{(n+1)}(\xi)}{(n+1)!}h^{n+1}$log–log slope $n+1$
Gradient descent$\mathbf{x}-\eta\nabla f$ (from the first-order model)step against the slope
Newton$\mathbf{x}-H^{-1}\nabla f$ (from the second-order model)jump to the bowl's bottom
Code it · NumPy

import numpy as np
from math import factorial

# 1) Taylor polynomial of e^x around 0
def exp_taylor(x, n):
    return sum(x**k / factorial(k) for k in range(n + 1))

for n in range(5):
    print(n, exp_taylor(1.0, n))        # 1.0, 2.0, 2.5, 2.6667, 2.7083  ->  e = 2.71828...

# 2) the error shrinks like h^2 (first order) and h^3 (second order)
for h in (0.2, 0.1, 0.05):
    e1 = abs(np.exp(h) - (1 + h))
    e2 = abs(np.exp(h) - (1 + h + h*h/2))
    print(h, e1 / h**2, e2 / h**3)      # ratios settle near 0.5 and 0.1667

hs = np.logspace(-2, -0.5, 20)
err1 = np.abs(np.exp(hs) - (1 + hs))
err2 = np.abs(np.exp(hs) - (1 + hs + hs**2/2))
print(np.polyfit(np.log10(hs), np.log10(err1), 1)[0])   # about 2.0 (slope of the log-log line)
print(np.polyfit(np.log10(hs), np.log10(err2), 1)[0])   # about 3.0

# 3) multivariate: f(x, y) = x^2 y + y^2 around (1, 2), step delta = (0.1, -0.1)
f = lambda p: p[0]**2 * p[1] + p[1]**2
grad = lambda p: np.array([2*p[0]*p[1], p[0]**2 + 2*p[1]])
hess = lambda p: np.array([[2*p[1], 2*p[0]], [2*p[0], 2.0]])
p0 = np.array([1.0, 2.0]); d = np.array([0.1, -0.1])
first = f(p0) + grad(p0) @ d
second = first + 0.5 * d @ hess(p0) @ d
print(first, second, f(p0 + d))         # 5.9  5.91  5.909

# 4) one gradient step vs the first-order prediction, f = 0.5*(x^2 + 4 y^2) at (2, 1)
A = np.diag([1.0, 4.0]); g = A @ np.array([2.0, 1.0]); eta = 0.1
q = lambda z: 0.5 * z @ A @ z
z0 = np.array([2.0, 1.0])
print(-eta * g @ g, q(z0 - eta*g) - q(z0))   # predicted -2.0, actual -1.66 (curvature term is +0.34)

# 5) why log1p exists
print(np.log(1 + 1e-17), np.log1p(1e-17))    # 0.0  1e-17

# 6) sigmoid near 0: 1/2 + x/4 - x^3/48
sig = lambda x: 1 / (1 + np.exp(-x))
x = 0.4
print(sig(x), 0.5 + x/4 - x**3/48)           # 0.598688  0.598667
Test yourself

1. The Taylor series of $\sin x$ around $0$ starts…

The derivatives of $\sin$ at $0$ are $0,1,0,-1,0,1,\dots$, so only odd powers appear, with alternating signs, divided by the factorials. (The first option is $\cos x$, the last is $e^x$.)

2. A first-order Taylor estimate has error $0.08$ for a step $h = 0.4$. About what error do you expect for $h = 0.2$?

First-order error scales like $h^2$. Halving $h$ divides the error by $4$: $0.08/4 = 0.02$.

3. Which expression is the second-order multivariate Taylor approximation?

The slope term is the gradient dotted with the step; the curvature term is $\tfrac12$ times the step on both sides of $H$, giving a single number.

4. Why can the Taylor series of $\ln(1+x)$ not be used at $x = 2$?

$\ln 3 = 1.0986$ exists and $\ln(1+x)$ is smooth, but the series around $0$ only converges for $|x|\lt1$ (the trouble spot is at $x=-1$, at distance $1$ from the base point).

5. Newton's method is derived by…

The second-order model $f + \mathbf{g}^\top\boldsymbol\delta + \tfrac12\boldsymbol\delta^\top H\boldsymbol\delta$ has a flat spot where $\mathbf{g} + H\boldsymbol\delta = 0$. A first-order (flat-sheet) model has no minimum at all.

6. A model's first-order Taylor expansion predicts that a gradient step lowers the loss by $2.0$, but the loss goes up. The most likely reason is…

The true change is $-\eta\|\mathbf{g}\|^2 + \tfrac12\eta^2\mathbf{g}^\top H\mathbf{g}$. For large $\eta$ the quadratic term wins. Reduce the learning rate (below $2/\lambda_{\max}$ it always goes down on a quadratic).

Practice problems

A. Find the degree-3 Taylor polynomial of $f(x)=\ln(1+x)$ at $0$ and use it to estimate $\ln 1.2$.

$f(0)=0$, $f'=\frac1{1+x}\to1$, $f''=-\frac1{(1+x)^2}\to-1$, $f'''=\frac2{(1+x)^3}\to2$. So $P_3 = x - \frac{x^2}{2} + \frac{2x^3}{6} = x - \frac{x^2}2 + \frac{x^3}3$. At $x=0.2$: $0.2 - 0.02 + 0.002667 = 0.182667$. True: $\ln1.2 = 0.182322$; error $0.00034$.

B. Use the tangent line of $f(x)=\sqrt{x}$ at $a=25$ to estimate $\sqrt{26}$, and say whether the estimate is above or below the truth.

$f(25)=5$, $f'(x)=\frac1{2\sqrt x}$, $f'(25)=\frac1{10}$. So $\sqrt{26}\approx5+0.1=5.1$. True: $5.09902$. The estimate is slightly above: $f''=-\frac14x^{-3/2}\lt0$ (a frown), so the tangent line lies above the curve.

C. For $f(x,y)=e^{x}\cos y$ find the second-order Taylor approximation at $(0,0)$ and estimate $f(0.1,\,0.2)$.

$f=e^x\cos y$: at $(0,0)$ $f=1$; $f_x=e^x\cos y=1$; $f_y=-e^x\sin y=0$; $f_{xx}=1$; $f_{xy}=-e^x\sin y=0$; $f_{yy}=-e^x\cos y=-1$. So $f\approx1+x+\tfrac12x^2-\tfrac12y^2$. At $(0.1,0.2)$: $1+0.1+0.005-0.02=1.085$. True: $e^{0.1}\cos0.2=1.10517\times0.98007=1.08314$. Error $0.0019$.

D. Derive the Taylor series of $\cos x$ at $0$ from that of $\sin x$ by differentiating.

$\sin x = x - \frac{x^3}{3!} + \frac{x^5}{5!} - \cdots$. Differentiate term by term: $\cos x = 1 - \frac{3x^2}{3!} + \frac{5x^4}{5!} - \cdots = 1 - \frac{x^2}{2!} + \frac{x^4}{4!} - \cdots$, because $\frac{3}{3!}=\frac1{2!}$, $\frac5{5!}=\frac1{4!}$, and so on.

E. Take one gradient step and one Newton step on $f(x)=x^4$ from $x=1$ with $\eta=0.05$. Compare.

$f'=4x^3=4$, $f''=12x^2=12$. Gradient step: $1-0.05\cdot4=0.8$. Newton step: $1-4/12=0.667$. Newton moves further because it knows the curvature: it uses an effective $\eta=1/12=0.083$ here. (The true minimum is $0$; Newton shrinks the distance by a factor $2/3$ each step for this function.)

F. By what factor does the error of a degree-3 Taylor polynomial drop when the step is halved? And for the tangent line?

Degree 3 means error $O(h^4)$: halving $h$ divides the error by $2^4=16$. The tangent line is $O(h^2)$: by $2^2=4$. (When the leading coefficient happens to vanish, e.g. $f''(a)=0$, the observed factor is larger.)

Chapter 2.12

Linearization

Zoom in far enough on any smooth curve and it looks like a straight line. That one fact lets us swap a hard, bendy function for an easy, flat one, as long as we stay close to where we stand. Almost every optimisation method in machine learning is built on it.

  • See a smooth function as "a straight line (or flat plane) up close"
  • Write the linear approximation $f(x+\delta) \approx f(x) + f'(x)\,\delta$ and its many-variable forms with the gradient and the Jacobian
  • Understand the Jacobian as a local linear transformation: near a point, a curvy map behaves like a matrix
  • Measure the error: it is about $\tfrac12\delta^\top H\delta$, so it shrinks like $\delta^2$ (halve the step, quarter the error)
  • Know when linearization fails: kinks and very high curvature
  • Use it: error propagation, Newton's method, linearizing a sigmoid neuron, gradient descent with small learning rates, and why ReLU networks are piecewise linear

This chapter pulls together the derivative (Chapter 2.3), the gradient (2.4), the Jacobian (2.5), the Hessian (2.10) and Taylor series (2.11). If a symbol looks unfamiliar, jump back to that chapter. Conventions, as everywhere in this guide: the gradient $\nabla f$ is a column vector, and the Jacobian of $F:\mathbb{R}^n\to\mathbb{R}^m$ is an $m\times n$ matrix whose entry in row $i$, column $j$ is $\partial F_i/\partial x_j$. When we talk about matrices as "things that move arrows around", see Linear transformations in the Linear Algebra guide.

Linear approximation: zoom in until it is straight core

The Earth is round, yet a map of your town is flat, and nobody notices the mistake. Zoom in on a curved surface far enough and it looks flat.

A curve on a graph behaves the same way. Pick a point and zoom in. The curve bends less and less, until it is hard to tell apart from a straight line. That line is the tangent line from Chapter 2.3.

So near a point we can pretend the curve is the line. Lines are easy: you only need a starting value and a slope. This pretend-it-is-a-line trick is called linear approximation or linearization.

Take $f(x) = x^2$ near $x = 1$. At $x=1$ the value is $f(1)=1$ and the slope is $f'(1) = 2$. So the tangent line is "start at 1, rise by 2 for every 1 step to the right".

  1. Go $0.1$ to the right, to $x = 1.1$. The line says: $1 + 2\cdot 0.1 = 1.2$. The true value is $1.1^2 = 1.21$. The gap is $0.01$.
  2. Go only $0.01$ to the right, to $x = 1.01$. The line says: $1 + 2\cdot 0.01 = 1.02$. The true value is $1.01^2 = 1.0201$. The gap is $0.0001$.

We moved 10 times closer and the gap became 100 times smaller. That is the "zoom in and it gets straighter" effect in numbers.

The linear approximation (or linearization) of $f$ at the point $a$ is the function

$$L(x) = f(a) + f'(a)\,(x - a).$$

Writing the step as $h = x - a$, this says

$$f(a + h) \;\approx\; f(a) + f'(a)\,h.$$
  • $f(a)$ is where we start (the height of the curve at $a$).
  • $f'(a)$ is the slope (how fast the height changes per step).
  • $h$ is the step. "Height change $\approx$ slope $\times$ step."

Where does it come from? The derivative is defined by $f'(a) = \lim_{h\to 0}\dfrac{f(a+h)-f(a)}{h}$. For a small (but not zero) $h$ the fraction is close to $f'(a)$:

$$\frac{f(a+h)-f(a)}{h} \approx f'(a) \;\;\Longrightarrow\;\; f(a+h) - f(a) \approx f'(a)\,h \;\;\Longrightarrow\;\; f(a+h) \approx f(a) + f'(a)\,h.$$

We only multiplied both sides by $h$ and moved $f(a)$ across. That is the whole idea.

Why do we need it?

Most functions in machine learning are curved and messy, but straight lines are easy to reason about and to compute with. Linearization lets us do easy maths and still be nearly right, as long as we stay near the point we know.

Where is it used?

Gradient descent (each step trusts a straight-line model of the loss), Newton's method, error bars on measurements, sensitivity analysis, the Gauss–Newton method in least squares, and the extended Kalman filter in robotics and GPS.

How is it used?

Compute the value $f(a)$ and the slope $f'(a)$ once. Then estimate $f$ at any nearby point with one multiplication and one addition, instead of evaluating the hard function again.

Start at zoom ×1: the blue curve clearly bends away from the orange dashed line (the red gap is the difference). Now slide zoom to the right. Watch the percentage in the readout fall towards 0 and the curve melt into the line. Then change the function and the point a and repeat: it works for every smooth curve.

"Approximately equal" is not "equal". The line is only an approximation. It is exactly right at the point $a$ and slowly gets worse as you walk away. The next section studies how fast.

Zoom is relative. A curve that looks straight in a window of width 0.01 may look very curved in a window of width 10. "Straight" always means "straight at this scale".

Quick check: linearize $f(x)=x^2$ at $a=3$, then estimate $3.1^2$.

$f(3)=9$ and $f'(3)=6$, so $L(x) = 9 + 6(x-3)$. At $x = 3.1$: $L(3.1) = 9 + 6\cdot0.1 = 9.6$. The true value is $9.61$, so we are off by only $0.01$.

Local approximation: how near is "near"?

A street map of your town is perfect for walking to the shop and useless for flying to another continent. A linear approximation is a street map: very good close to the spot where you drew it, worse and worse as you travel away.

"Local" simply means "close to the point $a$". The word that matters is how close. That depends on two things: how much accuracy you need (your tolerance), and how sharply the curve bends. A gentle curve keeps the line useful for a long way. A sharp curve spoils it quickly.

Linearize $f(x)=e^x$ at $a=0$. Here $f(0)=1$ and $f'(0)=e^0=1$, so $L(x) = 1 + x$. Compare the line with the truth at three distances:

distance $h$true $e^h$line $1+h$gap
$0.1$$1.10517$$1.1$$0.0052$
$0.5$$1.64872$$1.5$$0.1487$
$2$$7.38906$$3$$4.389$

Close by, the line is excellent. Two steps away it is far off. Moving 5 times further (from $0.1$ to $0.5$) made the gap about 29 times bigger. The gap does not grow in step with the distance. It grows faster.

A function $f$ is differentiable at $a$ exactly when its linear approximation is good locally, in this precise sense:

$$f(a+h) = f(a) + f'(a)\,h + \underbrace{(\text{error})}_{\text{shrinks faster than } h}, \qquad \frac{\text{error}}{h}\to 0 \text{ as } h\to 0.$$

"The error shrinks faster than $h$" is the careful way to say "the curve really does become a line". Given a tolerance $\varepsilon$ (the biggest mistake you will accept), the trust zone is the set of $x$ near $a$ where $|f(x) - L(x)| \le \varepsilon$. Inside it you may use the line. Outside it you may not.

Why do we need it?

A tangent line is only a promise about the neighbourhood of one point. To use it safely we must know how big that neighbourhood is, otherwise we may trust it far from home and get nonsense.

Where is it used?

The learning rate in gradient descent (how far a step may go before the straight-line model of the loss stops being true), trust-region optimisers, step-size control in ODE solvers, and numerical derivative checks.

How is it used?

Decide the error you can accept, then keep every step small enough that you stay in the trust zone. If your steps seem to "overshoot", the zone was smaller than you thought: shrink the step.

Drag the blue dot along the curve. The green band is the trust zone: every $x$ where the orange tangent line is within the tolerance of the curve. (1) Make the tolerance smaller and the band narrows. (2) Pick 4 tanh x and drag to the flat right-hand side: the curve barely bends there, so the zone becomes huge. (3) Pick x³ − 3x and drag to x = 2, where it bends strongly: the zone is tiny.

The trust zone is not symmetric in general. For $e^x$ the curve bends upward, so on one side the line stays good for longer than on the other. Do not assume "plus or minus the same amount".

Quick check: the tangent line to $\sin x$ at $0$ is $L(x)=x$. How wrong is it at $x=0.1$?

$\sin 0.1 = 0.099833\ldots$, so the line's answer $0.1$ is off by about $0.000167$. That is less than two parts in a thousand, so the line is excellent at this distance.

First-order Taylor approximation (many inputs) core

Now let the input be several numbers, such as the weights of a model. Nudge each input a little. How does the output change?

Because things look flat up close, each nudge acts on its own and the effects simply add up. A nudge of $0.1$ in the first input changes the output by "slope in that direction times $0.1$". A nudge of $-0.2$ in the second input does the same with its own slope. The total change is the sum.

Picture a surface (a landscape). Near your feet it looks like a tilted flat sheet, the tangent plane. The linear approximation just reads heights off that sheet instead of off the real, curved ground.

Let $f(x,y) = x^2y$. We stand at $(x,y)=(1,2)$ and step by $\delta = (0.1,\,-0.2)$.

  1. Value at the start: $f(1,2) = 1^2\cdot 2 = 2$.
  2. Partial derivatives: $\partial f/\partial x = 2xy = 4$ and $\partial f/\partial y = x^2 = 1$ at $(1,2)$. So $\nabla f = [4,\,1]^\top$.
  3. First-order change: $\nabla f^\top\delta = 4\cdot 0.1 + 1\cdot(-0.2) = 0.4 - 0.2 = 0.2$.
  4. Linear estimate: $f(1.1,\,1.8) \approx 2 + 0.2 = 2.2$.
  5. True value: $1.1^2\cdot 1.8 = 1.21\cdot 1.8 = 2.178$. The estimate is off by $0.022$.

The estimate is within about 1 percent after a step of size roughly $0.22$. Not bad for two multiplications and an addition.

Let $\mathbf{x}$ be a point and $\boldsymbol{\delta}$ a small step. The first-order Taylor approximation (the linearization) is:

functionapproximation
$f:\mathbb{R}\to\mathbb{R}$$f(x+\delta) \approx f(x) + f'(x)\,\delta$
$f:\mathbb{R}^n\to\mathbb{R}$ (a loss)$f(\mathbf{x}+\boldsymbol\delta) \approx f(\mathbf{x}) + \nabla f(\mathbf{x})^\top\boldsymbol\delta = f(\mathbf{x}) + \sum_i \dfrac{\partial f}{\partial x_i}\,\delta_i$
$F:\mathbb{R}^n\to\mathbb{R}^m$ (a layer)$F(\mathbf{x}+\boldsymbol\delta) \approx F(\mathbf{x}) + J(\mathbf{x})\,\boldsymbol\delta$

Here $\nabla f^\top\boldsymbol\delta$ is the dot product of the gradient (a column) with the step. The Taylor series of Chapter 2.11 keeps more terms. Keeping only the first gives the linearization.

Derivation for the middle row. Walk along the straight path $g(t) = f(\mathbf{x} + t\boldsymbol\delta)$, so $g(0) = f(\mathbf{x})$ and $g(1) = f(\mathbf{x}+\boldsymbol\delta)$. This is a one-variable function, so the line rule gives $g(1)\approx g(0) + g'(0)$. The chain rule (Chapter 2.8) gives $g'(0) = \nabla f(\mathbf{x})^\top\boldsymbol\delta$. Putting these together:

$$f(\mathbf{x}+\boldsymbol\delta) \approx f(\mathbf{x}) + \nabla f(\mathbf{x})^\top\boldsymbol\delta. \qquad\blacksquare$$
Why do we need it?

A model can have millions of weights. We cannot afford to re-run it for every possible change. The gradient is one cheap number-list that predicts the effect of any small change in all the weights at once.

Where is it used?

The gradient-descent update (we choose $\boldsymbol\delta=-\eta\nabla f$ because the formula says it lowers the loss), saliency maps (which pixels matter), adversarial examples, influence functions, and Gauss–Newton.

How is it used?

Evaluate $f$ and $\nabla f$ at the current point once. Then the predicted change for any small step $\boldsymbol\delta$ is the dot product $\nabla f^\top\boldsymbol\delta$. Trust it only while $\boldsymbol\delta$ is small.

Rotate by dragging the background. Drag the purple dot P (where we stand) and the green dot Q (where we step to) across the surface. The orange sheet is the tangent plane at P. The red bar is the error: the vertical gap between the true surface (at the green dot Q) and the plane (the orange dot above or below it). Make Q close to P and the error all but vanishes. Look from the Side view: the plane just touches the surface at P. Try the other functions too.

The step must be small, in every direction. The formula does not say "linear is good". It says "linear is good for small $\boldsymbol\delta$". A step that is tiny in one input and huge in another is not small.

Transpose matters. $\nabla f$ is a column, so we write $\nabla f^\top\boldsymbol\delta$ (row times column, a single number). Writing $\nabla f\,\boldsymbol\delta$ would multiply a column by a column, which does not make sense.

Quick check: for $f(x,y)=x^2+y^2$ at $(3,4)$, estimate $f(3.1,\,4)$ with the gradient.

$f(3,4)=25$ and $\nabla f = [2x,\,2y]^\top = [6,\,8]^\top$. The step is $\boldsymbol\delta=(0.1,\,0)$, so $\nabla f^\top\boldsymbol\delta = 6\cdot0.1 + 8\cdot 0 = 0.6$. Estimate: $25.6$. The truth is $3.1^2 + 16 = 25.61$.

The Jacobian as a local linear transformation core

A function $F:\mathbb{R}^2\to\mathbb{R}^2$ takes a point of the plane and sends it to another point. Imagine it as a rubber sheet with a grid drawn on it: $F$ pushes, stretches, twists and bends the sheet. Far away the grid may look wildly curved.

Now zoom in on one point. The same trick as before happens: the bent grid lines straighten, and the little grid squares become little parallelograms of equal size. That is exactly what a matrix does to a grid. (See Chapter 1.5 of the Linear Algebra guide: a matrix sends the grid to a tilted, stretched grid, and its columns say where the two basic arrows land.)

So near a point, every smooth map behaves like a matrix. That matrix is the Jacobian. Its first column says "where does a nudge in $x_1$ go?" and its second column says the same for $x_2$.

Let $F(x,y) = (x^2 - y^2,\; 2xy)$. We stand at $(1,1)$.

  1. Where it lands: $F(1,1) = (1-1,\; 2) = (0,\,2)$.
  2. The partial derivatives: $\partial F_1/\partial x = 2x$, $\partial F_1/\partial y = -2y$, $\partial F_2/\partial x = 2y$, $\partial F_2/\partial y = 2x$. At $(1,1)$ the Jacobian is $$J = \begin{bmatrix} 2 & -2 \\ 2 & 2 \end{bmatrix}.$$
  3. Take a small step $\boldsymbol\delta = (0.1,\,0.05)$. Then $J\boldsymbol\delta = (2\cdot0.1 - 2\cdot0.05,\; 2\cdot0.1 + 2\cdot0.05) = (0.1,\,0.3)$.
  4. Linear estimate: $F(1.1,\,1.05) \approx (0,2) + (0.1,0.3) = (0.1,\,2.3)$.
  5. True value: $(1.1^2 - 1.05^2,\; 2\cdot1.1\cdot1.05) = (1.21 - 1.1025,\; 2.31) = (0.1075,\,2.31)$.

The estimate is off by only $(0.0075,\,0.01)$. The determinant is $\det J = 2\cdot2 - (-2)\cdot 2 = 8$: tiny patches of area get stretched by a factor of 8 near this point.

For $F:\mathbb{R}^n\to\mathbb{R}^m$, the Jacobian at $\mathbf{x}$ is the $m\times n$ matrix of all first partial derivatives, with entry $\partial F_i/\partial x_j$ in row $i$, column $j$. Linearization says:

$$F(\mathbf{x}+\boldsymbol\delta) \;\approx\; F(\mathbf{x}) + J(\mathbf{x})\,\boldsymbol\delta.$$

So near $\mathbf{x}$, $F$ acts like the affine map "shift to $F(\mathbf{x})$, then apply the matrix $J(\mathbf{x})$ to the step". Reading the matrix:

  • Column $j$ of $J$ is where the little step $\mathbf{e}_j$ (a nudge of input $j$) goes, per unit of nudge.
  • $|\det J|$ (for $n=m$) is the factor by which tiny areas (or volumes) are scaled. A negative determinant means the map flips orientation. See the determinant.
  • If $m=1$ (a loss), $J$ is a single row, which is just $\nabla f^\top$. So the gradient is the Jacobian of a scalar function.
Why do we need it?

A neural-network layer is a curved map from vectors to vectors, which is hard to analyse. Its Jacobian replaces it, locally, by one matrix, and we know a lot about matrices (multiplying, inverting, eigenvalues, determinants).

Where is it used?

Backpropagation (each layer's local Jacobian is multiplied along the chain), normalising flows (the determinant of the Jacobian tracks how density is stretched), coordinate changes such as polar to Cartesian, and Gauss–Newton and Levenberg–Marquardt solvers.

How is it used?

Compute (or let autodiff compute) the Jacobian at the current point. Multiply it by a small input change to predict the output change, or look at its determinant or singular values to see how the map stretches space.

Left: the input plane. Drag the blue dot; the little square around it has size s. Right: what the map $F$ does to that little square, magnified so it stays readable. The orange curved shape is the true image. The dashed purple parallelogram is the Jacobian's prediction. (1) Shrink s and the two shapes become the same. (2) The teal and pink arrows are the Jacobian's two columns: where the $x$-edge and the $y$-edge go. (3) Try the three maps.

The Jacobian depends on the point. Move the dot and the matrix changes. It is a local picture: a different matrix at every location.

Rows or columns? Row $i$ is the gradient of output $i$. Column $j$ is the effect of input $j$ on all the outputs. With the convention used throughout this guide the shape is outputs $\times$ inputs.

Quick check: what is the Jacobian of the linear map $F(\mathbf{x}) = A\mathbf{x}$, and how good is its linearization?

$J = A$ at every point, because the partial derivative of $\sum_j A_{ij}x_j$ with respect to $x_j$ is $A_{ij}$. The linearization $F(\mathbf{x}) + A\boldsymbol\delta$ equals $A(\mathbf{x}+\boldsymbol\delta)$ exactly. A linear map has no curvature, so there is no error at all, for any step size.

How big is the error? It shrinks like $\delta^2$ core

The line ignores one thing: the curve bends. The error is the amount of bending you collected while walking the step. Two facts follow from that picture.

  • The bendier the curve (more curvature, the second derivative), the bigger the error.
  • If you walk half as far, you collect less than half of the bending. In fact you collect a quarter: the bend builds up with the square of the distance.

So a small step does not just give a small error. It gives a very small error. Halve the step and the error drops 4 times.

One variable. Linearize $e^x$ at $0$ again ($L(h) = 1 + h$). The curvature is $f''(0) = e^0 = 1$, so we guess the error is $\tfrac12 f''\,h^2 = \tfrac12h^2$.

step $h$true error $e^h - 1 - h$guess $\tfrac12h^2$
$0.1$$0.005171$$0.005$
$0.05$$0.001271$$0.00125$

Halving the step shrinks the error by $0.005171/0.001271 \approx 4.07$. Almost exactly 4.

Two variables. Back to $f(x,y)=x^2y$ at $(1,2)$ with $\boldsymbol\delta=(0.1,-0.2)$. The Hessian (the matrix of second derivatives) is $H = \begin{bmatrix} 2y & 2x\\ 2x & 0\end{bmatrix} = \begin{bmatrix} 4 & 2\\ 2 & 0\end{bmatrix}$.

  1. $H\boldsymbol\delta = [\,4\cdot0.1 + 2\cdot(-0.2),\;\; 2\cdot0.1 + 0\,]^\top = [\,0,\; 0.2\,]^\top$.
  2. $\boldsymbol\delta^\top H\boldsymbol\delta = 0.1\cdot 0 + (-0.2)\cdot0.2 = -0.04$.
  3. Half of it: $\tfrac12\boldsymbol\delta^\top H\boldsymbol\delta = -0.02$.

The true error (earlier) was $2.178 - 2.2 = -0.022$. The guess $-0.02$ is very close.

Keep one more term of the Taylor series (Chapter 2.11):

$$f(\mathbf{x}+\boldsymbol\delta) = f(\mathbf{x}) + \nabla f^\top\boldsymbol\delta + \underbrace{\tfrac12\,\boldsymbol\delta^\top H\,\boldsymbol\delta}_{\text{the error of the linearization}} + (\text{terms of size }\|\boldsymbol\delta\|^3).$$

So error $\approx \tfrac12\boldsymbol\delta^\top H\boldsymbol\delta$, where $H$ is the Hessian (see Chapter 2.10, and quadratic forms for the shape $\boldsymbol\delta^\top H\boldsymbol\delta$). In one variable it reads $\tfrac12f''(x)\,\delta^2$.

Where does it come from? Walk the path $g(t) = f(\mathbf{x}+t\boldsymbol\delta)$ again. A one-variable Taylor series gives $g(1) = g(0) + g'(0) + \tfrac12 g''(0) + \dots$. We saw $g'(0) = \nabla f^\top\boldsymbol\delta$. Applying the chain rule once more gives $g''(0) = \boldsymbol\delta^\top H\boldsymbol\delta$. Substituting gives the formula above.

A safe bound. If the curvature never exceeds $M$ on the way (in one variable $|f''|\le M$; in many, the Hessian's eigenvalues have size at most $M$), then $\;|\text{error}| \le \tfrac12 M\,\|\boldsymbol\delta\|^2$.

Two consequences: the error is second order (it scales like $\|\boldsymbol\delta\|^2$), and the sign of $\boldsymbol\delta^\top H\boldsymbol\delta$ tells you whether the true function sits above the tangent plane (bowl-like, positive) or below it (hill-like, negative).

Why do we need it?

An approximation without an error estimate is a guess. This formula tells us how wrong the line is, how small a step must be, and whether the true value lies above or below the line.

Where is it used?

Choosing learning rates and trust-region sizes in optimisation, proving that gradient descent converges, deciding step sizes in numerical solvers, and checking a gradient with a finite difference (that check relies on the error shrinking like $\delta^2$ or $\delta$).

How is it used?

To keep the error below a tolerance $\varepsilon$, pick a step with $\tfrac12 M\|\boldsymbol\delta\|^2 \le \varepsilon$. Or test your gradient code: compute the gap $f(\mathbf{x}+\boldsymbol\delta) - f(\mathbf{x}) - \nabla f^\top\boldsymbol\delta$ and shrink the step by 10. A correct gradient makes the gap drop by about 100 (second order). If it only drops by 10, the gradient is wrong.

The blue line shows the true error of the tangent line for steps from $10^{-3}$ to $1$. On this log-log plot a straight line with slope 2 means "error $\propto h^2$". The orange dashed line is the prediction $\tfrac12f''h^2$. (1) For eˣ the blue line is parallel to the orange one. (2) Choose |x| (a kink at 0): the slope drops to 1, so the error shrinks only as fast as the step does. (3) Choose sin 10x: same slope 2, but the line sits much higher (a bigger curvature). Move the step slider to read off the numbers.

The blue surface is the true error $f(\mathbf{x}+\boldsymbol\delta) - f(\mathbf{x}) - \nabla f^\top\boldsymbol\delta$ as a function of the step $\boldsymbol\delta$ (the horizontal position). The faint orange mesh is the estimate $\tfrac12\boldsymbol\delta^\top H\boldsymbol\delta$. Near the middle the two coincide. Drag the green dot around, rotate the view, then press Halve δ and see the error fall to a quarter. Switch to saddle: the error is positive in some directions and negative in others.

"Second order" is about the step, not about the function. A function with big curvature still has error $\propto\delta^2$. The constant out front is just bigger. That is why the blue lines in the plot all run parallel at slope 2 but sit at different heights.

Quick check: a linearization has error $0.04$ for a step of $0.2$. Roughly what error do you expect for a step of $0.05$?

The step shrank by a factor of $4$, so the error shrinks by $4^2 = 16$: $0.04/16 = 0.0025$.

When linearization fails: kinks and high curvature

"Zoom in and it gets straight" is true for smooth curves. Two kinds of curve break the promise.

  • A kink (corner). Zoom in on the tip of the letter V. The tip is still a tip, however far you zoom. There is no single straight line it turns into, because the left side and the right side have different slopes.
  • Very high curvature. A tight bend does flatten if you zoom, but only after you have zoomed in enormously. So the trust zone is tiny and the "small step" you need may be unrealistically small.

A kink. $f(x)=|x|$ at $a=0$. To the right the slope is $+1$, to the left it is $-1$. Try the best flat line, $L(x)=0$. The error at step $h$ is $|h|$, the same size as the step. The ratio error$/h$ stays at $1$ and never tends to $0$. So $|x|$ is not differentiable at $0$, and linearization has nothing good to offer there.

High curvature. From the error formula, the error is about $\tfrac12|f''|h^2$. To keep it below a tolerance $\varepsilon$ we need $h \le \sqrt{2\varepsilon/|f''|}$.

  1. For $e^x$ at $0$ we have $|f''|=1$. With $\varepsilon = 0.01$: $h\le\sqrt{0.02}\approx 0.14$. (The true error at $h=0.141$ is $0.0105$, just slightly over: we ignored the third-order terms.)
  2. For $\sin(10x)$ at its peak, $|f''| = 100$. Now $h\le\sqrt{0.02/100}\approx 0.014$.

A curvature 100 times bigger gives a trust zone 10 times smaller.

Linearization at $a$ is trustworthy when:

  1. $f$ is differentiable at $a$ (no kink, no jump, no vertical tangent), and
  2. the step satisfies $\tfrac12\,|f''|\,h^2 \le \varepsilon$, i.e. $\;|h| \le \sqrt{2\varepsilon/|f''|}$, where $|f''|$ is the curvature near $a$ (in many variables, the largest size of an eigenvalue of the Hessian).

Condition 1 is a yes/no question. Condition 2 is a how small is small enough question, answered by the curvature.

Why do we need it?

Every method that relies on a line or a plane silently assumes smoothness and a small step. Knowing the two ways it breaks tells you why training sometimes blows up, and what to change.

Where is it used?

Explaining why gradient descent diverges when the learning rate is too large (sharp curvature), why ReLU networks have undefined gradients exactly at a kink (frameworks pick 0), why gradient clipping and warm-up help, and why optimisers use trust regions.

How is it used?

If a step makes things worse than the line predicted, shrink the step (the learning rate). If a function has kinks, use a rule at the kink (a subgradient, such as "ReLU′(0)=0"), or a smooth stand-in.

Pick |x| with the point at 0 and slide the zoom all the way up: the percentage never falls, the corner never straightens. Now move point a to 0.5: the corner is out of view and the line is perfect. Then try sin 10x and compare with sin x at the same zoom: the curvier one needs far more zoom.

Kinks are everywhere in modern networks (ReLU, max-pooling, absolute-value losses). In practice you almost never land exactly on a kink, and away from it the function is perfectly linear. The danger is only at (or very near) the corner.

Quick check: a loss has curvature $|f''|\approx 400$. For tolerance $\varepsilon = 0.02$, how big a step can you trust?

$h \le \sqrt{2\cdot0.02/400} = \sqrt{0.0001} = 0.01$.

Sensitivity and error propagation core

Every measurement is a little wrong. You measure a table as 2.00 m but it might really be 2.01 m. If you then compute something from your measurements (an area, a speed, a model's prediction), how wrong is the answer?

Linearization answers instantly. The output wobbles by the slope times the input wobble. A steep slope means the output is very sensitive to that input: a tiny error in it makes a big error in the result. A flat slope means the input hardly matters.

With several inputs, each one adds its own wobble. Whichever input has the biggest "slope times wobble" is the one worth measuring more carefully.

One input. The area of a circle is $A=\pi r^2$. We measure $r = 10$ with an error of $\Delta r = 0.1$.

  1. Slope: $A'(r) = 2\pi r = 20\pi \approx 62.83$.
  2. Error in the area: $\Delta A \approx A'(r)\,\Delta r = 62.83\cdot 0.1 = 6.283$.
  3. Check: $\pi(10.1^2 - 10^2) = \pi\cdot 2.01 = 6.315$. The linear estimate is within 0.5 percent.
  4. Relative error: $\Delta A/A = 6.283/314.16 = 2\%$, while $\Delta r/r = 1\%$. The radius is squared, so its relative error is doubled.

Two inputs. A rectangle has width $w=5\pm0.1$ and height $h=3\pm0.2$. Its area is $A = wh = 15$.

  1. Slopes: $\partial A/\partial w = h = 3$ and $\partial A/\partial h = w = 5$.
  2. Contributions: $3\cdot0.1 = 0.3$ from the width, and $5\cdot 0.2 = 1.0$ from the height.
  3. Worst case (both errors push the same way): $0.3 + 1.0 = 1.3$. So $A = 15\pm1.3$.
  4. Independent random errors: they sometimes cancel, so combine them like the sides of a right triangle: $\sqrt{0.3^2 + 1.0^2} = \sqrt{1.09} \approx 1.04$.

The height error dominates, so measuring the height more carefully is what would help.

From $f(\mathbf{x}+\boldsymbol\delta)\approx f(\mathbf{x}) + \nabla f^\top\boldsymbol\delta$ we get, with input errors $\Delta x_i$:

$$\Delta f \;\approx\; \sum_i \frac{\partial f}{\partial x_i}\,\Delta x_i \qquad\text{(one input: } \Delta y \approx f'(x)\,\Delta x\text{)}.$$
  • Worst case (each error is at most $|\Delta x_i|$): $\;|\Delta f| \le \sum_i \left|\dfrac{\partial f}{\partial x_i}\right|\,|\Delta x_i|$. This is the "sensitivity" of $f$ to each input, times how wrong that input can be.
  • Independent random errors with typical sizes $\sigma_i$: $\;\sigma_f \approx \sqrt{\sum_i \left(\dfrac{\partial f}{\partial x_i}\sigma_i\right)^2}$. (Independent random wobbles add as squares, because they partly cancel.)

Both formulas only hold while the errors are small, so that the function is nearly linear over them. That is the linearization assumption, again.

Why do we need it?

To put honest error bars on anything we compute from imperfect numbers, and to find which input matters most. Without it we would have to recompute the answer for many possible input values.

Where is it used?

Lab measurements and engineering tolerances, saliency maps (which input pixels move the output most), adversarial examples such as FGSM (the loss changes by about $\nabla L^\top\boldsymbol\delta$, which is largest for $\boldsymbol\delta=\varepsilon\,\mathrm{sign}(\nabla L)$, giving $\varepsilon\|\nabla L\|_1$), and the condition number of a computation.

How is it used?

Compute the partial derivative for each input, multiply by that input's possible error, and add the pieces (straight sum for the worst case, square-root-of-squares for random errors). The biggest piece tells you what to fix first.

The blue box is the uncertainty around your measurement point (drag the dot to move it; the sliders set the sizes $\Delta x$ and $\Delta y$). The grey curves are level lines of the result $f$. The green arrow is the gradient: the output changes most along it. (1) Make $\Delta y$ large and $\Delta x$ small: the bar for $y$ dominates. (2) Compare the first-order worst case with the exact worst corner of the box: they almost agree. (3) Compare the straight sum with the random-error combination.

Errors only add up like this when they are small. For a huge error the function bends, the linear estimate is off, and (as the previous sections showed) the error grows like the square of the input error.

Quick check: $v = d/t$ with $d=100$ m (error 1 m) and $t=20$ s (error 0.2 s). Which error hurts more?

$\partial v/\partial d = 1/t = 0.05$, so the distance contributes $0.05\cdot1 = 0.05$. $\partial v/\partial t = -d/t^2 = -0.25$, so the time contributes $0.25\cdot0.2 = 0.05$. They contribute equally: the worst case is $0.05+0.05 = 0.1$ m/s on a speed of $5$ m/s.

Newton's method: repeated linearization

Suppose you want a number $x$ where $f(x)=0$ (a root), but $f$ is curvy and you cannot solve it by hand. Here is a neat trick.

  1. Stand at a guess $x_0$. Replace the curve by its tangent line. (Linearize.)
  2. A line is easy: find where the line crosses zero. Jump there. That is your next guess.
  3. Repeat from the new spot.

Each jump lands much closer to the true root, because the line fits the curve better and better as you get nearer. After a few jumps the number of correct digits doubles each time.

Find $\sqrt2$ by solving $f(x) = x^2 - 2 = 0$. The slope is $f'(x) = 2x$. Start at $x_0=1$.

  1. $f(1) = -1$, $f'(1)=2$. The line crosses zero at $x_1 = 1 - \dfrac{-1}{2} = 1.5$.
  2. $f(1.5) = 0.25$, $f'(1.5)=3$. So $x_2 = 1.5 - \dfrac{0.25}{3} = 1.41667$.
  3. $f(1.41667) = 0.006944$, $f' = 2.83333$. So $x_3 = 1.41667 - 0.002451 = 1.414216$.
  4. One more step gives $x_4 = 1.414213562375$. The true $\sqrt2 = 1.414213562373\ldots$

The errors are $0.414,\; 0.086,\; 0.0025,\; 0.0000021,\; 0.0000000000016$. The number of correct digits roughly doubles each step.

Newton's method. Linearize $f$ at the current guess $x_n$: $L(x) = f(x_n) + f'(x_n)(x - x_n)$. The new guess is where the line is zero:

$$0 = f(x_n) + f'(x_n)(x_{n+1} - x_n) \;\;\Longrightarrow\;\; x_{n+1} = x_n - \frac{f(x_n)}{f'(x_n)}.$$

(We just solved a line equation for $x_{n+1}$: move $f(x_n)$ across, then divide by $f'(x_n)$.)

  • Near a root where $f'\neq0$, the error is roughly squared each step: $\varepsilon_{n+1}\approx \dfrac{f''}{2f'}\,\varepsilon_n^2$. This is "quadratic convergence". It comes straight from the $\tfrac12 f''h^2$ error of linearization.
  • It can fail: if $f'(x_n)=0$ the line is flat and never crosses zero, and a bad start can bounce between points or fly away.
  • For optimisation, minimise $f$ by solving $f'(x)=0$. Newton then reads $x_{n+1} = x_n - f'(x_n)/f''(x_n)$, and in many variables $\mathbf{x}\leftarrow\mathbf{x} - H^{-1}\nabla f$ with the Hessian $H$ (Chapter 2.10).
Why do we need it?

Many equations cannot be solved with algebra. Newton's method turns a hard equation into a short list of easy line problems, and it is very fast once you are close.

Where is it used?

Computing square roots inside your calculator or library, solving equations in physics and finance, logistic regression (the IRLS algorithm is Newton's method), second-order optimisers, and L-BFGS (a cheaper cousin).

How is it used?

Pick a starting guess, compute $f$ and $f'$, update $x\leftarrow x - f/f'$, and stop when $|f(x)|$ is tiny. Always check the starting point: if $f'$ is near 0, choose a different start.

Press Step repeatedly. Each orange segment is a tangent line, and the green dot on the axis is the new guess. Watch the table: the correct digits column roughly doubles each step. Drag the starting point to try other starts. Then pick the third function and start at 0: Newton bounces between 0 and 1 forever and never finds the root. Linearization can mislead you when you are far from the answer.

Newton's method is a local method. It trusts the tangent line, so it needs a start that is reasonably near a root, and a function that is not too flat there. Far from the root, the line can send you somewhere worse.

Quick check: do one Newton step for $f(x)=x^2-9$ starting at $x_0=5$.

$f(5)=16$ and $f'(5)=10$, so $x_1 = 5 - 16/10 = 3.4$. The root is $3$, so we already moved from 2 away to $0.4$ away.

Linearizing a sigmoid / neuron

A single neuron computes a weighted sum $z = w_1x_1 + w_2x_2 + b$ and squashes it with the sigmoid $\sigma(z) = 1/(1+e^{-z})$. The sum is already linear. All the bending comes from the S-shaped squash.

Look at the S-curve at the middle: it is almost a straight tilted line. So for small changes around a point, the neuron is a tilted flat sheet: change the inputs a little, and the output changes by (steepness of the S) times (the change in $z$). Out in the flat tails, the S has no steepness, so the neuron barely reacts to anything. That is the saturated neuron.

The sigmoid has the neat slope $\sigma'(z) = \sigma(z)\,(1-\sigma(z))$. At $z=0$: $\sigma(0)=0.5$ and $\sigma'(0) = 0.5\cdot0.5 = 0.25$. So near zero,

$$\sigma(z) \approx 0.5 + 0.25\,z.$$
  1. At $z=0.4$: line gives $0.5 + 0.25\cdot0.4 = 0.6$. True: $\sigma(0.4) = 0.59869$. Off by $0.0013$.
  2. At $z = 2$: line gives $0.5+0.5 = 1.0$. True: $\sigma(2)=0.8808$. Off by $0.12$. The line is already poor, because the S bends over.

At $z=2$ the slope is $0.8808\cdot0.1192 = 0.105$, less than half of $0.25$. At $z=5$ it is only $0.0066$.

For the neuron $y = \sigma(\mathbf{w}^\top\mathbf{x} + b)$, let $z_0=\mathbf{w}^\top\mathbf{x}_0 + b$ and $y_0=\sigma(z_0)$. The linearization around the input $\mathbf{x}_0$ is

$$y(\mathbf{x}_0+\boldsymbol\delta) \approx y_0 + \sigma'(z_0)\;\mathbf{w}^\top\boldsymbol\delta, \qquad \nabla_{\mathbf{x}}\,y = \sigma'(z_0)\,\mathbf{w} = y_0(1-y_0)\,\mathbf{w}.$$

Derivation. Chain rule (Chapter 2.8): the inner function $z=\mathbf{w}^\top\mathbf{x}+b$ has gradient $\mathbf{w}$, the outer function $\sigma$ has slope $\sigma'(z_0)$. Multiply them: $\sigma'(z_0)\,\mathbf{w}$. Then the first-order formula gives the line above.

The level lines (where $y$ is constant) of the linearized neuron are straight, parallel lines perpendicular to $\mathbf{w}$. The real neuron's level lines are also straight lines (since $y$ depends only on $z$), but they are unevenly spaced: crowded in the middle, far apart in the tails.

Why do we need it?

It shows exactly how sensitive a neuron is to its inputs and weights: the slope $\sigma'$. It explains saturation: when $|z|$ is large the slope is nearly zero, so the gradient that flows backwards almost vanishes.

Where is it used?

The vanishing-gradient problem in deep sigmoid and tanh networks, weight-initialisation rules (keep $z$ near the steep middle), logistic regression, and analysing how a small input change moves a prediction.

How is it used?

Compute $z_0$ and $y_0$ in the forward pass. The local slope is $y_0(1-y_0)$ (free, no extra work). The backward pass multiplies by it. If it is close to 0, expect almost no learning signal through that neuron.

Left: the neuron's output over the input plane (darker blue = closer to 1). The grey lines are its true level lines. The orange lines are the level lines of the linearized neuron at the blue dot x₀: straight and evenly spaced. Drag the green dot x₁ close to x₀ and the error vanishes. Drag it far and the orange lines no longer match. Right: the same thing along the single number $z$. Slide b or drag x₀ out into a flat tail and watch the slope $\sigma'$ collapse.

The sigmoid's biggest slope is only $0.25$ (at $z=0$). So even a perfectly placed sigmoid multiplies the backward signal by at most $0.25$ per layer (this is the sigmoid's own factor; the weights multiply it too). Stack ten such layers and the sigmoid factors alone shrink the gradient to at most $0.25^{10}\approx 10^{-6}$ of its starting size. This is why deep sigmoid networks were hard to train, and why ReLU took over.

Quick check: at $z=1$, $\sigma(1)=0.7311$. What is the slope $\sigma'(1)$, and what does the line predict for $\sigma(1.2)$?

$\sigma'(1) = 0.7311\cdot(1-0.7311) = 0.7311\cdot0.2689 = 0.1966$. Line: $0.7311 + 0.1966\cdot0.2 = 0.7704$. The true $\sigma(1.2)=0.7685$.

Why small learning rates: gradient descent is a chain of linear approximations core

Stand on a hilly landscape in thick fog. You can only feel the slope under your feet. So you picture the ground as a flat tilted sheet (that is the linearization!) and walk a little way downhill on that sheet. Then you feel the slope again, draw a new sheet, and repeat.

There is a catch. A tilted flat sheet has no bottom: if you believed it forever you would walk downhill for ever. The real ground curves back up. So the sheet is only trustworthy for a short step. The learning rate $\eta$ is exactly "how far do I trust the sheet?". Too big, and you walk off the end of the sheet and may even end up higher than you started.

Let the loss be $L(x,y) = x^2 + 4y^2$ and stand at $\boldsymbol\theta=(2,1)$, where $L = 4+4=8$.

  1. Gradient: $\mathbf{g} = [2x,\,8y]^\top = [4,\,8]^\top$, so $\|\mathbf{g}\|^2 = 16+64 = 80$.
  2. Hessian: $H=\mathrm{diag}(2, 8)$, so $\mathbf{g}^\top H\mathbf{g} = 2\cdot16 + 8\cdot64 = 32+512=544$.
  3. Step with $\eta = 0.1$: the line predicts a drop of $\eta\|\mathbf{g}\|^2 = 8$ (all the way to 0!).
  4. Reality: the new point is $(2-0.4,\,1-0.8) = (1.6,\,0.2)$, with $L = 2.56 + 0.16 = 2.72$. The true drop is $8 - 2.72 = 5.28$. The curvature "ate" $\tfrac12\eta^2\,\mathbf{g}^\top H\mathbf{g} = 0.5\cdot0.01\cdot544 = 2.72$ of the predicted drop: $8 - 2.72 = 5.28$ ✓.
  5. Step with $\eta=0.3$: the new point is $(0.8,\,-1.4)$ and $L = 0.64 + 4\cdot1.96 = 8.48$. The loss went up, even though we stepped in the downhill direction.

One gradient-descent step is $\boldsymbol\theta_{\text{new}} = \boldsymbol\theta - \eta\,\mathbf{g}$ with $\mathbf{g}=\nabla L(\boldsymbol\theta)$. Put $\boldsymbol\delta = -\eta\mathbf{g}$ into the linearization and the second-order error term:

$$L(\boldsymbol\theta - \eta\mathbf{g}) \;\approx\; \underbrace{L(\boldsymbol\theta)}_{\text{now}} \;-\; \underbrace{\eta\,\|\mathbf{g}\|^2}_{\text{what the line promises}} \;+\; \underbrace{\tfrac12\,\eta^2\,\mathbf{g}^\top H\mathbf{g}}_{\text{curvature spoils it}}.$$

(The middle term: $\nabla L^\top\boldsymbol\delta = \mathbf{g}^\top(-\eta\mathbf{g}) = -\eta\|\mathbf{g}\|^2$. The last term: $\tfrac12\boldsymbol\delta^\top H\boldsymbol\delta$ with $\boldsymbol\delta=-\eta\mathbf{g}$.) The loss goes down exactly when the promise beats the spoiler:

$$\eta\,\|\mathbf{g}\|^2 > \tfrac12\eta^2\,\mathbf{g}^\top H\mathbf{g} \;\;\Longleftrightarrow\;\; \eta < \frac{2\,\|\mathbf{g}\|^2}{\mathbf{g}^\top H\mathbf{g}}.$$

In the example the limit is $2\cdot80/544 \approx 0.294$, which is why $\eta=0.3$ failed. Since $\mathbf{g}^\top H\mathbf{g}\le\lambda_{\max}\|\mathbf{g}\|^2$ (with $\lambda_{\max}$ the largest Hessian eigenvalue), any $\eta<2/\lambda_{\max}$ is safe. Sharper curvature forces a smaller learning rate. See the Hessian and convexity.

Why do we need it?

It explains what the learning rate really is, and why it must be small: it is the distance we trust a straight-line model of the loss. It also explains why training sometimes diverges.

Where is it used?

Choosing learning rates for SGD and Adam, learning-rate warm-up and decay, line search, trust-region methods, and the analysis showing gradient descent converges for $\eta<2/L$ (where $L$ bounds the curvature).

How is it used?

If the loss goes up or oscillates, cut the learning rate. A common rule of thumb: stay below $1/\lambda_{\max}$ of the Hessian. If you cannot compute the Hessian, lower $\eta$ by factors of 3 or 10 until the loss falls smoothly.

Left: the loss contours and your path. Drag the blue dot to choose the start. Press Step or Run 15 steps. Right: a slice of the loss along the downhill step. The orange line is the linear model ("the loss keeps falling forever"), the blue curve is the truth. The purple dashed line marks your learning rate $\eta$. (1) With a small $\eta$ the blue and orange curves almost agree and the loss falls. (2) Slide $\eta$ right until the marker passes where the blue curve turns upwards: the step now overshoots. (3) On the stretched bowl, push $\eta$ above 0.25 and Run: the path zig-zags more and more.

The safe-step formula uses the curvature at the current point. Curvature can change as you move (see the flat-bottom bowl), so a learning rate that is safe now can be too big or too small later. That is why schedules and adaptive methods exist.

Quick check: for $L=\tfrac12\lambda\theta^2$ (one variable), what is the largest safe learning rate?

Here $g=\lambda\theta$ and $H=\lambda$, so $\|g\|^2=\lambda^2\theta^2$ and $g^\top Hg=\lambda^3\theta^2$. The limit is $2\lambda^2\theta^2/(\lambda^3\theta^2) = 2/\lambda$. You can also see it directly: each step multiplies $\theta$ by $(1-\eta\lambda)$, which shrinks $|\theta|$ exactly when $0<\eta<2/\lambda$.

Why neural networks look piecewise linear

A ReLU neuron outputs $\max(0, z)$: it is either off (outputs 0) or on (passes $z$ straight through). Inside either state it is perfectly linear. So a whole network built from ReLUs is a patchwork: it is a different straight-line (or flat-sheet) function in each region, with bends where some neuron switches on or off.

Now the link to this chapter. Stand at a point and look at the activation pattern (which neurons are on). Inside that patch the network is a linear map, so its linearization is exact there, and the Jacobian is just a product of the weight matrices with the "off" neurons switched out. Even for smooth networks (sigmoid, tanh) the picture is the same, only with the corners rounded.

A tiny one-input network with three ReLU neurons: $y = 1\cdot\mathrm{ReLU}(x+1.5) - 2\cdot\mathrm{ReLU}(x-0.5) + 3\cdot\mathrm{ReLU}(x-2)$. The neurons switch on at $x=-1.5,\;0.5,\;2$. The slope is the sum of $v_i$ over the neurons that are on:

regionneurons onslope
$x < -1.5$none$0$
$-1.5 < x < 0.5$first$1$
$0.5 < x < 2$first, second$1 - 2 = -1$
$x > 2$all three$1 - 2 + 3 = 2$

Check some values: $y(0.5) = 2$, $y(2) = 3.5 - 2\cdot1.5 = 0.5$, $y(3) = 4.5 - 5 + 3 = 2.5$. The graph is four straight pieces joined at the three kinks.

For a ReLU network $F(\mathbf{x}) = W_2\,\mathrm{ReLU}(W_1\mathbf{x} + \mathbf{b}_1) + \mathbf{b}_2$, let $D(\mathbf{x})$ be the diagonal matrix with a $1$ for each neuron that is on at $\mathbf{x}$ and a $0$ for each that is off. In the region where $D$ does not change,

$$F(\mathbf{x}) = W_2\,D\,W_1\,\mathbf{x} + \text{(constant)}, \qquad J = W_2\,D\,W_1.$$

This is exactly linear (affine) there, so linearization has zero error inside the region, and the error appears only when you cross a kink. For smooth activations, $D$ is replaced by the diagonal matrix of slopes $\mathrm{diag}(\sigma'(z_i))$, and $J = W_2\,\mathrm{diag}(\sigma'(z))\,W_1$ is still the exact Jacobian at that point (it is the chain rule of Chapter 2.8, a product of Jacobians), but the straight-line model it gives is only good close to the point. (ReLU's slope at exactly 0 is not defined; software uses 0 by convention.)

Why do we need it?

It tells us what a ReLU network is at any point: one linear map. That makes its gradients and Jacobians easy to understand, and explains why the gradient is piecewise constant.

Where is it used?

Counting the linear regions of a network (a measure of its expressiveness), explaining backprop in ReLU networks (gates pass or block the gradient), adversarial robustness analysis, and verifying networks by checking each linear region.

How is it used?

To understand a prediction at an input, look at which neurons are on, multiply the on-paths of weights, and read off the local linear model. To verify or test a gradient, remember it only changes at the kinks.

Drag the blue dot along the curve. In ReLU mode the green band is the region where the same neurons are on: inside it the orange tangent line is the network. The purple rings mark the kinks. Move the three output weights and watch the slopes of the pieces change (the readout adds up the weights of the neurons that are on). Then switch to smooth (softplus): the corners round off, and the tangent line only hugs the curve locally.

"Piecewise linear" does not mean "simple". A modest network can have an enormous number of pieces, so the whole function can be extremely wiggly. But up close, at any one input, it is just a matrix.

Quick check: in the example network, what is the slope at $x=1$, and what is $y(1)$?

At $x=1$ the first two neurons are on (since $1>-1.5$ and $1>0.5$) and the third is off ($1<2$). Slope $=1-2=-1$. Value: $y(1) = 1\cdot(2.5) - 2\cdot(0.5) + 0 = 1.5$.

Recap, cheat sheet and practice

  • Zoom in on a smooth function and it looks linear: $f(x+\delta)\approx f(x) + f'(x)\,\delta$. For many inputs it is $f(\mathbf{x}+\boldsymbol\delta)\approx f(\mathbf{x}) + \nabla f^\top\boldsymbol\delta$, and for vector outputs $F(\mathbf{x}+\boldsymbol\delta)\approx F(\mathbf{x}) + J\boldsymbol\delta$.
  • This is a local statement. The trust zone depends on the tolerance and on the curvature.
  • The Jacobian is the matrix a curvy map turns into near a point: its columns say where nudges of each input go, and $|\det J|$ is the local area stretch.
  • The error is about $\tfrac12\boldsymbol\delta^\top H\boldsymbol\delta$: it shrinks like $\delta^2$ (halve the step, quarter the error). Linearization fails at kinks and needs tiny steps where the curvature is high: $|h|\lesssim\sqrt{2\varepsilon/|f''|}$.
  • Uses: error propagation $\Delta f\approx\sum \frac{\partial f}{\partial x_i}\Delta x_i$; Newton's method $x\leftarrow x-f/f'$ (repeated linearization, digits double each step); the sigmoid slope $\sigma(1-\sigma)\le0.25$; gradient descent trusts a plane for a step of length $\eta\|\nabla L\|$ and is safe when $\eta<2/\lambda_{\max}$; ReLU networks are piecewise linear, with $J=W_2DW_1$ inside each region.

Cheat sheet

IdeaFormulaPicture
Linearization (1D)$L(x) = f(a) + f'(a)(x-a)$tangent line
First-order, scalar loss$f(\mathbf{x}+\boldsymbol\delta)\approx f + \nabla f^\top\boldsymbol\delta$tangent plane
First-order, vector map$F(\mathbf{x}+\boldsymbol\delta)\approx F + J\boldsymbol\delta$grid becomes parallelograms
Error$\approx \tfrac12\boldsymbol\delta^\top H\boldsymbol\delta$, bounded by $\tfrac12 M\|\boldsymbol\delta\|^2$gap between curve and line
Trust radius$|h|\le\sqrt{2\varepsilon/|f''|}$width of the green band
Error propagationworst case $\sum|\partial_i f||\Delta x_i|$; random $\sqrt{\sum(\partial_i f\,\sigma_i)^2}$box over level lines
Newton step$x_{n+1} = x_n - f(x_n)/f'(x_n)$tangent hits zero
Sigmoid$\sigma'=\sigma(1-\sigma)$, $\sigma\approx0.5+0.25z$ near 0S-curve slope
Gradient-descent drop$\eta\|\mathbf{g}\|^2 - \tfrac12\eta^2\mathbf{g}^\top H\mathbf{g}$; safe if $\eta<2\|\mathbf{g}\|^2/\mathbf{g}^\top H\mathbf{g}$line vs bowl
ReLU network$J = W_2\,D\,W_1$ in each regionstraight pieces
Code it · NumPy

import numpy as np

# 1. Linearization of exp at a = 0, and how the error shrinks with the step
f = np.exp
a = 0.0
for h in [0.1, 0.05, 0.025]:
    err = f(a + h) - (f(a) + f(a) * h)       # f'(a) = e^a = 1
    print(h, err, 0.5 * h**2)                # error vs the guess (1/2) f'' h^2
# 0.1   0.005170918...  0.005
# 0.05  0.001271096...  0.00125     (half the step: about 4 times smaller error)
# 0.025 0.000315120...  0.0003125

# 2. Jacobian as a local linear map: F(x, y) = (x^2 - y^2, 2xy) at (1, 1)
def F(v):
    x, y = v
    return np.array([x**2 - y**2, 2 * x * y])

def jacobian(F, v, eps=1e-6):                # finite-difference Jacobian (m x n)
    v = np.asarray(v, float)
    cols = [(F(v + eps * e) - F(v - eps * e)) / (2 * eps) for e in np.eye(len(v))]
    return np.stack(cols, axis=1)

p = np.array([1.0, 1.0])
J = jacobian(F, p)
d = np.array([0.1, 0.05])
print(J.round(4))                            # [[ 2. -2.]  [ 2.  2.]]
print(F(p) + J @ d)                          # linear estimate  [0.1  2.3]
print(F(p + d))                              # true value       [0.1075  2.31]
print(np.linalg.det(J))                      # about 8 (the local area stretch)

# 3. Newton's method for sqrt(2): the error is roughly squared each step
x = 1.0
for n in range(4):
    x = x - (x**2 - 2) / (2 * x)
    print(n + 1, x, abs(x - np.sqrt(2)))
# 1 1.5                0.0858
# 2 1.41666...         0.00245
# 3 1.4142156862...    2.1e-06
# 4 1.4142135623746... 1.6e-12

# 4. Error propagation for A = w * h  (w = 5 +- 0.1, h = 3 +- 0.2)
w, h, dw, dh = 5.0, 3.0, 0.1, 0.2
worst = h * dw + w * dh                      # sum of |slope| * |error|
rand = np.hypot(h * dw, w * dh)              # independent random errors
print(worst, rand)                           # 1.3   1.0440...

# 5. Gradient descent: line promises vs the truth, for L = x^2 + 4 y^2 at (2, 1)
L = lambda t: t[0]**2 + 4 * t[1]**2
t = np.array([2.0, 1.0]); g = np.array([2 * t[0], 8 * t[1]]); H = np.diag([2.0, 8.0])
for eta in [0.1, 0.3]:
    promise = eta * g @ g
    spoil = 0.5 * eta**2 * g @ H @ g
    print(eta, promise, promise - spoil, L(t) - L(t - eta * g))
# 0.1  8.0   5.28   5.28      (the loss falls by 5.28)
# 0.3  24.0 -0.48  -0.48      (negative drop: the loss went UP)
print(2 * (g @ g) / (g @ H @ g))             # largest safe eta: 0.2941...
Test yourself

1. The linearization of $f$ at $a$ is…

Start at the height $f(a)$, then add slope times step: $f'(a)(x-a)$. It is the tangent line at $a$.

2. A linearization has error $0.08$ for a step of $0.4$. What error do you expect for a step of $0.2$?

The error is second order: it scales with the square of the step. Half the step means one quarter of the error: $0.08/4 = 0.02$.

3. At which point does linearization fail no matter how far you zoom in?

$|x|$ has a corner at 0: the left slope is $-1$ and the right slope is $+1$, so there is no single tangent line. The other three are smooth.

4. One Newton step for $f(x)=x^2-25$ starting at $x_0=10$ gives $x_1=$…

$f(10)=75$ and $f'(10)=20$. So $x_1 = 10 - 75/20 = 10 - 3.75 = 6.25$. The root is 5, so we are much closer.

5. A loss near a point is $L=5\theta^2$ (so the curvature is $L''=10$). Gradient descent lowers the loss for every learning rate below a limit. What is that limit?

The limit is $2/\lambda$ with $\lambda = 10$: each step multiplies $\theta$ by $1-10\eta$, which has size below 1 exactly when $\eta<0.2$.

6. What is the Jacobian of $F(x,y) = (x+y,\; xy)$ at the point $(2,3)$?

Row 1 is the gradient of $x+y$: $[1,1]$. Row 2 is the gradient of $xy$: $[y,\,x] = [3,\,2]$. Rows are outputs, columns are inputs.

Practice problems

A. Linearize $f(x)=\sqrt{x}$ at $a=4$ and use it to estimate $\sqrt{4.2}$. How big is the error?

$f(4)=2$ and $f'(x)=\dfrac{1}{2\sqrt x}$, so $f'(4)=\dfrac14$. Then $L(x) = 2 + \tfrac14(x-4)$. At $x=4.2$: $L = 2 + 0.25\cdot0.2 = 2.05$. The true value is $2.04939\ldots$, so the error is about $0.0006$. (Check with $\tfrac12|f''|h^2$: $f''(x) = -\tfrac14x^{-3/2}$, so $f''(4) = -\tfrac14\cdot\tfrac18 = -\tfrac1{32}$, giving $\tfrac12\cdot\tfrac1{32}\cdot0.04 = 0.000625$ ✓.)

B. For $f(x,y)=x^2+3xy$ at $(1,2)$ with $\boldsymbol\delta=(0.05,\,-0.1)$: find the linear estimate, the true value, and $\tfrac12\boldsymbol\delta^\top H\boldsymbol\delta$.

$f(1,2) = 1 + 6 = 7$. $\nabla f = [2x+3y,\;3x]^\top = [8,\,3]^\top$. So $\nabla f^\top\boldsymbol\delta = 8\cdot0.05 + 3\cdot(-0.1) = 0.4 - 0.3 = 0.1$, and the linear estimate is $7.1$. The truth: $1.05^2 + 3\cdot1.05\cdot1.9 = 1.1025 + 5.985 = 7.0875$, so the error is $-0.0125$.

The Hessian is $H = \begin{bmatrix}2&3\\3&0\end{bmatrix}$. $H\boldsymbol\delta = [0.1 - 0.3,\;0.15]^\top = [-0.2,\,0.15]^\top$, and $\boldsymbol\delta^\top H\boldsymbol\delta = 0.05\cdot(-0.2) + (-0.1)(0.15) = -0.01 - 0.015 = -0.025$. Half is $-0.0125$: exactly the error. (For a quadratic function the third-order terms are zero, so the formula is exact.)

C. A density is $\rho = m/V$ with $m = 50\pm1$ g and $V = 20\pm0.5$ cm³. Give $\rho$, the worst-case error, and the random-error estimate.

$\rho = 50/20 = 2.5$. Slopes: $\partial\rho/\partial m = 1/V = 0.05$ and $\partial\rho/\partial V = -m/V^2 = -50/400 = -0.125$. Contributions: $0.05\cdot1 = 0.05$ and $0.125\cdot0.5 = 0.0625$. Worst case: $0.05 + 0.0625 = 0.1125$ (so $\rho = 2.5\pm0.11$, a 4.5 percent error: the relative errors $1/50 = 2\%$ and $0.5/20 = 2.5\%$ add). Random: $\sqrt{0.05^2 + 0.0625^2} = \sqrt{0.00640625}\approx 0.080$.

D. Do two Newton steps for $f(x) = x^3 - 2$ from $x_0=1$ (to approximate $\sqrt[3]{2}$).

$f'(x) = 3x^2$. Step 1: $f(1) = -1$, $f'(1) = 3$, so $x_1 = 1 + 1/3 = 1.3333$. Step 2: $f(1.3333) = 2.3704 - 2 = 0.3704$, $f'(1.3333) = 3\cdot1.7778 = 5.3333$, so $x_2 = 1.3333 - 0.0694 = 1.2639$. The true value is $1.2599$. Two steps already give 2 correct digits.

E. The polar-to-Cartesian map is $F(r,\theta) = (r\cos\theta,\; r\sin\theta)$. Find its Jacobian and determinant at $(r,\theta)=(2,\pi/2)$ and say what they mean.

$J = \begin{bmatrix}\cos\theta & -r\sin\theta\\ \sin\theta & r\cos\theta\end{bmatrix}$. At $(2,\pi/2)$: $\cos=0$, $\sin=1$, so $J = \begin{bmatrix}0 & -2\\ 1 & 0\end{bmatrix}$ and $\det J = 0\cdot 0 - (-2)(1) = 2$ (in general $\det J = r$). The first column says that a nudge of $r$ moves the point straight up (here $(0,1)$ per unit). The second column says a nudge of $\theta$ moves it left by $r=2$ per radian. The determinant $r$ says small patches in $(r,\theta)$ space are stretched by a factor of $r$: bigger circles have bigger patches.

F. For $L(\theta) = 3\theta^2$ at $\theta = 2$, find the safe range of learning rates, and compare $\eta=0.1$ with $\eta=0.5$.

$L'' = 6$, so the safe range is $0<\eta<2/6\approx 0.333$. The gradient is $g = 6\theta = 12$ and $L = 12$. With $\eta=0.1$: $\theta_{\text{new}} = 2 - 1.2 = 0.8$ and $L = 3\cdot0.64 = 1.92$: a big drop. With $\eta = 0.5$: $\theta_{\text{new}} = 2 - 6 = -4$ and $L = 3\cdot16 = 48$: the loss quadrupled. The step overshot past the minimum and landed higher up the other side of the bowl.

Chapter 2.13

Multivariable Calculus

Fields, flows, spin and flux. This chapter names the tools that describe "a quantity at every point of space": a temperature map, a wind map, the way a flow expands or swirls. It is a tour for awareness: you should know what each tool means and where it appears.

Lower priority for ML (awareness). For machine learning, the tools that matter most are the ones you have already met: gradients (2.4), Jacobians (2.5), Hessians (2.10) and Taylor expansions (2.11). Divergence, curl, line integrals and surface integrals come from physics. They appear in a few corners of ML (diffusion models, normalising flows, physics-informed networks) but you can train most models without ever computing one. So this chapter is deliberately short: one clear picture per idea, and the honest place each one shows up. Sections marked awareness are for recognition, not mastery.

  • Tell a scalar field (a number at every point) from a vector field (an arrow at every point)
  • Recognise a gradient field and test whether a field is conservative
  • Recap the directional derivative, the Jacobian and the Hessian as questions about fields
  • Read divergence as "net outflow" and curl as "spin"
  • Know what a line integral (work along a path) and a surface integral (flux) measure
  • Recognise the big theorems (fundamental theorem for line integrals, Green, Stokes, Divergence) and where the ideas show up in ML

Scalar fields: a number at every point awareness

Look at a weather map showing the temperature. At every spot on the map there is one number: 18 degrees here, 25 degrees there. That whole map is a scalar field. ("Scalar" just means "a single number".)

You can draw it with colours (hot = orange, cold = blue) or with contour lines joining points of equal value (isotherms, or "level curves"). You met both pictures in Chapter 2.4. A mountain map is the same idea: the number is the height.

Let $T(x,y) = 30 - x^2 - y^2$: a hot spot in the middle that cools as you move away.

  • $T(0,0) = 30$ (the hottest point).
  • $T(1,2) = 30 - 1 - 4 = 25$.
  • $T(3,0) = 30 - 9 = 21$. And also $T(0,-3) = 21$, $T(2.12, 2.12)\approx 21$: points the same distance from the middle share a value, so the level curves are circles.

A scalar field is a function $f:\mathbb{R}^n\to\mathbb{R}$ that gives one number $f(\mathbf{x})$ to every point $\mathbf{x}$ of a region. Its level set for the value $c$ is $\{\mathbf{x} : f(\mathbf{x}) = c\}$ (a curve in 2D, a surface in 3D).

You already know examples: every loss function $L(\boldsymbol\theta)$ is a scalar field over "weight space". Its gradient $\nabla f$ turns it into a vector field (next sections).

Why do we need it?

We need a name and a picture for "a quantity that depends on position". It is the setting where gradients, contours and optimisation live.

Where is it used?

Loss surfaces over model weights, probability density functions $p(\mathbf{x})$, image brightness as a function of pixel position, temperature, pressure and potential energy in physics, and energy-based models.

How is it used?

Sample it on a grid and draw a heat map or contours to see its shape (hills, valleys, ridges). Then take its gradient to find which way it climbs.

Drag the dot to read the field's value at any point. Switch the field. Turn the level curves on and off: each curve joins points with the same number. Notice that where the curves are packed tightly, the value changes quickly when you move.

Quick check: with $T(x,y)=30-x^2-y^2$, is $T(2,1)$ bigger or smaller than $T(0,2)$?

$T(2,1) = 30 - 4 - 1 = 25$ and $T(0,2) = 30 - 0 - 4 = 26$. So $T(0,2)$ is bigger: that point is slightly closer to the hot centre ($\sqrt{4}=2$ versus $\sqrt{5}\approx2.24$).

Vector fields: an arrow at every point awareness

Now picture a wind map. At every spot there is an arrow: which way the air is moving, and how fast (the arrow's length or darkness). That is a vector field: a vector attached to every point.

Drop a leaf into the wind and it travels along a path that is always tangent to the arrows. That path is a flow line (or streamline). Vector fields are how we describe flowing water, moving air, and force fields such as gravity.

The field $\mathbf{F}(x,y) = (-y,\;x)$ attaches to the point $(x,y)$ the arrow $(-y, x)$.

  • At $(1,0)$ the arrow is $(0,1)$: straight up.
  • At $(0,2)$ the arrow is $(-2,0)$: to the left, and twice as long.
  • At $(-1,0)$ the arrow is $(0,-1)$: straight down.

Going around the origin, the arrows turn counter-clockwise. A leaf released here would circle the origin. This is a pure rotation.

A vector field is a function $\mathbf{F}:\mathbb{R}^n\to\mathbb{R}^n$ that gives a vector $\mathbf{F}(\mathbf{x})$ to each point $\mathbf{x}$. In 2D: $\mathbf{F}(x,y) = (F_x(x,y),\,F_y(x,y))$, i.e. two scalar fields, one per component.

A flow line is a curve $\mathbf{x}(t)$ whose velocity is the field at every moment: $\dfrac{d\mathbf{x}}{dt} = \mathbf{F}(\mathbf{x})$. The Jacobian of a vector field is a square matrix that says how the arrows change from place to place.

Why do we need it?

Many things have a direction and a size at every location (wind, current, force, "which way is uphill"). A vector field describes them all at once, and its flow lines show where things end up.

Where is it used?

Fluid and weather simulation, electric and gravitational forces, the "score" $\nabla\log p(\mathbf{x})$ that diffusion models learn (a vector field pointing towards likely data), neural ODEs and continuous normalising flows (a learned field moves the data), and gradient descent (the field $-\nabla L$).

How is it used?

Draw arrows on a grid to see the pattern. To follow a particle, take small steps along the arrow at its current position and repeat. That is exactly how an ODE solver, or gradient descent, works.

Choose a field. The faint arrows show the field (darker = stronger). Drag the orange dot: the orange curve is the flow line starting there. Press Flow to watch a particle ride it. Turn on Many flow lines to see the whole pattern. Compare rotation (particles circle), source (they run away from the centre), sink (they fall into it) and saddle (in along one axis, out along the other).

Quick check: for $\mathbf{F}(x,y)=(x,\,y)$, what is the arrow at $(2,-1)$? What shape is the field?

$\mathbf{F}(2,-1) = (2,-1)$: the arrow equals the position, so it points straight away from the origin, and longer the further you are. This is a source: everything flows outward from the origin.

Gradient fields and conservative fields awareness

Take any scalar field, such as the height of a landscape. At every point, draw the arrow pointing straight uphill, with a length equal to the steepness. The result is a vector field: the gradient field $\nabla f$. It is the bridge between the two kinds of field. (See Chapter 2.4: the gradient is perpendicular to the level curves and points to steepest ascent.)

Not every vector field is the gradient of something. A rotation field circles around: if it were "uphill" arrows, you could walk with the arrows and climb for ever while coming back to where you started, which is impossible for a real height. A field that is a gradient is called conservative, and the scalar field it comes from is its potential.

Let $f(x,y) = x^2y$. Then $\nabla f = (2xy,\;x^2)$, so $\mathbf{F} = (2xy,\;x^2)$ is a gradient field.

A test. If $\mathbf{F} = (F_x,F_y)=\nabla f$, then $\partial F_x/\partial y = f_{xy}$ and $\partial F_y/\partial x = f_{yx}$, and mixed partials are equal. So a gradient field must pass the test $\;\dfrac{\partial F_x}{\partial y} = \dfrac{\partial F_y}{\partial x}$.

  • $\mathbf{F} = (2xy,\;x^2)$: $\partial F_x/\partial y = 2x$ and $\partial F_y/\partial x = 2x$. Equal, so it passes.
  • $\mathbf{F} = (-y,\;x)$: $\partial F_x/\partial y = -1$ but $\partial F_y/\partial x = +1$. Not equal, so this rotation field is not a gradient. No potential exists.

A vector field $\mathbf{F}$ is a gradient field (or conservative) if $\mathbf{F} = \nabla f$ for some scalar field $f$, the potential. Then:

  • $\mathbf{F}$ points perpendicular to the level curves of $f$ and has size $\|\nabla f\|$.
  • In 2D, $\partial F_x/\partial y = \partial F_y/\partial x$ (in any dimension: the Jacobian of $\mathbf{F}$ is symmetric; it is the Hessian of $f$, see Chapter 2.10). On a region with no holes this test is also sufficient.
  • Its line integrals do not depend on the path, and it has zero curl.
Why do we need it?

A conservative field is the best-behaved kind: everything about it is packed into one scalar potential. And optimisation is exactly "follow a gradient field downhill".

Where is it used?

Gradient descent (the loss and its field $-\nabla L$), energy-based models (the force is the gradient of an energy), potential energy and gravity in physics, and the score function in diffusion models, which is by definition the gradient of $\log p$.

How is it used?

To check whether a field is a gradient, test $\partial F_x/\partial y = \partial F_y/\partial x$. If it passes, integrate to find the potential. To optimise, follow $-\mathbf{F}$ downhill.

The grey curves are contours of the potential $f$. The faint arrows are $\nabla f$. Drag the dot: the green arrow is $\nabla f$ there. Check that it always crosses the contour at a right angle (the dashed purple line is the contour's direction), and that it is long where the contours are crowded. In the readout, the two mixed partial derivatives are always equal: the conservative test.

"Conservative" does not mean "constant" or "small". It only means "is the gradient of something". The test $\partial F_x/\partial y = \partial F_y/\partial x$ can fail to be sufficient on a region with a hole (a famous example is the swirl around a point), but for ordinary fields on the whole plane it is enough.

Quick check: is $\mathbf{F}=(3x^2y,\;x^3 + 2y)$ conservative? If so, find a potential.

$\partial F_x/\partial y = 3x^2$ and $\partial F_y/\partial x = 3x^2$: equal, so yes. A potential is $f = x^3y + y^2$: check $\partial f/\partial x = 3x^2y$ ✓ and $\partial f/\partial y = x^3 + 2y$ ✓.

Three old friends, seen as questions about fields

Three tools you already know turn out to be three natural questions about fields. No new maths here, only new names for old ideas.

  • Directional derivative. Standing in a scalar field and walking in a chosen direction: how fast does the number change? (Chapter 2.4)
  • Jacobian. In a vector field, if I nudge my position, how do the arrows change? (Chapter 2.5)
  • Hessian. In a scalar field, how does the gradient field itself change? It is the Jacobian of $\nabla f$. (Chapter 2.10)

Directional derivative. $f=x^2y$ at $(1,2)$ has $\nabla f=(4,1)$. Walk in the direction $\mathbf{u} = (3,4)/5$ (a unit vector). Then $D_{\mathbf{u}}f = \nabla f\cdot\mathbf{u} = (4\cdot3 + 1\cdot4)/5 = 16/5 = 3.2$.

Jacobian of a vector field. $\mathbf{F}=(-y,\,x)$ has $J = \begin{bmatrix}0&-1\\1&0\end{bmatrix}$. Its trace is $0+0=0$ and $J_{21}-J_{12} = 1-(-1)=2$. (These two numbers will be the divergence and curl in a moment.)

Hessian as a Jacobian. $\nabla f = (2xy,\,x^2)$ has Jacobian $\begin{bmatrix}2y&2x\\2x&0\end{bmatrix}$, which at $(1,2)$ is $\begin{bmatrix}4&2\\2&0\end{bmatrix}$: the Hessian of $f$.

toolof whatdefinition
Directional derivativescalar field $f$, unit direction $\mathbf{u}$$D_{\mathbf{u}}f = \nabla f\cdot\mathbf{u} = \|\nabla f\|\cos\theta$
Jacobianvector field $\mathbf{F}$$J_{ij} = \partial F_i/\partial x_j$
Hessianscalar field $f$$H_{ij} = \partial^2 f/\partial x_i\partial x_j$, i.e. $H = J(\nabla f)$

The directional derivative is largest ($=\|\nabla f\|$) when $\mathbf{u}$ points along $\nabla f$, and zero along a level curve.

Why do we need it?

These three are the core of everything else in this chapter: divergence and curl are just two numbers read off the Jacobian of a vector field. Seeing the connections means you only have one idea to remember.

Where is it used?

Directional derivatives: checking a gradient numerically, sharpness measures and line searches. Jacobians: backpropagation and normalising flows. Hessians: curvature, Newton's method and second-order optimisers.

How is it used?

For a step in direction $\mathbf{u}$, compute $\nabla f\cdot\mathbf{u}$. For a vector field, compute its Jacobian once and read divergence (trace) and curl (the antisymmetric part) from it.

Drag the dot, then turn the direction dial. The orange arrow is the direction $\mathbf{u}$ you walk; the green arrow is $\nabla f$. The readout shows $\nabla f\cdot\mathbf{u}$ and a nudge test (a tiny step really changes $f$ by that amount). Press Steepest ascent: the rate is as large as it can be. Press Along the contour: the rate is 0.

Quick check: $\nabla f = (3,4)$ at a point. What is the largest directional derivative, and in which direction?

The largest rate is $\|\nabla f\| = \sqrt{9+16} = 5$, in the direction of the gradient, $\mathbf{u} = (0.6,\,0.8)$. Check: $(3\cdot0.6 + 4\cdot0.8) = 1.8+3.2 = 5$ ✓.

Divergence: net outflow awareness

Imagine a tiny imaginary box sitting in a flowing fluid. Count how much fluid leaves through its walls and how much enters. If more leaves than enters, the box is fed from inside: there is a source (a tap), and a blob of fluid placed there would expand. If more enters than leaves, there is a sink (a drain) and the blob would shrink. If the two balance, the fluid just flows through (we call that incompressible).

The divergence is that net outflow, per unit of area (or volume). It is a scalar field made from a vector field.

Take $\mathbf{F}=(x,\,y)$ and a tiny square of half-size $s$ centred at the origin (side $2s$, area $4s^2$).

  1. Right wall ($x=s$): outward direction $(1,0)$, $F_x = s$. Outflow $= s\cdot 2s = 2s^2$.
  2. Left wall ($x=-s$): outward direction $(-1,0)$, $F_x=-s$, so $\mathbf{F}\cdot\mathbf{n} = s$. Outflow $=2s^2$.
  3. Top and bottom walls: the same, $2s^2$ each.
  4. Total outflow $=8s^2$. Per unit area: $8s^2/4s^2 = 2$.

And $\partial F_x/\partial x + \partial F_y/\partial y = 1 + 1 = 2$ ✓. The divergence of a source field $(x,y)$ is $+2$ everywhere. For the rotation field $(-y,x)$ it is $0+0=0$: fluid circles around and nothing expands.

The divergence of $\mathbf{F}=(F_1,\dots,F_n)$ is the sum of the matching partial derivatives:

$$\nabla\!\cdot\mathbf{F} \;=\; \sum_{i=1}^n \frac{\partial F_i}{\partial x_i} \;=\; \frac{\partial F_x}{\partial x} + \frac{\partial F_y}{\partial y}\;(+\tfrac{\partial F_z}{\partial z}\text{ in 3D}).$$

Where it comes from. For a small box of side $2s$ at $(x,y)$: outflow through the right wall $\approx F_x(x+s,y)\cdot2s$, through the left wall $\approx -F_x(x-s,y)\cdot 2s$. Their sum is $[F_x(x+s,y)-F_x(x-s,y)]\cdot2s\approx \dfrac{\partial F_x}{\partial x}\cdot 2s\cdot 2s$, which is $\dfrac{\partial F_x}{\partial x}\times$ area. The top and bottom walls give the $y$-term. Divide by the area to get the formula.

Note: the divergence is the trace of the Jacobian of $\mathbf{F}$ (the sum of its diagonal entries, see the trace), a scalar, not a vector.

Why do we need it?

It measures, at every point, whether a flow is creating or removing "stuff". That is how we express conservation laws such as "mass is neither made nor destroyed here".

Where is it used?

Fluid flow and electromagnetism (Maxwell's equations). In ML: continuous normalising flows (the log-density of a point changes at the rate minus the divergence of the flow), diffusion-model probability-flow equations, and physics-informed networks that penalise a nonzero divergence.

How is it used?

Take the derivative of each component with respect to its own coordinate and add them. In code, it is the trace of the Jacobian (autodiff, or a randomised trace estimator when the dimension is large).

The blue square is a small box at the dragged position. Green arrows on its walls show fluid leaving, red arrows show fluid entering. Move the time slider: the orange shape is that box carried along by the flow. For a small time its area ratio is close to $1 + (\text{divergence})\times t$ (for a constant divergence the exact ratio is $e^{\text{divergence}\times t}$, because the growth compounds). Try source (expands), sink (shrinks), rotation (no change in area, only turns) and varying source (expands on the right, shrinks on the left).

Quick check: what is the divergence of $\mathbf{F}=(x^2,\;3y)$ at $(2,1)$?

$\partial F_x/\partial x = 2x = 4$ and $\partial F_y/\partial y = 3$, so $\nabla\cdot\mathbf{F} = 4+3 = 7$: a strong source at that point.

Curl: rotation awareness

Dip a tiny paddle wheel into a flowing river. If the water pushes one side of the wheel harder than the other, the wheel spins. How fast and which way it spins tells you the curl of the flow at that spot.

Surprise: a flow can have perfectly straight streamlines and still spin the wheel. In a river that is fast in the middle and slow near the bank, the water on one side of the wheel moves faster than on the other, so it turns. And a flow that goes round in circles can have zero curl, if the speed drops off in just the right way. Curl measures local spin, not "does it go around".

In 2D, $\operatorname{curl}\mathbf{F} = \dfrac{\partial F_y}{\partial x} - \dfrac{\partial F_x}{\partial y}$ (positive = counter-clockwise).

  • Rotation $\mathbf{F}=(-y,x)$: $\partial F_y/\partial x = 1$ and $\partial F_x/\partial y = -1$. Curl $=1-(-1) = 2$. The wheel spins counter-clockwise.
  • Shear $\mathbf{F}=(y,0)$ (wind blows right, faster higher up): $\partial F_y/\partial x = 0$ and $\partial F_x/\partial y = 1$. Curl $= 0 - 1 = -1$. The top of the wheel is pushed right harder than the bottom, so it spins clockwise, although every arrow points right.
  • Source $\mathbf{F}=(x,y)$: $0 - 0 = 0$. No spin.

2D. $\operatorname{curl}\mathbf{F} = \dfrac{\partial F_y}{\partial x} - \dfrac{\partial F_x}{\partial y}$, a number: the local spin. It equals the circulation around a tiny loop divided by the loop's area. A paddle wheel turns at angular speed $\tfrac12\operatorname{curl}\mathbf{F}$.

3D. The curl is a vector:

$$\nabla\times\mathbf{F} = \left(\frac{\partial F_z}{\partial y} - \frac{\partial F_y}{\partial z},\;\; \frac{\partial F_x}{\partial z} - \frac{\partial F_z}{\partial x},\;\; \frac{\partial F_y}{\partial x} - \frac{\partial F_x}{\partial y}\right).$$

It points along the axis the paddle wheel turns about (right-hand rule: curl your right fingers the way it spins, and your thumb points along the vector), and its length is the spin rate. The last entry is the 2D curl. A gradient field has zero curl, because mixed partial derivatives are equal ($f_{xy}=f_{yx}$), so every difference above is $0$. This is the same test as in the conservative-field section. In terms of the Jacobian, the curl is read from its antisymmetric part ($J_{21}-J_{12}$ in 2D).

Why do we need it?

It tells us, at each point, how much a flow rotates, and whether a field could be a gradient (no spin) or must contain a swirl.

Where is it used?

Fluid vortices and weather, electromagnetism (Maxwell's equations), computer graphics (curl noise for smoke), and as a diagnostic in ML: learned "force" fields in energy-based or score models should have zero curl if they are real gradients.

How is it used?

Compute $\partial F_y/\partial x - \partial F_x/\partial y$ (2D). If it is zero everywhere, the field is conservative (on a region with no holes). If not, the field has swirl and no scalar potential.

Drag the wheel anywhere. Press Spin (or move the time slider): the wheel turns at half the curl. The readout adds the flow around a small circle (counter-clockwise is positive) and divides by the area: it matches the formula (exactly for fields with constant curl, closely for a small circle otherwise). Try rotation (counter-clockwise), shear (clockwise, even though all arrows point the same way!), source (no spin) and fading whirl (the spin changes with position).

Quick check: what is the 2D curl of $\mathbf{F}=(x^2y,\;xy^2)$ at $(1,2)$?

$\partial F_y/\partial x = y^2 = 4$ and $\partial F_x/\partial y = x^2 = 1$. Curl $=4-1 = 3$ (counter-clockwise spin).

Line integrals: work along a path awareness

Push a cart along a winding path while a wind (a vector field) pushes on it. At each tiny step, only the part of the wind pointing along your step helps or hinders you. Add up (wind along the step) $\times$ (step length) over the whole path. That total is the work, and it is called the line integral of the field along the path.

A natural question: does the total depend on which path you took between the same two points? For some fields it does not. Those are exactly the gradient (conservative) fields.

Go from $A=(0,0)$ to $B=(2,1)$. Compare two paths: the straight line, and "right to $(2,0)$, then up".

Field $\mathbf{F}=(y,\,x)$ (this is $\nabla(xy)$):

  1. Straight: $\mathbf{r}(t)=(2t,\,t)$, $d\mathbf{r}=(2,1)\,dt$, and $\mathbf{F}=(t,\,2t)$. So $\mathbf{F}\cdot d\mathbf{r} = (2t+2t)\,dt = 4t\,dt$, and $\int_0^1 4t\,dt = 2$.
  2. Corner path, first leg $(2t,0)$: $\mathbf{F}=(0,2t)$, $d\mathbf{r}=(2,0)\,dt$, $\mathbf{F}\cdot d\mathbf{r}=0$. Second leg $(2,t)$: $\mathbf{F}=(t,2)$, $d\mathbf{r}=(0,1)\,dt$, $\mathbf{F}\cdot d\mathbf{r}=2\,dt$, giving $2$. Total $0+2=2$.
  3. Same answer, $2$, and it equals $f(B)-f(A) = 2\cdot1 - 0 = 2$ for the potential $f=xy$.

Field $\mathbf{F}=(-y,\,x)$ (rotation): straight path: $\mathbf{F}=(-t,2t)$, $\mathbf{F}\cdot d\mathbf{r}=(-2t+2t)\,dt=0$, so the work is $0$. Corner path: first leg $0$, second leg $\mathbf{F}=(-t,2)$, $d\mathbf{r}=(0,1)\,dt$: work $2$. Different answers: the work depends on the path, so this field is not conservative.

The line integral of $\mathbf{F}$ along a curve $C$ given by $\mathbf{r}(t)$, $a\le t\le b$, is

$$\int_C \mathbf{F}\cdot d\mathbf{r} \;=\; \int_a^b \mathbf{F}(\mathbf{r}(t))\cdot\mathbf{r}'(t)\,dt \;\approx\; \sum_i \mathbf{F}(\mathbf{r}_i)\cdot\Delta\mathbf{r}_i.$$

(Add up "field dotted with each small step".) For a gradient field $\mathbf{F}=\nabla f$ it collapses to the fundamental theorem for line integrals:

$$\int_C \nabla f\cdot d\mathbf{r} = f(B) - f(A).$$

Why. By the chain rule, $\dfrac{d}{dt}f(\mathbf{r}(t)) = \nabla f\cdot\mathbf{r}'(t)$. So the integral above is $\int_a^b \dfrac{d}{dt}f(\mathbf{r}(t))\,dt = f(\mathbf{r}(b)) - f(\mathbf{r}(a))$, by the ordinary fundamental theorem of calculus. Only the endpoints matter, and the work around any closed loop is $0$.

Why do we need it?

It adds up an effect that varies along a path: work done by a force, flow along a pipe, or the change of a quantity along a trajectory. And it gives a clean test for "is this a gradient?": path independence.

Where is it used?

Physics (work and energy, circulation), and in ML the idea behind path-based attribution methods such as integrated gradients (which integrate the gradient of a model along a straight path from a baseline input to the real input), and thinking of training as a path through weight space.

How is it used?

Parameterise the path, take many small steps, and add $\mathbf{F}\cdot\Delta\mathbf{r}$. If the field is a gradient, skip all that and subtract the potential at the two ends.

Drag the two black end points A and B, and the teal middle point M. The orange path is the straight line A→B; the teal path goes A→M→B. The readout adds up $\mathbf{F}\cdot d\mathbf{r}$ numerically along each. For the two gradient fields the numbers always agree (and equal the potential difference). For rotation and wind they differ: the difference is the work around the closed loop A→M→B→A.

Quick check: for $\mathbf{F}=\nabla f$ with $f(x,y)=x^2+y$, find the work from $(0,0)$ to $(1,3)$ along any path.

It depends only on the ends: $f(1,3)-f(0,0) = (1+3) - 0 = 4$. Check the gradient: $\nabla f=(2x,\,1)$; along the straight path $(t,3t)$, $\mathbf{F}\cdot d\mathbf{r} = (2t\cdot1 + 1\cdot3)\,dt$ and $\int_0^1(2t+3)\,dt = 1+3 = 4$ ✓.

Surface integrals and flux awareness

Hold a net in a steady wind. How much air passes through it? That amount is the flux of the wind through the net. It depends on how big the net is, how strong the wind is, and how the net is tilted: face-on to the wind catches the most, edge-on catches none.

Cut the surface into tiny patches. For each, take the part of the wind that goes straight through it (the wind dotted with the patch's unit normal arrow) times the patch's area. Add them all. That sum is a surface integral of the field.

Uniform wind $\mathbf{F} = (0,0,2)$ (straight up, strength 2) through a flat $3\times3$ square lying horizontally. Its normal is $\mathbf{n}=(0,0,1)$, so $\mathbf{F}\cdot\mathbf{n}=2$ and the flux is $2\times 9 = 18$.

Tilt the square so that its normal is $\mathbf{n}=(0,\,0.6,\,0.8)$ (a unit vector, since $0.36+0.64=1$). Now $\mathbf{F}\cdot\mathbf{n} = 1.6$ and the flux drops to $1.6\times9 = 14.4$. Turn it edge-on ($\mathbf{n}\perp\mathbf{F}$) and the flux is $0$.

The flux of $\mathbf{F}$ through a surface $S$ (with a chosen unit normal $\mathbf{n}$) is the surface integral

$$\iint_S \mathbf{F}\cdot\mathbf{n}\;dS \;\approx\; \sum_{\text{patches}} (\mathbf{F}\cdot\mathbf{n})\,\Delta S.$$

For a surface that is a graph $z=g(x,y)$ with the normal pointing up, $\mathbf{n}\,dS = (-g_x,\,-g_y,\,1)\,dx\,dy$, so

$$\text{flux} = \iint \big(F_z - F_x\,g_x - F_y\,g_y\big)\,dx\,dy.$$

For a closed surface with the normal pointing outward, the flux is the net amount flowing out, and the divergence theorem says this equals the total divergence inside (see the table below).

Why do we need it?

It counts how much of a flow crosses a boundary: how much heat leaves a room, how much fluid goes through a membrane. It is the "per-surface" cousin of divergence.

Where is it used?

Physics (Gauss's law for electric flux, heat and fluid flow). In ML it is rare in everyday work: it shows up in the maths of continuity equations behind normalising flows and diffusion models, and in graphics (rendering integrates light over surfaces).

How is it used?

Describe the surface, find the normal on each patch, dot it with the field and add up area-weighted values (by formula, or numerically on a grid). Or use the divergence theorem to turn a hard surface sum into a volume integral.

Rotate the view (drag the background; try Side). The blue surface is a bump of height $h$ above a square of area 16. The thin orange arrows are the wind, the green arrows are the surface normals. Drag the dot up and down to change the bump height $h$ (or press Flatten). With the upward wind the flux stays 16 whatever $h$ is. With the sideways wind it stays 0 (what enters on one side leaves on the other). With the outward (radial) wind it grows with the bump.

Quick check: wind $\mathbf{F}=(1,0,0)$ blows through a flat vertical $2\times 5$ rectangle whose normal is $(1,0,0)$. What is the flux?

$\mathbf{F}\cdot\mathbf{n} = 1$ and the area is $2\cdot5 = 10$, so the flux is $10$. (If the rectangle were turned to face another way, e.g. $\mathbf{n}=(0,1,0)$, the flux would be $0$.)

The big theorems, in plain words awareness

All of these theorems say the same thing in different clothes: to add up a "change" over a region, you only need to look at the boundary. You already know the simplest one. To find the total change of a quantity from $a$ to $b$ you do not need to add up every tiny change in between: just subtract the value at $a$ from the value at $b$. The vector-calculus theorems extend this from a line segment to curves, flat regions, surfaces and solids.

Check Green's theorem on the unit circle with the rotation field $\mathbf{F}=(-y,x)$. Going round the circle $\mathbf{r}=(\cos\phi,\sin\phi)$, the velocity is $(-\sin\phi,\cos\phi)$ and $\mathbf{F}=(-\sin\phi,\cos\phi)$ too. So $\mathbf{F}\cdot d\mathbf{r} = (\sin^2\phi+\cos^2\phi)\,d\phi = d\phi$ and the circulation is $\int_0^{2\pi}d\phi = 2\pi$. On the other side: the curl is $2$ everywhere and the disc has area $\pi$, so $\iint \text{curl}\,dA = 2\pi$. They match.

TheoremFormulaIn plain words
Fundamental theorem of calculus (1D, the ancestor)$\displaystyle\int_a^b f'(x)\,dx = f(b)-f(a)$The total change is the sum of the small changes.
Fundamental theorem for line integrals$\displaystyle\int_C \nabla f\cdot d\mathbf{r} = f(B)-f(A)$The work of a gradient field depends only on where you start and end.
Green's theorem (2D)$\displaystyle\oint_C \mathbf{F}\cdot d\mathbf{r} = \iint_D \Big(\frac{\partial F_y}{\partial x}-\frac{\partial F_x}{\partial y}\Big)\,dA$The circulation around a loop equals the total spin inside it.
Stokes' theorem (3D)$\displaystyle\oint_{\partial S} \mathbf{F}\cdot d\mathbf{r} = \iint_S (\nabla\times\mathbf{F})\cdot\mathbf{n}\,dS$The circulation around the rim of a surface equals the total spin passing through it.
Divergence theorem (Gauss)$\displaystyle\oiint_{S}\mathbf{F}\cdot\mathbf{n}\,dS = \iiint_V \nabla\!\cdot\mathbf{F}\,dV$The total outflow through a closed surface equals the total source strength inside.

Conditions (each theorem is only true when they hold):

  • Line integrals: $f$ is smooth (continuously differentiable) and $C$ is a smooth curve from $A$ to $B$.
  • Green: $C$ is a simple closed curve (it does not cross itself) bounding the region $D$, travelled counter-clockwise (so $D$ is on your left), and $\mathbf{F}$ is smooth on all of $D$, with no holes or blow-ups inside.
  • Stokes: $S$ is an oriented surface with edge $\partial S$. The direction round the edge and the choice of normal $\mathbf{n}$ go together by the right-hand rule (curl the fingers of your right hand along the edge, and your thumb points along $\mathbf{n}$). $\mathbf{F}$ is smooth on $S$.
  • Divergence: $S$ is the closed surface that bounds the solid $V$, with $\mathbf{n}$ pointing outward, and $\mathbf{F}$ is smooth everywhere inside $V$.

Green's theorem also has a "flux form" that says the same thing for the divergence in 2D: $\oint_C \mathbf{F}\cdot\mathbf{n}\,ds = \iint_D \nabla\cdot\mathbf{F}\,dA$ (the flow out through a loop equals the total divergence inside). The second comparison in the widget below is exactly this.

Why do we need it?

They let us trade a hard sum for an easy one: a sum over a whole region can be replaced by a sum over just its edge (or the reverse). They also connect the local ideas (divergence, curl) to the global ones (flux, circulation).

Where is it used?

Deriving conservation laws in physics, electromagnetism, and fluid mechanics. In ML they sit in the background: the divergence theorem is used to derive the continuity equation (how a density is carried along by a flow), and that equation is the starting point for the theory of continuous normalising flows and diffusion models.

How is it used?

To compute a flux or circulation, check whether replacing it by a divergence or curl integral is easier. Mostly, for ML, it is enough to recognise the names and what they say.

Pick a field, drag the circle's centre and change its radius. Two pairs of numbers are compared. (1) The circulation around the circle (blue arrows follow it) against the total curl inside. (2) The flow out through the circle against the total divergence inside. Each pair agrees, whatever the field, because the boundary sees exactly what the inside adds up.

Where these ideas appear in ML, honestly.

  • Gradient flow. Gradient descent with very small steps follows the flow lines of the field $-\nabla L$. This is the most direct use of vector fields in ML, and it is fully covered by earlier chapters.
  • Normalising flows and neural ODEs. A learned vector field moves points around; for continuous flows the log-density changes at the rate $-\nabla\!\cdot\mathbf{F}$ (minus the trace of the Jacobian). Divergence is therefore a real working quantity there.
  • Diffusion and score models. The learned "score" $\nabla_{\mathbf{x}}\log p(\mathbf{x})$ is a gradient field, and sampling moves points along it (plus random noise in the stochastic version). The theory uses the divergence theorem, but practitioners rarely compute it by hand.
  • Physics-informed models. Physics-informed neural networks add terms such as "divergence $=0$" to the loss to force a flow to conserve mass.
  • Not needed for ordinary supervised learning. Training a classifier or a regressor uses gradients, Jacobians (backprop) and sometimes Hessians. Line and surface integrals almost never appear. That is why this chapter is for awareness.
Quick check: a vector field has divergence $3$ everywhere. How much flows out of a closed surface that encloses a region of volume $2$?

By the divergence theorem, outflow $=\iiint\nabla\cdot\mathbf{F}\,dV = 3\times 2 = 6$.

Recap, cheat sheet and practice

  • A scalar field gives a number at every point; a vector field gives an arrow. A flow line follows the arrows: $d\mathbf{x}/dt=\mathbf{F}(\mathbf{x})$.
  • $\nabla f$ is a gradient field, perpendicular to the level curves. A field is conservative if it is a gradient; in 2D test $\partial F_x/\partial y=\partial F_y/\partial x$.
  • Directional derivative $\nabla f\cdot\mathbf{u}$; Jacobian $\partial F_i/\partial x_j$; Hessian $=$ Jacobian of $\nabla f$.
  • Divergence $\nabla\cdot\mathbf{F}=\sum\partial F_i/\partial x_i$ = net outflow = trace of the Jacobian. Curl (2D) $\partial F_y/\partial x-\partial F_x/\partial y$ = spin; in 3D it is a vector along the spin axis.
  • Line integral $\int_C\mathbf{F}\cdot d\mathbf{r}$ = work along a path; for $\mathbf{F}=\nabla f$ it equals $f(B)-f(A)$. Flux $\iint\mathbf{F}\cdot\mathbf{n}\,dS$ = how much flows through a surface.
  • The big theorems all say "total of a derivative inside = value on the boundary". For ML, this chapter is awareness: the daily tools are gradients, Jacobians, Hessians and Taylor expansions.

Cheat sheet

ToolFormulaPicture
Gradient$\nabla f = (\partial f/\partial x_i)_i$arrow uphill, perpendicular to contours
Conservative test (2D)$\partial F_x/\partial y = \partial F_y/\partial x$no swirl, has a potential
Divergence$\partial F_x/\partial x + \partial F_y/\partial y$little box expands or shrinks
Curl (2D)$\partial F_y/\partial x - \partial F_x/\partial y$paddle wheel spins
Line integral$\int_C \mathbf{F}\cdot d\mathbf{r}$work along a path
Flux$\iint_S \mathbf{F}\cdot\mathbf{n}\,dS$wind through a net
Green / Stokescirculation $=\iint$ curlloop = spin inside
Divergence theoremflux out $=\iiint$ divoutflow = sources inside
Code it · NumPy

import numpy as np

# Numerical divergence and 2D curl of a field F(x, y) -> [Fx, Fy]
def div_curl(F, x, y, h=1e-5):
    dFx_dx = (F(x + h, y)[0] - F(x - h, y)[0]) / (2 * h)
    dFx_dy = (F(x, y + h)[0] - F(x, y - h)[0]) / (2 * h)
    dFy_dx = (F(x + h, y)[1] - F(x - h, y)[1]) / (2 * h)
    dFy_dy = (F(x, y + h)[1] - F(x, y - h)[1]) / (2 * h)
    return dFx_dx + dFy_dy, dFy_dx - dFx_dy        # (divergence, curl)

rot    = lambda x, y: np.array([-y, x])
source = lambda x, y: np.array([x, y])
shear  = lambda x, y: np.array([y, 0.0])
print(np.round(div_curl(rot, 1.0, 2.0), 4))       # [0. 2.]   spins, does not expand
print(np.round(div_curl(source, 1.0, 2.0), 4))    # [2. 0.]   expands, does not spin
print(np.round(div_curl(shear, 1.0, 2.0), 4))     # [ 0. -1.] straight flow that still spins

# Line integral of F along a straight segment P -> Q (midpoint rule)
def line_integral(F, P, Q, n=1000):
    P, Q = np.array(P, float), np.array(Q, float)
    t = (np.arange(n) + 0.5) / n
    dr = (Q - P) / n
    return sum(F(*(P + ti * (Q - P))) @ dr for ti in t)

def along(F, pts):                                 # a path made of several segments
    return sum(line_integral(F, pts[i], pts[i + 1]) for i in range(len(pts) - 1))

grad = lambda x, y: np.array([y, x])               # = gradient of f(x, y) = x*y
A, M, B = (0, 0), (2, 0), (2, 1)
print(along(grad, [A, B]), along(grad, [A, M, B])) # 2.0  2.0  same: path independent (= f(B) - f(A))
print(along(rot,  [A, B]), along(rot,  [A, M, B])) # 0.0  2.0  different: rot is not conservative

# Green's theorem on the unit circle for the rotation field: circulation = integral of curl
phi = np.linspace(0, 2 * np.pi, 2000, endpoint=False)
dphi = 2 * np.pi / 2000
pts = np.stack([np.cos(phi), np.sin(phi)], axis=1)
tangent = np.stack([-np.sin(phi), np.cos(phi)], axis=1)
circulation = sum(rot(*p) @ t for p, t in zip(pts, tangent)) * dphi
print(circulation, 2 * (np.pi * 1.0**2))           # 6.2831...  6.2831...  (curl 2 times area pi)
Test yourself

1. Which of these is a scalar field?

Temperature is one number per point. Wind and gradients attach an arrow (a vector) to each point, and a Jacobian attaches a matrix.

2. Is $\mathbf{F}=(2x+y,\;x+3)$ a gradient field?

$\partial F_x/\partial y = 1$ and $\partial F_y/\partial x = 1$ match, so it passes the test. A potential is $f = x^2 + xy + 3y$.

3. The divergence of $\mathbf{F}=(3x,\;-y)$ is…

$\partial F_x/\partial x = 3$ and $\partial F_y/\partial y = -1$. The sum is $3 + (-1) = 2$.

4. A paddle wheel sits in the field $\mathbf{F}=(y,\,0)$ (wind to the right, stronger higher up). It…

Curl $= \partial F_y/\partial x - \partial F_x/\partial y = 0 - 1 = -1$. The top of the wheel is pushed right harder than the bottom, which is a clockwise turn.

5. For $\mathbf{F}=\nabla f$ with $f(x,y)=x^2y$, the line integral along any path from $(0,0)$ to $(2,1)$ is…

It is a gradient field, so only the ends matter: $f(2,1)-f(0,0) = 4\cdot1 - 0 = 4$.

6. Where do divergence, curl and surface integrals matter most in everyday ML practice?

Ordinary training needs gradients, Jacobians and Hessians. Divergence appears in continuous normalising flows and physics-informed losses, and the theorems sit behind the theory of diffusion models. It is awareness-level knowledge.

Practice problems

A. For $f(x,y) = x^2 + 3y^2$, sketch (describe) the gradient field and find $D_{\mathbf{u}}f$ at $(1,1)$ for $\mathbf{u}=(1,0)$.

$\nabla f = (2x,\;6y)$. At $(1,1)$: $(2,6)$. The arrows point away from the origin (uphill), steeper in $y$, and are perpendicular to the elliptical contours. $D_{\mathbf{u}}f = (2,6)\cdot(1,0) = 2$.

B. Is $\mathbf{F}=(y\cos x,\;\sin x + 2y)$ conservative? Find a potential.

$\partial F_x/\partial y = \cos x$ and $\partial F_y/\partial x = \cos x$: equal, so yes. Integrate $F_x$ with respect to $x$: $f = y\sin x + g(y)$. Then $\partial f/\partial y = \sin x + g'(y)$ must equal $\sin x + 2y$, so $g'(y)=2y$ and $g=y^2$. Potential: $f = y\sin x + y^2$.

C. Compute the divergence and 2D curl of $\mathbf{F}=(x^2-y^2,\;2xy)$ at $(1,1)$.

$\partial F_x/\partial x = 2x = 2$, $\partial F_y/\partial y = 2x = 2$: divergence $4$. $\partial F_y/\partial x = 2y = 2$ and $\partial F_x/\partial y = -2y = -2$: curl $2-(-2)=4$. The field both expands and spins counter-clockwise at that point.

D. Compute $\int_C \mathbf{F}\cdot d\mathbf{r}$ for $\mathbf{F}=(-y,\,x)$ along the straight segment from $(1,0)$ to $(0,1)$.

$\mathbf{r}(t) = (1-t,\;t)$, $d\mathbf{r}=(-1,\,1)\,dt$, $\mathbf{F} = (-t,\;1-t)$. So $\mathbf{F}\cdot d\mathbf{r} = (t + 1 - t)\,dt = dt$, and the integral is $\int_0^1 dt = 1$. (Compare: along the quarter circle between the same points the work is $\pi/2\approx1.57$, since $\mathbf{F}\cdot d\mathbf{r}=d\phi$ there. Different paths, different work: not conservative.)

E. Wind $\mathbf{F}=(0,\,3,\,4)$ crosses a flat square of area $5$ whose unit normal is $\mathbf{n}=(0,\,0.6,\,0.8)$. Find the flux.

$\mathbf{F}\cdot\mathbf{n} = 3\cdot0.6 + 4\cdot0.8 = 1.8 + 3.2 = 5$. Flux $=5\cdot 5 = 25$. (The wind is exactly face-on to the square, since $\mathbf{F}=5\mathbf{n}$.)

F. Use Green's theorem to find the circulation of $\mathbf{F}=(-y,\,x)$ around a circle of radius $3$.

The curl is $2$ everywhere and the disc has area $\pi\cdot3^2 = 9\pi$. So the circulation is $2\cdot9\pi = 18\pi\approx 56.5$. Direct check: on the circle $\mathbf{F}\cdot d\mathbf{r} = 3\cdot3\,d\phi = 9\,d\phi$ (speed 3, radius 3), and $\int_0^{2\pi}9\,d\phi = 18\pi$ ✓.

Chapter 2.14

Calculus → ML Connection

This is where everything comes together. A model learns by one repeated move: measure how wrong it is (the loss), ask the derivative which way is downhill (the gradient), and take a small step that way. In this chapter we derive the gradient of every classic model by hand, link each one to a picture, and finish with the training loop that PyTorch runs.

  • See training as "walk downhill on the loss surface", and know exactly what the update rule does
  • Derive the gradients of mean squared error, linear regression, the sigmoid, logistic regression, cross-entropy and softmax, step by step
  • Understand why cross-entropy trains classifiers well and squared error struggles (saturation)
  • Choose learning rates, use momentum, and add L1 / L2 regularisation, knowing what each does to the gradient
  • Follow the chain rule through a small neural network and watch one train live
  • Explain full-batch, mini-batch and stochastic gradient descent, and write the "forward, loss, backward, update" loop

The big picture: learning is going downhill core

What we need from earlier chapters: the derivative as a slope (Chapter 2.3), the gradient as the direction of steepest ascent (Chapter 2.4) and linearization (Chapter 2.12: close to a point, a curve looks like its tangent line).

A model is a machine with knobs. The knobs are called weights. For every setting of the knobs the model makes some predictions, and we can add up how wrong they are. That total wrongness is the loss.

So the loss is a landscape: one knob gives a curve, two knobs give a surface with hills and valleys, a million knobs give a landscape too big to draw (but the maths is the same). Training means: find a low point of this landscape.

We cannot see the whole landscape. But we can feel the ground under our feet: the slope tells us which way is down. So we take a small step downhill, feel the slope again, and repeat. That is all of machine-learning training. The rest of this chapter is about computing the slope for each kind of model.

Take one knob $w$ and three data points $(x, y) = (1, 2),\ (2, 3),\ (3, 7)$. Our model is $\hat y = w\,x$ ("multiply the input by $w$"). Start with $w = 1$.

  1. Forward: predictions $\hat y = [1, 2, 3]$.
  2. Errors (prediction minus truth): $[1-2,\ 2-3,\ 3-7] = [-1, -1, -4]$.
  3. Loss (average of squared errors): $(1 + 1 + 16)/3 = 6$.
  4. Slope of the loss at $w = 1$: $\dfrac{dL}{dw} = \dfrac{2}{3}\big(1\cdot(-1) + 2\cdot(-1) + 3\cdot(-4)\big) = \dfrac23(-15) = -10$ (we derive this formula in the next section). Negative slope: the loss goes down when $w$ goes up.
  5. Update with step size $\eta = 0.1$: $w \leftarrow 1 - 0.1\cdot(-10) = 2$.

The new loss is $(0+1+1)/3$: with $w=2$ the predictions are $[2,4,6]$ and the errors are $[0, 1, -1]$, so $L = 2/3 \approx 0.67$. One step took the loss from 6 down to 0.67.

Four ingredients. A model $f(\mathbf{x};\mathbf{w})$ with weights $\mathbf{w}$. A loss per example $\ell(\hat y, y)$ that is small when the prediction is good. The training loss, the average over the $n$ examples:

$$L(\mathbf{w}) = \frac1n\sum_{i=1}^{n}\ell\big(f(\mathbf{x}_i;\mathbf{w}),\,y_i\big).$$

And the gradient descent update, where $\eta > 0$ is the learning rate (step size):

$$\mathbf{w} \leftarrow \mathbf{w} - \eta\,\nabla L(\mathbf{w}).$$

The gradient $\nabla L$ is a column vector with one slope $\partial L/\partial w_j$ for each weight. Why does this step go downhill? Linearization (Chapter 2.12) says that for a small step $\boldsymbol{\Delta}$, $L(\mathbf{w}+\boldsymbol{\Delta}) \approx L(\mathbf{w}) + \nabla L^\top\boldsymbol{\Delta}$. Choose $\boldsymbol{\Delta} = -\eta\nabla L$. Then the change is $-\eta\,\nabla L^\top\nabla L = -\eta\,\|\nabla L\|^2 \le 0$. A squared length is never negative, so the loss can only go down (for small enough $\eta$).

Why do we need it?

A model has far too many possible settings to try them all. The slope of the loss tells us, for free, which tiny change to every knob lowers the error. That turns "search everywhere" into "keep walking downhill".

Where is it used?

Training linear and logistic regression, every neural network (CNNs, Transformers, diffusion models), matrix factorisation for recommenders, and almost any model with a differentiable loss.

How is it used?

Repeat four steps: run the model forward to get predictions, compute the loss, compute the gradient of the loss with respect to every weight (backward), and move the weights a small step against the gradient.

The curve is the loss $L(w)$ for the three data points above. Drag the dot along it: the orange tangent line is the slope $dL/dw$. Press Take one step to move by $-\eta\cdot\text{slope}$ and watch the dot slide toward the green minimum. Raise $\eta$ to 0.2 and see the steps overshoot to the other side.

"Gradient descent" never promises the best possible answer. It walks to a low point. For the bowl-shaped losses in this chapter's first sections (linear and logistic regression) there is only one low point, so we reach the best answer. For neural networks there are many, and we usually accept a good one (see Chapter 2.10 on curvature and saddle points).

Quick check: at some $w$ the slope is $+6$ and $\eta = 0.1$. Which way does $w$ move, and by how much?

$w \leftarrow w - 0.1\cdot 6 = w - 0.6$. A positive slope means the loss rises as $w$ rises, so we move $w$ left (down), by $0.6$.

Mean squared error: why squared, and what its gradient looks like core

What we need from earlier chapters: the power rule and the chain rule for one variable (Chapters 2.3 and 2.8), and the idea of a length (norm) from the Linear Algebra guide: MSE is a squared length.

For one example, the error is how far the prediction is from the truth: $e = \hat y - y$. To score many examples with one number we cannot just add the errors: a $+3$ and a $-3$ would cancel to zero and look perfect. So we square each error first. Squares are never negative, and they punish big mistakes much more than small ones (an error of 4 costs 16, an error of 1 costs only 1).

There is a second reason, and it is the one calculus cares about: the slope of a squared error is proportional to the error itself. Big mistake, big push. Small mistake, small push. When the model is nearly right the steps shrink by themselves, so we glide into the bottom instead of jumping over it.

Same data as before: $(1,2),\ (2,3),\ (3,7)$ and $\hat y = w\,x$. Let us derive the slope of the loss step by step.

  1. Loss for one example: $\ell_i = (w x_i - y_i)^2$. Write $u = w x_i - y_i$, so $\ell_i = u^2$.
  2. Chain rule (Chapter 2.8): $\dfrac{d\ell_i}{dw} = \dfrac{d\ell_i}{du}\cdot\dfrac{du}{dw} = 2u\cdot x_i = 2\,(w x_i - y_i)\,x_i$.
  3. Average over the three examples: $\dfrac{dL}{dw} = \dfrac{2}{3}\sum_i (w x_i - y_i)\,x_i$. Each term is error × input.
  4. At $w = 1$ the errors are $-1, -1, -4$, so the sum is $(-1)(1) + (-1)(2) + (-4)(3) = -15$, and $\dfrac{dL}{dw} = \dfrac23(-15) = -10$.
  5. Numeric check: nudge $w$ by $0.001$. $L(1.001) - L(0.999)$ divided by $0.002$ gives $-10.0000$ (to 4 decimals), which matches.
  6. Set the slope to zero to find the best $w$: $\sum_i x_i(w x_i - y_i) = 0 \Rightarrow w = \dfrac{\sum x_i y_i}{\sum x_i^2} = \dfrac{29}{14} \approx 2.07$.
$$\text{MSE} = \frac1n\sum_{i=1}^{n}(\hat y_i - y_i)^2, \qquad \frac{\partial\,\text{MSE}}{\partial \hat y_i} = \frac{2}{n}(\hat y_i - y_i).$$

The second formula is the key fact: the gradient with respect to each prediction is proportional to that prediction's error. To get the gradient with respect to a weight, we then pass it back through the model with the chain rule: $\dfrac{\partial L}{\partial w_j} = \sum_i \dfrac{\partial L}{\partial \hat y_i}\,\dfrac{\partial \hat y_i}{\partial w_j}$.

  • Compare with absolute error $|e|$: its slope is $\pm 1$ whatever the size of the error, with a sharp kink at $e=0$. It treats all mistakes alike and the step does not shrink near the bottom.
  • RMSE $=\sqrt{\text{MSE}}$ is in the same units as $y$. It has the same minimum but a different gradient scale.
  • Many books write $\tfrac12\text{MSE}$ so the gradient has no "2". PyTorch's mse_loss has no $\tfrac12$. We keep the 2 here, and say so whenever it matters.
Why do we need it?

We need one number that says how wrong a whole set of predictions is, that never lets errors cancel, and whose slope is smooth so that gradient descent can use it. The squared error does all three.

Where is it used?

The loss for regression: house prices, temperature forecasts, the reconstruction loss of autoencoders, the "value" loss in reinforcement learning, and as a metric for almost every numeric prediction.

How is it used?

Compute the errors, square and average them. For training, take the gradient $\tfrac{2}{n}(\hat y - y)$ at the output and send it backwards. Watch out for outliers: one huge error dominates the sum.

Drag the blue handle at $x = 3$ to tilt the line $\hat y = wx$. The red bars are the errors. In the right picture the blue straight line is the MSE slope: it shrinks to zero smoothly as $w$ approaches the best value. Switch to absolute error and the slope (orange staircase) keeps the same size right up to the bottom, then jumps.

Outliers. One wild data point with error $20$ contributes $400$ to the sum, more than a hundred points with error $2$. MSE bends the fit toward outliers. If that is a problem, use absolute error or the Huber loss (square for small errors, straight line for big ones).

Where the "2" goes. The 2 from the power rule only rescales the gradient (it is the same as doubling $\eta$). It never changes where the minimum is.

Quick check: for one example with $\hat y = 5$ and $y = 8$, what is $\partial(\hat y - y)^2/\partial\hat y$, and which way should $\hat y$ move?

$2(\hat y - y) = 2(5 - 8) = -6$. The slope is negative, so increasing $\hat y$ lowers the loss: move $\hat y$ up toward 8. Note the size, $6$, is twice the error $3$.

Linear regression: derive the gradient, then solve it core

What we need from earlier chapters: the gradient (Chapter 2.4), matrix calculus and the identity $\nabla\|X\mathbf{w}-\mathbf{y}\|^2 = 2X^\top(X\mathbf{w}-\mathbf{y})$ (Chapters 2.6 and 2.7), the Hessian (Chapter 2.10) and, from the Linear Algebra guide, matrix–vector products, the normal equations and linear regression.

Linear regression draws the best straight line (or flat plane) through a cloud of points. The line has two knobs: a slope $w$ and an intercept $b$. Every pair $(b, w)$ is a different line, and each line has a loss (its mean squared error). Plot the loss above the $(b, w)$ plane and you get a smooth bowl. The bottom of the bowl is the best line.

Two ways down: (1) walk with gradient descent, or (2) notice that at the very bottom the ground is flat (the gradient is zero) and solve "gradient $= 0$" directly. That equation is the famous normal equations, the same ones the Linear Algebra guide found from a projection picture. Calculus gives the same answer by a different road.

Data: $(1,2),\ (2,3),\ (3,7)$. To include the intercept, give every input a leading 1. Rows of $X$ are $[1, x_i]$ and the weight vector is $\mathbf{w} = [b, w]^\top$:

$$X = \begin{bmatrix}1&1\\1&2\\1&3\end{bmatrix},\quad \mathbf{y} = \begin{bmatrix}2\\3\\7\end{bmatrix},\quad \mathbf{w}=\begin{bmatrix}0\\1\end{bmatrix}\ (\text{intercept } 0,\ \text{slope } 1).$$
  1. Predictions: $X\mathbf{w} = [1, 2, 3]^\top$. Residual $\mathbf{r} = X\mathbf{w}-\mathbf{y} = [-1, -1, -4]^\top$.
  2. Loss: $L = \frac13\mathbf{r}^\top\mathbf{r} = \frac13(1+1+16) = 6$.
  3. $X^\top\mathbf{r} = [\,1(-1)+1(-1)+1(-4),\ \ 1(-1)+2(-1)+3(-4)\,]^\top = [-6, -15]^\top$.
  4. Gradient $\nabla L = \frac23 X^\top\mathbf{r} = [-4, -10]^\top$. The second entry, $-10$, matches the slope we found in the last section.
  5. One step with $\eta = 0.1$: $\mathbf{w} \leftarrow [0,1] - 0.1\,[-4,-10] = [0.4,\ 2.0]$. New residual $[0.4, 1.4, -0.6]$, new loss $(0.16+1.96+0.36)/3 \approx 0.83$.
  6. Zero gradient instead. $X^\top X = \begin{bmatrix}3&6\\6&14\end{bmatrix}$ and $X^\top\mathbf{y} = [12, 29]^\top$. Solving $X^\top X\mathbf{w} = X^\top\mathbf{y}$ gives $\mathbf{w}^\star = [-1,\ 2.5]^\top$: the line $\hat y = -1 + 2.5x$. Its residual is $[-0.5, 1, -0.5]$, and $X^\top\mathbf{r} = [0, 0]^\top$ (the residual adds to zero and is perpendicular to the $x$ column). The loss is $0.5$.

Model $\hat{\mathbf{y}} = X\mathbf{w}$ ($X$ is $n\times d$, one row per example). Loss (the mean squared error, a squared length divided by $n$):

$$L(\mathbf{w}) = \frac1n\|X\mathbf{w}-\mathbf{y}\|^2 = \frac1n\,\mathbf{r}^\top\mathbf{r},\qquad \mathbf{r} = X\mathbf{w}-\mathbf{y}.$$

Derivation, one weight at a time. $L = \frac1n\sum_i r_i^2$ and $r_i = \sum_j X_{ij}w_j - y_i$, so $\partial r_i/\partial w_j = X_{ij}$. The chain rule gives

$$\frac{\partial L}{\partial w_j} = \frac1n\sum_i 2\,r_i\,\frac{\partial r_i}{\partial w_j} = \frac2n\sum_i X_{ij}\,r_i = \frac2n\,(X^\top\mathbf{r})_j.$$

Stack the $d$ partial derivatives into a column (the gradient is a column vector):

$$\boxed{\ \nabla L(\mathbf{w}) = \frac2n X^\top(X\mathbf{w}-\mathbf{y})\ }\qquad\text{(shape check: } d\times n \text{ times } n\times 1 = d\times 1\text{)}$$

Set it to zero: $X^\top X\mathbf{w} = X^\top\mathbf{y}$ (the normal equations), so $\mathbf{w}^\star = (X^\top X)^{-1}X^\top\mathbf{y}$ when $X^\top X$ is invertible. Read $X^\top\mathbf{r} = \mathbf{0}$ as a picture: the residual is perpendicular to every column of $X$, which is exactly the projection picture of Chapter 1.10.

Why the bottom is the bottom. The Hessian is $\frac2n X^\top X$, which is positive semi-definite (it never curves downward), so $L$ is a bowl: every flat point is a lowest point (and there is exactly one when $X^\top X$ is invertible). For gradient descent the update is $\mathbf{w} \leftarrow \mathbf{w} - \eta\,\frac2n X^\top(X\mathbf{w}-\mathbf{y})$: "move the weights against (the inputs weighted by the errors)".

Why do we need it?

We want the best line, plane or hyperplane through data. Calculus gives a recipe that works for any number of features: take the gradient of the loss, then either walk down it or set it to zero.

Where is it used?

Predicting prices and demand, calibrating sensors, the last layer of many networks (a linear layer fitted by MSE), the "baseline" model for almost any regression problem, and the building block of ridge regression.

How is it used?

For small $d$, solve the normal equations (np.linalg.solve or lstsq). For huge data, run gradient descent with the gradient $\frac2n X^\top(X\mathbf{w}-\mathbf{y})$. Always check: after fitting, $X^\top\mathbf{r}$ should be almost zero.

Drag the six blue points to reshape the cloud. Use the sliders to choose a line (orange) and see its loss and its gradient. Press Jump to closed form to jump to the green line where the gradient is exactly $[0, 0]$. Press Gradient step a few times and watch the orange line crawl toward green. Check that the readout's finite-difference slope matches the formula.

Rotate the surface. Its height is the loss (scaled to fit) above each pair (intercept $b$, slope $w$) for the data $(1,2),(2,3),(3,7)$. Drag the orange start point on the surface, then press Run 60 steps to see gradient descent (red path). With raw $x$ the valley is long and thin and the path crawls; switch on Centre x (use $x-2$) and the bowl becomes round. Look from Top to see the path as a map. The green dot is the closed-form answer.

Gradient zero ≠ always one answer. If two columns of $X$ are copies of each other, $X^\top X$ is singular and the bowl has a flat-bottomed trough: infinitely many best lines. Gradient descent still works (it lands somewhere on the trough). Ridge regularisation (below) fixes it.

Inverting is not how you solve it. Libraries solve $X^\top X\mathbf{w} = X^\top\mathbf{y}$ (or use QR / SVD) instead of forming $(X^\top X)^{-1}$: it is faster and more accurate. See Chapter 1.10.

Quick check: with one feature and no intercept, show that $\nabla L = 0$ gives $w = \sum x_iy_i/\sum x_i^2$.

Here $X$ is a single column, so $X^\top X = \sum x_i^2$ and $X^\top\mathbf{y} = \sum x_iy_i$. The normal equation $\big(\sum x_i^2\big)w = \sum x_iy_i$ gives the formula. With our data: $29/14 \approx 2.07$.

The sigmoid: turning a score into a probability core

What we need from earlier chapters: the sigmoid function and exponentials (Chapter 2.1), and the chain rule and derivative of $e^x$ (Chapters 2.3 and 2.8).

A model that predicts a yes/no outcome first computes a score $z$ that can be any number, positive or negative. But a probability must sit between 0 and 1. The sigmoid is a smooth S-shaped curve that squeezes every score into that range: a very negative score becomes almost 0, a very positive score becomes almost 1, and a score of exactly 0 becomes $\tfrac12$ ("no idea").

The slope of the S tells us how much the probability reacts to a small change in the score. In the middle the curve is steep: the model is unsure, so evidence matters. At the far ends the curve is almost flat: the model is already confident, and extra evidence barely moves it. That flatness is called saturation, and it will matter a lot in this chapter.

Let $z = 2$. Then $\sigma(2) = \dfrac{1}{1+e^{-2}} = \dfrac{1}{1+0.1353} \approx 0.8808$.

Derive the slope. Write $\sigma(z) = (1+e^{-z})^{-1}$.

  1. Chain rule with outer power $u^{-1}$ and inner $u = 1+e^{-z}$: $\sigma'(z) = -(1+e^{-z})^{-2}\cdot\dfrac{d}{dz}(1+e^{-z}) = -(1+e^{-z})^{-2}\cdot(-e^{-z}) = \dfrac{e^{-z}}{(1+e^{-z})^2}$.
  2. Split the fraction into two factors: $\dfrac{e^{-z}}{(1+e^{-z})^2} = \dfrac{1}{1+e^{-z}}\cdot\dfrac{e^{-z}}{1+e^{-z}}$.
  3. The first factor is $\sigma(z)$. The second is $1 - \sigma(z)$, because $1 - \dfrac{1}{1+e^{-z}} = \dfrac{e^{-z}}{1+e^{-z}}$.
  4. So $\boxed{\sigma'(z) = \sigma(z)\,\big(1-\sigma(z)\big)}$. The slope can be computed from the output alone.
  5. At $z=2$: $\sigma' = 0.8808\times 0.1192 \approx 0.1050$. Numeric check: $\big(\sigma(2.001)-\sigma(1.999)\big)/0.002 = 0.1050$ ✓.
$$\sigma(z) = \frac{1}{1+e^{-z}},\qquad \sigma'(z) = \sigma(z)\big(1-\sigma(z)\big),\qquad \sigma(-z) = 1-\sigma(z).$$
  • Output is in $(0, 1)$; $\sigma(0) = \tfrac12$.
  • The slope is biggest at $z=0$, where $\sigma' = \tfrac12\cdot\tfrac12 = \tfrac14$, and it is never larger than $\tfrac14$. It tends to 0 as $z\to\pm\infty$ (saturation).
  • The inverse is the logit (log-odds): $z = \ln\dfrac{p}{1-p}$.
  • For a vector of scores, apply $\sigma$ to each entry; the Jacobian is diagonal, $\mathrm{diag}(\sigma(1-\sigma))$ (Chapter 2.5).
Why do we need it?

Models compute unbounded scores, but probabilities live between 0 and 1. The sigmoid is the smooth bridge, and its simple derivative $\sigma(1-\sigma)$ keeps backpropagation cheap.

Where is it used?

The output of logistic regression and of binary classifiers, gates inside LSTMs and GRUs, "is this token a yes?" heads, and (historically) hidden layers of early neural networks.

How is it used?

Compute $p=\sigma(z)$ and predict "yes" when $p\ge\frac12$ (that is, $z\ge0$). In backprop, multiply the gradient arriving at $p$ by $p(1-p)$ to get the gradient at $z$. Beware: when $p$ is near 0 or 1 this factor is tiny.

Drag the purple dot along the S-curve. The orange tangent is the slope $\sigma'$, plotted on the right as a bell. Put the dot at $z=0$: the slope is its largest, $0.25$. Drag far to either end: the slope collapses toward zero. That is saturation: the model barely reacts to the score there.

Vanishing gradients. In a deep network with sigmoids, the backward pass multiplies by $\sigma'\le0.25$ at every layer. The product shrinks fast, so early layers barely learn. This is one reason ReLU replaced the sigmoid in hidden layers (see vanishing gradients in Chapter 2.9). The sigmoid is still the right choice for a final yes/no probability.

Cousin. $\tanh z = 2\sigma(2z) - 1$ has range $(-1,1)$ and derivative $1-\tanh^2 z$ (maximum 1, so less shrinking).

Quick check: what is $\sigma'(0)$, and what is $\sigma'(5)$ roughly?

$\sigma'(0) = 0.5\times0.5 = 0.25$. At $z=5$, $\sigma\approx0.9933$, so $\sigma' \approx 0.9933\times0.0067 \approx 0.0066$: about 38 times smaller.

Logistic regression: the gradient and its beautiful cancellation core

What we need from earlier chapters: the chain rule (Chapter 2.8), the sigmoid and $\sigma'=\sigma(1-\sigma)$ (the previous section), the gradient of a dot product (Chapter 2.6) and, from the Linear Algebra guide, logistic regression and its decision boundary.

Now the answer is yes or no: is this email spam? We compute a score $z = \mathbf{w}^\top\mathbf{x}$ (a weighted sum of the features, with the bias folded in as a weight on a constant 1), squeeze it into a probability $p = \sigma(z)$, and predict spam when $p \ge \frac12$. The place where $z = 0$ is the decision boundary: a line in 2D.

How do we score a guess? By the probability the model gave to the true answer. If the email really is spam, a good model gave $p$ close to 1. We turn that into a loss by taking $-\ln$ of it: $-\ln(\text{probability of the truth})$ is near 0 when the model was right and confident, and huge when it was confident and wrong. (The next section shows where this loss comes from.)

The gradient has a lovely surprise: after the chain rule, almost everything cancels, and what is left is "prediction minus truth, times the input", the same shape as in linear regression.

Two examples with one feature plus a bias: $x = -1$ has label $y=0$ and $x = 1$ has label $y=1$. So

$$X=\begin{bmatrix}1&-1\\1&1\end{bmatrix},\quad \mathbf{y}=\begin{bmatrix}0\\1\end{bmatrix},\quad \mathbf{w} = \begin{bmatrix}0\\0\end{bmatrix}\ (\text{bias},\ \text{weight}).$$
  1. Scores $X\mathbf{w} = [0, 0]$, probabilities $\sigma(0) = [0.5, 0.5]$.
  2. Loss: $-\ln(1-0.5)$ for the first, $-\ln 0.5$ for the second; the mean is $\ln 2 \approx 0.693$.
  3. Prediction minus truth: $\mathbf{p}-\mathbf{y} = [0.5,\ -0.5]$.
  4. $X^\top(\mathbf{p}-\mathbf{y}) = [\,0.5 - 0.5,\ \ (-1)(0.5) + (1)(-0.5)\,] = [0, -1]$. Divide by $n=2$: $\nabla L = [0, -0.5]$.
  5. Step with $\eta = 1$: $\mathbf{w}\leftarrow [0, 0.5]$. New probabilities $[\sigma(-0.5), \sigma(0.5)] = [0.378, 0.622]$, new loss $\approx 0.474$. The model now leans the right way for both points.

Model and loss for one example (label $y\in\{0,1\}$):

$$z = \mathbf{w}^\top\mathbf{x},\qquad p = \sigma(z),\qquad \ell = -\,y\ln p-(1-y)\ln(1-p).$$

Derivation, four small derivatives multiplied (the chain rule):

  1. Loss with respect to $p$: $\dfrac{\partial\ell}{\partial p} = -\dfrac{y}{p}+\dfrac{1-y}{1-p} = \dfrac{-y(1-p) + (1-y)p}{p(1-p)} = \dfrac{p-y}{p(1-p)}$.
  2. Probability with respect to score: $\dfrac{\partial p}{\partial z} = \sigma'(z) = p(1-p)$.
  3. Score with respect to weights: $\dfrac{\partial z}{\partial\mathbf{w}} = \mathbf{x}$.
  4. Multiply: $\dfrac{\partial\ell}{\partial\mathbf{w}} = \dfrac{p-y}{\cancel{p(1-p)}}\cdot\cancel{p(1-p)}\cdot\mathbf{x} = (p-y)\,\mathbf{x}$. The awkward $p(1-p)$ cancels exactly.

Average over $n$ examples and stack into matrix form (the gradient is a column):

$$\boxed{\ \nabla L(\mathbf{w}) = \frac1n X^\top\big(\sigma(X\mathbf{w}) - \mathbf{y}\big)\ }$$
  • Compare linear regression: $\frac2n X^\top(X\mathbf{w}-\mathbf{y})$. Same pattern: inputs times (prediction − target). Only the prediction changed from $X\mathbf{w}$ to $\sigma(X\mathbf{w})$.
  • The Hessian is $\frac1n X^\top\,\mathrm{diag}\big(p_i(1-p_i)\big)\,X$, positive semi-definite, so the loss is a convex bowl: one valley, no false minima.
  • Setting the gradient to zero gives equations with $\sigma$ inside them. There is no closed form, so we use gradient descent (or Newton's method, Chapter 2.10).
Why do we need it?

Many problems are yes/no, and a straight line is a bad fit for 0/1 answers. Logistic regression gives a probability and a clean gradient, so a classifier can be trained with plain gradient descent.

Where is it used?

Spam filters, credit and medical risk scores, click-through prediction in online ads, and the final layer of every binary neural-network classifier (where the same $\hat p - y$ appears in backprop).

How is it used?

Compute $\mathbf{p} = \sigma(X\mathbf{w})$, then the gradient $\frac1nX^\top(\mathbf{p}-\mathbf{y})$, then step. Watch the loss go down and the boundary rotate. Scikit-learn and PyTorch do exactly this internally.

Drag the round handle on the boundary to slide the line, and the arrow tip to rotate it or make it steeper (a longer arrow means a more confident model). The background shows $p=\sigma(\mathbf{w}^\top\mathbf{x})$: blue means "class 1", orange means "class 0". Points circled in red are on the wrong side. Press Run 100 steps and watch the loss curve fall while the boundary settles.

Perfectly separable data. If a line can split the classes with no mistakes, the loss keeps shrinking as the weights grow (the model gets ever more confident), so the "best" weights are infinite. In practice we stop early or add a regulariser (see the regularisation section).

Forgetting the bias. Without the constant-1 column the boundary must pass through the origin.

Quick check: one example $\mathbf{x} = [1, 3]$ (bias, feature) with $y=1$ and current $p = 0.8$. What is its gradient $\partial\ell/\partial\mathbf{w}$?

$(p - y)\,\mathbf{x} = (0.8-1)[1, 3] = [-0.2, -0.6]$. Both entries are negative, so gradient descent will increase both weights, raising the score and so the probability of class 1.

Cross-entropy: where it comes from and why it beats MSE for classifying core

What we need from earlier chapters: the logarithm and its properties (Chapter 2.1: $\ln(ab)=\ln a+\ln b$), the derivative of $\ln x$ (Chapter 2.3, $1/x$), and the logistic-regression gradient from the previous section.

Think of the loss as surprise. If the model says "99% sure it is a cat" and it is a cat, you are not surprised: tiny loss. If it says "1% cat" and it is a cat, you are very surprised: huge loss. The right way to measure surprise is $-\ln(\text{probability given to what really happened})$.

Why this and not squared error? Squared error says a wrong answer costs at most $1$ (since $p$ is between 0 and 1). So a confidently wrong model gets only a mild loss and a tiny gradient, and it stays stuck. Cross-entropy says a confidently wrong answer costs a lot, and gives a big gradient to fix it.

From likelihood to loss. Three predictions $p = [0.9, 0.2, 0.6]$ (probability of class 1) with true labels $y = [1, 0, 1]$. The probability the model gave to what actually happened: $0.9,\ 1-0.2 = 0.8,\ 0.6$.

  1. Likelihood (probability of all three outcomes, if examples are independent): $0.9\times0.8\times0.6 = 0.432$. Training should maximise this.
  2. Take the log, so the product becomes a sum (sums are easier to differentiate, and do not underflow): $\ln 0.432 = \ln0.9+\ln0.8+\ln0.6 = -0.105-0.223-0.511 = -0.839$.
  3. Flip the sign (we minimise losses) and average: $\dfrac{0.105+0.223+0.511}{3} \approx 0.280$. That is the binary cross-entropy.

Punishing confident mistakes (true label is 1; the model says $p$ for class 1):

$p$ given to the truth0.90.50.10.010.001
cross-entropy $-\ln p$0.1050.6932.3034.6056.908
squared error $(1-p)^2$0.010.250.810.9800.998

Cross-entropy keeps growing without limit. Squared error flattens out near 1.

Derivation. A model that outputs $p_i = P(y_i=1\mid\mathbf{x}_i)$ assigns probability $p_i^{\,y_i}(1-p_i)^{1-y_i}$ to the label that happened (for $y=1$ this is $p$, for $y=0$ it is $1-p$). The likelihood of all the data is the product of these. Taking $\ln$ (which keeps the same maximiser, because $\ln$ only goes up) and flipping the sign gives

$$L = -\frac1n\sum_i\Big[y_i\ln p_i + (1-y_i)\ln(1-p_i)\Big]\quad(\text{binary cross-entropy}),\qquad \ell = -\sum_{k}y_k\ln p_k\quad(K\text{ classes, one-hot } \mathbf{y}).$$

("One-hot" means $\mathbf{y}$ has a 1 for the true class and 0 for every other class, for example $[0,1,0]$.)

For one-hot labels the $K$-class loss is just $-\ln p_{\text{true class}}$. (Name: $-\sum_k q_k\ln p_k$ is the cross-entropy of the model's distribution $\mathbf{p}$ against the true distribution $\mathbf{q}$. Minimising it makes $\mathbf{p}$ match $\mathbf{q}$.)

A neat form in terms of the score $z$ (this is what BCEWithLogitsLoss computes, and it is numerically safe): since $-\ln\sigma(z) = \ln(1+e^{-z})$ and $-\ln(1-\sigma(z)) = \ln(1+e^{z})$,

$$\ell(z, y) = \ln(1+e^{z}) - y\,z\qquad\Longrightarrow\qquad \frac{\partial\ell}{\partial z} = \frac{e^z}{1+e^z} - y = \sigma(z) - y.$$

Gradient sizes for $y = 1$. Squared error through a sigmoid: $\ell=(p-y)^2$, so $\dfrac{\partial\ell}{\partial z} = 2(p-y)\cdot\sigma'(z) = 2(p-y)\,p(1-p)$. The extra factor $p(1-p)$ is the culprit: it is almost 0 whenever the model is very sure, including when it is very wrong. Cross-entropy's gradient is just $p-y$, with no such factor.

score $z$ (true label 1)$-6$ (very wrong)$-3$$0$$3$$6$ (very right)
cross-entropy gradient $p-1$$-0.9975$$-0.9526$$-0.5$$-0.0474$$-0.0025$
MSE gradient $2(p-1)p(1-p)$$-0.0049$$-0.0861$$-0.25$$-0.0043$$-0.00001$

At $z=-6$ the cross-entropy push is about 200 times bigger. MSE saturates: it gives almost no signal exactly where a fix is most needed.

Why do we need it?

A classifier outputs probabilities, so we need a loss that measures how well probabilities match reality, and that pushes hard on confident mistakes. The negative log-likelihood does both and has the clean gradient $p-y$.

Where is it used?

Almost every classifier: logistic regression, image classifiers, and language models (next-token prediction is cross-entropy over the vocabulary). "Perplexity" is $e^{\text{cross-entropy}}$.

How is it used?

Use the built-in cross_entropy / BCEWithLogitsLoss on raw scores (logits), not on probabilities you computed yourself: it is faster and avoids $\ln 0$. Report the average loss per example.

The true label is 1 (switch to 0 if you like). Drag the purple dot along the score axis. On the right, the blue line is the cross-entropy gradient and the orange line is the MSE gradient. Drag the score to $-6$ (confidently wrong): cross-entropy still shouts "fix me" (about 1) while MSE whispers (about 0.005). Both agree when the model is right.

One sigmoid unit $p=\sigma(wx+b)$ on four points ($x=-2,-1$ labelled 0 and $x=1,2$ labelled 1). The surface is the loss over $(b, w)$ (height scaled to fit). The orange start point begins at $w=-4$: the model is confidently backwards. Press Run 30 steps with cross-entropy: the red path slides down the bowl. Switch to MSE and run again: in 30 steps the path barely moves, because the surface is a flat plateau there (press Run 100 steps and it does escape, but very slowly). Rotate and look from Side to see the plateau.

Never feed probabilities you computed with sigmoid/softmax into a log yourself. If $p$ rounds to exactly 0 you get $\ln 0 = -\infty$. Use the combined "with logits" functions, which use the stable form $\ln(1+e^z)-yz$.

MSE is not "wrong" for regression, where the output is not squeezed by a sigmoid. The saturation problem is specific to bounded outputs.

Quick check: the true class gets probability $0.25$. What is the cross-entropy, and what is it if the model had given $0.5$?

$-\ln0.25 = \ln4 \approx 1.386$. With $0.5$: $-\ln 0.5 = \ln 2 \approx 0.693$. Halving the probability doubled the loss here.

Gradient descent: convergence, the learning rate, and momentum core

What we need from earlier chapters: the gradient (Chapter 2.4), the second derivative and the Hessian as curvature (Chapter 2.10), linearization (Chapter 2.12) and, from the Linear Algebra guide, eigenvalues and gradient descent on the loss bowl.

Imagine walking down a foggy hill. You cannot see the valley, but you can feel the slope under your boots. The rule is: step in the direction that feels steepest downhill, then feel again. The only choice is how long each step is. Tiny steps are safe but you will be there all day. Giant steps are fast, but you may jump clean across the valley and land higher up the other side, or even fly off the map.

On a bowl-shaped valley the "just right" step lands you near the bottom quickly. The steepness of the bowl (its curvature, from the second derivative) decides what "too big" means: a steep, narrow bowl needs small steps. That is why the learning rate is the most important setting in training.

The simplest bowl: $L(w) = w^2$, with slope $L'(w) = 2w$ and minimum at $w=0$. Start at $w_0 = 4$. The update is $w \leftarrow w - \eta\cdot 2w = (1-2\eta)\,w$.

$\eta$multiplier $1-2\eta$$w_0,\ w_1,\ w_2,\ w_3$what happens
0.10.84, 3.2, 2.56, 2.05smooth but slow
0.504, 0, 0, 0perfect: lands on the minimum in one step
0.9$-0.8$4, $-3.2$, 2.56, $-2.05$jumps over the bottom each time, but shrinks
1.0$-1$4, $-4$, 4, $-4$bounces forever, no progress
1.1$-1.2$4, $-4.8$, 5.76, $-6.91$diverges: each step is bigger than the last

Everything is decided by that one multiplier.

Update rule: $\ \mathbf{w}_{t+1} = \mathbf{w}_t - \eta\,\nabla L(\mathbf{w}_t)$.

Convergence on a quadratic (derived). Take $L(w) = \tfrac12 a\,(w-w^\star)^2$ with curvature $a>0$ (the second derivative). Then $L'(w) = a(w-w^\star)$ and

$$w_{t+1}-w^\star = w_t - w^\star - \eta\,a\,(w_t - w^\star) = (1-\eta a)\,(w_t-w^\star)\ \Longrightarrow\ w_t - w^\star = (1-\eta a)^t\,(w_0-w^\star).$$

So the error shrinks to 0 exactly when $|1-\eta a| < 1$, that is, $0 < \eta < 2/a$. The best step is $\eta = 1/a$ (multiplier 0). A multiplier in $(0,1)$ gives a smooth approach, in $(-1,0)$ a zig-zag that still converges, exactly $-1$ bounces forever, and below $-1$ a blow-up.

Many weights. For $L=\tfrac12(\mathbf{w}-\mathbf{w}^\star)^\top H(\mathbf{w}-\mathbf{w}^\star)$ with Hessian $H$, the same argument works along each eigenvector of $H$ with $a=\lambda_i$. Stability needs $\eta < 2/\lambda_{\max}$, but the slowest direction (smallest $\lambda_{\min}$) shrinks by only $1-\eta\lambda_{\min}$ per step. The ratio $\kappa=\lambda_{\max}/\lambda_{\min}$ (the condition number) says how long and thin the valley is, and so how slow descent will be.

Momentum (awareness: a standard extra). Keep a "velocity" $\mathbf{v}$ that remembers past gradients:

$$\mathbf{v}\leftarrow\beta\,\mathbf{v}+\nabla L(\mathbf{w}),\qquad \mathbf{w}\leftarrow\mathbf{w}-\eta\,\mathbf{v},\qquad \beta\in[0,1)\ (\text{often }0.9).$$

Unrolled, $\mathbf{v}_t = \mathbf{g}_t+\beta\mathbf{g}_{t-1}+\beta^2\mathbf{g}_{t-2}+\dots$, a running average of recent gradients. Where gradients agree (down the valley floor) they add up and speed us along: a steady gradient $\mathbf{g}$ gives $\mathbf{v}\to\mathbf{g}/(1-\beta)$, ten times larger for $\beta=0.9$. Where they flip sign (across the valley) they cancel and the zig-zag fades. Related ideas: Nesterov momentum (look ahead before taking the gradient) and Adam (also rescales each weight by a running size of its gradient).

Why do we need it?

For almost every model there is no formula for the best weights. Gradient descent only needs the slope, so it works for any smooth loss. But it needs a good step size, or it crawls or explodes.

Where is it used?

Training every neural network. Plain gradient descent for convex models (logistic regression), SGD with momentum for image models, Adam and AdamW for Transformers and language models.

How is it used?

Try learning rates on a log scale (0.1, 0.01, 0.001) and watch the loss curve. If the loss rises or jumps wildly, $\eta$ is too big. If it creeps, it is too small. Add momentum 0.9, and decay $\eta$ later in training.

The bowl is $L(w)=w^2$, starting at $w_0=4$ (the same numbers as the table). Press the four presets, or drag the slider through $\eta=0.5$ (perfect) up to $1.0$ (bounces forever) and beyond. On the right, watch $w$ against the step number: a smooth decay, a straight drop to zero, a shrinking zig-zag, or a growing zig-zag.

The valley is long in $x$ and steep in $y$. Orange: plain gradient descent. Blue: with momentum $\beta$. Both start at the purple point (drag it) and take 40 steps with the same $\eta$. Raise $\beta$ from 0 to 0.8: the blue path gets far further along the valley floor (past about 0.9 the heavy ball overshoots and swings). Make the valley steeper with $\kappa$ and raise $\eta$ until orange starts zig-zagging.

The best learning rate changes during training. Early on a large step makes quick progress. Near the bottom the same step makes you bounce around the minimum, so many training runs decay the learning rate (see the batch-optimisation section).

Too large a learning rate shows up as a loss that goes UP or turns into NaN. That is the "diverges" row of the table, not a bug in your gradients.

Quick check: $L(w)=3(w-1)^2$. Which learning rates converge, and which is perfect?

Here $L'(w)=6(w-1)$, so the curvature is $a=6$ and the multiplier is $1-6\eta$. It converges when $|1-6\eta|<1$, i.e. $0<\eta<2/6 \approx 0.333$. The perfect step is $\eta=1/6\approx0.167$.

Regularisation: how a penalty changes the gradient core

What we need from earlier chapters: the derivative of $x^2$ and of $|x|$ (and its kink) (Chapters 2.2 and 2.3), the gradient and level sets (Chapter 2.4) and, from the Linear Algebra guide, the L1 and L2 norms and ridge and lasso.

A model with a lot of freedom can fit the training data perfectly by using huge weights that cancel each other out, and then fail on new data. Regularisation adds a price for large weights to the loss: "fit the data, but keep the weights small if you can".

In the gradient this price is a constant pull toward zero, in addition to the usual pull toward fitting the data. There are two flavours, and the difference is in the shape of the pull:

  • L2 (ridge): the pull is proportional to the weight. A big weight is pulled hard, a small weight only gently. So weights shrink toward zero but almost never reach it.
  • L1 (lasso): the pull has a fixed size, whatever the weight. Even a tiny weight keeps getting pushed, so it can be driven all the way to exactly zero. Those zero weights switch features off: a sparse model.

L2 as "weight decay". One weight $w = 2$, $\lambda = 0.1$, $\eta = 0.1$, and the data gradient is $\nabla L = 0.5$. The penalty $\lambda w^2$ has slope $2\lambda w = 2(0.1)(2) = 0.4$.

  1. Total gradient: $0.5 + 0.4 = 0.9$.
  2. Update: $w \leftarrow 2 - 0.1\times0.9 = 1.91$.
  3. Same thing, rearranged: $w\leftarrow(1-2\eta\lambda)\,w - \eta\nabla L = 0.98\times 2 - 0.05 = 1.91$ ✓. Each step first shrinks the weight by 2% ("decay"), then applies the data gradient.

L1 in one dimension. Minimise $\tfrac12(w-a)^2+\lambda|w|$, where $a$ is the weight the data alone would choose.

  1. If $w>0$, the slope is $(w-a)+\lambda$. Setting it to 0 gives $w = a-\lambda$, which is valid only if $a>\lambda$.
  2. If $w<0$, the slope is $(w-a)-\lambda$, giving $w=a+\lambda$, valid only if $a<-\lambda$.
  3. Otherwise the minimum sits at the kink $w=0$: just left of 0 the slope is $-a-\lambda\le0$ and just right it is $-a+\lambda\ge0$ when $|a|\le\lambda$.

So $w^\star = \operatorname{sign}(a)\max(|a|-\lambda,\,0)$ (the soft threshold). With $\lambda=0.5$: $a=0.8\Rightarrow w^\star=0.3$, and $a=0.4\Rightarrow w^\star=0$ exactly. Ridge with the same $a=0.8$, minimising $\tfrac12(w-a)^2+\lambda w^2$: $w^\star = a/(1+2\lambda) = 0.4$, never exactly 0.

$$L_{\text{ridge}}(\mathbf{w}) = L(\mathbf{w}) + \lambda\|\mathbf{w}\|_2^2,\qquad \nabla L_{\text{ridge}} = \nabla L + 2\lambda\mathbf{w}.$$ $$L_{\text{lasso}}(\mathbf{w}) = L(\mathbf{w}) + \lambda\|\mathbf{w}\|_1,\qquad \partial L_{\text{lasso}}/\partial w_j = \partial L/\partial w_j + \lambda\,\operatorname{sign}(w_j)\ \ (w_j\ne0).$$
  • Ridge gradient step: $\mathbf{w}\leftarrow(1-2\eta\lambda)\mathbf{w}-\eta\nabla L$, which is why L2 regularisation is called weight decay.
  • Ridge for linear regression, closed form: from $\frac2nX^\top(X\mathbf{w}-\mathbf{y})+2\lambda\mathbf{w}=\mathbf{0}$ we get $(X^\top X+n\lambda I)\mathbf{w}=X^\top\mathbf{y}$. The matrix $X^\top X+n\lambda I$ is always invertible (every eigenvalue is at least $n\lambda>0$), which cures the flat trough of copied features. See Chapter 1.10.
  • Lasso: $|w|$ has a kink at 0, where the derivative does not exist (Chapter 2.2: the left and right slopes, $-1$ and $+1$, disagree). Any slope in $[-1,1]$ is allowed there (a subgradient). A weight sits at exactly 0 when the data's pull on it is smaller than $\lambda$.
  • The picture. Minimising $L$ subject to $\|\mathbf{w}\|\le t$ is (for a matching $t$) the same as the penalised problem. The loss contours (ellipses) grow until they first touch the constraint region. A round L2 ball is touched at a generic point; an L1 diamond has corners on the axes, and the ellipses usually hit a corner first, where a weight is exactly 0.

Usually the bias is not penalised. (Awareness: L2 equals a Gaussian prior on the weights, L1 a Laplace prior; AdamW applies the decay separately from the gradient.)

Why do we need it?

With few examples, many features, or copied features, the plain loss has wild or non-unique solutions that fit noise. A penalty keeps the weights small and the answer stable, which usually generalises better.

Where is it used?

Ridge regression, lasso for choosing features (genomics, finance), weight decay in nearly every deep-learning optimiser (SGD, AdamW), and the L2 term in logistic regression (C in scikit-learn).

How is it used?

Add $\lambda\|\mathbf{w}\|^2$ or $\lambda\|\mathbf{w}\|_1$ to the loss (or set weight_decay in the optimiser). Pick $\lambda$ by checking performance on held-out data. Larger $\lambda$ means smaller weights and a simpler model.

Drag the purple dot along the bottom of the right picture: it is $a$, the weight the data alone would choose. The left picture shows the total loss for ridge (blue) and lasso (orange); the dots mark their minima. Slide $\lambda$ up. Ridge slides smoothly toward 0 but never lands. Lasso's minimum sticks to the kink at 0 whenever $|a|\le\lambda$. On the right, the orange curve has a flat piece at 0: that is sparsity.

The blue dot is the unconstrained best weights. The purple region is the allowed set $\|\mathbf{w}\|\le t$ (diamond for L1, circle for L2). The green dot is the best point inside it, where the red loss contour just touches the region. Shrink $t$ and drag the blue dot into different places: with the diamond the green dot often sits exactly on an axis, so one weight is 0. With the circle it almost never does.

The surface is a data loss plus a penalty, over two weights. Choose none, L2 or L1 and slide $\lambda$: the green minimum moves toward the origin. Drag the orange start point and press Run 80 steps. With L1 and $\lambda$ about 1.5 or more, the green minimum sits exactly on the $w_2 = 0$ line (the readout says so). Look from Top: the path is pulled onto that line and then chatters around the kink. Rotate to see the sharp crease that L1 adds along each axis.

Standardise features before regularising. The penalty treats all weights alike. If one feature is measured in thousands, its weight is tiny and escapes the penalty, so features on different scales get unfair treatment.

L1 and gradient descent. Plain gradient descent on an L1 loss chatters around 0 because the slope flips sign at the kink. Real lasso solvers use soft-thresholding (the formula derived above) after every step, an idea called the proximal step.

Quick check: $w=-3$, $\lambda=0.05$, $\eta=0.2$, data gradient $\nabla L=1$. One L2 step?

Penalty slope $2\lambda w = 2(0.05)(-3) = -0.3$, so the total gradient is $1-0.3 = 0.7$ and $w\leftarrow -3 - 0.2(0.7) = -3.14$. Without the penalty the step would have gone to $-3.2$. The penalty held the weight back, so it ended closer to zero. Check with the decay form: $1-2\eta\lambda = 0.98$, and $0.98\cdot(-3) - 0.2\cdot 1 = -3.14$ ✓.

Softmax and the gradient $\mathbf{p}-\mathbf{y}$ core

What we need from earlier chapters: the softmax function (Chapter 2.1), the Jacobian matrix (Chapter 2.5), the chain rule with vectors (Chapter 2.8) and, from the Linear Algebra guide, softmax regression and the dot product.

With more than two classes (cat, dog, bird) the model produces one score per class, called logits. Softmax turns the list of scores into a list of probabilities that are all positive and add to 1. It does this in two moves: exponentiate every score (this makes them positive and exaggerates the winner), then divide by the total (so they sum to 1).

Because the probabilities all share one total, they compete: raising one class's score must lower the others' probabilities. That competition shows up as the minus signs in the derivative. And once again, when we combine softmax with the cross-entropy loss, the messy parts cancel and the final gradient is as simple as it can be: predicted probabilities minus the true answer.

Logits $\mathbf{z} = [2, 1, 0]$, and the true class is class 2 (so $\mathbf{y} = [0, 1, 0]$).

  1. Exponentiate: $[e^2, e^1, e^0] = [7.389,\ 2.718,\ 1]$. Their total is $11.107$.
  2. Divide: $\mathbf{p} = [0.665,\ 0.245,\ 0.090]$ (these add to 1).
  3. Loss: $-\ln p_2 = -\ln0.245 \approx 1.408$.
  4. Gradient with respect to the logits: $\mathbf{p}-\mathbf{y} = [0.665,\ -0.755,\ 0.090]$. (Derived below.) Gradient descent will raise the true class's logit (negative entry) and lower the others.
  5. Numeric check: nudge $z_2$ by $0.001$ and recompute the loss: the change per unit is $-0.7553$ ✓.
$$p_i = \frac{e^{z_i}}{\sum_k e^{z_k}}.$$

Step 1: the Jacobian. Take logs: $\ln p_i = z_i - \ln\sum_k e^{z_k}$. Differentiate with respect to $z_j$ (the derivative of the second term is $e^{z_j}/\sum_k e^{z_k} = p_j$):

$$\frac{\partial\ln p_i}{\partial z_j} = \delta_{ij} - p_j\ \Longrightarrow\ \frac{\partial p_i}{\partial z_j} = p_i\,(\delta_{ij}-p_j),\qquad J = \operatorname{diag}(\mathbf{p}) - \mathbf{p}\mathbf{p}^\top,$$

where $\delta_{ij}$ is 1 if $i=j$ and 0 otherwise. For our example $J = \begin{bmatrix}0.223&-0.163&-0.060\\-0.163&0.185&-0.022\\-0.060&-0.022&0.082\end{bmatrix}$. It is symmetric, and every row sums to zero: adding the same number to all logits changes nothing, so $J\mathbf{1}=\mathbf{0}$ (and $J$ is singular). That is also why we subtract $\max_k z_k$ before exponentiating (it avoids overflow and changes nothing).

Step 2: add the cross-entropy loss. With one-hot $\mathbf{y}$, $\ell = -\sum_k y_k\ln p_k$, so $\partial\ell/\partial p_k = -y_k/p_k$. Chain rule over all $k$:

$$\frac{\partial\ell}{\partial z_j} = \sum_k\Big(-\frac{y_k}{p_k}\Big)p_k(\delta_{kj}-p_j) = -\sum_k y_k(\delta_{kj}-p_j) = -y_j + p_j\sum_k y_k = p_j - y_j,$$

because $\sum_k y_k = 1$. The factor $p_k$ from the Jacobian cancels the $1/p_k$ from the log, exactly like $p(1-p)$ cancelled for the sigmoid.

$$\boxed{\ \nabla_{\mathbf{z}}\,\ell = \mathbf{p}-\mathbf{y}\ }\qquad\text{Next layer back: if } \mathbf{z}=W\mathbf{h}+\mathbf{b}:\ \ \frac{\partial\ell}{\partial W} = (\mathbf{p}-\mathbf{y})\,\mathbf{h}^\top,\ \ \frac{\partial\ell}{\partial\mathbf{h}} = W^\top(\mathbf{p}-\mathbf{y}).$$
  • Why so clean: $\ell = -z_c + \ln\sum_k e^{z_k}$ for true class $c$. The first part has slope $-1$ on $z_c$ only; the second is the "soft maximum" whose slope on $z_j$ is its own probability $p_j$.
  • The sigmoid is the 2-class case: $\sigma(z) = \operatorname{softmax}([z, 0])_1$, and $\sigma(z)-y$ is the same formula.
  • Temperature (awareness): $\operatorname{softmax}(\mathbf{z}/T)$. Large $T$ gives nearly equal probabilities, small $T$ gives nearly one-hot.
Why do we need it?

A classifier with many classes needs probabilities that are positive and sum to 1. Softmax provides them, and paired with cross-entropy it has the simplest possible gradient, $\mathbf{p}-\mathbf{y}$, so training is stable and fast.

Where is it used?

The last layer of image classifiers, the next-word prediction layer of every language model (a softmax over tens of thousands of words), and attention weights in Transformers (a softmax over positions).

How is it used?

Feed raw logits to cross_entropy (it applies a stable log-softmax inside). In the backward pass, the gradient at the logits is just softmax(logits) - one_hot(label), then it flows back through the layer by $W^\top$.

Move the three logit sliders and choose the true class. The top bars are the probabilities $\mathbf{p}$; the lower bars are the gradient $\mathbf{p}-\mathbf{y}$ ("raise" for negative entries, "lower" for positive ones). Make the true class's logit small (a confident mistake): its gradient entry goes to about $-1$. Press One gradient step on the logits to see the loss drop. The matrix is the Jacobian $\operatorname{diag}(\mathbf{p})-\mathbf{p}\mathbf{p}^\top$: check that each row adds to 0.

Softmax is shift-invariant but not scale-invariant. Adding 100 to every logit changes nothing; doubling every logit makes the output more confident (it is the temperature).

Do not apply softmax and then $\ln$ yourself. A probability can underflow to 0 and give $-\infty$. Use log_softmax (it computes $z_i-\ln\sum e^{z_k}$ stably) or cross_entropy on logits.

Quick check: logits $[0, 0]$ (two classes), true class 1. What are $\mathbf{p}$, the loss and the gradient?

$\mathbf{p} = [0.5, 0.5]$. Loss $= -\ln0.5 = \ln 2 \approx 0.693$. Gradient $\mathbf{p}-\mathbf{y} = [0.5-1,\ 0.5-0] = [-0.5, 0.5]$: raise logit 1, lower logit 2.

The neural-network loss: the chain rule through the layers core

What we need from earlier chapters: the chain rule and computational graphs (Chapter 2.8), backpropagation and automatic differentiation (Chapter 2.9: this section is a worked recap), the Jacobian (Chapter 2.5) and, from the Linear Algebra guide, backprop as Jacobian products.

A neural network is a chain of simple steps: multiply by a weight matrix, add a bias, bend with a simple function (like ReLU), and repeat. The loss at the end is a single number that depends on every weight in every layer at once. A big network has millions or billions of weights, so its loss is a landscape in that many dimensions.

To train it we need the slope of the loss with respect to each of those weights. The chain rule gives it: change a weight early in the network, and the change travels forward through the later layers to the loss. So the slope of the loss with respect to an early weight is a product of the local slopes along the path. Backpropagation is simply the chain rule, organised so each local slope is computed once and reused. We start from the loss and walk backwards, carrying "how much does the loss care?" with us.

Tiny network: 2 inputs, 2 hidden ReLU units, 1 sigmoid output, cross-entropy loss. Input $\mathbf{x}=[1,2]$, label $y=1$. Weights: $W_1=\begin{bmatrix}1&0.5\\-0.5&1\end{bmatrix}$, $\mathbf{b}_1=\mathbf{0}$, $\mathbf{w}_2=[0.5,\,-1]$, $b_2=0.25$.

Forward (compute and remember):

  1. $\mathbf{z}_1 = W_1\mathbf{x}+\mathbf{b}_1 = [1\cdot1+0.5\cdot2,\ -0.5\cdot1+1\cdot2] = [2,\ 1.5]$.
  2. $\mathbf{h} = \mathrm{ReLU}(\mathbf{z}_1) = [2,\ 1.5]$ (both positive, so unchanged).
  3. $z_{\text{out}} = \mathbf{w}_2\cdot\mathbf{h}+b_2 = 0.5\cdot2 - 1\cdot1.5 + 0.25 = -0.25$, so $p=\sigma(-0.25) = 0.4378$.
  4. Loss: $-\ln p = 0.826$ (the model gave only 44% to the truth).

Backward (the arrows reversed, each step is one local derivative):

  1. At the output: $\delta_{\text{out}} = \partial L/\partial z_{\text{out}} = p - y = -0.5622$ (the sigmoid + cross-entropy cancellation again).
  2. Output weights: $\partial L/\partial\mathbf{w}_2 = \delta_{\text{out}}\,\mathbf{h} = [-1.1244,\ -0.8433]$ and $\partial L/\partial b_2 = -0.5622$.
  3. Into the hidden layer: $\partial L/\partial\mathbf{h} = \delta_{\text{out}}\,\mathbf{w}_2 = [-0.2811,\ 0.5622]$.
  4. Through ReLU (slope 1 where $z>0$, else 0): $\boldsymbol{\delta}_1 = [-0.2811,\ 0.5622]\odot[1,1] = [-0.2811,\ 0.5622]$.
  5. First-layer weights: $\partial L/\partial W_1 = \boldsymbol{\delta}_1\mathbf{x}^\top = \begin{bmatrix}-0.2811&-0.5622\\0.5622&1.1244\end{bmatrix}$, and $\partial L/\partial\mathbf{b}_1=\boldsymbol{\delta}_1$.
  6. Numeric check of one entry: nudge $W_1[1,2]$ by $\pm10^{-6}$ and recompute the loss: the slope is $-0.5622$ ✓.

Network with layers $l=1,\dots,\mathcal{L}$: $\ \mathbf{a}_0=\mathbf{x}$, $\ \mathbf{z}_l = W_l\mathbf{a}_{l-1}+\mathbf{b}_l$, $\ \mathbf{a}_l = \varphi(\mathbf{z}_l)$. Training loss $L(\theta)=\frac1n\sum_i\ell(\mathbf{a}_{\mathcal{L}}(\mathbf{x}_i),\mathbf{y}_i)$, where $\theta$ is the list of all $W_l,\mathbf{b}_l$. Define the error signal $\boldsymbol{\delta}_l=\partial L/\partial\mathbf{z}_l$ (a column vector). By the chain rule with Jacobians:

$$\frac{\partial\mathbf{z}_{l+1}}{\partial\mathbf{z}_l} = W_{l+1}\,\operatorname{diag}\big(\varphi'(\mathbf{z}_l)\big)\ \Longrightarrow\ \boxed{\ \boldsymbol{\delta}_l = \varphi'(\mathbf{z}_l)\odot\big(W_{l+1}^\top\boldsymbol{\delta}_{l+1}\big)\ },\qquad \frac{\partial L}{\partial W_l}=\boldsymbol{\delta}_l\,\mathbf{a}_{l-1}^\top,\quad \frac{\partial L}{\partial\mathbf{b}_l}=\boldsymbol{\delta}_l.$$

$\odot$ means entry-by-entry multiplication. The recipe: forward pass stores every $\mathbf{z}_l,\mathbf{a}_l$; backward pass starts with $\boldsymbol{\delta}_{\mathcal{L}}$ (for sigmoid or softmax + cross-entropy, simply $\mathbf{p}-\mathbf{y}$) and applies the boxed rule layer by layer. The backward pass costs about as much as the forward pass, no matter how many weights there are: that is the power of reverse mode (Chapter 2.9).

  • The loss is generally not convex in $\theta$: it has valleys, ridges and saddle points (Chapter 2.10), and training finds a good valley, not a provably best one.
  • Each backward step multiplies by $\varphi'$ and by $W^\top$. With $\varphi'\le0.25$ (sigmoid) the signal shrinks layer after layer (vanishing gradients); with large weights it can grow (exploding). ReLU, good initialisation and normalisation exist for this reason.
  • Gradient checking: compare backprop with a finite difference, as in the numeric check above. It is the standard way to test hand-written backward code.
Why do we need it?

A network has millions of weights and we need the slope for every one, every step. Trying a nudge per weight would take millions of forward passes. The chain rule, run backwards, gets all of them in about one extra pass.

Where is it used?

Training every deep network: CNNs, Transformers, diffusion models. It is what loss.backward() does in PyTorch and what tf.GradientTape does in TensorFlow.

How is it used?

You write only the forward model and the loss. The framework records the computational graph, then walks it backwards multiplying local derivatives and fills a .grad for each weight. You use these gradients in the update step.

Switch between Forward (numbers flow left to right) and Backward (the error signal $\delta$ flows right to left; each edge shows its weight gradient). Change the inputs or the label and watch the numbers update. Make an input negative so a ReLU unit goes dead (its $\delta$ becomes 0). Press One update step and see the loss fall. Check the finite-difference line against backprop.

A network with 3 tanh units was trained on a toy problem (the next section). Here we cut a 2D slice through its 13-dimensional weight space: the surface is the loss as we move the weights along two random directions $(\alpha,\beta)$ away from the trained solution. It is not a simple bowl. Drag the orange point, press Run 60 steps to follow the (finite-difference) gradient downhill, and press New random slice to cut another way. Look from Top for the valley shapes.

A weight's gradient needs two things: the error signal $\boldsymbol{\delta}$ arriving at its layer's output and the activation $\mathbf{a}_{l-1}$ that went in (the outer product $\boldsymbol{\delta}\mathbf{a}^\top$). That is why the forward pass must store its values: backward needs them.

Dead ReLU units. If $z\le0$ for every example, $\varphi'=0$, so no gradient reaches that unit's weights and it never recovers.

Quick check: in the worked example, what would $\partial L/\partial W_1[1,1]$ be if the label were $y=0$ instead?

Then $\delta_{\text{out}} = p-y = 0.4378-0 = 0.4378$, $\partial L/\partial\mathbf{h} = 0.4378\cdot[0.5,-1] = [0.2189,-0.4378]$, and $\boldsymbol\delta_1 = [0.2189,-0.4378]$ (both ReLUs still on). The entry is $\delta_1[1]\cdot x_1 = 0.2189\cdot1 = 0.2189$. The sign flipped, because now we want to make $p$ smaller.

A tiny network, trained live core

What we need from earlier chapters: the previous section's backward recipe (Chapters 2.8 and 2.9), the logistic-regression gradient, and the derivative of $\tanh$ (Chapter 2.3).

Logistic regression can only draw one straight line. Some problems need more. Picture four blobs of points where opposite corners share a class (the "XOR" pattern): no single line separates the classes.

A hidden layer fixes this. Each hidden unit $\tanh(\mathbf{w}\cdot\mathbf{x}+b)$ is a soft line: it is about $-1$ on one side and $+1$ on the other. The output unit then mixes several soft lines into a more complicated shape. Training slides all the lines around, and bends the shape, until the classes are separated. You are about to watch that happen, one gradient step at a time.

The model has $H$ hidden units. It has $2H$ first-layer weights, $H$ first-layer biases, $H$ output weights and 1 output bias: $4H+1$ parameters (13 for $H=3$). The data is 48 points. Every step does the following for the whole data set:

  1. Forward: for each point compute $h_j = \tanh(W_{j1}x+W_{j2}y+b_j)$, then $z=\sum_j w_{2j}h_j+b_2$ and $p=\sigma(z)$.
  2. Loss: the average cross-entropy over the 48 points. At the start the model is nearly clueless, so the loss is about $\ln2\approx0.69$.
  3. Backward: for each point $\delta=p-y$, $\ \delta_j=\delta\,w_{2j}(1-h_j^2)$. Average the gradients over the 48 points.
  4. Update every parameter by $-\eta\times$ its average gradient.

With $H=3$ and $\eta=0.5$, about 600 steps bring the loss from about $0.7$ down to about $0.02$.

Network: $h_j=\tanh(\mathbf{W}_{j}\cdot\mathbf{x}+b_j)$, $\ z=\sum_j w_{2j}h_j+b_2$, $\ p=\sigma(z)$, loss $\ell=\ln(1+e^z)-yz$. The gradients (the chain rule, as in the last section; $\tanh'(u)=1-\tanh^2(u)$ by the quotient rule on $\frac{e^u-e^{-u}}{e^u+e^{-u}}$):

$$\delta = p-y,\quad \frac{\partial\ell}{\partial w_{2j}}=\delta\,h_j,\quad \frac{\partial\ell}{\partial b_2}=\delta,\quad \delta_j=\delta\,w_{2j}\big(1-h_j^2\big),\quad \frac{\partial\ell}{\partial\mathbf{W}_j}=\delta_j\,\mathbf{x},\quad \frac{\partial\ell}{\partial b_j}=\delta_j.$$

Training is the loop "average these over the data, subtract $\eta$ times them", repeated. The loss curve typically drops quickly, then slowly: the network first finds the coarse pattern and then fine-tunes the boundary.

Why do we need it?

Real data is rarely separated by one straight line. Hidden layers let the model build curved boundaries by combining many simple ones, and the gradient recipe still tells every weight how to move.

Where is it used?

This exact loop, at a much bigger scale, trains image classifiers, speech recognisers and language models. A Transformer has the same ingredients: linear layers, a simple non-linearity, softmax, cross-entropy and backprop.

How is it used?

Pick the number of hidden units and a learning rate, then run the loop while watching the loss curve and the decision regions. A different random start can end in a different valley, so try several.

Press Run 300 steps. The background is the probability of class 1 (blue) versus class 0 (orange); the dashed lines are the hidden units' "soft lines". Watch the loss curve fall. Try 2, 3 and 4 hidden units: more units give more flexibility. Press New random start a few times with 2 units: some starts get stuck with a worse loss. Raise $\eta$ to 1.5: the loss usually falls faster, but with some starts (try 2 hidden units) the curve gets small bumps where a step overshoots.

Different start, different valley. Because the loss is not convex, two random starts can end in different places. Small networks (2 units) can get stuck; larger ones usually find a good solution more reliably.

A falling training loss is not the whole story. We have only shown the training loss. Real projects also measure loss on held-out data to check the model generalises.

Quick check: why can a network with only linear hidden units (no tanh) not solve the XOR pattern?

A linear unit followed by a linear output is still one linear function of $\mathbf{x}$, so the decision boundary is still a single straight line. The bend in $\tanh$ is what lets the units combine into curved boundaries.

Batch optimisation: full-batch, mini-batch and stochastic gradient descent core

What we need from earlier chapters: the training loss as an average over examples (the first section of this chapter), the learning rate, and the linearity of derivatives (Chapter 2.3: the gradient of an average is the average of the gradients).

The training loss is an average over all $n$ examples, so its gradient is an average too. With a million examples, one honest gradient means looking at a million examples, and then we take just one small step. That is wasteful.

Think of a poll. To learn the country's average opinion you do not ask everyone: a random sample of a few hundred gives a good estimate. Likewise we can estimate the gradient from a small random mini-batch. The estimate is a bit noisy, but it points roughly the right way, and it is hundreds of times cheaper, so we can take many more steps in the same time. Noisy steps in the right direction beat rare perfect steps.

The estimate is unbiased: a tiny check. Four examples have one-weight gradients $g=[1, 3, 5, 7]$, so the full gradient is the mean, $\frac{1+3+5+7}{4} = 4$. Draw a mini-batch of 2. There are 6 equally likely batches:

batch{1,3}{1,5}{1,7}{3,5}{3,7}{5,7}
mini-batch gradient234456

One batch can be off (2 or 6), but the average over all batches is $(2+3+4+4+5+6)/6 = 4$, exactly the full gradient. Counting steps: $n=1000$ examples, batch size $m=50$: one pass over the data (an epoch) makes $1000/50 = 20$ updates. Full-batch makes just 1 update per epoch, so ten epochs give 10 updates versus 200.

Let $\mathbf{g}_i=\nabla\ell_i(\mathbf{w})$ be the gradient from example $i$. Three flavours, all with the same update $\mathbf{w}\leftarrow\mathbf{w}-\eta\,\mathbf{g}$:

  • Full-batch GD: $\mathbf{g}=\frac1n\sum_{i=1}^n\mathbf{g}_i=\nabla L$.
  • Mini-batch SGD: pick a random batch $B$ of $m$ examples, $\mathbf{g}_B=\frac1m\sum_{i\in B}\mathbf{g}_i$.
  • Stochastic GD (SGD) proper: $m=1$. (In practice, people say "SGD" for any mini-batch version.)

Unbiased: a randomly chosen example has expected gradient $E[\mathbf{g}_i] = \frac1n\sum_i\mathbf{g}_i=\nabla L$, and an average of $m$ such terms has the same expectation: $E[\mathbf{g}_B]=\nabla L$. Noisy: if one example's gradient has variance $\sigma^2$, the average of $m$ roughly independent ones has variance $\sigma^2/m$, so the noise (standard deviation) shrinks like $1/\sqrt m$. A batch 4 times bigger halves the noise but costs 4 times more per step.

Practicalities. Shuffle the data each epoch and walk through it in batches. Smaller batches: cheaper and more frequent steps, noisier path (the noise can help escape sharp valleys). Larger batches: smoother and better for GPUs, but fewer updates per epoch, and often they need a larger learning rate (rule of thumb: scale $\eta$ with the batch size, with a short warm-up). Near a minimum the noise keeps us bouncing around instead of settling, so we decay the learning rate late in training.

Awareness: smoothed gradients and schedules. Momentum (previous section) averages recent gradients. Adam keeps two running averages, $\mathbf{m}\leftarrow\beta_1\mathbf{m}+(1-\beta_1)\mathbf{g}$ (a smoothed gradient) and $\mathbf{v}\leftarrow\beta_2\mathbf{v}+(1-\beta_2)\mathbf{g}^2$ (a smoothed squared gradient), and steps $\eta\,\hat{\mathbf{m}}/(\sqrt{\hat{\mathbf{v}}}+\epsilon)$ (here $\hat{\mathbf{m}},\hat{\mathbf{v}}$ are $\mathbf{m},\mathbf{v}$ corrected for starting at zero, and $\epsilon$ is a tiny number that avoids dividing by 0), so each weight gets its own step size. Common learning-rate schedules: step decay, cosine decay, and linear warm-up followed by decay.

Why do we need it?

Modern datasets have millions of examples, and models need thousands of updates. Using a small random batch per step makes each update cheap while staying pointed in the right direction on average.

Where is it used?

Every deep-learning training run: batch sizes from 32 to several thousand, with SGD+momentum for vision models and Adam/AdamW with warm-up and cosine decay for Transformers and language models.

How is it used?

Use a DataLoader with batch_size and shuffle=True. Choose the batch size by memory and speed, tune the learning rate, and add a schedule. If you double the batch size, try raising the learning rate too.

The map shows the loss of a line fit (intercept $b$ horizontal, slope $w$ vertical) on 60 data points. Orange: full-batch gradient descent. Blue: mini-batch SGD with the chosen batch size, same learning rate, 60 steps. With batch size 1 the blue path jitters wildly; with 16 it is a mild wobble; with 60 it matches orange. Press New noise for another random path. The readout shows that the average of many mini-batch gradients equals the full gradient.

Left: the grey dots are noisy mini-batch gradients of one weight, the green dashed line is the true gradient, and the blue curve is a running average $m\leftarrow\beta m+(1-\beta)g$, the "smoothed gradient" behind momentum and Adam. Raise $\beta$: smoother, but it lags behind the truth. Raise the noise to see why smoothing helps. Right: pick a learning-rate schedule and see $\eta$ change over the 100 steps.

"Batch" is overloaded. In "full-batch" it means all the data; in "batch size 32" it means the mini-batch. Papers and code usually mean the latter.

A noisy loss curve is normal with mini-batches. Each step's loss is measured on a different small batch. Look at a moving average, or at the loss over a whole epoch.

Quick check: $n=1200$ examples and batch size 32. How many updates in one epoch?

$1200/32 = 37.5$, so 38 updates (37 full batches plus a last smaller batch of 16). Libraries either keep the small last batch or drop it (drop_last), giving 37.

The training loop on one page: forward, loss, backward, update core

What we need from earlier chapters: every section above: the gradient of a loss, backpropagation (Chapter 2.9), the update rule and mini-batches. From the Linear Algebra guide: machine-learning applications.

All of this chapter fits in a four-beat rhythm, repeated thousands of times:

  1. Forward: push a batch of data through the model to get predictions.
  2. Loss: compare predictions with the truth and get one number.
  3. Backward: use the chain rule to get the slope of that number for every weight.
  4. Update: move every weight a little downhill.

That is exactly what a PyTorch training loop does. There is no more to it: bigger models change what happens inside each beat, never the rhythm.

Our running example: $(1,2),(2,3),(3,7)$, a line $\hat y=b+wx$, mean squared error, $\eta=0.1$, starting at $b=0,\ w=1$. Iteration 1:

  1. Forward: $\hat{\mathbf{y}} = X\mathbf{w} = [1, 2, 3]$.
  2. Loss: residual $\mathbf{r}=\hat{\mathbf{y}}-\mathbf{y} = [-1,-1,-4]$, $L=\frac13(1+1+16)=6$.
  3. Backward: $\nabla L=\frac23X^\top\mathbf{r}=\frac23[-6,-15]=[-4,-10]$.
  4. Update: $\mathbf{w}\leftarrow[0,1]-0.1[-4,-10]=[0.4,\ 2.0]$.

Iteration 2 starts from the new weights: $\hat{\mathbf{y}}=[2.4,4.4,6.4]$, $L\approx0.827$, gradient $[0.8, 0.933]$, and so on. The loss falls fast at first, then creeps along the floor of the long thin valley (the condition number is about 46, as in the 3D bowl earlier). It takes about 200 iterations at this $\eta$ to get within 0.01 of the closed-form $[-1, 2.5]$ from the earlier section, and about 1000 to match it to ten decimals.

Pseudo-code, with the PyTorch name for each beat:


for epoch in range(num_epochs):                 # one pass over the data
    for xb, yb in loader:                       # one mini-batch at a time
        pred = model(xb)                        # 1. FORWARD   (build the computational graph)
        loss = loss_fn(pred, yb)                # 2. LOSS      (one number)
        optimizer.zero_grad()                   #    clear old gradients (they ADD UP otherwise)
        loss.backward()                         # 3. BACKWARD  (chain rule fills p.grad for every weight)
        optimizer.step()                        # 4. UPDATE    (SGD: p -= lr * p.grad)
  • zero_grad() is needed because backward() adds to the stored gradients. That is the right behaviour whenever one weight is used in several places (its gradient is a sum over paths, Chapter 2.8), but it means you must clear them each step.
  • Swapping optimizer (SGD, SGD with momentum, Adam) changes only beat 4. Swapping loss_fn changes only beat 2. Adding regularisation changes beat 2 (a penalty term) or beat 4 (weight_decay).
  • Typical extras: a learning-rate scheduler (scheduler.step()), evaluation on held-out data each epoch, gradient clipping before step().
Why do we need it?

Knowing the loop turns "the model trains itself" into four understandable calls. When training misbehaves (loss not falling, NaN, slow), you know which beat to inspect.

Where is it used?

Every PyTorch, TensorFlow and JAX training script, from a 10-line regression to a large language model. Libraries such as Hugging Face Trainer and PyTorch Lightning wrap this exact loop.

How is it used?

Write the model and the loss; keep the four beats in order. Debug by printing the loss each step, checking one gradient against a finite difference, and trying to overfit a single tiny batch first.

Press Next beat repeatedly to walk through one iteration: the highlighted beat shows its numbers. Compare the first iteration with the example above ($[0,1]\to[0.4,2.0]$). Then press Run 20 iterations twice: the loss drops quickly (6 to under 1) and then creeps, while the orange line slowly turns toward the green one, whose weights are $[-1,\,2.5]$. Raise $\eta$ above about 0.18 and the loss explodes.

The same loop written by hand in plain NumPy for our linear model (no framework). Run it and compare with the widget:


import numpy as np

X = np.array([[1., 1.], [1., 2.], [1., 3.]])   # first column of ones = intercept
y = np.array([2., 3., 7.])
w = np.array([0., 1.])                         # [intercept b, slope w]
eta = 0.1

for it in range(1, 4):
    pred = X @ w                               # 1. forward
    r = pred - y
    loss = np.mean(r ** 2)                     # 2. loss
    grad = 2 / len(y) * X.T @ r                # 3. backward (the gradient we derived)
    w = w - eta * grad                         # 4. update
    print(it, round(loss, 4), grad.round(4), w.round(4))
# 1 6.0 [ -4. -10.] [0.4 2. ]
# 2 0.8267 [0.8    0.9333] [0.32   1.9067]
# 3 0.7525 [ 0.2667 -0.2578] [0.2933 1.9324]   (then a slow creep toward [-1, 2.5])

Forgetting zero_grad() makes gradients pile up across steps: effectively a huge, wrong learning rate. Updating inside no_grad: the update itself must not be recorded in the graph (the optimiser handles that).

Overfit one batch first. A good test of the whole loop: train on a single tiny batch until the model has memorised it. The loss should go to almost zero. If it does not, the bug is in the loop, not the data.

Quick check: what goes wrong if you call loss.backward() twice without zero_grad() in between?

The second call adds its gradients to the stored ones, so p.grad becomes the sum of two gradients (twice as large if nothing changed), and the next update takes a wrongly large step.

Recap, cheat sheet and practice

  • Training = going downhill on the loss. $\mathbf{w}\leftarrow\mathbf{w}-\eta\nabla L(\mathbf{w})$. The step lowers the loss because $L(\mathbf{w}-\eta\nabla L)\approx L-\eta\|\nabla L\|^2$ (linearization).
  • MSE: the gradient with respect to a prediction is $\frac2n(\hat y-y)$, proportional to the error. Linear regression: $\nabla L=\frac2nX^\top(X\mathbf{w}-\mathbf{y})$; setting it to zero gives the normal equations $X^\top X\mathbf{w}=X^\top\mathbf{y}$ (same as Chapter 1.10).
  • Sigmoid: $\sigma'=\sigma(1-\sigma)\le\frac14$. Logistic regression: $\nabla L=\frac1nX^\top(\sigma(X\mathbf{w})-\mathbf{y})$; the $p(1-p)$ factors cancel.
  • Cross-entropy is the negative log-likelihood. It punishes confident mistakes ($-\ln p\to\infty$) and its gradient with respect to the score is $p-y$; MSE through a sigmoid has an extra factor $p(1-p)$ and saturates.
  • Softmax: Jacobian $\operatorname{diag}(\mathbf{p})-\mathbf{p}\mathbf{p}^\top$; with cross-entropy the gradient at the logits is $\mathbf{p}-\mathbf{y}$.
  • Learning rate: on a quadratic the error is multiplied by $1-\eta\lambda$ each step: converge if $\eta<2/\lambda_{\max}$, perfect at $1/\lambda$. Momentum averages past gradients.
  • Regularisation: L2 adds $2\lambda\mathbf{w}$ to the gradient (weight decay: shrink by $1-2\eta\lambda$ each step); L1 adds $\lambda\,\mathrm{sign}(\mathbf{w})$, a constant pull with a kink, which gives exact zeros (sparsity).
  • Neural networks: the loss depends on all weights; backprop is the chain rule: $\boldsymbol\delta_l=\varphi'(\mathbf{z}_l)\odot W_{l+1}^\top\boldsymbol\delta_{l+1}$, $\ \partial L/\partial W_l=\boldsymbol\delta_l\mathbf{a}_{l-1}^\top$.
  • Mini-batches: the batch gradient is an unbiased, noisy estimate of the full gradient (noise $\propto1/\sqrt m$). The loop is always: forward, loss, backward (after zero_grad), update.

Cheat sheet

ModelLossGradient at the outputGradient for the weights
Linear regression$\frac1n\|X\mathbf{w}-\mathbf{y}\|^2$$\frac2n(\hat y_i-y_i)$$\frac2nX^\top(X\mathbf{w}-\mathbf{y})$
Ridge (L2)$L+\lambda\|\mathbf{w}\|^2$unchanged$\nabla L+2\lambda\mathbf{w}$
Lasso (L1)$L+\lambda\|\mathbf{w}\|_1$unchanged$\nabla L+\lambda\,\mathrm{sign}(\mathbf{w})$
Logistic regression$-\frac1n\sum[y\ln p+(1-y)\ln(1-p)]$$\partial\ell/\partial z=p-y$$\frac1nX^\top(\sigma(X\mathbf{w})-\mathbf{y})$
Softmax regression$-\sum_ky_k\ln p_k$$\partial\ell/\partial\mathbf{z}=\mathbf{p}-\mathbf{y}$$(\mathbf{p}-\mathbf{y})\mathbf{h}^\top$ per example
Layer $\mathbf{a}=\varphi(W\mathbf{a}'+\mathbf{b})$any$\boldsymbol\delta=\varphi'(\mathbf{z})\odot(\text{incoming gradient})$$\boldsymbol\delta\,\mathbf{a}'^\top$, and $W^\top\boldsymbol\delta$ goes back
Sigmoid$\sigma'=\sigma(1-\sigma)$
Softmax$J=\operatorname{diag}(\mathbf{p})-\mathbf{p}\mathbf{p}^\top$
Gradient descent$\mathbf{w}\leftarrow\mathbf{w}-\eta\,\mathbf{g}$; stable if $\eta<2/\lambda_{\max}$
Momentum$\mathbf{v}\leftarrow\beta\mathbf{v}+\mathbf{g}$, $\ \mathbf{w}\leftarrow\mathbf{w}-\eta\mathbf{v}$
Mini-batch$\mathbf{g}_B=\frac1m\sum_{i\in B}\nabla\ell_i$, $E[\mathbf{g}_B]=\nabla L$
Code it · NumPy

import numpy as np

def num_grad(f, w, h=1e-6):               # finite-difference gradient, for checking
    g = np.zeros_like(w)
    for j in range(len(w)):
        e = np.zeros_like(w); e[j] = h
        g[j] = (f(w + e) - f(w - e)) / (2 * h)
    return g

# ---------- 1. linear regression: gradient, closed form, gradient descent ----------
X = np.array([[1., 1.], [1., 2.], [1., 3.]])
y = np.array([2., 3., 7.])
mse = lambda w: np.mean((X @ w - y) ** 2)
grad_mse = lambda w: 2 / len(y) * X.T @ (X @ w - y)

w0 = np.array([0., 1.])
print(grad_mse(w0))                       # [ -4. -10.]
print(num_grad(mse, w0).round(4))         # [ -4. -10.]   (finite differences agree)
print(np.linalg.solve(X.T @ X, X.T @ y))  # [-1.   2.5]   normal equations: gradient = 0
w = w0.copy()
for _ in range(3000):
    w -= 0.1 * grad_mse(w)
print(w.round(3))                         # [-1.   2.5]   gradient descent reaches the same point

# ---------- 2. logistic regression: gradient = X^T (sigmoid(Xw) - y) / n ----------
sigmoid = lambda z: 1 / (1 + np.exp(-z))
Xc = np.array([[1., -1.], [1., 1.]])
yc = np.array([0., 1.])
def ce(w):                                # stable form: log(1 + e^z) - y z
    z = Xc @ w
    return np.mean(np.logaddexp(0, z) - yc * z)
grad_ce = lambda w: Xc.T @ (sigmoid(Xc @ w) - yc) / len(yc)
wc = np.zeros(2)
print(round(ce(wc), 4), grad_ce(wc))      # 0.6931 [ 0.  -0.5]   (ln 2, and the gradient we derived)
print(num_grad(ce, wc).round(4))          # [ 0.  -0.5]
wc = wc - 1.0 * grad_ce(wc)
print(wc, round(ce(wc), 4))               # [0.  0.5] 0.4741

# ---------- 3. softmax + cross-entropy: gradient = p - y ----------
def softmax(z):
    e = np.exp(z - z.max())               # subtract the max for stability
    return e / e.sum()
z = np.array([2., 1., 0.]); yo = np.array([0., 1., 0.])
loss = lambda z: -np.log(softmax(z)[1])
p = softmax(z)
print(p.round(3), round(loss(z), 3))      # [0.665 0.245 0.09 ] 1.408
print((p - yo).round(3))                  # [ 0.665 -0.755  0.09 ]
print(num_grad(loss, z).round(3))         # [ 0.665 -0.755  0.09 ]
print((np.diag(p) - np.outer(p, p)).sum(axis=1).round(10))   # [0. 0. 0.]  Jacobian rows sum to zero

# ---------- 4. ridge: the gradient adds 2*lambda*w, and the closed form changes ----------
lam, n = 0.1, len(y)
ridge_w = np.linalg.solve(X.T @ X + n * lam * np.eye(2), X.T @ y)
ridge_grad = lambda w: grad_mse(w) + 2 * lam * w
print(ridge_w.round(4), ridge_grad(ridge_w).round(8))        # [-0.2145  2.118 ] [0. 0.]  gradient is 0 there
print(np.linalg.norm(ridge_w[1:]) < np.linalg.norm(np.linalg.solve(X.T @ X, X.T @ y)[1:]))   # True: the slope shrank

# ---------- 5. mini-batches: unbiased, but noisy ----------
g = np.array([1., 3., 5., 7.])            # per-example gradients
from itertools import combinations
means = [float(np.mean(c)) for c in combinations(g, 2)]
print(means, float(np.mean(means)))       # [2.0, 3.0, 4.0, 4.0, 5.0, 6.0] 4.0
Test yourself

1. For $L(\mathbf{w})=\frac1n\|X\mathbf{w}-\mathbf{y}\|^2$, the gradient is…

Each partial is $\frac2n\sum_iX_{ij}r_i=\frac2n(X^\top\mathbf{r})_j$. Stacked, that is $\frac2nX^\top(X\mathbf{w}-\mathbf{y})$ with shape $d\times1$, matching the shape of $\mathbf{w}$. Setting it to zero gives the normal equations.

2. In the logistic-regression gradient, the factors $\dfrac{1}{p(1-p)}$ (from the loss) and $p(1-p)$ (from the sigmoid) multiply to give…

$\partial\ell/\partial p=(p-y)/(p(1-p))$ and $\partial p/\partial z=p(1-p)$. They cancel, so $\partial\ell/\partial z=p-y$, and multiplying by $\partial z/\partial\mathbf{w}=\mathbf{x}$ gives $(p-y)\mathbf{x}$.

3. Softmax gives $\mathbf{p}=[0.7, 0.2, 0.1]$ and the true class is class 1. The gradient of the cross-entropy loss with respect to the logits is…

$\mathbf{p}-\mathbf{y}=[0.7-1,\ 0.2-0,\ 0.1-0]=[-0.3, 0.2, 0.1]$. The negative entry means "raise the true class's logit"; the entries sum to 0, as the softmax Jacobian rows do.

4. A sigmoid classifier is confidently wrong ($z=-6$, true label 1). Compared with cross-entropy, the gradient of squared error with respect to $z$ is…

Cross-entropy: $|p-1|\approx0.9975$. Squared error: $2|p-1|\,p(1-p)\approx0.0049$. The sigmoid is flat at $z=-6$, so almost no gradient flows back. This is why classifiers use cross-entropy.

5. Which statement about L1 and L2 regularisation is correct?

L2's pull is proportional to $w$, so it vanishes as $w\to0$ and the weight only shrinks. L1's pull has constant size $\lambda$ and its kink at 0 holds the weight there whenever the data's pull is below $\lambda$ (soft-thresholding).

6. Minimising $L(w)=w^2$ by gradient descent from $w_0=1$ with $\eta=1.2$. What happens?

$w\leftarrow w-\eta\cdot2w=(1-2\eta)w=-1.4\,w$. A multiplier below $-1$ means the error grows each step. It converges only for $0<\eta<1$, and $\eta=0.5$ lands exactly on 0.

Practice problems

A. Data $(1,3),(2,5)$, model $\hat y=b+wx$, MSE, start $b=w=0$, $\eta=0.1$. Do one gradient-descent step.

$X=\begin{bmatrix}1&1\\1&2\end{bmatrix}$, $\mathbf{y}=[3,5]$. Residual $\mathbf{r}=\hat{\mathbf{y}}-\mathbf{y}=[-3,-5]$. $X^\top\mathbf{r}=[-3-5,\ -3-10]=[-8,-13]$. Gradient $=\frac2nX^\top\mathbf{r}=\frac22[-8,-13]=[-8,-13]$. Update: $[0,0]-0.1[-8,-13]=[0.8,\ 1.3]$. The loss went from $(9+25)/2=17$ to $\hat{\mathbf{y}}=[2.1,3.4]$, $\mathbf{r}=[-0.9,-1.6]$, $L=(0.81+2.56)/2=1.685$.

B. Logistic regression, one example $\mathbf{x}=[1,2]$ (bias, feature) with $y=1$, weights $\mathbf{w}=[0,0]$, $\eta=1$. Find the gradient, the new weights and the new loss.

$z=0$, $p=0.5$. Gradient $=(p-y)\mathbf{x}=-0.5\,[1,2]=[-0.5,-1]$. New weights $[0,0]-1\cdot[-0.5,-1]=[0.5,1]$. New score $z=0.5+1\cdot2=2.5$, $p=\sigma(2.5)\approx0.924$, loss $=-\ln0.924\approx0.079$ (down from $\ln2\approx0.693$).

C. Softmax with logits $[1,1,1]$ and true class 3. Give $\mathbf{p}$, the loss and the gradient.

All logits are equal, so $\mathbf{p}=[\tfrac13,\tfrac13,\tfrac13]$. Loss $=-\ln\tfrac13=\ln3\approx1.099$. Gradient $\mathbf{p}-\mathbf{y}=[\tfrac13,\tfrac13,-\tfrac23]$ (sums to 0).

D. One weight with data-best value $a=3$. Find the minimiser of $\tfrac12(w-a)^2$ with an L2 penalty $\lambda w^2$ and with an L1 penalty $\lambda|w|$, for $\lambda=0.25$. What if $a=0.2$?

Ridge: $(w-a)+2\lambda w=0\Rightarrow w=a/(1+2\lambda)=3/1.5=2$. Lasso: $w=\mathrm{sign}(a)\max(|a|-\lambda,0)=3-0.25=2.75$. For $a=0.2$: ridge gives $0.2/1.5\approx0.133$ (small but not 0), lasso gives $\max(0.2-0.25,0)=0$ exactly, since $|a|\le\lambda$.

E. For $L(w)=4(w-2)^2$, find the range of learning rates that converge and the learning rate that lands on the minimum in one step.

$L'(w)=8(w-2)$, curvature $a=8$. The error multiplier is $1-8\eta$. It converges when $|1-8\eta|<1$, i.e. $0<\eta<2/8=0.25$. The perfect step is $\eta=1/a=0.125$: $w\leftarrow w-0.125\cdot8(w-2)=2$.

F. A data set has 50,000 examples and the batch size is 128. How many updates does one epoch make, and how many for 5 epochs? How does the noise of the batch gradient change if the batch size grows to 512?

$50000/128=390.6$, so 391 updates per epoch (the last batch is smaller), and $5\times391=1955$ updates. The gradient noise shrinks like $1/\sqrt m$; going from 128 to 512 multiplies $m$ by 4, so the noise is $\sqrt4=2$ times smaller, but each step costs 4 times as much and there are 4 times fewer updates per epoch.

Appendix A

Glossary

Every important word in this guide, in one place, explained in plain English. Type in the box to filter. Each entry links to the chapter that teaches it.

Antiderivative
A function whose derivative is the one you started with. Finding it is the reverse of differentiating, and the integral uses it. 2.13
Asymptote
A line that a curve gets closer and closer to but never quite reaches (for example the x-axis for $e^{-x}$). 2.2
Automatic differentiation
Computing exact derivatives of a program by applying the chain rule to every small step it performs. Backpropagation is its reverse mode. 2.9
Backpropagation
The way neural networks compute all their gradients: one forward pass, then one backward pass that applies the chain rule layer by layer. 2.9
Batch / mini-batch
The group of training examples used to estimate the gradient in one step. Full batch uses all of them, a mini-batch uses a small random subset. 2.14
Chain rule
The derivative of a function inside another function is the product of the two derivatives: $(f\circ g)'(x) = f'(g(x))\,g'(x)$. 2.3 2.8
Computational graph
A diagram of a calculation: each node is a simple operation, each arrow carries a value forward and a derivative backward. 2.8 2.9
Concave / convex
A curve that bends downward (second derivative negative) is concave. One that bends upward (second derivative positive) is convex, like a bowl. 2.10
Continuity
A function is continuous at a point if you can draw it through that point without lifting your pencil: the limit equals the value. 2.2
Contour line
A line joining all points where a function has the same value, like the height lines on a map. Also called a level set. 2.4
Critical point
A point where the gradient is zero. It can be a minimum, a maximum or a saddle. 2.3 2.10
Cross-entropy
The standard loss for classification: $-\sum y_k \log p_k$. It punishes confident wrong answers very strongly. 2.14
Curl
How much a vector field rotates around a point, like a tiny paddle wheel placed in flowing water. 2.13
Curvature
How quickly the slope itself is changing. For a function it is measured by the second derivative, for a surface by the Hessian. 2.10
Derivative
The instantaneous rate of change of a function: the slope of its tangent line, $f'(x) = \lim_{h\to0}\frac{f(x+h)-f(x)}{h}$. 2.3
Differentiable
Having a derivative at every point. A smooth curve is differentiable, a sharp corner (like $|x|$ at 0) is not. 2.3
Directional derivative
The slope of a surface in a chosen direction $\mathbf{u}$: $D_{\mathbf{u}}f = \nabla f\cdot\mathbf{u}$. 2.4
Discontinuity
A point where a function breaks: a hole (removable), a jump, or a blow-up to infinity. 2.2
Divergence
How much a vector field spreads out from a point (a source) or flows in (a sink): $\nabla\cdot\mathbf{F} = \sum \partial F_i/\partial x_i$. 2.13
Domain
The set of inputs a function accepts. 2.1
Dual number
A number $a + b\varepsilon$ with $\varepsilon^2 = 0$. Calculating with dual numbers gives value and derivative together, which is forward-mode differentiation. 2.9
e (Euler's number)
About 2.71828. The base for which $e^x$ has slope equal to its own height. 2.1 2.2
Exploding gradient
Gradients that grow larger and larger as they are passed back through many layers, making training unstable. 2.9
Finite difference
Estimating a derivative by nudging the input a tiny bit: $f'(x) \approx [f(x+h) - f(x-h)]/2h$. Used for checking gradients. 2.3 2.9
First-order approximation
Replacing a function by its tangent line (or tangent plane): $f(x+\delta)\approx f(x)+f'(x)\delta$. 2.11 2.12
Forward mode
Differentiation that carries derivatives forward with the values. One pass per input variable. 2.8 2.9
Function
A rule that gives exactly one output for each input. 2.1
Gradient
The vector of all partial derivatives, $\nabla f$. It points in the direction of steepest increase, and its length is the steepness. 2.4
Gradient checking
Comparing a hand-derived or backprop gradient with a finite-difference estimate to catch mistakes. 2.9
Gradient clipping
Shrinking a gradient whose length is too big, to keep a training step under control. 2.14
Gradient descent
Repeatedly stepping against the gradient: $\mathbf{w}\leftarrow\mathbf{w}-\eta\nabla L(\mathbf{w})$. The basic way models learn. 2.3 2.14
Gradient field
The vector field made of the gradient of a scalar field. Always perpendicular to the level lines. 2.4 2.13
Hessian
The matrix of all second partial derivatives. It describes the curvature of a surface. 2.10
Higher-order derivative
The derivative of a derivative, and so on: $f''$, $f'''$, … 2.3 2.10
Integral
The total accumulated amount, shown as the area under a curve between two points. 2.13
Inverse function
A function that undoes another one: $f^{-1}(f(x)) = x$. Its graph is the mirror image across the line $y=x$. 2.1
Jacobian
The matrix of all first partial derivatives of a vector-valued function. It is the best local linear approximation of the function. 2.5
L1 / L2 regularization
Adding $\lambda\|\mathbf{w}\|_1$ (L1) or $\lambda\|\mathbf{w}\|_2^2$ (L2) to the loss to keep weights small. L1 produces zeros, L2 shrinks smoothly. 2.14
Learning curve
A plot of the loss against training steps (or training-set size). Its slope tells you if learning has stalled. 2.3 2.14
Learning rate
The step size $\eta$ in gradient descent. Too small is slow, too large overshoots. 2.3 2.14
Level set
All the inputs where a function has the same value. In 2D it is a contour line. 2.4
Limit
The value a function gets closer and closer to as the input approaches a point. 2.2
Line integral
Adding up a field along a path, for example the work done by a force along a route. 2.13
Linear function
A function with a constant slope: $f(x)=mx+b$ (or a matrix times a vector, plus a constant). 2.1
Linearization
Approximating a smooth function near a point by its tangent line, plane or Jacobian map. 2.12
Local derivative
The derivative of a single small operation in a computational graph. Backpropagation multiplies local derivatives. 2.8 2.9
Logarithm
The inverse of an exponential: $\ln x$ answers 'what power of $e$ gives $x$?'. It turns products into sums. 2.1
Loss function
A number that says how wrong a model is. Training means making it small. 2.3 2.14
Mean squared error (MSE)
The average of the squared differences between predictions and targets. 2.14
Mixed partial derivative
A second derivative taken with respect to two different variables, like $\partial^2 f/\partial x\,\partial y$. For smooth functions the order does not matter. 2.10
Newton's method
Using the first and second derivative to jump to where the local quadratic model is lowest (or to where the tangent crosses zero). 2.10 2.12
Numerator layout
The convention where the Jacobian of $\mathbf{F}:\mathbb{R}^n\to\mathbb{R}^m$ is an $m\times n$ matrix and the gradient of a scalar function is a column vector. 2.5 2.6
Partial derivative
The derivative with respect to one variable while all the others are held fixed: $\partial f/\partial x$. 2.4
Piecewise function
A function defined by different formulas on different pieces of its domain, like ReLU or $|x|$. 2.1
Polynomial
A sum of powers of $x$ with constant coefficients, such as $3x^2-2x+1$. 2.1
Power rule
$\dfrac{d}{dx}x^n = n x^{n-1}$. 2.3
Product rule
$(fg)' = f'g + fg'$. 2.3
Quotient rule
$(f/g)' = (f'g - fg')/g^2$. 2.3
Range
The set of outputs a function can produce. 2.1
ReLU
$\max(0,x)$: zero for negative inputs, the identity for positive ones. The most common activation function. 2.1 2.3
Remainder (error term)
What is left over when a Taylor polynomial is used instead of the true function. 2.11
Reverse mode
Differentiation that runs backward from the output, giving the gradient with respect to all inputs in one pass. This is backpropagation. 2.8 2.9
Riemann sum
Approximating the area under a curve by adding the areas of thin rectangles. 2.13
Saddle point
A critical point that is a minimum in one direction and a maximum in another, like a mountain pass. 2.10
Scalar field
A number attached to every point in space, like temperature. 2.13
Secant line
A line through two points of a curve. As the points merge it becomes the tangent line. 2.3
Second derivative
The derivative of the derivative. Positive means curving up, negative means curving down. 2.3 2.10
Sensitivity
How much the output changes when an input is nudged: $\Delta y\approx f'(x)\,\Delta x$. 2.3 2.12
Sigmoid
$\sigma(x)=1/(1+e^{-x})$: squashes any number into the range 0 to 1. Used for probabilities. 2.1 2.3
Softmax
Turns a list of scores into probabilities that add to 1: $p_i = e^{z_i}/\sum_j e^{z_j}$. 2.1 2.14
Stochastic gradient descent (SGD)
Gradient descent using a small random mini-batch per step, so each step is cheap but noisy. 2.14
Subgradient
A generalised slope at a corner, such as any value from −1 to 1 for $|x|$ at 0. 2.7
Surface integral
Adding up a field over a surface, for example the flow through it (flux). 2.13
Tangent line
The straight line that just touches a curve at a point and has the same slope there. 2.3
Taylor series
Writing a function as an infinite sum of powers built from its derivatives at one point. 2.11
Trace trick
Using $\operatorname{tr}(AB)=\operatorname{tr}(BA)$ and that a number equals its own trace to rearrange matrix derivatives. 2.6 2.7
Vanishing gradient
Gradients that shrink toward zero as they are passed back through many layers, so early layers stop learning. 2.9
Vector field
A vector (an arrow) attached to every point in space, like wind. 2.13
Vector-valued function
A function whose output is a vector, such as a neural-network layer. 2.5
Weight decay
Shrinking weights a little at every step. It is the gradient step of L2 regularization. 2.14
Appendix B

Formula cheat sheet

The formulas you will use most, on one page. If a formula does not make sense, the chapter link tells you where to relearn the idea. Remember: learn to derive them, do not just memorise them.

Derivative rules (one variable)

RuleFormulaChapter
Definition$f'(x)=\lim_{h\to0}\dfrac{f(x+h)-f(x)}{h}$2.3
Power$(x^n)'=nx^{n-1}$2.3
Sum, constant multiple$(f+g)'=f'+g'$,   $(cf)'=cf'$2.3
Product$(fg)'=f'g+fg'$2.3
Quotient$(f/g)'=\dfrac{f'g-fg'}{g^2}$2.3
Chain$(f\circ g)'(x)=f'(g(x))\,g'(x)$2.3 2.8
Exponential, log$(e^x)'=e^x$,   $(a^x)'=a^x\ln a$,   $(\ln x)'=1/x$2.3
Trigonometric$(\sin x)'=\cos x$,   $(\cos x)'=-\sin x$,   $(\tan x)'=1/\cos^2 x$2.3
Sigmoid, tanh, ReLU$\sigma'=\sigma(1-\sigma)$,   $\tanh'=1-\tanh^2$,   $\mathrm{ReLU}'=\mathbb{1}[x>0]$2.3

Gradients, Jacobians and Hessians

IdeaFormulaChapter
Gradient (column vector)$\nabla f=\left[\dfrac{\partial f}{\partial x_1},\dots,\dfrac{\partial f}{\partial x_n}\right]^\top$2.4
Directional derivative$D_{\mathbf u}f=\nabla f\cdot\mathbf u$, maximal along $\nabla f$2.4
Jacobian ($m\times n$)$J_{ij}=\partial F_i/\partial x_j$2.5
Hessian ($n\times n$)$H_{ij}=\partial^2 f/\partial x_i\partial x_j$ (symmetric)2.10
Chain rule with Jacobians$J_{f\circ g}=J_f\,J_g$;   $\nabla_{\mathbf x}f(g(\mathbf x))=J_g^\top\nabla f$2.5 2.8
Taylor (1st, 2nd order)$f(\mathbf x+\boldsymbol\delta)\approx f+\nabla f^\top\boldsymbol\delta+\tfrac12\boldsymbol\delta^\top H\boldsymbol\delta$2.11
Gradient descent, Newton$\mathbf w\leftarrow\mathbf w-\eta\nabla L$;   $\mathbf w\leftarrow\mathbf w-H^{-1}\nabla L$2.10 2.14

Matrix-calculus identities

FunctionGradient / derivativeChapter
$\mathbf a^\top\mathbf x$$\mathbf a$2.6
$A\mathbf x$Jacobian $A$2.6
$\mathbf x^\top\mathbf x=\|\mathbf x\|^2$$2\mathbf x$2.6
$\mathbf x^\top A\mathbf x$$(A+A^\top)\mathbf x$   ($=2A\mathbf x$ if $A$ symmetric)2.6 2.7
$\|A\mathbf x-\mathbf b\|^2$$2A^\top(A\mathbf x-\mathbf b)$2.6
$\|\mathbf x\|_2$, $\|\mathbf x\|_1$$\mathbf x/\|\mathbf x\|$,   $\operatorname{sign}(\mathbf x)$ (subgradient)2.7
$\|X\|_F^2$,   $\operatorname{tr}(AX)$$2X$,   $A^\top$2.7
$\log\det X$,   $d(X^{-1})$$X^{-\top}$,   $-X^{-1}\,dX\,X^{-1}$2.7

The gradients behind machine learning

Model / lossGradientChapter
Linear regression, MSE$\dfrac{2}{n}X^\top(X\mathbf w-\mathbf y)$2.14
Logistic regression, cross-entropy$\dfrac{1}{n}X^\top(\sigma(X\mathbf w)-\mathbf y)$2.14
Softmax + cross-entropy (logits $\mathbf z$)$\mathbf p-\mathbf y$2.14
Softmax Jacobian$\operatorname{diag}(\mathbf p)-\mathbf p\mathbf p^\top$2.14
L2 / L1 regularization$2\lambda\mathbf w$,   $\lambda\operatorname{sign}(\mathbf w)$2.14
Backprop through a dense layer ($\mathbf y=W\mathbf x+\mathbf b$)$\partial L/\partial W=\boldsymbol\delta\,\mathbf x^\top$,   $\partial L/\partial\mathbf x=W^\top\boldsymbol\delta$,   $\partial L/\partial\mathbf b=\boldsymbol\delta$2.9
Appendix C

Where to go next

You finished the tour. Here is how to make it stick, and where to dig deeper.

How to make it stick

  • Check every gradient numerically. After you derive a gradient, nudge each input by a tiny amount and compare. This one habit catches almost every mistake.
  • Derive, don't memorise. If you can derive $\nabla\|A\mathbf x-\mathbf b\|^2$ in two minutes, you will never forget it.
  • Write a tiny autograd. A few dozen lines that build a graph, run forward and run backward will make backpropagation feel obvious.
  • Train something. Fit a logistic regression with gradient descent using only NumPy, and watch the loss curve.

Free resources

  • 3Blue1Brown, "Essence of Calculus" (YouTube): the best visual intuition for derivatives and the chain rule.
  • Mathematics for Machine Learning (Deisenroth, Faisal, Ong), chapter 5 "Vector Calculus": free PDF.
  • The Matrix Calculus You Need For Deep Learning (Parr & Howard): free article.
  • Andrej Karpathy, "The spelled-out intro to neural networks and backpropagation" (micrograd): builds backprop from scratch.
  • The Matrix Cookbook (Petersen & Pedersen): free reference of matrix-derivative identities.
  • Deep Learning (Goodfellow, Bengio, Courville), chapters 4 and 6: numerical computation and backpropagation.

Next in the series

Calculus tells you which way is downhill. Optimization is the art of actually getting to the bottom, quickly and safely: gradient descent, momentum, Adam, constraints and convexity. Continue with the Optimization for Machine Learning guide.

Companion guide

Linear algebra and calculus work together in every model. Revisit the Linear Algebra guide whenever a vector or matrix feels shaky, in particular least squares, eigenvalues and the SVD.