Linear Algebra for Machine Learning
Learn the maths behind machine learning from zero. Every idea starts with a plain-English picture, then a worked example, then the formal definition, and then something you can drag, slide and play with.
What is linear algebra, in one sentence? It is the maths of lists of numbers (vectors), tables of numbers (matrices), and the things you can do with them: add them, stretch them, rotate them, and combine them.
Why care? A photo is a table of numbers. A sentence turns into a list of numbers. A neural network is a chain of tables of numbers that transform lists of numbers. If you understand these objects, machine learning stops feeling like magic.
How every topic is taught
Each concept follows the same six steps, in this order. The order matters: understanding comes before formulas.
1 · Intuition
A picture or everyday story. No symbols yet. If you only read this part, you will already get the idea.
2 · Example
A small problem with real numbers, solved slowly, step by step. You can redo it with pen and paper.
3 · Definition
Now the precise statement and the notation. Because you have seen the idea, the symbols just give it a name.
4 · Play
An interactive visual. Drag points, move sliders and watch the numbers change. In 3D you can also rotate the whole scene. Each one tells you what to try.
5 · Why · Where · How
Every concept answers three questions in plain English: Why do we need it? Where is it used? How is it used? So you always know why the idea is worth your time.
6 · Check
Short questions to test yourself, a recap of the key points, a NumPy code block to type out, and practice problems with worked answers.
Try your first interactive
The boxes marked Interactive react to you. Drag the dot below.
The roadmap
Seventeen chapters, in an order where each one builds on the last. Click any card to jump there.
- 1.1Mathematical FoundationsThe notation and algebra you need first
- 1.2VectorsArrows, lists, length, angle, dot product
- 1.3Vector SpacesSpan, independence, basis, dimension
- 1.4MatricesTables of numbers and how to multiply them
- 1.5Matrix GeometryMatrices as rotations, stretches and flips
- 1.6Systems of Linear EquationsSolving Ax = b by elimination
- 1.7Matrix PropertiesRank, determinant, inverse, condition number
- 1.8Four Fundamental SubspacesThe complete structure of any matrix
- 1.9OrthogonalityPerpendicularity, projection, Gram–Schmidt, QR
- 1.10Least SquaresThe best answer when no exact answer exists
- 1.11Eigenvalues & EigenvectorsThe special directions of a matrix
- 1.12Quadratic Forms & DefinitenessBowls, saddles and positive definite matrices
- 1.13Matrix DecompositionsLU, QR, Cholesky (in depth) and the star: SVD
- 1.14Matrix CalculusGradients, Jacobians, Hessians, backprop
- 1.15Numerical Linear AlgebraWhy computers get maths slightly wrong
- 1.16Tensors & Array ProgrammingShapes, broadcasting, einsum, batching
- 1.17ML ApplicationsRegression, PCA, embeddings, attention, LoRA…
How to read the maths symbols
You will meet these symbols again and again. Do not memorise them now. Come back to this table whenever one looks strange.
| Symbol | Say it as | Meaning |
|---|---|---|
| $x \in S$ | "x is in S" | $x$ is one of the things inside the set $S$. |
| $\mathbb{R}$ | "the real numbers" | Every number on the number line: 3, −2.5, π… |
| $\mathbb{R}^n$ | "R n" | All lists of $n$ real numbers. $\mathbb{R}^2$ is the flat plane, $\mathbb{R}^3$ is 3D space. |
| $\mathbf{v}$, $\vec v$ | "vector v" | A list of numbers. We write vectors in bold. |
| $A$ | "matrix A" | A table of numbers. Capital letters are matrices. |
| $\sum_{i=1}^{n}$ | "sum from i = 1 to n" | Add up the thing after it, for $i = 1, 2, \dots, n$. |
| $\|\mathbf{v}\|$ | "the norm of v" | The length of $\mathbf{v}$. |
| $A^\top$ | "A transpose" | Flip a matrix over its diagonal (rows become columns). |
| $A^{-1}$ | "A inverse" | The matrix that undoes $A$. |
| $\forall$, $\exists$ | "for all", "there exists" | Shorthand used in definitions. |
| $\approx$ | "approximately equal" | Close, but not exactly equal. |
| $\iff$ | "if and only if" | The two statements are either both true or both false. |
Tips for studying
- Play first, read second. If a paragraph confuses you, go to the interactive below it and move things around. Then re-read.
- Do the examples by hand once. It feels slow, but nothing builds intuition faster.
- Say "why" out loud. After each concept, try to explain it to an imaginary friend using only the intuition box.
- Don't rush. One chapter a week is a good pace. Come back to earlier chapters whenever you need to.
- Use the tools. Press / to search, use the sidebar to jump around, switch dark mode with the button at the top, and mark chapters complete as you go.
If a section feels too hard, it is almost always because a word from an earlier section is fuzzy. Go back, find that word in the sidebar, and reread it. Nobody gets linear algebra in a single pass. Seeing it twice is normal.
Mathematical Foundations
Before we meet vectors and matrices, we need the small toolkit that every page of machine-learning maths uses: sets, functions, the Σ symbol, exponents and logarithms. If school maths feels far away, relax. We will rebuild each idea slowly, with pictures you can play with.
- Read and write the basic language of sets, number systems and functions
- Read Σ (sum) and Π (product) notation, and translate it into code
- Do the basic algebra that appears everywhere: expanding, factoring, equations, inequalities, absolute value
- Use exponents and logarithms, and understand why machine learning adds logs instead of multiplying probabilities
This chapter is a refresher. If a section feels easy, play with its interactive and move on. The sections marked core (functions, Σ notation, exponents and logs) are the ones you will use again and again.
Sets: collections of things
A set is just a collection of things, like a bag of marbles or the list of people in your family. The only question a set answers is: "is this thing in the collection, yes or no?"
- The order does not matter. $\{1,2,3\}$ and $\{3,2,1\}$ are the same set.
- Repeats do not matter. $\{1,2,2,3\}$ is the same set as $\{1,2,3\}$.
Sets are the grammar of mathematics. When we later say "a vector in $\mathbb{R}^3$", we are using a set.
Let $A = \{1, 2, 3, 4, 5\}$ and $B = \{4, 5, 6, 7\}$, and let the whole universe of numbers we care about be $U = \{1, 2, \dots, 10\}$.
- Membership. $3 \in A$ ("3 is in A") and $6 \notin A$ ("6 is not in A").
- Union (in A or B or both): $A \cup B = \{1,2,3,4,5,6,7\}$.
- Intersection (in A and B): $A \cap B = \{4, 5\}$.
- Difference (in A but not in B): $A \setminus B = \{1, 2, 3\}$.
- Complement (everything in $U$ that is not in A): $A^c = \{6,7,8,9,10\}$.
- Subset. $\{1, 2\} \subseteq A$ because every element of $\{1,2\}$ is also in $A$.
- $x \in A$: "$x$ is an element of $A$". $x \notin A$: it is not.
- Set-builder notation describes a set by a rule: $\{x \in \mathbb{R} : x > 0\}$ reads "the set of all real numbers $x$ such that $x > 0$". The colon (or a bar $|$) means "such that".
- $A \subseteq B$ ("A is a subset of B"): every element of $A$ is in $B$.
- $A \cup B$ (union), $A \cap B$ (intersection), $A \setminus B$ (difference), $A^c$ (complement, relative to a universe $U$).
- The empty set $\varnothing = \{\}$ has no elements. It is a subset of every set.
- The cardinality $|A|$ is the number of elements. Here $|A| = 5$ and $|\varnothing| = 0$. (Sets like $\mathbb{R}$ have infinitely many elements.)
Why do we need it?
Machine learning talks about collections all the time: all the training examples, all the words, all the possible outcomes. Sets give us one clear language for "which things are in the collection".
Where is it used?
Probability (an event is a set of outcomes), train/test splits, vocabularies of words, database queries (AND is intersection, OR is union), and the set notation that fills every ML paper.
How is it used?
Write the collection with braces, then ask membership questions (is x in A?) or combine collections with union, intersection and difference. In Python, set objects do this with |, & and -.
"Subset" is not "element". $2 \in \{1,2,3\}$ is true, but $\{2\} \subseteq \{1,2,3\}$ is a different statement: one is a number in a set, the other is a small set inside a bigger set. Also, "or" in maths is inclusive: the union includes things that are in both.
Quick check: if $A = \{1,2,3\}$ and $B = \{3,4\}$, what are $A \cap B$ and $B \setminus A$?
$A \cap B = \{3\}$ (only 3 is in both). $B \setminus A = \{4\}$ (4 is in B but not in A).
Cartesian product and $\mathbb{R}^n$
You have 3 shirts and 2 pairs of trousers. How many different outfits can you make? Every outfit is a pair: (a shirt, a pair of trousers). The collection of all such pairs is the Cartesian product of the two sets.
The same idea, with two sets of numbers, is how we build the flat page we draw graphs on: one number for "across", one number for "up".
Let $A = \{1, 2\}$ and $B = \{a, b, c\}$. Then
$$A \times B = \{(1,a), (1,b), (1,c), (2,a), (2,b), (2,c)\}.$$There are $2 \times 3 = 6$ pairs. Order matters inside a pair: $(1, a)$ is a pair, but $(a, 1)$ is a different kind of object that belongs to $B \times A$.
If both sets are the real numbers, $\mathbb{R} \times \mathbb{R}$ is every pair $(x, y)$ of real numbers: the whole plane. We write it $\mathbb{R}^2$.
Repeating the product $n$ times gives lists of length $n$:
$$\mathbb{R}^n = \underbrace{\mathbb{R} \times \mathbb{R} \times \dots \times \mathbb{R}}_{n \text{ times}} = \{(x_1, x_2, \dots, x_n) : x_i \in \mathbb{R}\}.$$So an element of $\mathbb{R}^n$ is an ordered list of $n$ real numbers. That is exactly what a vector is (Chapter 1.2).
Why do we need it?
We need a way to say "a list of n numbers" precisely. The Cartesian product builds such lists out of simple sets, one slot per number.
Where is it used?
Every vector, data point and image lives in some $\mathbb{R}^n$. Coordinates on a graph, RGB colours (a point in $\mathbb{R}^3$) and a dataset's "feature space" are all products of number lines.
How is it used?
To describe a data point with n measurements, say it is an element of $\mathbb{R}^n$. Count the slots to get the size. For finite sets, multiply the sizes: |A × B| = |A| · |B|.
One more number gives space. A point on a page needs two numbers (across, up). A point in a room needs three: how far right, how far forward, how far up. That is an element of $\mathbb{R} \times \mathbb{R} \times \mathbb{R} = \mathbb{R}^3$. Drag the point below and watch the three numbers.
Number systems and intervals
Numbers came in stages, each one invented to fix a problem the last one could not solve.
- Counting numbers $1, 2, 3, \dots$ let you count apples.
- "I owe you 3 apples" needs negatives and zero, so we get the integers.
- "Share 1 apple between 2 people" needs fractions, so we get the rationals.
- "How long is the diagonal of a unit square?" needs numbers that are not fractions, like $\sqrt{2}$. The rationals have tiny gaps; filling every gap gives the real numbers.
- "Which number squared gives $-1$?" has no real answer. We invent a new number $i$ and get the complex numbers.
Each system sits inside the next, like nested boxes.
- $5$ is a natural number, so it is also an integer, a rational ($5 = \tfrac51$), a real and a complex number.
- $-3$ is an integer (not natural), and also rational, real, complex.
- $\tfrac12 = 0.5$ is rational. Its decimals stop or repeat forever, like $\tfrac13 = 0.333\ldots$
- $\sqrt{2} = 1.41421\ldots$ and $\pi = 3.14159\ldots$ are real but not rational: their decimals never repeat.
| Symbol | Name | What is in it |
|---|---|---|
| $\mathbb{N}$ | natural numbers | $1, 2, 3, \dots$ (some books also include 0) |
| $\mathbb{Z}$ | integers | $\dots, -2, -1, 0, 1, 2, \dots$ |
| $\mathbb{Q}$ | rationals | fractions $\tfrac{p}{q}$ with $p, q$ integers, $q \ne 0$ |
| $\mathbb{R}$ | reals | every point on the number line |
| $\mathbb{C}$ | complex | $a + bi$ with $a, b$ real and $i^2 = -1$ |
Why do we need it?
Different jobs need different kinds of numbers: counting, money owed, fractions, lengths, and numbers that square to a negative. Knowing which set a number lives in tells you what you may do with it.
Where is it used?
Integer labels and indices (ℤ), model weights and features (ℝ, stored as floats), probabilities (the interval [0, 1]), and complex numbers in the Fourier transform and in eigenvalues.
How is it used?
Check which set your quantity belongs to before using it. Write ranges as intervals, for example a probability is in [0, 1] and a ReLU output is in [0, ∞). Open or closed ends tell you whether the end value is allowed.
Intervals: pieces of the number line
An interval is an unbroken stretch of the number line, like "all the numbers from 1 to 4". The only question is whether the two end points are included. On a drawing, a filled dot means included and an empty ring means left out.
- $(1, 4)$ is every number strictly between 1 and 4: $1 \lt x \lt 4$. Neither end is included, so $1$ and $4$ are not in it, but $1.0001$ is.
- $[1, 4]$ includes both ends: $1 \le x \le 4$. So $4 \in [1,4]$.
- $[1, 4)$ includes 1 but not 4. $(1, 4]$ includes 4 but not 1 (these are half-open).
- $(2, \infty)$ is everything bigger than 2. Infinity is never "included", so it always gets a round bracket.
For real numbers $a \lt b$:
$$(a,b) = \{x : a \lt x \lt b\}, \quad [a,b] = \{x : a \le x \le b\}, \quad [a,b) = \{x : a \le x \lt b\}, \quad (a,b] = \{x : a \lt x \le b\}.$$Round bracket = end excluded (open). Square bracket = end included (closed).
Quick check: which of these is a natural number, which is rational but not an integer, which is real but not rational?
$7$ is natural. $\tfrac{3}{4}$ is rational but not an integer. $\sqrt{3}$ is real but not rational.
Complex numbers (a first look)
A complex number is a number with two parts: a real part and an "imaginary" part. The name sounds strange, but the idea is friendly: you can draw it as a point on a flat page. Go across by the real part, then up by the imaginary part. A complex number is a 2D point that you are allowed to multiply.
You only need a taste for now. Complex numbers return in Chapter 1.11, when some matrices have eigenvalues that are not real (for example, matrices that spin things around).
Let $z = 3 + 4i$. Its real part is $3$ and its imaginary part is $4$.
- The modulus (size, distance from the origin) is $|z| = \sqrt{3^2 + 4^2} = \sqrt{25} = 5$.
- The conjugate flips the sign of the imaginary part: $\bar{z} = 3 - 4i$.
- Multiply: $z\bar{z} = (3+4i)(3-4i) = 9 - 12i + 12i - 16i^2 = 9 + 16 = 25$, because $i^2 = -1$.
So $z\bar{z} = |z|^2$: the imaginary parts cancel and a plain real number is left.
A complex number is $z = a + bi$ where $a, b \in \mathbb{R}$ and $i^2 = -1$. Here $a = \operatorname{Re}(z)$ and $b = \operatorname{Im}(z)$.
$$|z| = \sqrt{a^2 + b^2}, \qquad \bar{z} = a - bi, \qquad z\bar{z} = a^2 + b^2 = |z|^2.$$On the complex plane, $z$ is the point $(a, b)$: horizontal axis = real part, vertical axis = imaginary part. Every real number is a complex number with $b = 0$.
Why do we need it?
Some everyday equations, like x² + 1 = 0, have no real answer. Complex numbers fix that, and later they let us describe rotations and waves with one tidy number.
Where is it used?
Eigenvalues of rotation-like matrices, the Fourier transform (audio, images, signal processing), and quantum computing. Most day-to-day ML code stays real, but these tools need ℂ.
How is it used?
Treat a + bi as the point (a, b) on a plane. Get its size with the modulus √(a² + b²), flip the sign of b for the conjugate, and remember that i² = −1. In NumPy, np.abs(z) and np.conj(z) do it.
$i$ is not a mystery: it is simply the name of a number that squares to $-1$. The "imaginary" part is a real number $b$ attached to $i$. Do not worry about complex arithmetic beyond $i^2 = -1$ and the three formulas above.
Functions: a rule from inputs to outputs core
A function is a machine. You put something in, the machine follows a fixed rule, and exactly one thing comes out.
- A vending machine: press B3 (input), get one particular snack (output). Press B3 again, get the same snack.
- A taxi meter: kilometres in, price out.
Two rules make something a function: every allowed input gives an output, and each input gives only one output. (Different inputs may share an output. That is allowed.)
Three words describe the machine. The domain is the set of inputs it accepts. The codomain is the set of outputs it is allowed to give. The range is the outputs it actually gives.
Take the rule "square it": $f(x) = x^2$.
- $f(3) = 3^2 = 9$.
- $f(-2) = (-2)^2 = 4$. (Different inputs, here $2$ and $-2$, can give the same output.)
- Domain: any real number can be squared, so the domain is $\mathbb{R}$.
- Range: a square is never negative, so the range is $\{y : y \ge 0\}$. If we declared the codomain to be $\mathbb{R}$, the range is a smaller piece of it.
Two more: $f(x) = \sqrt{x}$ only accepts $x \ge 0$ (no real square root of $-4$), and $f(x) = 1/x$ refuses $x = 0$ (you cannot divide by zero).
A function $f \colon A \to B$ assigns to every $a \in A$ exactly one element $f(a) \in B$.
- Domain $A$: the allowed inputs. Codomain $B$: the set the outputs are declared to live in.
- Range (or image): $\{f(a) : a \in A\} \subseteq B$, the outputs that really happen.
When inputs and outputs are lists of numbers we write $f \colon \mathbb{R}^n \to \mathbb{R}^m$: "takes $n$ numbers in, gives $m$ numbers out". Examples:
| Function | Type | In words |
|---|---|---|
| $f(x) = x^2$ | $\mathbb{R} \to \mathbb{R}$ | one number in, one out |
| $f(x, y) = x + y$ | $\mathbb{R}^2 \to \mathbb{R}$ | two numbers in, their sum out |
| $f(x, y) = (x,\, y,\, x+y)$ | $\mathbb{R}^2 \to \mathbb{R}^3$ | a 2D vector becomes a 3D vector |
| a loss function $L(\mathbf{w})$ | $\mathbb{R}^n \to \mathbb{R}$ | $n$ weights in, one "badness" number out |
Why do we need it?
A model, a loss and a layer are all "something that takes input and gives output". Functions give that idea a precise form, so we can state exactly what goes in, what comes out, and what is allowed.
Where is it used?
A trained model (features to prediction), a loss function (weights to one error number), activation functions such as ReLU and sigmoid, and every layer of a neural network.
How is it used?
Write down the type f : ℝⁿ → ℝᵐ first (how many numbers in, how many out). Then check the domain (inputs that are allowed) and the range (outputs that can happen). Plot or evaluate it at a few inputs to build a feel.
The function is the rule plus its domain. $x^2$ on all of $\mathbb{R}$ and $x^2$ on $x \ge 0$ are different functions with different properties (see the next section). Also, $f(x)$ means "the output of $f$ at $x$", not "$f$ times $x$".
Quick check: what is the domain of $f(x) = \dfrac{1}{x - 2}$?
Everything except $x = 2$, because that makes the denominator $0$. In interval notation: $(-\infty, 2) \cup (2, \infty)$.
One-to-one, onto and bijective functions core
Imagine matching students (inputs) to seats (outputs).
- Injective ("one-to-one"): no two students share a seat. Different inputs always give different outputs.
- Surjective ("onto"): no seat is left empty. Every possible output is actually produced.
- Bijective: both. A perfect matching, one student per seat and no empty seats. You can reverse it perfectly.
Why care? An injective function loses no information (you can tell which input was used). A function that is not injective merges inputs, and a merge cannot be undone.
Inputs $A = \{1, 2, 3\}$.
- To $B = \{a, b, c, d\}$ with $1 \to a,\ 2 \to b,\ 3 \to c$: injective (nobody shares), but not surjective ($d$ is empty).
- To $B = \{a, b\}$ with $1 \to a,\ 2 \to b,\ 3 \to b$: surjective (both seats used), but not injective ($2$ and $3$ share $b$).
- To $B = \{a, b, c\}$ with $1 \to b,\ 2 \to c,\ 3 \to a$: bijective.
With real functions. $f(x) = 2x + 1$ from $\mathbb{R}$ to $\mathbb{R}$ is bijective. $f(x) = x^2$ is neither: $f(2) = f(-2)$ (not injective) and no input gives $-1$ (not surjective). $f(x) = e^x$ is injective but misses every number $\le 0$.
- Injective: if $f(a) = f(a')$ then $a = a'$. (Equal outputs force equal inputs.)
- Surjective: for every $b \in B$ there is at least one $a \in A$ with $f(a) = b$. (Range = codomain.)
- Bijective: injective and surjective. Exactly one input for every output.
Horizontal line test. Draw any horizontal line across the graph of $f \colon \mathbb{R} \to \mathbb{R}$. Injective means every such line cuts the graph at most once. Surjective means every such line cuts it at least once.
Why do we need it?
We often need to know whether a process loses information or can be undone. Injective, surjective and bijective are the exact words for "nothing merged", "everything reachable" and "perfect undo".
Where is it used?
ReLU and pooling layers (not injective, so they forget information), embeddings and tokenisers (word to ID, a bijection on the vocabulary), and invertible matrices and normalising flows.
How is it used?
Ask two questions: can two different inputs give the same output (injective fails), and is some output never produced (surjective fails)? On a graph, sweep a horizontal line and count how often it cuts the curve.
It depends on the domain and codomain you declare. $x^2$ on all of $\mathbb{R}$ is not injective. But $x^2$ on $[0, \infty)$ is injective (the negative inputs are gone), and onto $[0, \infty)$ it is a bijection. This trick of "restricting the domain" is how $\sqrt{x}$ becomes the inverse of $x^2$.
Quick check: is $f(x) = |x|$ (from $\mathbb{R}$ to $\mathbb{R}$) injective? Surjective?
Neither. $|2| = |-2|$, so two inputs share an output (not injective). And no input gives $-1$, so $-1$ is never hit (not surjective onto $\mathbb{R}$).
Composition and inverse functions core
Composition means chaining machines: the output of the first machine goes straight into the second. Putting on socks, then shoes is a composition. Order matters: shoes, then socks gives a very different result.
An inverse is the "undo" button. If $f$ turns a Celsius temperature into Fahrenheit, then $f^{-1}$ turns it back. Only a perfect matching (a bijection) can be undone for sure, because otherwise you would not know which input to go back to.
Let $g(x) = 2x$ ("double") and $f(x) = x + 3$ ("add 3"). Start with $x = 5$.
- $(f \circ g)(5) = f(g(5)) = f(10) = 13$. (Double first, then add 3.)
- $(g \circ f)(5) = g(f(5)) = g(8) = 16$. (Add 3 first, then double.)
- As formulas: $(f \circ g)(x) = 2x + 3$ but $(g \circ f)(x) = 2(x+3) = 2x + 6$. Different!
Inverse of $f(x) = 2x + 1$: write $y = 2x + 1$ and solve for $x$: $x = \dfrac{y - 1}{2}$. So $f^{-1}(y) = \dfrac{y-1}{2}$. Check: $f(3) = 7$ and $f^{-1}(7) = \tfrac{6}{2} = 3$. ✓
Composition. $(f \circ g)(x) = f(g(x))$. Read it right to left: apply $g$ first, then $f$. In general $f \circ g \ne g \circ f$.
Inverse. If $f \colon A \to B$ is a bijection, its inverse $f^{-1} \colon B \to A$ satisfies
$$f^{-1}(f(x)) = x \quad\text{and}\quad f(f^{-1}(y)) = y.$$Geometrically, the graph of $f^{-1}$ is the graph of $f$ mirrored across the line $y = x$. Also $(f \circ g)^{-1} = g^{-1} \circ f^{-1}$: undo the steps in reverse order (take off shoes, then socks).
Why do we need it?
Big operations are built by chaining small ones, and often we want to undo a step. Composition describes chaining, and the inverse describes undoing.
Where is it used?
A neural network is a composition of layers, and backpropagation is the chain rule applied to that chain. Log and exp undo each other in softmax, and a matrix inverse undoes a matrix.
How is it used?
To compose, feed the output of g into f and read right to left: f(g(x)). To invert, write y = f(x), solve for x, and check that f⁻¹(f(x)) = x. Remember that order matters.
$f^{-1}$ does not mean $1/f$. The small $-1$ means "undo". For $f(x) = 2x+1$, $f^{-1}(x) = (x-1)/2$, while $1/f(x) = 1/(2x+1)$. These are completely different functions.
Quick check: find $f^{-1}$ for $f(x) = 3x - 6$.
Write $y = 3x - 6$, add $6$: $y + 6 = 3x$, divide by $3$: $x = \dfrac{y + 6}{3}$. So $f^{-1}(y) = \dfrac{y+6}{3}$. Check: $f(4) = 6$ and $f^{-1}(6) = 12/3 = 4$. ✓
Linear and non-linear functions (a preview)
A linear function is a fair, proportional machine.
- Double the input and the output doubles.
- Feed in two things added together, and you get the two outputs added together.
- Nothing in, nothing out: an input of $0$ gives an output of $0$.
"3 dollars per kilogram" is linear: 2 kg costs twice as much as 1 kg. But "3 dollars per kilogram plus a 5 dollar delivery fee" is not: 0 kg still costs 5. (Such "linear plus a shift" functions are called affine.) And "area of a square from its side", $x^2$, is not linear either: doubling the side quadruples the area.
Linear algebra is the study of exactly these well-behaved functions. This is only a preview. We make it precise in Chapter 1.5.
Test $f(x) = 2x$ with $x = 3$, $y = 4$, $c = 5$:
- Adding: $f(3 + 4) = f(7) = 14$ and $f(3) + f(4) = 6 + 8 = 14$ ✓.
- Scaling: $f(5 \cdot 3) = f(15) = 30$ and $5 \cdot f(3) = 5 \cdot 6 = 30$ ✓.
Test $f(x) = x^2$ with $x = 1$, $y = 2$: $f(1 + 2) = 9$ but $f(1) + f(2) = 1 + 4 = 5$. ✗ So $x^2$ is not linear. And $f(x) = x + 1$ fails because $f(0) = 1 \ne 0$.
A function $f$ is linear if for all inputs $\mathbf{x}, \mathbf{y}$ and every scalar $c$:
$$f(\mathbf{x} + \mathbf{y}) = f(\mathbf{x}) + f(\mathbf{y}) \quad\text{(additivity)}, \qquad f(c\,\mathbf{x}) = c\, f(\mathbf{x}) \quad\text{(scaling)}.$$Both tests must pass for all choices. A single failing example proves the function is not linear. A consequence: $f(\mathbf{0}) = \mathbf{0}$.
For $f \colon \mathbb{R} \to \mathbb{R}$, the linear functions are exactly $f(x) = m x$: straight lines through the origin. For $f \colon \mathbb{R}^n \to \mathbb{R}^m$, they are exactly matrix multiplications $f(\mathbf{x}) = A\mathbf{x}$ (Chapter 1.5).
Why do we need it?
Linear functions are the only ones we can analyse completely, and they are the heart of this whole guide. A quick test tells us whether a function is one of them.
Where is it used?
A neuron's weighted sum, every matrix multiplication, linear regression, and the first half of every layer in a network. Non-linear activations are added on purpose to make networks more powerful.
How is it used?
Test additivity f(x + y) = f(x) + f(y) and scaling f(cx) = c f(x). One failing example proves a function is not linear. Also check f(0) = 0: a function with a constant offset is only affine.
In everyday speech, "linear" is used for any straight line, like $y = 2x + 1$. In linear algebra it is stricter: the line must pass through the origin. $y = 2x + 1$ is affine, not linear. (We study affine maps in Chapters 1.3 and 1.5.)
Quick check: is $f(x) = 5x$ linear? Is $f(x) = 5x - 2$?
$5x$ is linear. $5x - 2$ is not, since $f(0) = -2 \ne 0$ (it is affine: a line that misses the origin).
Scalars and notation conventions core
Mathematics is partly a language, and good notation is like good handwriting: you stop noticing it. The big idea is that the shape of a letter tells you what kind of object it is.
- A plain slanted letter ($a$, $x$, $\lambda$) is a single number, called a scalar.
- A bold lowercase letter ($\mathbf{x}$, $\mathbf{w}$) is a list of numbers, a vector.
- A capital letter ($A$, $W$, $X$) is a grid of numbers, a matrix.
Small numbers written below a letter are indices: the address of one entry, like a house number.
Let $\mathbf{x} = \begin{bmatrix} 5 \\ 8 \\ 2 \end{bmatrix}$ and $A = \begin{bmatrix} 1 & 2 & 3 \\ 4 & 5 & 6 \end{bmatrix}$.
- $x_2 = 8$ (the second entry of $\mathbf{x}$).
- $A_{2,3} = 6$: row 2, column 3. Row first, then column.
- $A[1, :] = [1, 2, 3]$ means "row 1, every column". $A[:, 3] = [3, 6]^\top$ means "every row, column 3".
| What | Usual notation | Examples |
|---|---|---|
| Scalar (one number) | plain italic, often $a, b, c, \lambda, \eta, \alpha$ | learning rate $\eta = 0.01$ |
| Vector (list) | bold lowercase $\mathbf{x}$, or arrow $\vec{x}$ | weights $\mathbf{w}$, features $\mathbf{x}$ |
| Matrix (grid) | capital $A$, or bold capital $\mathbf{W}$ | data matrix $X$, weight matrix $W$ |
| Entry of a vector | $x_i$ (the $i$-th entry) | $x_3$ |
| Entry of a matrix | $A_{ij}$ or $A_{i,j}$ (row $i$, column $j$) | $A_{2,3}$ |
| Row / column of a matrix | $A[i, :]$ and $A[:, j]$ (code-style) | $A[2, :]$ |
| Transpose | $\mathbf{x}^\top$, $A^\top$ (flip rows and columns) | $\mathbf{w}^\top\mathbf{x}$ |
Maths counts from 1, but Python and NumPy count from 0. So $A_{2,3}$ is A[1, 2] in code.
Why do we need it?
Without a shared notation, every formula would need a paragraph of explanation. A letter's shape tells you at a glance whether it is a number, a list or a grid.
Where is it used?
Every ML paper, textbook and library doc: x for a feature vector, W for a weight matrix, η for a learning rate, A[i, j] for a matrix entry, and NumPy indexing in code.
How is it used?
Read the shape first (plain = scalar, bold lowercase = vector, capital = matrix), then the indices (row first, then column). When you code it, remember that maths counts from 1 but NumPy counts from 0.
- Subscript vs superscript. In ML, $x_j$ is usually the $j$-th entry of a vector, while $\mathbf{x}^{(i)}$ (with brackets) is the $i$-th example in a dataset. And $x_j^{(i)}$ is feature $j$ of example $i$. The brackets mean "this is a label, not an exponent".
- Not every book follows the bold/capital convention. When in doubt, look for the sentence that says "let $\mathbf{x}$ be a vector".
Quick check: in $A = \begin{bmatrix} 7 & 8 \\ 9 & 10 \end{bmatrix}$, what is $A_{2,1}$, and how do you write it in NumPy?
Row 2, column 1 is $9$. In NumPy: A[1, 0].
Summation notation Σ core
Σ is the capital Greek letter sigma, the "S" of Sum. It is shorthand for "add up a whole list of things". Writing $1 + 2 + 3 + \dots + 100$ is tiring, so we squeeze it into one symbol.
The best way to read Σ is as a for-loop that adds:
- Underneath: the loop counter and where it starts ($i = 1$).
- On top: where the counter stops ($n$).
- To the right: what to add each time round.
Compute $\displaystyle\sum_{i=1}^{4} i^2$.
- Let $i$ run through $1, 2, 3, 4$.
- For each $i$, compute $i^2$: $1,\ 4,\ 9,\ 16$.
- Add them: $1 + 4 + 9 + 16 = 30$.
Data version. If $\mathbf{x} = [3, 1, 4]$, then $\sum_{i=1}^{3} x_i = 3 + 1 + 4 = 8$, and the average is $\dfrac{1}{3}\sum_{i=1}^{3} x_i = \dfrac{8}{3} \approx 2.67$.
- $i$ is the index (the loop counter); $m$ is the lower bound and $n$ the upper bound. Both ends are included.
- The number of terms is $n - m + 1$. So $\sum_{i=0}^{n}$ has $n+1$ terms.
- The index is a dummy name: $\sum_i a_i$, $\sum_j a_j$ and $\sum_k a_k$ mean the same thing, just like a loop variable.
- An empty sum (upper bound below lower bound) equals $0$. Adding a constant $n$ times: $\sum_{i=1}^{n} c = n\,c$.
Why do we need it?
Losses and averages add up thousands or millions of terms. Σ lets us write all of that in one short expression instead of an endless list.
Where is it used?
Mean squared error, cross-entropy, the dot product, L1 and L2 norms, and every average over a dataset all contain a Σ.
How is it used?
Read a sum like a for-loop that adds: start at the lower bound, stop at the upper bound, add the expression each time. In code this is total += ... in a loop, or one NumPy call like x.sum().
- Off-by-one. $\sum_{i=1}^{n}$ has $n$ terms, but $\sum_{i=0}^{n}$ has $n+1$. Python's
range(1, n+1)stops before the second number, so to include $n$ you must write $n+1$. - The index belongs to the sum only. After the Σ, $i$ no longer exists, just as a loop variable is local to the loop.
Quick check: how many terms does $\displaystyle\sum_{k=3}^{7} k$ have, and what is it?
Terms: $k = 3, 4, 5, 6, 7$, so $7 - 3 + 1 = 5$ terms. The sum is $3+4+5+6+7 = 25$.
Manipulating sums: rules and double sums
Because a sum is just a long addition, all the usual rules of addition still work. You may:
- Pull out a common multiplier: if every term is multiplied by 3, multiply the total by 3 once.
- Split a sum into two sums, if each term is itself a sum of two pieces.
- Add in any order. That is why, for a grid of numbers, you can add rows first then columns, or columns first then rows, and get the same total.
What you may not do is pretend a sum of products is a product of sums.
Let $\mathbf{a} = [1, 2, 3]$ and $\mathbf{b} = [4, 5, 6]$.
- Split: $\sum (a_i + b_i) = (5 + 7 + 9) = 21$, and $\sum a_i + \sum b_i = 6 + 15 = 21$ ✓.
- Constant out: $\sum 3a_i = 3 + 6 + 9 = 18 = 3 \cdot 6$ ✓.
- Not allowed: $\sum a_ib_i = 4 + 10 + 18 = 32$, but $\left(\sum a_i\right)\left(\sum b_i\right) = 6 \cdot 15 = 90$. ✗ Different!
For finite sums:
$$\sum_{i} c\,a_i = c\sum_{i} a_i \qquad \sum_{i}(a_i + b_i) = \sum_{i} a_i + \sum_{i} b_i \qquad \sum_{i=1}^{n} a_i = \sum_{i=1}^{k} a_i + \sum_{i=k+1}^{n} a_i$$Double sums (a sum of sums) add every entry of a grid, and the order does not matter:
$$\sum_{i=1}^{m}\sum_{j=1}^{n} A_{ij} = \sum_{j=1}^{n}\sum_{i=1}^{m} A_{ij}.$$Read a double sum from the inside out: first sum over $j$ for a fixed $i$ (one row), then over $i$.
Warning examples: $\sum a_ib_i \ne \big(\sum a_i\big)\big(\sum b_i\big)$ and $\sum a_i^2 \ne \big(\sum a_i\big)^2$.
Why do we need it?
Raw sums can be long and clumsy. Rules for splitting them, pulling constants out and swapping the order let us simplify formulas and spot mistakes.
Where is it used?
Deriving the gradient of a loss, simplifying a mean, writing matrix products as (AB)ᵢⱼ = Σₖ AᵢₖBₖⱼ, and summing over rows or columns of a data table.
How is it used?
Pull constants outside the sum, split a sum of two parts into two sums, and swap the order of a double sum (rows first or columns first). Check by testing small numbers. Never turn a sum of products into a product of sums.
The classic mistake: $\sum_i a_i^2$ is the sum of the squares ($1 + 4 + 9 = 14$ for $[1,2,3]$), while $\left(\sum_i a_i\right)^2$ is the square of the sum ($6^2 = 36$). The square must be inside or outside the Σ on purpose.
Quick check: simplify $\displaystyle\sum_{i=1}^{10} (2x_i + 5)$ in terms of $\sum x_i$.
Split into two sums and pull the 2 out: $2\sum_{i=1}^{10} x_i + \sum_{i=1}^{10} 5 = 2\sum x_i + 10 \cdot 5 = 2\sum x_i + 50$.
Product notation Π core
Π is the capital Greek letter pi, the "P" of Product. It works exactly like Σ, but multiplies instead of adds. Think of a for-loop that starts with result = 1 and does result *= term.
Why would a data scientist multiply? Because probabilities of independent events (events that do not affect each other) multiply. If a coin lands heads with probability $0.5$, the chance of three heads in a row is $0.5 \times 0.5 \times 0.5$.
- Factorial: $\displaystyle\prod_{i=1}^{4} i = 1 \cdot 2 \cdot 3 \cdot 4 = 24 = 4!$.
- Three independent coin flips, each with probability $0.5$: $\displaystyle\prod_{i=1}^{3} 0.5 = 0.5^3 = 0.125$.
- A model gives probabilities $0.9,\ 0.8,\ 0.5$ to three observed data points. The probability that it "predicted" all of them is $0.9 \times 0.8 \times 0.5 = 0.36$.
The pieces are the same as for Σ (index, lower and upper bound). An empty product equals $1$ (the "do nothing" value for multiplication, just as an empty sum is $0$). Two handy rules: $\prod c\,a_i = c^n \prod a_i$ (the constant is multiplied in each of the $n$ factors), and $\prod (a_ib_i) = \big(\prod a_i\big)\big(\prod b_i\big)$.
Why do we need it?
The chance that several independent things all happen is the product of their chances. Π is the short way to write a long product.
Where is it used?
The likelihood of a dataset in maximum likelihood training (logistic regression, language models), factorials n! in counting and probability, and naive Bayes classifiers.
How is it used?
Read it like a loop that starts at 1 and multiplies each term in. In code, np.prod(p). With many probabilities, switch to logs (see the log-likelihood section at the end of this chapter) before the product underflows.
The constant rule is different from Σ: $\sum_{i=1}^{n} c\,a_i = c \sum a_i$, but $\prod_{i=1}^{n} c\,a_i = c^{\,n}\prod a_i$. The constant gets multiplied in every one of the $n$ factors.
Quick check: what is $\displaystyle\prod_{i=1}^{3} (2i)$?
$(2)(4)(6) = 48$. (Using the constant rule: $2^3 \cdot 3! = 8 \cdot 6 = 48$ ✓.)
Translating between Σ, loops and vectorised code core
A sum can be written in three languages that all mean the same thing:
- Maths with Σ: short and precise.
- A loop: easy to understand, but slow in Python for big data.
- Vectorised NumPy: one line, no visible loop, and it runs in fast compiled code.
Good ML code is written vectorised. Good understanding comes from being able to see the loop hiding inside. Practise reading a formula by taking it apart piece by piece, from the inside out.
Squared errors summed over $n$ examples, $\displaystyle\sum_{i=1}^{n} (y_i - \hat y_i)^2$. Here $y_i$ is the true value and $\hat y_i$ ("y-hat") is the model's prediction:
import numpy as np
y = np.array([3.0, 5.0, 2.0]) # true values
y_hat = np.array([2.5, 5.5, 4.0]) # predictions
n = len(y)
# 1) the loop
total = 0.0
for i in range(n):
total += (y[i] - y_hat[i]) ** 2
# 2) vectorised (same number, one line)
total = np.sum((y - y_hat) ** 2)
print(total) # 4.5 (0.25 + 0.25 + 4), the same from both ways
The Σ's "for each $i$" becomes the loop; or, in NumPy, the arrays are handled all at once and np.sum plays the role of Σ.
| Maths | NumPy (vectorised) |
|---|---|
| $\sum_i x_i$ | x.sum() |
| $\tfrac{1}{n}\sum_i x_i$ (mean) | x.mean() |
| $\sum_i x_i^2$ | (x**2).sum() or x @ x |
| $\sum_i a_ib_i$ (dot product) | a @ b |
| $\sum_i\sum_j A_{ij}$ | A.sum() |
| $\sum_j A_{ij}$ for each row $i$ | A.sum(axis=1) |
| $\prod_i p_i$ | np.prod(p) |
| $\sum_i \log p_i$ | np.log(p).sum() |
Why do we need it?
Maths on paper, plain loops and fast NumPy code all describe the same sum. Being able to move between them lets you read a paper and turn it into working, fast code.
Where is it used?
Implementing any loss from a paper, writing vectorised training loops, and checking a fast NumPy version against a slow, obviously-correct loop.
How is it used?
Take a formula apart from the inside out: predict, subtract, square, add up, average. Write the loop first to be safe, then replace it by an array expression such as np.mean((y - w*x)**2) and compare the answers.
- Vectorised does not mean "no loop". The loop is still there, just done inside fast code. You should always be able to write the loop in your head.
- In NumPy,
*multiplies entry by entry, while@is the dot product (a sum of products). Mixing them up gives a vector where you wanted a number.
Quick check: write $\displaystyle\sum_{i=1}^{n} |x_i|$ in NumPy.
np.abs(x).sum() (this is the L1 norm, also np.linalg.norm(x, 1)).
Basic algebra: expanding, factoring, simplifying
Algebra is arithmetic with letters standing in for numbers. Three moves do most of the work:
- Expand: multiply out brackets. $3(x + 2) = 3x + 6$.
- Factor: the reverse, pulling a common piece out into brackets. $3x + 6 = 3(x+2)$.
- Simplify: combine like terms. $2x + 5x = 7x$, but $2x + 5$ cannot be combined.
A picture makes expanding obvious: $(x+a)(x+b)$ is the area of a rectangle with sides $x + a$ and $x + b$. Cut it into four pieces and add the pieces.
- Expand $(x + 2)(x + 3)$: multiply every term in the first bracket by every term in the second: $x\cdot x + x\cdot 3 + 2\cdot x + 2\cdot 3 = x^2 + 3x + 2x + 6$.
- Combine like terms: $x^2 + 5x + 6$.
- Check with $x = 1$: $(1+2)(1+3) = 12$ and $1 + 5 + 6 = 12$ ✓.
- Factor $x^2 + 5x + 6$ backwards: find two numbers that multiply to $6$ and add to $5$. That is $2$ and $3$. So it equals $(x+2)(x+3)$.
Useful identities (each is a rectangle cut into pieces):
$$(x+a)(x+b) = x^2 + (a+b)x + ab$$ $$(a+b)^2 = a^2 + 2ab + b^2, \qquad (a-b)^2 = a^2 - 2ab + b^2, \qquad (a+b)(a-b) = a^2 - b^2.$$The key law behind expanding is the distributive law: $a(b + c) = ab + ac$.
Why do we need it?
Formulas are rarely in the form we need. Expanding and factoring let us rewrite an expression in a more useful shape without changing its value.
Where is it used?
Expanding a squared error (a − b)² = a² − 2ab + b² to derive least squares, expanding ‖a − b‖², and simplifying algebra in gradient derivations.
How is it used?
To expand, multiply every term in one bracket by every term in the other, then collect like terms. To factor, look for a common piece or two numbers with the right sum and product. Check by plugging in a number.
- $(a + b)^2 \ne a^2 + b^2$. The missing piece is the middle term $2ab$.
- $\sqrt{a^2 + b^2} \ne a + b$, and $\dfrac{a + b}{c} = \dfrac{a}{c} + \dfrac{b}{c}$ but $\dfrac{a}{b + c} \ne \dfrac{a}{b} + \dfrac{a}{c}$.
- A minus sign in front of a bracket flips every term: $-(x - 3) = -x + 3$.
Quick check: expand $(x - 4)(x + 4)$ and factor $x^2 + 7x + 12$.
$(x-4)(x+4) = x^2 - 16$ (the middle terms cancel). For $x^2 + 7x + 12$ find two numbers with product $12$ and sum $7$: $3$ and $4$. So it is $(x+3)(x+4)$.
Linear equations and inequalities in one variable
An equation is a balanced scale: "this side weighs the same as that side". To solve it, do the same thing to both sides so the scale stays balanced, until $x$ is alone.
An inequality is a scale that is tipped: "this side is heavier than that side". Most of the same moves work, with one surprise: if you multiply or divide both sides by a negative number, the tip flips direction.
The difference in the answer: an equation usually has one solution (a point). An inequality has a whole range of solutions (an interval).
Equation: $2x + 3 = 11$.
- Subtract 3 from both sides: $2x = 8$.
- Divide both sides by 2: $x = 4$.
- Check: $2(4) + 3 = 11$ ✓.
Inequality: $2x + 3 \lt 11$ gives $x \lt 4$: every number below 4 works, and the answer is the interval $(-\infty, 4)$.
The flip: $-2x + 3 \lt 11$. Subtract 3: $-2x \lt 8$. Divide by $-2$ and flip: $x \gt -4$. Check with $x = 0$: $3 \lt 11$ ✓ and $0 \gt -4$ ✓.
A linear equation in one variable has the form $ax + b = c$. If $a \ne 0$, the single solution is $$x = \frac{c - b}{a}.$$ A linear inequality replaces "$=$" by $\lt, \le, \gt$ or $\ge$. Allowed moves: add or subtract the same number on both sides; multiply or divide by a positive number. Multiplying or dividing by a negative number reverses the inequality sign.
Special case $a = 0$: then no $x$ is left, and the statement "$b$ vs $c$" is either true for every $x$ or for none.
Why do we need it?
Training and fitting are about finding unknown numbers that satisfy conditions. Solving equations finds exact values; inequalities describe whole allowed ranges.
Where is it used?
Constraints in optimisation (weights must be non-negative), decision boundaries of classifiers (wᵀx + b = 0 versus > 0), margins in support vector machines, and tolerance checks in code.
How is it used?
Do the same thing to both sides until x is alone. For inequalities do the same, but flip the sign when you multiply or divide by a negative number. Test one value from your answer to be safe.
- Forgetting to flip when dividing by a negative number is the most common inequality mistake. Test one value to be safe.
- Multiplying by something that might be zero or negative (like a variable) is unsafe in inequalities. Stick to numbers whose sign you know.
- "At least 5" means $\ge 5$. "More than 5" means $\gt 5$.
Quick check: solve $5 - 3x \ge 14$.
Subtract 5: $-3x \ge 9$. Divide by $-3$ and flip: $x \le -3$. Check $x = -4$: $5 + 12 = 17 \ge 14$ ✓.
Solving two equations by hand
Sometimes you have two unknowns and two clues. "Two apples and one banana cost 8 dollars. One apple and one banana cost 5 dollars. What does each cost?" One clue is not enough, but together they pin down the answer.
Each equation is a straight line of possibilities. The answer must satisfy both clues, so it is the point where the two lines cross. Two lines can cross once, never (parallel), or lie on top of each other (infinitely many answers).
Solve $\ x + y = 5\ $ and $\ x - y = 1$.
- Eliminate $y$ by adding the two equations: $(x + y) + (x - y) = 5 + 1$, so $2x = 6$ and $x = 3$.
- Substitute back into the first: $3 + y = 5$, so $y = 2$.
- Check the second: $3 - 2 = 1$ ✓.
Apples and bananas: $2a + b = 8$ and $a + b = 5$. Subtract: $a = 3$, then $b = 2$.
A system of two linear equations:
$$\begin{aligned} a_1 x + b_1 y &= c_1 \\ a_2 x + b_2 y &= c_2 \end{aligned}$$Two methods: elimination (add or subtract multiples of the equations to cancel one unknown) and substitution (solve one equation for one unknown and plug it into the other). With $D = a_1b_2 - a_2b_1$:
- $D \ne 0$: exactly one solution, $x = \dfrac{c_1b_2 - c_2b_1}{D}$, $y = \dfrac{a_1c_2 - a_2c_1}{D}$.
- $D = 0$: the lines are parallel. Either no solution (different lines) or infinitely many (the same line).
This is the beginning of Chapter 1.6, where we solve systems with thousands of unknowns.
Why do we need it?
One clue rarely pins down two unknowns, but two clues often do. Systems of equations are how several conditions are combined to find unknown values.
Where is it used?
Linear regression (finding weights that match the data), balancing and mixing problems, solving for circuit currents, and the very start of Chapter 1.6, where the same idea scales to thousands of unknowns.
How is it used?
Use elimination: add or subtract multiples of the equations to cancel one unknown, solve for the other, then substitute back. Always check the answer in both equations. In code: np.linalg.solve(A, b).
When you eliminate, apply the operation to both sides of the equation, including the right-hand number. And always check your answer in both original equations.
Quick check: solve $3x + y = 11$ and $x - y = 1$.
Add the equations: $4x = 12$, so $x = 3$. Then $3 - y = 1$ gives $y = 2$. Check the first: $9 + 2 = 11$ ✓.
Absolute value
The absolute value $|x|$ is the distance from 0 on the number line. Distance never has a sign: $-3$ and $3$ are both 3 steps from zero, so $|-3| = |3| = 3$.
More generally, $|x - c|$ is the distance between $x$ and $c$. So "$|x - c| \lt r$" says "x is closer to c than r": all the points inside a circle (in 1D, an interval) of radius $r$ around $c$.
- $|5| = 5$, $|-5| = 5$, $|0| = 0$.
- $|x| \lt 2$ means "within 2 of zero", i.e. $-2 \lt x \lt 2$.
- $|x - 3| \lt 2$ means "within 2 of 3", so $1 \lt x \lt 5$.
- $|x| \gt 2$ means "farther than 2 from zero": $x \lt -2$ or $x \gt 2$ (two separate pieces).
- $|x| \ge 0$, and $|x| = 0$ only for $x = 0$.
- $|xy| = |x|\,|y|$ and $|-x| = |x|$.
- Triangle inequality: $|x + y| \le |x| + |y|$.
- For $a \gt 0$: $\;|x| \lt a \iff -a \lt x \lt a$ and $|x| \gt a \iff x \lt -a \text{ or } x \gt a$.
Why do we need it?
We often care how far apart two numbers are, not which one is bigger. Absolute value measures that distance, and |x − c| < r says "close to c".
Where is it used?
The L1 norm and mean absolute error, tolerance checks such as |a − b| < 1e-6 for comparing floats, robust losses (Huber), and ε-balls in optimisation and nearest-neighbour methods.
How is it used?
Read |x − c| as "distance from x to c". To solve |x − c| < r, turn it into the interval c − r < x < c + r. For |x − c| > r you get two outer pieces joined by "or".
- $|x| = -x$ is true only when $x \le 0$. Absolute value is not "just drop the sign of the letter": $|x|$ for negative $x$ is a positive number.
- $|x + y| \ne |x| + |y|$ in general: $|3 + (-5)| = 2$ but $|3| + |-5| = 8$. Only $\le$ is always true.
- "$|x| \gt 2$" is two pieces joined by "or", not one interval.
Quick check: solve $|x - 1| \le 3$.
Within 3 of 1: $-3 \le x - 1 \le 3$, so $-2 \le x \le 4$, the closed interval $[-2, 4]$.
Exponents core
An exponent counts repeated multiplication. $2^5$ means "multiply five 2s together". The small raised number is how many times to use the base.
The rules all come from that one idea. If you multiply three 2s by two more 2s, you get five 2s: $2^3 \cdot 2^2 = 2^5$. Multiplying adds the exponents.
Extending the idea to zero, negative and fractional exponents is just a matter of keeping the rules working:
- $b^0 = 1$ (because $b^3 / b^3 = b^{0}$ must equal $1$).
- $b^{-n} = 1/b^n$ (a negative exponent means "divide instead").
- $b^{1/2} = \sqrt{b}$ (because $\sqrt{b}\cdot\sqrt{b} = b = b^{1/2 + 1/2}$).
- $2^3 = 2\cdot2\cdot2 = 8$.
- $2^3 \cdot 2^2 = 8 \cdot 4 = 32 = 2^5$ ✓ (exponents add).
- $2^3 / 2^2 = 8 / 4 = 2 = 2^1$ ✓ (exponents subtract).
- $(2^3)^2 = 8^2 = 64 = 2^6$ ✓ (exponents multiply).
- $2^{-3} = 1/8 = 0.125$ and $9^{1/2} = 3$ and $10^0 = 1$.
$b$ is the base and $n$ the exponent (or power). For real exponents we take $b \gt 0$.
Why do we need it?
Many quantities grow or shrink by repeating a multiplication: interest, learning-rate decay, the number of binary combinations. Exponents write that repeated multiplication compactly.
Where is it used?
Learning-rate schedules (0.9ᵗ), scientific notation such as 1e-3, the size 2ⁿ of a space of n binary features, the squared norm Σxᵢ², and e^x inside softmax and sigmoid.
How is it used?
Apply the rules: add exponents when multiplying powers of the same base, multiply exponents for a power of a power. Remember b⁰ = 1, b⁻ⁿ = 1/bⁿ and b^(1/2) = √b. In code use ** or np.power.
- $(a + b)^2 \ne a^2 + b^2$ (see the area picture earlier). Exponent rules work for products and powers, never for sums.
- $-3^2 = -9$ but $(-3)^2 = 9$. The exponent binds tighter than the minus sign.
- $2^3 \cdot 3^3 = 6^3$ works (same exponent), but $2^3 \cdot 2^2 = 2^5$ needs the same base. $2^3 \cdot 3^2$ does not simplify.
- $(b^m)^n = b^{mn}$ multiplies, while $b^m \cdot b^n = b^{m+n}$ adds. Do not mix them up.
Quick check: simplify $\dfrac{x^5 \cdot x^{-2}}{x^{4}}$.
Top: $x^{5 + (-2)} = x^3$. Then $x^3 / x^4 = x^{3-4} = x^{-1} = 1/x$.
Logarithms, $e^x$ and the natural log core
A logarithm answers the question "what exponent do I need?" Since $2^3 = 8$, we say $\log_2 8 = 3$: "to get 8 from 2, use the exponent 3". The logarithm is simply the inverse (the "undo") of the exponent, just as subtraction undoes addition.
Another view: $\log_{10}$ roughly counts the digits. $\log_{10} 1000 = 3$ (three zeros), $\log_{10} 1{,}000{,}000 = 6$. Logs squash enormous ranges into small, manageable numbers.
And because exponents add when you multiply, logarithms turn multiplication into addition. This is the single most useful fact about them.
- $\log_2 8 = 3$ because $2^3 = 8$. And $\log_{10} 0.01 = -2$ because $10^{-2} = 0.01$.
- Product: $\log_2(8 \cdot 4) = \log_2 32 = 5$, and $\log_2 8 + \log_2 4 = 3 + 2 = 5$ ✓.
- Quotient: $\log_2(8 / 4) = \log_2 2 = 1$, and $3 - 2 = 1$ ✓.
- Power: $\log_2(8^2) = \log_2 64 = 6$, and $2 \cdot \log_2 8 = 2 \cdot 3 = 6$ ✓.
For a base $b \gt 0$, $b \ne 1$ and a number $x \gt 0$:
$$\log_b x = y \iff b^y = x.$$ $$\log_b(xy) = \log_b x + \log_b y \qquad \log_b\!\left(\frac{x}{y}\right) = \log_b x - \log_b y \qquad \log_b(x^k) = k\log_b x$$ $$\log_b 1 = 0 \qquad \log_b b = 1 \qquad b^{\log_b x} = x \qquad \log_b(b^x) = x \qquad \log_b x = \frac{\ln x}{\ln b}$$The special number $e$. $e \approx 2.71828$ is a constant that appears whenever something grows at a rate proportional to its size (interest, populations, probabilities). The function $e^x$ (also written $\exp(x)$) is the natural exponential. Its inverse is the natural logarithm $\ln x = \log_e x$. Most ML maths uses $\ln$, and many people write plain $\log$ to mean $\ln$.
$\log$ only accepts positive inputs: its domain is $(0, \infty)$, and its range is all of $\mathbb{R}$. It is increasing: bigger input, bigger output.
Why do we need it?
Numbers in ML range from 0.000001 to a billion, and products of many numbers are hard to handle. Logarithms squash huge ranges and turn multiplication into addition.
Where is it used?
Cross-entropy and log-likelihood losses, softmax and log-softmax, information measured in bits (−log₂ p), and log-scale plots for loss curves and learning-rate sweeps.
How is it used?
Ask "which exponent gives this number?". Use log(xy) = log x + log y, log(x/y) = log x − log y and log(xᵏ) = k log x to simplify. Only feed positive numbers. In NumPy: np.log, np.exp.
- $\log(x + y) \ne \log x + \log y$. The rule is for products: $\log(xy) = \log x + \log y$. There is no neat formula for the log of a sum.
- $\log(0)$ and the log of a negative number are not defined. Code returns
-infornan. - $\log(x^k) = k \log x$, but $(\log x)^k$ is something else.
- $\log_{10}$, $\log_2$ and $\ln$ differ only by a constant factor ($\log_b x = \ln x / \ln b$), so they have the same shape.
Quick check: simplify $\ln(e^3 \cdot e^2)$ and $\log_{10} 1{,}000{,}000 - \log_{10} 1000$.
$\ln(e^3 e^2) = \ln(e^5) = 5$. And $6 - 3 = 3$, which is also $\log_{10}(10^6/10^3) = \log_{10} 1000 = 3$ ✓.
Why logs turn products into sums: log-likelihood and underflow core
A model gives each of your data points a probability, say $0.3$. The probability of the whole dataset is the product of all of these. With thousands of data points you are multiplying thousands of numbers below 1.
That product becomes unimaginably tiny. A computer stores numbers with a fixed amount of room, and below a certain size it gives up and just stores 0. This is called underflow. Once the product is $0$, you can no longer tell a good model from a bad one.
The fix is lovely: take the log of the product. Log turns the product into a sum of moderate numbers, which a computer handles easily. And since log is increasing, the model that makes the product biggest also makes the log of the product biggest. Nothing is lost.
Small case. Probabilities $0.9,\ 0.8,\ 0.5$:
- Product: $0.9 \times 0.8 \times 0.5 = 0.36$.
- Logs: $\ln 0.9 = -0.105,\ \ln 0.8 = -0.223,\ \ln 0.5 = -0.693$.
- Sum: $-0.105 - 0.223 - 0.693 = -1.022$. And $e^{-1.022} = 0.36$ ✓: same information.
Big case. 400 data points, each with probability $0.1$. The true product is $0.1^{400} = 10^{-400}$. The smallest positive number a standard 64-bit float can hold is about $10^{-324}$, so the computer says 0. But $\sum \ln 0.1 = 400 \times (-2.303) = -921.0$, a perfectly ordinary number.
For data points with model probabilities $p_1, \dots, p_n$ (independent), the likelihood and log-likelihood are
$$L = \prod_{i=1}^{n} p_i, \qquad \ell = \log L = \sum_{i=1}^{n} \log p_i.$$Because $\log$ is increasing, maximising $L$ and maximising $\ell$ give the same answer. In practice we minimise the negative log-likelihood $-\sum_i \log p_i$, which is the same thing as the cross-entropy loss used for classification. As a bonus, the derivative of a sum is much simpler than the derivative of a product.
Why do we need it?
Multiplying thousands of probabilities makes a number so small that a computer rounds it to zero (underflow), which ruins every comparison. Adding logs avoids this and finds the same best model.
Where is it used?
Training by maximum likelihood: logistic regression, classifiers (cross-entropy loss), language models scoring text with log-probabilities, and the log-sum-exp trick inside softmax libraries.
How is it used?
Never compute np.log(np.prod(p)) for long lists. Compute np.sum(np.log(p)) instead, and minimise the negative of it. If a probability can be 0, clip it first, for example np.clip(p, 1e-12, 1).
- Never compute
np.log(np.prod(p))for long lists. Computenp.sum(np.log(p)). - Log-probabilities are negative (or zero), because probabilities are at most 1. Do not be alarmed by large negative numbers; "less negative" means "more likely".
- If some $p_i = 0$, then $\log p_i = -\infty$. Models avoid this by clipping probabilities away from exactly 0 (for example
np.clip(p, 1e-12, 1)).
Quick check: 100 data points each with probability $0.5$. What is the log-likelihood (natural log)?
$\sum \ln 0.5 = 100 \times (-0.693) = -69.3$. The product itself is $0.5^{100} \approx 7.9 \times 10^{-31}$, tiny but still fine in float64. With 2000 points it would underflow, while the sum of logs would just be $-1386.3$.
Recap, cheat sheet and practice
- Sets: collections; $\in$, $\subseteq$, $\cup$, $\cap$, $\setminus$, complement, $A \times B$, and $\mathbb{R}^n$ is lists of $n$ reals.
- Number systems nest: $\mathbb{N} \subset \mathbb{Z} \subset \mathbb{Q} \subset \mathbb{R} \subset \mathbb{C}$. Intervals use round (excluded) and square (included) brackets.
- A function $f \colon A \to B$ gives one output per input. Know domain, range, injective/surjective/bijective, composition $f \circ g$, inverse $f^{-1}$, and the two tests for linear.
- Notation: plain = scalar, bold lowercase = vector, capital = matrix, $A_{ij}$ is row $i$, column $j$. Maths counts from 1, NumPy from 0.
- Σ is a loop that adds, Π a loop that multiplies. Pull constants out of Σ, split sums, swap double sums. Vectorise in NumPy.
- Algebra: expand, factor, solve $ax + b = c$, flip an inequality when dividing by a negative, solve two equations by elimination. $|x - c| \lt r$ means "within $r$ of $c$".
- Exponents add when you multiply; logs are their inverse and turn products into sums. That is why we use the log-likelihood and avoid underflow.
Cheat sheet
| Topic | Key facts |
|---|---|
| Sets | $A \cup B$ (or), $A \cap B$ (and), $A \setminus B$ (not in B), $|A \times B| = |A||B|$ |
| Functions | injective: no shared outputs · surjective: all outputs hit · bijective: both · $(f\circ g)(x) = f(g(x))$ |
| Linear | $f(x+y) = f(x)+f(y)$ and $f(cx) = cf(x)$ |
| Σ rules | $\sum c a_i = c\sum a_i$ · $\sum (a_i + b_i) = \sum a_i + \sum b_i$ · $\sum_{i=1}^n c = nc$ |
| Π rules | $\prod c a_i = c^n \prod a_i$ · empty product = 1 |
| Inequalities | divide by a negative: flip the sign |
| Absolute value | $|x| \lt a \iff -a \lt x \lt a$ · $|x + y| \le |x| + |y|$ |
| Exponents | $b^mb^n = b^{m+n}$ · $(b^m)^n = b^{mn}$ · $b^0 = 1$ · $b^{-n} = 1/b^n$ |
| Logs | $\log(xy) = \log x + \log y$ · $\log(x^k) = k\log x$ · $\log_b x = \ln x / \ln b$ |
| Log-likelihood | $\log\prod p_i = \sum \log p_i$ (never multiply thousands of probabilities) |
import numpy as np
# ---- sets (Python has them built in)
A = {1, 2, 3, 4, 5}
B = {4, 5, 6, 7}
print(A | B, A & B, A - B) # union, intersection, difference
print(len(A)) # cardinality = 5
# ---- functions and composition
f = lambda x: x + 3
g = lambda x: 2 * x
print(f(g(5)), g(f(5))) # 13 16 -> order matters
# ---- Sum of squares: loop first, then NumPy, then compare
def sum_of_squares_loop(x):
total = 0.0
for xi in x:
total += xi ** 2
return total
def sum_of_squares_np(x):
return float(np.sum(x ** 2)) # or x @ x
x = np.arange(1, 1001)
print(sum_of_squares_loop(x), sum_of_squares_np(x)) # 333833500.0 333833500.0 (the same number)
# ---- indexing: maths counts from 1, NumPy from 0
M = np.array([[11, 12, 13], [21, 22, 23]])
print(M[1, 2], M[1, :], M[:, 2]) # A_{2,3}=23, row 2, column 3
# ---- two equations, two unknowns: x + y = 5, x - y = 1
print(np.linalg.solve(np.array([[1, 1], [1, -1]]), np.array([5, 1]))) # [3. 2.]
# ---- the underflow problem: log of a product, two ways
p = np.full(400, 0.1) # 400 probabilities, each 0.1
print(np.prod(p)) # 0.0 (underflow! true value is 1e-400)
print(np.log(np.prod(p))) # -inf (log of 0, with a warning)
print(np.sum(np.log(p))) # -921.03... (sum of logs is fine)
# ---- exp and log are inverses
v = 3.7
print(np.exp(np.log(v)), np.log(np.exp(v))) # both give 3.7 (up to tiny rounding error, e.g. 3.7000000000000006)
print(np.log(8) / np.log(2)) # log base 2 of 8 = 3
1. If $A = \{1, 2, 3, 4\}$ and $B = \{3, 4, 5\}$, what is $A \cap B$?
2. Which function from $\mathbb{R}$ to $\mathbb{R}$ is injective?
3. What is $\displaystyle\sum_{i=1}^{4} 2i$?
4. Solve $-2x \lt 6$.
5. What is $\log_2 8 + \log_2 4$?
6. Why do we maximise the log-likelihood instead of the likelihood?
Practice problems
A. Write $\{x \in \mathbb{R} : x^2 \lt 4\}$ as an interval.
$x^2 \lt 4$ means $|x| \lt 2$, so $-2 \lt x \lt 2$: the open interval $(-2, 2)$.
B. Compute $\displaystyle\sum_{i=1}^{5} (i^2 - i)$ two ways.
Directly: $(1-1) + (4-2) + (9-3) + (16-4) + (25-5) = 0 + 2 + 6 + 12 + 20 = 40$. With the rules: $\sum i^2 - \sum i = 55 - 15 = 40$ ✓.
C. Solve the system $2x + 3y = 12$ and $x - y = 1$.
From the second, $x = y + 1$. Substitute: $2(y+1) + 3y = 12$, so $5y + 2 = 12$, $y = 2$, and $x = 3$. Check: $2\cdot3 + 3\cdot2 = 12$ ✓.
D. Write $3\ln 2 + \ln 5$ as a single logarithm, and simplify $\log_2 32 - \log_2 4$.
$3\ln 2 = \ln 8$, so the sum is $\ln(8 \cdot 5) = \ln 40$. And $\log_2 32 - \log_2 4 = 5 - 2 = 3 = \log_2 8$.
E. A model gives probability $0.9$ to each of 400 data points. Find the log-likelihood and the likelihood.
$\ell = 400 \ln 0.9 = 400 \times (-0.10536) \approx -42.14$. The likelihood is $e^{-42.14} \approx 5 \times 10^{-19}$: tiny but representable in float64. (float32 can also hold this, but with about a thousand points float32 would underflow to 0.)
F. Is $f(x) = x^2 + 1$ injective on $[0, \infty)$? Find its range and inverse.
On $x \ge 0$ different inputs give different outputs, so yes. The range is $[1, \infty)$. Solve $y = x^2 + 1$: $x = \sqrt{y - 1}$, so $f^{-1}(y) = \sqrt{y-1}$ for $y \ge 1$.
Vectors
A vector is the most basic object in linear algebra. Learn it well and every later chapter gets easier. We will treat a vector two ways at once: as an arrow you can see, and as a list of numbers a computer can store.
- See a vector as both an arrow and a list of numbers
- Add vectors, stretch them, and mix them (linear combinations)
- Measure length (norms) and distance
- Use the dot product to measure agreement, angle and similarity
- Understand perpendicular (orthogonal) vectors and projection
What is a vector? core
Imagine giving someone directions: "walk 3 steps east, then 2 steps north." Two numbers describe the whole trip. That trip is a vector.
You can draw the trip as an arrow that starts at the origin (the point where you started) and ends at the place you reached. Or you can write it as a list: $[3, 2]$. The arrow and the list are the same thing seen two ways.
The list idea is the powerful one. A list does not have to describe a walk. It can describe anything with several measurements.
A house as a vector. Suppose we describe a house with three numbers: its area in square metres, its number of bedrooms, and its age in years.
A house with 120 m², 3 bedrooms and 10 years of age becomes the vector $[120, 3, 10]$.
You can't draw this one as an arrow on paper (it needs three directions at once, one for each measurement), but the maths still works exactly the same. This is how machine learning sees data: every example is a vector of measurements.
A vector in $\mathbb{R}^n$ is an ordered list of $n$ real numbers:
$$\mathbf{v} = \begin{bmatrix} v_1 \\ v_2 \\ \vdots \\ v_n \end{bmatrix}$$The numbers $v_1, v_2, \dots$ are called the components (or entries). The number $n$ is the vector's dimension. We say "$\mathbf{v}$ lives in $\mathbb{R}^n$".
- A column vector is written standing up, as above. This is the default in linear algebra.
- A row vector is written lying down: $\mathbf{v}^\top = [v_1, v_2, \dots, v_n]$. The small $\top$ (a "transpose") flips a column into a row.
- The zero vector $\mathbf{0}$ has every component equal to $0$. It is the "stay where you are" arrow.
Why do we need it?
Computers cannot work with "a house" or "a word". They can only work with numbers. A vector turns anything into an ordered list of numbers, so that maths can be done on it.
Where is it used?
Every row of a dataset, every image, every word embedding and every set of learned model weights is a vector.
How is it used?
Write the measurements down in a fixed order, then use the vector rules in this chapter (add, scale, dot product) on them. A model's prediction is just maths done on feature vectors.
Order matters. $[3, 2]$ and $[2, 3]$ are different vectors (the first goes 3 right and 2 up; the second goes 2 right and 3 up). A vector is a list, not a bag.
Arrow or point? An arrow from the origin and the point where it ends carry exactly the same numbers, so people use the words almost interchangeably. In machine learning we mostly say "point" or "data vector".
Quick check: what is the dimension of $[4, -1, 0, 7]$?
It has four numbers, so the dimension is 4, and the vector lives in $\mathbb{R}^4$. Zeros count as entries too.
Adding and subtracting vectors core
Add vectors the way you chain trips. First do trip $\mathbf{a}$. Then, from wherever you ended up, do trip $\mathbf{b}$. Where you finish compared with where you began is the sum $\mathbf{a}+\mathbf{b}$.
Drawing it: put the tail of $\mathbf{b}$ at the tip of $\mathbf{a}$ ("tip-to-tail"). The sum is the arrow from the very first tail to the very last tip.
Let $\mathbf{a} = [2, 1]$ and $\mathbf{b} = [1, 3]$.
- Add the first components: $2 + 1 = 3$.
- Add the second components: $1 + 3 = 4$.
- So $\mathbf{a} + \mathbf{b} = [3, 4]$.
Data version. Store A sold $[10, 4, 7]$ apples, bananas and cherries on Monday, and store B sold $[3, 6, 2]$. Together they sold $[13, 10, 9]$. Same rule: add matching positions.
For two vectors of the same dimension, add them component by component:
$$\mathbf{a} + \mathbf{b} = \begin{bmatrix} a_1 + b_1 \\ a_2 + b_2 \\ \vdots \\ a_n + b_n \end{bmatrix}, \qquad \mathbf{a} - \mathbf{b} = \begin{bmatrix} a_1 - b_1 \\ \vdots \\ a_n - b_n \end{bmatrix}$$Order does not matter: $\mathbf{a}+\mathbf{b} = \mathbf{b}+\mathbf{a}$. Subtraction is adding the opposite: $\mathbf{a}-\mathbf{b} = \mathbf{a} + (-\mathbf{b})$.
Why do we need it?
We often need to combine two things that are described in the same way: the sales of two stores, or "where I am" plus "how far I moved".
Where is it used?
Adding a bias in a neural-network layer, residual connections ($\mathbf{x} + f(\mathbf{x})$), averaging embeddings, and every training step (new weights = old weights + a step).
How is it used?
Line up matching positions and add each pair. Subtraction does the same and gives the difference: the arrow that goes from one point to another.
You can only add vectors of the same dimension. $[1,2]+[1,2,3]$ makes no sense, because there is nothing to pair the third number with.
Multiplying a vector by a number core
A plain number (like $2$ or $-0.5$) is called a scalar, because it scales things. Multiply a vector by a scalar and you stretch or shrink its arrow:
- Times $2$: twice as long, same direction.
- Times $\tfrac12$: half as long, same direction.
- Times $-1$: same length, opposite direction.
- Times $0$: shrinks to nothing, the zero vector.
Take $\mathbf{v} = [3, 1]$.
- $2\mathbf{v} = [6, 2]$ (double each entry)
- $-1\cdot\mathbf{v} = [-3, -1]$
- $0.5\,\mathbf{v} = [1.5, 0.5]$
For a scalar $c$ and a vector $\mathbf{v}$, multiply every component by $c$:
$$c\,\mathbf{v} = \begin{bmatrix} c\,v_1 \\ c\,v_2 \\ \vdots \\ c\,v_n \end{bmatrix}$$All the multiples of one vector lie on a single straight line through the origin.
Why do we need it?
To change how strong something is without changing what it points at: make an arrow longer, shorter, or turn it around.
Where is it used?
The learning rate times the gradient in training, scaling features to a similar size, and the weights that multiply the inputs of a neuron.
How is it used?
Multiply every entry by the same number. A number above 1 stretches, between 0 and 1 shrinks, and a negative number flips.
Linear combinations core
Think of vectors as ingredients and numbers as how many scoops of each to use. A linear combination is a recipe: "3 scoops of $\mathbf{u}$ plus 2 scoops of $\mathbf{v}$". Stretch each ingredient, then add the results.
This one idea sits underneath almost everything in linear algebra. Matrix multiplication, solving equations and neural-network layers are all linear combinations in disguise.
The two simplest arrows are $\mathbf{e}_1 = [1, 0]$ (one step right) and $\mathbf{e}_2 = [0, 1]$ (one step up). Then $3\mathbf{e}_1 + 2\mathbf{e}_2 = [3, 0] + [0, 2] = [3, 2]$. So every vector in the plane is a linear combination of these two. The numbers in a vector are just "how many scoops" of each.
A less obvious one: let $\mathbf{u} = [2, 1]$ and $\mathbf{v} = [-1, 1]$. Then
$$2\mathbf{u} + 1\mathbf{v} = [4, 2] + [-1, 1] = [3, 3].$$A linear combination of vectors $\mathbf{v}_1, \dots, \mathbf{v}_k$ is any vector of the form
$$c_1\mathbf{v}_1 + c_2\mathbf{v}_2 + \dots + c_k\mathbf{v}_k,$$where $c_1, \dots, c_k$ are scalars (the weights or coefficients).
Why do we need it?
Almost everything in linear algebra is "mix some ingredients". Linear combinations tell us exactly which results we can build from the vectors we have.
Where is it used?
Every neuron computes a weighted sum of its inputs. A linear model's prediction is one. Attention outputs are weighted mixes of value vectors. Solving $A\mathbf{x} = \mathbf{b}$ asks "which mix of the columns gives $\mathbf{b}$?".
How is it used?
Choose a weight for each vector, stretch each vector by its weight, then add the results.
Quick check: is $[5, 5]$ a linear combination of $[1, 1]$ alone? What about $[5, 0]$?
$[5,5] = 5\cdot[1,1]$, so yes (5 scoops). But $[5,0]$ is not: every multiple of $[1,1]$ has equal entries, and $5 \ne 0$. One ingredient can only reach its own line.
The dot product core
The dot product answers one question: how much do two arrows agree about direction?
- Pointing the same way → a big positive number.
- At a right angle (perpendicular) → exactly zero.
- Pointing against each other → a negative number.
Picture the sun directly above. Arrow $\mathbf{b}$ casts a shadow onto the line of arrow $\mathbf{a}$. The dot product is (length of $\mathbf{a}$) × (length of that shadow), counted as negative when the shadow points backwards.
A shopping bill is a dot product. Prices (per item) are $\mathbf{p} = [2, 5, 1]$, and you buy quantities $\mathbf{q} = [3, 1, 4]$. The bill is
$$\mathbf{p}\cdot\mathbf{q} = 2\cdot3 + 5\cdot1 + 1\cdot4 = 6 + 5 + 4 = 15.$$Multiply matching entries, then add everything up.
Geometric one. $\mathbf{a} = [2, 3]$ and $\mathbf{b} = [4, -1]$ give $2\cdot4 + 3\cdot(-1) = 8 - 3 = 5$. Positive, so they roughly agree. And $[1,0]\cdot[0,1] = 0 + 0 = 0$: east and north are perpendicular.
The dot product (or inner product) of two vectors of the same dimension is
$$\mathbf{a}\cdot\mathbf{b} \;=\; \mathbf{a}^\top\mathbf{b} \;=\; a_1b_1 + a_2b_2 + \dots + a_nb_n \;=\; \sum_{i=1}^{n} a_i b_i.$$It has a second, geometric form that gives the same number:
$$\mathbf{a}\cdot\mathbf{b} = \|\mathbf{a}\|\,\|\mathbf{b}\|\cos\theta,$$where $\theta$ is the angle between the arrows and $\|\mathbf{a}\|$ means "the length of $\mathbf{a}$" (next section). The answer is a single number, not a vector.
Rules: $\mathbf{a}\cdot\mathbf{b} = \mathbf{b}\cdot\mathbf{a}$; $(c\mathbf{a})\cdot\mathbf{b} = c\,(\mathbf{a}\cdot\mathbf{b})$; $\mathbf{a}\cdot(\mathbf{b}+\mathbf{c}) = \mathbf{a}\cdot\mathbf{b} + \mathbf{a}\cdot\mathbf{c}$.
Why do we need it?
We need one single number that says how much two things agree. The dot product is that number.
Where is it used?
The weighted sum inside every neuron, attention scores in Transformers, similarity search, a linear model's prediction, and projections.
How is it used?
Multiply matching entries and add them up. A large positive answer means "pointing the same way", zero means "perpendicular", negative means "opposite".
Optional: why do the two formulas always agree?
Draw the triangle with sides $\mathbf{a}$, $\mathbf{b}$ and $\mathbf{a}-\mathbf{b}$. The law of cosines says $\|\mathbf{a}-\mathbf{b}\|^2 = \|\mathbf{a}\|^2 + \|\mathbf{b}\|^2 - 2\|\mathbf{a}\|\|\mathbf{b}\|\cos\theta$.
Now expand the left side using components: $\|\mathbf{a}-\mathbf{b}\|^2 = \sum (a_i-b_i)^2 = \|\mathbf{a}\|^2 + \|\mathbf{b}\|^2 - 2\sum a_ib_i$.
Comparing the two, $\sum a_i b_i = \|\mathbf{a}\|\|\mathbf{b}\|\cos\theta$. ∎
Vector length: the L2 norm core
How long is the arrow $[3, 4]$? Walk 3 steps east and 4 steps north. A bird flying straight to that spot covers less than $3 + 4 = 7$ steps. It flies along the long side of a right triangle, and Pythagoras tells us that side is $\sqrt{3^2 + 4^2} = 5$.
That straight-line length is the vector's norm. It is the "size" of the vector.
- $\|[3, 4]\| = \sqrt{9 + 16} = \sqrt{25} = 5$.
- $\|[1, 1]\| = \sqrt{2} \approx 1.414$.
- It works in 3D too: $\|[1, 2, 2]\| = \sqrt{1 + 4 + 4} = \sqrt{9} = 3$.
The L2 norm (or Euclidean length) of $\mathbf{v}\in\mathbb{R}^n$ is
$$\|\mathbf{v}\|_2 = \sqrt{v_1^2 + v_2^2 + \dots + v_n^2} = \sqrt{\mathbf{v}\cdot\mathbf{v}}.$$Usually we drop the small 2 and just write $\|\mathbf{v}\|$. It is never negative, and it is $0$ only for the zero vector. Stretching by $c$ stretches the length by $|c|$: $\|c\mathbf{v}\| = |c|\,\|\mathbf{v}\|$.
Why do we need it?
We need a way to say how big a vector is: the size of an error, or of a whole set of weights.
Where is it used?
Loss functions (mean squared error is a squared length), weight decay, gradient clipping, and checking whether training has converged.
How is it used?
Square every entry, add, and take the square root. A smaller length of the error vector means a better model.
Other ways to measure size: L1, Lp and L∞ core
"Length" depends on how you are allowed to move.
- A bird flies in a straight line: that is the L2 norm.
- A taxi in a city of square blocks can only drive along streets, so it adds up the horizontal and vertical parts: that is the L1 norm.
- A chess king moves one square in any direction, diagonals included, so the number of moves is the larger of the two parts: that is the L∞ norm.
All three are reasonable ways to say "how big is this vector?". Each one is useful for a different job.
Let $\mathbf{v} = [3, -4]$.
- L1: $|3| + |-4| = 7$
- L2: $\sqrt{9 + 16} = 5$
- L∞: $\max(|3|, |-4|) = 4$
Notice the order: $\|\mathbf{v}\|_\infty \le \|\mathbf{v}\|_2 \le \|\mathbf{v}\|_1$. This is true for every vector.
For $p \ge 1$, the $L_p$ norm is
$$\|\mathbf{v}\|_p = \left(|v_1|^p + |v_2|^p + \dots + |v_n|^p\right)^{1/p}.$$- $p = 1$: $\|\mathbf{v}\|_1 = \sum |v_i|$ (the sum of absolute values)
- $p = 2$: the usual length
- $p \to \infty$: $\|\mathbf{v}\|_\infty = \max_i |v_i|$ (the biggest entry)
Any norm must obey three rules: (1) it is $\ge 0$, and $=0$ only for $\mathbf{0}$; (2) $\|c\mathbf{v}\| = |c|\,\|\mathbf{v}\|$; (3) the triangle inequality $\|\mathbf{a}+\mathbf{b}\| \le \|\mathbf{a}\|+\|\mathbf{b}\|$.
The "L0 norm" counts the non-zero entries. It is handy, but it breaks rule (2), so it is not a true norm.
Why do we need it?
"Size" can mean different things, and the choice changes how a model behaves: the straight-line size, the total size, or the worst single entry.
Where is it used?
L1 (Lasso) gives sparse models with many zero weights. L2 (Ridge) shrinks weights smoothly. L∞ is used for "largest error" and for adversarial examples.
How is it used?
Add a penalty such as $\lambda\|\mathbf{w}\|_1$ to the loss, or measure errors with the norm that matches what you care about.
- L1 regularisation (Lasso) adds $\lambda\|\mathbf{w}\|_1$. Look at the diamond: it has sharp corners on the axes, so the best solutions often land exactly there, with some weights equal to 0. L1 creates sparse models.
- L2 regularisation (Ridge) uses the smooth circle, so weights shrink gently but rarely hit exactly 0.
- L∞ appears in "the largest error" metrics and in adversarial examples ("change each pixel by at most ε").
Distance between two vectors
The distance between two points is the length of the arrow that goes from one to the other. To get that arrow, subtract: the trip from $\mathbf{a}$ to $\mathbf{b}$ is $\mathbf{b} - \mathbf{a}$. Then measure it with any norm.
Points $\mathbf{a} = [1, 2]$ and $\mathbf{b} = [4, 6]$.
- Subtract: $\mathbf{b} - \mathbf{a} = [3, 4]$.
- Measure with L2: $\sqrt{9 + 16} = 5$.
With the L1 norm the same trip is $3 + 4 = 7$ (the taxi distance), and with L∞ it is $\max(3, 4) = 4$.
The order inside doesn't matter, because $\|\mathbf{a}-\mathbf{b}\| = \|\mathbf{b}-\mathbf{a}\|$. Using the L2 norm gives Euclidean distance, L1 gives Manhattan distance, and L∞ gives Chebyshev distance.
Why do we need it?
To find "similar" examples we must say how far apart two examples are.
Where is it used?
k-nearest neighbours, k-means clustering, search, anomaly detection, and every loss that compares a prediction to the truth.
How is it used?
Subtract the two vectors to get the arrow between them, then measure that arrow with a norm.
Unit vectors and normalisation
Sometimes you care only about which way an arrow points, not how long it is. Then rescale it until its length is exactly $1$. That keeps the direction and throws away the size. An arrow of length 1 is a unit vector, and the rescaling is called normalisation.
$\mathbf{v} = [3, 4]$ has length $5$. Divide every entry by $5$:
$$\hat{\mathbf{v}} = \frac{[3, 4]}{5} = [0.6,\; 0.8].$$Check: $\sqrt{0.6^2 + 0.8^2} = \sqrt{0.36 + 0.64} = 1$ ✓. It points the same way as $\mathbf{v}$, just shorter.
A unit vector has $\|\mathbf{u}\| = 1$. To normalise any non-zero vector $\mathbf{v}$:
$$\hat{\mathbf{v}} = \frac{\mathbf{v}}{\|\mathbf{v}\|}.$$You cannot normalise the zero vector (it would mean dividing by 0). The standard unit vectors $\mathbf{e}_1 = [1,0,\dots]$, $\mathbf{e}_2 = [0,1,\dots]$, … each have a single 1 and zeros elsewhere.
Why do we need it?
Sometimes only the direction matters, and the length just gets in the way. A unit vector keeps the direction and removes the size.
Where is it used?
Normalising embeddings before a similarity search, "unit-length" weight or feature vectors, and the direction of a gradient (the gradient divided by its length).
How is it used?
Divide the vector by its own length. The result has length exactly 1 and points the same way.
Angle and cosine similarity core
Two documents: a short tweet about cats, and a long essay about cats. Their word-count vectors have very different lengths (the essay has far more words) but they point in nearly the same direction (both are about cats). So compare the direction and ignore the size. The measure for that is the cosine of the angle between the vectors.
Cosine similarity is $+1$ when they point the same way, $0$ when perpendicular (nothing in common), and $-1$ when exactly opposite.
Vocabulary: [cat, dog, car]. Document A has counts $[3, 1, 0]$ and document B has counts $[2, 2, 0]$.
- Dot product: $3\cdot2 + 1\cdot2 + 0\cdot0 = 8$.
- Lengths: $\|\mathbf{A}\| = \sqrt{10} \approx 3.162$ and $\|\mathbf{B}\| = \sqrt{8} \approx 2.828$.
- Cosine: $8 / (3.162 \times 2.828) \approx 0.894$.
Quite similar. Now a document C about cars with counts $[0, 0, 4]$: its dot product with A is $0$, so the cosine is $0$ (nothing in common).
Rearranging the geometric formula $\mathbf{a}\cdot\mathbf{b} = \|\mathbf{a}\|\|\mathbf{b}\|\cos\theta$ gives
$$\cos\theta = \frac{\mathbf{a}\cdot\mathbf{b}}{\|\mathbf{a}\|\,\|\mathbf{b}\|}, \qquad \text{cosine similarity}(\mathbf{a},\mathbf{b}) = \cos\theta \in [-1, 1].$$The angle itself is $\theta = \arccos(\cdot)$. Both vectors must be non-zero.
Scaling a vector (by a positive number) changes its length but not its angle, so cosine similarity ignores size.
Why do we need it?
To compare the topic or meaning of two things regardless of their size, such as a short and a long document about the same subject.
Where is it used?
Semantic search and RAG (finding the most relevant text), recommenders, duplicate detection, and clustering of text.
How is it used?
Divide the dot product by the two lengths (or take the dot product of two unit vectors). Close to 1 means very similar.
Orthogonal (perpendicular) vectors core
Orthogonal is the mathematician's word for perpendicular: a right angle. East and north are orthogonal. Walking east tells you nothing about how far north you went. Orthogonal directions are completely independent, with zero overlap or redundancy.
Are $[1, 2]$ and $[-2, 1]$ orthogonal? Compute the dot product: $1\cdot(-2) + 2\cdot1 = -2 + 2 = 0$. Yes.
A handy trick in 2D: to get a vector perpendicular to $[x, y]$, swap the entries and flip one sign: $[-y, x]$.
Because $\cos 90^\circ = 0$, the formula $\mathbf{a}\cdot\mathbf{b} = \|\mathbf{a}\|\|\mathbf{b}\|\cos\theta$ gives exactly this. The zero vector is orthogonal to everything. If, in addition, both vectors have length 1, they are called orthonormal.
Why do we need it?
Perpendicular directions never interfere with each other. They carry independent information, which makes calculations simpler and more stable.
Where is it used?
The axes found by PCA, orthonormal bases and QR, orthogonal weight initialisation, and least squares (the leftover error is perpendicular to the features).
How is it used?
Test with the dot product: zero means perpendicular. To build perpendicular vectors from ordinary ones we use the Gram–Schmidt recipe (Chapter 1.9).
Orthogonal does not mean opposite. Opposite arrows have dot product $-\|\mathbf{a}\|\|\mathbf{b}\|$, which is as non-zero as it gets. Orthogonal means a dot product of exactly zero: the arrows are at a right angle.
Vector projection core
The projection of $\mathbf{b}$ onto $\mathbf{a}$ is the shadow of $\mathbf{b}$ on the line of $\mathbf{a}$, with the light coming from directly above (at a right angle to $\mathbf{a}$). It answers: "How much of $\mathbf{b}$ points along $\mathbf{a}$?"
What is left over, the part of $\mathbf{b}$ that sticks out sideways, is always perpendicular to $\mathbf{a}$. So projection splits $\mathbf{b}$ into two pieces: one along $\mathbf{a}$ and one at a right angle to it.
Project $\mathbf{b} = [3, 1]$ onto $\mathbf{a} = [1, 1]$.
- $\mathbf{a}\cdot\mathbf{b} = 3 + 1 = 4$ and $\mathbf{a}\cdot\mathbf{a} = 1 + 1 = 2$.
- The scaling factor is $4/2 = 2$.
- The shadow is $2\cdot[1, 1] = [2, 2]$.
- The leftover is $\mathbf{b} - [2,2] = [1, -1]$. Check: $[1,-1]\cdot[1,1] = 0$ ✓ perpendicular.
The number $\dfrac{\mathbf{a}\cdot\mathbf{b}}{\|\mathbf{a}\|}$ is the scalar projection: the signed length of the shadow. The leftover $\mathbf{b} - \operatorname{proj}_{\mathbf{a}}(\mathbf{b})$ is called the residual or error, and it is orthogonal to $\mathbf{a}$.
Why do we need it?
To find the best approximation of a vector using only one direction, and to know exactly what is left over.
Where is it used?
Least squares and linear regression (project the targets onto what the features can express), PCA (project data onto a few directions), removing an unwanted component such as bias from an embedding.
How is it used?
Compute $(\mathbf{a}\cdot\mathbf{b})/(\mathbf{a}\cdot\mathbf{a})$ and stretch $\mathbf{a}$ by it. The leftover is perpendicular to $\mathbf{a}$.
Two famous inequalities
- Triangle inequality. A detour can never be shorter than the straight path. Going along $\mathbf{a}$ then $\mathbf{b}$ cannot beat going straight to the end: $\|\mathbf{a}+\mathbf{b}\| \le \|\mathbf{a}\| + \|\mathbf{b}\|$.
- Cauchy–Schwarz. The shadow can never be longer than the arrow that casts it. So the dot product can never exceed the product of the lengths: $|\mathbf{a}\cdot\mathbf{b}| \le \|\mathbf{a}\|\,\|\mathbf{b}\|$.
$\mathbf{a} = [3, 0]$, $\mathbf{b} = [0, 4]$. Then $\|\mathbf{a}+\mathbf{b}\| = \|[3,4]\| = 5$, while $\|\mathbf{a}\|+\|\mathbf{b}\| = 3 + 4 = 7$, and $5 \le 7$ ✓.
Also $|\mathbf{a}\cdot\mathbf{b}| = 0 \le 3\cdot4 = 12$ ✓.
For all vectors $\mathbf{a}, \mathbf{b}$:
$$|\mathbf{a}\cdot\mathbf{b}| \le \|\mathbf{a}\|\,\|\mathbf{b}\| \quad\text{(Cauchy–Schwarz)} \qquad \|\mathbf{a}+\mathbf{b}\| \le \|\mathbf{a}\|+\|\mathbf{b}\| \quad\text{(triangle)}$$Equality in the triangle inequality holds exactly when the vectors point the same way (or one is $\mathbf{0}$). Equality in Cauchy–Schwarz holds exactly when the vectors lie on one line: the same way or opposite ways (or one is $\mathbf{0}$). Cauchy–Schwarz is why cosine similarity always lands between $-1$ and $1$.
Why do we need it?
These two rules guarantee that our measures of size and similarity behave sensibly: cosine cannot exceed 1, and a detour cannot beat the straight path. That is what lets us prove that algorithms work.
Where is it used?
Proofs for k-NN and clustering, error bounds, the fact that cosine similarity stays between −1 and 1, and convergence proofs in optimisation.
How is it used?
Use them as safety checks and to put an upper bound on a quantity: for example, $|\mathbf{a}\cdot\mathbf{b}|$ is never more than $\|\mathbf{a}\|\|\mathbf{b}\|$.
Vectors in 3D (and beyond)
So far the arrows lived on a flat page. The real world has one more direction: up. A drone flying "3 metres east, 1 metre north and 2 metres up" needs three numbers, so its move is a vector with three entries.
Everything you learned still works. You add vectors by adding matching entries, you stretch them by multiplying every entry, and you find the dot product by multiplying matching entries and adding. The only new thing is the picture: arrows now live in a room, not on a page.
And what about 100 numbers, like an embedding? You cannot draw 100 directions, but the same formulas work unchanged. A 3D picture is the best "stand-in" we have for how high-dimensional vectors behave.
Let $\mathbf{a} = [3, 1, 1]$ and $\mathbf{b} = [0, 3, 2]$.
- Sum: $\mathbf{a} + \mathbf{b} = [3+0,\; 1+3,\; 1+2] = [3, 4, 3]$.
- Dot product: $3\cdot0 + 1\cdot3 + 1\cdot2 = 5$ (positive, so they roughly agree).
- Length of $\mathbf{a}$: $\sqrt{9 + 1 + 1} = \sqrt{11} \approx 3.32$.
A vector in $\mathbb{R}^3$ is a list of three numbers $[x, y, z]$: how far along the $x$-axis (east), the $y$-axis (north) and the $z$-axis (up). Every formula from this chapter holds with one extra term:
$$\mathbf{a}\cdot\mathbf{b} = a_1b_1 + a_2b_2 + a_3b_3, \qquad \|\mathbf{a}\| = \sqrt{a_1^2 + a_2^2 + a_3^2}.$$In $\mathbb{R}^n$ you have $n$ terms.
Why do we need it?
Real data lives in many dimensions. Learning to see three of them gives you the right instincts for hundreds.
Where is it used?
3D graphics and robotics (positions and directions), colour (red, green, blue is a 3-entry vector), and every picture of "a plane inside a bigger space" in later chapters.
How is it used?
Use exactly the same formulas, with more entries. Use pictures like these to build a feel for what the formulas mean.
Quick check: what is $[1, 2, 3]\cdot[4, -5, 2]$?
$1\cdot4 + 2\cdot(-5) + 3\cdot2 = 4 - 10 + 6 = 0$. They are perpendicular, even though you can't tell by looking at the numbers.
Recap, cheat sheet and practice
- A vector is a list of numbers, and also an arrow. Machine-learning data points are vectors.
- Add vectors entry by entry (tip-to-tail). Multiply by a number to stretch, shrink or flip.
- A linear combination is "some scoops of each vector, added up". It is the core move of linear algebra.
- The dot product $\mathbf{a}\cdot\mathbf{b} = \sum a_ib_i = \|\mathbf{a}\|\|\mathbf{b}\|\cos\theta$ measures how much two vectors agree.
- A norm measures size: L1 (taxi), L2 (straight line), L∞ (biggest entry). Distance is the norm of the difference.
- Normalising gives length 1. Cosine similarity compares direction only. Orthogonal means dot product = 0. Projection is the shadow.
Cheat sheet
| Idea | Formula | Picture |
|---|---|---|
| Dot product | $\sum a_ib_i$ | how much they agree |
| L2 length | $\sqrt{\mathbf{v}\cdot\mathbf{v}}$ | straight-line size |
| L1 / L∞ | $\sum|v_i|$ / $\max|v_i|$ | taxi / chess-king size |
| Distance | $\|\mathbf{a}-\mathbf{b}\|$ | arrow from b to a |
| Normalise | $\mathbf{v}/\|\mathbf{v}\|$ | shrink to length 1 |
| Cosine similarity | $\dfrac{\mathbf{a}\cdot\mathbf{b}}{\|\mathbf{a}\|\|\mathbf{b}\|}$ | cos of the angle |
| Orthogonal | $\mathbf{a}\cdot\mathbf{b}=0$ | right angle |
| Projection of b on a | $\dfrac{\mathbf{a}\cdot\mathbf{b}}{\mathbf{a}\cdot\mathbf{a}}\mathbf{a}$ | shadow of b on a |
import numpy as np
a = np.array([2, 3])
b = np.array([4, -1])
print(a + b) # [6 2] entry-by-entry addition
print(3 * a) # [6 9] scaling
print(a @ b) # 5 dot product (same as np.dot(a, b)): 2*4 + 3*(-1)
print(np.linalg.norm(a)) # L2 length = sqrt(13) = 3.6055...
print(np.linalg.norm(a, 1)) # L1 norm = 5
print(np.linalg.norm(a, np.inf)) # L-infinity = 3
dist = np.linalg.norm(a - b) # Euclidean distance
unit = a / np.linalg.norm(a) # normalise
cos = a @ b / (np.linalg.norm(a) * np.linalg.norm(b)) # cosine similarity
proj = (a @ b) / (a @ a) * a # projection of b onto a
print(dist) # 4.4721... = sqrt(20), the length of [-2, 4]
print(unit) # [0.5547 0.8321] (length 1, same direction as a)
print(cos) # 0.3363... (the angle is about 70.3 degrees)
print(proj) # [0.7692 1.1538] = (5/13) * a
1. What is $[1, 2, 3] + [4, 0, -1]$?
2. The dot product of $[2, 5]$ and $[-5, 2]$ is…
3. Which statement about cosine similarity is true?
4. For $\mathbf{v} = [-6, 8]$, which is correct?
5. The projection of $[3, 4]$ onto $[1, 0]$ is…
Practice problems
A. Normalise $[2, -1, 2]$.
Length: $\sqrt{4+1+4} = 3$. So $\hat{\mathbf{v}} = [2/3,\,-1/3,\,2/3]$.
B. Find the cosine similarity of $[1, 0, 1]$ and $[1, 1, 0]$, and the angle.
Dot $= 1$. Lengths are $\sqrt2$ each, so $\cos\theta = 1/2$ and $\theta = 60^\circ$.
C. Find $c$ so that $[c, 3]$ is orthogonal to $[4, -2]$.
$4c - 6 = 0$, so $c = 1.5$.
D. Project $[4, 2]$ onto $[3, 1]$ and give the leftover.
$\mathbf{a}\cdot\mathbf{b} = 14$, $\mathbf{a}\cdot\mathbf{a} = 10$, so $k = 1.4$ and the projection is $[4.2, 1.4]$. The leftover is $[4,2]-[4.2,1.4] = [-0.2, 0.6]$, and $[-0.2,0.6]\cdot[3,1] = -0.6+0.6 = 0$ ✓.
E. Why is it always true that $\|\mathbf{v}\|_\infty \le \|\mathbf{v}\|_2 \le \|\mathbf{v}\|_1$?
The biggest entry can't exceed the length of the whole vector (the other entries only add to it). And squaring then square-rooting the entries can't exceed adding their absolute values: $(\sum|v_i|)^2 = \sum v_i^2 + (\text{non-negative cross terms}) \ge \sum v_i^2$.
Vector Spaces
In Chapter 1.2 we met vectors one at a time. Now we zoom out and look at the room they live in. We will learn which collections of vectors form a tidy "space", how to describe a whole space with a few well-chosen vectors (a basis), and how the same vector can be written in different "languages".
- Know what a vector space is, and spot things that are not one
- Decide whether a set is a subspace (lines and planes through the origin)
- Describe everything a set of vectors can reach (span), and detect redundancy (linear dependence)
- Choose a basis, find a space's dimension, and write vectors in coordinates for any basis
- Change from one basis to another, and meet shifted subspaces (affine spaces)
What is a vector space?
In Chapter 1.2 the two basic moves were adding vectors and stretching them. A vector space is any collection of objects where those two operations make sense, never throw you out of the collection, and behave the way you expect.
Think of a room whose walls you can never leave by mixing. Take any two things in the room, add them: still in the room. Stretch one by any number: still in the room.
The surprise: the objects do not have to be arrows or lists of numbers. Polynomials, functions, matrices and images can all be "vectors", as long as you can add them and scale them. That is why linear algebra is so widely useful.
- Lists: $\mathbb{R}^2$. $[1,2] + [3,-1] = [4,1]$ and $2\cdot[1,2] = [2,4]$.
- Polynomials of degree at most 2, like $p(x) = 1 + x$ and $q(x) = -2x + x^2$. Their sum is $1 - x + x^2$, still a polynomial of degree at most 2. And $2p(x) = 2 + 2x$. Each polynomial is described by three coefficients $[c_0, c_1, c_2]$, so $p = [1, 1, 0]$ and $q = [0, -2, 1]$. Adding polynomials is adding these lists.
- Functions: all functions $\mathbb{R} \to \mathbb{R}$. Add them point by point: $(\sin + x^2)(t) = \sin t + t^2$. Scale: $(3\sin)(t) = 3\sin t$.
- Matrices of one fixed size, e.g. $2 \times 2$: $\begin{bmatrix}1&2\\3&4\end{bmatrix} + \begin{bmatrix}1&0\\0&1\end{bmatrix} = \begin{bmatrix}2&2\\3&5\end{bmatrix}$. A $2\times2$ matrix is secretly a list of 4 numbers.
Non-examples (the room has an exit):
- Vectors $[x, y]$ with $x \ge 0$ and $y \ge 0$ (the "first quadrant"). Stretching $[1,1]$ by $-1$ gives $[-1,-1]$: you left the set.
- Vectors of length exactly 1. $[1,0] + [0,1] = [1,1]$ has length $\sqrt2$: you left the set.
- Polynomials of degree exactly 2. $x^2 + (-x^2 + x) = x$ has degree 1: you left.
- Invertible matrices (square matrices that can be "undone"; you will meet them in Chapter 1.7). $I + (-I) = 0$, and the zero matrix is not invertible.
A vector space (over the real numbers) is a set $V$ of objects called vectors, with addition and scalar multiplication, such that:
- Closure: if $\mathbf{u}, \mathbf{v} \in V$ and $c$ is a real number, then $\mathbf{u} + \mathbf{v} \in V$ and $c\mathbf{u} \in V$.
and these eight "common sense" rules hold for all $\mathbf{u}, \mathbf{v}, \mathbf{w} \in V$ and scalars $c, d$:
| # | Rule | In plain words |
|---|---|---|
| 1 | $\mathbf{u} + \mathbf{v} = \mathbf{v} + \mathbf{u}$ | order of adding does not matter |
| 2 | $(\mathbf{u} + \mathbf{v}) + \mathbf{w} = \mathbf{u} + (\mathbf{v} + \mathbf{w})$ | grouping does not matter |
| 3 | there is $\mathbf{0}$ with $\mathbf{v} + \mathbf{0} = \mathbf{v}$ | there is a "do nothing" vector |
| 4 | for each $\mathbf{v}$ there is $-\mathbf{v}$ with $\mathbf{v} + (-\mathbf{v}) = \mathbf{0}$ | every vector can be undone |
| 5 | $c(\mathbf{u} + \mathbf{v}) = c\mathbf{u} + c\mathbf{v}$ | stretching a sum = sum of the stretches |
| 6 | $(c + d)\mathbf{v} = c\mathbf{v} + d\mathbf{v}$ | scoops can be combined |
| 7 | $c(d\mathbf{v}) = (cd)\mathbf{v}$ | stretching twice = stretching by the product |
| 8 | $1\,\mathbf{v} = \mathbf{v}$ | stretching by 1 changes nothing |
For things like $\mathbb{R}^n$, polynomials, functions and matrices, the eight rules come for free from ordinary arithmetic. In practice, you only need to check closure.
Why do we need it?
We want to add and scale many kinds of things, not only arrows: data rows, polynomials, images, functions. A vector space names exactly the setting where those two operations always work.
Where is it used?
The space of weight vectors of a model, the space of images (ℝ^784 for 28×28 pixels), function spaces behind kernel methods and Fourier series, and spaces of polynomial features.
How is it used?
To decide whether something is a vector space, check that adding two members and scaling a member always stays in the set, and that the zero vector is there. The other eight rules usually come free from ordinary arithmetic.
- The zero vector must be in the space. Every vector space has a $\mathbf{0}$ (for example $0\cdot\mathbf{v}$). A set that does not contain it cannot be a vector space.
- Closure is the usual point of failure. When a set fails to be a vector space, it is almost always because adding or scaling can leave it.
- "Vector" now means "any element of a vector space". In this chapter we mostly use $\mathbb{R}^n$, but everything is true for polynomials and functions too.
Quick check: is the set of all $2\times2$ matrices with trace $0$ (diagonal adds to 0) closed under addition and scaling?
Yes. If two matrices each have diagonal sum $0$, their sum has diagonal sum $0 + 0 = 0$, and scaling by $c$ gives $c\cdot 0 = 0$. It also contains the zero matrix. (It is a vector space, in fact a subspace of all $2 \times 2$ matrices.)
Subspaces core
A subspace is a smaller vector space sitting inside a bigger one. In the plane $\mathbb{R}^2$, the subspaces are: just the origin, any line through the origin, or the whole plane. In $\mathbb{R}^3$ you also get the planes through the origin.
Why the origin? Because stretching any vector by $0$ gives the zero vector. So a subspace must contain it, and a flat sheet that misses the origin cannot be one.
Picture a sheet of glass passing through the origin of a room. Add two arrows lying in the glass, and the result still lies in the glass. Stretch an arrow in the glass, and it stays in the glass. Shift the glass one metre to the side, and this stops working.
Is the line $y = 2x$ a subspace of $\mathbb{R}^2$? Take two points on it, $\mathbf{u} = [1, 2]$ and $\mathbf{v} = [3, 6]$.
- Contains $\mathbf{0}$? $[0,0]$ has $0 = 2\cdot0$ ✓.
- Sum: $\mathbf{u}+\mathbf{v} = [4, 8]$, and $8 = 2\cdot4$ ✓ still on the line.
- Scale: $-3\mathbf{u} = [-3,-6]$, and $-6 = 2\cdot(-3)$ ✓.
Now the shifted line $y = 2x + 1$. The origin is not on it: $0 \ne 2\cdot0 + 1$. That alone settles it. Also $[0, 1] + [1, 3] = [1, 4]$, but $2\cdot1 + 1 = 3 \ne 4$, so even the sum leaves the line.
A subset $S$ of a vector space $V$ is a subspace if
- $\mathbf{0} \in S$ (it contains the zero vector),
- $\mathbf{u}, \mathbf{v} \in S \implies \mathbf{u} + \mathbf{v} \in S$ (closed under addition),
- $\mathbf{v} \in S,\ c \in \mathbb{R} \implies c\mathbf{v} \in S$ (closed under scalar multiplication).
Equivalently: $S$ is closed under linear combinations. Every subspace is itself a vector space. The two extreme subspaces are $\{\mathbf{0}\}$ and $V$ itself.
Why do we need it?
Many important sets inside a bigger space are flat sheets through the origin: all outputs of a linear model, all solutions of Ax = 0. We need a simple test for "is this a proper space on its own?".
Where is it used?
The column space of a data matrix (what a linear model can predict), the null space (directions the model cannot see), and the low-dimensional subspace that PCA fits to data.
How is it used?
Run the three tests: the zero vector is in the set, the sum of two members is in the set, and every multiple of a member is in the set. One failing example proves it is not a subspace.
- Passing the sample tests is not a proof. To show a set is a subspace you must argue for all vectors in it. To show it is not, one failing example is enough.
- A subspace is "flat" and goes on forever. A circle, a disk and a square are not subspaces.
- The union of two subspaces (all vectors that lie in one or the other) is usually not a subspace: the two axes are an example. Their sum (all vectors $\mathbf{u}+\mathbf{v}$ with $\mathbf{u}$ in the first and $\mathbf{v}$ in the second) is a subspace.
Quick check: is the plane $x + y + z = 0$ in $\mathbb{R}^3$ a subspace? What about $x + y + z = 1$?
$x+y+z=0$: contains $\mathbf{0}$; if two vectors each have coordinates adding to 0, so does their sum, and so does any multiple. Yes. $x+y+z=1$: the origin gives $0 \ne 1$, so it is not a subspace (it is a shifted plane; see affine spaces at the end).
Span: everything you can reach core
Recall from linear combinations in Chapter 1.2: a linear combination is a recipe, "some scoops of each ingredient vector, added up". The span of some vectors is the collection of every vector you can make this way, using any amounts (including negative and fractional scoops).
It is the answer to the question "where can I get to with these ingredients?"
- One non-zero vector: you can only slide back and forth along its line.
- Two vectors that point in different directions: you can reach every point of a flat sheet (a plane).
- Three vectors in 3D, not lying in one flat sheet: you can reach every point in space.
- But if one vector is just a stretched copy of another (or lies in the sheet of the others), it adds nothing new. The span stays smaller. You saw this in the "Dependent pair" mode of the linear combinations widget in Chapter 1.2.
- In $\mathbb{R}^2$, $\operatorname{span}\{[1,0],[0,1]\} = \mathbb{R}^2$, since $[x,y] = x[1,0] + y[0,1]$.
- $\operatorname{span}\{[1,2]\}$ is the line $y = 2x$ (every multiple $c[1,2] = [c, 2c]$).
- $\operatorname{span}\{[1,2],[2,4]\}$ is also just the line $y = 2x$, because $[2,4] = 2[1,2]$ is redundant.
- In $\mathbb{R}^3$, $\operatorname{span}\{\mathbf{e}_1, \mathbf{e}_2\}$ is the flat "floor" (the $xy$-plane). Is $[3,4,0]$ in it? Yes: $3\mathbf{e}_1 + 4\mathbf{e}_2$. Is $[0,0,1]$? No: no combination of floor vectors lifts you off the floor.
- A vector $\mathbf{b}$ is in the span exactly when the equation $c_1\mathbf{v}_1 + \dots + c_k\mathbf{v}_k = \mathbf{b}$ has a solution.
- The span is always a subspace: it contains $\mathbf{0}$ (all $c_i = 0$), and sums or multiples of combinations are combinations. It is the smallest subspace containing the $\mathbf{v}_i$.
- The span of no vectors is $\{\mathbf{0}\}$.
Why do we need it?
We often ask "what can I build from these ingredients?". The span answers it: it is the full set of everything reachable by mixing the given vectors.
Where is it used?
The outputs a linear model can produce (the span of the data columns), what a neural-network layer W x can reach, and the question "can we fit the target exactly?" in least squares.
How is it used?
To test whether a vector b is in the span, try to solve c₁v₁ + … + cₖvₖ = b for the scoops. If there is a solution, b is reachable (the solution is the recipe). In code: np.linalg.solve or compare ranks.
- More vectors does not mean a bigger span. $k$ vectors span at most a $k$-dimensional space, and fewer if some are redundant.
- "The span of $\mathbf{v}_1, \mathbf{v}_2$" is a set of vectors (a line, a plane, ...), not a single vector.
- To test "is $\mathbf{b}$ in the span?", you must solve for the scoops. Staring at the numbers is not enough (this becomes routine in Chapter 1.6).
Quick check: is $[2, 4, 6]$ in $\operatorname{span}\{[1,2,3]\}$? Is $[1, 0, 0]$?
$[2,4,6] = 2[1,2,3]$, so yes. $[1,0,0]$ would need $c[1,2,3] = [1,0,0]$, so $c = 1$ from the first entry but $2c = 0$ from the second. Contradiction, so no.
Linear independence and dependence core
A list of vectors is independent if each one brings something new. It is dependent if there is redundancy: some vector can be built from the others, so it could be thrown away without shrinking the span.
Think of a recipe cupboard. "Flour, sugar, eggs" are independent: none can be made from the others. Add "flour mixed with sugar" and the list becomes dependent, since the new item is just a mix of two others. A data analogy: a table with columns "price in dollars" and "price in cents" has one column too many, because cents is just $100 \times$ dollars.
Another way to say it: independent vectors can only combine to the zero vector in the boring way, using zero scoops of everything.
- Dependent: $\mathbf{u} = [1, 2]$ and $\mathbf{v} = [2, 4]$. Here $2\mathbf{u} - \mathbf{v} = [2,4] - [2,4] = \mathbf{0}$ with scoops $(2, -1)$, not both zero.
- Independent: $\mathbf{e}_1 = [1,0]$, $\mathbf{e}_2 = [0,1]$. $c_1\mathbf{e}_1 + c_2\mathbf{e}_2 = [c_1, c_2] = [0,0]$ forces $c_1 = c_2 = 0$.
- Three vectors in $\mathbb{R}^2$ are always dependent: $[1,0], [0,1], [3,2]$ satisfy $3[1,0] + 2[0,1] - 1[3,2] = [0,0]$.
- In $\mathbb{R}^3$: $\mathbf{u} = [1,0,1]$, $\mathbf{v} = [0,1,1]$, $\mathbf{w} = [1,1,2]$. Since $\mathbf{w} = \mathbf{u} + \mathbf{v}$, we have $\mathbf{u} + \mathbf{v} - \mathbf{w} = \mathbf{0}$: dependent.
How to test by hand. Write $c_1\mathbf{u} + c_2\mathbf{v} + c_3\mathbf{w} = \mathbf{0}$ as equations and solve. For the last example: $c_1 + c_3 = 0$, $c_2 + c_3 = 0$, $c_1 + c_2 + 2c_3 = 0$. The first two give $c_1 = c_2 = -c_3$, and the third is then automatically true. So $c_3$ is free: choose $c_3 = -1$ to get $(1, 1, -1)$, a non-zero solution, hence dependent.
Vectors $\mathbf{v}_1, \dots, \mathbf{v}_k$ are linearly independent if
$$c_1\mathbf{v}_1 + c_2\mathbf{v}_2 + \dots + c_k\mathbf{v}_k = \mathbf{0} \quad\Longrightarrow\quad c_1 = c_2 = \dots = c_k = 0.$$If some non-zero choice of the $c_i$ gives $\mathbf{0}$, they are linearly dependent. Handy facts:
- Dependent $\iff$ at least one vector is a linear combination of the others (it is redundant).
- A list containing the zero vector is dependent. Two vectors are dependent exactly when they are parallel.
- More than $n$ vectors in $\mathbb{R}^n$ are always dependent.
- $k$ vectors are independent exactly when the matrix with those vectors as columns has rank $k$ (Chapter 1.7), or for $n$ vectors in $\mathbb{R}^n$, when its determinant is non-zero (a single number computed from a square matrix; Chapter 1.7). The rank is the number of independent columns.
Why do we need it?
Redundant ingredients waste effort and make answers ambiguous. Linear independence tells us whether every vector in a list brings something new.
Where is it used?
Detecting multicollinearity (duplicate or near-duplicate features) before linear regression, choosing a minimal set of features, and checking whether a set of vectors can serve as a basis.
How is it used?
Stack the vectors as columns of a matrix and compute its rank with np.linalg.matrix_rank. If the rank equals the number of vectors, they are independent. If not, the null space shows the recipe for zero.
- Pairwise is not enough. $[1,0,1], [0,1,1], [1,1,2]$ has no two vectors parallel, yet the three together are dependent ($\mathbf{w} = \mathbf{u}+\mathbf{v}$). Dependence can hide in a group.
- Independence is a property of the whole list, not of one vector. And order does not matter.
- In real data, columns are rarely exactly dependent: they are nearly dependent, which causes numerical trouble (Chapter 1.15).
Quick check: are $[1,2,3]$, $[4,5,6]$, $[7,8,9]$ independent?
No. Notice $[7,8,9] = 2[4,5,6] - [1,2,3]$, since $[8-1,\ 10-2,\ 12-3] = [7,8,9]$. So $[1,2,3] - 2[4,5,6] + [7,8,9] = \mathbf{0}$: dependent, even though no two vectors are parallel.
Basis: the perfect set of ingredients core
A basis of a space is a set of vectors with two properties at once:
- Enough (they span the space): you can reach every vector.
- No waste (they are independent): no vector can be built from the others.
Think of the three primary colours of paint: every colour is a mix of them, and none of the three can be mixed from the other two. A basis is the smallest complete toolkit.
The wonderful consequence: with a basis, every vector has exactly one recipe. The scoops become the vector's "address" (its coordinates, next section).
The most familiar basis is the standard basis: "one step east" and "one step north" in the plane. But a basis is a choice, and there are infinitely many good ones: tilted, stretched, squeezed.
- $\{[1,0],\ [0,1]\}$ is a basis of $\mathbb{R}^2$ (the standard basis $\mathbf{e}_1, \mathbf{e}_2$).
- $\{[1,1],\ [-1,1]\}$ is also a basis of $\mathbb{R}^2$: not parallel, so independent, and (as we saw with span) two non-parallel vectors in the plane reach everything.
- $\{[1,0],\ [0,1],\ [1,1]\}$ spans $\mathbb{R}^2$ but is not a basis: it is dependent ($[1,1] = [1,0] + [0,1]$), one vector too many.
- $\{[1,0]\}$ is independent but not a basis of $\mathbb{R}^2$: it only reaches the $x$-axis, one vector too few.
One recipe only. In the basis $\{[1,1],\ [-1,1]\}$, the vector $[3,1]$ needs $c_1[1,1] + c_2[-1,1] = [3,1]$, so $c_1 - c_2 = 3$ and $c_1 + c_2 = 1$. Adding: $2c_1 = 4$, so $c_1 = 2$ and $c_2 = -1$. No other pair works.
A basis of a vector space $V$ is a list of vectors $\mathbf{b}_1, \dots, \mathbf{b}_k$ that is
- linearly independent, and
- spans $V$.
Equivalently: every $\mathbf{v} \in V$ can be written as $c_1\mathbf{b}_1 + \dots + c_k\mathbf{b}_k$ in exactly one way.
The standard basis of $\mathbb{R}^n$ is $\mathbf{e}_1, \dots, \mathbf{e}_n$, where $\mathbf{e}_i$ has a 1 in position $i$ and 0 elsewhere. For example, in $\mathbb{R}^3$: $\mathbf{e}_1 = [1,0,0]$, $\mathbf{e}_2 = [0,1,0]$, $\mathbf{e}_3 = [0,0,1]$.
Non-uniqueness. A space has many bases (rescale or tilt the vectors). What stays fixed is how many vectors a basis has (next section). A handy test: $n$ vectors in $\mathbb{R}^n$ form a basis exactly when the matrix with them as columns has a non-zero determinant.
Why do we need it?
To describe every vector in a space we want the smallest toolkit that still reaches everything, with exactly one recipe for each vector. That is a basis.
Where is it used?
The standard feature-by-feature basis of any dataset, the principal directions found by PCA, Fourier and wavelet bases for signals, and polynomial bases in curve fitting.
How is it used?
Pick vectors that are independent and span the space; for n vectors in ℝⁿ check that the determinant is not zero. Then every vector has one set of coordinates in that basis.
- A basis is an ordered list when we talk about coordinates: $(\mathbf{b}_1, \mathbf{b}_2)$ and $(\mathbf{b}_2, \mathbf{b}_1)$ give coordinates in different orders.
- A basis does not need perpendicular or unit-length vectors. Perpendicular unit vectors are a special, extra-nice kind (an orthonormal basis, Chapter 1.9).
- Too few vectors cannot span, too many cannot be independent. A basis is the Goldilocks number.
Quick check: do $[1,2]$ and $[3,6]$ form a basis of $\mathbb{R}^2$? Do $[1,2]$ and $[3,5]$?
$[3,6] = 3[1,2]$: parallel, so no. For $[1,2]$ and $[3,5]$: $\det = 1\cdot5 - 3\cdot2 = -1 \ne 0$, so yes.
Dimension core
The dimension of a space is the number of independent directions in it, or equivalently, how many numbers you need to pin down a point. A line has 1, a flat sheet has 2, ordinary space has 3.
The key fact is that every basis of a space has the same number of vectors. You can swap one basis for another, but never change the count. That count is the dimension.
Beware one trap: a sheet of paper floating in a 3D room has points described by 3 numbers $[x,y,z]$, but the sheet itself is only 2-dimensional, because two numbers (two scoops) are enough to say where you are on the sheet.
- $\mathbb{R}^n$ has dimension $n$ (the standard basis has $n$ vectors).
- The plane through the origin spanned by two non-parallel vectors in $\mathbb{R}^3$ has dimension 2.
- $\operatorname{span}\{[1,2],[2,4]\}$ has dimension 1, not 2, since the vectors are dependent.
- Polynomials of degree at most 2: $\{1,\ x,\ x^2\}$ is a basis, so the dimension is 3. The polynomial $5 - 2x + 3x^2$ has "coordinates" $[5, -2, 3]$.
- $2\times2$ matrices: the four matrices with a single 1 and three 0s form a basis. Dimension 4.
- $\{\mathbf{0}\}$ has dimension 0 (its basis is the empty list). The space of all functions $\mathbb{R}\to\mathbb{R}$ is infinite-dimensional.
The dimension $\dim V$ of a vector space is the number of vectors in any basis of $V$. Consequences, in a space of dimension $n$:
- Any $n$ independent vectors form a basis. Any $n$ spanning vectors form a basis.
- More than $n$ vectors are always dependent. Fewer than $n$ vectors cannot span.
- A subspace $S$ of $V$ has $\dim S \le \dim V$, with equality only when $S = V$.
Why do we need it?
We need one number that says how "big" a space is, in the sense of how many independent directions it has. That number does not depend on which basis we choose.
Where is it used?
The embedding size of a model, the number of weights to learn, the effective dimension of data (rank), and dimensionality reduction with PCA and autoencoders.
How is it used?
Count the vectors in any basis, or compute the rank of a matrix whose columns span the space (np.linalg.matrix_rank). For noisy data, count the singular values that are clearly above zero.
- Dimension of a space is not the length of its vectors. A plane in $\mathbb{R}^3$ has vectors of length 3 but dimension 2.
- The dimension is a property of the space; a basis is just one way to describe it.
- "High-dimensional" in ML usually means many coordinates. Whether the data is effectively high-dimensional depends on the rank, as in the widget.
Quick check: what is the dimension of $\operatorname{span}\{[1,0,0],\ [1,1,0],\ [2,1,0]\}$?
The third vector is the sum of the first two, so it is redundant. The first two are independent, so the dimension is 2 (the floor, the $xy$-plane).
Coordinates relative to a basis core
A vector is a thing, like a place on a map. Its coordinates are just its address in a chosen language. The standard basis is one language ("3 steps east, 1 step north"). Another basis is another language ("2 steps along $\mathbf{b}_1$, then $-1$ steps along $\mathbf{b}_2$"). The place is the same; only the description changes. It is like giving a temperature in Celsius or Fahrenheit: one temperature, two numbers.
Because every vector has exactly one recipe in a basis, the scoops $(c_1, c_2, \dots)$ are a perfect, unambiguous address.
Basis $B = \{\mathbf{b}_1, \mathbf{b}_2\}$ with $\mathbf{b}_1 = [1,1]$ and $\mathbf{b}_2 = [-1,1]$. The vector $\mathbf{v} = [3, 1]$ (standard coordinates) equals $2\mathbf{b}_1 - 1\mathbf{b}_2$ (we solved this in the last section). So
$$[\mathbf{v}]_B = \begin{bmatrix} 2 \\ -1 \end{bmatrix}.$$Check: $2[1,1] - [-1,1] = [2+1,\ 2-1] = [3,1]$ ✓.
How to find coordinates. Put the basis vectors as the columns of a matrix $P = \begin{bmatrix}1&-1\\1&1\end{bmatrix}$. The equation "$c_1\mathbf{b}_1 + c_2\mathbf{b}_2 = \mathbf{v}$" is the linear system $P\mathbf{c} = \mathbf{v}$. Solve it.
Let $B = (\mathbf{b}_1, \dots, \mathbf{b}_n)$ be a basis of $V$. For $\mathbf{v} \in V$, the coordinate vector $[\mathbf{v}]_B = (c_1, \dots, c_n)$ is the unique list with
$$\mathbf{v} = c_1\mathbf{b}_1 + \dots + c_n\mathbf{b}_n = P_B\,[\mathbf{v}]_B, \qquad P_B = \begin{bmatrix} | & & | \\ \mathbf{b}_1 & \cdots & \mathbf{b}_n \\ | & & | \end{bmatrix}.$$So $[\mathbf{v}]_B = P_B^{-1}\mathbf{v}$. Here $P_B^{-1}$ is the "undo" matrix of $P_B$ (inverses are in Chapter 1.7). In practice you just solve $P_B\mathbf{c} = \mathbf{v}$. The coordinates in the standard basis are just the entries of $\mathbf{v}$ itself.
Why do we need it?
A vector is a thing; to compute with it we need numbers. Coordinates turn a vector into an unambiguous list of numbers in whichever basis we like.
Where is it used?
PCA scores (coordinates along principal directions), pixel values (coordinates in the pixel basis), word embeddings (coordinates in a learned basis of meaning), and Fourier coefficients.
How is it used?
Put the basis vectors as columns of a matrix P and solve P c = v, for example with np.linalg.solve(P, v). The solution c is the coordinate list [v]_B.
- Always say which basis the numbers refer to. $[3,1]$ means $[3,1]$ in the standard basis, but "the vector with coordinates $(3,1)$ in $B$" is a different arrow.
- Coordinates depend on the order of the basis vectors. Swap $\mathbf{b}_1$ and $\mathbf{b}_2$ and the two numbers swap too.
- Do not confuse the vector with its coordinates. The arrow does not move when you change basis; only its numbers change.
Quick check: find the coordinates of $[7, 4]$ in the basis $\{[2,1],\ [1,1]\}$.
Solve $2c_1 + c_2 = 7$ and $c_1 + c_2 = 4$. Subtract: $c_1 = 3$, then $c_2 = 1$. So $[\mathbf{v}]_B = (3, 1)$. Check: $3[2,1] + 1[1,1] = [7,4]$ ✓.
Change of basis core
Same vector, different "language". A change of basis is a translator between two descriptions of the same arrow, like converting metres to feet. The arrow itself never moves. Only the numbers we use to describe it change.
The translator is a matrix:
- The matrix $P$ whose columns are the basis vectors turns B-coordinates into standard coordinates ("spend $c_1$ scoops of $\mathbf{b}_1$ and $c_2$ of $\mathbf{b}_2$ and see where you land").
- Its inverse $P^{-1}$ translates the other way: standard to B.
- To go between two non-standard bases $B$ and $C$: translate to standard, then on to $C$.
With $B = \{[1,1],\ [-1,1]\}$, $P = \begin{bmatrix}1&-1\\1&1\end{bmatrix}$ and its inverse $P^{-1} = \tfrac12\begin{bmatrix}1&1\\-1&1\end{bmatrix}$ (check: $PP^{-1} = I$).
- B to standard. If $[\mathbf{v}]_B = (2,-1)$, then $\mathbf{v} = P\begin{bmatrix}2\\-1\end{bmatrix} = [2\cdot1 + (-1)(-1),\ 2\cdot1 + (-1)\cdot1] = [3, 1]$ ✓.
- Standard to B. If $\mathbf{v} = [5,3]$, then $[\mathbf{v}]_B = P^{-1}\begin{bmatrix}5\\3\end{bmatrix} = \tfrac12[5+3,\ -5+3] = [4, -1]$. Check: $4[1,1] - [-1,1] = [5,3]$ ✓.
- B to another basis $C$. Let $C = \{[1,0],\ [1,1]\}$, with matrix $P_C = \begin{bmatrix}1&1\\0&1\end{bmatrix}$ and inverse $\begin{bmatrix}1&-1\\0&1\end{bmatrix}$. The B-to-C translator is $Q = P_C^{-1}P_B = \begin{bmatrix}0&-2\\1&1\end{bmatrix}$. Apply it to $(2,-1)$: $[0\cdot2 + (-2)(-1),\ 1\cdot2 + 1\cdot(-1)] = [2, 1]$. Check: $2[1,0] + 1[1,1] = [3,1]$, the same vector ✓.
Let $P_B$ and $P_C$ be the matrices whose columns are the vectors of bases $B$ and $C$ (written in standard coordinates). Then for every vector $\mathbf{v}$:
$$\mathbf{v} = P_B[\mathbf{v}]_B, \qquad [\mathbf{v}]_B = P_B^{-1}\mathbf{v}, \qquad [\mathbf{v}]_C = \underbrace{P_C^{-1}P_B}_{\text{change-of-basis matrix}}\,[\mathbf{v}]_B.$$The matrix $P_C^{-1}P_B$ converts $B$-coordinates to $C$-coordinates. It is invertible, and its inverse converts back.
(Matrix multiplication and inverses are covered properly in Chapters 1.4 and 1.7. For now, read $P\mathbf{c}$ as "take $c_1$ scoops of column 1, $c_2$ scoops of column 2, and add", exactly the linear combination idea.)
Why do we need it?
The same data can be easier to understand in a different basis. A change of basis is the translator that converts coordinates from one description to another without moving the vector.
Where is it used?
PCA (re-express data along the principal axes), whitening and feature scaling, converting between colour spaces, and reading the weights of a layer in a more meaningful basis.
How is it used?
Build P_B and P_C from the basis vectors, then use [v]_C = P_C⁻¹ P_B [v]_B. In code: np.linalg.solve(PC, PB @ c). Remember that P (columns as basis) maps basis coordinates to standard coordinates.
- Which direction? $P$ (basis vectors as columns) maps B-coordinates to standard. People often expect the opposite. Look at what each side of the equation is.
- A change of basis changes the description, not the vector. (A transformation, Chapter 1.5, moves vectors. These two ideas are related but different.)
- Order matters: $P_C^{-1}P_B$ and $P_B^{-1}P_C$ are different, and in fact inverses of each other.
Quick check: if $P = \begin{bmatrix}2&0\\0&1\end{bmatrix}$ has columns forming the basis $B$, what are the standard coordinates of the vector with $[\mathbf{v}]_B = (3, 4)$?
$\mathbf{v} = P[\mathbf{v}]_B = [2\cdot3 + 0\cdot4,\ 0\cdot3 + 1\cdot4] = [6, 4]$. (It is $3\mathbf{b}_1 + 4\mathbf{b}_2 = 3[2,0] + 4[0,1]$.)
Affine spaces: shifted subspaces
We saw that a line must pass through the origin to be a subspace. But lines and planes that miss the origin are everywhere (think of the surface of a table). Take a subspace and slide it away from the origin: you get an affine subspace. It is flat and straight like a subspace, only it has been picked up and moved.
Everything is described by one point plus one subspace: "start at a point $\mathbf{p}$, then move around inside the subspace $S$". Every affine subspace looks the same from any of its points; it just has no special origin.
A second view: an affine combination is a mix whose scoops add up to exactly 1, like "0.3 of $\mathbf{a}$ and 0.7 of $\mathbf{b}$". These mixes never need the origin: they depend only on the points themselves.
- The solutions of $x + y = 3$ form a line missing the origin. Take the subspace $S = \operatorname{span}\{[1,-1]\}$ (the line $x + y = 0$) and the point $\mathbf{p} = [1, 2]$. Then $\mathbf{p} + S = \{[1+t,\ 2-t]\}$: for $t = 1$ we get $[2,1]$, for $t = -1$ we get $[0,3]$. All satisfy $x+y=3$ ✓.
- It is not a subspace: $\mathbf{0}$ is not in it ($0 + 0 \ne 3$). But the difference of any two of its points lies in $S$: $[2,1] - [1,2] = [1,-1] \in S$ ✓.
- Affine combination. For points $\mathbf{a} = [1,1]$ and $\mathbf{b} = [3,2]$, the mix $(1-t)\mathbf{a} + t\mathbf{b}$ has weights summing to $(1-t) + t = 1$. At $t = 0.5$ it is the midpoint $[2, 1.5]$. At $t = 2$ it is $2\mathbf{b} - \mathbf{a} = [5, 3]$, a point beyond $\mathbf{b}$. As $t$ varies you trace out the whole line through $\mathbf{a}$ and $\mathbf{b}$.
- An affine subspace of $V$ is a set $A = \mathbf{p} + S = \{\mathbf{p} + \mathbf{s} : \mathbf{s} \in S\}$, where $\mathbf{p} \in V$ is a point and $S$ is a subspace. $S$ is its direction space, and $\dim A = \dim S$. If $\mathbf{p} \in S$, then $A = S$ is itself a subspace.
- An affine combination of points $\mathbf{x}_1, \dots, \mathbf{x}_k$ is $c_1\mathbf{x}_1 + \dots + c_k\mathbf{x}_k$ with $c_1 + \dots + c_k = 1$. Affine subspaces are exactly the sets closed under affine combinations.
- If we also require every $c_i \ge 0$, we get a convex combination: the points between them (a segment for two points, a triangle for three).
- A hyperplane $\{\mathbf{x} : \mathbf{w}^\top\mathbf{x} + b = 0\}$ in $\mathbb{R}^n$ is an affine subspace of dimension $n - 1$ (as long as $\mathbf{w} \ne \mathbf{0}$). It is a subspace only if $b = 0$.
Why do we need it?
Real flat things such as decision boundaries and solution sets rarely pass through the origin. Affine spaces describe shifted subspaces and mixes whose weights add to 1.
Where is it used?
The bias b in w·x + b (an affine function), decision boundaries of linear classifiers (hyperplanes), solution sets of Ax = b, and interpolation between two embeddings or model weights.
How is it used?
Describe a shifted flat as one point plus a subspace: p + S. For mixes, use weights that sum to 1, and optionally non-negative weights to stay between the points. In code: (1 - t) * a + t * b.
- An affine subspace is not a vector space (unless it passes through the origin). Adding two of its points leaves it; subtracting them lands in the direction space $S$.
- Affine combinations make sense for points; the "weights add to 1" rule is what removes the dependence on where the origin is.
- "Linear" and "affine" are different words with different meanings. $y = 2x + 1$ is affine; $y = 2x$ is linear (Chapter 1.1, linear functions; Chapter 1.5, affine transformations).
Quick check: is $\{[x,y] : y = 3x - 2\}$ an affine subspace? A subspace? Give its point and direction.
It is a line, so it is an affine subspace: $\mathbf{p} + \operatorname{span}\{\mathbf{d}\}$ with $\mathbf{p} = [0, -2]$ and $\mathbf{d} = [1, 3]$. It is not a subspace, since $\mathbf{0}$ is not on it ($0 \ne 3\cdot0 - 2$).
Recap, cheat sheet and practice
- A vector space is any set where adding and scaling never leave the set (plus the eight common-sense rules). $\mathbb{R}^n$, polynomials, functions and matrices all qualify.
- A subspace contains $\mathbf{0}$ and is closed under $+$ and scaling: lines and planes through the origin.
- The span of some vectors is everything their linear combinations reach. It is always a subspace.
- Vectors are independent if the only combination giving $\mathbf{0}$ is the trivial one; otherwise there is redundancy.
- A basis is independent and spanning: every vector has exactly one recipe. The dimension is the number of vectors in any basis.
- Coordinates $[\mathbf{v}]_B = P_B^{-1}\mathbf{v}$ are the vector's address in basis $B$. A change of basis $P_C^{-1}P_B$ translates between addresses; the vector itself never moves.
- An affine subspace $\mathbf{p} + S$ is a shifted subspace; affine combinations have weights summing to 1.
Cheat sheet
| Idea | Test or formula | Picture |
|---|---|---|
| Subspace | $\mathbf{0} \in S$; closed under $+$ and $c\cdot$ | line / plane through the origin |
| Span | $\{c_1\mathbf{v}_1 + \dots + c_k\mathbf{v}_k\}$ | everything you can reach |
| Independent | $\sum c_i\mathbf{v}_i = \mathbf{0} \Rightarrow$ all $c_i = 0$ · rank $= k$ | no redundancy |
| Basis | independent + spanning · $n$ vectors in $\mathbb{R}^n$ with $\det \ne 0$ | new graph paper |
| Dimension | number of basis vectors | independent directions |
| Coordinates | $[\mathbf{v}]_B = P_B^{-1}\mathbf{v}$ (solve $P_B\mathbf{c} = \mathbf{v}$) | address in a language |
| Change of basis | $[\mathbf{v}]_C = P_C^{-1}P_B[\mathbf{v}]_B$ | translator |
| Affine subspace | $\mathbf{p} + S$; weights sum to 1 | shifted line/plane |
import numpy as np
# ---- test independence with the rank
def is_independent(vectors):
A = np.column_stack(vectors) # each vector becomes a column
return np.linalg.matrix_rank(A) == A.shape[1]
u, v, w = np.array([1, 0, 1]), np.array([0, 1, 1]), np.array([1, 1, 2])
print(is_independent([u, v])) # True
print(is_independent([u, v, w])) # False (w = u + v)
# ---- coordinates in a non-standard basis: solve P c = x
P = np.array([[1, -1],
[1, 1]]) # columns are b1 = (1,1) and b2 = (-1,1)
x = np.array([3, 1])
c = np.linalg.solve(P, x) # [ 2. -1.]
print(c, P @ c) # [ 2. -1.] [3. 1.]
# ---- change of basis from B (matrix P) to C (matrix PC)
PC = np.array([[1, 1],
[0, 1]])
Q = np.linalg.solve(PC, P) # PC^{-1} P
print(Q) # [[ 0. -2.] [ 1. 1.]]
print(Q @ c) # [2. 1.] = coordinates of x in C
# ---- is a vector in the span? the rank does not grow when you add it
def in_span(vectors, t):
A = np.column_stack(vectors)
return np.linalg.matrix_rank(np.column_stack([A, t])) == np.linalg.matrix_rank(A)
print(in_span([u, v], np.array([2, 3, 5]))) # True (2u + 3v)
print(in_span([u, v], np.array([0, 0, 1]))) # False
# ---- dimension of data = numerical rank (5 features, but only 2 real directions)
rng = np.random.default_rng(0)
X = rng.standard_normal((100, 2)) @ rng.standard_normal((2, 5))
print(np.linalg.matrix_rank(X)) # 2
# ---- affine combination (weights add to 1)
a, b = np.array([1.0, 1.0]), np.array([3.0, 2.0])
print(0.5 * a + 0.5 * b, 2 * b - 1 * a) # [2. 1.5] [5. 3.]
1. Which of these is a subspace of $\mathbb{R}^2$?
2. What is $\operatorname{span}\{[1,2],\ [2,4]\}$?
3. Are $[1,0,1]$, $[0,1,1]$ and $[1,1,2]$ linearly independent?
4. What is the dimension of the plane $x + y + z = 0$ in $\mathbb{R}^3$?
5. In the basis $B = \{[1,1],\ [-1,1]\}$, what are the coordinates of $[5, 3]$?
6. Which is an affine combination of two points $\mathbf{a}$ and $\mathbf{b}$?
Practice problems
A. Is $\{[x,y,z] : x + y + z = 1\}$ a subspace of $\mathbb{R}^3$? What kind of set is it?
No: $[0,0,0]$ gives $0 \ne 1$. It is a plane that misses the origin, an affine subspace, equal to $[1,0,0] + \{x+y+z=0\}$.
B. Do $[1,2,3]$, $[4,5,6]$, $[7,8,9]$ form a basis of $\mathbb{R}^3$?
No. $[7,8,9] = 2[4,5,6] - [1,2,3]$, so they are dependent (rank 2). They span only a plane.
C. Find the coordinates of $[7,4]$ in the basis $\{[2,1],\ [1,1]\}$.
$2c_1 + c_2 = 7$ and $c_1 + c_2 = 4$ give $c_1 = 3$, $c_2 = 1$. So $[\mathbf{v}]_B = (3,1)$. Check: $3[2,1] + [1,1] = [7,4]$ ✓.
D. Give a basis of the polynomials of degree at most 2, and the coordinates of $3 - 2x + 5x^2$.
A basis is $\{1,\ x,\ x^2\}$ (dimension 3). The coordinates of $3 - 2x + 5x^2$ are $(3, -2, 5)$.
E. Write the line through $\mathbf{a} = [1,1]$ and $\mathbf{b} = [3,2]$ as $\mathbf{p} + t\mathbf{d}$, and check that $[5,3]$ is on it.
Take $\mathbf{p} = \mathbf{a} = [1,1]$ and $\mathbf{d} = \mathbf{b} - \mathbf{a} = [2,1]$. At $t = 2$: $[1,1] + 2[2,1] = [5,3]$ ✓ (this is also the affine combination $2\mathbf{b} - \mathbf{a}$).
F. Vectors $\mathbf{v}_1, \mathbf{v}_2, \mathbf{v}_3, \mathbf{v}_4$ are in $\mathbb{R}^3$. Can they be independent? Can they span $\mathbb{R}^3$?
They cannot be independent: more than 3 vectors in $\mathbb{R}^3$ are always dependent. They can span $\mathbb{R}^3$ (for example, if three of them form a basis, the fourth is redundant).
Matrices
A matrix is a table of numbers. It is also a stack of vectors, and a machine that moves arrows around. In this chapter you will learn to read matrices, add them, multiply them in four different ways, and meet the special matrices that appear again and again in machine learning.
- See a matrix as a table, a stack of vectors, a transformation and a data set
- Read sizes like $m \times n$ and entries like $A_{ij}$
- Add, scale and multiply matrices element by element (Hadamard product)
- Multiply a matrix by a vector, and see it as "a mix of the columns"
- Multiply two matrices in four equivalent ways, and know why order matters
- Use the transpose, the Gram matrix, and the special matrices: identity, diagonal, symmetric, triangular, orthogonal, permutation, sparse and block
- Recognise a neural-network layer as $W\mathbf{x} + \mathbf{b}$
What is a matrix? core
Think of a spreadsheet. It has rows going across and columns going down, and a number in every cell. A matrix is exactly that: a rectangular table of numbers.
But a matrix wears four different hats, and good machine-learning thinking means switching between them freely:
- A table. Numbers in rows and columns, like a school timetable or a price list.
- A stack of vectors. Each row is a vector. Or: each column is a vector. Same numbers, two ways to slice them.
- A transformation. A machine: feed in a vector, get a new vector out. (You will study this properly in Chapter 1.5. Here is a first taste.)
- A data set. Each row is one example (one house, one photo, one customer). Each column is one measurement.
Take the table $A = \begin{bmatrix} 2 & -1 \\ 1 & 1 \end{bmatrix}$.
- As a stack of rows: the vectors $[2, -1]$ and $[1, 1]$, one on top of the other.
- As a stack of columns: the vectors $\begin{bmatrix}2\\1\end{bmatrix}$ and $\begin{bmatrix}-1\\1\end{bmatrix}$, side by side.
- As a machine: feed in the arrow $[1, 0]$ (one step right) and out comes $[2, 1]$, which is the first column. Feed in $[0, 1]$ (one step up) and out comes $[-1, 1]$, the second column. So the columns tell you where the two basic arrows go.
A matrix is a rectangular array of numbers arranged in rows and columns. We write matrices with a capital letter and square brackets:
$$A = \begin{bmatrix} a_{11} & a_{12} & \cdots & a_{1n} \\ a_{21} & a_{22} & \cdots & a_{2n} \\ \vdots & \vdots & & \vdots \\ a_{m1} & a_{m2} & \cdots & a_{mn} \end{bmatrix}$$The numbers inside are the entries (or elements). A matrix with $m$ rows and $n$ columns, whose entries are real numbers, is written $A \in \mathbb{R}^{m \times n}$.
Why do we need it?
People need one tool for tables of numbers, for moving arrows around, and for whole data sets. A matrix is that one tool, so one set of rules covers all three jobs.
Where is it used?
Spreadsheets and CSV files, grayscale images, the weight matrix of every neural-network layer, graph adjacency tables, and the data matrix X that scikit-learn and PyTorch expect.
How is it used?
Pick the hat that fits your question. To store numbers, think table. To ask what it does to a vector, think transformation. To lay out data, put one sample in each row and one feature in each column.
A matrix is not just "a long vector". The same six numbers arranged as $2 \times 3$ or as $3 \times 2$ are two different matrices, and they behave differently. The shape is part of the object.
Quick check: the matrix $\begin{bmatrix} 3 & 0 \\ 0 & 3 \end{bmatrix}$ is fed the arrow $[1, 0]$. Using the "columns" idea, where does it land?
On the first column, $[3, 0]$. The arrow is stretched to three times its length. (The second column $[0, 3]$ tells you the up-arrow also gets stretched by 3.)
Size and notation: $m \times n$, $A_{ij}$, rows and columns core
To find a seat in a cinema you need two numbers: which row, and which seat in that row. A matrix entry works the same way. Every entry has an address made of two numbers, and the row always comes first.
The size of a matrix is also "rows first, then columns". A table with 3 rows and 4 columns is a "3 by 4" matrix.
Let $A = \begin{bmatrix} 5 & 1 & 7 \\ 2 & 0 & 4 \end{bmatrix}$.
- Count the rows: 2. Count the columns: 3. So $A$ is a $2 \times 3$ matrix, with $2 \cdot 3 = 6$ entries.
- The entry in row 1, column 2 is $A_{12} = 1$.
- The entry in row 2, column 1 is $A_{21} = 2$.
- Row 1 is the vector $[5, 1, 7]$. Column 3 is the vector $\begin{bmatrix}7\\4\end{bmatrix}$.
$A \in \mathbb{R}^{m \times n}$ means $m$ rows and $n$ columns. Notation we will use:
- $A_{ij}$ (or $a_{ij}$) is the entry in row $i$, column $j$.
- The $i$-th row is written $\mathbf{r}_i^\top$ (a row vector with $n$ numbers).
- The $j$-th column is written $\mathbf{a}_j$ (a column vector with $m$ numbers).
- If $m = n$ the matrix is square.
- A vector is a matrix with one column ($n \times 1$) or, as a row vector, one row ($1 \times n$). A single number is a $1 \times 1$ matrix.
You can think of $A$ as the columns placed side by side: $A = [\,\mathbf{a}_1 \;\; \mathbf{a}_2 \;\; \cdots \;\; \mathbf{a}_n\,]$.
Why do we need it?
To talk about one number inside a big table, and to check that two tables can be combined, we need a clear address for each entry and a clear size for the whole table.
Where is it used?
Every line of NumPy or PyTorch that reads A[i, j] or prints A.shape, and every shape error message in deep learning, such as 'mat1 and mat2 shapes cannot be multiplied'.
How is it used?
Write the shape (rows, columns) next to each matrix before you compute. Use A[i, j] for one entry, A[i] for a row and A[:, j] for a column. Remember that Python counts from 0, not from 1.
Rows first, then columns. $A_{23}$ means row 2, column 3, never the other way round. And "$2 \times 3$" has 2 rows and 3 columns.
Some books number entries from 1 (like us) and some from 0 (like Python). In NumPy, A[0, 1] is the entry we call $A_{12}$.
Quick check: a matrix is $4 \times 7$. How many entries does it have, and how many numbers are in one column?
It has $4 \cdot 7 = 28$ entries. One column runs down through all the rows, so it has 4 numbers. (One row has 7.)
The data matrix $X$ core
In Chapter 1.2 we saw that one house can be described by a vector of measurements. Now stack many houses. If you write each house as one row, the whole collection becomes a matrix.
Rows answer "which example?". Columns answer "which measurement?". This table is the standard way a machine-learning model sees its training data.
Four houses, three features each: area in m², bedrooms, age in years.
$$X = \begin{bmatrix} 120 & 3 & 10 \\ 85 & 2 & 25 \\ 200 & 4 & 5 \\ 60 & 1 & 40 \end{bmatrix}$$- $X$ is $4 \times 3$: 4 samples, 3 features.
- Row 3 is house number 3: the feature vector $[200, 4, 5]$.
- Column 2 holds the bedrooms of every house: $[3, 2, 4, 1]$.
The prices we want to predict are stored separately, as a vector $\mathbf{y}$ with one number per house.
A data matrix is $X \in \mathbb{R}^{n \times d}$:
- $n$ = the number of samples (rows),
- $d$ = the number of features (columns),
- row $i$ is the feature vector $\mathbf{x}_i^\top$ of sample $i$, and entry $X_{ij}$ is feature $j$ of sample $i$.
In general matrices the letters are $m \times n$. In data science people say $n \times d$. They are just different names for "rows $\times$ columns".
Why do we need it?
A model must look at many examples at once. Putting one example in each row lets us store and process the whole data set as one single object.
Where is it used?
Linear and logistic regression, k-means, PCA, scikit-learn's fit(X, y), and every deep-learning mini-batch, for example 32 images flattened into a 32 by 784 matrix.
How is it used?
Write the measurements of each example as one row, stack the rows, and check that X.shape is (n_samples, n_features). Column statistics such as the mean are taken down each column, and models multiply X by their weights.
Some books put samples in columns instead. Then the data matrix is $d \times n$. Both conventions exist. In this guide (and in NumPy, scikit-learn and PyTorch) one sample = one row. Always check the shape before you multiply.
Quick check: 1000 customers, each described by 20 numbers. What is the shape of $X$, and what is row 7?
$X$ is $1000 \times 20$. Row 7 is the feature vector of customer 7: a vector with 20 numbers.
Adding, scaling and the Hadamard product core
Operations that treat every cell on its own are the easy ones. Lay two tables of the same size on top of each other, and combine the numbers that sit at the same address.
- Add / subtract: add (or subtract) matching cells.
- Scale: multiply every cell by the same number.
- Hadamard product: multiply matching cells.
Think of two sales tables (store A and store B, items in columns, days in rows). Adding them gives the combined sales. Scaling by 1.1 adds ten percent to everything.
Let $A = \begin{bmatrix} 1 & 2 \\ 3 & 4 \end{bmatrix}$ and $B = \begin{bmatrix} 5 & 6 \\ 7 & 8 \end{bmatrix}$.
- $A + B = \begin{bmatrix} 1+5 & 2+6 \\ 3+7 & 4+8 \end{bmatrix} = \begin{bmatrix} 6 & 8 \\ 10 & 12 \end{bmatrix}$
- $A - B = \begin{bmatrix} -4 & -4 \\ -4 & -4 \end{bmatrix}$
- $3A = \begin{bmatrix} 3 & 6 \\ 9 & 12 \end{bmatrix}$
- $A \odot B = \begin{bmatrix} 1\cdot5 & 2\cdot6 \\ 3\cdot7 & 4\cdot8 \end{bmatrix} = \begin{bmatrix} 5 & 12 \\ 21 & 32 \end{bmatrix}$
For matrices $A, B$ of the same size and a scalar $c$:
$$(A + B)_{ij} = A_{ij} + B_{ij}, \qquad (A - B)_{ij} = A_{ij} - B_{ij}, \qquad (cA)_{ij} = c\,A_{ij}$$The Hadamard product (element-wise product) is
$$(A \odot B)_{ij} = A_{ij}\,B_{ij}.$$Addition is commutative ($A + B = B + A$) and associative. Scaling spreads over addition: $c(A + B) = cA + cB$. The matrices must have the same size, or the operation is not defined.
Why do we need it?
Combining two tables of the same shape cell by cell is the simplest operation of all, and it shows up everywhere: totals, updates and masks.
Where is it used?
Residual connections in ResNets and Transformers (X + F(X)), the gradient-descent update W - lr * G, dropout and attention masks (A times a 0/1 matrix), and the gates of LSTM networks.
How is it used?
Make sure both matrices have the same shape (NumPy will also stretch a single row or number to fit). Use + and - for sums, a number times a matrix for scaling, and * (not @) for the Hadamard product.
"Multiply" can mean two different things. The Hadamard product $A \odot B$ multiplies matching cells. The matrix product $AB$ (next sections) mixes whole rows with whole columns. They give different answers. In NumPy, A * B is Hadamard and A @ B is the matrix product.
Quick check: can you add a $2 \times 3$ matrix and a $3 \times 2$ matrix?
No. Addition pairs up cells at the same address, and the shapes do not match, so some cells have no partner. Both matrices must have exactly the same size.
Matrix times vector: two ways to read $A\mathbf{x}$ core
A matrix is a machine. A vector goes in, a vector comes out. The output is written $A\mathbf{x}$. There are two honest ways to see what the machine does, and you should be able to switch between them without thinking.
- Row view: a set of scores. Each row of $A$ is a "question". Take the dot product of that row with $\mathbf{x}$ (see the dot product). That gives one output number. Do it for every row.
- Column view: a recipe. Each column of $A$ is an ingredient. The numbers in $\mathbf{x}$ say how many scoops of each. Stretch each column by its scoop count and add them all up. That is a linear combination of the columns.
Both views give the same answer. The column view is the one that explains what a matrix does, so it is the one to remember.
Let $A = \begin{bmatrix} 2 & 1 \\ 0 & 3 \\ 1 & 1 \end{bmatrix}$ (3 rows, 2 columns) and $\mathbf{x} = \begin{bmatrix} 3 \\ 2 \end{bmatrix}$.
Row view (dot product of each row with $\mathbf{x}$):
- Row 1: $2\cdot3 + 1\cdot2 = 6 + 2 = 8$.
- Row 2: $0\cdot3 + 3\cdot2 = 0 + 6 = 6$.
- Row 3: $1\cdot3 + 1\cdot2 = 3 + 2 = 5$.
Column view (3 scoops of column 1, 2 scoops of column 2):
$$3\begin{bmatrix}2\\0\\1\end{bmatrix} + 2\begin{bmatrix}1\\3\\1\end{bmatrix} = \begin{bmatrix}6\\0\\3\end{bmatrix} + \begin{bmatrix}2\\6\\2\end{bmatrix} = \begin{bmatrix}8\\6\\5\end{bmatrix}.$$Same result: $A\mathbf{x} = [8, 6, 5]$. Notice the shapes: $\mathbf{x}$ has 2 numbers, and the output has 3. A $3 \times 2$ matrix turns 2-dimensional vectors into 3-dimensional ones.
For $A \in \mathbb{R}^{m \times n}$ and $\mathbf{x} \in \mathbb{R}^n$, the product $A\mathbf{x}$ is a vector in $\mathbb{R}^m$.
Row view: the $i$-th entry is the dot product of row $i$ with $\mathbf{x}$:
$$(A\mathbf{x})_i = \sum_{j=1}^{n} A_{ij}\,x_j = \mathbf{r}_i^\top \mathbf{x}.$$Column view: a combination of the columns of $A$, with the entries of $\mathbf{x}$ as the weights:
$$A\mathbf{x} = x_1\,\mathbf{a}_1 + x_2\,\mathbf{a}_2 + \cdots + x_n\,\mathbf{a}_n.$$Shape rule: the number of columns of $A$ must equal the number of entries of $\mathbf{x}$. (One scoop-count for every column.)
Why do we need it?
We need a way to feed a vector into a matrix machine and get a vector out. Reading the product two ways, by rows and by columns, lets us pick the view that explains the question.
Where is it used?
Every prediction of a linear model (the vector of outputs is W times x), every neuron layer before its bias, embedding lookup (a matrix times a one-hot vector), and solving Ax = b.
How is it used?
To compute by hand, use the row view: one dot product per row. To understand what happens, use the column view: scale each column by the matching entry of x and add. In NumPy write A @ x and check A.shape[1] equals x.shape[0].
Shapes must fit. A $3 \times 2$ matrix can multiply a vector with 2 entries, not 3. If $A$ has $n$ columns, $\mathbf{x}$ must have $n$ entries.
$A\mathbf{x}$ is not $\mathbf{x}A$. We always write the matrix on the left of a column vector. (Writing a row vector on the left, $\mathbf{x}^\top A$, is a different calculation that mixes the rows of $A$.)
Quick check: compute $\begin{bmatrix} 1 & 2 \\ 3 & 4 \end{bmatrix}\begin{bmatrix} 2 \\ 1 \end{bmatrix}$ both ways.
Row view: $[1\cdot2 + 2\cdot1,\; 3\cdot2 + 4\cdot1] = [4, 10]$. Column view: $2\begin{bmatrix}1\\3\end{bmatrix} + 1\begin{bmatrix}2\\4\end{bmatrix} = \begin{bmatrix}2\\6\end{bmatrix} + \begin{bmatrix}2\\4\end{bmatrix} = \begin{bmatrix}4\\10\end{bmatrix}$. Same answer.
Matrix times matrix: four views of one product core
Matrix-times-vector feeds one vector to the machine $A$. Matrix-times-matrix feeds a whole stack of vectors at once. If $B$ is a stack of column vectors, then $AB$ is "run every one of them through $A$ and keep the outputs in a stack".
That is the column view below. There are four ways to organise the same arithmetic. They all give the identical matrix, but each one answers a different question:
- Entry view: "What is this one number?" (a dot product)
- Column view: "What happens to each input vector?" ($A$ times each column of $B$)
- Row view: "How is each output row built?" (a row of $A$ mixes the rows of $B$)
- Outer-product view: "Can I build the answer from simple pieces?" (a sum of simple "rank-1" matrices: each piece is one column of $A$ times one row of $B$. Chapter 1.7 explains rank.)
Let $A = \begin{bmatrix} 1 & 2 \\ 3 & 4 \end{bmatrix}$ and $B = \begin{bmatrix} 5 & 6 \\ 7 & 8 \end{bmatrix}$. We want $C = AB$.
Entry view (row of $A$ dotted with column of $B$):
- $C_{11} = 1\cdot5 + 2\cdot7 = 5 + 14 = 19$
- $C_{12} = 1\cdot6 + 2\cdot8 = 6 + 16 = 22$
- $C_{21} = 3\cdot5 + 4\cdot7 = 15 + 28 = 43$
- $C_{22} = 3\cdot6 + 4\cdot8 = 18 + 32 = 50$
So $AB = \begin{bmatrix} 19 & 22 \\ 43 & 50 \end{bmatrix}$. Now the other three views give the same thing:
- Column view. Column 1 of $AB$ is $A$ times column 1 of $B$: $A\begin{bmatrix}5\\7\end{bmatrix} = 5\begin{bmatrix}1\\3\end{bmatrix} + 7\begin{bmatrix}2\\4\end{bmatrix} = \begin{bmatrix}19\\43\end{bmatrix}$ ✓.
- Row view. Row 1 of $AB$ is row 1 of $A$ times $B$: $[1, 2]\,B = 1\,[5, 6] + 2\,[7, 8] = [19, 22]$ ✓.
- Outer-product view. (column 1 of $A$)(row 1 of $B$) + (column 2 of $A$)(row 2 of $B$): $$\begin{bmatrix}1\\3\end{bmatrix}[5,\,6] + \begin{bmatrix}2\\4\end{bmatrix}[7,\,8] = \begin{bmatrix}5 & 6\\15 & 18\end{bmatrix} + \begin{bmatrix}14 & 16\\28 & 32\end{bmatrix} = \begin{bmatrix}19 & 22\\43 & 50\end{bmatrix} ✓$$
For $A \in \mathbb{R}^{m \times n}$ and $B \in \mathbb{R}^{n \times p}$, the product $C = AB$ is an $m \times p$ matrix. Four equivalent descriptions:
- Entry: $C_{ij} = \sum_{k=1}^{n} A_{ik}B_{kj} = (\text{row } i \text{ of } A)\cdot(\text{column } j \text{ of } B)$.
- Column: $C = [\,A\mathbf{b}_1 \;\; A\mathbf{b}_2 \;\; \cdots \;\; A\mathbf{b}_p\,]$, where $\mathbf{b}_j$ are the columns of $B$.
- Row: row $i$ of $C$ equals $\mathbf{r}_i^\top B$, where $\mathbf{r}_i^\top$ is row $i$ of $A$.
- Outer product: $AB = \sum_{k=1}^{n} \mathbf{a}_k\,\mathbf{b}_k^\top$, where $\mathbf{a}_k$ is column $k$ of $A$ and $\mathbf{b}_k^\top$ is row $k$ of $B$. Each term is an $m \times p$ matrix.
Why do we need it?
Doing one transformation and then another is something we do constantly. Matrix multiplication packs a whole chain into one matrix and processes many vectors in one go.
Where is it used?
A mini-batch passing through a layer (X @ W.T), attention scores Q times K-transpose, the covariance matrix X-transpose times X, low-rank updates such as LoRA, and chained rotations in graphics.
How is it used?
Check the inner shapes first. Then choose the view: entry view to compute one number, column view to see what happens to each input, row view to see how each output row is built, outer-product view to build the answer from simple pieces. In code, write A @ B.
The entry view is the one you compute with, but not the one you think with. Practise the column view until it feels natural. It says: "$AB$ means $A$ acts on each column of $B$."
Do not multiply cell by cell. $(AB)_{ij}$ is not $A_{ij}B_{ij}$. That would be the Hadamard product.
Quick check: use the column view to get column 2 of $AB$ for the example $A = \begin{bmatrix}1&2\\3&4\end{bmatrix}$, $B = \begin{bmatrix}5&6\\7&8\end{bmatrix}$.
Column 2 of $B$ is $[6, 8]$. So column 2 of $AB$ is $6\begin{bmatrix}1\\3\end{bmatrix} + 8\begin{bmatrix}2\\4\end{bmatrix} = \begin{bmatrix}6+16\\18+32\end{bmatrix} = \begin{bmatrix}22\\50\end{bmatrix}$. It matches the second column of $\begin{bmatrix}19&22\\43&50\end{bmatrix}$.
Which matrices can be multiplied? The shape rule core
In the entry view, each output number is a dot product of a row of $A$ with a column of $B$. A dot product only works if both vectors have the same length. A row of $A$ has as many numbers as $A$ has columns. A column of $B$ has as many numbers as $B$ has rows.
So the inner numbers must match, like two puzzle pieces. The outer numbers give the shape of the answer.
- $(2 \times \mathbf{3})(\mathbf{3} \times 4)$: inner numbers $3 = 3$, so it works. The result is $2 \times 4$.
- $(2 \times \mathbf{3})(\mathbf{2} \times 3)$: inner numbers $3 \ne 2$, so it is not defined.
- $(1 \times \mathbf{3})(\mathbf{3} \times 1)$: works, result $1 \times 1$. A single number: this is the dot product $\mathbf{x}^\top\mathbf{y}$.
- $(3 \times \mathbf{1})(\mathbf{1} \times 3)$: works, result $3 \times 3$. This is the outer product $\mathbf{x}\mathbf{y}^\top$.
$AB$ is defined exactly when (number of columns of $A$) = (number of rows of $B$). The answer has (rows of $A$) rows and (columns of $B$) columns.
Why do we need it?
Most bugs in matrix code are shape bugs. A quick rule tells you, before you run anything, whether a product exists and what shape it will have.
Where is it used?
Every deep-learning model definition, written as (batch, in) times (in, out) gives (batch, out), the error messages of PyTorch, NumPy and TensorFlow, and the choice of layer sizes when you design a network.
How is it used?
Write the two shapes side by side, (m by n)(n by p). If the two inner numbers match, the answer has the outer shape m by p. If they do not, transpose or reshape something. Put the shapes in comments next to your code.
It is possible that $AB$ exists but $BA$ does not. Take $A$ of size $2\times3$ and $B$ of size $3\times4$. Then $AB$ is $2\times4$, but $BA$ would need $4 = 2$, which is false.
Quick check: $A$ is $5 \times 3$ and $B$ is $3 \times 2$. Which products exist, and what are their shapes?
$AB$ exists: inner numbers $3 = 3$, so it is $5 \times 2$. $BA$ would pair $2$ with $5$, so it does not exist.
Rules of matrix multiplication: order matters core
Putting on socks and then shoes is not the same as shoes and then socks. Each matrix is an action, and doing actions in a different order usually gives a different result.
But some friendly rules still hold. You may regroup a long product (do the left pair first or the right pair first) as long as you keep the order. And multiplication spreads over addition, just like ordinary numbers.
With $A = \begin{bmatrix} 1 & 2 \\ 3 & 4 \end{bmatrix}$ and $B = \begin{bmatrix} 5 & 6 \\ 7 & 8 \end{bmatrix}$ we found $AB = \begin{bmatrix} 19 & 22 \\ 43 & 50 \end{bmatrix}$. Now swap the order:
$$BA = \begin{bmatrix} 5\cdot1 + 6\cdot3 & 5\cdot2 + 6\cdot4 \\ 7\cdot1 + 8\cdot3 & 7\cdot2 + 8\cdot4 \end{bmatrix} = \begin{bmatrix} 23 & 34 \\ 31 & 46 \end{bmatrix} \neq AB.$$Even though both products exist and have the same shape, the answers differ.
For matrices of compatible shapes and a scalar $c$:
- Associative: $(AB)C = A(BC)$. (So we can write $ABC$.)
- Distributive: $A(B + C) = AB + AC$ and $(A + B)C = AC + BC$.
- Scalars move freely: $(cA)B = c(AB) = A(cB)$.
- Identity: $IA = A = AI$ (the identity matrix is met below).
- Not commutative: in general $AB \neq BA$. (Sometimes they are equal, but you may never assume it.)
Why do we need it?
To simplify long expressions and to avoid illegal moves, we need to know which familiar rules from ordinary numbers still hold for matrices, and which ones break.
Where is it used?
Deriving gradients by hand, simplifying chains such as ABC in backpropagation, reasoning about the order of layers in a network, and spotting bugs in matrix code.
How is it used?
Regroup freely using associativity and expand using distributivity. Never swap the order of two matrices unless you know they commute. When you expand (A + B) squared, keep both the AB and the BA terms.
- No cancelling. $AB = AC$ does not let you conclude $B = C$.
- Zero products. $AB = 0$ does not mean $A = 0$ or $B = 0$. For example $\begin{bmatrix}1&0\\0&0\end{bmatrix}\begin{bmatrix}0&0\\0&1\end{bmatrix} = \begin{bmatrix}0&0\\0&0\end{bmatrix}$.
- Expanding squares. $(A + B)^2 = A^2 + AB + BA + B^2$, not $A^2 + 2AB + B^2$, because $AB \ne BA$ in general.
Quick check: does $(A + B)(A - B) = A^2 - B^2$ hold for matrices?
Expand: $(A + B)(A - B) = A^2 - AB + BA - B^2$. The middle terms cancel only if $AB = BA$. In general they do not, so the identity fails for matrices.
How much work is a matrix product? $O(mnp)$
Count the work in the entry view. The answer $C = AB$ has $m \cdot p$ entries. Each entry is a dot product of two lists of length $n$, which needs $n$ multiplications (and about $n$ additions).
So the total is about $m \cdot p \cdot n$ multiplications. Double any one of the three sizes and the work doubles. That is why big matrices are expensive, and why GPUs (which do thousands of multiplications at once) matter so much.
- $A$ is $2 \times 3$ and $B$ is $3 \times 4$. $C$ has $2 \cdot 4 = 8$ entries, each needs $3$ multiplications: $8 \cdot 3 = 24$ in total.
- Two $1000 \times 1000$ matrices: $1000 \cdot 1000 \cdot 1000 = 10^9$ multiplications. One billion.
Order matters for the cost. Take $A$ of size $2 \times 1000$, $B$ of size $1000 \times 2$, $C$ of size $2 \times 1000$. The answer $ABC$ is the same either way, but
- $(AB)C$: $AB$ is $2 \times 2$ (cost $2\cdot1000\cdot2 = 4000$), then $(2\times2)(2\times1000)$ costs $2\cdot2\cdot1000 = 4000$. Total 8000.
- $A(BC)$: $BC$ is $1000 \times 1000$ (cost $1000\cdot2\cdot1000 = 2{,}000{,}000$), then $A(BC)$ costs another $2{,}000{,}000$. Total 4 million.
Multiplying $A \in \mathbb{R}^{m \times n}$ by $B \in \mathbb{R}^{n \times p}$ with the standard method takes
$$\text{cost} = m \cdot n \cdot p \text{ multiplications} \;\; (\text{and about the same number of additions}), \quad \text{written } O(mnp).$$For two $n \times n$ matrices this is $O(n^3)$. Cleverer algorithms exist (they are a little faster for huge matrices), but $O(n^3)$ is what you should expect in practice.
Why do we need it?
Training a model means billions of multiplications. Knowing the cost lets us predict how slow something will be and choose a cheaper way to bracket a product.
Where is it used?
Estimating the cost of dense layers, explaining why GPUs and TPUs exist, low-rank tricks such as LoRA, attention (cost grows with the square of the sequence length), and ordering long matrix chains.
How is it used?
Work out m times n times p for each product. For a chain, compare the cost of each possible bracketing and pick the cheapest. With a single vector, multiply right to left: A @ (B @ x), not (A @ B) @ x.
Quick check: how many multiplications does $(3 \times 5)(5 \times 2)$ need?
$3 \cdot 5 \cdot 2 = 30$. The result is $3 \times 2$ (6 entries), each a dot product of length 5.
The transpose $A^\top$ core
The transpose flips a table over its main diagonal (the line from the top-left corner going down to the right). Rows become columns and columns become rows. It is like turning a spreadsheet on its side.
You have already met it in a small way: the transpose of a column vector is a row vector. Now we do it for a whole matrix.
- $A$ is $2 \times 3$, so $A^\top$ is $3 \times 2$.
- Row 1 of $A$, $[1, 2, 3]$, is now column 1 of $A^\top$.
- The entry $A_{12} = 2$ moved to position $(2, 1)$ in $A^\top$.
The transpose of $A \in \mathbb{R}^{m \times n}$ is the matrix $A^\top \in \mathbb{R}^{n \times m}$ with
$$(A^\top)_{ij} = A_{ji}.$$Basic rules: $(A^\top)^\top = A$ (flip twice and you are back), $(A + B)^\top = A^\top + B^\top$, and $(cA)^\top = cA^\top$. The most important rule, the one for products, comes in the section after next.
Why do we need it?
Sometimes rows and columns are the wrong way round for the multiplication we want. The transpose flips them without changing any of the numbers.
Where is it used?
The batch layer X times W-transpose, backpropagation (errors travel backwards through W-transpose), Gram and covariance matrices X-transpose times X, and any formula that needs a row instead of a column vector.
How is it used?
In NumPy write A.T. An m by n matrix becomes n by m, and rows become columns. Use it to make shapes fit, and remember that transposing twice gives back the original matrix.
The shape changes. If $A$ is $2 \times 3$, then $A^\top$ is $3 \times 2$. Only square matrices keep their shape.
The transpose is not the inverse. $A^\top$ just rearranges the same numbers. (The inverse of $A$ is the matrix that undoes $A$; you will meet it in Chapter 1.7. For some very special matrices, called orthogonal, the transpose happens to be the inverse. See below.)
Quick check: what is the transpose of the $1 \times 3$ matrix $[7, 8, 9]$?
It is the $3 \times 1$ column $\begin{bmatrix}7\\8\\9\end{bmatrix}$. (A row vector and a column vector are transposes of each other.)
The transpose of a product: $(AB)^\top = B^\top A^\top$ core
In the morning you put on socks, then shoes. In the evening you reverse it: take off shoes, then socks. Transposing a product reverses the order.
There is a shape reason too. If $A$ is $2 \times 3$ and $B$ is $3 \times 2$, then $A^\top B^\top$ is $(3 \times 2)(2 \times 3)$, which is $3 \times 3$. But $(AB)^\top$ is $2 \times 2$. Only $B^\top A^\top$, which is $(2\times3)(3\times2)$, has the right shape.
Use $A = \begin{bmatrix} 1 & 2 \\ 3 & 4 \end{bmatrix}$ and $B = \begin{bmatrix} 5 & 6 \\ 7 & 8 \end{bmatrix}$. Earlier, $AB = \begin{bmatrix} 19 & 22 \\ 43 & 50 \end{bmatrix}$, so
$$(AB)^\top = \begin{bmatrix} 19 & 43 \\ 22 & 50 \end{bmatrix}.$$Now the other side: $B^\top = \begin{bmatrix} 5 & 7 \\ 6 & 8 \end{bmatrix}$ and $A^\top = \begin{bmatrix} 1 & 3 \\ 2 & 4 \end{bmatrix}$.
$$B^\top A^\top = \begin{bmatrix} 5\cdot1 + 7\cdot2 & 5\cdot3 + 7\cdot4 \\ 6\cdot1 + 8\cdot2 & 6\cdot3 + 8\cdot4 \end{bmatrix} = \begin{bmatrix} 19 & 43 \\ 22 & 50 \end{bmatrix} ✓$$And $A^\top B^\top = \begin{bmatrix} 1\cdot5+3\cdot6 & 1\cdot7+3\cdot8 \\ 2\cdot5+4\cdot6 & 2\cdot7+4\cdot8 \end{bmatrix} = \begin{bmatrix} 23 & 31 \\ 34 & 46 \end{bmatrix} \ne (AB)^\top$. The wrong order fails.
Why it is true: entry $(i, j)$ of $(AB)^\top$ is entry $(j, i)$ of $AB$, which is (row $j$ of $A$)$\cdot$(column $i$ of $B$). Entry $(i, j)$ of $B^\top A^\top$ is (row $i$ of $B^\top$)$\cdot$(column $j$ of $A^\top$), and that is the very same two lists of numbers: column $i$ of $B$ and row $j$ of $A$. The dot product does not care about order, so they agree.
Why do we need it?
When we transpose a long expression we need to know what happens to the order. Getting it wrong is the most common slip in hand-derived gradients.
Where is it used?
Deriving the gradient of the squared error ||Xw - y||^2, the normal equations of linear regression, backpropagation formulas, and the proof that A-transpose times A is symmetric.
How is it used?
Reverse the order and transpose each piece: the transpose of ABC is C-transpose, B-transpose, A-transpose. As a quick check, compare shapes: only the reversed order fits. In NumPy test it with np.allclose((A @ B).T, B.T @ A.T).
It is a very common slip to write $(AB)^\top = A^\top B^\top$. The order must reverse. The same happens for the inverse of a product, $(AB)^{-1} = B^{-1}A^{-1}$ (Chapter 1.7).
Quick check: simplify $(A B \mathbf{x})^\top$.
Reverse the order and transpose each piece: $\mathbf{x}^\top B^\top A^\top$.
Revisited: the inner product $\mathbf{x}^\top\mathbf{y}$ and the outer product $\mathbf{x}\mathbf{y}^\top$
In Chapter 1.2 the dot product was written $\mathbf{x}\cdot\mathbf{y}$. Now we can see it as a matrix product and explain the notation $\mathbf{x}^\top\mathbf{y}$.
A column vector is an $n \times 1$ matrix, so $\mathbf{x}^\top$ is $1 \times n$. Put the row on the left and the column on the right, and the inner numbers match: $(1\times n)(n\times1) = 1\times1$, a single number. Swap the order and you get $(n\times1)(1\times n) = n\times n$, a whole table.
The two products have the same ingredients, and they look completely different. The first one squeezes two vectors into one number ("how much do they agree?"). The second one spreads them out into a grid of all possible products.
Let $\mathbf{x} = [1, 2, 3]$ and $\mathbf{y} = [4, 5, 6]$.
Inner: $\mathbf{x}^\top\mathbf{y} = 1\cdot4 + 2\cdot5 + 3\cdot6 = 4 + 10 + 18 = 32$.
Outer: every entry is (a number of $\mathbf{x}$) times (a number of $\mathbf{y}$):
$$\mathbf{x}\mathbf{y}^\top = \begin{bmatrix} 1\cdot4 & 1\cdot5 & 1\cdot6 \\ 2\cdot4 & 2\cdot5 & 2\cdot6 \\ 3\cdot4 & 3\cdot5 & 3\cdot6 \end{bmatrix} = \begin{bmatrix} 4 & 5 & 6 \\ 8 & 10 & 12 \\ 12 & 15 & 18 \end{bmatrix}$$Notice that every row is a multiple of $[4, 5, 6]$ (times 1, 2, 3). The outer product always has this "one pattern, repeated" look.
For $\mathbf{x}, \mathbf{y} \in \mathbb{R}^n$:
$$\mathbf{x}^\top\mathbf{y} = \sum_{i} x_iy_i \in \mathbb{R} \quad (1\times1), \qquad (\mathbf{x}\mathbf{y}^\top)_{ij} = x_i\,y_j \in \mathbb{R}^{n\times n}.$$For different lengths ($\mathbf{x} \in \mathbb{R}^m$, $\mathbf{y} \in \mathbb{R}^p$), $\mathbf{x}\mathbf{y}^\top$ is an $m \times p$ matrix. Every row is a multiple of $\mathbf{y}^\top$ and every column is a multiple of $\mathbf{x}$. A matrix like this is called rank 1 (Chapter 1.7 explains rank).
Why do we need it?
Two vectors can be combined in two very different ways. Seeing both as matrix products explains why the dot product is one number and the outer product is a whole table.
Where is it used?
A neuron's pre-activation (weights transposed times inputs), the weight gradient of a layer (error times input-transpose), covariance matrices (averages of x times x-transpose), and rank-1 updates such as in LoRA.
How is it used?
To score how much two vectors agree, use x-transpose times y (x @ y in NumPy). To build a table of all products x_i times y_j, use x times y-transpose (np.outer(x, y)). Check the shapes: (1 by n)(n by 1) is 1 by 1, and (m by 1)(1 by p) is m by p.
Quick check: $\mathbf{x}$ has 4 entries and $\mathbf{y}$ has 2. What is the shape of $\mathbf{x}\mathbf{y}^\top$? Is $\mathbf{x}^\top\mathbf{y}$ defined?
$\mathbf{x}\mathbf{y}^\top$ is $(4\times1)(1\times2) = 4\times2$. But $\mathbf{x}^\top\mathbf{y}$ would be $(1\times4)(2\times1)$, and $4 \ne 2$, so it is not defined. The dot product needs equal lengths.
The Gram matrix $A^\top A$ core
Take a matrix and look at its columns as separate vectors. Now build a table of how much every column agrees with every other column, using dot products. That table is the Gram matrix.
It is a table of "similarities between columns". Entry $(i, j)$ answers: "how much do column $i$ and column $j$ point the same way?" The table is automatically symmetric, because column $i$ against column $j$ is the same dot product as column $j$ against column $i$.
Let $A = \begin{bmatrix} 1 & 2 \\ 2 & 0 \\ 2 & 1 \end{bmatrix}$. Its columns are $\mathbf{a}_1 = [1, 2, 2]$ and $\mathbf{a}_2 = [2, 0, 1]$.
- $\mathbf{a}_1\cdot\mathbf{a}_1 = 1 + 4 + 4 = 9$ (the squared length of column 1).
- $\mathbf{a}_1\cdot\mathbf{a}_2 = 2 + 0 + 2 = 4$.
- $\mathbf{a}_2\cdot\mathbf{a}_1 = 4$ (same number).
- $\mathbf{a}_2\cdot\mathbf{a}_2 = 4 + 0 + 1 = 5$.
The matrix is $2 \times 2$ (one row and one column for each column of $A$), and it is symmetric.
The Gram matrix of $A \in \mathbb{R}^{m\times n}$ is $G = A^\top A \in \mathbb{R}^{n \times n}$, with entries
$$G_{ij} = \mathbf{a}_i^\top \mathbf{a}_j = (\text{column } i)\cdot(\text{column } j).$$- Always symmetric: $(A^\top A)^\top = A^\top (A^\top)^\top = A^\top A$.
- The diagonal entries are the squared lengths of the columns: $G_{ii} = \|\mathbf{a}_i\|^2$.
- $G_{ij} = 0$ exactly when column $i$ and column $j$ are orthogonal.
- $AA^\top$ is the Gram matrix of the rows. It is $m \times m$.
Why do we need it?
We often want to know how all the columns (the features) relate to each other at once. The Gram matrix collects every pairwise dot product into one symmetric table.
Where is it used?
The normal equations of linear regression, covariance matrices and PCA, kernel methods (the table X times X-transpose of sample similarities), and checking whether features are redundant.
How is it used?
Compute G = A.T @ A. Read the diagonal as the squared lengths of the columns, and each off-diagonal entry as how much two columns agree. A zero entry means those two columns are orthogonal.
$A^\top A$ and $AA^\top$ are different matrices (different sizes unless $A$ is square). Which one you want depends on the question: columns (features) or rows (samples).
Quick check: $A$ is $100 \times 5$. What shapes are $A^\top A$ and $AA^\top$?
$A^\top A$ is $(5\times100)(100\times5) = 5\times5$. $AA^\top$ is $(100\times5)(5\times100) = 100\times100$. Both are symmetric, but they have very different sizes.
Identity, zero and diagonal matrices core
Some matrices are so simple that you can read what they do at a glance.
- Identity $I$: the "do nothing" machine. It is to matrices what the number 1 is to numbers.
- Zero matrix $0$: the "squash everything to zero" machine, like multiplying by 0.
- Diagonal matrix $D$: numbers only on the main diagonal. It stretches each coordinate on its own: the first entry by $d_1$, the second by $d_2$, and so on. Think of a mixing desk where each slider controls one channel.
Let $D = \begin{bmatrix} 2 & 0 & 0 \\ 0 & 3 & 0 \\ 0 & 0 & -1 \end{bmatrix}$ and $\mathbf{x} = [4, 5, 6]$. Then $D\mathbf{x} = [2\cdot4,\; 3\cdot5,\; (-1)\cdot6] = [8, 15, -6]$.
And $I\mathbf{x} = \mathbf{x}$ for the identity $I = \begin{bmatrix}1&0&0\\0&1&0\\0&0&1\end{bmatrix}$.
Diagonal matrices also act on other matrices. With $D = \begin{bmatrix}2&0\\0&3\end{bmatrix}$ and $A = \begin{bmatrix}1&2\\3&4\end{bmatrix}$:
$$DA = \begin{bmatrix}2&4\\9&12\end{bmatrix} \;(\text{rows scaled by } 2, 3), \qquad AD = \begin{bmatrix}2&6\\6&12\end{bmatrix} \;(\text{columns scaled by } 2, 3).$$- The identity matrix $I_n$ is $n\times n$ with $1$ on the diagonal and $0$ elsewhere. For every matrix of a compatible shape: $IA = A$ and $AI = A$. Its columns are the standard unit vectors $\mathbf{e}_1, \dots, \mathbf{e}_n$.
- The zero matrix $0$ has every entry $0$. $A + 0 = A$ and $0\cdot A = 0$.
- A diagonal matrix has $D_{ij} = 0$ whenever $i \ne j$. We write $D = \operatorname{diag}(d_1, \dots, d_n)$. Then $D\mathbf{x} = [d_1x_1, \dots, d_nx_n]$, $DA$ scales the rows of $A$, and $AD$ scales the columns of $A$.
Diagonal matrices are very easy to multiply: $\operatorname{diag}(a_i)\operatorname{diag}(b_i) = \operatorname{diag}(a_ib_i)$, and they commute with each other. Powers are easy too: $D^k = \operatorname{diag}(d_1^k, \dots, d_n^k)$.
Why do we need it?
Some operations are as simple as 'do nothing', 'make everything zero' or 'scale each coordinate separately'. These matrices give such operations a name and a fast recipe.
Where is it used?
Feature standardisation (X times a diagonal matrix), per-parameter step sizes in the Adam optimiser, the lambda times I term in ridge regression, skip connections, and the diagonal matrices inside the SVD (Chapter 1.13).
How is it used?
Build them with np.eye(n), np.zeros((m, n)) and np.diag([d1, d2, d3]). Remember that D @ A scales the rows of A and A @ D scales its columns. In practice multiply by the diagonal entries directly (d * x) instead of building a big matrix.
The identity is always square. If $A$ is $2\times3$, then $I_2A = A$ and $AI_3 = A$, and you need two different identities (one on each side).
Quick check: what does $\operatorname{diag}(1, 0, 1)$ do to $[7, 8, 9]$?
It gives $[1\cdot7, 0\cdot8, 1\cdot9] = [7, 0, 9]$. The middle coordinate is switched off.
Symmetric and skew-symmetric matrices core
- Symmetric = a mirror image across the diagonal. A table of distances between cities is symmetric: the distance from A to B is the distance from B to A.
- Skew-symmetric = a mirror image with the sign flipped. A table of "who owes whom" is skew: if A owes B 5 dollars, then B owes A $-5$. The diagonal has to be 0, because nobody owes themselves.
Here is a neat fact: every square matrix is a symmetric part plus a skew part.
$S = \begin{bmatrix} 2 & 3 \\ 3 & 5 \end{bmatrix}$ is symmetric ($S_{12} = S_{21} = 3$). $K = \begin{bmatrix} 0 & 4 \\ -4 & 0 \end{bmatrix}$ is skew ($K_{12} = 4$, $K_{21} = -4$, zero diagonal).
Split $A = \begin{bmatrix} 1 & 2 \\ 4 & 3 \end{bmatrix}$. First $A^\top = \begin{bmatrix} 1 & 4 \\ 2 & 3 \end{bmatrix}$. Then
$$\tfrac12(A + A^\top) = \begin{bmatrix} 1 & 3 \\ 3 & 3 \end{bmatrix}, \qquad \tfrac12(A - A^\top) = \begin{bmatrix} 0 & -1 \\ 1 & 0 \end{bmatrix},$$and indeed $\begin{bmatrix} 1 & 3 \\ 3 & 3 \end{bmatrix} + \begin{bmatrix} 0 & -1 \\ 1 & 0 \end{bmatrix} = \begin{bmatrix} 1 & 2 \\ 4 & 3 \end{bmatrix} = A$ ✓.
- $A$ is symmetric if $A^\top = A$, i.e. $A_{ij} = A_{ji}$ for all $i, j$.
- $A$ is skew-symmetric if $A^\top = -A$, i.e. $A_{ij} = -A_{ji}$. Then $A_{ii} = -A_{ii}$, so every diagonal entry is $0$.
- Both must be square.
- Every square $A$ equals $S + K$ with $S = \tfrac12(A + A^\top)$ symmetric and $K = \tfrac12(A - A^\top)$ skew.
Why do we need it?
Mirror symmetry makes a matrix much easier to work with and gives it very nice properties. Splitting any matrix into a symmetric part and a skew part separates balanced behaviour from rotation-like behaviour.
Where is it used?
Covariance matrices, Gram matrices A-transpose times A, the Hessian (second derivatives) in optimisation, distance and similarity tables, and kernel matrices. Skew matrices appear when studying rotations.
How is it used?
Test with np.allclose(A, A.T). Split any square matrix with S = (A + A.T) / 2 and K = (A - A.T) / 2. When you know a matrix is symmetric, you can use faster and more stable routines such as np.linalg.eigh.
A symmetric matrix must be square. And "symmetric" is about the numbers mirroring across the main diagonal, not about a left-right flip of the whole picture.
Quick check: can $\begin{bmatrix} 1 & 2 \\ -2 & 3 \end{bmatrix}$ be skew-symmetric?
No. A skew matrix needs a zero diagonal, but the diagonal here is $1$ and $3$. (The off-diagonal pair $2$ and $-2$ is fine.)
Upper and lower triangular matrices
Imagine drawing the main diagonal and then wiping out everything below it. What is left looks like a staircase, with numbers on the diagonal and above. That is an upper triangular matrix. Wipe out everything above instead, and you have a lower triangular one.
Why care? Equations with a triangular matrix are easy to solve. With an upper triangular matrix, the last equation has only one unknown, so you solve it first and work upwards. (With a lower triangular matrix you start at the first equation and work downwards.)
Solve $U\mathbf{x} = \mathbf{b}$ with $U = \begin{bmatrix} 2 & 1 \\ 0 & 3 \end{bmatrix}$ and $\mathbf{b} = [8, 6]$.
- Row 2 says $3x_2 = 6$, so $x_2 = 2$.
- Row 1 says $2x_1 + 1\cdot x_2 = 8$. Put in $x_2 = 2$: $2x_1 + 2 = 8$, so $x_1 = 3$.
So $\mathbf{x} = [3, 2]$. This "solve from the bottom up" method is called back-substitution.
Also, multiplying two upper triangular matrices gives another upper triangular matrix: $\begin{bmatrix}1&2\\0&3\end{bmatrix}\begin{bmatrix}4&5\\0&6\end{bmatrix} = \begin{bmatrix}4&17\\0&18\end{bmatrix}$.
- Upper triangular: $U_{ij} = 0$ whenever $i \gt j$ (everything below the diagonal is zero).
- Lower triangular: $L_{ij} = 0$ whenever $i \lt j$ (everything above the diagonal is zero).
The product of two upper (or two lower) triangular matrices is again upper (or lower) triangular. Diagonal matrices are both at once. Elimination (Chapter 1.6) turns a general matrix into triangular pieces.
Why do we need it?
Equations with a triangular matrix can be solved one unknown at a time, which is very cheap. That is why computers turn hard problems into triangular ones.
Where is it used?
The LU and Cholesky factorisations (ways to split a matrix into triangular pieces, see Chapter 1.13) inside np.linalg.solve, least squares through QR, computing determinants, and the fast back-substitution in numerical libraries.
How is it used?
Use np.triu and np.tril to build or extract triangular parts. To solve Ux = b, start at the last row, solve for the last unknown, then work upwards (back-substitution). This takes about n squared steps instead of n cubed.
Quick check: is $\begin{bmatrix} 3 & 0 & 0 \\ 1 & 2 & 0 \\ 4 & 5 & 6 \end{bmatrix}$ upper or lower triangular?
Lower triangular: everything above the diagonal is zero.
Orthogonal matrices core
An orthogonal matrix is a rigid motion: it can turn or flip space, but it never stretches, squashes or shears it. Lengths stay the same. Angles stay the same. Think of rotating a photo or flipping it in a mirror.
You can spot one by its columns: they are perpendicular unit vectors (orthonormal, as in the orthogonal vectors section). And its transpose undoes it.
$Q = \begin{bmatrix} 0 & -1 \\ 1 & 0 \end{bmatrix}$ turns every arrow by $90^\circ$ anticlockwise.
- Columns: $[0, 1]$ and $[-1, 0]$. Each has length 1, and their dot product is $0\cdot(-1) + 1\cdot0 = 0$.
- $Q^\top Q = \begin{bmatrix} 0 & 1 \\ -1 & 0 \end{bmatrix}\begin{bmatrix} 0 & -1 \\ 1 & 0 \end{bmatrix} = \begin{bmatrix} 1 & 0 \\ 0 & 1 \end{bmatrix} = I$ ✓.
- Length check: $Q[3, 4] = [-4, 3]$, and both $[3,4]$ and $[-4,3]$ have length $5$ ✓.
A square matrix $Q$ is orthogonal if
$$Q^\top Q = I.$$This is the same as saying the columns of $Q$ are orthonormal. It also means $Q^\top$ is the matrix that undoes $Q$ (its inverse, Chapter 1.7): $QQ^\top = Q^\top Q = I$. Key properties:
- It keeps lengths: $\|Q\mathbf{x}\| = \|\mathbf{x}\|$.
- It keeps dot products and angles: $(Q\mathbf{x})\cdot(Q\mathbf{y}) = \mathbf{x}\cdot\mathbf{y}$.
- Rotations and reflections are orthogonal. So are permutation matrices.
Why do we need it?
We want a transformation that turns or flips space without distorting it. Orthogonal matrices do exactly that, and they are easy to undo: the inverse is just the transpose.
Where is it used?
Rotations and reflections in graphics and robotics, PCA and the SVD (U and V are orthogonal; see Chapter 1.13), orthogonal weight initialisation in deep and recurrent networks, QR factorisation, and rotary position embeddings.
How is it used?
Check Q-transpose times Q equals I with np.allclose(Q.T @ Q, np.eye(n)). To undo Q use Q.T instead of an inverse. Use an orthogonal Q when you want to rotate or change coordinates without changing any length or angle.
The word is confusing: an orthogonal matrix has orthonormal columns (perpendicular and length 1). A matrix with perpendicular columns of other lengths is not orthogonal.
Quick check: is $\begin{bmatrix} 1 & 1 \\ -1 & 1 \end{bmatrix}$ orthogonal?
The columns $[1, -1]$ and $[1, 1]$ are perpendicular ($1 - 1 = 0$), but each has length $\sqrt2$, not 1. So $Q^\top Q = 2I$, not $I$. It is not orthogonal. Dividing by $\sqrt2$ would fix it.
Permutation matrices
A permutation matrix is a shuffler. Take the identity matrix and shuffle its rows. The result reorders anything you multiply it with: entries of a vector, rows of a matrix, or columns of a matrix.
$P = \begin{bmatrix} 0 & 1 & 0 \\ 0 & 0 & 1 \\ 1 & 0 & 0 \end{bmatrix}$ and $\mathbf{x} = [10, 20, 30]$.
- Row 1 of $P$ has its 1 in column 2, so it picks $x_2 = 20$.
- Row 2 has its 1 in column 3, so it picks $x_3 = 30$.
- Row 3 has its 1 in column 1, so it picks $x_1 = 10$.
So $P\mathbf{x} = [20, 30, 10]$: the entries were shifted around.
A permutation matrix has exactly one $1$ in every row and every column, and $0$ everywhere else. There are $n!$ of them of size $n \times n$ (6 for $3\times3$).
- $P\mathbf{x}$ reorders the entries of $\mathbf{x}$. $PA$ reorders the rows of $A$. $AP$ reorders the columns of $A$.
- Every permutation matrix is orthogonal: $P^\top P = I$. So $P^\top$ undoes the shuffle.
Why do we need it?
Reordering things is a very common need: shuffling data, swapping rows during elimination, relabelling nodes. A permutation matrix writes any reordering as a multiplication.
Where is it used?
Shuffling data sets and mini-batches, row pivoting in LU factorisation (PA = LU), relabelling graph nodes (and permutation equivariance in graph neural networks), and sorting-based algorithms.
How is it used?
Build one by shuffling the rows of the identity: np.eye(n)[perm]. Then P @ x reorders entries, P @ A reorders rows, A @ P reorders columns, and P.T undoes it. In practice you simply index, A[perm], instead of multiplying.
Quick check: which of $\begin{bmatrix}0&1\\1&0\end{bmatrix}$ and $\begin{bmatrix}1&1\\0&0\end{bmatrix}$ is a permutation matrix?
The first one: it has one 1 in each row and each column (it swaps two things). The second has two 1s in row 1 and none in row 2, so it is not.
Sparse matrices (a preview)
Picture a table with one row and one column for every person on a social network. Put a 1 if two people are friends and a 0 if not. Almost every entry is 0, because each person knows a few hundred people out of millions.
A matrix that is mostly zeros is called sparse. It is wasteful to store all those zeros, and wasteful to multiply by them. A smart program stores only the non-zero entries and their addresses, and skips the rest.
A matrix of size $1{,}000{,}000 \times 1{,}000{,}000$ with only 5 non-zeros per row:
- Dense storage: $10^{12}$ numbers, around 8 terabytes. It does not fit in any normal computer.
- Sparse storage: $5 \times 10^6$ non-zeros, each stored with its address (about 3 numbers each): around 120 megabytes (at 8 bytes per number). It fits easily.
A small one: $\begin{bmatrix} 0 & 0 & 3 \\ 0 & 0 & 0 \\ 7 & 0 & 0 \end{bmatrix}$ can be stored as the list $\{(1, 3, \mathbf{3}),\ (3, 1, \mathbf{7})\}$: (row, column, value).
A matrix is sparse when most entries are zero. The number of non-zero entries is $\text{nnz}$, and the sparsity is the fraction of zeros: $1 - \text{nnz}/(mn)$.
- Common storage formats: COO (a list of row, column, value) and CSR (rows packed one after another).
- Multiplying a sparse matrix by a vector costs about $\text{nnz}$ operations instead of $mn$.
- Typical sparse shapes: diagonal, banded (a few diagonals), block-diagonal, or just random scatter. (A full triangular matrix is only about half zeros, so it is not really sparse.)
Why do we need it?
Many real tables are almost all zeros. Storing and multiplying all those zeros wastes memory and time, and can make a problem impossible to solve at all.
Where is it used?
Bag-of-words and one-hot text features, graph adjacency matrices, recommender-system user-by-item tables, pruned or sparse-attention neural networks, and scientific simulations.
How is it used?
Count how many entries are non-zero. If most are zero, store only (row, column, value) with scipy.sparse (CSR or COO) and multiply as usual. The cost then grows with the number of non-zeros, not with the full size.
The inverse or product of sparse matrices is often not sparse. A big sparse matrix is cheap to store and multiply, but you rarely want to invert it.
Quick check: a $1000 \times 1000$ matrix has 3000 non-zeros. How sparse is it?
It has $10^6$ entries, and $3000/10^6 = 0.3\%$ of them are non-zero. So $99.7\%$ are zeros: very sparse.
Block matrices and block multiplication
A big matrix can be cut into rectangular tiles by drawing a few lines through it, like a chocolate bar. Each tile is itself a small matrix. The nice surprise: you can multiply big matrices tile by tile, treating the tiles as if they were single numbers. You just have to keep the order of each product.
Blocks are how we think about structure: "the first half of the features and the second half", or "two independent sub-problems placed side by side".
The $4\times4$ matrix below is cut into four $2\times2$ blocks, and two of them are zero blocks:
$$M = \begin{bmatrix} 1 & 2 & 0 & 0 \\ 3 & 4 & 0 & 0 \\ 0 & 0 & 5 & 6 \\ 0 & 0 & 7 & 8 \end{bmatrix} = \begin{bmatrix} A & 0 \\ 0 & B \end{bmatrix}, \quad A = \begin{bmatrix}1&2\\3&4\end{bmatrix},\; B = \begin{bmatrix}5&6\\7&8\end{bmatrix}.$$This is a block-diagonal matrix. Multiply it by $\mathbf{x} = [1, 1, 1, 1]$, cut into two halves $[1,1]$ and $[1,1]$: the top half only sees $A$ and the bottom half only sees $B$.
$$M\mathbf{x} = \begin{bmatrix} A[1,1]^\top \\ B[1,1]^\top \end{bmatrix} = \begin{bmatrix} 3 \\ 7 \\ 11 \\ 15 \end{bmatrix}.$$If $A$ and $B$ are cut into blocks so that the block sizes fit (the column cuts of $A$ match the row cuts of $B$), then
$$\begin{bmatrix} A_{11} & A_{12} \\ A_{21} & A_{22} \end{bmatrix}\begin{bmatrix} B_{11} & B_{12} \\ B_{21} & B_{22} \end{bmatrix} = \begin{bmatrix} A_{11}B_{11} + A_{12}B_{21} & A_{11}B_{12} + A_{12}B_{22} \\ A_{21}B_{11} + A_{22}B_{21} & A_{21}B_{12} + A_{22}B_{22} \end{bmatrix}.$$It is exactly the entry-view rule, with blocks in place of numbers: block $(I, J)$ of the answer is $\sum_K A_{IK}B_{KJ}$. Keep the order, because blocks do not commute either.
Why do we need it?
Big matrices often have structure: two groups of features, or two independent problems side by side. Cutting a matrix into blocks lets us reason about, and compute, each piece separately.
Where is it used?
Block-diagonal weights in grouped convolutions and multi-head attention, tiled matrix multiplication on GPUs, adding a bias by appending a column of ones, and block matrices in Kalman filters and optimisation.
How is it used?
Cut A and B so the block sizes fit, then multiply the blocks as if they were numbers, keeping the order: C_IJ is the sum over K of A_IK times B_KJ. In NumPy, np.block builds a matrix from blocks and slices like A[:2, 2:] pick one out.
The blocks must have compatible sizes, and each block product must be written in the right order: $A_{11}B_{11}$, never $B_{11}A_{11}$.
Quick check: what is $\begin{bmatrix} A & 0 \\ 0 & B \end{bmatrix}\begin{bmatrix} C & 0 \\ 0 & D \end{bmatrix}$ for compatible blocks?
$\begin{bmatrix} AC & 0 \\ 0 & BD \end{bmatrix}$. The zero blocks kill the cross terms, so the two halves never interact.
Special-matrix detective: what kind of matrix is this?
You now know a whole family of special matrices. Each has a simple test. In practice you will often meet a matrix and ask: "Which kind is it? Does it have structure I can use?" Structure means speed, stability and good theory.
Test $\begin{bmatrix} 0 & -1 & 0 \\ 1 & 0 & 0 \\ 0 & 0 & 1 \end{bmatrix}$.
- Symmetric? $A_{12} = -1$ but $A_{21} = 1$. No.
- Columns $[0,1,0]$, $[-1,0,0]$, $[0,0,1]$: each has length 1 and they are pairwise perpendicular. Orthogonal ✓ (a rotation by $90^\circ$ about the vertical axis).
- More than half of the entries are zero: sparse ✓.
| Type | Test |
|---|---|
| Zero | every entry is $0$ |
| Identity | 1s on the diagonal, 0s elsewhere |
| Diagonal | $A_{ij} = 0$ for $i \ne j$ |
| Symmetric | $A^\top = A$ |
| Skew-symmetric | $A^\top = -A$ |
| Upper / lower triangular | zeros below / above the diagonal |
| Orthogonal | $A^\top A = I$ |
| Permutation | exactly one 1 in each row and column, 0 elsewhere |
| Sparse | most entries are $0$ |
A matrix can belong to several types at once. The identity is diagonal, symmetric, triangular, orthogonal, a permutation, and sparse.
Why do we need it?
Spotting structure in a matrix tells you which faster and safer method to use, and which theory applies to it.
Where is it used?
Choosing a solver in numerical libraries (np.linalg.eigh for symmetric matrices, a triangular solve for triangular ones), debugging a model (is this weight matrix nearly orthogonal or nearly diagonal?), and choosing sparse or dense storage.
How is it used?
Run the simple tests from the table, with a tolerance in real code: np.allclose(A, A.T), np.allclose(Q.T @ Q, np.eye(n)), np.count_nonzero(A) / A.size. Then pick the method that matches the structure you found.
The widget allows a tiny rounding error when it tests for zeros and ones. In your own code, decimal (floating-point) numbers are rarely exactly 0 or 1, so always use a tolerance: np.allclose(Q.T @ Q, np.eye(3)) instead of ==.
Quick check: is the zero matrix symmetric? Skew-symmetric? Orthogonal?
Symmetric: yes ($0^\top = 0$). Skew-symmetric: yes ($0^\top = -0$). Orthogonal: no, because $0^\top 0 = 0 \ne I$.
Putting it together: a neural-network layer $W\mathbf{x} + \mathbf{b}$ core
A dense (fully connected) layer takes a vector of $d$ input numbers and produces a vector of $k$ output numbers. Each output is a weighted sum of all inputs, plus a small shift (a bias). That is exactly matrix times vector, plus a vector:
- Row view: output $j$ is the dot product of the weight row $j$ with the input, plus $b_j$.
- Column view: the output is a mix of the weight columns, with the input numbers as scoops.
To process a whole batch of samples at once, stack them as the rows of a data matrix $X$ and do one matrix product. That single product is why GPUs make neural networks fast.
Let $W = \begin{bmatrix} 1 & 0 & -1 \\ 2 & 1 & 0 \end{bmatrix}$ ($2\times3$: 3 inputs, 2 outputs), $\mathbf{b} = [1, -1]$, and one sample $\mathbf{x} = [1, 2, 3]$.
- $W\mathbf{x} = [1\cdot1 + 0\cdot2 + (-1)\cdot3,\; 2\cdot1 + 1\cdot2 + 0\cdot3] = [-2, 4]$.
- Add the bias: $[-2 + 1,\; 4 - 1] = [-1, 3]$.
For a batch $X$ with one sample per row, the same layer is $Y = XW^\top + \mathbf{1}\mathbf{b}^\top$: row $i$ of $Y$ is $(W\mathbf{x}_i + \mathbf{b})^\top$. Shapes: $(n\times d)(d\times k) = n \times k$.
A dense layer with weights $W \in \mathbb{R}^{k\times d}$ and bias $\mathbf{b} \in \mathbb{R}^k$ maps
$$\mathbf{y} = W\mathbf{x} + \mathbf{b} \quad\text{(one sample)}, \qquad Y = XW^\top + \mathbf{1}\mathbf{b}^\top \quad\text{(a batch of } n \text{ samples as rows)}.$$The bias vector is added to every row (this is called broadcasting). Usually a non-linear function (like ReLU, $\max(0, z)$) is applied after. You will see in Chapter 1.5 why that last step is essential.
Why do we need it?
A neural network needs a rule that turns d input numbers into k output numbers. Weighted sums plus a bias are the simplest such rule, and a matrix writes all the weighted sums at once.
Where is it used?
Every fully connected layer of an MLP, the query, key and value projections in a Transformer, the final classifier layer, and PyTorch's nn.Linear or Keras Dense.
How is it used?
Stack your samples as the rows of X and compute Y = X @ W.T + b. W has shape (outputs, inputs), and b is added to every row (broadcasting). Check that the batch size is kept: (n, d) becomes (n, k).
$W$ is $k \times d$ in the maths, but frameworks differ. PyTorch's nn.Linear stores W as (out, in) and computes X @ W.T + b. Some libraries store the transpose and compute X @ W. Always check the shapes.
Quick check: a layer maps 784 inputs to 128 outputs. What are the shapes of $W$, $\mathbf{b}$, and of $Y$ for a batch of 32?
$W$ is $128 \times 784$, $\mathbf{b}$ has 128 entries, $X$ is $32 \times 784$, and $Y = XW^\top + \mathbf{b}$ is $(32\times784)(784\times128) = 32 \times 128$.
Recap, cheat sheet and practice
- A matrix is a table, a stack of vectors, a transformation, or a data set (rows = samples, columns = features). Size is rows $\times$ columns.
- Add, subtract and scale cell by cell. The Hadamard product $A \odot B$ multiplies cell by cell too, and is not the matrix product.
- $A\mathbf{x}$ = dot products of rows with $\mathbf{x}$ (row view) = a mix of the columns of $A$ (column view).
- $AB$: inner numbers must match, outer numbers give the shape. Four views: entry, column, row, outer-product sum. Associative, distributive, but not commutative. Cost $O(mnp)$.
- Transpose flips rows and columns. $(AB)^\top = B^\top A^\top$. The Gram matrix $A^\top A$ is always symmetric.
- Special matrices: identity, zero, diagonal, symmetric, skew, triangular, orthogonal ($Q^\top Q = I$), permutation, sparse, block.
- A dense layer is $W\mathbf{x} + \mathbf{b}$; a batch is $XW^\top + \mathbf{b}$.
Cheat sheet
| Idea | Formula | Remember |
|---|---|---|
| Shape | $A \in \mathbb{R}^{m\times n}$ | rows first, then columns |
| Hadamard | $(A \odot B)_{ij} = A_{ij}B_{ij}$ | cell by cell, same shape |
| Matrix-vector | $A\mathbf{x} = \sum_j x_j\mathbf{a}_j$ | mix of columns |
| Matrix-matrix (entry) | $(AB)_{ij} = \text{row}_i(A)\cdot\text{col}_j(B)$ | $(m\times n)(n\times p) = m\times p$ |
| Outer-product view | $AB = \sum_k \mathbf{a}_k\mathbf{b}_k^\top$ | sum of rank-1 pieces |
| Not commutative | $AB \ne BA$ | order matters |
| Cost | $O(mnp)$ | bracket chains wisely |
| Transpose of product | $(AB)^\top = B^\top A^\top$ | reverse the order |
| Gram matrix | $G = A^\top A$ | column dot products, symmetric |
| Orthogonal | $Q^\top Q = I$ | keeps lengths and angles |
| Dense layer | $\mathbf{y} = W\mathbf{x} + \mathbf{b}$, $Y = XW^\top + \mathbf{b}$ | one product per batch |
import numpy as np, time
A = np.array([[1, 2, 3], [4, 5, 6]]) # shape (2, 3)
B = np.array([[1, 0], [2, 1], [0, 3]]) # shape (3, 2)
print(A.shape, A[0, 1]) # (2, 3) 2 (row 1, column 2 in our notation)
print(A + A, 3 * A) # add and scale cell by cell
print(A * A) # Hadamard (element-wise) product
x = np.array([2, 1, 0])
print(A @ x) # matrix-vector: [4 13]
# ---- matrix multiplication with triple loops, and timing against @ ----
def matmul_loops(A, B):
m, n = A.shape
n2, p = B.shape
assert n == n2, "inner numbers must match"
C = np.zeros((m, p))
for i in range(m):
for j in range(p):
for k in range(n):
C[i, j] += A[i, k] * B[k, j] # entry view: m*n*p multiplications
return C
R1, R2 = np.random.rand(120, 120), np.random.rand(120, 120)
t0 = time.time(); C1 = matmul_loops(R1, R2); t1 = time.time()
C2 = R1 @ R2; t2 = time.time()
print(np.allclose(C1, C2), "loops:", round(t1 - t0, 3), "s @:", round(t2 - t1, 5), "s")
# ---- all four views give the same answer ----
C = A @ B
entry = np.array([[A[i] @ B[:, j] for j in range(2)] for i in range(2)])
column = np.column_stack([A @ B[:, j] for j in range(2)]) # A times each column of B
row = np.vstack([A[i] @ B for i in range(2)]) # each row of A times B
outer = sum(np.outer(A[:, k], B[k]) for k in range(3)) # sum of rank-1 pieces
print(all(np.allclose(C, V) for V in (entry, column, row, outer))) # True
# ---- transpose, Gram matrix ----
print(np.allclose((A @ B).T, B.T @ A.T)) # True: (AB)^T = B^T A^T
G = A.T @ A # Gram matrix, shape (3, 3)
print(np.allclose(G, G.T)) # True: always symmetric
# ---- special matrices ----
I = np.eye(3); Z = np.zeros((3, 3)); D = np.diag([2, 3, -1])
U = np.triu(np.ones((3, 3))); L = np.tril(np.ones((3, 3)))
P = np.eye(3)[[1, 2, 0]] # permutation: shuffled rows of I
t = np.pi / 6
Q = np.array([[np.cos(t), -np.sin(t)], [np.sin(t), np.cos(t)]])
print(np.allclose(Q.T @ Q, np.eye(2)), np.allclose(P.T @ P, np.eye(3))) # True True
# ---- sparse and block ----
from scipy import sparse
S = sparse.csr_matrix(np.eye(1000)) # stores only the 1000 non-zeros
print(S.nnz, (S @ np.ones(1000)).sum()) # 1000 1000.0
M = np.block([[A[:, :2], np.zeros((2, 2))], [np.zeros((2, 2)), np.ones((2, 2))]]) # block matrix
print(M.shape) # (4, 4)
# ---- a dense layer on a batch: Y = X W^T + b ----
W = np.array([[1, 0, -1], [2, 1, 0]]) # (out=2, in=3)
b = np.array([1, -1])
X = np.array([[1, 2, 3], [0, 1, 0], [2, 0, 1]]) # (n=3 samples, d=3 features)
Y = X @ W.T + b # b is broadcast to every row
print(Y) # first row: [-1 3]
1. $A$ is $3 \times 4$ and $B$ is $4 \times 2$. What is $AB$?
2. Using the column view, $\begin{bmatrix}1&2\\3&4\end{bmatrix}\begin{bmatrix}2\\1\end{bmatrix}$ equals…
3. Which is equal to $(AB)^\top$?
4. Which statement about $G = A^\top A$ is true for every matrix $A$?
5. How many multiplications does the standard method need for $(2 \times 3)(3 \times 4)$?
6. $Q$ is an orthogonal matrix. Which statement is true?
Practice problems
A. For $A = \begin{bmatrix}1&2\\0&1\end{bmatrix}$ and $B = \begin{bmatrix}3&0\\1&1\end{bmatrix}$, compute $AB$ and $BA$.
$AB = \begin{bmatrix} 1\cdot3+2\cdot1 & 1\cdot0+2\cdot1 \\ 0\cdot3+1\cdot1 & 0\cdot0+1\cdot1 \end{bmatrix} = \begin{bmatrix}5&2\\1&1\end{bmatrix}$. $BA = \begin{bmatrix} 3\cdot1+0\cdot0 & 3\cdot2+0\cdot1 \\ 1\cdot1+1\cdot0 & 1\cdot2+1\cdot1 \end{bmatrix} = \begin{bmatrix}3&6\\1&3\end{bmatrix}$. They are different.
B. Compute $A\mathbf{x}$ with $A = \begin{bmatrix}1&0&2\\0&1&1\end{bmatrix}$ and $\mathbf{x} = [3, -1, 2]$ using the column view.
$3\begin{bmatrix}1\\0\end{bmatrix} - 1\begin{bmatrix}0\\1\end{bmatrix} + 2\begin{bmatrix}2\\1\end{bmatrix} = \begin{bmatrix}3\\0\end{bmatrix} + \begin{bmatrix}0\\-1\end{bmatrix} + \begin{bmatrix}4\\2\end{bmatrix} = \begin{bmatrix}7\\1\end{bmatrix}$. Check with the row view: $1\cdot3 + 0 + 2\cdot2 = 7$ and $0 + 1\cdot(-1) + 1\cdot2 = 1$ ✓.
C. $A$ is $5 \times 3$ and $B$ is $3 \times 2$. Which of $AB$, $BA$, $A^\top B$, $B^\top A^\top$ are defined, and what are their shapes?
$AB$: $(5\times3)(3\times2) = 5\times2$ ✓. $BA$: $(3\times2)(5\times3)$, and $2 \ne 5$, so ✗. $A^\top B$: $(3\times5)(3\times2)$, and $5 \ne 3$, so ✗. $B^\top A^\top$: $(2\times3)(3\times5) = 2\times5$ ✓ (it equals $(AB)^\top$).
D. Compute the Gram matrix $X^\top X$ for $X = \begin{bmatrix}1&1\\1&2\\1&3\end{bmatrix}$. Is it symmetric?
Columns: $[1,1,1]$ and $[1,2,3]$. Dot products: $1+1+1 = 3$, $1+2+3 = 6$, $1+4+9 = 14$. So $X^\top X = \begin{bmatrix}3&6\\6&14\end{bmatrix}$. Yes, it is symmetric.
E. Show that $Q = \begin{bmatrix}0.6&-0.8\\0.8&0.6\end{bmatrix}$ is orthogonal.
Column 1 has length $\sqrt{0.36+0.64} = 1$, column 2 has length $\sqrt{0.64+0.36} = 1$, and their dot product is $0.6\cdot(-0.8) + 0.8\cdot0.6 = -0.48 + 0.48 = 0$. So the columns are orthonormal, and $Q^\top Q = I$.
F. $A$ is $2 \times 1000$, $B$ is $1000 \times 2$, $C$ is $2 \times 1000$. Which bracketing of $ABC$ is cheaper, and by how much?
$(AB)C$: $2\cdot1000\cdot2 + 2\cdot2\cdot1000 = 4000 + 4000 = 8000$. $A(BC)$: $1000\cdot2\cdot1000 + 2\cdot1000\cdot1000 = 2{,}000{,}000 + 2{,}000{,}000 = 4{,}000{,}000$. So $(AB)C$ is $500$ times cheaper.
Matrix Geometry & Linear Transformations
Every matrix is a function that moves space. In this chapter you will see matrices rotate, stretch, flip, shear and flatten the plane, and you will learn the one fact that makes it all easy: the columns of a matrix show where the basic arrows land.
- Say exactly what makes a transformation "linear", and see why grid lines stay straight
- Read a $2 \times 2$ matrix as "where do $\mathbf{e}_1$ and $\mathbf{e}_2$ land?"
- Build scaling, rotation, reflection, shear and projection matrices
- Understand composition: multiplying matrices means doing transformations one after another (and order matters)
- Undo a transformation with an inverse
- Add a shift to get affine transformations, and see why a shift is not linear
- Describe one transformation in a different coordinate system (similar matrices)
- Preview kernel and image, and see why a neural-network layer needs a non-linearity
Linear transformations core
A transformation is a rule that takes every arrow in the plane and moves it to a new arrow. Picture the whole plane drawn on a rubber sheet with a square grid printed on it. A transformation grabs the sheet and deforms it.
A transformation is linear when the deformation is "fair" in three ways:
- The origin stays where it is.
- Grid lines stay straight lines (no bending).
- Grid lines stay parallel and evenly spaced (no uneven stretching).
Rotating, stretching, flipping and shearing the sheet are all fair. Sliding the whole sheet sideways is not (the origin moves). Bending it like a wave is not (the lines curve).
Take the rule $T(x, y) = (2x + y,\; x - y)$. Test the fairness idea on $\mathbf{x} = [1, 2]$ and $\mathbf{y} = [3, 1]$.
- $T(\mathbf{x}) = (2\cdot1 + 2,\; 1 - 2) = (4, -1)$ and $T(\mathbf{y}) = (2\cdot3 + 1,\; 3 - 1) = (7, 2)$. Their sum is $(11, 1)$.
- Add first: $\mathbf{x} + \mathbf{y} = (4, 3)$, so $T(\mathbf{x}+\mathbf{y}) = (2\cdot4 + 3,\; 4 - 3) = (11, 1)$. The same. ✓
Now a rule that slides everything: $S(x, y) = (x + 1,\; y)$. Then $S(\mathbf{x}) + S(\mathbf{y}) = (2, 2) + (4, 1) = (6, 3)$ but $S(\mathbf{x}+\mathbf{y}) = S(4, 3) = (5, 3)$. Different. ✗ So $S$ is not linear.
A function $T$ that takes vectors to vectors is a linear transformation (or linear map) if for all vectors $\mathbf{x}, \mathbf{y}$ and all numbers $a, b$:
$$T(a\mathbf{x} + b\mathbf{y}) = a\,T(\mathbf{x}) + b\,T(\mathbf{y}).$$This one rule packs two: additivity $T(\mathbf{x}+\mathbf{y}) = T(\mathbf{x}) + T(\mathbf{y})$ and homogeneity $T(c\mathbf{x}) = c\,T(\mathbf{x})$. In words: "transforming a linear combination is the same as taking the linear combination of the transformed vectors." (Linear combinations are in Chapter 1.2.)
A consequence: $T(\mathbf{0}) = \mathbf{0}$ (take $c = 0$). A map that moves the origin can never be linear.
Why do we need it?
We want transformations that are simple to describe, combine and compute. Linear ones are exactly those: they keep sums and scalings, so one small table of numbers captures them completely.
Where is it used?
The matrix part of every neural-network layer, rotations and scalings in graphics, data augmentation, PCA projections, and the theory behind linear regression.
How is it used?
To test whether a map is linear, check three things: the origin stays put, T(x + y) = T(x) + T(y), and T(cx) = cT(x). If any check fails (a shift, squaring, ReLU), no single matrix can describe the map.
"Linear" in school means a straight line $y = mx + c$. Here it means something stricter. The function $f(x) = 2x + 3$ draws a straight line, but it is not linear in the linear-algebra sense, because $f(0) = 3 \ne 0$. Maps like that have a special name: affine (see the section on affine transformations).
Squaring, taking absolute values, and applying $\max(0, \cdot)$ (ReLU) are not linear either.
Quick check: is $T(x, y) = (x, 0)$ linear? What about $T(x, y) = (x, 1)$?
$T(x, y) = (x, 0)$ is linear: it keeps the origin, and $T(a\mathbf{u} + b\mathbf{v}) = (au_1 + bv_1, 0) = aT(\mathbf{u}) + bT(\mathbf{v})$. (It flattens the plane onto the x-axis.) But $T(x, y) = (x, 1)$ sends the origin to $(0, 1)$, so it is not linear.
Every linear map is a matrix: the columns tell you where the basis vectors land core
Here is the big secret. You do not need to know what a linear map does to every arrow. You only need to know what it does to two arrows: $\mathbf{e}_1 = [1, 0]$ (one step right) and $\mathbf{e}_2 = [0, 1]$ (one step up).
Why? Every arrow is a recipe made from these two: $[3, 2]$ means "3 scoops of $\mathbf{e}_1$ and 2 scoops of $\mathbf{e}_2$". A linear map keeps recipes intact, so the new arrow is "3 scoops of the new $\mathbf{e}_1$ and 2 scoops of the new $\mathbf{e}_2$".
So: write down where $\mathbf{e}_1$ lands as column 1, and where $\mathbf{e}_2$ lands as column 2. That table is the matrix.
A linear map sends $\mathbf{e}_1 = [1, 0]$ to $[2, 1]$ and $\mathbf{e}_2 = [0, 1]$ to $[-1, 1]$. Where does $[3, 2]$ go?
- Put the landing spots in as columns: $A = \begin{bmatrix} 2 & -1 \\ 1 & 1 \end{bmatrix}$.
- $[3, 2] = 3\mathbf{e}_1 + 2\mathbf{e}_2$, so it lands on $3\begin{bmatrix}2\\1\end{bmatrix} + 2\begin{bmatrix}-1\\1\end{bmatrix} = \begin{bmatrix}6\\3\end{bmatrix} + \begin{bmatrix}-2\\2\end{bmatrix} = \begin{bmatrix}4\\5\end{bmatrix}$.
- Check with the matrix-times-vector rule from Chapter 1.4: $A\begin{bmatrix}3\\2\end{bmatrix} = \begin{bmatrix}2\cdot3 + (-1)\cdot2 \\ 1\cdot3 + 1\cdot2\end{bmatrix} = \begin{bmatrix}4\\5\end{bmatrix}$ ✓.
Same answer: "matrix times vector" is "apply the transformation".
Theorem. Every linear transformation $T:\mathbb{R}^n \to \mathbb{R}^m$ is multiplication by a unique $m \times n$ matrix $A$, namely
$$A = \big[\; T(\mathbf{e}_1) \;\; T(\mathbf{e}_2) \;\; \cdots \;\; T(\mathbf{e}_n) \;\big], \qquad T(\mathbf{x}) = A\mathbf{x}.$$Why: write $\mathbf{x} = x_1\mathbf{e}_1 + \dots + x_n\mathbf{e}_n$. By linearity, $T(\mathbf{x}) = x_1T(\mathbf{e}_1) + \dots + x_nT(\mathbf{e}_n)$, which is the column view of $A\mathbf{x}$. Conversely, every matrix gives a linear map, because $A(a\mathbf{x} + b\mathbf{y}) = aA\mathbf{x} + bA\mathbf{y}$.
So "linear map" and "matrix" are two names for one thing. In $\mathbb{R}^2$ the first column is where $\mathbf{e}_1$ lands and the second is where $\mathbf{e}_2$ lands.
Why do we need it?
Without this fact we would have to remember what a map does to every single vector. With it, a handful of columns describes everything.
Where is it used?
Building rotation and scaling matrices by hand, reading what a trained weight matrix does, converting between coordinate systems, and designing data-augmentation transforms.
How is it used?
To find the matrix of a linear map, apply the map to e1, e2, and so on, and write the results as columns. To read a matrix someone gives you, look at its columns: they show where the unit arrows land. Then A @ x mixes those columns.
The same idea works in 3D. A $3 \times 3$ matrix has three columns: where $\mathbf{e}_1$, $\mathbf{e}_2$ and $\mathbf{e}_3$ land. The unit cube (built from those three arrows) becomes a slanted box, called a parallelepiped. Its volume is the absolute value of a number called the determinant (Chapter 1.7), and the sign tells you if space was flipped like a mirror.
Columns, not rows. It is a very common mistake to read the landing spots off the rows. The first column is where $\mathbf{e}_1$ goes. (Check: $A\mathbf{e}_1$ picks out column 1.)
Everything is relative to the origin. The matrix tells you how arrows move, and the origin never moves.
Quick check: a linear map sends $\mathbf{e}_1 \mapsto [0, 3]$ and $\mathbf{e}_2 \mapsto [1, 0]$. What is its matrix, and where does $[2, 5]$ go?
$A = \begin{bmatrix} 0 & 1 \\ 3 & 0 \end{bmatrix}$ (columns are the landing spots). $[2, 5] \mapsto 2[0, 3] + 5[1, 0] = [5, 6]$. Check by rows: $[0\cdot2 + 1\cdot5,\; 3\cdot2 + 0\cdot5] = [5, 6]$ ✓.
Scaling core
Scaling stretches or shrinks space along the axes. Pull the rubber sheet wider and squeeze it shorter: every shape gets wider and shorter too.
To build the matrix, ask where $\mathbf{e}_1$ and $\mathbf{e}_2$ go. If we stretch $x$ by $s_x$, then $\mathbf{e}_1 = [1, 0]$ becomes $[s_x, 0]$. If we stretch $y$ by $s_y$, then $\mathbf{e}_2$ becomes $[0, s_y]$. Put those in as columns and you get a diagonal matrix.
Stretch $x$ by 2 and $y$ by 3: $S = \begin{bmatrix} 2 & 0 \\ 0 & 3 \end{bmatrix}$.
- $S\begin{bmatrix}1\\1\end{bmatrix} = \begin{bmatrix}2\\3\end{bmatrix}$. The corner of the unit square moves to $(2, 3)$.
- The unit square becomes a $2 \times 3$ rectangle, with area $6 = 2\cdot3$.
- Using a negative number, $s_x = -1$, flips left and right.
If $s_x = s_y = s$, it is uniform scaling (a zoom): $S = sI$. Areas are multiplied by $|s_xs_y|$. In 3D: $\operatorname{diag}(s_x, s_y, s_z)$. Scaling is the geometric meaning of a diagonal matrix from Chapter 1.4.
Why do we need it?
Zooming, or stretching one direction more than another, is the simplest change we make to data and pictures.
Where is it used?
Feature standardisation, resizing images, zoom augmentation, weight scaling and decay, and the diagonal matrix inside the SVD.
How is it used?
Put the stretch factors on the diagonal: np.diag([sx, sy]). A factor above 1 stretches, a factor between 0 and 1 shrinks, a negative factor flips, and 0 flattens that direction. Areas change by the product of the factors.
Quick check: what matrix halves every coordinate? By what factor does it change areas?
$\begin{bmatrix} 0.5 & 0 \\ 0 & 0.5 \end{bmatrix}$ (that is $0.5\,I$). Areas shrink by $0.5 \cdot 0.5 = 0.25$, to a quarter.
Rotation core
A rotation spins the whole sheet around the origin by an angle $\theta$ (anticlockwise is positive). Nothing is stretched, so lengths, angles and areas are all kept.
Let us find the matrix by asking where $\mathbf{e}_1$ and $\mathbf{e}_2$ go. After a turn by $\theta$, the arrow $[1, 0]$ lands on the point of the unit circle at angle $\theta$, which is $[\cos\theta, \sin\theta]$. The arrow $[0, 1]$ was already $90^\circ$ ahead, so it lands at angle $\theta + 90^\circ$, which is $[-\sin\theta, \cos\theta]$. Those are the two columns.
- $\theta = 90^\circ$: $\cos = 0$, $\sin = 1$, so $R = \begin{bmatrix} 0 & -1 \\ 1 & 0 \end{bmatrix}$. It sends $[1, 0] \mapsto [0, 1]$ (right becomes up).
- $\theta = 30^\circ$: $\cos30^\circ \approx 0.866$ and $\sin30^\circ = 0.5$, so $R \approx \begin{bmatrix} 0.866 & -0.5 \\ 0.5 & 0.866 \end{bmatrix}$. It sends $[1, 0] \mapsto [0.866, 0.5]$.
- $\theta = 180^\circ$: $R = \begin{bmatrix} -1 & 0 \\ 0 & -1 \end{bmatrix} = -I$. Every arrow turns around.
Properties: the columns are orthonormal (perpendicular to each other, each of length 1), so $R$ is an orthogonal matrix ($R^\top R = I$, see Chapter 1.4). It keeps lengths and angles. Undoing a rotation by $\theta$ means rotating by $-\theta$, so $R(\theta)^{-1} = R(-\theta) = R(\theta)^\top$. Two rotations add their angles: $R(\alpha)R(\beta) = R(\alpha + \beta)$.
In 3D you rotate about an axis. For example, about the $z$-axis:
$$R_z(\theta) = \begin{bmatrix} \cos\theta & -\sin\theta & 0 \\ \sin\theta & \cos\theta & 0 \\ 0 & 0 & 1 \end{bmatrix}.$$The $z$-arrow stays put (third column is $[0, 0, 1]$) and the $xy$-plane turns. The matrices for the other axes are built the same way.
Why do we need it?
We often need to turn data or objects without changing their size or shape. A rotation matrix does this exactly, with the angle as the only choice.
Where is it used?
Rotation augmentation of images, rotary position embeddings (RoPE) in language models, robotics and 3D vision, and rotating data to line up with its main axes.
How is it used?
Build R(theta) from cos and sin: the columns are [cos, sin] and [-sin, cos]. Apply it with R @ x. To undo it use the transpose (rotate by minus theta). In 3D choose an axis and put the 2D rotation in the other two coordinates.
Quick check: what does $R(90^\circ)$ do to $[2, 3]$?
$\begin{bmatrix} 0 & -1 \\ 1 & 0 \end{bmatrix}\begin{bmatrix}2\\3\end{bmatrix} = \begin{bmatrix}0\cdot2 - 3 \\ 2 + 0\end{bmatrix} = \begin{bmatrix}-3\\2\end{bmatrix}$. The arrow turned a quarter-turn anticlockwise, and its length ($\sqrt{13}$) is unchanged.
Reflection
A reflection flips the plane over a mirror line through the origin. Every point jumps to the other side of the mirror, the same distance away. The letter F turns into a backwards F: a mirror image.
Where do $\mathbf{e}_1$ and $\mathbf{e}_2$ go? For the mirror being the x-axis, $\mathbf{e}_1$ is already on the mirror, so it stays $[1, 0]$. And $\mathbf{e}_2 = [0, 1]$ jumps to $[0, -1]$.
- Mirror = the x-axis: columns $[1, 0]$ and $[0, -1]$, so $\begin{bmatrix} 1 & 0 \\ 0 & -1 \end{bmatrix}$ (the sign of $y$ flips).
- Mirror = the y-axis: $\begin{bmatrix} -1 & 0 \\ 0 & 1 \end{bmatrix}$.
- Mirror = the line $y = x$: $\mathbf{e}_1 = [1, 0]$ lands on $[0, 1]$ and $\mathbf{e}_2$ lands on $[1, 0]$, so $\begin{bmatrix} 0 & 1 \\ 1 & 0 \end{bmatrix}$. It simply swaps the two coordinates: $[3, 5] \mapsto [5, 3]$.
Reflection across the line through the origin at angle $\alpha$ (measured from the x-axis):
$$F(\alpha) = \begin{bmatrix} \cos 2\alpha & \sin 2\alpha \\ \sin 2\alpha & -\cos 2\alpha \end{bmatrix}.$$(Check: $\alpha = 0$ gives the x-axis case, and $\alpha = 45^\circ$ gives the swap.) A reflection is orthogonal and symmetric. Doing it twice gives back the original: $F^2 = I$. Because it makes a mirror image, it reverses orientation (the determinant is $-1$; see Chapter 1.7). In 3D you reflect across a plane, for example $\operatorname{diag}(1, 1, -1)$ flips $z$.
Why do we need it?
Mirror images appear in data (left versus right) and inside numerical methods. A reflection matrix flips space across a line or a plane while keeping all sizes.
Where is it used?
Horizontal-flip augmentation of images, Householder reflections inside QR factorisation, and symmetry in physics and graphics.
How is it used?
Choose the mirror line at angle alpha and build [[cos 2a, sin 2a], [sin 2a, -cos 2a]]. Apply it with F @ x. Applying it twice gives back the original, and it flips orientation (its determinant is minus 1).
Quick check: apply the reflection across $y = x$ to $[2, -1]$.
$\begin{bmatrix}0&1\\1&0\end{bmatrix}\begin{bmatrix}2\\-1\end{bmatrix} = \begin{bmatrix}-1\\2\end{bmatrix}$. The two coordinates were swapped.
Shear
Take a deck of cards stacked in a neat pile, and push the top of the deck sideways. The bottom card stays, the top card moves most, and every card in between slides a bit. The pile leans over, but each card keeps its width and the deck keeps its height. That is a shear.
A horizontal shear slides every point sideways by an amount that grows with its height $y$. The x-axis (height 0) stays fixed.
Horizontal shear with $k = 1$: $(x, y) \mapsto (x + y,\; y)$.
- $\mathbf{e}_1 = [1, 0]$ stays $[1, 0]$ (height 0, no slide).
- $\mathbf{e}_2 = [0, 1]$ slides right by 1 and becomes $[1, 1]$.
So the matrix is $\begin{bmatrix} 1 & 1 \\ 0 & 1 \end{bmatrix}$. The unit square becomes a leaning parallelogram with the same area (same base 1, same height 1).
A shear keeps areas the same (determinant 1), but it does not keep lengths or angles. It is easy to undo: shear by $-k$.
Why do we need it?
Some changes lean shapes over without changing their area, like italic text or a pushed deck of cards. Shear matrices describe exactly that.
Where is it used?
Shear augmentation of images, italic and perspective effects in graphics, and row operations in Gaussian elimination, where 'add k times row 1 to row 2' is a shear.
How is it used?
Put k beside the diagonal of the identity: [[1, k], [0, 1]] is a horizontal shear. Apply it with S @ x and undo it with minus k. The area stays the same, but lengths and angles change.
Quick check: where does the vertical shear $\begin{bmatrix} 1 & 0 \\ 2 & 1 \end{bmatrix}$ send $[1, 1]$?
$[1\cdot1 + 0\cdot1,\; 2\cdot1 + 1\cdot1] = [1, 3]$. The point slid up by $2 \times x = 2$.
Projection core
A projection drops every point straight onto a line, like a shadow cast by a lamp directly overhead. The whole plane gets flattened onto that line. This is the first transformation that loses information: many different points end up with the same shadow.
It is the transformation version of the projection you met in Chapter 1.2.
- Project onto the x-axis: $(x, y) \mapsto (x, 0)$. The matrix is $\begin{bmatrix} 1 & 0 \\ 0 & 0 \end{bmatrix}$. The points $(3, 2)$ and $(3, -7)$ both land on $(3, 0)$.
- Project onto the line along $\mathbf{u} = \tfrac{1}{\sqrt2}[1, 1]$. Using the formula below, $P = \begin{bmatrix} 0.5 & 0.5 \\ 0.5 & 0.5 \end{bmatrix}$, and $P\begin{bmatrix}3\\1\end{bmatrix} = \begin{bmatrix}0.5\cdot3 + 0.5\cdot1 \\ 0.5\cdot3 + 0.5\cdot1\end{bmatrix} = \begin{bmatrix}2\\2\end{bmatrix}$. (The same answer as the projection example in Chapter 1.2.)
Projection onto the line through the origin in the direction of a unit vector $\mathbf{u}$:
$$P = \mathbf{u}\mathbf{u}^\top \quad(\text{an outer product}), \qquad P\mathbf{x} = (\mathbf{u}\cdot\mathbf{x})\,\mathbf{u}.$$For a general (not unit) direction $\mathbf{a}$: $P = \dfrac{\mathbf{a}\mathbf{a}^\top}{\mathbf{a}^\top\mathbf{a}}$. Properties:
- $P^2 = P$: projecting twice is the same as projecting once ("idempotent").
- $P$ is symmetric.
- $P$ squashes the plane to a line, so it has no inverse: you cannot recover $(x, y)$ from its shadow.
Projecting onto a plane in 3D works the same way: for the $xy$-plane, $P = \operatorname{diag}(1, 1, 0)$. For a plane spanned by the columns of a matrix $A$ the formula is $P = A(A^\top A)^{-1}A^\top$ (you will meet this in Chapter 1.9).
Why do we need it?
We often want only the part of a vector that lies along some direction, and want to throw the rest away. A projection matrix does that in one multiplication.
Where is it used?
PCA (project onto the main directions), least squares and linear regression (project y onto the space the features can reach), dimensionality reduction, and bottleneck layers.
How is it used?
For a unit direction u use P = np.outer(u, u); for a general direction a use np.outer(a, a) / (a @ a). Then P @ x is the shadow and x - P @ x is the leftover. Applying P twice changes nothing more, and P has no inverse.
A projection is not a reflection or a rotation. Those can be undone. A projection destroys the sideways part of every vector.
Quick check: the matrix $\begin{bmatrix} 1 & 0 & 0 \\ 0 & 1 & 0 \\ 0 & 0 & 0 \end{bmatrix}$ acts on 3D space. What does it do?
It keeps $x$ and $y$ and sets $z$ to 0: a projection onto the flat $xy$-plane (the "floor"). Applying it twice changes nothing more.
Composition: doing one transformation after another core
What if you rotate the sheet and then stretch it? The combined effect is itself a linear transformation. So it must have a matrix. Which one? It is the matrix product.
This is the real reason matrix multiplication is defined the way it is. In the column view of $AB$ from Chapter 1.4, "$B$ comes first, then $A$ acts on what $B$ produced".
And order matters. Putting on socks then shoes is not shoes then socks. In the same way, rotating then stretching is not the same as stretching then rotating.
Let $A$ = rotate by $90^\circ$, so $A = \begin{bmatrix} 0 & -1 \\ 1 & 0 \end{bmatrix}$, and let $B$ = stretch $x$ by 2, so $B = \begin{bmatrix} 2 & 0 \\ 0 & 1 \end{bmatrix}$.
Rotate first, then stretch. The input $\mathbf{x}$ goes through $A$ and then $B$: $B(A\mathbf{x}) = (BA)\mathbf{x}$.
$$BA = \begin{bmatrix} 2 & 0 \\ 0 & 1 \end{bmatrix}\begin{bmatrix} 0 & -1 \\ 1 & 0 \end{bmatrix} = \begin{bmatrix} 0 & -2 \\ 1 & 0 \end{bmatrix}.$$Stretch first, then rotate: $A(B\mathbf{x}) = (AB)\mathbf{x}$.
$$AB = \begin{bmatrix} 0 & -1 \\ 1 & 0 \end{bmatrix}\begin{bmatrix} 2 & 0 \\ 0 & 1 \end{bmatrix} = \begin{bmatrix} 0 & -1 \\ 2 & 0 \end{bmatrix}.$$Check with $\mathbf{e}_1 = [1, 0]$. Rotate then stretch: $[1,0] \to [0,1] \to [0,1]$, which is the first column of $BA$ ✓. Stretch then rotate: $[1,0] \to [2,0] \to [0,2]$, which is the first column of $AB$ ✓. Different results.
If $T_A(\mathbf{x}) = A\mathbf{x}$ and $T_B(\mathbf{x}) = B\mathbf{x}$, the composition "first $T_A$, then $T_B$" is
$$(T_B \circ T_A)(\mathbf{x}) = B(A\mathbf{x}) = (BA)\mathbf{x}.$$The matrix of a composition is the product, written in reverse order: the transformation applied first is on the right, next to $\mathbf{x}$. For three steps, $C(B(A\mathbf{x})) = (CBA)\mathbf{x}$. Because matrix multiplication is associative, you may bracket it any way you like. But, since $AB \ne BA$ in general, you may not swap.
Why do we need it?
Real systems do many steps one after another. Composition tells us that a whole chain of transformations is itself one matrix, so we can combine it once and apply it many times.
Where is it used?
Layers stacked in a network, chains of rotations, scalings and shifts in graphics and robotics, and the order of preprocessing steps in a data pipeline.
How is it used?
For 'first A, then B' multiply B @ A, because the first step sits next to x. For three steps use C @ B @ A. Never swap the order without checking: A @ B and B @ A are usually different.
Right to left. In $BA\mathbf{x}$ the matrix nearest to $\mathbf{x}$ acts first. It reads like function notation $g(f(x))$, where $f$ acts first.
Some pairs do commute: two rotations about the same point (so 2D rotations commute), two diagonal scalings, or any matrix with the identity. But you cannot count on it.
Quick check: first rotate by $90^\circ$, then flip over the x-axis, $\begin{bmatrix}1&0\\0&-1\end{bmatrix}$. Which matrix is that, and where does $\mathbf{e}_1$ end up?
The matrix is (flip)(rotate) $= \begin{bmatrix}1&0\\0&-1\end{bmatrix}\begin{bmatrix}0&-1\\1&0\end{bmatrix} = \begin{bmatrix}0&-1\\-1&0\end{bmatrix}$. Its first column is $[0, -1]$: $\mathbf{e}_1 \to [0, 1]$ (rotate) $\to [0, -1]$ (flip) ✓.
Inverse transformations core
An inverse is the "undo" button. If a transformation turned the sheet, the inverse turns it back. If it stretched, the inverse squeezes.
- Undo a rotation by $\theta$: rotate by $-\theta$.
- Undo a stretch by 2: stretch by $\tfrac12$.
- Undo a shear by $k$: shear by $-k$.
But some actions cannot be undone. If you flatten the sheet onto a line (a projection), then many different points sit on the same spot, and there is no way to tell where each came from. A transformation that squashes space has no inverse.
Let $A = \begin{bmatrix} 2 & 1 \\ 1 & 1 \end{bmatrix}$. Its inverse is $A^{-1} = \begin{bmatrix} 1 & -1 \\ -1 & 2 \end{bmatrix}$. Check by multiplying:
$$A^{-1}A = \begin{bmatrix} 1 & -1 \\ -1 & 2 \end{bmatrix}\begin{bmatrix} 2 & 1 \\ 1 & 1 \end{bmatrix} = \begin{bmatrix} 1\cdot2 - 1\cdot1 & 1\cdot1 - 1\cdot1 \\ -1\cdot2 + 2\cdot1 & -1\cdot1 + 2\cdot1 \end{bmatrix} = \begin{bmatrix} 1 & 0 \\ 0 & 1 \end{bmatrix} = I \;✓$$$A$ sends $[1, 0]$ to $[2, 1]$, and $A^{-1}$ sends $[2, 1]$ straight back: $A^{-1}[2, 1] = [2 - 1,\; -2 + 2] = [1, 0]$ ✓.
The inverse of a square matrix $A$ is the matrix $A^{-1}$ with
$$A^{-1}A = AA^{-1} = I.$$As a transformation: first $A$, then $A^{-1}$, and you are back where you started ("do nothing" = the identity). Facts:
- It exists only if $A$ does not squash space (then $A$ is called invertible or non-singular).
- $(AB)^{-1} = B^{-1}A^{-1}$ (socks and shoes again: undo the last step first).
- For $A = \begin{bmatrix} a & b \\ c & d \end{bmatrix}$: $A^{-1} = \dfrac{1}{ad - bc}\begin{bmatrix} d & -b \\ -c & a \end{bmatrix}$, as long as $ad - bc \ne 0$. (The number $ad - bc$ is the determinant, which measures the area scale. Chapter 1.7 explains it fully, and how to invert bigger matrices.)
Why do we need it?
To recover the original input from a result, we need an 'undo' for the transformation. The inverse is that undo, when one exists.
Where is it used?
Solving equations Ax = b, undoing a rotation or a change of basis, invertible networks and normalising flows, and diagnosing information loss in layers that squash space.
How is it used?
First check that the map does not squash space (a flat result means there is no inverse). If it is fine, use np.linalg.inv(A), or better np.linalg.solve(A, b) when you only need the product. Verify with A inverse times A = I.
Only square matrices can have an inverse, and not every square matrix does. Do not write "$A^{-1}$" for $1/A$: matrices cannot be divided that way. Also $(A + B)^{-1}$ is not $A^{-1} + B^{-1}$.
Quick check: what is the inverse of the scaling $\begin{bmatrix} 4 & 0 \\ 0 & 5 \end{bmatrix}$? And of $\begin{bmatrix} 1 & 0 \\ 0 & 0 \end{bmatrix}$?
The first: $\begin{bmatrix} 1/4 & 0 \\ 0 & 1/5 \end{bmatrix}$ (shrink by the same amounts). The second is a projection onto the x-axis: it throws the $y$-information away, so it has no inverse.
Affine transformations: a linear map plus a shift core
Linear maps always keep the origin fixed. But real life has shifts: moving a photo to the right, or adding a bias to a score. The fix is simple. First do a linear map, then slide everything by a fixed vector $\mathbf{b}$.
Doing "linear map + shift" gives an affine transformation: $\mathbf{y} = A\mathbf{x} + \mathbf{b}$. The origin lands on $\mathbf{b}$.
A shift on its own is not linear, because it moves the origin. It breaks the "linear combinations are kept" rule. But it still maps straight lines to straight lines and keeps parallel lines parallel. That is why affine maps are such a well-behaved family.
Let $A = \begin{bmatrix} 0 & -1 \\ 1 & 0 \end{bmatrix}$ (rotate $90^\circ$) and $\mathbf{b} = [2, 1]$, so $T(\mathbf{x}) = A\mathbf{x} + \mathbf{b}$.
- $T([1, 0]) = [0, 1] + [2, 1] = [2, 2]$.
- $T([0, 0]) = [0, 0] + [2, 1] = [2, 1]$. The origin moved, so $T$ is not linear.
Why is it not linear? Test homogeneity with $\mathbf{x} = [1, 0]$: $T(2\mathbf{x}) = A(2\mathbf{x}) + \mathbf{b} = [0, 2] + [2, 1] = [2, 3]$, but $2\,T(\mathbf{x}) = [4, 4]$. They differ, because the shift $\mathbf{b}$ is added once on the left and twice on the right.
An affine transformation is $T(\mathbf{x}) = A\mathbf{x} + \mathbf{b}$ with a matrix $A$ and a vector $\mathbf{b}$. It is linear exactly when $\mathbf{b} = \mathbf{0}$.
Homogeneous coordinates (awareness). There is a neat trick to turn the shift into a matrix too. Add an extra coordinate that is always $1$: write $\mathbf{x} = [x, y]$ as $[x, y, 1]$. Then
$$\begin{bmatrix} A & \mathbf{b} \\ \mathbf{0}^\top & 1 \end{bmatrix}\begin{bmatrix} \mathbf{x} \\ 1 \end{bmatrix} = \begin{bmatrix} A\mathbf{x} + \mathbf{b} \\ 1 \end{bmatrix}.$$For our example: $\begin{bmatrix} 0 & -1 & 2 \\ 1 & 0 & 1 \\ 0 & 0 & 1 \end{bmatrix}\begin{bmatrix} 1 \\ 0 \\ 1 \end{bmatrix} = \begin{bmatrix} 2 \\ 2 \\ 1 \end{bmatrix}$ ✓. Now a chain of rotations, scalings and shifts is just a product of $3 \times 3$ matrices. Graphics and robotics use this all the time.
Why do we need it?
Linear maps can never move the origin, but real tasks need shifts: biases, translations and offsets. Affine maps add the shift while keeping straight lines straight.
Where is it used?
The Wx + b of every dense layer, normalisation layers (scale times x plus shift), image registration and spatial-transformer networks, and graphics and robotics, which use 3 by 3 or 4 by 4 matrices.
How is it used?
Compute A @ x + b. To fold the shift into one matrix, append a 1 to x and build [[A, b], [0, 1]] (homogeneous coordinates). Then a chain of affine steps becomes a product of matrices of one kind.
"Linear" and "affine" are not the same word. $f(x) = 2x + 3$ is affine but not linear. Many ML texts say "linear layer" for $W\mathbf{x} + \mathbf{b}$. That is an affine map, and it is convenient to keep the bias, but only the $W\mathbf{x}$ part obeys the strict linear rules.
Quick check: $T(\mathbf{x}) = 2\mathbf{x} + [1, -1]$. What is $T([3, 2])$, and what is its homogeneous matrix?
$T([3, 2]) = [6, 4] + [1, -1] = [7, 3]$. The matrix is $\begin{bmatrix} 2 & 0 & 1 \\ 0 & 2 & -1 \\ 0 & 0 & 1 \end{bmatrix}$, and it sends $[3, 2, 1]$ to $[7, 3, 1]$.
Coordinate transformations and similar matrices
An arrow is an arrow. But the numbers we write down for it depend on the ruler we use. Describe a trip as "3 blocks east, 3 blocks north", or in the language of a city whose streets run diagonally: "2 blocks along Main Street and 1 along Side Street". Same trip, different numbers.
In Chapter 1.3 you saw a basis: a set of arrows used as the rulers. Changing the basis changes the coordinates. Here we do this with matrices:
- Put the new basis arrows as the columns of a matrix $P$.
- If a vector has coordinates $\mathbf{c}$ in the new basis, its ordinary coordinates are $\mathbf{x} = P\mathbf{c}$ (a mix of the columns of $P$).
- Going the other way: $\mathbf{c} = P^{-1}\mathbf{x}$.
Now, a transformation described in the standard basis has matrix $A$. What is the matrix of the same transformation described in the new basis? Read $P^{-1}AP$ from right to left: convert the new coordinates into ordinary ones ($P$), apply the transformation ($A$), then convert back ($P^{-1}$).
Change of basis. Let the new basis be $\mathbf{p}_1 = [2, 1]$ and $\mathbf{p}_2 = [-1, 1]$, so $P = \begin{bmatrix} 2 & -1 \\ 1 & 1 \end{bmatrix}$. The vector $\mathbf{x} = [3, 3]$ has new coordinates $\mathbf{c} = P^{-1}\mathbf{x}$. Here $P^{-1} = \tfrac13\begin{bmatrix} 1 & 1 \\ -1 & 2 \end{bmatrix}$, so $\mathbf{c} = \tfrac13[6, 3] = [2, 1]$. Check: $2\,[2, 1] + 1\,[-1, 1] = [3, 3]$ ✓.
Similar matrices. Take $A = \begin{bmatrix} 3 & 1 \\ 1 & 3 \end{bmatrix}$ and the basis $\mathbf{p}_1 = [1, 1]$, $\mathbf{p}_2 = [1, -1]$, so $P = \begin{bmatrix} 1 & 1 \\ 1 & -1 \end{bmatrix}$ and $P^{-1} = \tfrac12\begin{bmatrix} 1 & 1 \\ 1 & -1 \end{bmatrix}$.
- $AP = \begin{bmatrix} 3\cdot1 + 1\cdot1 & 3\cdot1 + 1\cdot(-1) \\ 1\cdot1 + 3\cdot1 & 1\cdot1 + 3\cdot(-1) \end{bmatrix} = \begin{bmatrix} 4 & 2 \\ 4 & -2 \end{bmatrix}$.
- $P^{-1}(AP) = \tfrac12\begin{bmatrix} 4 + 4 & 2 - 2 \\ 4 - 4 & 2 + 2 \end{bmatrix} = \begin{bmatrix} 4 & 0 \\ 0 & 2 \end{bmatrix}$.
In this basis the transformation is just stretch the first basis arrow by 4 and the second by 2. A messy matrix became a clean diagonal one, because we looked at the same transformation with better rulers.
- Change of basis: if the columns of an invertible $P$ are the new basis vectors, then $\mathbf{x} = P\mathbf{c}$ and $\mathbf{c} = P^{-1}\mathbf{x}$.
- If $A$ is a transformation in standard coordinates, its matrix in the new basis is $B = P^{-1}AP$.
- Matrices $A$ and $B$ related this way are called similar. They describe the same transformation in different coordinates. They share many properties (for example the determinant and trace, Chapter 1.7, and the eigenvalues, Chapter 1.11).
Why do we need it?
The numbers that describe a vector or a transformation depend on the axes we pick. A good choice of axes can turn a messy matrix into a simple one.
Where is it used?
PCA (data written along its main axes), diagonalisation and the SVD, whitening of features, and graphics (object coordinates versus world coordinates).
How is it used?
Put the new basis vectors as the columns of P. Convert with c = P inverse times x (to new coordinates) and x = P times c (back). A transformation A becomes B = P inverse A P in the new basis. Look for a P that makes B diagonal.
Mind the order: $P^{-1}AP$, not $PAP^{-1}$. Here $P$'s columns are the new basis written in the old coordinates, so $P$ converts new → old and $P^{-1}$ converts old → new. Reading right to left: new→old, apply $A$, old→new.
Similar matrices are the same map in different clothes. They are not equal as tables of numbers.
Quick check: a transformation has the matrix $B = \begin{bmatrix} 4 & 0 \\ 0 & 2 \end{bmatrix}$ in some basis. What does it do to the first basis arrow?
It sends $\mathbf{p}_1$ to $4\mathbf{p}_1$: it stretches that arrow by 4 and keeps its direction (and stretches the second by 2). A diagonal matrix always means "scale each basis arrow separately".
A preview: kernel and image
Two natural questions about any transformation:
- What gets sent to zero? Which inputs does the machine completely ignore? Those arrows form the kernel (also called the null space).
- Where can the outputs land? What is the set of all possible results? That is the image (also called the column space, because outputs are mixes of the columns).
A rotation has a kernel of just $\{\mathbf{0}\}$ (nothing but the zero arrow is lost) and its image is the whole plane. A projection onto a line sends the whole perpendicular direction to zero, and its image is just that line.
Take $A = \begin{bmatrix} 1 & 2 \\ 2 & 4 \end{bmatrix}$.
- Kernel: $A\mathbf{x} = \mathbf{0}$ means $x_1 + 2x_2 = 0$ (and the second row says the same thing, times 2). So $\mathbf{x} = t\,[2, -1]$. Check: $A[2, -1] = [2 - 2,\; 4 - 4] = [0, 0]$ ✓. The kernel is the line through $[2, -1]$.
- Image: every output is $x_1[1, 2] + x_2[2, 4] = (x_1 + 2x_2)[1, 2]$, a multiple of $[1, 2]$. The image is the line through $[1, 2]$.
The matrix flattens the plane onto a line, and one whole line of inputs disappears.
For $A \in \mathbb{R}^{m \times n}$:
$$\ker(A) = \{\mathbf{x} : A\mathbf{x} = \mathbf{0}\}, \qquad \operatorname{im}(A) = \{A\mathbf{x} : \mathbf{x} \in \mathbb{R}^n\}.$$Both are subspaces (Chapter 1.3). The image is the span of the columns of $A$. There is a balance: the bigger the kernel, the smaller the image. The two dimensions add up to $n$. This "rank–nullity" law gets its full treatment in Chapter 1.8.
Why do we need it?
To understand what a matrix can and cannot do, we ask: which inputs does it ignore, and which outputs can it reach? The kernel and the image answer exactly that.
Where is it used?
Rank and low-rank weights (LoRA), redundant (copied) features in regression, deciding whether a linear system has a solution and whether it is unique, and PCA and the SVD.
How is it used?
Find the kernel by solving Ax = 0 (scipy.linalg.null_space). Find the image as the span of the columns (scipy.linalg.orth). Use np.linalg.matrix_rank to count how many independent directions survive.
The kernel lives in the input space and the image lives in the output space. For a non-square matrix these are different spaces ($\mathbb{R}^n$ and $\mathbb{R}^m$).
Quick check: what is the kernel of $\begin{bmatrix} 1 & -1 \\ -1 & 1 \end{bmatrix}$?
$A\mathbf{x} = \mathbf{0}$ means $x_1 - x_2 = 0$, so $x_1 = x_2$. The kernel is the line through $[1, 1]$. (The image is the line through $[1, -1]$.)
Why a layer needs a non-linearity core
One layer of a neural network does two things: an affine map $W\mathbf{x} + \mathbf{b}$ (rotate, stretch, shear, shift) and then a non-linear function such as ReLU, $\max(0, z)$, applied to each number.
What if we skip the second step? Then a stack of layers is a composition of affine maps, and a composition of affine maps is one affine map. A hundred layers would be no more powerful than a single layer. They would just multiply out into one matrix.
The non-linearity is what bends space. ReLU folds the plane along the axes (everything negative is clipped to 0). Many bends in a row let the network carve out curved, complicated shapes that no single matrix can produce.
Let $W_1 = \begin{bmatrix} 1 & 2 \\ 0 & 1 \end{bmatrix}$, $\mathbf{b}_1 = [1, 0]$, $W_2 = \begin{bmatrix} 1 & 0 \\ 1 & 1 \end{bmatrix}$, $\mathbf{b}_2 = [0, 1]$, and the input $\mathbf{x} = [1, 1]$.
No non-linearity.
- Layer 1: $W_1\mathbf{x} + \mathbf{b}_1 = [3, 1] + [1, 0] = [4, 1]$.
- Layer 2: $W_2[4, 1] + \mathbf{b}_2 = [4, 5] + [0, 1] = [4, 6]$.
- One-step version: $W_2W_1 = \begin{bmatrix} 1 & 2 \\ 1 & 3 \end{bmatrix}$ and $W_2\mathbf{b}_1 + \mathbf{b}_2 = [1, 1] + [0, 1] = [1, 2]$. So the single layer gives $\begin{bmatrix} 1 & 2 \\ 1 & 3 \end{bmatrix}[1, 1] + [1, 2] = [3, 4] + [1, 2] = [4, 6]$ ✓. The same.
With ReLU in between, try $\mathbf{x} = [-3, 0]$. Layer 1 gives $[-3, 0] + [1, 0] = [-2, 0]$, and ReLU clips it to $[0, 0]$. Layer 2 gives $[0, 0] + [0, 1] = [0, 1]$. But the single-layer shortcut gives $\begin{bmatrix} 1 & 2 \\ 1 & 3 \end{bmatrix}[-3, 0] + [1, 2] = [-3, -3] + [1, 2] = [-2, -1]$. Different. The ReLU made the network genuinely deeper than one layer.
Two stacked affine layers collapse into one:
$$W_2\,(W_1\mathbf{x} + \mathbf{b}_1) + \mathbf{b}_2 = \underbrace{(W_2W_1)}_{\text{one matrix}}\mathbf{x} + \underbrace{(W_2\mathbf{b}_1 + \mathbf{b}_2)}_{\text{one bias}}.$$By repeating this, $L$ affine layers become one. A standard layer is therefore $\mathbf{h} = \sigma(W\mathbf{x} + \mathbf{b})$, with a non-linear activation $\sigma$ (ReLU, sigmoid, tanh, GELU, …) applied entry by entry. Because $\sigma$ is not linear, $\sigma(W_2\sigma(W_1\mathbf{x}))$ cannot, in general, be rewritten as one matrix times $\mathbf{x}$.
Why do we need it?
Stacking matrix layers alone gains nothing, because they collapse into one matrix. We need to know why a non-linear step is essential to give a network real depth.
Where is it used?
Every multilayer perceptron (ReLU, GELU or tanh between layers), the feed-forward blocks of Transformers, and the explanation of why a deep network without activations is just linear regression.
How is it used?
Write each layer as sigma(Wx + b). To audit a model, look for two Linear layers with no activation between them: they merge into one (W2 W1). Pick an activation such as ReLU so the network can bend space.
Making a network deeper only adds power if there is a non-linearity between the layers. Also, the non-linearity is applied to each entry separately; it is not a matrix. That is why it can do something a matrix cannot.
Quick check: a network has 50 layers, each $\mathbf{x} \mapsto W\mathbf{x}$ with no bias and no activation. How many matrices describe the whole network?
One: the product $W_{50}\cdots W_2W_1$. The 50 linear layers collapse into a single linear layer (with one large matrix, whatever its depth).
Recap, cheat sheet and practice
- A linear transformation keeps the origin fixed and keeps linear combinations: $T(a\mathbf{x} + b\mathbf{y}) = aT(\mathbf{x}) + bT(\mathbf{y})$. Grid lines stay straight, parallel and evenly spaced.
- Every linear map is a matrix, and the columns are where the basis vectors land: $A = [T(\mathbf{e}_1)\; T(\mathbf{e}_2)\;\cdots]$.
- Standard 2D maps: scaling (diagonal), rotation $\begin{bmatrix}\cos\theta & -\sin\theta\\ \sin\theta & \cos\theta\end{bmatrix}$, reflection (flips orientation), shear (keeps area), projection $\mathbf{u}\mathbf{u}^\top$ (flattens, no inverse).
- Composition is matrix multiplication, right to left: "first $A$, then $B$" is $BA$. Order matters. The inverse undoes a map and exists only if space is not squashed.
- Affine $A\mathbf{x} + \mathbf{b}$ = linear + shift. It is linear only when $\mathbf{b} = \mathbf{0}$ (otherwise the origin moves). Homogeneous coordinates turn it into one $3\times3$ matrix.
- Similar matrices $B = P^{-1}AP$ are the same transformation in a different basis. The kernel is what is sent to $\mathbf{0}$; the image is where outputs can land.
- A neural-network layer is affine + non-linearity. Without the non-linearity, stacked layers collapse into one matrix.
Cheat sheet
| Transformation | Matrix (2D) | Key fact |
|---|---|---|
| Scaling | $\begin{bmatrix} s_x & 0 \\ 0 & s_y \end{bmatrix}$ | area $\times |s_xs_y|$ |
| Rotation by $\theta$ | $\begin{bmatrix} \cos\theta & -\sin\theta \\ \sin\theta & \cos\theta \end{bmatrix}$ | orthogonal; inverse is $R(-\theta) = R^\top$ |
| Reflection over line at $\alpha$ | $\begin{bmatrix} \cos2\alpha & \sin2\alpha \\ \sin2\alpha & -\cos2\alpha \end{bmatrix}$ | $F^2 = I$; flips orientation |
| Horizontal shear | $\begin{bmatrix} 1 & k \\ 0 & 1 \end{bmatrix}$ | keeps area; inverse has $-k$ |
| Projection on unit $\mathbf{u}$ | $\mathbf{u}\mathbf{u}^\top$ | $P^2 = P$; not invertible |
| Composition (A then B) | $BA$ | right to left; $AB \ne BA$ |
| Inverse | $A^{-1}A = I$ | $(AB)^{-1} = B^{-1}A^{-1}$ |
| Affine map | $\begin{bmatrix} A & \mathbf{b} \\ \mathbf{0}^\top & 1 \end{bmatrix}$ on $[\mathbf{x}; 1]$ | origin goes to $\mathbf{b}$ |
| Same map, new basis | $B = P^{-1}AP$ | $\mathbf{c} = P^{-1}\mathbf{x}$ |
import numpy as np
import matplotlib.pyplot as plt
square = np.array([[0, 1, 1, 0, 0], [0, 0, 1, 1, 0]]) # corners as columns (closed loop)
F = np.array([[0, .25, .25, .6, .6, .25, .25, .75, .75, 0, 0],
[0, 0, .4, .4, .6, .6, .85, .85, 1.05, 1.05, 0]]) # the letter F (closed loop)
def rot(deg):
t = np.radians(deg)
return np.array([[np.cos(t), -np.sin(t)], [np.sin(t), np.cos(t)]])
u = np.array([1, 1]) / np.sqrt(2)
maps = {"rotation 45": rot(45), "scaling": np.diag([2, 0.5]),
"reflection y=x": np.array([[0, 1], [1, 0]]),
"shear": np.array([[1, 1], [0, 1]]), "projection": np.outer(u, u)}
# apply each matrix to the unit square and to the F, then plot before / after
fig, axes = plt.subplots(1, 5, figsize=(17, 3.5))
for ax, (name, A) in zip(axes, maps.items()):
ax.plot(*square, "k--"); ax.plot(*F, "k--") # before (dashed)
ax.plot(*(A @ square), "C1"); ax.plot(*(A @ F), "C1") # after (orange)
ax.set_title(name); ax.set_aspect("equal"); ax.grid(True)
plt.show()
# the columns of A are the images of the basis vectors
A = np.array([[2, -1], [1, 1]])
print(A @ np.array([1, 0]), A @ np.array([0, 1])) # [2 1] [-1 1]
# composition = product (right to left), and order matters
R, S = rot(90), np.diag([2, 1])
x = np.array([1.0, 0.0])
print(S @ (R @ x), (S @ R) @ x) # the same vector twice, about [0 1] (the 1e-16 is rounding)
print(np.allclose(S @ R, R @ S)) # False
# inverse
M = np.array([[2, 1], [1, 1]])
print(np.linalg.inv(M), np.allclose(np.linalg.inv(M) @ M, np.eye(2)))
# affine map as one 3x3 matrix (homogeneous coordinates)
b = np.array([2, 1])
H = np.block([[R, b[:, None]], [np.zeros((1, 2)), np.ones((1, 1))]])
print(H @ np.array([1, 0, 1])) # about [2 2 1]: R @ x + b
# change of basis and similar matrices
A2 = np.array([[3, 1], [1, 3]])
P = np.array([[1, 1], [1, -1]])
B = np.linalg.inv(P) @ A2 @ P
print(np.round(B, 10)) # [[4 0] [0 2]]
print(np.trace(A2), np.trace(B), np.linalg.det(A2), np.linalg.det(B)) # 6 6.0 and 8.0 8.0 (up to rounding)
# kernel and image
from scipy.linalg import null_space, orth
K = np.array([[1, 2], [2, 4]])
print(np.linalg.matrix_rank(K)) # 1
print(null_space(K).ravel(), orth(K).ravel()) # kernel along [2 -1], image along [1 2] (signs may differ)
# two stacked layers with no activation equal ONE layer
rng = np.random.default_rng(0)
W1, W2 = rng.normal(size=(4, 3)), rng.normal(size=(2, 4))
b1, b2 = rng.normal(size=4), rng.normal(size=2)
X = rng.normal(size=(5, 3)) # 5 samples as rows
two = (X @ W1.T + b1) @ W2.T + b2
one = X @ (W2 @ W1).T + (W2 @ b1 + b2)
print(np.allclose(two, one)) # True
relu = lambda z: np.maximum(0, z)
print(np.allclose(relu(X @ W1.T + b1) @ W2.T + b2, one)) # False: ReLU makes it genuinely deeper
1. The matrix $\begin{bmatrix} 0 & -1 \\ 1 & 0 \end{bmatrix}$ acts on the arrow $[1, 0]$. Where does the arrow land?
2. Which of these transformations is not linear?
3. You apply transformation $A$ first and then transformation $B$. The matrix of the combined transformation is…
4. $P$ is a projection matrix (projecting onto a line). What is $P^2$?
5. The affine map $T(\mathbf{x}) = A\mathbf{x} + \mathbf{b}$ is a linear transformation exactly when…
6. Two layers $\mathbf{x} \mapsto W_1\mathbf{x}$ and then $\mathbf{h} \mapsto W_2\mathbf{h}$, with no activation in between, equal…
Practice problems
A. Write the matrix of the rotation by $180^\circ$. Where does $[3, -2]$ go?
$\cos180^\circ = -1$ and $\sin180^\circ = 0$, so $R = \begin{bmatrix} -1 & 0 \\ 0 & -1 \end{bmatrix} = -I$. It sends $[3, -2]$ to $[-3, 2]$ (every arrow turns around).
B. First reflect across the line $y = x$, then stretch $x$ by 2. Find the single matrix, and check it on $\mathbf{e}_1$.
Reflection $F = \begin{bmatrix}0&1\\1&0\end{bmatrix}$, stretch $S = \begin{bmatrix}2&0\\0&1\end{bmatrix}$. First $F$, then $S$: $SF = \begin{bmatrix}2&0\\0&1\end{bmatrix}\begin{bmatrix}0&1\\1&0\end{bmatrix} = \begin{bmatrix}0&2\\1&0\end{bmatrix}$. Check: $\mathbf{e}_1 \to [0, 1]$ (swap) $\to [0, 1]$ (stretch $x$ of $0$); first column of $SF$ is $[0, 1]$ ✓. (The other order, $FS = \begin{bmatrix}0&1\\2&0\end{bmatrix}$, is different.)
C. Show that the inverse of the shear $\begin{bmatrix}1&2\\0&1\end{bmatrix}$ is $\begin{bmatrix}1&-2\\0&1\end{bmatrix}$.
Multiply: $\begin{bmatrix}1&2\\0&1\end{bmatrix}\begin{bmatrix}1&-2\\0&1\end{bmatrix} = \begin{bmatrix}1\cdot1 + 2\cdot0 & 1\cdot(-2) + 2\cdot1 \\ 0 & 1\end{bmatrix} = \begin{bmatrix}1&0\\0&1\end{bmatrix}$ ✓. Shearing right by 2 per unit of height, then left by 2, puts everything back.
D. Let $T(\mathbf{x}) = A\mathbf{x} + \mathbf{b}$ with $A = \begin{bmatrix}1&0\\0&-1\end{bmatrix}$ and $\mathbf{b} = [0, 3]$. Find $T([2, 1])$, $T(\mathbf{0})$, and the homogeneous matrix.
$T([2, 1]) = [2, -1] + [0, 3] = [2, 2]$. $T(\mathbf{0}) = [0, 3] \ne \mathbf{0}$, so $T$ is not linear. The homogeneous matrix is $\begin{bmatrix}1&0&0\\0&-1&3\\0&0&1\end{bmatrix}$; check: it sends $[2, 1, 1]$ to $[2,\; -1 + 3,\; 1] = [2, 2, 1]$ ✓.
E. Find the kernel and the image of $\begin{bmatrix}2&4\\1&2\end{bmatrix}$.
Kernel: $2x + 4y = 0$ (and $x + 2y = 0$, the same condition), so $x = -2y$. The kernel is the line through $[-2, 1]$ (equivalently $[2, -1]$). Image: the columns are $[2, 1]$ and $[4, 2] = 2[2, 1]$, so every output is a multiple of $[2, 1]$: the image is the line through $[2, 1]$.
F. Let $A = \begin{bmatrix}2&0\\0&3\end{bmatrix}$ and $P = \begin{bmatrix}0&1\\1&0\end{bmatrix}$ (the basis is $\mathbf{e}_2, \mathbf{e}_1$). Compute $B = P^{-1}AP$ and explain it.
$P^{-1} = P$. $AP = \begin{bmatrix}0&2\\3&0\end{bmatrix}$, then $P(AP) = \begin{bmatrix}3&0\\0&2\end{bmatrix}$. So $B = \operatorname{diag}(3, 2)$. We listed the basis arrows in the opposite order, so the stretch factors swap places. It is the same transformation, with its numbers written in a different order. Trace ($5$) and determinant ($6$) are unchanged.
Systems of Linear Equations
Many questions in maths and machine learning have the same shape: "find the numbers that make several conditions true at once." This chapter shows how to write such a puzzle as $A\mathbf{x}=\mathbf{b}$, how to see it in two different pictures, and how to solve it step by step, or to prove that it cannot be solved.
- Write a system of equations as $A\mathbf{x}=\mathbf{b}$ and as an augmented matrix $[A\,|\,\mathbf{b}]$
- See the same system as crossing lines (row picture) and as a mix of columns (column picture)
- Solve it by Gaussian elimination, back substitution and Gauss–Jordan elimination (REF and RREF)
- Read pivots and free variables, and decide: one solution, infinitely many, or none
- Understand tall (overdetermined) and wide (underdetermined) systems, and why ML cares
- Write every solution as one particular solution plus the null space
Linear equations core
You are at a market. Apples cost 2 dollars each and bananas cost 1 dollar each. You have exactly 5 dollars and want to spend all of it. How many apples and bananas can you buy?
Call the number of apples $x$ and the number of bananas $y$. Your spending is "$2$ dollars per apple times $x$, plus $1$ dollar per banana times $y$". Spending all 5 dollars gives one sentence in maths:
$2x + y = 5$
That is a linear equation. It has several answers: 2 apples and 1 banana, or 1 apple and 3 bananas, or no apples and 5 bananas. If you draw every answer as a point $(x, y)$, all the points sit on one straight line. That is where the word "linear" comes from.
Check the equation $2x + y = 5$ for a few points. Put the numbers in and see if the left side equals 5.
- $(x, y) = (2, 1)$: $2\cdot 2 + 1 = 4 + 1 = 5$. Yes, it is a solution.
- $(x, y) = (1, 3)$: $2\cdot 1 + 3 = 2 + 3 = 5$. Yes.
- $(x, y) = (3, 1)$: $2\cdot 3 + 1 = 7$. Not 5, so this point is not a solution.
A linear equation in the unknowns $x_1, x_2, \dots, x_n$ has the form
$$a_1x_1 + a_2x_2 + \dots + a_nx_n = b,$$where $a_1, \dots, a_n$ are known numbers (the coefficients) and $b$ is a known number (the right-hand side). A solution is a list of values for the unknowns that makes the equation true.
"Linear" means each unknown is only multiplied by a number and then everything is added. No squares ($x^2$), no unknown times unknown ($xy$), no $\sin x$, no $\sqrt{x}$, no $1/x$.
Why do we need it?
Many questions ask for unknown numbers that must obey a rule: a budget, a balance, a prediction. A linear equation writes that rule exactly, so we can search for the numbers instead of guessing.
Where is it used?
Linear regression (every training example gives one equation in the unknown weights), a neuron's sum $w_1x_1 + \dots + b$, balancing chemical equations, circuit laws, and budget or mixing problems.
How is it used?
Name the unknowns. Write each condition as (number)·(unknown) + … = (number). Test any candidate answer by putting it in. In machine learning the unknowns are the weights and the known numbers are your data.
Which of these are linear? $3x - 2y = 7$ is linear. $xy = 4$ is not (two unknowns multiplied). $x^2 + y = 1$ is not (a square). $\sin x + y = 0$ is not. Also, $x + 2y + 0z = 4$ is linear: a missing unknown just has coefficient 0.
Quick check: is $5x - y = 0$ linear? Is $(1, 5)$ a solution?
Yes, it is linear (the right-hand side may be 0). Put in $(1, 5)$: $5\cdot 1 - 5 = 0$. Yes, it is a solution.
Systems of equations and the matrix form $A\mathbf{x}=\mathbf{b}$ core
Real puzzles come with several conditions at once. Add a second rule to the market story: apples and bananas together must be 3 pieces of fruit. Now you need numbers that make both sentences true. A group of linear equations that must all hold together is a system of linear equations.
Writing the system out is slow. So we use a grid. Each row of the grid is one equation. Each column belongs to one unknown. We put the unknowns in a vector $\mathbf{x}$ and the right-hand sides in a vector $\mathbf{b}$. The whole system then shrinks to one short line:
$A\mathbf{x} = \mathbf{b}$
In words: "the matrix $A$ turns the unknown vector $\mathbf{x}$ into $\mathbf{b}$. Which $\mathbf{x}$ does that?"
The system
$$\begin{aligned} 2x + y &= 5 \\ x - y &= 1 \end{aligned} \qquad\text{becomes}\qquad \underbrace{\begin{bmatrix} 2 & 1 \\ 1 & -1 \end{bmatrix}}_{A} \underbrace{\begin{bmatrix} x \\ y \end{bmatrix}}_{\mathbf{x}} = \underbrace{\begin{bmatrix} 5 \\ 1 \end{bmatrix}}_{\mathbf{b}}$$Check that $\mathbf{x} = [2, 1]$ works. Multiply row by row (each row of $A$ dot $\mathbf{x}$, as in the dot product):
- Row 1: $2\cdot 2 + 1\cdot 1 = 5$. It matches the first entry of $\mathbf{b}$.
- Row 2: $1\cdot 2 + (-1)\cdot 1 = 1$. It matches the second entry of $\mathbf{b}$.
So $\mathbf{x} = [2, 1]$ is a solution.
A system of $m$ linear equations in $n$ unknowns is written
$$A\mathbf{x} = \mathbf{b}, \qquad A \in \mathbb{R}^{m\times n},\; \mathbf{x} \in \mathbb{R}^{n},\; \mathbf{b} \in \mathbb{R}^{m}.$$- $A$ is the coefficient matrix. Entry $a_{ij}$ is the coefficient of unknown $j$ in equation $i$.
- $\mathbf{x}$ is the vector of unknowns. $\mathbf{b}$ is the right-hand side.
- Row $i$ of the product $A\mathbf{x}$ is (row $i$ of $A$) $\cdot\,\mathbf{x}$, and it must equal $b_i$.
- A solution is any $\mathbf{x}$ with $A\mathbf{x} = \mathbf{b}$. The solution set is the collection of all of them. It may be empty, have one member, or have infinitely many.
Why do we need it?
Writing out ten equations in ten unknowns is slow and hides the pattern. One short line, $A\mathbf{x}=\mathbf{b}$, says the same thing and lets us use matrix tools and computer libraries.
Where is it used?
Linear regression $X\mathbf{w}=\mathbf{y}$ (rows are examples, columns are features), node voltages in electric circuits, finite-element simulations, least-squares fitting, and the huge linear system behind Google's PageRank.
How is it used?
Stack the coefficients into $A$ (one row per equation, one column per unknown) and the right-hand sides into $\mathbf{b}$. Check the shapes. In code, np.linalg.solve(A, b) returns $\mathbf{x}$.
Shapes. $A$ is $m\times n$: $m$ rows (equations) and $n$ columns (unknowns). So $\mathbf{x}$ has $n$ entries and $\mathbf{b}$ has $m$ entries. The number of unknowns is the number of columns, not rows. Mixing this up is the most common beginner slip.
Quick check: a system has 4 equations and 3 unknowns. What is the shape of $A$, $\mathbf{x}$ and $\mathbf{b}$?
$A$ is $4\times 3$, $\mathbf{x}$ has 3 entries and $\mathbf{b}$ has 4 entries.
Two pictures of one system: rows and columns core
A system can be seen in two different ways. Both give the same answer.
- Row picture (roads). Look at one equation at a time. In two unknowns, each equation is a line. In three unknowns, it is a plane. A solution must lie on every line at once, so it is where the roads cross.
- Column picture (recipe). Look at the columns of $A$ as ingredient arrows. The unknowns are how many scoops of each. The question becomes: "can I mix the columns to build the target arrow $\mathbf{b}$?" This is the linear combination idea from Chapter 1.2.
Take the same system: $2x + y = 5$ and $x - y = 1$.
Row picture. The first line passes through $(0, 5)$ and $(2.5, 0)$. The second passes through $(0, -1)$ and $(1, 0)$. They cross at $(2, 1)$.
Column picture. The columns are $\mathbf{a}_1 = [2, 1]$ and $\mathbf{a}_2 = [1, -1]$. The target is $\mathbf{b} = [5, 1]$. Two scoops of $\mathbf{a}_1$ and one scoop of $\mathbf{a}_2$:
$$2\begin{bmatrix} 2 \\ 1 \end{bmatrix} + 1\begin{bmatrix} 1 \\ -1 \end{bmatrix} = \begin{bmatrix} 4 \\ 2 \end{bmatrix} + \begin{bmatrix} 1 \\ -1 \end{bmatrix} = \begin{bmatrix} 5 \\ 1 \end{bmatrix} = \mathbf{b}.$$The scoop counts $(2, 1)$ are the same numbers as the crossing point. Same answer, two views.
Row picture. Equation $i$ is $\mathbf{r}_i \cdot \mathbf{x} = b_i$, where $\mathbf{r}_i$ is row $i$ of $A$. Its solutions form a line (in $\mathbb{R}^2$), a plane (in $\mathbb{R}^3$), or a "hyperplane" in higher dimensions. The solution set of the system is the intersection of all of them.
Column picture. If $\mathbf{a}_1, \dots, \mathbf{a}_n$ are the columns of $A$, then
$$A\mathbf{x} = x_1\mathbf{a}_1 + x_2\mathbf{a}_2 + \dots + x_n\mathbf{a}_n.$$So $A\mathbf{x} = \mathbf{b}$ asks: is $\mathbf{b}$ a linear combination of the columns, and with which weights $x_1, \dots, x_n$?
Why do we need it?
Two views of one problem make it easier to understand. The row picture tells you whether the conditions agree (do the lines meet?). The column picture tells you whether the target can be built from the ingredients.
Where is it used?
Row picture: constraints on weights and feasibility questions. Column picture: regression ("can the target be built from the feature columns?"), least squares as a projection, PCA, and the idea of a column space.
How is it used?
When stuck, switch picture. For a small system, sketch the lines or planes and ask where they meet, or draw the columns and ask which mix reaches $\mathbf{b}$. The scoop counts are the solution $\mathbf{x}$.
The two planes live in different spaces. In the row picture, the axes are the unknowns $(x, y)$ and a point is a candidate answer. In the column picture, the axes are the outputs $(b_1, b_2)$ and a point is a possible right-hand side. Do not mix their coordinates up.
Quick check: in the column picture, what do the unknowns $x$ and $y$ mean?
They are the scoop counts: how much of column 1 and how much of column 2 to add up to reach $\mathbf{b}$.
The augmented matrix $[A\,|\,\mathbf{b}]$ core
When you solve a system by hand, you keep rewriting $x$, $y$, "$+$" and "$=$". That is a lot of ink for no new information. The augmented matrix keeps only the numbers, like a tidy spreadsheet of the system. A thin vertical bar stands in for the "$=$" signs.
Take the system
$$\begin{aligned} x + 2y - z &= 3 \\ 2x - y + z &= 0 \\ y + 3z &= 5 \end{aligned}$$Notice that the third equation has no $x$. We write its $x$-coefficient as $0$, so the columns stay lined up:
$$[A\,|\,\mathbf{b}] = \left[\begin{array}{ccc|c} 1 & 2 & -1 & 3 \\ 2 & -1 & 1 & 0 \\ 0 & 1 & 3 & 5 \end{array}\right]$$Row 2 reads "2 of $x$, minus 1 of $y$, plus 1 of $z$, equals 0". Nothing is lost.
For $A\mathbf{x}=\mathbf{b}$, the augmented matrix is $A$ with $\mathbf{b}$ attached as one extra last column:
$$[A\,|\,\mathbf{b}] \in \mathbb{R}^{m\times(n+1)}.$$Row $i$ is equation $i$. Column $j$ (for $j \le n$) belongs to unknown $x_j$. The last column is the right-hand side. The bar is only a reminder; the numbers are all that matter. Doing an operation on a row of $[A\,|\,\mathbf{b}]$ is the same as doing it on the whole equation, both sides at once.
Why do we need it?
While solving by hand you rewrite the same system many times. Dropping the letters and plus signs keeps only the numbers, so each step is quick and easy to check.
Where is it used?
Every hand elimination, every textbook exercise, and the inside of software: LU decomposition and sympy's rref work on exactly this array of numbers.
How is it used?
Write each equation as a row of coefficients (a 0 for a missing unknown, always the same column order), then attach $\mathbf{b}$ as the last column. Do row operations on whole rows. In code: np.column_stack([A, b]).
Keep every equation in the same order of unknowns ($x$, then $y$, then $z$). Write $0$ for a missing unknown. And do not forget the last column: it is the right-hand side, not another unknown.
Quick check: write $3x - y = 4$ and $x + 2y = -1$ as an augmented matrix.
$\left[\begin{array}{cc|c} 3 & -1 & 4 \\ 1 & 2 & -1 \end{array}\right]$
Elementary row operations core
To solve a system, we rewrite it into a simpler system with the same answer. Think of simplifying a fraction: $\frac{6}{8}$ and $\frac{3}{4}$ look different but are the same number. We need rewriting moves that never change the solutions. There are exactly three:
- Swap two equations. (The order of rules does not matter.)
- Scale an equation: multiply both sides by a number that is not zero.
- Add a multiple of one equation to another. (If $A = B$ and $C = D$ are both true, then $A + 2C = B + 2D$ is true too.)
Each move can be undone by another move, so nothing is ever lost. That is why the solutions stay the same. (Multiplying by $0$ is forbidden because it cannot be undone: it throws the equation away.)
Start with $[A\,|\,\mathbf{b}] = \left[\begin{array}{cc|c} 2 & 1 & 5 \\ 1 & -1 & 1 \end{array}\right]$, which is $2x + y = 5$ and $x - y = 1$.
- Swap $R_1 \leftrightarrow R_2$: $\left[\begin{array}{cc|c} 1 & -1 & 1 \\ 2 & 1 & 5 \end{array}\right]$. Same two equations, other order.
- Scale $R_1 \leftarrow 3R_1$ (from the start): $\left[\begin{array}{cc|c} 6 & 3 & 15 \\ 1 & -1 & 1 \end{array}\right]$. The first equation $6x + 3y = 15$ is just $2x + y = 5$ again, multiplied by 3.
- Add a multiple $R_2 \leftarrow R_2 - 2R_1$ (on the swapped matrix): the second row becomes $[2 - 2,\; 1 - (-2),\; 5 - 2] = [0, 3, 3]$. That says $3y = 3$, so $y = 1$. The $x$ has vanished from that row, which is exactly the point of this move.
All three versions still have the solution $(x, y) = (2, 1)$.
The three elementary row operations on a matrix are:
- Swap: $R_i \leftrightarrow R_j$.
- Scale: $R_i \leftarrow c\,R_i$, with $c \neq 0$.
- Replace (add a multiple): $R_i \leftarrow R_i + c\,R_j$, with $i \neq j$. Only row $i$ changes.
Two matrices linked by a chain of these operations are row equivalent. Row-equivalent augmented matrices describe systems with exactly the same solution set.
Why do we need it?
To solve a system we must simplify it, but any change that alters the answer is useless. The three row operations are exactly the safe moves: they never change the solutions.
Where is it used?
Inside every linear solver: LU decomposition (np.linalg.solve), computing rank and inverses, and finding null spaces. Also used by hand in exams and in symbolic tools such as sympy.
How is it used?
Swap rows to bring a non-zero number to the pivot place. Add a multiple of one row to another to create zeros. Scale a row to tidy it. Never multiply by 0. After each move the solutions are the same.
Never multiply a row by 0. It turns an equation into $0 = 0$, and the information is gone for ever. Also, in "add a multiple" only the target row changes. The helper row $R_j$ stays exactly as it was.
Quick check: after $R_2 \leftarrow R_2 - 3R_1$ on $\left[\begin{array}{cc|c} 1 & 2 & 4 \\ 3 & 4 & 10 \end{array}\right]$, what is row 2?
$[3 - 3,\; 4 - 6,\; 10 - 12] = [0, -2, -2]$.
Gaussian elimination, row echelon form and back substitution
Picture a row of dominoes where each one only depends on the ones after it. If the last equation has just one unknown, you can solve it at once. Put that answer into the equation above (which now also has just one new unknown), and so on, climbing up. This is back substitution.
So the plan is: reshape the system into a staircase where each equation has fewer unknowns than the one above. We do this column by column, left to right:
- Pick a pivot: a non-zero number in the current column (if the spot holds a 0, swap in a row that does not).
- Use the pivot row to wipe out every number below the pivot (add a multiple of the pivot row to each lower row).
- Move one row down and one column right, and repeat.
The result is the staircase shape. This method is called Gaussian elimination (forward elimination).
Solve
$$\begin{aligned} x + y + z &= 6 \\ 2x + y - z &= 1 \\ x - y + 2z &= 5 \end{aligned} \qquad [A\,|\,\mathbf{b}] = \left[\begin{array}{ccc|c} 1 & 1 & 1 & 6 \\ 2 & 1 & -1 & 1 \\ 1 & -1 & 2 & 5 \end{array}\right]$$Forward elimination. The first pivot is the $1$ in row 1, column 1.
- $R_2 \leftarrow R_2 - 2R_1$: $[2-2,\; 1-2,\; -1-2,\; 1-12] = [0, -1, -3, -11]$.
- $R_3 \leftarrow R_3 - R_1$: $[1-1,\; -1-1,\; 2-1,\; 5-6] = [0, -2, 1, -1]$.
- The next pivot is the $-1$ in row 2, column 2. $R_3 \leftarrow R_3 - 2R_2$: $[0,\; -2+2,\; 1+6,\; -1+22] = [0, 0, 7, 21]$.
Back substitution. Start from the bottom row.
- Row 3: $7z = 21$, so $z = 3$.
- Row 2: $-y - 3z = -11$. Put in $z = 3$: $-y - 9 = -11$, so $-y = -2$ and $y = 2$.
- Row 1: $x + y + z = 6$. Put in $y = 2$, $z = 3$: $x + 5 = 6$, so $x = 1$.
The solution is $(x, y, z) = (1, 2, 3)$. Check in the second original equation: $2\cdot1 + 2 - 3 = 1$ ✓.
A pivot is the first non-zero entry of a row (reading from the left), once the matrix has been tidied. A matrix is in row echelon form (REF) when:
- Every all-zero row is at the bottom.
- Each row's pivot is strictly to the right of the pivot in the row above (a staircase).
- Every entry below a pivot is 0.
Gaussian elimination uses row swaps and "add a multiple" operations to reach REF. Back substitution then solves the system from the last non-zero row upward. (Some books also demand that every pivot equals 1 in REF. We do not; that extra tidying belongs to the reduced form, next.)
For an $n\times n$ system, Gaussian elimination needs roughly $n^3/3$ multiplications. That is cheap enough for thousands of unknowns.
Why do we need it?
A tangled system has no easy answer, but a staircase-shaped system can be solved from the bottom up in a few lines. Elimination is the method that turns the first into the second.
Where is it used?
np.linalg.solve (LU with row swaps), the normal equations of linear regression, finite-element solvers, the Newton step in optimisation, and computing determinants and ranks.
How is it used?
Go column by column: pick a pivot (swap rows if it is 0), subtract multiples of its row from the rows below, then move down and right. Finish with back substitution from the last row, and check the answer in the original equations.
- A zero in the pivot place is not a disaster. Swap with a lower row that has a non-zero entry in that column. Never divide by zero.
- Check your answer by putting it back into the original equations. One arithmetic slip early on spoils everything after it.
- On a computer, a pivot that is non-zero but tiny causes big rounding errors. Real software picks the biggest available entry as the pivot ("partial pivoting", Chapter 1.15).
Quick check: why is $\left[\begin{array}{cc|c} 2 & 1 & 5 \\ 0 & 3 & 3 \end{array}\right]$ already in REF, and what is the solution?
Its pivots are 2 (row 1, column 1) and 3 (row 2, column 2): a staircase with a 0 below the first pivot. Back substitution: $3y = 3$ gives $y = 1$; then $2x + 1 = 5$ gives $x = 2$.
Gauss–Jordan elimination and the reduced row echelon form
Gaussian elimination stops at a staircase and then makes you climb back up by substitution. Gauss–Jordan elimination keeps cleaning instead. It makes each pivot equal to 1, and it wipes out the numbers above each pivot too. When it is finished, nothing is left to substitute: the answer can be read off the last column.
It is more work than back substitution, but the final form is tidy and unique, and it is the form we use to read pivots, free variables and null spaces. (We also use it to find inverses, in Chapter 1.7.)
Continue from the REF of the previous example. Work from the last pivot upward.
$$\left[\begin{array}{ccc|c} 1 & 1 & 1 & 6 \\ 0 & -1 & -3 & -11 \\ 0 & 0 & 7 & 21 \end{array}\right]$$- $R_3 \leftarrow R_3 \div 7$: row 3 becomes $[0, 0, 1, 3]$.
- Clear above that pivot. $R_2 \leftarrow R_2 + 3R_3$: $[0, -1, 0, -11 + 9] = [0, -1, 0, -2]$. And $R_1 \leftarrow R_1 - R_3$: $[1, 1, 0, 3]$.
- $R_2 \leftarrow R_2 \div (-1)$: row 2 becomes $[0, 1, 0, 2]$.
- Clear above that pivot. $R_1 \leftarrow R_1 - R_2$: $[1, 0, 0, 1]$.
Read it: $x = 1$, $y = 2$, $z = 3$. The left side became the identity matrix, and the right column is the answer.
A matrix is in reduced row echelon form (RREF) when it is in REF and:
- Every pivot is exactly $1$.
- Every pivot is the only non-zero entry in its column.
Gauss–Jordan elimination is Gaussian elimination followed by this extra cleaning. Fact: every matrix has exactly one RREF, however you order the row operations. (A REF is not unique, but the RREF is.) Software such as sympy.Matrix.rref() returns it.
Why do we need it?
Back substitution still takes work. Gauss–Jordan cleaning gives one tidy form from which the pivots, the free variables and the answer can be read directly. It is the same no matter how you got there.
Where is it used?
Finding rank, null space and column space; computing inverses (Gauss–Jordan on $[A\,|\,I]$); sympy's Matrix.rref(); classifying systems. For noisy floating-point data, ML code uses the SVD instead.
How is it used?
After reaching REF, go from the last pivot upward: divide the pivot row to make the pivot 1, then clear the numbers above it. Read the answer from the last column. Remember that the RREF is unique, the REF is not.
The RREF is unique, the REF is not. Two people can reach different staircases (for example, one swaps rows and the other does not) and both are right. But if both keep cleaning to RREF, they must end up with the identical matrix.
Quick check: is $\left[\begin{array}{cc|c} 1 & 2 & 3 \\ 0 & 1 & 4 \end{array}\right]$ in RREF?
No. It is in REF, and the pivots are 1. But the pivot in column 2 has a 2 above it, so rule 5 fails. Do $R_1 \leftarrow R_1 - 2R_2$ to get $\left[\begin{array}{cc|c} 1 & 0 & -5 \\ 0 & 1 & 4 \end{array}\right]$, which is RREF.
Pivots, free variables and solving
After row reduction, look at the columns of the coefficient part.
- A column with a pivot belongs to an unknown that some equation pins down.
- A column without a pivot belongs to an unknown that nothing pins down. It is free: you may choose it to be anything, and the pinned unknowns then adjust to match.
Think of friends splitting a bill under fewer rules than there are people. Some friends can pay whatever they like (free). Once they decide, the rules fix what everybody else pays.
Suppose row reduction gives (columns are $x, y, z$, then the right side):
$$\left[\begin{array}{ccc|c} 1 & 0 & 2 & 3 \\ 0 & 1 & -1 & 1 \\ 0 & 0 & 0 & 0 \end{array}\right]$$The pivots are in columns 1 and 2. Column 3 has no pivot, so $z$ is free. Read the equations:
- Row 1: $x + 2z = 3$, so $x = 3 - 2z$.
- Row 2: $y - z = 1$, so $y = 1 + z$.
- Row 3: $0 = 0$. It says nothing new (a free "spare" row).
Pick $z = 0$ and you get the solution $(3, 1, 0)$. Pick $z = 1$ and you get $(1, 2, 1)$. Check the second one in the equations above: $1 + 2 = 3$ ✓ and $2 - 1 = 1$ ✓. There is one solution for every value of $z$: infinitely many.
- A pivot column is a column that contains a pivot; its unknown is a pivot variable (or basic variable).
- A column with no pivot is a free column; its unknown is a free variable.
- The number of pivots is the rank of $A$ (Chapter 1.7). So number of free variables $= n - \text{rank}(A)$, where $n$ is the number of unknowns.
How to solve any system $A\mathbf{x}=\mathbf{b}$:
- Write $[A\,|\,\mathbf{b}]$ and row reduce it.
- If a row reads $[0\ \cdots\ 0\,|\,\text{non-zero}]$, stop: no solution.
- Otherwise, call each free variable by a free name (such as $z = t$).
- Solve each pivot row for its pivot variable, in terms of the free ones.
Why do we need it?
After row reduction we need a quick way to say whether the answer is one point, a line or nothing, and how many choices are free. Pivots and free variables tell us.
Where is it used?
Rank and null-space computations, spotting redundant (collinear) features in a dataset, over-parameterised models with more weights than data, and describing all solutions of an under-determined system.
How is it used?
Reduce $[A\,|\,\mathbf{b}]$ to RREF. Mark the pivot columns; the other columns are free. Look for a row $0 = $ non-zero. Give each free variable a name (such as $t$), then write each pivot variable in terms of them.
"Free" does not mean "does not matter". It means you are free to choose it, and then the other unknowns depend on your choice. Also, a pivot column is a column of the original matrix $A$: the pivot positions are found during elimination, but the column number is what counts.
Quick check: $[A\,|\,\mathbf{b}]$ for 4 unknowns reduces to rank 3 with no contradiction. How many free variables, and how many solutions?
Free variables $= 4 - 3 = 1$. One free variable means infinitely many solutions (a line of them).
Types of systems: consistent or not, one, none or infinitely many core
A linear system can end in only three ways: no solution, exactly one, or infinitely many. It can never have exactly two or exactly five. Why? Two straight lines cannot cross twice. If they touch in two places, they must be the same line, and then they touch everywhere.
- Lines cross once: one solution.
- Lines are parallel: none.
- Lines lie on top of each other: infinitely many.
A system with at least one solution is called consistent. A system with no solution is inconsistent (its equations contradict each other).
- $x + y = 2$ and $x - y = 0$. Adding them gives $2x = 2$, so $x = 1$, $y = 1$. One solution.
- $x + y = 2$ and $x + y = 3$. The same left side cannot equal both 2 and 3. No solution (inconsistent).
- $x + y = 2$ and $2x + 2y = 4$. The second is the first doubled. Infinitely many solutions: every point on the line.
A square system has as many equations as unknowns ($m = n$). It is the "fair" case, and usually it has one solution. But as example 2 shows (it is square), it can still fail.
Let $r$ be the number of pivots of $A$ (its rank), and $n$ the number of unknowns. Row reduce $[A\,|\,\mathbf{b}]$ and compare:
| What you see | Verdict |
|---|---|
| A pivot in the last column (row $0 = $ non-zero), i.e. $\operatorname{rank}[A\,|\,\mathbf{b}] \gt \operatorname{rank}A$ | No solution (inconsistent) |
| Consistent, and $r = n$ (every unknown has a pivot) | Exactly one solution |
| Consistent, and $r \lt n$ | Infinitely many, with $n - r$ free variables |
Why never exactly two? If $\mathbf{x}_1 \ne \mathbf{x}_2$ both solve $A\mathbf{x}=\mathbf{b}$, then so does every point $\mathbf{x}_1 + t(\mathbf{x}_2 - \mathbf{x}_1)$ on the line through them, because $A\mathbf{x}_1 + t(A\mathbf{x}_2 - A\mathbf{x}_1) = \mathbf{b} + t(\mathbf{b} - \mathbf{b}) = \mathbf{b}$.
Why do we need it?
Before spending time solving, you want to know whether an answer exists and whether it is unique. A linear system has no solution, one, or infinitely many, and ranks tell you which.
Where is it used?
Checking whether a linear model can fit data exactly, finding contradictory constraints in optimisation, asking whether regression weights are uniquely determined, and the "singular matrix" errors that numerical libraries raise.
How is it used?
Row reduce $[A\,|\,\mathbf{b}]$. A row $0 = $ non-zero means none. Otherwise compare rank with the number of unknowns $n$: equal means one solution, smaller means infinitely many ($n - \text{rank}$ free variables). In code, compare np.linalg.matrix_rank(A) with the rank of $[A\,|\,\mathbf{b}]$.
- More unknowns than equations does not guarantee solutions. $x + y + z = 1$ together with $x + y + z = 2$ has 2 equations and 3 unknowns, yet no solution.
- Many equations do not rule out solutions. If every extra equation is a combination of earlier ones, a tall system can still have one.
- Only the ranks decide the verdict, not the shape alone.
Quick check: after reduction you see $\left[\begin{array}{cc|c} 1 & 0 & 2 \\ 0 & 1 & 3 \\ 0 & 0 & 0 \end{array}\right]$. What type is the system?
Consistent (the last row is $0 = 0$), and both unknowns have a pivot: exactly one solution, $(x, y) = (2, 3)$. This is a tall system (3 equations, 2 unknowns) that works because the third equation was redundant.
Underdetermined systems: too few equations core
"I am thinking of two numbers that add up to 10. What are they?" You cannot tell. There is not enough information: $(5, 5)$, $(0, 10)$ and $(-3, 13)$ all work. When a system has fewer equations than unknowns, the rules do not pin everything down, so there is usually a whole family of answers.
The matrix $A$ is wide: it has more columns than rows.
- One equation, two unknowns: $x + y = 4$. Solutions: $(4, 0)$, $(0, 4)$, $(2, 2)$, $(5, -1)$, ... a whole line.
- Two equations, three unknowns: $x + y + z = 6$ and $x - z = 0$. From the second, $x = z$. Put that into the first: $2z + y = 6$, so $y = 6 - 2z$. Every value of $z$ gives a solution: $(z,\; 6 - 2z,\; z)$. For $z = 1$ that is $(1, 4, 1)$: check $1 + 4 + 1 = 6$ ✓ and $1 - 1 = 0$ ✓. This is a line in 3D.
A system is underdetermined when $m \lt n$ (fewer equations than unknowns). Then $\operatorname{rank} A \le m \lt n$, so there is always at least $n - m$ free variables. Therefore:
- If the system is consistent, it has infinitely many solutions.
- It can still be inconsistent (for example $x + y + z = 1$ and $x + y + z = 2$).
- It can never have exactly one solution.
Why do we need it?
Sometimes you have fewer equations than unknowns. We need to understand what that means: not "impossible" but "many answers", and we need a rule to pick one of them.
Where is it used?
Over-parameterised neural networks (more weights than training examples), regression with more features than samples, compressed sensing, and robot arms with more joints than needed (many ways to reach one point).
How is it used?
Reduce, find the free variables and write the family of solutions. If you need just one, pick the smallest solution (np.linalg.lstsq or pinv), or add a regulariser such as ridge to choose for you.
Quick check: a system has 3 equations and 5 unknowns and is consistent. At least how many free variables?
Rank is at most 3, so there are at least $5 - 3 = 2$ free variables, and infinitely many solutions.
Overdetermined systems: too many equations core
Three friends each tell you how far it is from home to the shop. If the shop is at one exact spot, all three answers should agree. But people measure with small errors. With more information than unknowns, the answers usually disagree a little.
Geometrically: three roads that "should" meet at one junction in fact form a tiny triangle. No point lies on all three. The matrix $A$ is tall: more rows than columns.
Solve $x + y = 3$, $\; x + 2y = 5$, $\; x + 3y = 6$.
- Subtract the first from the second: $y = 2$. Then $x = 3 - 2 = 1$.
- Test the third equation: $1 + 3\cdot 2 = 7$, but it needs to be $6$. It fails.
Two equations fix both unknowns; the third is a witness that disagrees. No exact solution.
A system is overdetermined when $m \gt n$ (more equations than unknowns). Typically $\operatorname{rank}[A\,|\,\mathbf{b}] \gt \operatorname{rank}A$, so it is inconsistent. It has a solution only when $\mathbf{b}$ is special (a combination of the columns of $A$).
When there is no exact solution we do the next best thing: find the $\mathbf{x}$ that makes the error $\mathbf{r} = \mathbf{b} - A\mathbf{x}$ as small as possible, by minimising $\|\mathbf{b} - A\mathbf{x}\|^2$. This is least squares, the subject of Chapter 1.10. Its answer solves the "normal equations" $A^\top A\,\mathbf{x} = A^\top\mathbf{b}$.
Why do we need it?
In practice we collect more measurements than unknowns, and noise makes them disagree. We need a way to deal with a system that has no exact solution.
Where is it used?
Linear regression with many data points, fitting a line or curve to measurements, GPS positioning (more satellites than unknowns), camera calibration and sensor calibration.
How is it used?
Do not try to make every equation exact. Find the $\mathbf{x}$ that makes the total squared miss $\|A\mathbf{x}-\mathbf{b}\|^2$ smallest (least squares). In code: np.linalg.lstsq(A, b). Chapter 1.10 explains why this works.
"No solution" here does not mean the problem is hopeless. It means no $\mathbf{x}$ makes the equation exactly true. We replace the question "is $A\mathbf{x} = \mathbf{b}$?" by "how small can $\|A\mathbf{x} - \mathbf{b}\|$ be?" That question always has an answer.
Quick check: you fit a line $y = mx + c$ to 100 points. What is the shape of $A$, and is the system usually solvable exactly?
$A$ is $100\times 2$ (100 equations, 2 unknowns): tall. Unless all 100 points lie on one line, there is no exact solution, so we use least squares.
Homogeneous systems $A\mathbf{x}=\mathbf{0}$ core
A homogeneous system is one where every right-hand side is $0$: "all the equations balance to zero". Every equation is a line (or plane) through the origin.
So there is always one obvious solution: do nothing, $\mathbf{x} = \mathbf{0}$. The interesting question is: is that the only one? Or is there a whole line or plane of other solutions through the origin?
- $x + y = 0$ and $x - y = 0$. Add them: $2x = 0$, so $x = 0$ and then $y = 0$. Only the zero solution.
- $x + y = 0$ and $2x + 2y = 0$. The second is the first doubled. Solutions: $(t, -t)$ for any $t$. A whole line.
- $x + y + z = 0$ and $y - z = 0$. Then $y = z$ and $x = -2z$: solutions $z\cdot(-2, 1, 1)$. A line through the origin in 3D. Check $z = 1$: $-2 + 1 + 1 = 0$ ✓ and $1 - 1 = 0$ ✓.
A system $A\mathbf{x} = \mathbf{0}$ is homogeneous. It is always consistent, since $\mathbf{x} = \mathbf{0}$ (the trivial solution) works. Facts:
- A non-trivial solution exists exactly when there is a free variable, i.e. when $\operatorname{rank}A \lt n$ (the columns of $A$ are linearly dependent, Chapter 1.3).
- If $m \lt n$ (a wide matrix), there is always a non-trivial solution.
- The set of all solutions is called the null space of $A$ (Chapter 1.8). It is closed under adding and scaling: if $A\mathbf{u} = \mathbf{0}$ and $A\mathbf{v} = \mathbf{0}$ then $A(\mathbf{u} + \mathbf{v}) = \mathbf{0}$ and $A(c\mathbf{u}) = \mathbf{0}$. So it is a subspace (Chapter 1.3).
This is just the definition of independence in disguise: the columns are independent exactly when $x_1\mathbf{a}_1 + \dots + x_n\mathbf{a}_n = \mathbf{0}$ forces every $x_i = 0$.
Why do we need it?
The question "is some non-trivial mix of the columns equal to zero?" is the test for dependence. $A\mathbf{x}=\mathbf{0}$ asks exactly that, and its solutions are the directions the matrix ignores.
Where is it used?
Spotting redundant or collinear features, finding eigenvectors (solve $(A-\lambda I)\mathbf{x}=\mathbf{0}$), steady states of Markov chains, kernel (null-space) computations, and Kirchhoff's laws in circuits.
How is it used?
Row reduce $A$ (a zero right-hand side never changes). If every column has a pivot, only $\mathbf{x}=\mathbf{0}$ works. Each free variable gives one non-zero direction. In code: scipy.linalg.null_space(A).
"Always has the solution $\mathbf{0}$" is true but boring. The useful question is whether there are other solutions. Also note: a system with $\mathbf{b} \ne \mathbf{0}$ is not homogeneous, so it does not always have a solution.
Quick check: does $x + 2y + 3z = 0$ (one equation, three unknowns) have non-trivial solutions?
Yes. It is wide ($m = 1 \lt n = 3$), so there are at least 2 free variables. For example $(-2, 1, 0)$ works: $-2 + 2 + 0 = 0$.
The complete solution: $\mathbf{x} = \mathbf{x}_p + \mathbf{x}_n$
Suppose you want to park on a long straight street. If you know one house number where parking is allowed, then you can reach any allowed spot by sliding along the street. The solutions to $A\mathbf{x}=\mathbf{b}$ work the same way:
- Find one solution, any one. Call it the particular solution $\mathbf{x}_p$.
- Then add anything that $A$ sends to zero (a solution of $A\mathbf{x}=\mathbf{0}$, from the null space). Adding it does not change the output.
So the solution set is the null space, slid over by $\mathbf{x}_p$. It is a parallel copy of the null space that does not pass through the origin (unless $\mathbf{b} = \mathbf{0}$).
Solve $x + 2y = 4$ (one equation, two unknowns).
- Particular solution: set $y = 0$, then $x = 4$. So $\mathbf{x}_p = [4, 0]$.
- Null space: solve the homogeneous equation $x + 2y = 0$. Then $x = -2y$, so the solutions are $t\,[-2, 1]$.
- Complete solution: $\mathbf{x} = [4, 0] + t\,[-2, 1]$.
Check: $t = 1$ gives $(2, 1)$, and $2 + 2 = 4$ ✓. $t = 2$ gives $(0, 2)$ and $0 + 4 = 4$ ✓.
From the earlier free-variable example: $x = 3 - 2z$, $y = 1 + z$ means $[x, y, z] = [3, 1, 0] + z\,[-2, 1, 1]$. Here $[3, 1, 0]$ is a particular solution (take $z = 0$), and $[-2, 1, 1]$ solves the homogeneous system.
If $\mathbf{x}_p$ is any one solution of $A\mathbf{x}=\mathbf{b}$, then every solution has the form
$$\mathbf{x} = \mathbf{x}_p + \mathbf{x}_n, \qquad A\mathbf{x}_n = \mathbf{0}.$$Why: $A(\mathbf{x}_p + \mathbf{x}_n) = \mathbf{b} + \mathbf{0} = \mathbf{b}$, so each such vector is a solution. Conversely, if $A\mathbf{x} = \mathbf{b}$ then $A(\mathbf{x} - \mathbf{x}_p) = \mathbf{b} - \mathbf{b} = \mathbf{0}$, so $\mathbf{x} - \mathbf{x}_p$ is in the null space.
Parametric form. With $k = n - \operatorname{rank}A$ free variables, the null space has $k$ basis directions $\mathbf{n}_1, \dots, \mathbf{n}_k$ (one per free variable: set that free variable to 1 and the others to 0 in the homogeneous system). Then
$$\mathbf{x} = \mathbf{x}_p + t_1\mathbf{n}_1 + \dots + t_k\mathbf{n}_k, \qquad t_1, \dots, t_k \in \mathbb{R}.$$A unique solution means $k = 0$: the null space is just $\{\mathbf{0}\}$, and there is nothing to add.
Why do we need it?
When there are many solutions we want to describe all of them with a short formula instead of listing them. One particular solution plus the null space does exactly that.
Where is it used?
Underdetermined regression (all weight vectors that fit the data), minimum-norm and regularised solutions, robots with redundant joints, and linear differential equations (the same "particular + homogeneous" pattern).
How is it used?
Find any one solution $\mathbf{x}_p$ (set the free variables to 0). Find the null-space directions from $A\mathbf{x}=\mathbf{0}$ (set one free variable to 1 at a time). Write $\mathbf{x}=\mathbf{x}_p+t_1\mathbf{n}_1+\dots$ and check by multiplying by $A$.
Different people pick different particular solutions, and all are right. The set $\mathbf{x}_p + \text{null space}$ is the same. Also, if $A\mathbf{x}=\mathbf{b}$ has no solution at all, there is no $\mathbf{x}_p$, and the formula does not apply.
Quick check: $\mathbf{x}_p = [1, 1]$ solves $A\mathbf{x}=\mathbf{b}$, and $[2, -1]$ solves $A\mathbf{x}=\mathbf{0}$. Name three more solutions of $A\mathbf{x}=\mathbf{b}$.
Add multiples of $[2,-1]$ to $[1,1]$. With $t = 1$: $[3, 0]$. With $t = 2$: $[5, -1]$. With $t = -1$: $[-1, 2]$. In general $[1 + 2t,\; 1 - t]$ works for every $t$.
Recap, cheat sheet and practice
- A system of linear equations is written $A\mathbf{x}=\mathbf{b}$, or as the augmented matrix $[A\,|\,\mathbf{b}]$. Rows are equations; columns are unknowns.
- Row picture: lines/planes that must cross. Column picture: can $\mathbf{b}$ be mixed from the columns?
- Three row operations (swap, scale by non-zero, add a multiple) never change the solutions.
- Gaussian elimination gives REF (a staircase); back substitution finishes. Gauss–Jordan gives the unique RREF, with pivots equal to 1 and alone in their columns.
- Pivot columns are pinned; free variables are the columns without pivots. Free variables $= n - \operatorname{rank}A$.
- Only three outcomes: no solution (a row $0 = $ non-zero), exactly one (rank $= n$), or infinitely many (rank $\lt n$).
- Tall (overdetermined) systems usually have no exact solution, which leads to least squares. Wide (underdetermined) systems have infinitely many.
- $A\mathbf{x}=\mathbf{0}$ always has $\mathbf{x}=\mathbf{0}$. All solutions of $A\mathbf{x}=\mathbf{b}$ are $\mathbf{x}_p + \mathbf{x}_n$ (one solution plus the null space).
Cheat sheet
| Idea | Rule | Picture |
|---|---|---|
| Matrix form | $A\mathbf{x}=\mathbf{b}$, $A$ is $m\times n$ | $m$ equations, $n$ unknowns |
| Row ops | swap; scale ($c\ne0$); $R_i \leftarrow R_i + cR_j$ | rewrite without changing the answer |
| REF | staircase, zeros below pivots | ready for back substitution |
| RREF | pivots are 1, alone in column | answer can be read off |
| Free variables | $n - \operatorname{rank}A$ | columns without a pivot |
| No solution | row $[0\ \cdots\ 0\,|\,k]$, $k\ne0$ | parallel lines |
| One solution | consistent and $\operatorname{rank}A=n$ | lines cross once |
| Infinitely many | consistent and $\operatorname{rank}A\lt n$ | same line / shared line |
| Complete solution | $\mathbf{x}_p + t_1\mathbf{n}_1+\dots+t_k\mathbf{n}_k$ | null space shifted by $\mathbf{x}_p$ |
import numpy as np
def gaussian_solve(A, b):
"""Solve Ax = b for a square, invertible A: forward elimination + back substitution."""
A = A.astype(float).copy()
b = b.astype(float).copy()
n = len(b)
for k in range(n):
p = k + np.argmax(np.abs(A[k:, k])) # pivot row (largest entry: partial pivoting)
if abs(A[p, k]) < 1e-12:
raise ValueError("no pivot in column %d: matrix is singular" % k)
if p != k:
A[[k, p]] = A[[p, k]] # swap rows
b[[k, p]] = b[[p, k]]
for i in range(k + 1, n):
f = A[i, k] / A[k, k] # multiplier
A[i, k:] -= f * A[k, k:] # R_i <- R_i - f * R_k
b[i] -= f * b[k]
x = np.zeros(n)
for i in range(n - 1, -1, -1): # back substitution, bottom row first
x[i] = (b[i] - A[i, i + 1:] @ x[i + 1:]) / A[i, i]
return x
A = np.array([[1, 1, 1], [2, 1, -1], [1, -1, 2]])
b = np.array([6, 1, 5])
print(gaussian_solve(A, b)) # [1. 2. 3.]
print(np.linalg.solve(A, b)) # same answer
print(np.allclose(gaussian_solve(A, b), np.linalg.solve(A, b))) # True
# Classify systems with ranks and sympy's RREF (install it first with: pip install sympy)
from sympy import Matrix
aug = Matrix([[1, 1, 1, 3], [1, -1, 1, 1], [2, 0, 2, 5]])
R, pivots = aug.rref()
print(R)
print(pivots) # (0, 1, 3): a pivot in the LAST column, so no solution
def classify(A, b):
r = np.linalg.matrix_rank(A)
ra = np.linalg.matrix_rank(np.column_stack([A, b]))
if ra > r:
return "no solution"
return "one solution" if r == A.shape[1] else "infinitely many (%d free)" % (A.shape[1] - r)
print(classify(np.array([[1, 1, 1], [1, -1, 1], [2, 0, 2]]), np.array([3, 1, 4]))) # infinitely many (1 free)
# Complete solution: particular + null space; and least squares for a tall system
from scipy.linalg import null_space
A = np.array([[1.0, 2.0, 1.0], [0.0, 1.0, 1.0]])
b = np.array([4.0, 1.0])
x_p = np.linalg.lstsq(A, b, rcond=None)[0] # one particular (smallest) solution
N = null_space(A) # columns span the null space
print(A @ (x_p + 3 * N[:, 0])) # still [4. 1.]
tall = np.array([[1.0, 1], [1, 2], [1, 3]])
y = np.array([3.0, 5, 6])
print(np.linalg.lstsq(tall, y, rcond=None)[0]) # about [1.667 1.5]: best compromise (least squares)
1. Which number of solutions is impossible for a linear system?
2. A system has 5 equations and 3 unknowns. It is called…
3. After row reduction you see a row $[0\ 0\ 0\,|\,4]$. What does it mean?
4. A consistent system has 4 unknowns and the reduced matrix has 3 pivots. What is the solution set?
5. Which is not an allowed elementary row operation?
6. $\mathbf{x}_p$ solves $A\mathbf{x}=\mathbf{b}$ and $\mathbf{n}$ solves $A\mathbf{x}=\mathbf{0}$. Then $\mathbf{x}_p + 5\mathbf{n}$…
Practice problems
A. Write $3x - y = 4$ and $x + 2y = -1$ as $A\mathbf{x}=\mathbf{b}$, and check that $(1, -1)$ is a solution.
$A = \begin{bmatrix} 3 & -1 \\ 1 & 2 \end{bmatrix}$, $\mathbf{x} = [x, y]$, $\mathbf{b} = [4, -1]$. Check: $3\cdot1 - (-1) = 4$ ✓ and $1 + 2\cdot(-1) = -1$ ✓.
B. Solve $x + 2y = 5$, $3x + 4y = 11$ by elimination.
$R_2 \leftarrow R_2 - 3R_1$: $[0,\; 4 - 6,\; 11 - 15] = [0, -2, -4]$. So $-2y = -4$, $y = 2$. Then $x + 4 = 5$, $x = 1$. Check: $3 + 8 = 11$ ✓.
C. Describe all solutions of $\left[\begin{array}{ccc|c} 1 & 2 & 0 & 3 \\ 0 & 0 & 1 & 1 \end{array}\right]$.
Pivots in columns 1 and 3, so $y$ is free. Row 1: $x = 3 - 2y$. Row 2: $z = 1$. Solutions: $[3, 0, 1] + y\,[-2, 1, 0]$. Infinitely many.
D. Find the complete solution of $x + 2y + z = 4$, $\; y + z = 1$.
$R_1 \leftarrow R_1 - 2R_2$ gives $x - z = 2$ and $y + z = 1$. So $x = 2 + z$, $y = 1 - z$, $z$ free: $[2, 1, 0] + z\,[1, -1, 1]$. Check $z = 1$: $(3, 0, 1)$ gives $3 + 0 + 1 = 4$ ✓ and $0 + 1 = 1$ ✓.
E. Why can a linear system never have exactly two solutions?
If $\mathbf{x}_1 \ne \mathbf{x}_2$ both solve it, then $\mathbf{x}_1 + t(\mathbf{x}_2 - \mathbf{x}_1)$ solves it for every number $t$, because $A\mathbf{x}_1 + t(A\mathbf{x}_2 - A\mathbf{x}_1) = \mathbf{b} + t\cdot\mathbf{0}$. That is infinitely many solutions.
F. You fit a plane $y = w_1x_1 + w_2x_2 + w_0$ to 50 data points. What shape is the system, and what type is it likely to be?
Three unknowns $(w_1, w_2, w_0)$ and 50 equations: a $50\times 3$ tall, overdetermined system. Unless the points are exactly on a plane, it is inconsistent, so we use least squares.
Matrix Properties
A matrix can hold millions of numbers, but a few summary numbers tell you what it can do: how many directions it keeps (rank), how it scales area (determinant), whether it can be undone (inverse), how big it is (norms), and how touchy it is to small errors (condition number).
- Compute the rank of a matrix and say whether it is full rank or rank-deficient
- Use the trace and its cyclic property
- Compute 2×2 and 3×3 determinants and explain them as signed area and volume scaling
- Decide when a matrix is invertible, find an inverse, and know why we rarely compute one
- Measure the size of a matrix with the Frobenius, spectral and nuclear norms
- Understand the condition number and why tiny errors can blow up
Rank core
A cookbook lists many dishes. But suppose dish C is "half of dish A plus one portion of dish B". Dish C adds nothing new: you could already make it. If you only count the dishes that are genuinely new, you get the number that matters.
The columns of a matrix are like those dishes (arrows in space). The rank counts how many columns are truly independent: how many different directions the matrix really has. It is also the dimension of everything you can build by mixing the columns: a line (rank 1), a flat plane (rank 2), full 3D space (rank 3).
Take $A = \begin{bmatrix} 1 & 2 & 3 \\ 2 & 4 & 6 \\ 1 & 1 & 1 \end{bmatrix}$. Its columns are $\mathbf{a}_1 = [1,2,1]$, $\mathbf{a}_2 = [2,4,1]$, $\mathbf{a}_3 = [3,6,1]$.
- Look for a dependency: $-\mathbf{a}_1 + 2\mathbf{a}_2 = [-1+4,\; -2+8,\; -1+2] = [3, 6, 1] = \mathbf{a}_3$. So the third column is a mix of the first two.
- Count the pivots (Chapter 1.6). $R_2 \leftarrow R_2 - 2R_1$ gives $[0,0,0]$. $R_3 \leftarrow R_3 - R_1$ gives $[0,-1,-2]$. Swap rows 2 and 3: the pivots are in columns 1 and 2.
- Two pivots, so the rank is 2. The matrix is $3\times3$, but it only keeps a plane's worth of directions.
The rank of a matrix $A$, written $\operatorname{rank}(A)$, is the number of pivots after row reduction. Equivalently, it is the number of linearly independent columns, which equals the dimension of the span of the columns (the column space, Chapter 1.8).
A key fact: the number of independent rows is the same number. So $\operatorname{rank}(A) = \operatorname{rank}(A^\top)$.
- For an $m \times n$ matrix, $\operatorname{rank}(A) \le \min(m, n)$.
- Full rank: the rank reaches that maximum, $\operatorname{rank}(A) = \min(m,n)$. (A square $n\times n$ matrix is full rank when its rank is $n$.)
- Rank-deficient: the rank is smaller than $\min(m, n)$. Some column (and some row) is redundant.
- Free variables of $A\mathbf{x}=\mathbf{b}$ number $n - \operatorname{rank}(A)$.
Why do we need it?
A matrix can look big but contain only a few truly different directions. Rank counts them, so we know whether a system has enough information and whether some columns are redundant.
Where is it used?
Detecting collinear features in regression, low-rank approximation (PCA, recommender systems, LoRA fine-tuning of language models), image compression, and deciding whether a system has a unique solution.
How is it used?
Row reduce and count pivots, or call np.linalg.matrix_rank(A) (it uses the SVD and a tolerance). Compare with $\min(m,n)$: equal means full rank, smaller means some columns or rows are redundant.
- Rank counts independent columns, not non-zero columns. A matrix of 100 columns that are all multiples of one vector has rank 1.
- On a computer, "exactly dependent" almost never happens, because of rounding. Libraries compute a numerical rank: they treat tiny pivots (or tiny singular values) as zero, using a tolerance (
np.linalg.matrix_rank).
Quick check: what is the largest possible rank of a $5\times 3$ matrix? If it reaches it, is it full rank?
$\min(5, 3) = 3$. Yes: rank 3 is full (column) rank, meaning all 3 columns are independent.
Trace
The trace is the simplest summary of a square matrix: add up the numbers on the main diagonal (top-left to bottom-right). That is all. It is easy to compute, and it hides a surprise: it also equals the sum of the matrix's eigenvalues (the special stretch factors you will meet in Chapter 1.11).
$A = \begin{bmatrix} 2 & 5 \\ 1 & 3 \end{bmatrix}$ has $\operatorname{tr}(A) = 2 + 3 = 5$. And $B = \begin{bmatrix} 0 & 1 \\ 4 & 2 \end{bmatrix}$ has $\operatorname{tr}(B) = 0 + 2 = 2$.
- $A + B = \begin{bmatrix} 2 & 6 \\ 5 & 5 \end{bmatrix}$ has trace $2 + 5 = 7 = 5 + 2$.
- $AB = \begin{bmatrix} 20 & 12 \\ 12 & 7 \end{bmatrix}$, trace $27$.
- $BA = \begin{bmatrix} 1 & 3 \\ 10 & 26 \end{bmatrix}$, trace $27$.
The matrices $AB$ and $BA$ are different, yet their traces are the same. That is the cyclic property.
- Linear: $\operatorname{tr}(A + B) = \operatorname{tr}(A) + \operatorname{tr}(B)$ and $\operatorname{tr}(cA) = c\,\operatorname{tr}(A)$.
- Transpose: $\operatorname{tr}(A^\top) = \operatorname{tr}(A)$.
- Swap rule: $\operatorname{tr}(AB) = \operatorname{tr}(BA)$ whenever both products make sense (even if $AB$ and $BA$ have different sizes).
- Cyclic property: you may move the last matrix to the front: $\operatorname{tr}(ABC) = \operatorname{tr}(BCA) = \operatorname{tr}(CAB)$. But $\operatorname{tr}(ACB)$ is generally different.
- Eigenvalues (preview): $\operatorname{tr}(A) = \lambda_1 + \dots + \lambda_n$.
- Also $\|A\|_F^2 = \operatorname{tr}(A^\top A)$, which connects the trace to matrix size (see Matrix norms below).
Why do we need it?
We want one cheap number that summarises a square matrix and behaves nicely. The trace is that number: easy to compute, linear, and equal to the sum of the eigenvalues.
Where is it used?
The total variance of data (trace of a covariance matrix), the Frobenius norm $\|W\|_F^2=\operatorname{tr}(W^\top W)$ in weight decay, matrix calculus for ML losses (the cyclic property), and physics formulas.
How is it used?
Add the diagonal: np.trace(A). Use $\operatorname{tr}(ABC)=\operatorname{tr}(BCA)$ to rearrange products inside a derivation. Use "trace = sum of eigenvalues" as a quick check on a computation.
The trace is only defined for square matrices. And "$\operatorname{tr}(AB) = \operatorname{tr}(BA)$" does not mean $AB = BA$. Also, $\operatorname{tr}(AB) \ne \operatorname{tr}(A)\operatorname{tr}(B)$ in general; the trace is not multiplicative.
Quick check: $\operatorname{tr}(A)=3$ and $\operatorname{tr}(B)=-1$ for $2\times2$ matrices. What is $\operatorname{tr}(2A - B)$?
By linearity: $2\cdot3 - (-1) = 7$.
The determinant core
Draw a unit square (area 1) on a rubber sheet. Now apply a matrix: the sheet stretches, and the square becomes a parallelogram. The determinant is the number that says how much the area was multiplied.
- $\det = 2$: every area doubles.
- $\det = 0.5$: every area halves.
- $\det = 0$: the sheet is squashed flat onto a line, so the area becomes 0.
- $\det \lt 0$: the sheet is also flipped over (like turning a page over to see its back), so left and right swap. The size is the number without the minus sign.
In 3D the same idea holds for volume: the unit cube becomes a slanted box (a parallelepiped), and the determinant is its signed volume.
2×2. $A = \begin{bmatrix} 3 & 1 \\ 2 & 4 \end{bmatrix}$. The columns $[3,2]$ and $[1,4]$ are the images of the two unit arrows.
- Multiply along the main diagonal: $3\cdot 4 = 12$.
- Multiply along the other diagonal: $1\cdot 2 = 2$.
- Subtract: $\det A = 12 - 2 = 10$. So areas become 10 times bigger.
Swap the two columns: $\begin{bmatrix} 1 & 3 \\ 4 & 2 \end{bmatrix}$ has $\det = 1\cdot2 - 3\cdot4 = -10$. Same size, but flipped.
3×3. $B = \begin{bmatrix} 2 & 0 & 1 \\ 1 & 3 & 2 \\ 0 & 1 & 1 \end{bmatrix}$. Walk along the top row. Each entry multiplies the determinant of the smaller $2\times2$ matrix left when you cross out its row and column, with signs $+,-,+$:
- $+\,2\cdot\det\begin{bmatrix} 3 & 2 \\ 1 & 1 \end{bmatrix} = 2\cdot(3 - 2) = 2$
- $-\,0\cdot\det\begin{bmatrix} 1 & 2 \\ 0 & 1 \end{bmatrix} = 0$
- $+\,1\cdot\det\begin{bmatrix} 1 & 3 \\ 0 & 1 \end{bmatrix} = 1\cdot(1 - 0) = 1$
Add: $\det B = 2 + 0 + 1 = 3$. The box built from the three columns has volume 3.
For a $2\times2$ matrix:
$$\det\begin{bmatrix} a & b \\ c & d \end{bmatrix} = ad - bc.$$For a $3\times 3$ matrix (expanding along the first row):
$$\det\begin{bmatrix} a & b & c \\ d & e & f \\ g & h & i \end{bmatrix} = a(ei - fh) - b(di - fg) + c(dh - eg).$$Geometric meaning: for a square matrix $A$, $|\det A|$ is the factor by which $A$ scales areas (2D), volumes (3D) or $n$-dimensional volumes, and the sign says whether orientation is kept ($+$) or flipped ($-$). The columns of $A$ are the edges of the box.
And $\det A = 0$ exactly when the columns are linearly dependent, i.e. the box is flat.
Why do we need it?
We need one number that tells whether a square matrix squashes space flat (so it cannot be undone) and by how much it scales area or volume.
Where is it used?
Gaussian likelihoods (the $\log\det$ of the covariance), change of variables for probability densities and normalising flows, testing invertibility, orientation tests in graphics, and Jacobians in calculus.
How is it used?
For $2\times2$ use $ad-bc$. For bigger matrices row reduce, multiply the pivots, and flip the sign once for every row swap (np.linalg.det does this for you). Read $|\det|$ as the scale factor and the sign as "flipped or not". For large matrices use slogdet to avoid overflow.
- The determinant is defined only for square matrices.
- Do not forget the minus: $ad - bc$, not $ad + bc$.
- A negative determinant is not "bad". It only means the orientation flipped (a reflection is included). The size of the scaling is $|\det A|$.
Quick check: compute $\det\begin{bmatrix} 5 & 2 \\ 7 & 3 \end{bmatrix}$, and say what it does to areas.
$5\cdot3 - 2\cdot7 = 15 - 14 = 1$. Areas keep their size, and orientation is kept.
Determinant rules, and cofactor expansion
Because the determinant is "the area scale factor", its rules follow from common sense about stretching.
- Do transformation $B$ first, then $A$: the area is scaled by $\det B$, then by $\det A$. Total: $\det A\cdot\det B$. So $\det(AB) = \det A\det B$.
- Undoing a transformation that scaled area by 5 must scale it by $\tfrac15$. So $\det(A^{-1}) = 1/\det A$.
- Swapping two columns (or rows) flips the sheet over: the sign changes.
- A flat transformation cannot be undone, so no inverse exists when $\det = 0$.
Let $A = \begin{bmatrix} 1 & 2 \\ 3 & 4 \end{bmatrix}$ ($\det A = 4 - 6 = -2$) and $B = \begin{bmatrix} 2 & 0 \\ 1 & 3 \end{bmatrix}$ ($\det B = 6 - 0 = 6$).
- $AB = \begin{bmatrix} 1\cdot2 + 2\cdot1 & 1\cdot0 + 2\cdot3 \\ 3\cdot2 + 4\cdot1 & 3\cdot0 + 4\cdot3 \end{bmatrix} = \begin{bmatrix} 4 & 6 \\ 10 & 12 \end{bmatrix}$.
- $\det(AB) = 4\cdot12 - 6\cdot10 = 48 - 60 = -12$.
- And $\det A\cdot\det B = (-2)(6) = -12$ ✓.
A faster way for big matrices. Row reduce to a triangle. For $\begin{bmatrix} 2 & 1 & 1 \\ 4 & 3 & 3 \\ 8 & 7 & 9 \end{bmatrix}$: $R_2 \leftarrow R_2 - 2R_1$, $R_3 \leftarrow R_3 - 4R_1$, $R_3 \leftarrow R_3 - 3R_2$ give a triangle with pivots $2, 1, 2$. These operations do not change the determinant, so $\det = 2\cdot1\cdot2 = 4$.
Key properties (all matrices $n\times n$):
- $\det(AB) = \det(A)\det(B)$.
- $\det(A^\top) = \det(A)$.
- $\det(A^{-1}) = 1/\det(A)$ (when $A$ is invertible).
- $\det(cA) = c^n\det(A)$ (each of the $n$ rows gets multiplied by $c$).
- Swap two rows: the sign flips. Multiply one row by $c$: the determinant is multiplied by $c$. Add a multiple of one row to another: no change.
- For a triangular (or diagonal) matrix, $\det$ = the product of the diagonal entries. So $\det A = \pm$ (product of the pivots), with a minus sign for each row swap used.
- $\det A = 0 \iff A$ is singular (not invertible).
- Eigenvalues (preview): $\det A = \lambda_1\lambda_2\cdots\lambda_n$, the product of the eigenvalues (Chapter 1.11).
Cofactor expansion (awareness). Pick any row $i$. Let $M_{ij}$ be the determinant of the matrix left after deleting row $i$ and column $j$ (a minor). Then
$$\det A = \sum_{j=1}^{n} (-1)^{i+j}\,a_{ij}\,M_{ij}.$$The signs follow a checkerboard: $\begin{smallmatrix} + & - & + \\ - & + & - \\ + & - & + \end{smallmatrix}$. You get the same answer along any row, or any column. This is a beautiful definition but a slow way to compute: it takes about $n!$ steps. Elimination takes about $n^3$.
Why do we need it?
Rules let us compute and reason without starting from scratch. For products, transposes, inverses and row operations we know exactly how the determinant changes.
Where is it used?
Jacobian determinants in probability (change of variables), the log-determinant in multivariate Gaussian densities, proofs of invertibility, and eigenvalue formulas ($\det$ = product of the eigenvalues).
How is it used?
Use $\det(AB)=\det A\det B$ to split a problem. To compute $\det$, row reduce and keep track of swaps and scalings. Remember $\det(cA)=c^n\det A$. Use cofactor expansion only by hand on small matrices.
- $\det(A + B)$ is not $\det A + \det B$. The determinant is not linear in the whole matrix.
- $\det(cA) = c^n\det A$, not $c\det A$. Doubling a $3\times3$ matrix multiplies the determinant by $8$.
- Do not use cofactor expansion for big matrices. Use elimination.
Quick check: $A$ is $3\times3$ with $\det A = 2$. What are $\det(3A)$, $\det(A^{-1})$ and $\det(A^\top A)$?
$\det(3A) = 3^3\cdot2 = 54$. $\det(A^{-1}) = 1/2$. $\det(A^\top A) = \det(A^\top)\det(A) = 2\cdot2 = 4$.
The inverse of a matrix core
A matrix is a machine that moves the plane around. The inverse is the undo button: apply $A$, then apply $A^{-1}$, and every point returns to where it started.
Some machines can be undone (rotate, stretch, shear). Others cannot: if a machine squashes the whole plane flat onto a line, many different points land on the same spot, and you can never tell where each one came from. Such a matrix is singular. Its determinant is $0$.
Another everyday picture: you put on socks, then shoes. To undo, you remove the shoes first, then the socks: the reverse order. This is why $(AB)^{-1} = B^{-1}A^{-1}$.
$A = \begin{bmatrix} 4 & 7 \\ 2 & 6 \end{bmatrix}$. Use the $2\times 2$ formula.
- $\det A = 4\cdot6 - 7\cdot2 = 24 - 14 = 10$ (not zero, so an inverse exists).
- Swap the diagonal entries and flip the signs of the other two: $\begin{bmatrix} 6 & -7 \\ -2 & 4 \end{bmatrix}$.
- Divide by the determinant: $A^{-1} = \tfrac{1}{10}\begin{bmatrix} 6 & -7 \\ -2 & 4 \end{bmatrix} = \begin{bmatrix} 0.6 & -0.7 \\ -0.2 & 0.4 \end{bmatrix}$.
- Check: row 1 times column 1: $4\cdot0.6 + 7\cdot(-0.2) = 2.4 - 1.4 = 1$. Row 1 times column 2: $4\cdot(-0.7) + 7\cdot0.4 = 0$. The other two entries give $0$ and $1$, so $AA^{-1} = I$ ✓.
A singular example. $\begin{bmatrix} 1 & 2 \\ 2 & 4 \end{bmatrix}$ has $\det = 4 - 4 = 0$. Its second row is twice the first; it squashes the plane onto the line $y = 2x$. No inverse.
Reverse order. Let $A = \begin{bmatrix} 1 & 1 \\ 0 & 1 \end{bmatrix}$ (a shear) and $B = \begin{bmatrix} 2 & 0 \\ 0 & 1 \end{bmatrix}$ (stretch $x$ by 2). Then $AB = \begin{bmatrix} 2 & 1 \\ 0 & 1 \end{bmatrix}$ and $(AB)^{-1} = \begin{bmatrix} 0.5 & -0.5 \\ 0 & 1 \end{bmatrix}$. Also $B^{-1}A^{-1} = \begin{bmatrix} 0.5 & 0 \\ 0 & 1\end{bmatrix}\begin{bmatrix} 1 & -1 \\ 0 & 1\end{bmatrix} = \begin{bmatrix} 0.5 & -0.5 \\ 0 & 1\end{bmatrix}$ ✓. But $A^{-1}B^{-1} = \begin{bmatrix} 0.5 & -1 \\ 0 & 1\end{bmatrix}$, which is wrong.
A square matrix $A$ is invertible (also called non-singular) if there is a matrix $A^{-1}$ with
$$AA^{-1} = A^{-1}A = I.$$If no such matrix exists, $A$ is singular. Only square matrices can have an inverse, and the inverse, when it exists, is unique.
The $2\times2$ formula:
$$\begin{bmatrix} a & b \\ c & d \end{bmatrix}^{-1} = \frac{1}{ad - bc}\begin{bmatrix} d & -b \\ -c & a \end{bmatrix}, \qquad \text{valid only if } ad - bc \ne 0.$$Rules: $(AB)^{-1} = B^{-1}A^{-1}$ (reverse order); $(A^\top)^{-1} = (A^{-1})^\top$; $(A^{-1})^{-1} = A$; $\det(A^{-1}) = 1/\det A$. If $A$ is invertible, $A\mathbf{x}=\mathbf{b}$ has the unique solution $\mathbf{x} = A^{-1}\mathbf{b}$.
Why do we need it?
When a transformation or system must be undone, we need an "undo button". The inverse is that button, and whether it exists tells us whether the problem has a unique answer.
Where is it used?
Writing down the normal equations $(X^\top X)^{-1}X^\top\mathbf{y}$, Gaussian distributions (the precision matrix), undoing transformations in graphics and robotics, and Kalman filters.
How is it used?
Check that the matrix is square with $\det\ne0$. For $2\times2$ use the formula; otherwise np.linalg.inv(A), but to solve $A\mathbf{x}=\mathbf{b}$ prefer solve. Verify with $A A^{-1}\approx I$. Reverse the order for products.
- Only square matrices have inverses.
- Reverse the order: $(AB)^{-1} = B^{-1}A^{-1}$, not $A^{-1}B^{-1}$.
- $(A+B)^{-1}$ is not $A^{-1} + B^{-1}$.
- "Almost singular" is also trouble (see the condition number below). A tiny determinant is not itself the problem, though; see that section.
Quick check: find the inverse of $\begin{bmatrix} 1 & 2 \\ 3 & 4 \end{bmatrix}$.
$\det = 4 - 6 = -2$. So $A^{-1} = \tfrac{1}{-2}\begin{bmatrix} 4 & -2 \\ -3 & 1 \end{bmatrix} = \begin{bmatrix} -2 & 1 \\ 1.5 & -0.5 \end{bmatrix}$.
The Invertible Matrix Theorem
For a square matrix, a whole list of statements are secretly the same statement in different clothes. If one is true, they are all true. If one is false, they are all false. They all say: "this matrix loses no information and collapses no direction."
That is very useful. To show a matrix is invertible, you may check whichever item is easiest. To show it is not, one failing item is enough.
$A = \begin{bmatrix} 1 & 2 \\ 3 & 4 \end{bmatrix}$: $\det = -2 \ne 0$. The columns $[1,3]$ and $[2,4]$ are not multiples of each other, so they are independent. Row reduction finds 2 pivots. $A\mathbf{x}=\mathbf{0}$ has only $\mathbf{x}=\mathbf{0}$. All the statements below hold.
$B = \begin{bmatrix} 1 & 2 \\ 2 & 4 \end{bmatrix}$: $\det = 0$. Column 2 is twice column 1, so the columns are dependent. Only 1 pivot. And $B\begin{bmatrix} 2 \\ -1 \end{bmatrix} = \begin{bmatrix} 0 \\ 0 \end{bmatrix}$, a non-zero solution of $B\mathbf{x}=\mathbf{0}$. All the statements fail together.
Invertible Matrix Theorem. For a square $n\times n$ matrix $A$, the following are equivalent (all true, or all false):
- $A$ is invertible (a matrix $A^{-1}$ exists).
- $\det A \ne 0$.
- $\operatorname{rank}A = n$ (there are $n$ pivots).
- The reduced row echelon form of $A$ is the identity $I$.
- The columns of $A$ are linearly independent.
- The columns of $A$ span $\mathbb{R}^n$.
- $A\mathbf{x}=\mathbf{0}$ has only the trivial solution $\mathbf{x}=\mathbf{0}$.
- $A\mathbf{x}=\mathbf{b}$ has exactly one solution for every $\mathbf{b}$.
- $A^\top$ is invertible (so the rows are independent too).
- $0$ is not an eigenvalue of $A$ (you will meet eigenvalues in Chapter 1.11).
Why do we need it?
Invertibility appears in many disguises (rank, determinant, solvability, independence). The theorem lets us pick whichever test is easiest, and know it answers all of them.
Where is it used?
Checking that $X^\top X$ can be inverted in regression (independent features), confirming that a set of vectors is a valid basis, and asking whether a layer's weight matrix loses information.
How is it used?
Choose the cheapest test: $\det\ne0$ for small matrices, counting pivots by hand, np.linalg.matrix_rank in code. If any one statement fails, the matrix is singular and all the others fail too.
The theorem is only about square matrices. A tall or wide matrix is never "invertible", and the list does not apply. For those, you use least squares and the pseudo-inverse (Chapters 1.10 and 1.13).
Checkpoint: list five equivalent conditions for a square matrix to be invertible.
Any five of: $\det\ne0$; rank $=n$; RREF is $I$; columns independent; columns span $\mathbb{R}^n$; $A\mathbf{x}=\mathbf{0}$ only for $\mathbf{x}=\mathbf{0}$; $A\mathbf{x}=\mathbf{b}$ has a unique solution for every $\mathbf{b}$; $A^\top$ invertible; no zero eigenvalue.
Computing an inverse with Gauss–Jordan elimination
Finding $A^{-1}$ means solving $A\mathbf{x} = \mathbf{e}_1$, then $A\mathbf{x} = \mathbf{e}_2$, and so on (each unit vector gives one column of the inverse). We can solve all of them at once: put $A$ on the left, the identity matrix $I$ on the right, and row reduce (Gauss–Jordan, Chapter 1.6) until the left side becomes $I$. Whatever the right side has turned into is $A^{-1}$.
Invert $A = \begin{bmatrix} 2 & 1 \\ 5 & 3 \end{bmatrix}$. Start with $\left[\begin{array}{cc|cc} 2 & 1 & 1 & 0 \\ 5 & 3 & 0 & 1 \end{array}\right]$.
- $R_1 \leftarrow R_1 \div 2$: $\left[\begin{array}{cc|cc} 1 & \tfrac12 & \tfrac12 & 0 \\ 5 & 3 & 0 & 1 \end{array}\right]$.
- $R_2 \leftarrow R_2 - 5R_1$: row 2 becomes $[0,\; 3 - \tfrac52,\; 0 - \tfrac52,\; 1] = [0, \tfrac12, -\tfrac52, 1]$.
- $R_2 \leftarrow 2R_2$: $[0, 1, -5, 2]$.
- $R_1 \leftarrow R_1 - \tfrac12R_2$: row 1 becomes $[1,\; 0,\; \tfrac12 + \tfrac52,\; 0 - 1] = [1, 0, 3, -1]$.
The left side is $I$. So $A^{-1} = \begin{bmatrix} 3 & -1 \\ -5 & 2 \end{bmatrix}$. Check the first column: $2\cdot3 + 1\cdot(-5) = 1$ ✓ and $5\cdot3 + 3\cdot(-5) = 0$ ✓.
If row operations turn $[A\,|\,I]$ into $[I\,|\,B]$, then $B = A^{-1}$. If the left side cannot reach $I$ (a column has no pivot), then $A$ is singular.
Why it works: each row operation is a multiplication by some matrix. Together they multiply to a matrix $E$ with $EA = I$. So $E = A^{-1}$. The same operations applied to $I$ produce $EI = A^{-1}$. This costs about $n^3$ operations.
Why do we need it?
Sometimes you really do need the inverse matrix itself. Gauss–Jordan gives a systematic hand method, and it also shows you when no inverse exists.
Where is it used?
Teaching and exams, small symbolic calculations, inverting tiny covariance matrices, and as the idea behind library routines that build an inverse from an LU factorisation.
How is it used?
Write $[A\,|\,I]$ and row reduce until the left side is $I$. The right side is then $A^{-1}$. If a zero row appears on the left, stop: $A$ is singular. Check by multiplying $AA^{-1}$.
Work with fractions or keep many decimals, and check your answer by multiplying $AA^{-1}$: you must get $I$. A singular matrix shows itself by a column with no pivot, not by a wrong answer.
Quick check: while row reducing $[A\,|\,I]$ the left half ends up with a row of all zeros. What does that mean?
The left half cannot become $I$, because $A$ has fewer than $n$ pivots. So $A$ is singular and has no inverse.
Why you almost never compute an inverse (use solve)
Suppose you want one cake. Computing $A^{-1}$ is like building a whole bakery that can bake every possible cake, and then using it once. You only needed the answer $\mathbf{x}$ for one right-hand side $\mathbf{b}$.
The formula $\mathbf{x} = A^{-1}\mathbf{b}$ is wonderful for writing and reasoning. But as a recipe for the computer, it is wasteful, a little less accurate, and it can destroy useful structure. Read "$A^{-1}\mathbf{b}$" as "solve $A\mathbf{x}=\mathbf{b}$", and call a solver.
Suppose $A$ is $1000\times1000$.
np.linalg.solve(A, b)does Gaussian elimination: roughly $\tfrac23 n^3 \approx 0.7$ billion operations.np.linalg.inv(A) @ bbuilds all 1000 columns of the inverse: roughly $2n^3 = 2$ billion operations, about three times the work, and then multiplies once more.
And if $A$ is mostly zeros (a "sparse" matrix with a few non-zeros per row), $A^{-1}$ is typically completely full of non-zeros. Storing it can be impossible, while a sparse solver would be fast.
Rules of thumb.
- To solve $A\mathbf{x}=\mathbf{b}$: use
np.linalg.solve(A, b), notinv(A) @ b. - To solve for many right-hand sides with the same $A$: factor $A$ once (LU or Cholesky, Chapter 1.13) and reuse the factors.
- For tall systems (more equations than unknowns): use
np.linalg.lstsq(Chapter 1.10), never an explicit inverse of $A^\top A$. - If you need a quantity like $\mathbf{u}^\top A^{-1}\mathbf{v}$, solve $A\mathbf{z}=\mathbf{v}$ first, then compute $\mathbf{u}^\top\mathbf{z}$.
- Explicit inverses are fine for tiny matrices (2×2 by hand) and when you truly need the inverse's entries (for example a covariance matrix of weights).
Why do we need it?
Computing a whole inverse to get one answer wastes time, can lose accuracy, and destroys structure such as sparsity. Solving directly avoids all three problems.
Where is it used?
Every serious numerical library (np.linalg.solve, SciPy, LAPACK), Gaussian-process code (solve with a Cholesky factor), least squares (lstsq), Newton steps in optimisation, and simulations.
How is it used?
Whenever you see $A^{-1}\mathbf{b}$ in a formula, write np.linalg.solve(A, b). For many right-hand sides, factor once (LU or Cholesky). For tall systems use np.linalg.lstsq. Form an explicit inverse only for tiny matrices.
This is not magic: both methods suffer when the matrix is ill-conditioned (see below). The point is that the inverse route does more work for a result that is no more accurate. In the demo above its residual is far worse, while the error in $\mathbf{x}$ itself can go either way. And "A is invertible" is a statement about existence; it does not say you should compute $A^{-1}$.
Quick check: how would you compute $\mathbf{u}^\top A^{-1}\mathbf{v}$ without forming $A^{-1}$?
Solve $A\mathbf{z} = \mathbf{v}$ for $\mathbf{z}$ (one solve), then take the dot product $\mathbf{u}^\top\mathbf{z}$.
Matrix norms
How "big" is a matrix? There are two natural answers.
- Count every entry. Lay all the entries in one long list and take its ordinary length. This is the Frobenius norm.
- Measure the biggest stretch. Feed the matrix every vector of length 1 (the unit circle) and see how far the longest output reaches. A matrix turns the unit circle into an ellipse. The longest half-width of that ellipse is the spectral norm: the most the matrix can stretch any vector. It is one of the induced (operator) norms.
The ellipse's half-widths are called the matrix's singular values $\sigma_1 \ge \sigma_2 \ge \dots$ (stretch factors; full story in Chapter 1.13). The spectral norm is the biggest, $\sigma_1$. Adding all of them up gives the nuclear norm.
$D = \begin{bmatrix} 3 & 0 \\ 0 & 1 \end{bmatrix}$ stretches $x$ by 3 and leaves $y$. The unit circle becomes an ellipse with half-widths 3 and 1.
- Frobenius: $\sqrt{3^2 + 0^2 + 0^2 + 1^2} = \sqrt{10} \approx 3.162$.
- Spectral: $\sigma_1 = 3$ (the longest half-width).
- Nuclear: $3 + 1 = 4$.
A less simple one, $A = \begin{bmatrix} 1 & 2 \\ 3 & 4 \end{bmatrix}$: Frobenius $= \sqrt{1 + 4 + 9 + 16} = \sqrt{30} \approx 5.477$. Its singular values are about $5.465$ and $0.366$. So the spectral norm is $\approx 5.465$ and the nuclear norm is $\approx 5.831$. Notice $5.465^2 + 0.366^2 = 29.87 + 0.13 = 30$: the Frobenius norm squared is the sum of the squared singular values.
- Frobenius norm: $\|A\|_F = \sqrt{\sum_{i,j} a_{ij}^2} = \sqrt{\operatorname{tr}(A^\top A)} = \sqrt{\sigma_1^2 + \sigma_2^2 + \cdots}$.
- Induced (operator) norm: $\|A\|_p = \max_{\mathbf{x}\ne\mathbf{0}} \dfrac{\|A\mathbf{x}\|_p}{\|\mathbf{x}\|_p}$ = "the largest stretch factor" measured with the vector $p$-norm.
- $p = 2$: spectral norm $\|A\|_2 = \sigma_{\max}$, the largest singular value.
- $p = 1$: the largest absolute column sum. $p = \infty$: the largest absolute row sum.
- Nuclear norm (awareness): $\|A\|_* = \sigma_1 + \sigma_2 + \cdots$, the sum of the singular values.
Useful rules: $\|A\mathbf{x}\| \le \|A\|\,\|\mathbf{x}\|$ and $\|AB\| \le \|A\|\,\|B\|$ (for induced norms and the Frobenius norm). Also $\|A\|_2 \le \|A\|_F \le \sqrt{r}\,\|A\|_2$, where $r$ is the rank.
Why do we need it?
We need a number for "how big is this matrix" and "how much can it stretch a vector", so we can compare matrices, bound errors and keep weights under control.
Where is it used?
Weight decay (Frobenius norm), spectral normalisation in GANs, matrix-factorisation losses $\|X-UV^\top\|_F^2$, nuclear-norm penalties for low-rank matrix completion, and error bounds in numerical linear algebra.
How is it used?
Call np.linalg.norm(A, 'fro'), np.linalg.norm(A, 2) (spectral) or 'nuc'. Use $\|A\mathbf{x}\|\le\|A\|\|\mathbf{x}\|$ to bound how much a layer can amplify an input.
- The Frobenius norm is not an induced norm. It treats the matrix as a long list. It is always at least as big as the spectral norm.
- "The norm of a matrix" with no label usually means the spectral norm (in
np.linalg.norm(A, 2)) or the Frobenius norm (the NumPy default for matrices). Always check which. - A matrix norm measures size, not invertibility: a matrix can be huge and singular.
Quick check: for $\begin{bmatrix} 3 & 0 \\ 0 & 4 \end{bmatrix}$, what are the Frobenius, spectral and nuclear norms?
Singular values 4 and 3. Frobenius: $\sqrt{9 + 16} = 5$. Spectral: $4$. Nuclear: $3 + 4 = 7$.
The condition number (introduction)
Two roads that cross at a wide angle meet at a clear junction. Nudge one road a little and the junction barely moves.
Two roads that are almost parallel also meet somewhere, but that junction is slippery: shift one road a tiny bit and the crossing point slides a long way. A system $A\mathbf{x}=\mathbf{b}$ with almost-parallel rows (or almost-dependent columns) is like that. A small change in the data $\mathbf{b}$, such as a measurement error or a rounding error, causes a large change in the answer $\mathbf{x}$.
The condition number $\kappa$ is the number that measures this: how many times a relative error in $\mathbf{b}$ can be magnified in $\mathbf{x}$.
$x + y = 2$ and $x + 1.01y = 2.01$. Subtract: $0.01y = 0.01$, so $y = 1$, $x = 1$.
Now change the last right-hand side from $2.01$ to $2.02$ (a change of 0.01, only about 0.5% of that number). Subtract again: $0.01y = 0.02$, so $y = 2$ and $x = 0$.
A 0.5% change in the data moved the answer from $(1, 1)$ to $(0, 2)$: a change as large as the answer itself. This matrix is ill-conditioned. Its condition number is about 400.
For an invertible matrix $A$,
$$\kappa(A) = \|A\|\,\|A^{-1}\| \;=\; \frac{\sigma_{\max}}{\sigma_{\min}} \quad\text{(for the 2-norm).}$$It is the ratio of the biggest stretch to the smallest stretch (the long half-width over the short half-width of the ellipse). Always $\kappa \ge 1$. A singular matrix has $\kappa = \infty$.
- Sensitivity: $\dfrac{\|\Delta\mathbf{x}\|}{\|\mathbf{x}\|} \le \kappa(A)\,\dfrac{\|\Delta\mathbf{b}\|}{\|\mathbf{b}\|}$.
- Well-conditioned: $\kappa$ is small (near 1). Ill-conditioned: $\kappa$ is huge (say $10^8$ or more).
- Digits lost: a rule of thumb is that you lose about $\log_{10}\kappa$ correct digits in the answer.
The best possible value is $\kappa = 1$, reached by rotations and reflections (orthogonal matrices, Chapter 1.9).
Why do we need it?
Some problems magnify tiny input errors into huge output errors. The condition number measures this in advance, so you know how far to trust an answer.
Where is it used?
Choosing least-squares solvers (the normal equations square $\kappa$), ridge regularisation for correlated features, diagnosing unstable training, and checking linear systems in engineering simulations.
How is it used?
Compute np.linalg.cond(A). Near 1 is healthy. Around $10^8$ you lose about half of your 16 digits, and near $10^{16}$ you lose all of them. Fix ill-conditioning by regularising (add $\lambda I$), rescaling features, or using QR or SVD instead of the normal equations.
- A small determinant is not the same as ill-conditioned. $\begin{bmatrix} 10^{-6} & 0 \\ 0 & 10^{-6}\end{bmatrix}$ has $\det = 10^{-12}$ but $\kappa = 1$ (it just shrinks everything evenly). The condition number, not the determinant, measures sensitivity.
- Conditioning belongs to the problem. No algorithm can fully remove it. (Whether an algorithm adds more error is a separate question: stability, next.)
Quick check: $\kappa(A) = 10^8$ and your data $\mathbf{b}$ has about 8 correct digits. How many correct digits can you trust in $\mathbf{x}$?
Roughly $8 - \log_{10}(10^8) = 0$. Possibly none. That is why ill-conditioned systems are dangerous.
A first look at numerical stability
A computer stores numbers with about 16 digits, like a calculator with a small screen. Every calculation is rounded a little. Usually that is harmless. But two things make it dangerous:
- Subtracting nearly equal numbers. $1.0000000000000001 - 1$ should be tiny but non-zero. If the computer could only keep the "1.000000000000000" part, the answer becomes $0$: all the digits that mattered were thrown away. This is called cancellation.
- Ill-conditioned problems take those tiny rounding errors and magnify them by $\kappa$.
An algorithm is called stable if it does not make the rounding errors worse than the problem forces. This chapter only gives a taste; the full treatment is in Chapter 1.15.
In Python, 0.1 + 0.2 gives 0.30000000000000004, not 0.3. So 0.1 + 0.2 == 0.3 is False.
Take $A = \begin{bmatrix} 1 & 1 \\ 1 & 1+\varepsilon \end{bmatrix}$. Its determinant is exactly $(1)(1+\varepsilon) - (1)(1) = \varepsilon$. For $\varepsilon = 10^{-8}$ the true value is $0.00000001$. But the computer first has to store $1 + 10^{-8}$, and then subtract $1$. For $\varepsilon = 10^{-16}$ the stored value of $1 + \varepsilon$ is just $1$, so it computes $\det = 0$ and thinks the matrix is singular, even though the true determinant is not zero.
Floating-point numbers (64-bit "doubles") have a machine epsilon of about $2.2\times10^{-16}$: the gap between 1 and the next number the computer can store. One operation therefore makes a relative rounding error of at most about $1.1\times10^{-16}$. That is why $1 + 10^{-16}$ is stored as exactly $1$. For $A\mathbf{x}=\mathbf{b}$, a stable solver gives an answer whose relative error is roughly
$$\text{error} \;\approx\; \kappa(A)\times 10^{-16}\quad\text{(a rough rule).}$$Practical rules:
- Never test
det(A) == 0orx == yfor floats. Use a tolerance (np.isclose,np.linalg.matrix_rank). - Judge near-singularity by the condition number, not the determinant.
- Prefer
solve, QR or SVD over explicit inverses and over normal equations. - Row swaps (partial pivoting) keep Gaussian elimination stable. That is why the pivot is chosen as the biggest entry.
Why do we need it?
Computers round every number, so results are only approximately right. Knowing when rounding is harmless and when it ruins an answer stops silent, wrong results.
Where is it used?
NaN or exploding losses in training, overflow in softmax and log-likelihoods (the log-sum-exp trick), float16 versus float32 choices (mixed precision), and any code that tests a determinant or forms an inverse.
How is it used?
Never compare floats with ==; use np.isclose. Judge near-singularity with cond or matrix_rank, not the determinant. Prefer solve, QR and SVD. If a result looks odd, repeat in higher precision.
The computed determinant being $0$ does not prove the matrix is singular, and a computed determinant of $10^{-10}$ does not prove it is nearly singular. Use np.linalg.cond or np.linalg.matrix_rank (which use the SVD) to judge.
Quick check: why does 0.1 + 0.2 == 0.3 fail in Python, and how should you compare?
Neither 0.1 nor 0.2 can be stored exactly in binary, so each carries a tiny error, and the sum is off in the 17th digit. Compare with a tolerance: np.isclose(0.1 + 0.2, 0.3) is True.
Recap, cheat sheet and practice
- Rank = number of pivots = number of independent columns (and of independent rows). Full rank means it reaches $\min(m,n)$; otherwise the matrix is rank-deficient.
- Trace = sum of the diagonal. It is linear, $\operatorname{tr}(AB)=\operatorname{tr}(BA)$, and it is cyclic. It also equals the sum of the eigenvalues.
- Determinant = signed area/volume scale factor. $2\times2$: $ad-bc$. Rules: $\det(AB)=\det A\det B$, $\det A^\top=\det A$, $\det A^{-1}=1/\det A$. It is zero exactly when the matrix is singular.
- Inverse: $AA^{-1}=I$. It exists exactly when the matrix is square and full rank. $(AB)^{-1}=B^{-1}A^{-1}$. Find it with $[A|I]\to[I|A^{-1}]$.
- The Invertible Matrix Theorem lists many equivalent conditions: $\det\ne0$, rank $n$, independent columns, only $\mathbf{x}=\mathbf{0}$ solves $A\mathbf{x}=\mathbf{0}$, and so on.
- In practice, solve, don't invert:
np.linalg.solve(A, b). - Norms: Frobenius (all entries), spectral (largest stretch $=\sigma_{\max}$), nuclear (sum of $\sigma$'s).
- Condition number $\kappa=\sigma_{\max}/\sigma_{\min}$ measures how errors are magnified. About $\log_{10}\kappa$ digits are lost. Judge sensitivity by $\kappa$, not by the determinant.
Cheat sheet
| Quantity | Formula / test | Meaning |
|---|---|---|
| Rank | number of pivots | number of independent directions |
| Trace | $\sum a_{ii}$, $\operatorname{tr}(AB)=\operatorname{tr}(BA)$ | sum of eigenvalues |
| Determinant (2×2) | $ad-bc$ | signed area scale factor |
| Singular? | $\det=0$, rank $\lt n$ | squashes space flat; no inverse |
| Inverse (2×2) | $\frac{1}{ad-bc}\begin{bmatrix} d&-b\\-c&a\end{bmatrix}$ | the undo button |
| $(AB)^{-1}$ | $B^{-1}A^{-1}$ | socks and shoes |
| Frobenius norm | $\sqrt{\sum a_{ij}^2}$ | length of the entry list |
| Spectral norm | $\sigma_{\max}$ | biggest stretch |
| Nuclear norm | $\sum\sigma_i$ | convex proxy for rank |
| Condition number | $\sigma_{\max}/\sigma_{\min}$ | error magnification |
import numpy as np
A = np.array([[1., 2., 3.], [2., 4., 6.], [1., 1., 1.]])
print(np.linalg.matrix_rank(A)) # 2 (rank-deficient: column 3 = -col 1 + 2 * col 2)
print(np.trace(A)) # 6.0
print(np.linalg.det(A)) # about 0 (maybe 0.0 or 1e-16: never test det == 0)
B = np.array([[4., 7.], [2., 6.]])
print(np.linalg.det(B)) # 10.0 (up to rounding)
print(np.linalg.inv(B)) # [[ 0.6 -0.7] [-0.2 0.4]]
print(B @ np.linalg.inv(B)) # the identity (up to rounding)
# trace is cyclic: tr(AB) == tr(BA) even for non-square factors
P = np.array([[1., 2., 0.], [0., 1., 3.]]); Q = np.array([[1., 0.], [2., 1.], [0., 1.]])
print(np.trace(P @ Q), np.trace(Q @ P)) # 9.0 9.0
# determinant = area scaling: push the unit square through M and measure its area
M = np.array([[3., 1.], [2., 4.]])
square = np.array([[0, 1, 1, 0], [0, 0, 1, 1]]) # corners, one per column
x, y = M @ square
area = 0.5 * abs(np.dot(x, np.roll(y, -1)) - np.dot(y, np.roll(x, -1))) # shoelace formula
print(area, np.linalg.det(M)) # 10.0 10.0
# norms
C = np.array([[1., 2.], [3., 4.]])
print(np.linalg.norm(C, 'fro')) # 5.477 (Frobenius)
print(np.linalg.norm(C, 2)) # 5.465 (spectral = largest singular value)
print(np.linalg.norm(C, 'nuc')) # 5.831 (nuclear = sum of singular values)
print(np.linalg.svd(C, compute_uv=False)) # [5.465 0.366]
# a nearly singular matrix: watch the condition number explode
for eps in [1e-1, 1e-4, 1e-8, 1e-12]:
N = np.array([[1., 1.], [1., 1. + eps]])
print(eps, np.linalg.cond(N)) # grows like 4 / eps
# solve vs inv: speed and accuracy
rng = np.random.default_rng(0)
n = 1000
G = rng.standard_normal((n, n)); b = rng.standard_normal(n)
import time
t = time.perf_counter(); x1 = np.linalg.solve(G, b); t1 = time.perf_counter() - t
t = time.perf_counter(); x2 = np.linalg.inv(G) @ b; t2 = time.perf_counter() - t
print(t1, t2) # solve is usually faster
print(np.linalg.norm(G @ x1 - b), np.linalg.norm(G @ x2 - b)) # and its residual is usually smaller
1. What is the rank of $\begin{bmatrix} 1 & 2 \\ 2 & 4 \end{bmatrix}$?
2. A $2\times2$ matrix has $\det = -3$. What does it do to areas?
3. If $A$ and $B$ are invertible, $(AB)^{-1}$ equals…
4. What is the best way to compute the solution of $A\mathbf{x}=\mathbf{b}$ in code?
solve uses LU elimination: less work, usually better accuracy, and it keeps structure. Forming the inverse is the wasteful route. There is no meaningful "b / A" for matrices.5. A matrix has condition number $10^6$. Which statement is right?
6. For $D = \begin{bmatrix} 3 & 0 \\ 0 & 4 \end{bmatrix}$, which pair is (Frobenius norm, spectral norm)?
Practice problems
A. Find the rank and the determinant of $\begin{bmatrix} 1 & 2 & 3 \\ 4 & 5 & 6 \\ 7 & 8 & 9 \end{bmatrix}$.
$R_2 \leftarrow R_2 - 4R_1 = [0,-3,-6]$ and $R_3 \leftarrow R_3 - 7R_1 = [0,-6,-12]$. Then $R_3 \leftarrow R_3 - 2R_2 = [0,0,0]$. Two pivots, so the rank is 2 and the matrix is singular: $\det = 0$. (Indeed column 3 $= 2\cdot$ column 2 $-$ column 1.)
B. Compute $\det\begin{bmatrix} 2 & 1 & 0 \\ 1 & 3 & 1 \\ 0 & 1 & 4 \end{bmatrix}$ by expanding along the first row.
$2(3\cdot4 - 1\cdot1) - 1(1\cdot4 - 1\cdot0) + 0 = 2\cdot11 - 4 = 18$.
C. Find the inverse of $\begin{bmatrix} 3 & 1 \\ 5 & 2 \end{bmatrix}$ and check it.
$\det = 6 - 5 = 1$. So $A^{-1} = \begin{bmatrix} 2 & -1 \\ -5 & 3 \end{bmatrix}$. Check: $\begin{bmatrix} 3 & 1 \\ 5 & 2 \end{bmatrix}\begin{bmatrix} 2 & -1 \\ -5 & 3 \end{bmatrix} = \begin{bmatrix} 6-5 & -3+3 \\ 10-10 & -5+6 \end{bmatrix} = I$ ✓.
D. $A$ is $4\times4$ with $\det A = 5$. Find $\det(2A)$, $\det(A^{-1})$ and $\det(A^\top A)$.
$\det(2A) = 2^4\cdot5 = 80$. $\det(A^{-1}) = 1/5$. $\det(A^\top A) = 5\cdot5 = 25$.
E. For $P = \begin{bmatrix} 1 & 2 & 0 \\ 0 & 1 & 3 \end{bmatrix}$ and $Q = \begin{bmatrix} 1 & 0 \\ 2 & 1 \\ 0 & 1 \end{bmatrix}$, compute $\operatorname{tr}(PQ)$ and $\operatorname{tr}(QP)$.
$PQ = \begin{bmatrix} 5 & 2 \\ 2 & 4 \end{bmatrix}$, trace 9. $QP = \begin{bmatrix} 1 & 2 & 0 \\ 2 & 5 & 3 \\ 0 & 1 & 3 \end{bmatrix}$, trace $1 + 5 + 3 = 9$. They are equal, although the matrices have different sizes.
F. Give the Frobenius, spectral and nuclear norms and the condition number of $\begin{bmatrix} 3 & 0 \\ 0 & 1 \end{bmatrix}$.
Singular values 3 and 1. Frobenius $= \sqrt{10} \approx 3.16$; spectral $= 3$; nuclear $= 4$; condition number $= 3/1 = 3$ (well-conditioned).
The Four Fundamental Subspaces
Every matrix hides four special sets of vectors. Together they tell you everything about what the matrix can do: what it can reach, what it destroys, and when an equation has an answer. Once you see the four pieces and how they fit together, a matrix stops being a block of numbers and becomes a picture.
- Name the four subspaces of a matrix: column space, null space, row space and left null space
- Find a basis for each one and say which space each lives in
- Count dimensions: rank $r$, nullity $n-r$, and the rank–nullity theorem
- See that the subspaces come in perpendicular pairs, and draw Strang's "big picture"
- Decide when $A\mathbf{x}=\mathbf{b}$ has a solution, and when that solution is unique
- Connect it all to machine learning: why some targets cannot be fitted exactly, and which parameters the data cannot see
A matrix has two sides. Throughout the chapter $A$ has $m$ rows and $n$ columns. It takes an input vector $\mathbf{x}$ with $n$ numbers and produces an output $A\mathbf{x}$ with $m$ numbers. So inputs live in $\mathbb{R}^n$ and outputs live in $\mathbb{R}^m$. Two of the subspaces live on the input side and two on the output side.
In all drawings of this chapter we keep the same colours: column space = blue, row space = orange, null space = purple, left null space = teal. (Colour only appears inside the pictures. The words always say which space is meant.)
The column space $C(A)$ core
Feed every possible input into a matrix and collect all the outputs. That collection is the column space. It answers the question: "Where can this matrix send things?"
Why "column"? Because of the recipe view from Chapter 1.4: $A\mathbf{x}$ is a mix of the columns of $A$, where the entries of $\mathbf{x}$ say how many scoops of each column to use. All possible mixes of the columns is exactly the span of the columns (see Chapter 1.3).
Picture a bag of coloured paints. The columns are the paints you own. The column space is every colour you can mix. If two of your paints are really the same colour, owning both adds nothing new.
Take $A=\begin{bmatrix} 1 & 2 \\ 2 & 4 \end{bmatrix}$. Its columns are $[1,2]$ and $[2,4]$. Look at what happens to a few inputs:
- $\mathbf{x}=[1,0]$ gives $A\mathbf{x}=1\cdot[1,2]+0\cdot[2,4]=[1,2]$.
- $\mathbf{x}=[0,1]$ gives $A\mathbf{x}=[2,4]=2\cdot[1,2]$.
- $\mathbf{x}=[3,-1]$ gives $3\cdot[1,2]-1\cdot[2,4]=[3,6]-[2,4]=[1,2]$.
Every output is a multiple of $[1,2]$. So $C(A)$ is the line through the origin in the direction $[1,2]$, even though $A$ is a $2\times2$ matrix and could "in principle" fill the whole plane. The second column was redundant.
Now take the $3\times2$ matrix $B=\begin{bmatrix} 1 & 0 \\ 0 & 1 \\ 1 & 1 \end{bmatrix}$. Then $B\mathbf{x}=[x_1,\,x_2,\,x_1+x_2]$. The outputs live in $\mathbb{R}^3$, but they all satisfy "third entry = first + second". They fill a flat plane inside 3D space, not all of it.
The column space of an $m\times n$ matrix $A$ is the set of all linear combinations of its columns:
$$C(A)=\{\,A\mathbf{x} : \mathbf{x}\in\mathbb{R}^n\,\}\;\subseteq\;\mathbb{R}^m.$$- It is the image (the set of all outputs) of the transformation $\mathbf{x}\mapsto A\mathbf{x}$.
- It is a subspace of $\mathbb{R}^m$ (the output side): it contains $\mathbf{0}$, and adding or scaling outputs gives outputs.
- A basis for it is given by the pivot columns of $A$, the columns that bring a new direction (you find them with elimination, Chapter 1.6). Its dimension is the rank $r$.
Why do we need it?
We need to know which outputs a matrix can make and which it never can. Without this, we cannot say whether an equation Ax = b has an answer at all.
Where is it used?
Linear regression (can the model reproduce the targets?), the span of the features in PCA, the set of outputs a neural-network layer can produce, and any check of whether a system of equations can be solved.
How is it used?
Run elimination and find the pivot columns, then take those columns from the original A as a basis. In code, np.linalg.matrix_rank gives its size and scipy.linalg.orth gives an orthonormal basis. Then ask: is my b inside it?
The column space lives in $\mathbb{R}^m$, the output space. Its vectors have $m$ entries (one per row of $A$). Beginners often think "column space lives in $\mathbb{R}^n$ because there are $n$ columns". The count of columns tells you how many ingredients there are, but each ingredient has $m$ entries.
Use the columns of $A$, not of its row-reduced form. Elimination changes the column space. Use elimination only to find which columns are the pivot ones, then take those columns from the original $A$.
Quick check: a $5\times3$ matrix has rank 2. Which space is $C(A)$ in, and what is its dimension?
$C(A)\subseteq\mathbb{R}^5$ (outputs have 5 entries), and its dimension is the rank, 2. So it is a flat plane sitting inside 5-dimensional space.
The null space $N(A)$ core
The column space asks "what can the matrix reach?" The null space asks the opposite: "what does the matrix destroy?" It is the set of inputs that get squashed all the way down to the zero vector.
Think of a shadow. A lamp straight above flattens a 3D object into a 2D shadow on the floor. Lift the object straight up, or lower it, and its shadow does not change at all. Those "invisible" movements are the null space of the flattening: the direction the shadow cannot see.
Two different inputs that differ by a null-space vector give exactly the same output. That is why the null space matters so much: it measures what information the matrix throws away.
Use the same $A=\begin{bmatrix} 1 & 2 \\ 2 & 4 \end{bmatrix}$. We want every $\mathbf{x}=[x_1,x_2]$ with $A\mathbf{x}=\mathbf{0}$.
- Write the two equations: $x_1+2x_2=0$ and $2x_1+4x_2=0$.
- The second equation is just twice the first, so we only really have one equation: $x_1=-2x_2$.
- Let $x_2=t$ be anything. Then $\mathbf{x}=[-2t,\;t]=t\,[-2,\;1]$.
- Check with $t=1$: $A[-2,1]=[1\cdot(-2)+2\cdot1,\;2\cdot(-2)+4\cdot1]=[0,0]$ ✓.
So $N(A)$ is the line through the origin in direction $[-2,1]$. Notice that $A[3,-1]=A[1,0]$: the inputs $[3,-1]$ and $[1,0]$ differ by $[2,-1]=-1\cdot[-2,1]$, a null vector, so they produce the same output $[1,2]$ (we saw this in the last section).
The null space (also called the kernel) of an $m\times n$ matrix $A$ is
$$N(A)=\{\,\mathbf{x}\in\mathbb{R}^n : A\mathbf{x}=\mathbf{0}\,\}\;\subseteq\;\mathbb{R}^n.$$- It lives on the input side, in $\mathbb{R}^n$.
- It is always a subspace: $\mathbf{0}$ is in it, and if $A\mathbf{x}=\mathbf{0}$ and $A\mathbf{y}=\mathbf{0}$ then $A(\mathbf{x}+\mathbf{y})=\mathbf{0}$ and $A(c\mathbf{x})=\mathbf{0}$.
- You find a basis by solving $A\mathbf{x}=\mathbf{0}$ with elimination: each free column (a column with no pivot) gives one basis vector.
- If the only solution is $\mathbf{x}=\mathbf{0}$, we write $N(A)=\{\mathbf{0}\}$ (the smallest possible subspace).
Why do we need it?
We need to know what a matrix throws away. If two different inputs give the same output, we can never recover which input was used, and the null space measures exactly that loss.
Where is it used?
Redundant features in regression, non-unique solutions of equations, the hidden directions of a flattening step such as a shadow, a projection or a compression step, and the free variables found in Gaussian elimination.
How is it used?
Solve Ax = 0 with elimination: every free column gives one basis vector. In code, scipy.linalg.null_space(A) returns an orthonormal basis. If it returns no columns, only x = 0 is destroyed and the matrix loses nothing.
The null space is never empty. The zero vector is always in it, because $A\mathbf{0}=\mathbf{0}$. "Trivial null space" means $N(A)=\{\mathbf{0}\}$, a set with one element, not an empty set.
The null space lives in $\mathbb{R}^n$ (inputs). Do not confuse it with the left null space, which lives on the output side.
Quick check: find the null space of $A=\begin{bmatrix}1&1\\1&-1\end{bmatrix}$.
Solve $x_1+x_2=0$ and $x_1-x_2=0$. Adding the equations gives $2x_1=0$, so $x_1=0$ and then $x_2=0$. Only $\mathbf{0}$ works, so $N(A)=\{\mathbf{0}\}$.
The other two: row space and left null space core
We have the two spaces made from $A$. The other two come from the transpose $A^\top$ (the matrix with rows and columns swapped, Chapter 1.4). Just apply the same two ideas to $A^\top$:
- The column space of $A^\top$ is made from the rows of $A$. We call it the row space.
- The null space of $A^\top$ is called the left null space of $A$.
Why "left"? Because $A^\top\mathbf{y}=\mathbf{0}$ is the same as $\mathbf{y}^\top A=\mathbf{0}^\top$, where $\mathbf{y}$ stands on the left of $A$. It is a recipe for the rows: it says "mix these multiples of the rows and the total cancels to zero".
This will be our running example for the whole chapter (it is $3\times3$ with $m=n=3$):
$$A=\begin{bmatrix} 1 & 0 & 1 \\ 0 & 1 & 1 \\ 2 & 1 & 3 \end{bmatrix}.$$Look closely. Row 3 is $2\cdot(\text{row }1)+1\cdot(\text{row }2)$, because $[2,0,2]+[0,1,1]=[2,1,3]$. And column 3 is column 1 plus column 2. The matrix has a redundant row and a redundant column.
- Eliminate (subtract $2\times$ row 1 and $1\times$ row 2 from row 3): we get $R=\begin{bmatrix} 1&0&1\\0&1&1\\0&0&0\end{bmatrix}$. There are 2 pivots (in columns 1 and 2) and one free column (column 3). So the rank is $r=2$.
- Column space: the pivot columns of the original $A$, namely $[1,0,2]$ and $[0,1,1]$. This is a plane in $\mathbb{R}^3$.
- Row space: the non-zero rows of $R$, namely $[1,0,1]$ and $[0,1,1]$. This is also a plane in $\mathbb{R}^3$ (a different one).
- Null space: set the free variable $x_3=1$. The rows of $R$ say $x_1+x_3=0$ and $x_2+x_3=0$, so $x_1=-1$, $x_2=-1$. Basis: $[-1,-1,1]$. Check: $A[-1,-1,1]=[-1+0+1,\;0-1+1,\;-2-1+3]=[0,0,0]$ ✓.
- Left null space: find $\mathbf{y}$ with $\mathbf{y}^\top A=\mathbf{0}^\top$. Since row 3 $=2\,$row 1 $+$ row 2, we can take $\mathbf{y}=[-2,-1,1]$: $-2[1,0,1]-1[0,1,1]+1[2,1,3]=[0,0,0]$ ✓.
For an $m\times n$ matrix $A$, the four fundamental subspaces are:
| Subspace | Symbol | Definition | Lives in | Basis from |
|---|---|---|---|---|
| Column space | $C(A)$ | all $A\mathbf{x}$ | $\mathbb{R}^m$ | pivot columns of $A$ |
| Null space | $N(A)$ | all $\mathbf{x}$ with $A\mathbf{x}=\mathbf{0}$ | $\mathbb{R}^n$ | one vector per free column |
| Row space | $C(A^\top)$ | all combinations of the rows | $\mathbb{R}^n$ | non-zero rows of $R$ |
| Left null space | $N(A^\top)$ | all $\mathbf{y}$ with $A^\top\mathbf{y}=\mathbf{0}$ | $\mathbb{R}^m$ | solve for $A^\top$ the same way |
Two live on the input side $\mathbb{R}^n$ (row space and null space). Two live on the output side $\mathbb{R}^m$ (column space and left null space).
Why do we need it?
The column and null spaces describe the output side and the input side, but a matrix has two more pieces, one for each side. Without them the picture of what A does is incomplete.
Where is it used?
The row space is what PCA and the SVD work in (the directions the data actually uses). The left null space holds the constraints on the outputs and the residuals of least squares.
How is it used?
Take the non-zero rows of the reduced form R as a row-space basis. Get the left null space by finding the null space of the transpose: null_space(A.T). Check each basis by multiplying A (or Aᵀ) by it and expecting zeros.
Rows or columns, which one for which space? A quick rule: column space and left null space are made of vectors with $m$ entries (one per row of $A$). Row space and null space are made of vectors with $n$ entries (one per column of $A$).
For the row space, take rows of the reduced matrix $R$ (elimination only mixes rows, so it keeps the row space). For the column space, go back to the original columns.
Quick check: $A$ is $4\times6$. List the four spaces and the $\mathbb{R}^k$ each lives in.
Here $m=4$ and $n=6$. Column space: $\mathbb{R}^4$. Left null space: $\mathbb{R}^4$. Row space: $\mathbb{R}^6$. Null space: $\mathbb{R}^6$.
Rank, nullity and the rank–nullity theorem core
A matrix takes an $n$-dimensional input and produces an $m$-dimensional output. Think of each side as a budget of directions.
- Some input directions survive: the matrix turns each of them into a genuinely different output direction. The number of those is the rank $r$.
- The remaining input directions get squashed to zero. That is the null space. The number of those is the nullity, $n-r$.
- On the output side, the matrix only reaches $r$ directions. The $m-r$ directions it never reaches form the left null space.
Every direction is accounted for. Nothing is lost in the accounting: $r$ survive plus $n-r$ squashed equals all $n$ input directions.
Our running example has $m=n=3$ and after elimination it has $r=2$ pivots.
- Column space: dimension $r=2$ (a plane in $\mathbb{R}^3$).
- Row space: dimension $r=2$ (another plane in $\mathbb{R}^3$).
- Null space: dimension $n-r=3-2=1$ (a line, spanned by $[-1,-1,1]$).
- Left null space: dimension $m-r=3-2=1$ (a line, spanned by $[-2,-1,1]$).
Now try a bigger one without looking at the numbers: $A$ is $5\times8$ and has rank $3$. Then $\dim C(A)=3$, $\dim C(A^\top)=3$, $\dim N(A)=8-3=5$ and $\dim N(A^\top)=5-3=2$.
Why the row rank equals the column rank. Elimination only mixes rows, so it does not change the row space. After elimination, the row space has the $r$ non-zero rows of $R$ as a basis, so its dimension is $r$. At the same time the column space has the $r$ pivot columns as a basis, so its dimension is also $r$. One number, $r$, counts both.
For an $m\times n$ matrix $A$:
- The rank is $r=\dim C(A)=\dim C(A^\top)$ (the number of pivots). Always $r\le\min(m,n)$.
- The nullity is $\dim N(A)=n-r$.
- The left nullity is $\dim N(A^\top)=m-r$.
Rank–nullity theorem:
$$\underbrace{\operatorname{rank}(A)}_{r}+\underbrace{\operatorname{nullity}(A)}_{n-r}=n\qquad(\text{the number of columns}).$$Why: each column of $A$ is either a pivot column ($r$ of them) or a free column ($n-r$ of them), and each free column produces exactly one basis vector of the null space. Apply the same theorem to $A^\top$, which has $m$ columns and the same rank $r$, to get $r+(m-r)=m$.
| Space | Lives in | Dimension |
|---|---|---|
| Row space $C(A^\top)$ | $\mathbb{R}^n$ | $r$ |
| Null space $N(A)$ | $\mathbb{R}^n$ | $n-r$ |
| Column space $C(A)$ | $\mathbb{R}^m$ | $r$ |
| Left null space $N(A^\top)$ | $\mathbb{R}^m$ | $m-r$ |
Why do we need it?
We need numbers that say how big each subspace is. Dimensions let us predict, before solving anything, how many solutions exist and how much information the matrix keeps.
Where is it used?
Checking if a data matrix has redundant columns (rank less than the number of features), choosing how many components to keep in low-rank compression, and counting free parameters of a model.
How is it used?
Compute the rank r with np.linalg.matrix_rank(A). Then read off the others: the nullity is n − r, the left nullity is m − r. If r is smaller than the number of columns, some inputs are being squashed.
Rank is not "number of non-zero rows of $A$". It is the number of independent rows (equivalently, columns). A matrix can have no zero row and still have a low rank: in the running example no row is zero, yet the rank is 2 because row 3 is made from rows 1 and 2.
Full rank means $r=\min(m,n)$, the biggest possible. A matrix can have full rank without being square, e.g. a tall $5\times3$ matrix with $r=3$.
Quick check: a $7\times4$ matrix has rank 4. What are the dimensions of its four subspaces?
$\dim C(A)=4$, $\dim C(A^\top)=4$, $\dim N(A)=4-4=0$ (only the zero vector), $\dim N(A^\top)=7-4=3$.
The subspaces come in perpendicular pairs core
Recall that two vectors are orthogonal (perpendicular) when their dot product is zero (Chapter 1.2). Two subspaces are orthogonal when every vector of one is perpendicular to every vector of the other.
Here is the surprising and beautiful fact. Look at what "$A\mathbf{x}=\mathbf{0}$" really says. The first entry of $A\mathbf{x}$ is (row 1)$\,\cdot\,\mathbf{x}$, the second is (row 2)$\,\cdot\,\mathbf{x}$, and so on. So $A\mathbf{x}=\mathbf{0}$ means: $\mathbf{x}$ is perpendicular to every row of $A$. And then it is also perpendicular to every mix of rows. That is the whole row space!
So the null space (inputs destroyed) sits at a right angle to the row space (inputs that "matter"). The same argument on the output side shows the column space is perpendicular to the left null space.
Running example: row space basis $[1,0,1]$, $[0,1,1]$; null space basis $[-1,-1,1]$. Take all the dot products:
- $[1,0,1]\cdot[-1,-1,1]=-1+0+1=0$ ✓
- $[0,1,1]\cdot[-1,-1,1]=0-1+1=0$ ✓
Output side: column space basis $[1,0,2]$, $[0,1,1]$; left null space basis $[-2,-1,1]$:
- $[1,0,2]\cdot[-2,-1,1]=-2+0+2=0$ ✓
- $[0,1,1]\cdot[-2,-1,1]=0-1+1=0$ ✓
It is enough to test the basis vectors: if each basis vector of one space is perpendicular to each basis vector of the other, every combination is too (the dot product is linear).
Two subspaces $V$ and $W$ are orthogonal, written $V\perp W$, if $\mathbf{v}\cdot\mathbf{w}=0$ for every $\mathbf{v}\in V$ and every $\mathbf{w}\in W$.
Theorem. For every matrix $A$:
$$C(A^\top)\perp N(A)\quad(\text{in }\mathbb{R}^n),\qquad C(A)\perp N(A^\top)\quad(\text{in }\mathbb{R}^m).$$Proof, in one line each. If $\mathbf{x}\in N(A)$ then every row $\mathbf{r}_i$ of $A$ has $\mathbf{r}_i\cdot\mathbf{x}=0$, so $(c_1\mathbf{r}_1+\dots+c_m\mathbf{r}_m)\cdot\mathbf{x}=0$. If $\mathbf{y}\in N(A^\top)$ then $A^\top\mathbf{y}=\mathbf{0}$, which says every column $\mathbf{a}_j$ has $\mathbf{a}_j\cdot\mathbf{y}=0$.
Why do we need it?
We need a clean way to separate what the matrix sees from what it ignores. Perpendicular pairs guarantee the two parts never overlap and never interfere.
Where is it used?
Least squares (the error is perpendicular to the model's reachable outputs), splitting a weight vector into a part the data determines and a part it cannot see, and the proof behind the SVD.
How is it used?
To test it in practice, compute bases for two subspaces and check that every dot product between them is zero, for example np.allclose(row.T @ nul, 0). Use it as a quick sanity check on any subspace code.
Orthogonal subspaces are not the same as "planes at a right angle" in a room. Two walls of a room meet at a right angle, but they are not orthogonal subspaces: they share a line (the corner), and a vector along that line would have to be perpendicular to itself. In $\mathbb{R}^3$ a plane can only be orthogonal to a line (or to $\{\mathbf{0}\}$).
Also, the row space is generally not equal to the column space, even if both have dimension $r$. They live in different spaces, unless $m=n$, and even then they are usually different planes.
Quick check: $A=\begin{bmatrix}1&2&3\end{bmatrix}$. Is $[2,-1,0]$ in the null space? Is it perpendicular to the row?
$A[2,-1,0]=2-2+0=0$, so yes it is in the null space. And that is exactly the same calculation as the dot product of the row $[1,2,3]$ with $[2,-1,0]$: both give $0$. "In the null space" and "perpendicular to all rows" are the same statement.
Orthogonal complements core
Take a flat table top in a room. Every vector sticking straight up from the table is perpendicular to the table. All the "straight up and straight down" vectors together form a line. That line is the orthogonal complement of the table: everything that is perpendicular to the whole table.
The table and its complement are the two halves of a split of the room: sideways directions on the table, plus the one direction that is left over, up. Together they use up all 3 dimensions.
For the four subspaces this means more than "perpendicular": the null space is not just some subspace perpendicular to the row space, it is all of what is perpendicular to it.
In $\mathbb{R}^3$ let $V$ be the line spanned by $\mathbf{d}=[1,1,1]$.
- A vector $\mathbf{w}=[a,b,c]$ is perpendicular to $V$ when $\mathbf{d}\cdot\mathbf{w}=a+b+c=0$.
- That is one equation in three unknowns, so there are $3-1=2$ free choices: it is a plane. A basis is $[1,-1,0]$ and $[0,1,-1]$.
- So $V^\perp$ is a plane, and $\dim V+\dim V^\perp=1+2=3$.
Now go back: what is perpendicular to that plane? Only the line spanned by $[1,1,1]$. So $(V^\perp)^\perp=V$.
The orthogonal complement of a subspace $V\subseteq\mathbb{R}^n$ is
$$V^\perp=\{\,\mathbf{w}\in\mathbb{R}^n:\ \mathbf{w}\cdot\mathbf{v}=0\ \text{ for all }\mathbf{v}\in V\,\}.$$It is itself a subspace, and:
- $\dim V+\dim V^\perp=n$.
- $V\cap V^\perp=\{\mathbf{0}\}$ (the only vector perpendicular to itself is $\mathbf{0}$, since $\mathbf{v}\cdot\mathbf{v}=\|\mathbf{v}\|^2$).
- $(V^\perp)^\perp=V$.
- Every vector $\mathbf{x}\in\mathbb{R}^n$ splits in exactly one way as $\mathbf{x}=\mathbf{v}+\mathbf{w}$ with $\mathbf{v}\in V$ and $\mathbf{w}\in V^\perp$ (more in Chapter 1.9).
For a matrix $A$ the four subspaces are two complementary pairs:
$$N(A)=C(A^\top)^\perp,\qquad N(A^\top)=C(A)^\perp .$$The dimensions agree: $r+(n-r)=n$ and $r+(m-r)=m$.
Why do we need it?
We often need 'everything that is left over after removing a subspace'. The orthogonal complement gives that leftover space, and it always fits together with the original to fill the whole space.
Where is it used?
Splitting data into signal and noise in PCA, constructing the residual space in regression, and building a full orthonormal basis from just a few given directions.
How is it used?
To find V⊥ for the span of the columns of A, compute the null space of Aᵀ (every vector perpendicular to all columns). Its dimension is m minus the dimension of V (the columns live in $\mathbb{R}^m$), which gives a quick check.
A complement is relative to the whole space: the complement of a line in $\mathbb{R}^3$ is a plane, but the complement of the same line inside $\mathbb{R}^2$ is another line. Always say which space you are in.
"Perpendicular subspaces" is weaker than "complements". Two perpendicular lines in $\mathbb{R}^3$ are orthogonal subspaces, but they are not complements because $1+1\ne3$: there is a whole direction missing.
Quick check: $W$ is a 2D plane in $\mathbb{R}^5$. What is $\dim W^\perp$?
$\dim W^\perp=5-2=3$. The complement is a 3-dimensional subspace.
The big picture core
Put everything on one page. The matrix $A$ is a machine that takes things from the input space $\mathbb{R}^n$ (left) to the output space $\mathbb{R}^m$ (right).
The input space splits into two perpendicular pieces:
- the row space: the part of the input the machine actually listens to. Every vector there goes to a different, non-zero place in the output;
- the null space: the part the machine ignores. Everything there is sent to $\mathbf{0}$.
The output space also splits into two perpendicular pieces: the column space (everything the machine can produce) and the left null space (what it can never produce). This diagram is called Strang's big picture, after the teacher Gilbert Strang who made it famous.
For our running example ($m=n=3$, $r=2$):
- Input space $\mathbb{R}^3$ = row space (2D plane) $\oplus$ null space (1D line, perpendicular to the plane).
- Output space $\mathbb{R}^3$ = column space (2D plane) $\oplus$ left null space (1D line, perpendicular to that plane).
- $A$ takes the row-space plane one-to-one onto the column-space plane. It takes the null line to the single point $\mathbf{0}$.
Here is the nice surprise: the part of the input that matters (the row space) and the part of the output that is reachable (the column space) always have the same dimension $r$, and $A$ matches them up perfectly, with no repeats and nothing missing. Everything wasteful or unreachable is in the other two spaces.
The big picture for an $m\times n$ matrix of rank $r$:
- $\mathbb{R}^n=C(A^\top)\oplus N(A)$ with dimensions $r+(n-r)=n$, and $C(A^\top)\perp N(A)$.
- $\mathbb{R}^m=C(A)\oplus N(A^\top)$ with dimensions $r+(m-r)=m$, and $C(A)\perp N(A^\top)$.
- $A$ maps $C(A^\top)$ onto $C(A)$ in a one-to-one way, and maps $N(A)$ to $\{\mathbf{0}\}$.
The symbol $\oplus$ ("direct sum") means: every vector splits in exactly one way into a piece from each subspace.
Why do we need it?
The four subspaces are easy to mix up. One picture that holds all four, with their dimensions and right angles, lets you recall what any matrix does in a few seconds.
Where is it used?
Teaching and reasoning about linear regression, the SVD, pseudoinverses and conjugate-gradient-type solvers; it is the standard mental map of linear algebra.
How is it used?
Compute the rank r. Then sketch the row space (r), null space (n − r), column space (r) and left null space (m − r) with their right angles. Use it to decide at once if Ax = b is solvable and unique.
The pictures in the books are sketches. The row space and the null space are not two separate "boxes": they are subspaces that cross at the origin at right angles (like a floor and a vertical pole). The box drawing makes the dimension counts easy to see.
Notice what the diagram does not say: it does not say the row space equals the column space. They live in different spaces. $A$ simply gives a perfect matching between them.
Quick check: sketch from memory the boxes for a $2\times4$ matrix of rank 2. What are the four dimensions?
Input space $\mathbb{R}^4$: row space dimension 2 and null space dimension $4-2=2$. Output space $\mathbb{R}^2$: column space dimension 2 (the whole $\mathbb{R}^2$) and left null space dimension $2-2=0$ (just $\{\mathbf{0}\}$).
When does $A\mathbf{x}=\mathbf{b}$ have an answer? core
Solving $A\mathbf{x}=\mathbf{b}$ means: "Can I make $\mathbf{b}$ by mixing the columns of $A$?" (that was the recipe view). The set of everything you can make is the column space. So:
- If $\mathbf{b}$ is inside the column space, a solution exists.
- If $\mathbf{b}$ is outside, there is no solution. No amount of mixing hits it.
And when a solution exists, is it the only one? Remember the null space is the set of inputs that do nothing. If $\mathbf{x}_p$ is a solution, then $\mathbf{x}_p+\mathbf{n}$ is also a solution for any $\mathbf{n}$ in the null space, because the extra part is destroyed. So the answer is unique exactly when the null space has nothing but $\mathbf{0}$.
Take $A=\begin{bmatrix} 1 & 2 \\ 1 & 2 \end{bmatrix}$. Its column space is the line through $[1,1]$. Its null space is the line through $[-2,1]$. Its left null space is the line through $[-1,1]$ (since $-1\cdot\text{row}_1+1\cdot\text{row}_2=\mathbf{0}$).
- $\mathbf{b}=[3,3]$. This is $3\cdot[1,1]$, so it is in the column space: solvable. One solution is $\mathbf{x}=[1,1]$, because $1+2=3$ in both rows ✓.
- Other solutions: add any multiple of $[-2,1]$. With $t=1$: $\mathbf{x}=[-1,2]$, and $-1+4=3$ ✓. So there are infinitely many solutions (a whole line), because $N(A)$ is not $\{\mathbf{0}\}$.
- $\mathbf{b}=[3,1]$. It is not a multiple of $[1,1]$, so it is outside the column space: no solution. (The two equations would say $x_1+2x_2=3$ and $x_1+2x_2=1$ at once.)
- A second test for step 3, using the left null space $\mathbf{y}=[-1,1]$: $\mathbf{y}\cdot\mathbf{b}=-3+1=-2\ne0$. For $\mathbf{b}=[3,3]$ we get $-3+3=0$ ✓. A vector is in $C(A)$ exactly when it is perpendicular to the left null space.
Existence. $A\mathbf{x}=\mathbf{b}$ has a solution $\iff$ $\mathbf{b}\in C(A)$ $\iff$ $\mathbf{y}^\top\mathbf{b}=0$ for every $\mathbf{y}\in N(A^\top)$ (that is, $\mathbf{b}\perp N(A^\top)$, because $C(A)=N(A^\top)^\perp$).
All solutions. If $\mathbf{x}_p$ is one solution, then the complete set of solutions is $\{\mathbf{x}_p+\mathbf{n}:\ \mathbf{n}\in N(A)\}$ (a shifted copy of the null space).
Uniqueness. A solution (when it exists) is unique $\iff$ $N(A)=\{\mathbf{0}\}$ $\iff$ $r=n$ (full column rank).
| Shape of $A$ | Rank | $N(A)$ | $N(A^\top)$ | Solutions of $A\mathbf{x}=\mathbf{b}$ |
|---|---|---|---|---|
| square | $r=m=n$ | $\{\mathbf{0}\}$ | $\{\mathbf{0}\}$ | exactly 1, for every $\mathbf{b}$ |
| tall ($m\gt n$) | $r=n$ | $\{\mathbf{0}\}$ | not trivial | 0 or 1 |
| wide ($m\lt n$) | $r=m$ | not trivial | $\{\mathbf{0}\}$ | infinitely many, for every $\mathbf{b}$ |
| any | $r\lt m$ and $r\lt n$ | not trivial | not trivial | 0 or infinitely many |
Why do we need it?
Before spending time solving Ax = b we want to know if an answer exists, and if it is the only one. The subspaces answer both questions without solving anything.
Where is it used?
Linear regression (a target outside the column space means no exact fit, so we use least squares), solving circuits and balance equations, and checking if a model's parameters can be identified from the data.
How is it used?
Ask two questions. Is b in the column space (test it with matrix_rank of A against A with b added, or with the left null space)? Is the null space just {0}? Together they tell you: none, one or infinitely many.
"Solvable" and "unique" are two separate questions. Solvable depends on the column space ($\mathbf{b}$ inside or outside). Unique depends on the null space (trivial or not). A matrix can have a solution that is not unique (wide full-row-rank matrices), or can be unique when it exists but not always exist (tall full-column-rank matrices).
The solutions of $A\mathbf{x}=\mathbf{b}$ do not form a subspace when $\mathbf{b}\ne\mathbf{0}$: they miss the origin. They form a shifted copy of the null space (an "affine" set, Chapter 1.3).
Quick check: $A$ is $3\times5$ with rank 3. Is $A\mathbf{x}=\mathbf{b}$ solvable for every $\mathbf{b}$? Is the solution unique?
The rank is 3 = $m$, so the column space is all of $\mathbb{R}^3$ and every $\mathbf{b}$ is reachable: yes, always solvable. The null space has dimension $5-3=2$, which is not zero, so solutions are never unique: there is a whole plane of them.
Machine learning link: the directions the data cannot see
Suppose a model predicts a house price from two features: the floor area in square metres and the same floor area in half-square-metres (a silly, redundant duplicate). Two features, but they say the same thing: the second is always twice the first.
You can then trade weight between the two features without changing a single prediction. Raise one weight and lower the other by the right amount, and every prediction stays exactly the same. The data is blind to that change. Those blind changes are the null space of the data matrix.
The result: the weights are not uniquely determined. Many different weight vectors fit the data equally well. (People call this non-identifiability, or in regression, perfect multicollinearity.)
Three houses, two features $[\text{area},\ 2\cdot\text{area}]$:
$$X=\begin{bmatrix} 1 & 2 \\ 2 & 4 \\ 3 & 6 \end{bmatrix},\qquad X\mathbf{w}=(w_1+2w_2)\begin{bmatrix}1\\2\\3\end{bmatrix}.$$- The predictions only depend on the single number $w_1+2w_2$.
- Weights $\mathbf{w}=[1,1]$ give $w_1+2w_2=3$, so the predictions are $[3,6,9]$.
- Weights $[3,0]$ also give $3$. So do $[-1,2]$ and $[5,-1]$: predictions $[3,6,9]$ every time.
- The difference of two such weight vectors, like $[1,1]-[3,0]=[-2,1]$, is a null vector: $X[-2,1]=[-2+2,\ -4+4,\ -6+6]=\mathbf{0}$ ✓.
$N(X)$ is the line through $[-2,1]$. The "true" weights can be recovered only up to adding multiples of that vector.
For a data matrix $X$ ($m$ examples, $n$ features) and any weights $\mathbf{w}$:
$$X(\mathbf{w}+\mathbf{n})=X\mathbf{w}\quad\text{for every }\mathbf{n}\in N(X).$$The weights are uniquely determined by the predictions exactly when $N(X)=\{\mathbf{0}\}$, i.e. when the columns (features) are independent and the rank is $n$. If the rank is less than $n$, some weight directions are invisible to the data.
Why do we need it?
Some parameter changes leave every prediction unchanged. We need to recognise this, otherwise we may think a model has been learned uniquely when many different weights fit equally well.
Where is it used?
Regression with collinear or duplicated features, one-hot encodings that sum to one, and over-parameterised neural networks that have more weights than data. Ridge regression and weight decay exist to deal with it.
How is it used?
Compute the rank of the data matrix X. If it is smaller than the number of features, null_space(X) lists the invisible directions. Then remove redundant features or add regularisation to pick one solution.
Quick check: $X$ has 100 examples and 150 features. Can the weights be unique?
No. The rank is at most $\min(100,150)=100$, so the nullity is at least $150-100=50$. There are at least 50 independent directions of weights the data cannot see.
Recap, cheat sheet and practice
- An $m\times n$ matrix has four fundamental subspaces: column space $C(A)$ and left null space $N(A^\top)$ in $\mathbb{R}^m$ (outputs); row space $C(A^\top)$ and null space $N(A)$ in $\mathbb{R}^n$ (inputs).
- The column space is every output $A\mathbf{x}$. The null space is every input that is squashed to $\mathbf{0}$.
- The rank $r$ is the number of pivots: $\dim C(A)=\dim C(A^\top)=r$. The null space has dimension $n-r$ and the left null space has dimension $m-r$. Rank–nullity: $r+(n-r)=n$.
- The spaces are perpendicular in pairs, and are complements: $N(A)=C(A^\top)^\perp$ and $N(A^\top)=C(A)^\perp$.
- Big picture: $A$ matches the row space with the column space one-to-one, and sends the null space to $\mathbf{0}$. Nothing in the left null space, except $\mathbf{0}$, is ever an output.
- $A\mathbf{x}=\mathbf{b}$ is solvable iff $\mathbf{b}\in C(A)$. A solution is unique iff $N(A)=\{\mathbf{0}\}$. All solutions are $\mathbf{x}_p+N(A)$.
- In ML: a target outside $C(X)$ means no exact fit (so we project: least squares). The null space of $X$ holds the parameter directions the data cannot see.
Cheat sheet
| Subspace | In | Dimension | Basis | Perpendicular to |
|---|---|---|---|---|
| Column space $C(A)$ | $\mathbb{R}^m$ | $r$ | pivot columns of $A$ | $N(A^\top)$ |
| Null space $N(A)$ | $\mathbb{R}^n$ | $n-r$ | one vector per free column | $C(A^\top)$ |
| Row space $C(A^\top)$ | $\mathbb{R}^n$ | $r$ | non-zero rows of $R$ | $N(A)$ |
| Left null space $N(A^\top)$ | $\mathbb{R}^m$ | $m-r$ | null space of $A^\top$ | $C(A)$ |
| Question | Answer |
|---|---|
| Is $A\mathbf{x}=\mathbf{b}$ solvable? | $\mathbf{b}\in C(A)$ (equivalently $\mathbf{y}^\top\mathbf{b}=0$ for all $\mathbf{y}\in N(A^\top)$) |
| Is the solution unique? | $N(A)=\{\mathbf{0}\}$ (equivalently $r=n$) |
| What are all the solutions? | $\mathbf{x}_p+\mathbf{n}$ with $\mathbf{n}\in N(A)$ |
import numpy as np
from scipy.linalg import null_space, orth
A = np.array([[1, 0, 1],
[0, 1, 1],
[2, 1, 3]], dtype=float)
m, n = A.shape
r = np.linalg.matrix_rank(A) # rank = 2
# orthonormal bases for the four subspaces
col = orth(A) # C(A) in R^m, shape (3, 2)
nul = null_space(A) # N(A) in R^n, shape (3, 1)
row = orth(A.T) # C(A^T) in R^n, shape (3, 2)
left = null_space(A.T) # N(A^T) in R^m, shape (3, 1)
print(col.shape[1], nul.shape[1], row.shape[1], left.shape[1]) # 2 1 2 1 = r, n-r, r, m-r
print(r + nul.shape[1] == n) # rank-nullity theorem: True
# orthogonality of the subspaces: all products should be (almost) zero
print(np.allclose(row.T @ nul, 0)) # row space ⊥ null space -> True
print(np.allclose(col.T @ left, 0)) # column space ⊥ left null space -> True
print(np.allclose(A @ nul, 0)) # A really destroys the null space -> True
# solvability: is b in the column space?
def solvable(A, b):
x, *_ = np.linalg.lstsq(A, b, rcond=None) # best attempt at A x = b
return np.allclose(A @ x, b)
b_in = A[:, 0] + A[:, 1] # a mix of the columns: [1, 1, 3]
b_out = np.array([1.0, 0.0, 0.0])
print(solvable(A, b_in), solvable(A, b_out)) # True False
print(left.T @ b_in, left.T @ b_out) # ~0 for b_in, non-zero for b_out (the left-null test)
# a different way to get the same number: rank from the singular values
print(np.sum(np.linalg.svd(A, compute_uv=False) > 1e-10)) # 2
1. $A$ is a $4\times6$ matrix of rank 3. What is $\dim N(A)$?
2. $A$ is $3\times5$. The null space of $A$ is a subspace of…
3. Which pair of subspaces is always orthogonal?
4. $A\mathbf{x}=\mathbf{b}$ has at least one solution exactly when…
5. $A$ is a tall $5\times3$ matrix with rank 3. What can you say about $A\mathbf{x}=\mathbf{b}$?
6. In regression with data matrix $X$ and targets $\mathbf{y}$, what does "$\mathbf{y}$ is not in the column space of $X$" mean?
Practice problems
A. Find a basis and the dimension of all four subspaces of $A=\begin{bmatrix}1&2\\3&6\end{bmatrix}$.
Row 2 is $3\times$ row 1, so the rank is $r=1$. Column space: the first column $[1,3]$ (dimension 1). Row space: $[1,2]$ (dimension 1). Null space: $x_1+2x_2=0$ gives $[-2,1]$ (dimension $2-1=1$). Left null space: $\mathbf{y}^\top A=\mathbf{0}$ means $y_1+3y_2=0$, giving $[-3,1]$ (dimension $2-1=1$). Check: $[1,2]\cdot[-2,1]=0$ ✓ and $[1,3]\cdot[-3,1]=0$ ✓.
B. A $3\times4$ matrix has rank 2. Give the dimensions of the four subspaces and say where each lives.
$C(A)$: dimension 2 in $\mathbb{R}^3$. $N(A^\top)$: dimension $3-2=1$ in $\mathbb{R}^3$. $C(A^\top)$: dimension 2 in $\mathbb{R}^4$. $N(A)$: dimension $4-2=2$ in $\mathbb{R}^4$.
C. Let $B=\begin{bmatrix}1&0\\0&1\\1&1\end{bmatrix}$. Show that $B\mathbf{x}=[1,1,1]$ has no solution, using the left null space.
The left null space of $B$: solve $B^\top\mathbf{y}=\mathbf{0}$, i.e. $y_1+y_3=0$ and $y_2+y_3=0$. Taking $y_3=-1$ gives $\mathbf{y}=[1,1,-1]$. For $\mathbf{b}=[1,1,1]$: $\mathbf{y}\cdot\mathbf{b}=1+1-1=1\ne0$, so $\mathbf{b}\notin C(B)$. (Directly: $B\mathbf{x}=[x_1,x_2,x_1+x_2]$ would need $1+1=1$.)
D. $A$ is $3\times5$. Can $N(A)=\{\mathbf{0}\}$? If $A\mathbf{x}=\mathbf{b}$ is solvable, how many solutions are there?
No. The rank is at most 3, so the nullity is at least $5-3=2$. The null space always has a plane of vectors. So if the equation is solvable it has infinitely many solutions (a shifted copy of that null space).
E. For $A=\begin{bmatrix}1&2&3\\4&5&6\end{bmatrix}$ verify that the null space is perpendicular to the row space.
Eliminating gives $R=\begin{bmatrix}1&0&-1\\0&1&2\end{bmatrix}$, with a free column 3. Setting $x_3=1$: $x_1=1$, $x_2=-2$, so $N(A)$ is spanned by $[1,-2,1]$. Row space basis: $[1,2,3]$ and $[4,5,6]$. Dot products: $1-4+3=0$ ✓ and $4-10+6=0$ ✓. The rank is 2, and the nullity is $3-2=1$.
Orthogonality
Perpendicular directions are the nicest directions to work with. They do not interfere with each other, coordinates become one dot product, and "the closest point" becomes a shadow. This chapter turns that idea into tools: orthonormal bases, orthogonal matrices, projections, Gram–Schmidt and the QR decomposition.
- Use orthogonal and orthonormal sets, and see why orthogonal vectors are always independent
- Find coordinates in an orthonormal basis with a single dot product: $c_i=\mathbf{q}_i\cdot\mathbf{x}$
- Recognise orthogonal matrices ($Q^\top Q=I$): rotations and reflections that keep lengths and angles
- Project a vector onto a line and onto a subspace, and split $\mathbf{x}=\mathbf{x}_\parallel+\mathbf{x}_\perp$
- Turn any basis into an orthonormal one with Gram–Schmidt, and write $A=QR$
Orthogonal and orthonormal sets core
Think of the corner of a room. The two walls and the floor meet at right angles along three edges: left-right, front-back and up-down. Moving along one edge tells you nothing about how far you moved along another. These three directions are orthogonal: every pair is perpendicular.
A set of vectors is orthogonal when every pair in it is perpendicular. If, on top of that, each vector has length 1, the set is orthonormal ("ortho" for perpendicular, "normal" for length 1). We met perpendicular vectors in Chapter 1.2; now we use whole sets of them.
Here is a reassuring fact. Perpendicular directions can never be redundant. If you have walls, floor and nothing else, you can't make "up" out of "left" and "forward". So orthogonal vectors are automatically independent.
Let $\mathbf{v}_1=[1,1,0]$, $\mathbf{v}_2=[1,-1,0]$, $\mathbf{v}_3=[0,0,2]$.
- $\mathbf{v}_1\cdot\mathbf{v}_2=1-1+0=0$ ✓
- $\mathbf{v}_1\cdot\mathbf{v}_3=0+0+0=0$ ✓
- $\mathbf{v}_2\cdot\mathbf{v}_3=0+0+0=0$ ✓
All three pairs are perpendicular, so this is an orthogonal set. It is not orthonormal: the lengths are $\sqrt2$, $\sqrt2$ and $2$. Divide each vector by its length:
$$\mathbf{q}_1=\tfrac{1}{\sqrt2}[1,1,0],\quad \mathbf{q}_2=\tfrac{1}{\sqrt2}[1,-1,0],\quad \mathbf{q}_3=[0,0,1].$$Now every length is 1, and dividing by a positive number does not change a right angle, so the set is orthonormal.
Why independence is automatic. Suppose $c_1\mathbf{v}_1+c_2\mathbf{v}_2+c_3\mathbf{v}_3=\mathbf{0}$. Take the dot product of both sides with $\mathbf{v}_1$. The terms with $\mathbf{v}_2$ and $\mathbf{v}_3$ vanish, leaving $c_1\,(\mathbf{v}_1\cdot\mathbf{v}_1)=0$, i.e. $2c_1=0$. So $c_1=0$. The same trick with $\mathbf{v}_2,\mathbf{v}_3$ gives $c_2=c_3=0$. The only way to make zero is the trivial one: that is the definition of independence.
A set of non-zero vectors $\mathbf{v}_1,\dots,\mathbf{v}_k$ is orthogonal if $\mathbf{v}_i\cdot\mathbf{v}_j=0$ for all $i\ne j$. It is orthonormal if in addition $\|\mathbf{v}_i\|=1$. Both conditions in one line, for orthonormal vectors $\mathbf{q}_i$:
$$\mathbf{q}_i\cdot\mathbf{q}_j=\begin{cases}1 & i=j\\ 0 & i\ne j\end{cases}$$- Theorem. An orthogonal set of non-zero vectors is linearly independent (proof above, in general: dot the combination with $\mathbf{v}_i$).
- An orthogonal basis is a basis (of a space or subspace) whose vectors are orthogonal. An orthonormal basis is one whose vectors are also unit length.
- Because orthogonal vectors are independent, $n$ orthogonal non-zero vectors in $\mathbb{R}^n$ are automatically a basis of $\mathbb{R}^n$.
- To make an orthogonal set orthonormal, normalise each vector: $\mathbf{q}_i=\mathbf{v}_i/\|\mathbf{v}_i\|$ (Chapter 1.2).
- The standard basis $\mathbf{e}_1,\dots,\mathbf{e}_n$ is the most famous orthonormal basis.
Why do we need it?
Perpendicular directions never repeat each other. We need sets of them because they give the simplest possible basis: no redundancy and no interference.
Where is it used?
Coordinate axes, the principal components found by PCA, the Fourier basis in signal processing, and the orthonormal columns inside QR and the SVD.
How is it used?
Check pairs with the dot product (zero means perpendicular). Divide each vector by its length to make a set orthonormal. In code, test it with np.allclose(Q.T @ Q, np.eye(k)).
Orthogonal $\Rightarrow$ independent, but independent $\not\Rightarrow$ orthogonal. $[1,0]$ and $[1,1]$ are independent but not perpendicular.
Orthogonal is not orthonormal. Orthogonal only talks about angles. The lengths can be anything (but non-zero). Orthonormal means perpendicular and unit length.
The zero vector is perpendicular to everything, but we leave it out of orthogonal sets because it would break independence.
Quick check: is $\{[2,1],[-1,2]\}$ orthogonal? Orthonormal? A basis of $\mathbb{R}^2$?
Dot product: $2\cdot(-1)+1\cdot2=0$, so orthogonal. Lengths are $\sqrt5$ each, not 1, so not orthonormal. Two orthogonal non-zero vectors in $\mathbb{R}^2$ are independent, so yes, a basis of $\mathbb{R}^2$. Dividing each by $\sqrt5$ makes it orthonormal.
Why orthonormal bases are wonderful: easy coordinates core
Giving "coordinates" means saying how much of each basis vector to use (see Chapter 1.3). With a general basis you must solve a system of equations to find the scoops.
With an orthonormal basis there is no solving. To learn how much of $\mathbf{q}_1$ is in $\mathbf{x}$, just take the shadow of $\mathbf{x}$ on $\mathbf{q}_1$: the dot product $\mathbf{q}_1\cdot\mathbf{x}$. The other basis vectors are perpendicular to $\mathbf{q}_1$, so they cast no shadow on that direction and cannot interfere. Each coordinate is found on its own, like reading a measuring tape along each axis.
The vectors $\mathbf{q}_1=[0.6,\,0.8]$ and $\mathbf{q}_2=[0.8,\,-0.6]$ are orthonormal ($0.6\cdot0.8+0.8\cdot(-0.6)=0$ and each has length $\sqrt{0.36+0.64}=1$). Find the coordinates of $\mathbf{x}=[5,5]$.
- $c_1=\mathbf{q}_1\cdot\mathbf{x}=0.6\cdot5+0.8\cdot5=3+4=7$.
- $c_2=\mathbf{q}_2\cdot\mathbf{x}=0.8\cdot5+(-0.6)\cdot5=4-3=1$.
- Rebuild: $7\mathbf{q}_1+1\mathbf{q}_2=[4.2,\,5.6]+[0.8,\,-0.6]=[5,\,5]$ ✓.
So in this basis $\mathbf{x}$ has coordinates $[7,1]$. If the basis was orthogonal but not normalised, divide by the length squared: $c_i=\dfrac{\mathbf{v}_i\cdot\mathbf{x}}{\mathbf{v}_i\cdot\mathbf{v}_i}$ (that is just the projection formula from Chapter 1.2).
If $\mathbf{q}_1,\dots,\mathbf{q}_n$ is an orthonormal basis of $\mathbb{R}^n$, then every $\mathbf{x}$ satisfies
$$\mathbf{x}=c_1\mathbf{q}_1+\dots+c_n\mathbf{q}_n,\qquad c_i=\mathbf{q}_i\cdot\mathbf{x}=\mathbf{q}_i^\top\mathbf{x}.$$Why: dot both sides of the first equation with $\mathbf{q}_i$. Because $\mathbf{q}_i\cdot\mathbf{q}_j=0$ for $j\ne i$ and $\mathbf{q}_i\cdot\mathbf{q}_i=1$, only $c_i$ survives.
Bonus (the Pythagorean theorem in coordinates): $\|\mathbf{x}\|^2=c_1^2+\dots+c_n^2$. The length is the same whichever orthonormal basis you use.
Why do we need it?
With a general basis, finding the coordinates of a vector means solving equations. We want a way to read coordinates off directly, which is what an orthonormal basis allows.
Where is it used?
PCA scores, the Fourier and wavelet transforms, JPEG-style image compression, and any place a signal is rewritten in a better basis.
How is it used?
Stack the basis vectors as the columns of Q. The coordinates of x are then simply Q.T @ x (one dot product per basis vector), and Q @ c rebuilds x. No linear system needs to be solved.
Quick check: $\mathbf{q}_1=[1,0]$, $\mathbf{q}_2=[0,1]$ is orthonormal. What are the coordinates of $\mathbf{x}=[4,-3]$?
$c_1=\mathbf{q}_1\cdot\mathbf{x}=4$ and $c_2=\mathbf{q}_2\cdot\mathbf{x}=-3$. The ordinary entries of a vector are its coordinates in the standard basis.
Orthogonal matrices core
Put orthonormal vectors in the columns of a square matrix $Q$. What does $Q$ do as a transformation? It sends the first axis $\mathbf{e}_1$ to the column $\mathbf{q}_1$, the second axis to $\mathbf{q}_2$, and so on. The new axes are still unit length and still perpendicular. So the whole space is just turned or flipped, like a rigid object. Nothing is stretched, squashed or sheared.
There are only two kinds of rigid motion that fix the origin: a rotation (turn) and a reflection (mirror flip). Orthogonal matrices are exactly these (and combinations).
The rotation by angle $\theta$ in the plane is $Q=\begin{bmatrix}\cos\theta&-\sin\theta\\ \sin\theta&\cos\theta\end{bmatrix}$. Take $\cos\theta=0.6$, $\sin\theta=0.8$ (about $53^\circ$):
- $Q=\begin{bmatrix}0.6&-0.8\\0.8&0.6\end{bmatrix}$. Column 1 has length $\sqrt{0.36+0.64}=1$, column 2 too, and the columns' dot product is $0.6\cdot(-0.8)+0.8\cdot0.6=0$.
- Compute $Q^\top Q$: top-left $=0.6\cdot0.6+0.8\cdot0.8=1$; top-right $=0.6\cdot(-0.8)+0.8\cdot0.6=0$; and by symmetry the rest. So $Q^\top Q=I$ ✓.
- Apply it: $Q[3,4]=[0.6\cdot3-0.8\cdot4,\;0.8\cdot3+0.6\cdot4]=[-1.4,\;4.8]$. Length: $\sqrt{1.96+23.04}=\sqrt{25}=5=\|[3,4]\|$ ✓. The length is unchanged.
A reflection example: $S=\begin{bmatrix}0&1\\1&0\end{bmatrix}$ swaps the two coordinates (a mirror along the diagonal). $S^\top S=I$, and its determinant is $0\cdot0-1\cdot1=-1$.
A square matrix $Q$ is orthogonal if its columns are orthonormal, which is equivalent to
$$Q^\top Q=I\quad\Longleftrightarrow\quad Q^{-1}=Q^\top.$$(Here $Q^\top Q$ has entries $\mathbf{q}_i\cdot\mathbf{q}_j$, so it equals $I$ exactly when the columns are orthonormal. The rows are then orthonormal too: $QQ^\top=I$.) The inverse is free: just transpose.
- Preserves lengths: $\|Q\mathbf{x}\|=\|\mathbf{x}\|$, because $\|Q\mathbf{x}\|^2=\mathbf{x}^\top Q^\top Q\mathbf{x}=\mathbf{x}^\top\mathbf{x}$.
- Preserves dot products and angles: $(Q\mathbf{x})\cdot(Q\mathbf{y})=\mathbf{x}\cdot\mathbf{y}$.
- Determinant: $\det Q=\pm1$. It is $+1$ for a rotation (keeps orientation) and $-1$ for a reflection (flips orientation, like a mirror image).
- Products and inverses of orthogonal matrices are orthogonal. Permutation matrices are orthogonal too.
Why do we need it?
We need transformations that move things around without distorting them. Orthogonal matrices are exactly the rigid motions, and their inverse costs nothing, just a transpose.
Where is it used?
Rotating images and 3D models, rotary position embeddings in Transformers, orthogonal weight initialisation in deep networks, and the building blocks of QR, SVD and eigen-solvers.
How is it used?
Test a matrix with Q.T @ Q = I. Use Q.T instead of an inverse, and trust that lengths and angles will not change. The sign of the determinant tells you whether it includes a reflection.
Orthogonal matrices have orthonormal columns, not just orthogonal ones. The name is a historical accident ("orthonormal matrix" would be clearer). If the columns are perpendicular but have different lengths, $Q^\top Q$ is a diagonal matrix that is not $I$.
A rectangular matrix can have orthonormal columns ($Q^\top Q=I$) without being square, but then $QQ^\top\ne I$ and it is not called orthogonal. Such a $Q$ appears in QR below.
Quick check: $Q$ is orthogonal and $\|\mathbf{x}\|=7$. What is $\|Q\mathbf{x}\|$? What is $Q^{-1}$?
$\|Q\mathbf{x}\|=7$: orthogonal matrices preserve length. $Q^{-1}=Q^\top$, so you never need to run an inversion algorithm.
Orthogonal projection onto a line core
Recall the shadow picture from projection in Chapter 1.2: the projection of $\mathbf{b}$ onto the line of $\mathbf{a}$ is the shadow of $\mathbf{b}$ on that line, with the light coming from directly above (at a right angle to the line). It is the closest point on the line to $\mathbf{b}$.
New idea: the shadow-making is itself a matrix. There is one matrix $P$ such that for every $\mathbf{b}$, the shadow is $P\mathbf{b}$. You build it once, then it projects anything.
Project onto the line through $\mathbf{a}=[1,1]$.
- $\mathbf{a}\mathbf{a}^\top=\begin{bmatrix}1\\1\end{bmatrix}\begin{bmatrix}1&1\end{bmatrix}=\begin{bmatrix}1&1\\1&1\end{bmatrix}$ (a column times a row gives a matrix).
- $\mathbf{a}^\top\mathbf{a}=1+1=2$.
- So $P=\dfrac{1}{2}\begin{bmatrix}1&1\\1&1\end{bmatrix}=\begin{bmatrix}0.5&0.5\\0.5&0.5\end{bmatrix}$.
- Test with $\mathbf{b}=[3,1]$: $P\mathbf{b}=[0.5\cdot3+0.5\cdot1,\;0.5\cdot3+0.5\cdot1]=[2,2]$. This matches the answer we found in Chapter 1.2 ✓.
- Project again: $P[2,2]=[2,2]$. A point already on the line does not move.
The projection of $\mathbf{b}$ onto the line spanned by $\mathbf{a}\ne\mathbf{0}$ is
$$\operatorname{proj}_{\mathbf{a}}(\mathbf{b})=\frac{\mathbf{a}^\top\mathbf{b}}{\mathbf{a}^\top\mathbf{a}}\,\mathbf{a}=P\mathbf{b},\qquad P=\frac{\mathbf{a}\mathbf{a}^\top}{\mathbf{a}^\top\mathbf{a}}.$$(Move the number to the front: $\frac{\mathbf{a}^\top\mathbf{b}}{\mathbf{a}^\top\mathbf{a}}\mathbf{a}=\frac{\mathbf{a}(\mathbf{a}^\top\mathbf{b})}{\mathbf{a}^\top\mathbf{a}}=\frac{\mathbf{a}\mathbf{a}^\top}{\mathbf{a}^\top\mathbf{a}}\mathbf{b}$.) If $\mathbf{a}$ has unit length ($\mathbf{a}=\mathbf{q}$), this is just $P=\mathbf{q}\mathbf{q}^\top$.
Two properties: $P^2=P$ (projecting twice is the same as once) and $P^\top=P$ (it is symmetric). $P$ has rank 1: its column space is the line.
Why do we need it?
Often we want 'the part of b that points along a direction a', or the closest point on a line to b. A projection matrix gives that answer for every b at once.
Where is it used?
Removing the component of a vector along a direction (for example, subtracting the mean direction), the first step of Gram–Schmidt, and simple regression with one feature and no intercept.
How is it used?
Build P = a aᵀ / (aᵀ a) once, then apply it with P @ b. The result is the shadow of b on the line, and b − P @ b is the leftover perpendicular to a.
$P$ is not invertible (unless the "line" is the whole space). It crushes a whole line of $\mathbf{b}$'s onto each shadow point, so you cannot go back from the shadow to $\mathbf{b}$. Projection throws information away.
The formula needs $\mathbf{a}^\top\mathbf{a}\ne0$, i.e. $\mathbf{a}\ne\mathbf{0}$: the zero vector spans a point, not a line.
Quick check: what is the projection matrix onto the $x$-axis in $\mathbb{R}^2$?
Take $\mathbf{a}=[1,0]$: $\mathbf{a}\mathbf{a}^\top=\begin{bmatrix}1&0\\0&0\end{bmatrix}$ and $\mathbf{a}^\top\mathbf{a}=1$, so $P=\begin{bmatrix}1&0\\0&0\end{bmatrix}$. It keeps the first entry and deletes the second.
Orthogonal projection onto a subspace core
Now replace the line by a plane (or any subspace). Put a lamp directly above a flat table and hold a stick in the air: the shadow on the table is the projection of the stick onto the plane. It is the closest point on the table to the tip of the stick, and the segment from the shadow up to the tip is perpendicular to the table.
That one fact, "the error is perpendicular to the subspace", is all we need. It is the whole recipe: find the point $\mathbf{p}$ in the subspace so that $\mathbf{b}-\mathbf{p}$ is perpendicular to the subspace. Everything else is algebra.
Let the subspace be the column space of $A=\begin{bmatrix}1&0\\1&1\\0&1\end{bmatrix}$ (a plane in $\mathbb{R}^3$, spanned by $\mathbf{a}_1=[1,1,0]$ and $\mathbf{a}_2=[0,1,1]$). Project $\mathbf{b}=[1,2,-2]$.
- The shadow is some mix $\mathbf{p}=\hat{x}_1\mathbf{a}_1+\hat{x}_2\mathbf{a}_2=A\hat{\mathbf{x}}$. We need the error $\mathbf{e}=\mathbf{b}-A\hat{\mathbf{x}}$ to be perpendicular to both columns, i.e. $A^\top\mathbf{e}=\mathbf{0}$.
- That gives the normal equations $A^\top A\,\hat{\mathbf{x}}=A^\top\mathbf{b}$. Compute $A^\top A=\begin{bmatrix}2&1\\1&2\end{bmatrix}$ and $A^\top\mathbf{b}=[1+2,\;2-2]=[3,0]$.
- Solve $\begin{bmatrix}2&1\\1&2\end{bmatrix}\hat{\mathbf{x}}=\begin{bmatrix}3\\0\end{bmatrix}$. The inverse of $\begin{bmatrix}2&1\\1&2\end{bmatrix}$ is $\tfrac13\begin{bmatrix}2&-1\\-1&2\end{bmatrix}$, so $\hat{\mathbf{x}}=\tfrac13[6,-3]=[2,-1]$.
- Projection: $\mathbf{p}=2\mathbf{a}_1-1\mathbf{a}_2=[2,2,0]-[0,1,1]=[2,1,-1]$.
- Error: $\mathbf{e}=\mathbf{b}-\mathbf{p}=[1-2,\,2-1,\,-2+1]=[-1,1,-1]$. Check: $\mathbf{e}\cdot\mathbf{a}_1=-1+1=0$ ✓ and $\mathbf{e}\cdot\mathbf{a}_2=1-1=0$ ✓.
- Pythagoras: $\|\mathbf{b}\|^2=1+4+4=9$ and $\|\mathbf{p}\|^2+\|\mathbf{e}\|^2=6+3=9$ ✓.
Let $A$ have independent columns that span the subspace $V$. The projection of $\mathbf{b}$ onto $V$ is $\mathbf{p}=A\hat{\mathbf{x}}$, where $\hat{\mathbf{x}}$ solves the normal equations. In one line:
$$\mathbf{p}=P\mathbf{b},\qquad P=A\,(A^\top A)^{-1}A^\top .$$Where this comes from (the checkpoint derivation): the condition "error $\perp V$" means $A^\top(\mathbf{b}-A\hat{\mathbf{x}})=\mathbf{0}$. Rearranging, $A^\top A\hat{\mathbf{x}}=A^\top\mathbf{b}$. Since $A$ has independent columns, $A^\top A$ is invertible, so $\hat{\mathbf{x}}=(A^\top A)^{-1}A^\top\mathbf{b}$ and $\mathbf{p}=A\hat{\mathbf{x}}$.
- Orthonormal basis shortcut. If the columns of $Q$ are orthonormal then $Q^\top Q=I$ and the formula collapses to $P=QQ^\top$. No inverse needed: $\mathbf{p}=\sum_i(\mathbf{q}_i\cdot\mathbf{b})\,\mathbf{q}_i$, the sum of the shadows on each orthonormal direction.
- Properties: $P^2=P$ (idempotent: projecting twice changes nothing) and $P^\top=P$ (symmetric). Together they characterise orthogonal projections. Rank $P$ = dimension of the subspace.
- The error vector $\mathbf{e}=\mathbf{b}-P\mathbf{b}=(I-P)\mathbf{b}$ is perpendicular to the subspace, and $\|\mathbf{b}\|^2=\|\mathbf{p}\|^2+\|\mathbf{e}\|^2$.
Why do we need it?
When the target b is outside the set of outputs a model can produce, we need the closest reachable point. Projection onto a subspace finds it, and the error is perpendicular to the subspace.
Where is it used?
Linear regression and least squares (the hat matrix puts a hat on y), PCA (projecting data onto a few directions), and denoising by projecting onto a signal subspace.
How is it used?
Stack independent basis vectors as columns of A, solve AᵀA x̂ = Aᵀb for the weights, and take p = A x̂. In practice use np.linalg.lstsq or QR, not an explicit inverse. Check that Aᵀ(b − p) is zero.
The formula needs independent columns. If the columns of $A$ are dependent, $A^\top A$ is singular. Fix: keep a basis only (drop redundant columns), or use the pseudoinverse (Chapter 1.10).
Do not "cancel" the matrices: $A(A^\top A)^{-1}A^\top$ is not $AA^{-1}(A^\top)^{-1}A^\top=I$, because a rectangular $A$ has no inverse. It is only the identity if $A$ is square and invertible (then the "subspace" is the whole space).
In real code you never form $(A^\top A)^{-1}$: you solve the normal equations (or use QR below), which is faster and more accurate.
Quick check: $\mathbf{q}_1=[1,0,0]$, $\mathbf{q}_2=[0,1,0]$ span the floor. Project $\mathbf{b}=[3,4,5]$ using $P=QQ^\top$.
$\mathbf{p}=(\mathbf{q}_1\cdot\mathbf{b})\mathbf{q}_1+(\mathbf{q}_2\cdot\mathbf{b})\mathbf{q}_2=3[1,0,0]+4[0,1,0]=[3,4,0]$. The shadow on the floor simply drops the height. The error is $[0,0,5]$, pointing straight up.
Splitting a vector: $\mathbf{x}=\mathbf{x}_\parallel+\mathbf{x}_\perp$ core
Any arrow can be split into two pieces: the part that lies along a chosen subspace, and the part that sticks out perpendicular to it. Standing on a hillside: your movement splits into "along the slope" and "straight away from the slope". The first piece is the projection; the second is what is left over, which is the projection onto the orthogonal complement $V^\perp$ (Chapter 1.8).
The two pieces are perpendicular, so they obey Pythagoras: the squared lengths add up.
Let $V$ be the line through $[1,1]$ and $\mathbf{x}=[3,1]$. We already found $P\mathbf{x}=[2,2]$.
- $\mathbf{x}_\parallel=P\mathbf{x}=[2,2]$.
- $\mathbf{x}_\perp=\mathbf{x}-\mathbf{x}_\parallel=[1,-1]$, and also $(I-P)\mathbf{x}=\begin{bmatrix}0.5&-0.5\\-0.5&0.5\end{bmatrix}[3,1]=[1,-1]$ ✓.
- Perpendicular? $[2,2]\cdot[1,-1]=2-2=0$ ✓.
- Pythagoras: $\|\mathbf{x}\|^2=9+1=10$ and $\|\mathbf{x}_\parallel\|^2+\|\mathbf{x}_\perp\|^2=8+2=10$ ✓.
For a subspace $V\subseteq\mathbb{R}^n$ with projection matrix $P$, every $\mathbf{x}\in\mathbb{R}^n$ splits uniquely as
$$\mathbf{x}=\underbrace{P\mathbf{x}}_{\mathbf{x}_\parallel\in V}+\underbrace{(I-P)\mathbf{x}}_{\mathbf{x}_\perp\in V^\perp},\qquad \|\mathbf{x}\|^2=\|\mathbf{x}_\parallel\|^2+\|\mathbf{x}_\perp\|^2.$$$I-P$ is itself a projection matrix: it projects onto $V^\perp$. Also $P(I-P)=P-P^2=0$: the two parts never overlap.
Link to Chapter 1.8. This is the big picture in action: every input $\mathbf{x}$ splits as (row-space part) + (null-space part). $A$ only sees the row-space part, because it destroys the null-space part.
Why do we need it?
We often want to separate a vector into the part we can explain and the part we cannot. A split into 'inside the subspace' plus 'perpendicular to it' is the cleanest such separation.
Where is it used?
Signal and noise separation in PCA, fitted values and residuals in regression, and the split of an input into row-space part and null-space part for a matrix.
How is it used?
Compute the parallel part with P @ x, then the perpendicular part as x − P @ x. Check that their dot product is zero and that their squared lengths add up to the squared length of x.
The split is unique: there is only one way to write $\mathbf{x}$ as (something in $V$) + (something in $V^\perp$). If you allowed non-perpendicular leftovers you could split it in many ways, but then it would not be the closest-point projection.
"Parallel" is shorthand for "inside the subspace". $\mathbf{x}_\parallel$ lies in $V$, even when $V$ is a plane and "parallel" sounds odd.
Quick check: if $\mathbf{x}$ already lies in $V$, what are $\mathbf{x}_\parallel$ and $\mathbf{x}_\perp$?
$\mathbf{x}_\parallel=\mathbf{x}$ (projecting a vector that is already in the subspace changes nothing, since $P^2=P$) and $\mathbf{x}_\perp=\mathbf{0}$.
The Gram–Schmidt process core
Orthonormal bases are wonderful, but real data hands you an ordinary, slanted basis. Gram–Schmidt is a recipe that straightens it out, one vector at a time, without changing the space they span.
- Keep the first vector's direction. Scale it to length 1. That is $\mathbf{q}_1$.
- Take the second vector. It leans partly along $\mathbf{q}_1$. Subtract that leaning part (its projection on $\mathbf{q}_1$). What is left is perpendicular to $\mathbf{q}_1$. Scale it to length 1. That is $\mathbf{q}_2$.
- Take the third vector. Subtract its shadows on $\mathbf{q}_1$ and $\mathbf{q}_2$. The leftover is perpendicular to both. Scale it. That is $\mathbf{q}_3$.
- Keep going.
Each new vector is first "cleaned" of everything the earlier ones already explain, and only then kept. This is just projection (the previous sections) used over and over.
Start from $\mathbf{a}_1=[3,4]$ and $\mathbf{a}_2=[5,5]$ (independent, not perpendicular).
- $\|\mathbf{a}_1\|=\sqrt{9+16}=5$, so $\mathbf{q}_1=[3,4]/5=[0.6,\,0.8]$.
- How much of $\mathbf{a}_2$ lies along $\mathbf{q}_1$? $\mathbf{q}_1\cdot\mathbf{a}_2=0.6\cdot5+0.8\cdot5=7$. Its shadow is $7\mathbf{q}_1=[4.2,\,5.6]$.
- Subtract: $\mathbf{v}_2=\mathbf{a}_2-7\mathbf{q}_1=[5-4.2,\;5-5.6]=[0.8,\,-0.6]$.
- Check it is perpendicular to $\mathbf{q}_1$: $0.8\cdot0.6+(-0.6)\cdot0.8=0.48-0.48=0$ ✓.
- Its length is $\sqrt{0.64+0.36}=1$ already, so $\mathbf{q}_2=[0.8,\,-0.6]$.
We turned $\{[3,4],[5,5]\}$ into the orthonormal pair $\{[0.6,0.8],[0.8,-0.6]\}$, the same one used earlier in this chapter. They span the same plane (all of $\mathbb{R}^2$ here).
Given independent vectors $\mathbf{a}_1,\dots,\mathbf{a}_n$, classical Gram–Schmidt produces orthonormal $\mathbf{q}_1,\dots,\mathbf{q}_n$ with the same span at every step:
$$\mathbf{v}_k=\mathbf{a}_k-\sum_{i=1}^{k-1}(\mathbf{q}_i\cdot\mathbf{a}_k)\,\mathbf{q}_i,\qquad \mathbf{q}_k=\frac{\mathbf{v}_k}{\|\mathbf{v}_k\|}.$$The sum removes the shadows of $\mathbf{a}_k$ on all earlier directions (this is the projection onto $\text{span}(\mathbf{q}_1,\dots,\mathbf{q}_{k-1})$, using the $QQ^\top$ shortcut). $\mathbf{v}_k$ is the error vector: it is perpendicular to everything before it.
- If $\mathbf{v}_k=\mathbf{0}$, then $\mathbf{a}_k$ was already inside the span of the earlier vectors (the inputs were dependent). The process cannot continue for that vector: you must skip it.
- Modified Gram–Schmidt (a more stable version, see below) does the same maths but subtracts the shadows one at a time from the vector as it is being cleaned: $\mathbf{v}\leftarrow\mathbf{v}-(\mathbf{q}_i\cdot\mathbf{v})\mathbf{q}_i$ for $i=1,2,\dots$ (using the current $\mathbf{v}$ instead of the original $\mathbf{a}_k$).
Why do we need it?
Real bases come slanted. We need a reliable recipe that turns any basis into a perpendicular one that spans the same space, so that all the easy formulas above apply.
Where is it used?
Building orthonormal bases of feature spaces, computing QR decompositions, orthogonalising new directions in iterative solvers (Arnoldi and GMRES), and creating orthogonal weight matrices.
How is it used?
Go through the vectors in order. Subtract from each one its shadows on the earlier q's, then divide by the length of what is left. In code use the modified version, or just call np.linalg.qr.
Always normalise after subtracting. Without the last division you get an orthogonal set, not an orthonormal one.
Order matters. Gram–Schmidt keeps $\mathbf{a}_1$'s direction and treats later vectors as "corrections". Reorder the inputs and you get a different (equally valid) orthonormal basis of the same space.
Classical vs modified (awareness). On paper both give the same answer. On a computer, rounding errors make the classical version lose orthogonality when the input vectors are nearly dependent (see the widget). The modified version is much more stable. In practice, libraries use even more stable methods (Householder reflections) for QR.
Quick check: run Gram–Schmidt on $\mathbf{a}_1=[1,0]$, $\mathbf{a}_2=[1,1]$.
$\mathbf{q}_1=[1,0]$. Shadow of $\mathbf{a}_2$: $(\mathbf{q}_1\cdot\mathbf{a}_2)\mathbf{q}_1=1\cdot[1,0]$. $\mathbf{v}_2=[1,1]-[1,0]=[0,1]$, which has length 1, so $\mathbf{q}_2=[0,1]$. The slanted pair became the standard basis.
QR decomposition (an introduction) core
Gram–Schmidt takes the columns $\mathbf{a}_j$ of a matrix and produces orthonormal vectors $\mathbf{q}_j$. Nothing is thrown away: each original column can be rebuilt from the $\mathbf{q}$'s. Writing down the "recipe" for each column gives a second matrix, $R$. The result is a factorisation: $A=QR$ (like writing $12=3\times4$).
The recipe has a pleasant shape. Column 1 of $A$ is made only from $\mathbf{q}_1$. Column 2 is made from $\mathbf{q}_1$ and $\mathbf{q}_2$. Column 3 from $\mathbf{q}_1,\mathbf{q}_2,\mathbf{q}_3$. So the recipe matrix $R$ has zeros below the diagonal: it is upper triangular.
Use $A=\begin{bmatrix}3&5\\4&5\end{bmatrix}$ (columns $\mathbf{a}_1=[3,4]$, $\mathbf{a}_2=[5,5]$). We just ran Gram–Schmidt on these: $\mathbf{q}_1=[0.6,0.8]$, $\mathbf{q}_2=[0.8,-0.6]$.
- $\mathbf{a}_1=5\,\mathbf{q}_1+0\,\mathbf{q}_2$ (its length is 5). So the first column of $R$ is $[5,0]$.
- $\mathbf{a}_2=7\,\mathbf{q}_1+1\,\mathbf{q}_2$ (we found shadow coefficient 7, leftover length 1). So the second column of $R$ is $[7,1]$.
- Put them together: $Q=\begin{bmatrix}0.6&0.8\\0.8&-0.6\end{bmatrix}$ and $R=\begin{bmatrix}5&7\\0&1\end{bmatrix}$.
- Multiply to verify: $QR=\begin{bmatrix}0.6\cdot5&0.6\cdot7+0.8\cdot1\\0.8\cdot5&0.8\cdot7-0.6\cdot1\end{bmatrix}=\begin{bmatrix}3&5\\4&5\end{bmatrix}=A$ ✓.
Notice how to get $R$ without extra work: $R_{ij}=\mathbf{q}_i\cdot\mathbf{a}_j$ (the shadow coefficients) and the diagonal entry $R_{jj}=\|\mathbf{v}_j\|$ (the leftover length). Equivalently $R=Q^\top A$.
Every $m\times n$ matrix $A$ with independent columns can be written as
$$A=QR$$- $Q$ is $m\times n$ with orthonormal columns ($Q^\top Q=I$): the output of Gram–Schmidt.
- $R$ is $n\times n$ upper triangular with positive diagonal: $R_{ij}=\mathbf{q}_i\cdot\mathbf{a}_j$ for $i\le j$, and $R_{ij}=0$ for $i\gt j$.
- $C(Q)=C(A)$: the columns of $Q$ are an orthonormal basis for the column space of $A$.
- Why triangular: $\mathbf{q}_i$ is perpendicular to $\mathbf{a}_1,\dots,\mathbf{a}_{i-1}$ and so $R_{ij}=\mathbf{q}_i\cdot\mathbf{a}_j=0$ whenever $i\gt j$.
If $A$ is square then $Q$ is a true orthogonal matrix. For a tall $A$ ($m\gt n$), this is the "reduced" (or "thin") QR. Libraries can also give a "full" $Q$ that is $m\times m$.
Why do we need it?
We want to solve least-squares problems and find orthonormal bases without forming AᵀA, which squares the sensitivity to rounding errors. QR does this with a triangular system that is cheap to solve.
Where is it used?
Stable linear regression (np.linalg.lstsq), the QR algorithm for eigenvalues, orthonormal bases for column spaces, and many numerical libraries behind scipy and PyTorch.
How is it used?
Compute Q, R = np.linalg.qr(A). To solve Ax ≈ b, solve the triangular system R x = Qᵀ b by back-substitution. Check that Q @ R reproduces A and that Qᵀ Q is the identity.
Do not confuse this $R$ with the $R$ (reduced row echelon form) of Chapter 1.8. Here $R$ is the "recipe" matrix. Likewise $Q$ is not a general matrix: its columns are orthonormal.
The factorisation is unique if we insist that the diagonal of $R$ is positive. (Flip a column's sign in $Q$ and the matching row's sign in $R$ and you get another valid QR.)
Quick check: for $A=QR$ with $Q^\top Q=I$, how do you get $R$ from $A$ and $Q$?
Multiply both sides by $Q^\top$: $Q^\top A=Q^\top QR=IR=R$. So $R=Q^\top A$.
Recap, cheat sheet and practice
- A set of vectors is orthogonal if every pair has dot product 0, and orthonormal if each also has length 1. Orthogonal non-zero vectors are always independent.
- In an orthonormal basis, the coordinates are just dot products: $c_i=\mathbf{q}_i\cdot\mathbf{x}$, and $\mathbf{x}=\sum c_i\mathbf{q}_i$.
- An orthogonal matrix has orthonormal columns: $Q^\top Q=I$, $Q^{-1}=Q^\top$. It keeps lengths and angles, is a rotation ($\det=+1$) or reflection ($\det=-1$).
- Projection onto a line: $P=\mathbf{a}\mathbf{a}^\top/\mathbf{a}^\top\mathbf{a}$. Onto a subspace: $P=A(A^\top A)^{-1}A^\top$, or $P=QQ^\top$ with orthonormal columns. $P^2=P$, $P^\top=P$. The error is perpendicular to the subspace.
- Every vector splits uniquely as $\mathbf{x}=\mathbf{x}_\parallel+\mathbf{x}_\perp$ with $\|\mathbf{x}\|^2=\|\mathbf{x}_\parallel\|^2+\|\mathbf{x}_\perp\|^2$.
- Gram–Schmidt straightens a basis: subtract shadows on earlier $\mathbf{q}$'s, then normalise (modified GS is more stable). It gives $A=QR$, with $Q$ orthonormal columns and $R$ upper triangular.
Cheat sheet
| Idea | Formula | Picture |
|---|---|---|
| Orthonormal | $\mathbf{q}_i\cdot\mathbf{q}_j=\delta_{ij}$ | perpendicular unit arrows |
| Coordinates | $c_i=\mathbf{q}_i\cdot\mathbf{x}$ | shadow on each axis |
| Orthogonal matrix | $Q^\top Q=I,\ Q^{-1}=Q^\top,\ \det=\pm1$ | rotation or mirror |
| Projection on a line | $P=\mathbf{a}\mathbf{a}^\top/\mathbf{a}^\top\mathbf{a}$ | shadow on the line |
| Projection on a subspace | $P=A(A^\top A)^{-1}A^\top=QQ^\top$ | shadow on the plane |
| Projection properties | $P^2=P,\ P^\top=P$ | shadow of a shadow is itself |
| Split | $\mathbf{x}=P\mathbf{x}+(I-P)\mathbf{x}$ | along + perpendicular |
| Gram–Schmidt | $\mathbf{v}_k=\mathbf{a}_k-\sum(\mathbf{q}_i\cdot\mathbf{a}_k)\mathbf{q}_i$, $\mathbf{q}_k=\mathbf{v}_k/\|\mathbf{v}_k\|$ | subtract shadows, rescale |
| QR | $A=QR,\ R=Q^\top A$ | orthonormal basis + triangular recipe |
import numpy as np
# ---- projection matrix and its properties --------------------------------
A = np.array([[1., 0.],
[1., 1.],
[0., 1.]])
P = A @ np.linalg.inv(A.T @ A) @ A.T # P = A (A^T A)^-1 A^T
print(np.allclose(P @ P, P)) # P^2 = P -> True
print(np.allclose(P.T, P)) # P^T = P -> True
b = np.array([1., 2., -2.])
p = P @ b # [ 2. 1. -1.]
e = b - p # [-1. 1. -1.]
print(A.T @ e) # ~[0, 0]: the error is perpendicular to the columns
# ---- Gram-Schmidt: classical and modified ---------------------------------
def classical_gs(A):
m, n = A.shape
Q = np.zeros((m, n)); R = np.zeros((n, n))
for j in range(n):
v = A[:, j].copy()
for i in range(j):
R[i, j] = Q[:, i] @ A[:, j] # shadow computed from the ORIGINAL column
v -= R[i, j] * Q[:, i]
R[j, j] = np.linalg.norm(v)
Q[:, j] = v / R[j, j]
return Q, R
def modified_gs(A):
m, n = A.shape
Q = np.zeros((m, n)); R = np.zeros((n, n))
for j in range(n):
v = A[:, j].copy()
for i in range(j):
R[i, j] = Q[:, i] @ v # shadow computed from the CURRENT, partly cleaned v
v -= R[i, j] * Q[:, i]
R[j, j] = np.linalg.norm(v)
Q[:, j] = v / R[j, j]
return Q, R
# nearly dependent columns (Läuchli matrix)
eps = 1e-8
L = np.array([[1, 1, 1],
[eps, 0, 0],
[0, eps, 0],
[0, 0, eps]])
for name, gs in [("classical", classical_gs), ("modified", modified_gs)]:
Q, R = gs(L)
loss = np.abs(Q.T @ Q - np.eye(3)).max() # loss of orthogonality
print(name, loss) # classical: about 0.5 modified: about 1e-8
# ---- QR with NumPy ---------------------------------------------------------
B = np.array([[3., 5.], [4., 5.]])
Q, R = np.linalg.qr(B)
print(np.allclose(Q @ R, B), np.allclose(Q.T @ Q, np.eye(2))) # True True
print(R) # upper triangular (NumPy may flip signs of Q's columns and R's rows)
1. Is the set $\{[1,2],[2,-1]\}$ orthogonal or orthonormal?
2. Which statement about an orthogonal matrix $Q$ is false?
3. $Q$ has orthonormal columns and $P=QQ^\top$. Which pair of equations holds?
4. In Gram–Schmidt, how do we get $\mathbf{q}_2$ from $\mathbf{a}_2$?
5. In $A=QR$ (from Gram–Schmidt), what kind of matrix is $R$?
6. $\mathbf{q}_1=[0.6,0.8]$, $\mathbf{q}_2=[0.8,-0.6]$ is an orthonormal basis. What is the second coordinate $c_2$ of $\mathbf{x}=[10,0]$?
Practice problems
A. Show that $[1,1,1]$, $[1,-1,0]$, $[1,1,-2]$ are orthogonal, and normalise the first.
Dot products: $1-1+0=0$, $1+1-2=0$, $1-1+0=0$ ✓. The first has length $\sqrt3$, so $\mathbf{q}_1=[1,1,1]/\sqrt3\approx[0.577,0.577,0.577]$.
B. Using the orthogonal basis from A, write $\mathbf{x}=[3,1,2]$ in that basis.
Because the basis is orthogonal, $c_i=\dfrac{\mathbf{v}_i\cdot\mathbf{x}}{\mathbf{v}_i\cdot\mathbf{v}_i}$. $\mathbf{v}_1\cdot\mathbf{x}=6$, $\mathbf{v}_1\cdot\mathbf{v}_1=3$, so $c_1=2$. $\mathbf{v}_2\cdot\mathbf{x}=3-1=2$, $\mathbf{v}_2\cdot\mathbf{v}_2=2$, so $c_2=1$. $\mathbf{v}_3\cdot\mathbf{x}=3+1-4=0$, so $c_3=0$. Check: $2[1,1,1]+1[1,-1,0]=[3,1,2]$ ✓.
C. Find the matrix $P$ projecting onto the line through $\mathbf{a}=[2,1]$, verify $P^2=P$, and project $\mathbf{b}=[1,3]$.
$\mathbf{a}\mathbf{a}^\top=\begin{bmatrix}4&2\\2&1\end{bmatrix}$ and $\mathbf{a}^\top\mathbf{a}=5$, so $P=\tfrac15\begin{bmatrix}4&2\\2&1\end{bmatrix}$. $P^2=\tfrac1{25}\begin{bmatrix}16+4&8+2\\8+2&4+1\end{bmatrix}=\tfrac1{25}\begin{bmatrix}20&10\\10&5\end{bmatrix}=P$ ✓. $P\mathbf{b}=\tfrac15[4+6,\;2+3]=[2,1]$. The error $[1,3]-[2,1]=[-1,2]$ is perpendicular to $\mathbf{a}$: $-2+2=0$ ✓.
D. Run Gram–Schmidt on $\mathbf{a}_1=[1,1]$, $\mathbf{a}_2=[1,0]$ and write $A=QR$.
$\|\mathbf{a}_1\|=\sqrt2$, so $\mathbf{q}_1=[1,1]/\sqrt2$. $\mathbf{q}_1\cdot\mathbf{a}_2=1/\sqrt2$, so the shadow is $\tfrac1{\sqrt2}\cdot\tfrac1{\sqrt2}[1,1]=[0.5,0.5]$. $\mathbf{v}_2=[0.5,-0.5]$, with length $1/\sqrt2$, so $\mathbf{q}_2=[1,-1]/\sqrt2$. Then $R=\begin{bmatrix}\sqrt2&1/\sqrt2\\0&1/\sqrt2\end{bmatrix}$. Check the first column: $\sqrt2\,\mathbf{q}_1=[1,1]$ ✓. Second: $\tfrac1{\sqrt2}\mathbf{q}_1+\tfrac1{\sqrt2}\mathbf{q}_2=\tfrac12[1,1]+\tfrac12[1,-1]=[1,0]$ ✓.
E. If $P$ is a projection matrix ($P^2=P$), show that $I-P$ is also one.
$(I-P)^2=I-2P+P^2=I-2P+P=I-P$ ✓. (If $P$ is also symmetric, so is $I-P$.) $P$ projects onto $V$ and $I-P$ projects onto the orthogonal complement $V^\perp$, as in the splitting $\mathbf{x}=P\mathbf{x}+(I-P)\mathbf{x}$.
Least Squares
Real data never fits perfectly. There are more equations than unknowns, and no exact answer exists. Least squares finds the best answer anyway, and it turns out to be a projection. This is the heart of linear regression, so we will take it slowly and look at it from every side.
- Say what "best approximate solution" means: residuals and the squared-error objective
- See the answer as a projection: the best prediction is the closest point in the column space, and the error is perpendicular to it
- Derive the normal equations twice: once with geometry, once with calculus
- Solve least squares three ways (normal equations, QR, SVD) and know why the first one is risky
- Use the pseudoinverse, including the shortest solution of an underdetermined system
- Fix unstable problems with ridge (regularised) least squares, and meet weighted least squares
- Fit lines, planes and curves by building a design matrix
This chapter builds on Chapter 1.8 (the column space and the left null space) and Chapter 1.9 (orthogonal projection and QR). If a word like "column space" feels rusty, take a quick look back. We will also peek ahead at the SVD, which gets its own proper chapter later (Chapter 1.13).
The problem: more equations than unknowns core
Suppose you measure three points and you want the straight line through them. A line has only two knobs: where it crosses the y-axis (the intercept) and how steep it is (the slope). Two points fix both knobs. But a third point usually sits off that line.
Real measurements are noisy, so this is the normal situation, not a rare accident. We have more equations (one per point) than unknowns (two knobs). Demanding that every equation holds exactly is impossible.
So we change the question. Instead of "which line passes through every point?", we ask: "which line misses the points by the least, overall?" That is least squares.
Three data points: $(1, 1)$, $(2, 2)$ and $(3, 2)$. We want a line $y = c + m\,x$. Each point gives one equation:
$$\begin{aligned} c + 1m &= 1 \\ c + 2m &= 2 \\ c + 3m &= 2 \end{aligned} \qquad\Longleftrightarrow\qquad \underbrace{\begin{bmatrix} 1 & 1 \\ 1 & 2 \\ 1 & 3 \end{bmatrix}}_{A}\underbrace{\begin{bmatrix} c \\ m \end{bmatrix}}_{\mathbf{x}} = \underbrace{\begin{bmatrix} 1 \\ 2 \\ 2 \end{bmatrix}}_{\mathbf{b}}$$- Use the first two equations. Subtracting them gives $m = 1$, and then $c = 0$.
- Check the third equation: $0 + 3\cdot 1 = 3$, but we needed $2$. It fails. So $A\mathbf{x}=\mathbf{b}$ has no solution.
- Try the line $y = x$ anyway ($c = 0$, $m = 1$). It predicts $[1, 2, 3]$. The data is $[1, 2, 2]$. The misses are $[0, 0, -1]$.
- Add up the squares of the misses: $0^2 + 0^2 + (-1)^2 = 1$.
- Try $c = 1$, $m = 0.5$. It predicts $[1.5, 2, 2.5]$. Misses: $[-0.5, 0, -0.5]$. Sum of squares: $0.25 + 0 + 0.25 = 0.5$. Better!
The best line of all is $y = \tfrac23 + \tfrac12 x$, with total squared miss $\tfrac16 \approx 0.167$. By the end of the next two sections you will be able to find it yourself.
Let $A$ be an $m\times n$ matrix and $\mathbf{b}\in\mathbb{R}^m$. The system $A\mathbf{x}=\mathbf{b}$ is overdetermined when $m > n$ (more equations than unknowns). Usually it has no exact solution (Chapter 1.6: $\mathbf{b}$ is not in the column space of $A$).
For any guess $\mathbf{x}$, the residual is the leftover error:
$$\mathbf{r} = \mathbf{b} - A\mathbf{x}.$$Entry $r_i$ says how far the $i$-th prediction misses the $i$-th data value. The least-squares problem is to find the $\hat{\mathbf{x}}$ that makes the residual as short as possible:
$$\hat{\mathbf{x}} = \arg\min_{\mathbf{x}}\; \|A\mathbf{x}-\mathbf{b}\|^2 = \arg\min_{\mathbf{x}}\; \sum_{i=1}^{m} r_i^2.$$("$\arg\min$" means "the input that gives the smallest value".) The hat in $\hat{\mathbf{x}}$ is the usual way to say "the best estimate". The number $\|\mathbf{r}\|^2=\sum r_i^2$ is the sum of squared residuals.
Why do we need it?
Real measurements are noisy, so with more equations than unknowns an exact fit almost never exists. Without a rule for the best compromise, we could only say "no solution" and stop.
Where is it used?
Linear regression, fitting a trend line to sales or sensor data, calibrating an instrument, estimating a position from many noisy GPS readings, and the repeated "solve for one side" steps in recommender systems (alternating least squares, ALS).
How is it used?
Write the data as a matrix A and a target vector b. Pick the weights that make the squared residual ‖b − Ax‖² as small as possible. Then always look at the residuals: they show where the model misses and whether one outlier is dragging the fit.
Why squares? We could add up the plain misses, but then a miss of $+1$ and a miss of $-1$ would cancel and look like "no error". Squaring makes every miss positive. It also punishes big misses much more than small ones (a miss of 3 costs 9, a miss of 1 costs 1). And, very importantly, squares are smooth, so we can use calculus and linear algebra on them. Adding up absolute values instead (the L1 idea from Chapter 1.2) is a valid choice too, but it has no neat closed-form answer.
Squares have a cost. A single wild point (an outlier) has a huge squared miss, so it pulls the whole line toward itself. Try dragging one blue point far away in the widget above.
Quick check: you fit a straight line (intercept and slope) to 50 data points. What is the shape of $A$, and is the system overdetermined?
$A$ has one row per point and one column per unknown, so it is $50\times 2$. Since $50 > 2$ the system is overdetermined: there are far more equations than unknowns.
The picture: least squares is a projection core
Remember the column space $C(A)$ from Chapter 1.8: it is every vector you can build as $A\mathbf{x}$, that is, every possible mix of the columns. Picture it as a flat sheet of paper floating in space.
The data vector $\mathbf{b}$ usually hovers off the sheet. That is exactly why $A\mathbf{x}=\mathbf{b}$ has no solution: we can only ever reach points on the sheet.
So which point of the sheet is best? The closest one to $\mathbf{b}$. And the closest point on a flat sheet is found by dropping a straight line down from $\mathbf{b}$ at a right angle: the foot of the perpendicular, like the point directly beneath a lamp. This is the projection from Chapter 1.9.
The little arrow from that foot up to $\mathbf{b}$ is the leftover error. It points straight out of the sheet, so it is perpendicular to every direction in the sheet.
Our running example again: the columns are $\mathbf{a}_1 = [1,1,1]$ and $\mathbf{a}_2 = [1,2,3]$, and $\mathbf{b} = [1,2,2]$. The best weights are $\hat{\mathbf{x}} = [\tfrac23, \tfrac12]$.
- The closest point on the sheet is $\mathbf{p} = A\hat{\mathbf{x}} = \tfrac23[1,1,1] + \tfrac12[1,2,3] = [\tfrac76, \tfrac53, \tfrac{13}{6}]$.
- The residual is $\mathbf{e} = \mathbf{b} - \mathbf{p} = [1 - \tfrac76,\; 2-\tfrac53,\; 2-\tfrac{13}6] = [-\tfrac16, \tfrac13, -\tfrac16]$.
- Is it perpendicular to the first column? $\mathbf{e}\cdot\mathbf{a}_1 = -\tfrac16 + \tfrac13 - \tfrac16 = 0$. ✓
- And to the second? $\mathbf{e}\cdot\mathbf{a}_2 = -\tfrac16 + \tfrac23 - \tfrac36 = 0$. ✓
- Pythagoras for the right triangle $\mathbf{p}$, $\mathbf{e}$, $\mathbf{b}$: $\|\mathbf{p}\|^2 + \|\mathbf{e}\|^2 = \tfrac{318}{36} + \tfrac{6}{36} = 9$, and $\|\mathbf{b}\|^2 = 1+4+4 = 9$. ✓
The least-squares solution $\hat{\mathbf{x}}$ is characterised by two equivalent facts:
- $A\hat{\mathbf{x}} = \operatorname{proj}_{C(A)}(\mathbf{b})$: the prediction is the orthogonal projection of $\mathbf{b}$ onto the column space.
- $\mathbf{b} - A\hat{\mathbf{x}} \;\perp\; C(A)$: the residual is orthogonal to every column of $A$.
In the language of Chapter 1.8, the residual lives in the left null space $N(A^\top)$, the orthogonal complement of $C(A)$. So $\mathbf{b}$ splits into two perpendicular pieces: $\mathbf{b} = \underbrace{A\hat{\mathbf{x}}}_{\text{in } C(A)} + \underbrace{\mathbf{e}}_{\text{in } N(A^\top)}$.
Why is the projection the minimiser? Take any other point $A\mathbf{x}$ on the sheet. Then $\|\mathbf{b}-A\mathbf{x}\|^2 = \|\mathbf{e}\|^2 + \|A\hat{\mathbf{x}} - A\mathbf{x}\|^2 \ge \|\mathbf{e}\|^2$, by Pythagoras: the leg from $\mathbf{b}$ down to the sheet is perpendicular to the leg that runs along the sheet.
Why do we need it?
We need to know what "best" really means, so that we can find it and trust it. The picture answers it: the best prediction is the closest point the model can reach, and the leftover error is perpendicular to everything the model can reach.
Where is it used?
Behind every linear regression, behind PCA reconstruction (a projection onto a few directions), in signal denoising, and in the statistics rule that residuals are uncorrelated with the features (the orthogonality principle).
How is it used?
Use it as a test of your fit: multiply Aᵀ by the residual. If the result is (almost) zero, you have the least-squares answer. With a bias column, also check that the residuals add up to zero.
"Closest" depends on how you measure distance. Least squares uses the ordinary straight-line (L2) distance between the vectors $A\mathbf{x}$ and $\mathbf{b}$. Both have one entry per data point. The picture above lives in the space of data (3 numbers, one per data point), not in the space of the line's two knobs (intercept and slope) and not in the 2D plot of the points. This is the most common source of confusion, so it is worth repeating: the 2D scatter plot and the 3D projection picture are two different pictures of the same problem.
Quick check: if $\mathbf{b}$ already lies on the sheet $C(A)$, what is the residual?
Zero. The closest point to $\mathbf{b}$ is $\mathbf{b}$ itself, so the system $A\mathbf{x}=\mathbf{b}$ has an exact solution and the squared error is $0$.
The normal equations: two derivations core
We now know the target: find $\hat{\mathbf{x}}$ so that the residual points straight out of the sheet. "Straight out of the sheet" means perpendicular to every column. And perpendicular means "dot product equals zero".
So we just write down: (column 1) · (residual) = 0, (column 2) · (residual) = 0, … That is one equation per unknown. Stack them up and you get a square system we can solve. It is called the normal equations, because "normal" is another word for "perpendicular".
There is a second way to find the same equations. Think of the total squared error as a bowl whose height at each choice of weights is the error. The best weights are at the very bottom of the bowl, where the ground is flat in every direction (the slope is zero). Setting the slopes to zero gives exactly the same equations. Two completely different roads, one destination.
Running example: $A = \begin{bmatrix} 1&1\\1&2\\1&3 \end{bmatrix}$, $\mathbf{b} = [1,2,2]$. We need $A^\top A$ and $A^\top\mathbf{b}$.
- $A^\top A$ is a table of dot products of columns: $\mathbf{a}_1\!\cdot\!\mathbf{a}_1 = 3$, $\mathbf{a}_1\!\cdot\!\mathbf{a}_2 = 1+2+3 = 6$, $\mathbf{a}_2\!\cdot\!\mathbf{a}_2 = 1+4+9 = 14$. So $A^\top A = \begin{bmatrix} 3&6\\6&14 \end{bmatrix}$.
- $A^\top\mathbf{b}$ is each column dotted with $\mathbf{b}$: $\mathbf{a}_1\!\cdot\!\mathbf{b} = 1+2+2 = 5$ and $\mathbf{a}_2\!\cdot\!\mathbf{b} = 1+4+6 = 11$. So $A^\top\mathbf{b} = [5, 11]$.
- Solve $\begin{bmatrix} 3&6\\6&14 \end{bmatrix}\begin{bmatrix} c\\m \end{bmatrix} = \begin{bmatrix} 5\\11 \end{bmatrix}$. Double the first equation: $6c + 12m = 10$. Subtract it from the second ($6c + 14m = 11$): $2m = 1$, so $m = \tfrac12$.
- Back in the first equation: $3c + 6\cdot\tfrac12 = 5$, so $3c = 2$ and $c = \tfrac23$.
- So $\hat{\mathbf{x}} = [\tfrac23, \tfrac12]$: the line $y = \tfrac23 + \tfrac12 x$, exactly the "best line" promised earlier.
The normal equations for the least-squares problem $\min_{\mathbf{x}}\|A\mathbf{x}-\mathbf{b}\|^2$ are
$$\boxed{\;A^\top A\,\hat{\mathbf{x}} = A^\top\mathbf{b}\;}$$$A^\top A$ is an $n\times n$ symmetric matrix (the table of dot products of the columns). $A^\top \mathbf{b}$ is a vector with $n$ entries.
Derivation 1: geometry. The residual $\mathbf{b}-A\hat{\mathbf{x}}$ must be orthogonal to every column of $A$. "Every column dotted with the residual is zero" is written all at once as
$$A^\top(\mathbf{b} - A\hat{\mathbf{x}}) = \mathbf{0} \;\;\Longrightarrow\;\; A^\top A\hat{\mathbf{x}} = A^\top\mathbf{b}.$$Derivation 2: calculus. Write the error as a function of the weights and set its slopes to zero. First with two unknowns, using our example. With $E(c,m) = (c+m-1)^2 + (c+2m-2)^2 + (c+3m-2)^2$, the chain rule gives
$$\frac{\partial E}{\partial c} = 2\big[(c+m-1) + (c+2m-2) + (c+3m-2)\big] = 2\,(3c + 6m - 5),$$ $$\frac{\partial E}{\partial m} = 2\big[1(c+m-1) + 2(c+2m-2) + 3(c+3m-2)\big] = 2\,(6c + 14m - 11).$$Setting both to zero gives $3c+6m=5$ and $6c+14m=11$: the very same equations as before. Now the general case. Expand the squared length (using $(\mathbf{u}-\mathbf{v})^\top = \mathbf{u}^\top - \mathbf{v}^\top$):
$$E(\mathbf{x}) = (A\mathbf{x}-\mathbf{b})^\top(A\mathbf{x}-\mathbf{b}) = \mathbf{x}^\top A^\top A\,\mathbf{x} \;-\; 2\,\mathbf{b}^\top A\,\mathbf{x} \;+\; \mathbf{b}^\top\mathbf{b}.$$(The two middle terms are the same single number, one being the transpose of the other, so they add to $-2\mathbf{b}^\top A\mathbf{x}$.) Two standard gradient rules, which you will prove yourself in Chapter 1.14, say that the gradient of $\mathbf{x}^\top S\mathbf{x}$ is $2S\mathbf{x}$ (for symmetric $S$) and the gradient of $\mathbf{c}^\top\mathbf{x}$ is $\mathbf{c}$. Therefore
$$\nabla E(\mathbf{x}) = 2A^\top A\,\mathbf{x} - 2A^\top\mathbf{b} = \mathbf{0} \;\;\Longrightarrow\;\; A^\top A\,\mathbf{x} = A^\top\mathbf{b}.$$Same equations again. (Why is this flat spot a minimum and not a maximum? $E$ is a sum of squares, so it can never go below zero and it opens upward like a bowl. Chapters 1.12 and 1.14 make "bowl" precise.)
When $A^\top A$ is invertible (next section), the solution has a formula, and the fitted values come from the projection matrix of Chapter 1.9:
$$\hat{\mathbf{x}} = (A^\top A)^{-1}A^\top\mathbf{b}, \qquad A\hat{\mathbf{x}} = \underbrace{A(A^\top A)^{-1}A^\top}_{P}\,\mathbf{b}.$$Why do we need it?
A picture is not a program. We need an equation that a computer can solve to get the best weights directly, without searching around.
Where is it used?
The closed-form solution of linear and ridge regression, the Gauss–Newton step used for nonlinear curve fitting, and the update formulas of Kalman filters and Gaussian processes all contain AᵀA x = Aᵀb.
How is it used?
Form AᵀA and Aᵀb, then solve the small square system (do not invert). Use it for small, well-behaved problems, and as a recipe to derive new formulas: write the error, set its gradient to zero.
Do not memorise $\hat{\mathbf{x}}=(A^\top A)^{-1}A^\top\mathbf{b}$ as a recipe for computing. It is a great formula for thinking and for proofs. On a computer you should solve the system instead of inverting the matrix, and you will see in a moment that forming $A^\top A$ can itself lose accuracy.
Where the name comes from. The equations are "normal" because the residual is normal (perpendicular) to the column space. They have nothing to do with the "normal distribution".
A free fact. If $A$ has a column of ones, the first normal equation says $\mathbf{1}\cdot\mathbf{r}=0$: the residuals add up to zero. We saw this in the running example: $-\tfrac16+\tfrac13-\tfrac16 = 0$.
Quick check: a model has a single unknown, the constant $x$, so $A = [1,1,1]^\top$. Solve the normal equations for $\mathbf{b}=[2,4,9]$.
$A^\top A = 1+1+1 = 3$ and $A^\top\mathbf{b} = 2+4+9 = 15$, so $3\hat{x} = 15$ and $\hat{x}=5$. That is the average of the data. The mean is the least-squares answer to "what single number best represents these values?".
When does the best answer have to be unique?
Suppose a house-price model has two features: the area in square metres and the same area in square feet. The second column is just the first column times 10.76. The data can never tell the model how to split credit between them. "5 units on metres and 0 on feet" gives exactly the same predictions as "0 on metres and 0.46 on feet". There are infinitely many equally good answers, and the normal equations cannot pick one.
That is what goes wrong when features are copies (or blends) of each other. The fix is a requirement: the columns of $A$ must be independent, so that each one brings something new.
Take $A = \begin{bmatrix} 1&2\\2&4\\3&6 \end{bmatrix}$. The second column is twice the first.
- Dot products: $\mathbf{a}_1\!\cdot\!\mathbf{a}_1 = 14$, $\mathbf{a}_1\!\cdot\!\mathbf{a}_2 = 28$, $\mathbf{a}_2\!\cdot\!\mathbf{a}_2 = 56$. So $A^\top A = \begin{bmatrix} 14&28\\28&56 \end{bmatrix}$.
- Its determinant is $14\cdot56 - 28\cdot28 = 784 - 784 = 0$. The matrix is singular, so it has no inverse.
- Meaning: the prediction is $A\mathbf{x} = (x_1 + 2x_2)\,[1,2,3]$. Every pair with the same value of $x_1+2x_2$, like $(2,0)$, $(0,1)$ and $(1, 0.5)$, predicts the same thing and has the same error.
For an $m\times n$ matrix $A$, the following are all equivalent:
- $A^\top A$ is invertible.
- $A$ has full column rank: $\operatorname{rank}(A) = n$, so the columns are linearly independent.
- The only solution of $A\mathbf{x}=\mathbf{0}$ is $\mathbf{x}=\mathbf{0}$ (the null space is trivial).
- The least-squares solution $\hat{\mathbf{x}}$ is unique.
Why? The key is that $A\mathbf{x}$ and $A^\top A\mathbf{x}$ have the same zeros. If $A\mathbf{x}=\mathbf{0}$ then clearly $A^\top A\mathbf{x}=\mathbf{0}$. Conversely, if $A^\top A\mathbf{x}=\mathbf{0}$, multiply by $\mathbf{x}^\top$ to get $\mathbf{x}^\top A^\top A\mathbf{x} = \|A\mathbf{x}\|^2 = 0$, which forces $A\mathbf{x}=\mathbf{0}$. So $N(A^\top A) = N(A)$, and a matrix with a trivial null space is invertible.
If you have fewer rows than columns ($m\lt n$, fewer data points than unknowns), full column rank is impossible, so the solution is never unique. When the columns are dependent, least squares still has a best fit $A\hat{\mathbf{x}}$ (the projection is always unique), but many different $\hat{\mathbf{x}}$ produce it. The pseudoinverse and ridge sections below show two standard ways to pick one.
Why do we need it?
Before trusting an answer we must know it is the only one. If columns of A repeat each other, many weight vectors fit equally well, and the individual weights mean very little.
Where is it used?
Spotting the dummy-variable trap in one-hot encoding, perfectly correlated columns in a regression table, and having fewer data points than parameters.
How is it used?
Compute rank(A) (np.linalg.matrix_rank). If it is smaller than the number of columns, drop or merge redundant features, or add ridge. Also print the condition number: a huge value warns of near-duplicates.
Nearly dependent is almost as bad as dependent. If two features are almost copies, $A^\top A$ is technically invertible but extremely sensitive: tiny noise in the data flips the weights wildly (one huge positive, one huge negative). Computers also round off numbers, so "almost singular" can become "singular" inside the machine. The remedies are in the sections below: QR/SVD for accuracy, and ridge to tame the weights.
Quick check: a dataset has 3 features but only 2 examples. Can $A^\top A$ be invertible?
No. $A$ is $2\times3$, so its rank is at most 2, which is less than the number of columns 3. Then $A^\top A$ (a $3\times3$ matrix) has rank at most 2 and is singular.
Solving it, method 1: normal equations (and their weak spot) core
The normal equations are the "obvious" way to compute least squares: multiply out $A^\top A$ and $A^\top\mathbf{b}$, then solve a small square system. It is short, fast and fine for many problems.
But it has a hidden weakness. Think of making a photocopy of a photocopy: each copy loses a bit of quality. The matrix $A^\top A$ is like a "photocopy of $A$ against itself", and it is twice as sensitive to rounding errors and noise as $A$ is.
The number that measures sensitivity is the condition number $\kappa$ (from Chapter 1.7). A matrix with $\kappa = 10^k$ can cost you about $k$ digits of accuracy. A computer keeps roughly 16 digits. So a problem with $\kappa(A)=10^6$ should still leave 10 good digits, but $A^\top A$ has $\kappa = 10^{12}$ and leaves only about 4.
For our running matrix $A = \begin{bmatrix} 1&1\\1&2\\1&3 \end{bmatrix}$ a computer finds that $A$ stretches vectors by at most $4.079$ and at least $0.600$ (these are its singular values, see Chapter 1.13).
- $\kappa(A) = 4.079 / 0.600 \approx 6.79$.
- $A^\top A = \begin{bmatrix} 3&6\\6&14 \end{bmatrix}$ stretches by at most $4.079^2 \approx 16.64$ and at least $0.600^2\approx 0.36$.
- $\kappa(A^\top A) = 16.64 / 0.36 \approx 46.1$, and indeed $6.79^2 = 46.1$.
Here that is harmless. But if $\kappa(A)$ were $10^8$, then $\kappa(A^\top A)$ would be $10^{16}$, which is the size of the computer's whole precision: every digit of the answer would be lost.
Solving by normal equations: form $G = A^\top A$ and $\mathbf{h} = A^\top\mathbf{b}$, then solve $G\hat{\mathbf{x}}=\mathbf{h}$ (never by explicitly computing $G^{-1}$).
The squaring rule:
$$\kappa(A^\top A) = \kappa(A)^2.$$Why? $A$ stretches every direction by an amount between $\sigma_{\min}$ and $\sigma_{\max}$, so $\kappa(A) = \sigma_{\max}/\sigma_{\min}$. The matrix $A^\top A$ first applies $A$ and then applies $A^\top$, which stretches by the same amounts again. The stretches multiply, giving $\sigma_{\min}^2$ and $\sigma_{\max}^2$, so the ratio is squared. (We will prove this properly with the SVD in Chapter 1.13.)
Rule of thumb: you lose about $\log_{10}\kappa$ digits. Normal equations lose about $2\log_{10}\kappa(A)$ digits. The methods in the next section lose only $\log_{10}\kappa(A)$.
Why do we need it?
We want a fast, simple method for everyday problems, and we also need to know exactly when it silently loses accuracy.
Where is it used?
Quick regression on a few well-scaled features, big-data fitting where AᵀA is added up batch by batch, and teaching examples. Not for badly scaled polynomial features or very correlated columns.
How is it used?
Scale your features, form AᵀA and Aᵀb, and solve with np.linalg.solve. Print np.linalg.cond(A): if it is above about 10⁴ to 10⁵, switch to QR or SVD.
This is a numerical problem, not a math problem. On paper, with exact arithmetic, the normal equations always give the right answer. The danger appears only with rounding, and only when $\kappa(A)$ is large (nearly dependent columns, or polynomial features of high degree). For well-behaved data, normal equations are perfectly fine. Know the risk, then decide.
Quick check: if $\kappa(A) = 10^5$, what is $\kappa(A^\top A)$, and roughly how many of 16 digits survive in each approach?
$\kappa(A^\top A) = (10^5)^2 = 10^{10}$. Using $A$ directly (QR or SVD) you lose about 5 digits and keep about 11. Using the normal equations you lose about 10 digits and keep only about 6.
Solving it, methods 2 and 3: QR and SVD core
The normal equations are messy because the columns of $A$ lean on each other: $A^\top A$ is full of cross terms. If the columns were already perpendicular unit arrows (an orthonormal set from Chapter 1.9), the cross terms would vanish and the answer would be nothing but a few dot products.
QR does exactly that. It straightens the columns first (that is Gram–Schmidt) and keeps a record $R$ of how it straightened them. Then finding the best weights is a gentle back-substitution (solving the equations from the last one upward, one unknown at a time). We never form $A^\top A$, so the condition number is not squared.
SVD goes one step further. It describes $A$ as rotate, stretch, rotate. In those rotated coordinates the problem splits into independent one-number problems. It also copes gracefully when columns are dependent. It is the most robust tool, and the price is a little more computing.
QR on the running example. Orthogonalise the columns $\mathbf{a}_1=[1,1,1]$ and $\mathbf{a}_2=[1,2,3]$ (the Gram–Schmidt steps of Chapter 1.9):
- $\|\mathbf{a}_1\| = \sqrt3$, so $\mathbf{q}_1 = [1,1,1]/\sqrt3$. This gives $R_{11} = \sqrt3 \approx 1.732$.
- $R_{12} = \mathbf{q}_1\cdot\mathbf{a}_2 = 6/\sqrt3 = 2\sqrt3\approx 3.464$.
- Remove that part: $\mathbf{a}_2 - R_{12}\mathbf{q}_1 = [1,2,3] - 2[1,1,1] = [-1, 0, 1]$. Its length is $\sqrt2$, so $R_{22}=\sqrt2\approx1.414$ and $\mathbf{q}_2 = [-1,0,1]/\sqrt2$.
- Dot $\mathbf{b} = [1,2,2]$ with each $\mathbf{q}$: $\mathbf{q}_1\cdot\mathbf{b} = 5/\sqrt3 \approx 2.887$ and $\mathbf{q}_2\cdot\mathbf{b} = 1/\sqrt2 \approx 0.707$. So $Q^\top\mathbf{b} = [2.887,\; 0.707]$.
- Solve $R\mathbf{x} = Q^\top\mathbf{b}$ from the bottom up: $1.414\,m = 0.707$ gives $m = 0.5$. Then $1.732\,c + 3.464\cdot0.5 = 2.887$ gives $c = (2.887-1.732)/1.732 = 0.667$.
Same answer, $\hat{\mathbf{x}}=[\tfrac23,\tfrac12]$, with no $A^\top A$ in sight.
Least squares by QR. Write $A = QR$ (reduced form): $Q$ is $m\times n$ with orthonormal columns, $R$ is $n\times n$ upper triangular (invertible when $A$ has full column rank). Put this into the normal equations:
$$A^\top A\hat{\mathbf{x}} = A^\top\mathbf{b} \;\Rightarrow\; R^\top\underbrace{Q^\top Q}_{I}R\hat{\mathbf{x}} = R^\top Q^\top\mathbf{b} \;\Rightarrow\; R^\top R\hat{\mathbf{x}} = R^\top Q^\top\mathbf{b} \;\Rightarrow\; \boxed{R\hat{\mathbf{x}} = Q^\top\mathbf{b}}$$(In the last step we cancelled the invertible $R^\top$.) Solve the triangular system by back-substitution. Geometrically, $Q^\top\mathbf{b}$ lists the coordinates of the projection of $\mathbf{b}$ in the orthonormal basis $\mathbf{q}_1,\dots,\mathbf{q}_n$ of the column space. Software uses numerically careful variants (Householder reflections) rather than plain Gram–Schmidt.
Least squares by SVD. Any matrix can be written $A = U\Sigma V^\top$. For a matrix with independent columns, the thin form is: $U$ is $m\times n$ with orthonormal columns, $\Sigma$ is diagonal with numbers $\sigma_1\ge\dots\ge\sigma_n>0$, $V$ is an $n\times n$ rotation/reflection. Full story in Chapter 1.13; here we only need that shape. Since rotations do not change lengths (Chapter 1.9), splitting $\mathbf{b}$ into its part inside $C(A)$ and the rest gives
$$\|A\mathbf{x}-\mathbf{b}\|^2 = \|\Sigma V^\top\mathbf{x} - U^\top\mathbf{b}\|^2 + \|\mathbf{b} - UU^\top\mathbf{b}\|^2.$$The second term does not involve $\mathbf{x}$. The first becomes $0$ when $\Sigma\,(V^\top\mathbf{x}) = U^\top\mathbf{b}$, which is a diagonal system, solved one entry at a time:
$$\hat{\mathbf{x}} = V\,\Sigma^{-1}U^\top\mathbf{b} = \sum_{i=1}^{n}\frac{\mathbf{u}_i^\top\mathbf{b}}{\sigma_i}\,\mathbf{v}_i.$$Each term says: "how much of $\mathbf{b}$ points along $\mathbf{u}_i$, divided by how strongly $A$ stretches that direction, times the matching direction $\mathbf{v}_i$." Tiny $\sigma_i$ are the dangerous ones (dividing by almost zero), and SVD lets you see and control them.
Why do we need it?
We want the same answer without squaring the condition number, and a method that survives nearly dependent columns.
Where is it used?
np.linalg.lstsq and scikit-learn's LinearRegression (SVD inside), statistics packages (QR inside), polynomial fitting, and any regression with correlated features.
How is it used?
Call np.linalg.lstsq(A, b, rcond=None). For a manual QR: Q, R = np.linalg.qr(A), then solve R x = Qᵀb by back-substitution. Look at the returned rank and residual size.
Quick comparison. Normal equations: fastest, but squares $\kappa$ (loses about twice the digits). QR: a bit slower, accuracy limited by $\kappa(A)$ only; the standard choice for full-rank problems. SVD: slowest, most robust, and the only one of the three that handles dependent columns smoothly. NumPy's np.linalg.lstsq uses the SVD.
- Never write
inv(A.T @ A) @ A.T @ bin real code. Usenp.linalg.lstsq(A, b)(SVD-based) ornp.linalg.solveon the normal equations only when you know the problem is well-conditioned. - Plain classical Gram–Schmidt can lose orthogonality on nearly dependent columns (Chapter 1.9). Real QR routines use Householder reflections for exactly this reason.
- Basic QR still needs independent columns. For a rank-deficient $A$, use the SVD (or QR with column pivoting).
Quick check: in the QR route, why is solving $R\hat{\mathbf{x}}=Q^\top\mathbf{b}$ easy?
$R$ is upper triangular, so the last equation has only one unknown. Solve it, plug into the one above, and so on up to the top. This is back-substitution, and it costs very little.
The pseudoinverse $A^+$ core
An inverse $A^{-1}$ is an "undo button": $A^{-1}(A\mathbf{x}) = \mathbf{x}$. But only square, independent matrices have one. What if $A$ is tall, or wide, or has dependent columns? We still want the best possible undo button. That is the pseudoinverse (or Moore–Penrose inverse), written $A^+$.
- If $A$ is a normal invertible matrix, $A^+$ is just $A^{-1}$.
- If $A$ is tall (more equations than unknowns), you cannot undo exactly, but $A^+\mathbf{b}$ gives the least-squares answer.
- If $A$ is wide or has dependent columns (many answers fit), $A^+\mathbf{b}$ picks the shortest of the best answers.
In one sentence: $A^+\mathbf{b}$ is the shortest vector among all the vectors that make $\|A\mathbf{x}-\mathbf{b}\|$ smallest.
Tall matrix. For our running $A$ we found $(A^\top A)^{-1} = \tfrac16\begin{bmatrix} 14&-6\\-6&3 \end{bmatrix}$. Multiply by $A^\top = \begin{bmatrix} 1&1&1\\1&2&3 \end{bmatrix}$:
$$A^+ = (A^\top A)^{-1}A^\top = \frac16\begin{bmatrix} 8 & 2 & -4 \\ -3 & 0 & 3 \end{bmatrix} = \begin{bmatrix} \tfrac43 & \tfrac13 & -\tfrac23 \\[2pt] -\tfrac12 & 0 & \tfrac12 \end{bmatrix}.$$- Apply it to $\mathbf{b} = [1,2,2]$. First entry: $\tfrac43\cdot1 + \tfrac13\cdot2 - \tfrac23\cdot2 = \tfrac43+\tfrac23-\tfrac43 = \tfrac23$.
- Second entry: $-\tfrac12\cdot1 + 0\cdot2 + \tfrac12\cdot2 = \tfrac12$.
- So $A^+\mathbf{b} = [\tfrac23, \tfrac12] = \hat{\mathbf{x}}$, the least-squares solution. ✓
- $A^+$ is a left inverse: $A^+A = I$ (check the first row: $\tfrac43+\tfrac13-\tfrac23 = 1$ and $\tfrac43+\tfrac23-2 = 0$). But $AA^+\ne I$: it is the projection matrix $P$ of Chapter 1.9.
A matrix with a zero stretch. For $A = \begin{bmatrix} 2&0\\0&0 \end{bmatrix}$ the second direction is crushed to nothing, and nothing can bring it back. The pseudoinverse inverts what can be inverted and leaves the rest at zero: $A^+ = \begin{bmatrix} \tfrac12&0\\0&0 \end{bmatrix}$. That is the whole idea of the SVD formula below.
The (Moore–Penrose) pseudoinverse $A^+$ of any $m\times n$ matrix $A$ is the unique $n\times m$ matrix satisfying the four Penrose conditions:
$$AA^+A = A,\qquad A^+AA^+ = A^+,\qquad (AA^+)^\top = AA^+,\qquad (A^+A)^\top = A^+A.$$(You do not need to memorise these. They say "$A^+$ undoes $A$ as well as possible, and $AA^+$ and $A^+A$ are orthogonal projections".) Four ways to get it:
- Invertible square $A$: $A^+ = A^{-1}$.
- Tall, full column rank (independent columns): $A^+ = (A^\top A)^{-1}A^\top$. This is the least-squares formula.
- Wide, full row rank (independent rows): $A^+ = A^\top(AA^\top)^{-1}$. This gives the minimum-norm solution (next section).
- Any matrix, via the SVD: if $A = U\Sigma V^\top$, then
where $\Sigma^+$ is made from $\Sigma$ by flipping its shape (transpose) and replacing every non-zero $\sigma_i$ by $1/\sigma_i$ (zeros stay zero). Compare with the SVD least-squares formula in the previous section: it is the same thing, now allowed to skip zero singular values.
Useful facts: $AA^+$ is the orthogonal projection onto the column space $C(A)$; $A^+A$ is the orthogonal projection onto the row space $C(A^\top)$; and $(A^+)^+ = A$. For any $\mathbf{b}$, $\mathbf{x}^+ = A^+\mathbf{b}$ is the least-squares solution of smallest norm.
Why do we need it?
Matrices from real data are rarely square or invertible, yet we still want an "undo button". The pseudoinverse gives one clear answer in every case.
Where is it used?
np.linalg.pinv and lstsq, linear models with dependent features, inverse problems such as image deblurring and CT scans, recommender systems, and robot-arm control (pseudoinverse of the Jacobian).
How is it used?
Take the SVD, flip each non-zero singular value, ignore tiny ones (set a cut-off), and multiply: x = pinv(A) @ b. Check A⁺A and AA⁺ to see which directions were lost.
- Never compute $A^+$ for real work by forming $(A^\top A)^{-1}$ (you know why now). Use
np.linalg.pinv, which uses the SVD. - The SVD formula needs a cut-off to decide what counts as "zero". A singular value of $10^{-17}$ is zero up to rounding, and dividing by it would blow up. NumPy's
pinvhas anrcondsetting for this. - Even when a true inverse exists, solving $A\mathbf{x}=\mathbf{b}$ is better done with a solver than with $A^{+}\mathbf{b}$.
Quick check: what is $A^+$ for $A = \begin{bmatrix} 5&0\\0&0.1 \end{bmatrix}$? And for $A=\begin{bmatrix} 5&0\\0&0 \end{bmatrix}$?
For the first, $A$ is invertible, so $A^+ = A^{-1} = \operatorname{diag}(\tfrac15, 10)$. For the second, invert the non-zero entry and leave the zero alone: $A^+ = \operatorname{diag}(\tfrac15, 0)$.
Too few equations: the minimum-norm solution core
Now the opposite situation: fewer equations than unknowns. One equation with two unknowns, like $3x+4y=10$, has a whole line of exact solutions. Nothing is wrong, there is just no single best answer. How do we pick one?
A natural rule: choose the smallest solution, the one closest to the origin. It is the "laziest" answer: it uses the least total weight to do the job. On a line, the closest point to the origin is where you drop a perpendicular from the origin onto the line.
Solve $3x + 4y = 10$ with the smallest $\|[x,y]\|$.
- Some solutions: $(2,1)$ since $6+4=10$; $(\tfrac{10}3, 0)$; $(0, 2.5)$. Their lengths: $\sqrt5 \approx 2.24$, $3.33$, $2.5$.
- The shortest one points along the row $[3,4]$ itself (perpendicular to the solution line). Write it as $[x,y] = t[3,4]$.
- Plug in: $3(3t)+4(4t) = 25t = 10$, so $t = 0.4$ and $[x,y] = [1.2, 1.6]$.
- Its length is $\sqrt{1.44+2.56} = \sqrt4 = 2$. That beats all the others. ✓
In matrix words: $A=[3\;\;4]$, $AA^\top = 25$, and $\mathbf{x}^+ = A^\top(AA^\top)^{-1}\cdot10 = [3,4]\cdot\tfrac{10}{25} = [1.2, 1.6]$.
For an underdetermined consistent system $A\mathbf{x}=\mathbf{b}$ (full row rank, $m\lt n$), the minimum-norm solution is
$$\mathbf{x}^+ = A^\top(AA^\top)^{-1}\mathbf{b} = A^+\mathbf{b}, \qquad \text{the solution of } \min\|\mathbf{x}\| \text{ subject to } A\mathbf{x}=\mathbf{b}.$$Why. Every vector splits into a part in the row space $C(A^\top)$ and a part in the null space $N(A)$, and these two parts are perpendicular (Chapter 1.8). The null-space part is invisible to $A$ ($A\mathbf{n}=\mathbf{0}$), so every solution is $\mathbf{x}^+ + \mathbf{n}$ for some $\mathbf{n}\in N(A)$. By Pythagoras,
$$\|\mathbf{x}^+ + \mathbf{n}\|^2 = \|\mathbf{x}^+\|^2 + \|\mathbf{n}\|^2 \;\ge\; \|\mathbf{x}^+\|^2,$$so the shortest solution has no null-space part at all. It lies in the row space: $\mathbf{x}=A^\top\mathbf{y}$ for some $\mathbf{y}$. Then $A\mathbf{x}=\mathbf{b}$ reads $AA^\top\mathbf{y}=\mathbf{b}$, so $\mathbf{y}=(AA^\top)^{-1}\mathbf{b}$, which gives the formula.
For a general matrix (any shape, any rank), $A^+\mathbf{b}$ first picks the best-fitting predictions (least squares) and then, among all weights that give them, the shortest.
Why do we need it?
With fewer equations than unknowns there are infinitely many exact solutions. We need a rule to pick one, and "the smallest" is the simplest rule that gives a unique answer.
Where is it used?
Overparameterised models (more weights than examples), compressed sensing and image reconstruction, control (the least effort that reaches a target), and the solution that gradient descent finds when started from zero.
How is it used?
Compute x = np.linalg.pinv(A) @ b (or Aᵀ(AAᵀ)⁻¹b when the rows are independent). Check that Ax = b holds, and remember that adding any null-space vector can only make ‖x‖ bigger.
- "Shortest" depends on the norm you choose. With the L1 norm (Chapter 1.2) the shortest solution would be a different, sparse point sitting on an axis. L2, the one we used here, always gives the straight perpendicular foot.
- $A^\top(AA^\top)^{-1}$ requires independent rows. If the rows are dependent too, use the SVD form of $A^+$.
Quick check: find the minimum-norm solution of $x + y + z = 3$.
$A = [1\;1\;1]$ and $AA^\top = 3$, so $\mathbf{x}^+ = A^\top\cdot\tfrac33 = [1,1,1]$. By symmetry the shortest answer shares the total equally. Its length is $\sqrt3$, shorter than, say, $[3,0,0]$ with length $3$.
Regularised least squares: ridge core
When two features are nearly copies of each other, plain least squares can behave oddly. To fit the noise, it may give one feature a huge positive weight and the other a huge negative weight, so that they cancel. The fit looks fine on the training data, but the weights are absurd and a tiny change in the data flips them.
The cure is to put the weights on a leash. We still want a small error, but we also pay a price for large weights. The model must now balance two goals: fit the data and keep the weights small. A knob $\lambda$ (lambda) decides how short the leash is: $\lambda=0$ means no leash, and a huge $\lambda$ pulls all weights to zero.
This is ridge regression, also called L2 regularisation or Tikhonov regularisation.
One weight. Let $A = [1, 2]^\top$ and $\mathbf{b} = [2, 3]$. Then $A^\top A = 5$ and $A^\top\mathbf{b} = 8$. Ridge solves $(5+\lambda)\,x = 8$.
- $\lambda = 0$: $x = 8/5 = 1.6$ (ordinary least squares).
- $\lambda = 5$: $x = 8/10 = 0.8$.
- $\lambda = 15$: $x = 8/20 = 0.4$. A bigger $\lambda$ shrinks the weight toward $0$.
A broken problem repaired. Let $A = \begin{bmatrix} 1&1\\2&2 \end{bmatrix}$ (two identical columns). Then $A^\top A = \begin{bmatrix} 5&5\\5&5 \end{bmatrix}$ has determinant $0$: not invertible. Add $\lambda = 1$: $A^\top A + I = \begin{bmatrix} 6&5\\5&6 \end{bmatrix}$ has determinant $36-25 = 11$. It is invertible, so there is exactly one answer.
Ridge regression minimises the squared error plus a penalty on the squared size of the weights:
$$\hat{\mathbf{x}}_\lambda = \arg\min_{\mathbf{x}}\;\|A\mathbf{x}-\mathbf{b}\|^2 + \lambda\|\mathbf{x}\|^2, \qquad \lambda \ge 0.$$Setting the gradient to zero (the same calculus as before) gives $2A^\top(A\mathbf{x}-\mathbf{b}) + 2\lambda\mathbf{x} = \mathbf{0}$, so
$$\boxed{(A^\top A + \lambda I)\,\hat{\mathbf{x}}_\lambda = A^\top\mathbf{b}} \qquad\Longrightarrow\qquad \hat{\mathbf{x}}_\lambda = (A^\top A + \lambda I)^{-1}A^\top\mathbf{b}.$$Why this fixes singular or ill-conditioned $A^\top A$.
- It is always invertible when $\lambda>0$. For any non-zero $\mathbf{x}$: $\mathbf{x}^\top(A^\top A+\lambda I)\mathbf{x} = \|A\mathbf{x}\|^2 + \lambda\|\mathbf{x}\|^2 > 0$. So no non-zero vector is sent to zero, and the matrix has a trivial null space. Even duplicate columns are fine.
- It improves the condition number. If $A^\top A$ stretches by $s_{\max}$ and $s_{\min}$ (these are $\sigma^2$ of the singular values), then $A^\top A+\lambda I$ stretches by $s_{\max}+\lambda$ and $s_{\min}+\lambda$, so $\kappa = \dfrac{s_{\max}+\lambda}{s_{\min}+\lambda}$. Adding $\lambda$ lifts the bottom much more (in relative terms) than the top. In the example above: $\kappa = 10/0=\infty$ becomes $11/1 = 11$ with $\lambda=1$ and $2$ with $\lambda=10$.
- It shrinks the dangerous directions. In the SVD formula, each term $\frac{\mathbf{u}_i^\top\mathbf{b}}{\sigma_i}\mathbf{v}_i$ becomes $\frac{\sigma_i}{\sigma_i^2+\lambda}(\mathbf{u}_i^\top\mathbf{b})\mathbf{v}_i$. For big $\sigma_i$ nothing changes. For tiny $\sigma_i$ the huge $1/\sigma_i$ is replaced by something small. Exactly the noisy directions get damped.
Another view. Ridge is ordinary least squares on extra made-up data: stack $\sqrt\lambda\,I$ under $A$ and zeros under $\mathbf{b}$. The extra rows say "also, please make each weight about 0". Its normal equations are exactly $(A^\top A+\lambda I)\mathbf{x}=A^\top\mathbf{b}$. As $\lambda\to0^+$ the ridge solution tends to the pseudoinverse solution $A^+\mathbf{b}$.
Why do we need it?
When features are nearly copies, plain least squares gives huge, unstable weights that chase the noise. We need a way to keep the weights modest and the problem solvable.
Where is it used?
Ridge regression in scikit-learn, weight decay in neural-network training, regularised recommender models, and any regression with many correlated features (genetics, text, finance).
How is it used?
Standardise the features, leave the bias unpenalised, pick λ by cross-validation (try 10⁻⁴ up to 10³), and solve (AᵀA + λI)x = Aᵀb. Choose the λ where the error on held-out data is lowest.
- Do not penalise the bias (intercept). The leash should shrink the effect of features, not drag the average prediction to zero. In practice, centre the data or leave the bias column out of the penalty.
- Scale your features first. The penalty $\lambda\|\mathbf{x}\|^2$ treats all weights alike, so a feature measured in tiny units (needing a huge weight) is punished unfairly. Standardise the columns before using ridge.
- The best $\lambda$ is not known in advance. It is chosen with a validation set or cross-validation (trying several values and checking error on held-out data), exactly as the test-error curve above suggests.
Quick check: why can ridge never produce a singular matrix when $\lambda>0$?
For any non-zero $\mathbf{x}$, $\mathbf{x}^\top(A^\top A+\lambda I)\mathbf{x} = \|A\mathbf{x}\|^2+\lambda\|\mathbf{x}\|^2$, which is strictly positive because $\lambda\|\mathbf{x}\|^2>0$. So $(A^\top A+\lambda I)\mathbf{x}\neq\mathbf{0}$ for every non-zero $\mathbf{x}$, meaning the matrix has no null space and is invertible.
Weighted least squares
Imagine weighing an object twice. One scale is precise, the other is sloppy. Averaging the two readings equally treats them as equally trustworthy, which they are not. Better: give the precise scale more votes.
Weighted least squares does this for data points. Each point $i$ gets a weight $w_i\ge0$ saying how much we trust it, and a miss at a trusted point costs more. (This is an awareness topic: you should know it exists and what it looks like.)
Two measurements of the same length: $10$ (precise scale, weight 3) and $12$ (sloppy scale, weight 1). One unknown $x$, so $A=[1,1]^\top$, $\mathbf{b}=[10,12]$.
- Minimise $3(10-x)^2 + 1\cdot(12-x)^2$.
- In matrix form with $W=\operatorname{diag}(3,1)$: $A^\top WA = 3+1 = 4$ and $A^\top W\mathbf{b} = 30+12 = 42$.
- So $x = 42/4 = 10.5$: a weighted average, closer to the trusted reading. (Equal weights would give 11.)
With a diagonal weight matrix $W=\operatorname{diag}(w_1,\dots,w_m)$, weighted least squares minimises
$$\sum_i w_i\,r_i^2 = (A\mathbf{x}-\mathbf{b})^\top W (A\mathbf{x}-\mathbf{b}) \qquad\Longrightarrow\qquad (A^\top W A)\,\hat{\mathbf{x}} = A^\top W\mathbf{b}.$$Trick: multiply each row of $A$ and each entry of $\mathbf{b}$ by $\sqrt{w_i}$. Then ordinary least squares on the scaled data is the weighted problem, so every solver above still works. A common choice is $w_i = 1/\sigma_i^2$, where $\sigma_i$ is the noise level of measurement $i$ ("inverse-variance weighting").
Why do we need it?
Not all measurements are equally reliable. Plain least squares treats them all alike, so one noisy reading can wreck the fit. Weights let trusted points count more.
Where is it used?
Sensor fusion and Kalman filtering, regression with different noise levels, sample weights for imbalanced data, and iteratively reweighted least squares for logistic and robust regression.
How is it used?
Choose a weight per point (often 1 divided by its noise variance), multiply each row of A and each entry of b by the square root of its weight, then run ordinary least squares. In scikit-learn, pass sample_weight.
Quick check: three readings 4, 6, 8 of one quantity, with weights 1, 1, 2. What is the weighted estimate?
$A^\top WA = 1+1+2 = 4$, and $A^\top W\mathbf{b} = 4+6+16 = 26$. So $x = 26/4 = 6.5$ (the plain average would be 6). The last reading counts double and pulls the answer up.
Least-squares regression: the design matrix core
Here is how a table of data becomes a least-squares problem. Each row of the table is one example (one house, one patient). Each column is one measurement (a feature): area, number of bedrooms, and so on. The thing we want to predict goes in a separate column of targets $\mathbf{b}$.
A linear model predicts each target as a weighted sum of that row's features, plus a starting amount (the bias or intercept). To write "plus a starting amount" as a plain dot product, we add one extra column to the table that is all ones. The weight on that column is the bias: it gets multiplied by 1 every time.
With that trick, the predictions for all examples at once are simply the matrix times the weight vector: $A\mathbf{w}$. Then fitting the model is exactly the least-squares problem you have just learned.
Fitting a plane. Four examples, each with two features $(x_1, x_2)$ and a target $y$: $(0,0)\to1$, $(1,0)\to2$, $(0,1)\to2$, $(1,1)\to4$. We want $y\approx w_0 + w_1x_1 + w_2x_2$.
$$A = \begin{bmatrix} 1&0&0\\1&1&0\\1&0&1\\1&1&1 \end{bmatrix},\quad \mathbf{b} = \begin{bmatrix} 1\\2\\2\\4 \end{bmatrix},\quad A^\top A = \begin{bmatrix} 4&2&2\\2&2&1\\2&1&2 \end{bmatrix},\quad A^\top\mathbf{b} = \begin{bmatrix} 9\\6\\6 \end{bmatrix}.$$- The equations are $4w_0+2w_1+2w_2=9$, $2w_0+2w_1+w_2=6$, $2w_0+w_1+2w_2=6$.
- Subtract the third from the second: $w_1 - w_2 = 0$, so $w_1=w_2=:w$.
- Now $4w_0+4w=9$ and $2w_0+3w=6$. Double the second: $4w_0+6w=12$. Subtract the first: $2w=3$, so $w=1.5$. Then $4w_0 = 9-6 = 3$, so $w_0=0.75$.
- The plane is $y = 0.75 + 1.5x_1 + 1.5x_2$. Predictions: $[0.75,\,2.25,\,2.25,\,3.75]$. Residuals: $[0.25,\,-0.25,\,-0.25,\,0.25]$. They add to zero (there is a bias column) and are perpendicular to each column. ✓
The straight-line fit of the first section is the same recipe with a single feature: $A$ has columns $[1, x]$.
Given $n$ examples with $p$ features, put example $i$ in row $i$ of the design matrix
$$A = \begin{bmatrix} 1 & x_{11} & x_{12} & \cdots & x_{1p}\\ 1 & x_{21} & x_{22} & \cdots & x_{2p}\\ \vdots & & & & \vdots\\ 1 & x_{n1} & x_{n2} & \cdots & x_{np} \end{bmatrix}\in\mathbb{R}^{n\times(p+1)},\qquad \mathbf{w}=\begin{bmatrix} w_0\\w_1\\\vdots\\w_p \end{bmatrix},\qquad \hat{\mathbf{y}}=A\mathbf{w}.$$The first column of ones carries the bias $w_0$. Linear regression chooses $\hat{\mathbf{w}} = \arg\min\|A\mathbf{w}-\mathbf{y}\|^2$, i.e. $A^\top A\hat{\mathbf{w}}=A^\top\mathbf{y}$. Fitting a line ($p=1$) draws a line, fitting two features ($p=2$) draws a plane, and with more features it is a flat "hyperplane" that we cannot draw but can still compute.
Reading the weights: $w_j$ says how much the prediction changes when feature $j$ rises by 1 and the other features stay put. The column space $C(A)$ is "all predictions the model can make", and $\hat{\mathbf{y}}$ is the projection of the target vector $\mathbf{y}$ onto it.
Why do we need it?
We need a standard way to turn a data table into a matrix problem. The design matrix, with its column of ones for the bias, is that translation.
Where is it used?
Every linear regression: house-price models, dose–response curves in medicine, A/B-test analysis, calibration lines, and the linear layer of a neural network (weights plus bias).
How is it used?
Put examples in rows and features in columns, add a column of ones, solve with lstsq, then read the weights (effect per feature) and plot the residuals. Standardise the columns first if their scales differ.
- Do not forget the column of ones. Without it the model is forced through the origin, which is rarely what you want.
- Residuals are measured vertically (in the target direction), not as the shortest distance to the line. Least squares assumes the features are exact and the target is noisy.
- Features on very different scales (square metres versus number of bedrooms) make $A$ badly conditioned. Standardise columns first.
Quick check: you predict price from area, bedrooms and age using 200 houses. How many rows and columns does the design matrix have?
200 rows (one per house) and $3+1 = 4$ columns (three features plus the column of ones for the bias). Shape $200\times4$.
Curves from straight-line machinery: polynomial features core
"Linear" regression can fit curves. The word linear refers to the weights, not to $x$. A parabola $y = w_0 + w_1x + w_2x^2$ is still a weighted sum of "ingredients" $1$, $x$ and $x^2$. We simply add new columns to the design matrix: one column holding $x^2$ for every example, and so on. Least squares does not know or care that these columns came from the same $x$.
More columns give a more flexible curve. That is a double-edged sword: with enough columns the curve can pass through every training point exactly while wiggling wildly in between. This is overfitting.
Fit a parabola to the four points $(-1,2)$, $(0,1)$, $(1,2)$, $(2,5)$. The design matrix gets a third column with the squares:
$$A = \begin{bmatrix} 1&-1&1\\1&0&0\\1&1&1\\1&2&4 \end{bmatrix},\qquad \mathbf{b}=\begin{bmatrix} 2\\1\\2\\5 \end{bmatrix}.$$- Try $\mathbf{w}=[1,0,1]$, i.e. the curve $y = 1 + x^2$.
- $A\mathbf{w} = [1+0+1,\; 1+0+0,\; 1+0+1,\; 1+0+4] = [2,1,2,5] = \mathbf{b}$. A perfect fit.
- So $\mathbf{b}$ lies in the column space, the residual is zero, and since the three columns are independent this is the unique least-squares answer.
With noisy data the fit would not be perfect, and you would solve the same normal equations (or better, QR/SVD).
A feature map turns one input into several features. The polynomial feature map of degree $d$ is $\phi(x) = [1, x, x^2, \dots, x^d]$. Stacking it for all examples gives the Vandermonde design matrix
$$A = \begin{bmatrix} 1 & x_1 & x_1^2 & \cdots & x_1^d\\ \vdots & & & & \vdots\\ 1 & x_n & x_n^2 & \cdots & x_n^d \end{bmatrix},\qquad \hat{\mathbf{w}} = \arg\min\|A\mathbf{w}-\mathbf{y}\|^2.$$- Full column rank needs $d+1\le n$ (at least as many distinct $x$ values as weights). With $d+1=n$ the curve passes through every point (interpolation), and the training error is zero.
- Training error never goes up as $d$ grows, but the error on new data first falls (underfitting fixed) and then rises (overfitting).
- The columns $x^k$ and $x^{k+1}$ look alike on $[0,1]$, so $A$ becomes nearly dependent and $\kappa(A)$ explodes as $d$ grows. This is where QR/SVD and ridge earn their keep.
The same trick works for any features: $\sin(x)$, products of two features, one-hot columns. The model is still linear in the weights, so the whole chapter applies unchanged.
Why do we need it?
Straight lines cannot follow curved data. Adding new columns (x², x³, …) bends the fit while keeping the same easy solving method.
Where is it used?
Trend and curve fitting, calibration curves in physics and engineering, spline and Fourier-feature regression, and kernel methods, which are "clever columns" at heart.
How is it used?
Build the Vandermonde matrix with np.vander, fit with lstsq, and choose the degree by test or cross-validation error, never by training error. Scale x to [−1, 1] and consider ridge for high degrees.
- A lower training error does not mean a better model. Always judge by error on data the model has not seen.
- High-degree monomials on a wide range of $x$ are numerically awful. Centre and scale $x$ to $[-1,1]$, use orthogonal polynomials (Legendre, Chebyshev), or add ridge. And never use the normal equations here.
- Polynomials oscillate near the edges of the data and shoot off to $\pm\infty$ outside it. They are poor for extrapolation.
Quick check: you have 6 data points with distinct $x$ values. What is the smallest degree whose polynomial can pass through all of them exactly?
You need $d+1=6$ weights, so $d=5$. A degree-5 polynomial has 6 free coefficients, exactly enough to pass through 6 points. (It will usually wiggle wildly between them.)
Recap, cheat sheet and practice
- The problem. More equations than unknowns: $A\mathbf{x}=\mathbf{b}$ has no exact solution. Minimise the squared residual $\|\mathbf{r}\|^2 = \|\mathbf{b}-A\mathbf{x}\|^2$.
- The picture. $A\hat{\mathbf{x}}$ is the projection of $\mathbf{b}$ onto the column space $C(A)$. The residual is perpendicular to every column (it lives in $N(A^\top)$).
- Normal equations $A^\top A\hat{\mathbf{x}}=A^\top\mathbf{b}$, found two ways: "residual $\perp$ columns" (geometry) and "gradient of the error $=\mathbf{0}$" (calculus). A unique solution needs independent columns (full column rank).
- Solvers. Normal equations are quick but square the condition number. QR ($R\hat{\mathbf{x}}=Q^\top\mathbf{b}$) is the stable default. SVD is the most robust and handles rank deficiency.
- Pseudoinverse $A^+=V\Sigma^+U^\top$: $A^+\mathbf{b}$ is the shortest of the best-fit solutions. It equals $(A^\top A)^{-1}A^\top$ for tall full-rank $A$ and $A^\top(AA^\top)^{-1}$ for wide full-rank $A$.
- Ridge $(A^\top A+\lambda I)\mathbf{x}=A^\top\mathbf{b}$ is always solvable, shrinks weights and improves conditioning. Weighted least squares adds trust weights $W$.
- Regression. Build the design matrix (a column of ones for the bias, plus features). Polynomial features turn the same machinery into curve fitting, with a risk of overfitting and bad conditioning.
Cheat sheet
| Idea | Formula | Remember |
|---|---|---|
| Residual | $\mathbf{r}=\mathbf{b}-A\mathbf{x}$ | the miss at each data point |
| Objective | $\min\|A\mathbf{x}-\mathbf{b}\|^2$ | sum of squared misses |
| Geometry | $A\hat{\mathbf{x}}=\operatorname{proj}_{C(A)}\mathbf{b}$, $\;A^\top\mathbf{r}=\mathbf{0}$ | foot of the perpendicular |
| Normal equations | $A^\top A\hat{\mathbf{x}}=A^\top\mathbf{b}$ | unique iff independent columns |
| Condition number | $\kappa(A^\top A)=\kappa(A)^2$ | squared: avoid for ill-conditioned $A$ |
| QR solve | $R\hat{\mathbf{x}}=Q^\top\mathbf{b}$ | back-substitution, stable |
| SVD solve | $\hat{\mathbf{x}}=V\Sigma^{+}U^\top\mathbf{b}$ | most robust |
| Pseudoinverse | $(A^\top A)^{-1}A^\top$ tall; $A^\top(AA^\top)^{-1}$ wide | shortest best-fit solution |
| Ridge | $(A^\top A+\lambda I)^{-1}A^\top\mathbf{b}$ | leash on the weights |
| Weighted LS | $(A^\top WA)\hat{\mathbf{x}}=A^\top W\mathbf{b}$ | trusted points count more |
| Design matrix | $[\mathbf{1}\;\;x_1\;\;x_2\cdots]$ | column of ones = bias |
import numpy as np
# Data: the running example (three points) and its design matrix with a bias column
x = np.array([1., 2., 3.])
y = np.array([1., 2., 2.])
A = np.column_stack([np.ones_like(x), x]) # shape (3, 2): [1, x]
# 1) Normal equations (fast, but squares the condition number)
w_ne = np.linalg.solve(A.T @ A, A.T @ y)
# 2) QR: A = QR, then solve R w = Q^T y
Q, R = np.linalg.qr(A) # reduced QR
w_qr = np.linalg.solve(R, Q.T @ y) # R is triangular (scipy.linalg.solve_triangular is faster)
# 3) SVD: w = V diag(1/s) U^T y
U, s, Vt = np.linalg.svd(A, full_matrices=False)
w_svd = Vt.T @ ((U.T @ y) / s)
# 4) Pseudoinverse and the library routine (SVD inside)
w_pinv = np.linalg.pinv(A) @ y
w_lstsq, res, rank, sv = np.linalg.lstsq(A, y, rcond=None)
print(w_ne, w_qr, w_svd, w_pinv, w_lstsq) # all about [0.6667 0.5]: the four answers agree
# Residual is perpendicular to the columns: A^T r = 0
r = y - A @ w_lstsq
print(A.T @ r) # ~ [0, 0]
print(np.linalg.cond(A), np.linalg.cond(A.T @ A)) # 6.79 and 46.1: the square of the first
# Minimum-norm solution of an underdetermined system: 3x + 4y = 10
print(np.linalg.pinv(np.array([[3., 4.]])) @ np.array([10.])) # [1.2 1.6]
# Ridge regression: (A^T A + lam I) w = A^T y
def ridge(A, y, lam):
n = A.shape[1]
return np.linalg.solve(A.T @ A + lam * np.eye(n), A.T @ y)
# Coefficients versus lambda for two nearly identical features
rng = np.random.default_rng(0)
f1 = rng.normal(size=30)
F = np.column_stack([f1, f1 + 0.05 * rng.normal(size=30)])
t = F @ np.array([1., 1.]) + 0.5 * rng.normal(size=30)
for lam in [0, 1e-3, 1e-1, 1, 10, 100]:
wr = np.linalg.lstsq(F, t, rcond=None)[0] if lam == 0 else ridge(F, t, lam)
print(lam, wr) # weights shrink and stabilise as lam grows
# To plot: collect wr for many lam values and use plt.semilogx(lams, coefficients)
# Polynomial fits of increasing degree: watch the condition number of A^T A grow
xs = np.linspace(0, 1, 12)
ys = np.sin(2 * np.pi * xs) + 0.3 * rng.normal(size=12)
for d in [1, 3, 5, 7, 9, 11]:
V = np.vander(xs, d + 1, increasing=True) # columns 1, x, x^2, ...
w_d = np.linalg.lstsq(V, ys, rcond=None)[0] # use lstsq, never inv(V.T @ V)
print(d, "cond(A) = %.1e cond(A^T A) = %.1e" % (np.linalg.cond(V), np.linalg.cond(V.T @ V)),
"train MSE = %.4f" % np.mean((V @ w_d - ys) ** 2))
1. At the least-squares solution $\hat{\mathbf{x}}$, the residual $\mathbf{b}-A\hat{\mathbf{x}}$ is…
2. $A$ is $50\times 3$ with independent columns. What is the size of $A^\top A$, and is it invertible?
3. Why are the normal equations risky for badly conditioned $A$?
4. Which vector is the minimum-norm solution of $x + y = 2$?
5. What does adding $\lambda I$ (with $\lambda>0$) to $A^\top A$ do?
6. You fit a degree-11 polynomial to 12 noisy points. What do you expect?
Practice problems
A. Find the least-squares line $y=c+mx$ through $(0,1)$, $(1,3)$, $(2,4)$, and check that the residual is perpendicular to both columns.
$A=\begin{bmatrix}1&0\\1&1\\1&2\end{bmatrix}$, $\mathbf{b}=[1,3,4]$. Then $A^\top A=\begin{bmatrix}3&3\\3&5\end{bmatrix}$ and $A^\top\mathbf{b}=[8,11]$. The equations $3c+3m=8$, $3c+5m=11$ give (subtract) $2m=3$, so $m=1.5$, and $3c = 8-4.5=3.5$, so $c=\tfrac76$. Predictions: $[\tfrac76,\tfrac83,\tfrac{25}6]$. Residual $=[-\tfrac16,\tfrac13,-\tfrac16]$. Check: $\mathbf{1}\cdot\mathbf{r} = -\tfrac16+\tfrac13-\tfrac16=0$ and $[0,1,2]\cdot\mathbf{r}=\tfrac13-\tfrac26=0$ ✓.
B. Find the minimum-norm solution of $x+2y+2z=9$ and its length.
$A=[1\;2\;2]$, $AA^\top=1+4+4=9$. So $\mathbf{x}^+=A^\top\cdot\tfrac99=[1,2,2]$, with length $\sqrt{1+4+4}=3$.
C. One-weight ridge: $A=[3,4]^\top$, $\mathbf{b}=[5,10]$. Compute the ridge weight for $\lambda=0$ and $\lambda=25$.
$A^\top A=9+16=25$ and $A^\top\mathbf{b}=15+40=55$. For $\lambda=0$: $x=55/25=2.2$. For $\lambda=25$: $x=55/(25+25)=1.1$. The weight is cut in half.
D. Weighted average: readings $3$, $5$, $10$ with weights $2$, $2$, $1$. Find the weighted least-squares estimate.
$A^\top WA=2+2+1=5$ and $A^\top W\mathbf{b}=6+10+10=26$. So $x=26/5=5.2$ (the plain average is 6; the outlier 10 has less influence).
E. If $\kappa(A)=10^9$, estimate $\kappa(A^\top A)$ and say what happens to the normal equations in double precision (about 16 digits).
$\kappa(A^\top A)=10^{18}$, bigger than $10^{16}$. All digits of the answer can be lost (the computer may even see $A^\top A$ as singular). QR or SVD with $\kappa=10^9$ keeps about $16-9=7$ good digits.
F. A regression has 100 rows and 3 features, but features 2 and 3 are identical columns. What does np.linalg.lstsq return and why?
The rank is 2, not 3, so $A^\top A$ is singular and there are infinitely many equally good weight vectors: any split of the combined weight between the two copies. lstsq uses the SVD and returns the pseudoinverse solution, the minimum-norm one, which splits the shared weight equally between the two identical columns.
Eigenvalues & Eigenvectors
A matrix turns almost every arrow it touches. But a few special arrows only get stretched or shrunk, and stay on their own line. Finding those arrows is one of the most useful tricks in all of linear algebra: it explains PCA, PageRank, and why some neural networks blow up.
- See an eigenvector as a direction a matrix does not turn, and an eigenvalue as the stretch along it
- Find eigenvalues by hand with the characteristic equation, including complex ones
- Understand eigenspaces, multiplicity and "defective" matrices
- Use the key facts: trace = sum of eigenvalues, determinant = product, and what happens for $A^k$, $A^{-1}$, $A+cI$
- Diagonalise a matrix, take powers of it, and predict its long-term behaviour
- Know why symmetric matrices are special (the spectral theorem), and how computers find eigenvalues
- Connect it all to PCA, PageRank and exploding gradients
Eigenvectors and eigenvalues: directions that do not turn core
Remember from Chapter 1.5 that a matrix is a machine that moves every arrow in the plane to a new place. Most arrows get turned: after the machine, they point in a new direction.
But some special arrows are different. The machine pushes them along the same line they were already on. They get longer, shorter, or flipped, but they do not turn. Those special arrows are the eigenvectors of the matrix. How much each one is stretched is its eigenvalue.
Picture a rubber sheet that you pull with both hands. Lines along the pull direction stay on the same line, only longer. Or picture a spinning globe: the axis does not move at all. Both are "directions that survive".
("Eigen" is a German word meaning "own" or "characteristic". An eigenvector is a matrix's very own direction.)
Let $A = \begin{bmatrix} 2 & 1 \\ 1 & 2 \end{bmatrix}$. Try two arrows.
- Take $\mathbf{v} = [1, 1]$. Then $A\mathbf{v} = [2\cdot1 + 1\cdot1,\; 1\cdot1 + 2\cdot1] = [3, 3] = 3\,[1,1]$. The arrow stayed on its line and became 3 times longer. So $\mathbf{v}$ is an eigenvector with eigenvalue $3$.
- Take $\mathbf{w} = [1, -1]$. Then $A\mathbf{w} = [2 - 1,\; 1 - 2] = [1, -1] = 1\cdot\mathbf{w}$. Same line, same length. Eigenvalue $1$.
- Take $\mathbf{u} = [1, 0]$. Then $A\mathbf{u} = [2, 1]$. This is not a multiple of $[1,0]$ (it turned upward), so $\mathbf{u}$ is not an eigenvector.
Eigenvectors are only defined up to scale. Try $5\mathbf{v} = [5,5]$: $A[5,5] = [15,15] = 3\,[5,5]$. It still works, with the same eigenvalue 3. Every multiple of $[1,1]$ is an eigenvector, so we talk about "the direction $[1,1]$".
Let $A$ be a square matrix ($n\times n$). A non-zero vector $\mathbf{v}$ is an eigenvector of $A$, with eigenvalue $\lambda$ (the Greek letter "lambda", just a number), if
$$A\mathbf{v} = \lambda\,\mathbf{v}.$$The pair $(\lambda, \mathbf{v})$ is called an eigenpair. In words: "multiplying by $A$ does the same thing to $\mathbf{v}$ as multiplying by the plain number $\lambda$." A whole matrix acts on $\mathbf{v}$ like one number.
- Why square? $A\mathbf{v}$ must live in the same space as $\mathbf{v}$ so that the two sides can be equal.
- Up to scale. If $A\mathbf{v}=\lambda\mathbf{v}$ then $A(c\mathbf{v}) = cA\mathbf{v} = \lambda(c\mathbf{v})$ for any $c\neq0$. So $c\mathbf{v}$ is an eigenvector with the same $\lambda$.
- $\mathbf{v}=\mathbf{0}$ is not allowed. $A\mathbf{0}=\lambda\mathbf{0}$ is true for every $\lambda$, so it tells us nothing. (But $\lambda = 0$ is allowed.)
What the eigenvalue $\lambda$ tells you about the stretch along $\mathbf{v}$:
| Eigenvalue | What happens to the arrow |
|---|---|
| $\lambda \gt 1$ | stretched (longer, same direction) |
| $0 \lt \lambda \lt 1$ | shrunk (shorter, same direction) |
| $\lambda = 1$ | unchanged |
| $\lambda = 0$ | squashed to the zero vector (so $\mathbf{v}$ is in the null space) |
| $\lambda \lt 0$ | flipped to point the opposite way (and stretched or shrunk) |
Why do we need it?
Matrices can look messy because they push and turn every arrow. We need a way to find the simple part: the special directions where the matrix only stretches. Those directions make the whole matrix easy to understand.
Where is it used?
PCA (finding the main directions of your data), Google's PageRank, vibration analysis in engineering, stability checks for recurrent neural networks, and the curvature of loss surfaces (Hessian eigenvalues).
How is it used?
Ask: which non-zero $\mathbf{v}$ gives $A\mathbf{v} = \lambda\mathbf{v}$? In code call np.linalg.eig(A) (or eigh for symmetric matrices). The eigenvalues say how much each direction is stretched; the eigenvectors (columns) say which directions they are.
An eigenvector is a direction, not one special arrow. $[1,1]$, $[5,5]$ and $[-2,-2]$ are all "the same" eigenvector direction. Software usually returns one of them with length 1, and its sign can differ from your hand answer. Both are right.
Negative eigenvalues flip the arrow. It still lies on the same line (just pointing backwards), so it still counts as "not turned".
Not every matrix has real eigen-directions. A rotation by 90° turns every arrow, so it has none (we will see why, and what replaces them, in the complex-eigenvalues section below).
Quick check: $A=\begin{bmatrix}3&0\\0&5\end{bmatrix}$. Is $[0,1]$ an eigenvector? With which eigenvalue?
$A[0,1] = [3\cdot0 + 0\cdot1,\; 0\cdot0 + 5\cdot1] = [0,5] = 5\,[0,1]$. Yes, with eigenvalue $5$. (Likewise $[1,0]$ has eigenvalue $3$.)
Computing eigenvalues by hand: the characteristic equation core
We want a non-zero $\mathbf{v}$ with $A\mathbf{v} = \lambda\mathbf{v}$. Move everything to one side and it becomes a question about the null space:
$A\mathbf{v} - \lambda\mathbf{v} = \mathbf{0}$, which is the same as $(A - \lambda I)\mathbf{v} = \mathbf{0}$.
So we need a matrix, $A - \lambda I$, that sends some non-zero arrow to zero. That only happens when the matrix squashes space flat, which means it is singular, which means its determinant is 0 (see Chapter 1.7).
That gives a plan: first find the numbers $\lambda$ that make the determinant zero. Then, for each one, solve for the arrows $\mathbf{v}$.
Find the eigenvalues and eigenvectors of $A = \begin{bmatrix} 4 & 1 \\ 2 & 3 \end{bmatrix}$.
- Subtract $\lambda$ from the diagonal: $A - \lambda I = \begin{bmatrix} 4-\lambda & 1 \\ 2 & 3-\lambda \end{bmatrix}$.
- Take the determinant: $(4-\lambda)(3-\lambda) - 1\cdot2 = 12 - 7\lambda + \lambda^2 - 2 = \lambda^2 - 7\lambda + 10$.
- Set it to zero: $\lambda^2 - 7\lambda + 10 = 0$, which factors as $(\lambda - 5)(\lambda - 2) = 0$. So $\lambda_1 = 5$ and $\lambda_2 = 2$.
- For $\lambda = 5$: $A - 5I = \begin{bmatrix} -1 & 1 \\ 2 & -2 \end{bmatrix}$. The equations $-x + y = 0$ and $2x - 2y = 0$ say $y = x$. So $\mathbf{v}_1 = [1, 1]$.
- For $\lambda = 2$: $A - 2I = \begin{bmatrix} 2 & 1 \\ 2 & 1 \end{bmatrix}$. Both rows say $2x + y = 0$, so $y = -2x$. So $\mathbf{v}_2 = [1, -2]$.
- Check one: $A[1,-2] = [4 - 2,\; 2 - 6] = [2, -4] = 2\,[1,-2]$ ✓.
Shortcut for triangular and diagonal matrices. For $\begin{bmatrix} 3 & 1 \\ 0 & 2 \end{bmatrix}$ the determinant is $(3-\lambda)(2-\lambda) - 0$, whose roots are just $3$ and $2$. The eigenvalues of a triangular (or diagonal) matrix are its diagonal entries. For the $3\times3$ matrix with rows $[2,7,1]$, $[0,3,5]$, $[0,0,-4]$, no work is needed: the eigenvalues are $2$, $3$ and $-4$.
For an $n\times n$ matrix $A$:
$$\det(A - \lambda I) = 0 \qquad\text{(the characteristic equation)}$$The left side, $p(\lambda) = \det(A - \lambda I)$, is a polynomial of degree $n$ in $\lambda$ called the characteristic polynomial. Its roots are exactly the eigenvalues. An $n\times n$ matrix has $n$ eigenvalues, counting repeats and allowing complex numbers.
For a $2\times2$ matrix there is a handy formula. With $\operatorname{tr}(A) = a + d$ (the sum of the diagonal) and $\det A = ad - bc$:
$$\det(A - \lambda I) = \lambda^2 - \operatorname{tr}(A)\,\lambda + \det(A), \qquad \lambda = \frac{\operatorname{tr}(A) \pm \sqrt{\operatorname{tr}(A)^2 - 4\det(A)}}{2}.$$Recipe: (1) write $\det(A-\lambda I)$, (2) solve it for $\lambda$, (3) for each $\lambda$ solve $(A-\lambda I)\mathbf{v}=\mathbf{0}$ (row reduce, as in Chapter 1.6).
Why do we need it?
To find eigenvalues exactly we need an equation to solve. Asking for $A - \lambda I$ to squash space flat turns the search into a polynomial, and solving a polynomial is something we know how to do.
Where is it used?
Textbook and exam problems, small $2\times2$ and $3\times3$ models (population models, small Markov chains), and quick hand checks with trace and determinant. Big matrices use numerical methods instead.
How is it used?
Write $\det(A - \lambda I)$, expand it, set it to zero and solve for $\lambda$. For $2\times2$ use $\lambda^2 - \text{trace}\cdot\lambda + \det = 0$. Then, for each $\lambda$, solve $(A - \lambda I)\mathbf{v} = \mathbf{0}$. For triangular matrices just read the diagonal.
Sign slip. The polynomial is $\det(A-\lambda I)$, so $\lambda$ is subtracted from each diagonal entry. Writing $A+\lambda I$ by mistake flips the signs of your eigenvalues.
A matrix that is not singular can still have eigenvalues. It is $A-\lambda I$ that must be singular, not $A$. (And $A$ itself is singular exactly when $\lambda=0$ is an eigenvalue.)
Big matrices. The characteristic polynomial is a teaching tool. Nobody computes eigenvalues of a $1000\times1000$ matrix this way (see the numerical section below).
Quick check: find the eigenvalues of $\begin{bmatrix} 1 & 2 \\ 2 & 1 \end{bmatrix}$.
Trace $=2$, determinant $=1-4=-3$. So $\lambda^2 - 2\lambda - 3 = 0$, i.e. $(\lambda-3)(\lambda+1)=0$. The eigenvalues are $3$ and $-1$.
Complex eigenvalues: what a rotation is hiding
Spin the whole plane by 90°. Which arrow stays on its own line? None. Every arrow turns. So a rotation has no real eigenvectors.
But the algebra still gives an answer, if we allow a new kind of number. The imaginary unit $i$ is defined by $i\cdot i = -1$. A complex number looks like $a + bi$ (a real part $a$ and an imaginary part $b$). You can think of it as a point in a plane, which is exactly why it fits rotations.
The message: complex eigenvalues mean "this matrix turns things". The size of the eigenvalue says how much it stretches each turn, and its angle says how far it turns.
Rotate by 90°: $R = \begin{bmatrix} 0 & -1 \\ 1 & 0 \end{bmatrix}$.
- $\det(R - \lambda I) = (0-\lambda)(0-\lambda) - (-1)(1) = \lambda^2 + 1$.
- Set it to zero: $\lambda^2 = -1$, so $\lambda = i$ or $\lambda = -i$.
- Eigenvector for $\lambda = i$: $(R - iI)\mathbf{v} = \mathbf{0}$ gives $-ix - y = 0$, so $y = -ix$. Take $\mathbf{v} = [1, -i]$.
- Check: $R[1,-i] = [0\cdot1 + (-1)(-i),\; 1\cdot1 + 0\cdot(-i)] = [i, 1]$. And $i\cdot[1,-i] = [i, -i^2] = [i, 1]$ ✓.
For a rotation by a general angle $\theta$ the eigenvalues are $\cos\theta \pm i\sin\theta$. Each has size $1$ (a pure turn, no stretching). If we also scale by $r$, the matrix $r\begin{bmatrix}\cos\theta & -\sin\theta\\ \sin\theta & \cos\theta\end{bmatrix}$ has eigenvalues $r(\cos\theta \pm i\sin\theta)$, of size $r$.
- If a real matrix has a complex eigenvalue $\lambda = a + bi$, then its conjugate $\bar\lambda = a - bi$ is an eigenvalue too. They come in pairs. (A real quadratic with negative discriminant has two conjugate roots.)
- The size (modulus) is $|\lambda| = \sqrt{a^2+b^2}$. The angle is $\theta = \operatorname{atan2}(b, a)$. Then $\lambda = |\lambda|(\cos\theta + i\sin\theta)$.
- On the plane, such a matrix behaves like "turn by $\theta$, then scale by $|\lambda|$". Repeating it gives a spiral: inward if $|\lambda|\lt1$, a circle if $|\lambda|=1$, outward if $|\lambda|\gt1$.
- An $n\times n$ real matrix always has $n$ eigenvalues counted with repeats, in the complex numbers. Odd-sized real matrices always have at least one real eigenvalue.
Why do we need it?
Rotations turn every arrow, so they have no real eigenvectors, yet we still want to describe them. Complex eigenvalues capture exactly "turn by some angle and scale by some amount".
Where is it used?
Oscillating systems (springs, circuits, pendulums), recurrent neural networks whose hidden state spirals, signal processing, and time-series models that contain cycles.
How is it used?
Compute the eigenvalues. If they come as a pair $a \pm bi$, read the size $\sqrt{a^2+b^2}$ as the stretch per step and the angle $\operatorname{atan2}(b,a)$ as the turn per step. Size below 1 means a spiral inward, above 1 a spiral outward.
Quick check: what are the eigenvalues of $\begin{bmatrix}0&-2\\2&0\end{bmatrix}$?
Trace $=0$, determinant $=0\cdot0-(-2)(2)=4$. So $\lambda^2+4=0$, giving $\lambda=\pm2i$. This is a 90° rotation that also doubles the length each time (size $|\lambda|=2$).
Eigenspaces and multiplicity (and defective matrices)
If $\mathbf{v}$ is an eigenvector for $\lambda$, so is every multiple of it. And if two different eigenvectors share the same $\lambda$, so is any mix of them. So all the eigenvectors of one eigenvalue, together with the zero vector, form a subspace: a line, a plane, or more. We call it the eigenspace of $\lambda$.
A natural question: how many independent arrows does an eigenvalue give you? You hope the answer matches how many times the eigenvalue appears in the characteristic polynomial. Usually it does. Sometimes there are fewer, and the matrix is "short of eigenvectors".
Compare two matrices. Both have characteristic polynomial $(\lambda - 2)^2$, so $\lambda=2$ appears twice.
- $A = \begin{bmatrix} 2 & 0 \\ 0 & 2 \end{bmatrix}$. $A - 2I$ is the zero matrix, so $(A-2I)\mathbf{v}=\mathbf{0}$ holds for every $\mathbf{v}$. The eigenspace is the whole plane: 2 independent eigenvectors.
- $B = \begin{bmatrix} 2 & 1 \\ 0 & 2 \end{bmatrix}$. $B - 2I = \begin{bmatrix} 0 & 1 \\ 0 & 0 \end{bmatrix}$. The only equation is $y = 0$, with $x$ free. The eigenspace is just the x-axis: only 1 independent eigenvector, $[1,0]$.
The matrix $B$ is a shear: it slides everything sideways, and only the x-axis stays on its line. It is short of eigenvectors.
The eigenspace of an eigenvalue $\lambda$ is the null space of $A - \lambda I$:
$$E_\lambda = N(A - \lambda I) = \{\mathbf{v} : A\mathbf{v} = \lambda\mathbf{v}\}.$$- Algebraic multiplicity of $\lambda$: how many times $\lambda$ is a root of the characteristic polynomial.
- Geometric multiplicity of $\lambda$: the dimension of $E_\lambda$, i.e. the number of independent eigenvectors for $\lambda$ (use the null space tools of Chapter 1.8).
- Always: $1 \le \text{geometric} \le \text{algebraic}$.
- If geometric $\lt$ algebraic for some eigenvalue, the matrix is defective: it does not have enough eigenvectors to span the whole space. (Awareness topic: it matters mainly because defective matrices cannot be diagonalised.)
Why do we need it?
One eigenvalue can have several independent eigenvectors, or fewer than we hoped. We need to count them to know whether we have enough directions to describe the whole space.
Where is it used?
Deciding whether a matrix can be diagonalised, solving systems of differential equations, and analysing Markov chains with a repeated eigenvalue 1. Symmetric matrices (the main ones in ML) are never short of eigenvectors.
How is it used?
For each eigenvalue $\lambda$, row-reduce $A - \lambda I$ and find its null space. Its dimension is the geometric multiplicity. Compare with how often $\lambda$ repeats in the characteristic polynomial (the algebraic multiplicity). If it is smaller, the matrix is defective.
Repeated eigenvalue does not mean defective. The identity matrix has the eigenvalue 1 repeated $n$ times and is not defective at all. You must count the eigenvectors, not just the roots.
Defective matrices are rare in practice. A random matrix is almost never defective, and symmetric matrices (the ones that matter most in ML) are never defective. Still, they appear in theory, so it is good to know the word.
Quick check: $A=\begin{bmatrix}5&0&0\\0&5&0\\0&0&7\end{bmatrix}$. What are the algebraic and geometric multiplicities of $5$?
The characteristic polynomial is $(5-\lambda)^2(7-\lambda)$, so the algebraic multiplicity of $5$ is $2$. $A - 5I$ has rank $1$ (only the last row is non-zero), so its null space has dimension $3-1=2$: the geometric multiplicity is $2$. They match, so no problem.
Key properties of eigenvalues core
Once you know the eigenvalues, you can read off a lot about the matrix without any more work.
- Determinant = product of eigenvalues. The determinant is the factor by which the matrix scales areas. If it stretches by 5 along one direction and by 2 along another, areas grow by $5\times2=10$.
- Trace = sum of eigenvalues. The trace (sum of the diagonal) is the "total stretch".
- Powers. Applying $A$ twice stretches an eigenvector by $\lambda$ twice: $\lambda^2$.
- Inverse. The inverse undoes the stretch, so it stretches by $1/\lambda$.
- Shift. Adding $c$ to every diagonal entry just adds $c$ to every stretch.
Use $A = \begin{bmatrix} 4 & 1 \\ 2 & 3 \end{bmatrix}$, which has eigenvalues $5$ and $2$ with eigenvectors $[1,1]$ and $[1,-2]$ (from the earlier example).
- Trace and determinant. $\operatorname{tr}A = 4 + 3 = 7 = 5 + 2$ ✓. $\det A = 12 - 2 = 10 = 5\cdot2$ ✓.
- Square. $A^2 = \begin{bmatrix} 18 & 7 \\ 14 & 11 \end{bmatrix}$. Its trace is $29 = 5^2 + 2^2 = 25 + 4$ ✓ and its determinant is $198 - 98 = 100 = 25\cdot4$ ✓. So the eigenvalues of $A^2$ are $25$ and $4$.
- Inverse. $A^{-1} = \tfrac{1}{10}\begin{bmatrix} 3 & -1 \\ -2 & 4 \end{bmatrix}$. Its trace is $0.7 = 0.2 + 0.5 = \tfrac15 + \tfrac12$ ✓. Eigenvalues $\tfrac15$ and $\tfrac12$.
- Shift. $A + 2I = \begin{bmatrix} 6 & 1 \\ 2 & 5 \end{bmatrix}$ has trace $11 = 7 + 4$ and determinant $30 - 2 = 28 = 7\cdot4$. Eigenvalues $7$ and $4$, which are $5+2$ and $2+2$ ✓.
- Similar matrix. With $S = \begin{bmatrix}1&1\\0&1\end{bmatrix}$, the matrix $B = S^{-1}AS = \begin{bmatrix} 2 & 0 \\ 2 & 5 \end{bmatrix}$ is triangular, so its eigenvalues are $2$ and $5$: the same as $A$ ✓.
Let $A$ be an $n\times n$ matrix with eigenvalues $\lambda_1,\dots,\lambda_n$ (counted with repeats, possibly complex).
$$\operatorname{tr}(A) = \sum_{i=1}^n \lambda_i, \qquad \det(A) = \prod_{i=1}^n \lambda_i.$$If $A\mathbf{v}=\lambda\mathbf{v}$ then, with the same eigenvector $\mathbf{v}$:
- $A^k\mathbf{v} = \lambda^k\mathbf{v}$ (because $A^2\mathbf{v} = A(\lambda\mathbf{v}) = \lambda A\mathbf{v} = \lambda^2\mathbf{v}$, and so on)
- $A^{-1}\mathbf{v} = \lambda^{-1}\mathbf{v}$ (if $A$ is invertible; multiply $A\mathbf{v}=\lambda\mathbf{v}$ by $A^{-1}$ and divide by $\lambda$)
- $(A + cI)\mathbf{v} = (\lambda + c)\mathbf{v}$
Similar matrices ($B = S^{-1}AS$ for some invertible $S$) have the same eigenvalues. Reason: if $A\mathbf{v}=\lambda\mathbf{v}$, then $B(S^{-1}\mathbf{v}) = S^{-1}AS\,S^{-1}\mathbf{v} = S^{-1}A\mathbf{v} = \lambda(S^{-1}\mathbf{v})$. Similar matrices are the same transformation seen in a different coordinate system (Chapter 1.3, change of basis), and eigenvalues do not depend on the coordinates.
One more consequence: $A$ is invertible exactly when no eigenvalue is 0 (because $\det A = \prod\lambda_i$).
Why do we need it?
Often you do not need every eigenvalue, only a quick fact: is the matrix invertible, are all eigenvalues positive, by how much does it scale area? These rules answer that without solving anything.
Where is it used?
Checking invertibility (no zero eigenvalue), stability of iterations (all $|\lambda| \lt 1$), the ridge trick $A + cI$, and the QR algorithm (which relies on similar matrices sharing eigenvalues).
How is it used?
Add up the diagonal to get $\sum\lambda_i$ and take the determinant to get $\prod\lambda_i$. To get eigenvalues of $A^2$, $A^{-1}$ or $A + cI$, transform the eigenvalues of $A$ ($\lambda^2$, $1/\lambda$, $\lambda + c$); the eigenvectors stay the same.
The rules for $A+B$ and $AB$ do not work. The eigenvalues of $A+B$ are not the sums of eigenvalues of $A$ and $B$, and the eigenvalues of $AB$ are not the products. (They fail because $A$ and $B$ usually have different eigenvectors.) The shift $A+cI$ works only because $I$ shares every eigenvector with $A$.
Similar is not equal. $A$ and $S^{-1}AS$ share eigenvalues but have different entries and different eigenvectors.
Quick check: a $3\times3$ matrix has eigenvalues $1$, $2$ and $4$. What are its trace, its determinant, and the eigenvalues of $A^2$?
Trace $=1+2+4=7$. Determinant $=1\cdot2\cdot4=8$. Eigenvalues of $A^2$: $1, 4, 16$.
Diagonalization: $A = PDP^{-1}$ core
Multiplying by a matrix can look messy: the arrow gets pushed and turned all over the place. But the eigenvectors give us a new set of axes on which the matrix is easy: along each axis it only stretches.
So here is a three-step plan, like translating a hard job into a language where it is easy:
- Translate in. Write the arrow using the eigenvectors as axes (find "how many scoops of each eigenvector").
- Stretch. Multiply each scoop-count by its eigenvalue. This is easy: it is just a diagonal matrix.
- Translate back to ordinary coordinates.
That is what $A = PDP^{-1}$ says, read from right to left. The big payoff: powers become easy, because stretching twice is just stretching by $\lambda^2$.
Take $A = \begin{bmatrix} 4 & 1 \\ 2 & 3 \end{bmatrix}$ again, with eigenpairs $(5, [1,1])$ and $(2, [1,-2])$. Put the eigenvectors in the columns of $P$ and the eigenvalues (same order) on the diagonal of $D$:
$$P = \begin{bmatrix} 1 & 1 \\ 1 & -2 \end{bmatrix},\quad D = \begin{bmatrix} 5 & 0 \\ 0 & 2 \end{bmatrix},\quad P^{-1} = \frac13\begin{bmatrix} 2 & 1 \\ 1 & -1 \end{bmatrix}.$$Apply $A$ to $\mathbf{x} = [3, 0]$ in three steps:
- Translate in: $P^{-1}\mathbf{x} = \tfrac13[2\cdot3 + 0,\; 3 - 0] = [2, 1]$. Check: $2\,[1,1] + 1\,[1,-2] = [3, 0]$ ✓. So $\mathbf{x}$ is "2 scoops of $\mathbf{v}_1$ and 1 scoop of $\mathbf{v}_2$".
- Stretch: $D[2,1] = [5\cdot2,\; 2\cdot1] = [10, 2]$.
- Translate back: $P[10,2] = 10\,[1,1] + 2\,[1,-2] = [12, 6]$.
Directly: $A\mathbf{x} = [4\cdot3, 2\cdot3] = [12, 6]$ ✓. Same answer.
Powers. $A^3 = PD^3P^{-1}$ with $D^3 = \operatorname{diag}(125, 8)$:
$$A^3 = \begin{bmatrix}1&1\\1&-2\end{bmatrix}\begin{bmatrix}125&0\\0&8\end{bmatrix}\tfrac13\begin{bmatrix}2&1\\1&-1\end{bmatrix} = \tfrac13\begin{bmatrix}258 & 117\\ 234 & 141\end{bmatrix} = \begin{bmatrix}86 & 39\\ 78 & 47\end{bmatrix}.$$And repeated multiplication gives the same matrix: $A^2A = \begin{bmatrix}18&7\\14&11\end{bmatrix}\begin{bmatrix}4&1\\2&3\end{bmatrix} = \begin{bmatrix}86&39\\78&47\end{bmatrix}$ ✓. To get $A^{100}$ we only need $5^{100}$ and $2^{100}$, not 99 matrix multiplications.
An $n\times n$ matrix $A$ is diagonalizable if it can be written
$$A = P D P^{-1},$$where $D$ is a diagonal matrix of eigenvalues and the columns of $P$ are the matching eigenvectors (same order). Equivalent form: $AP = PD$, which just says "$A\mathbf{v}_j = \lambda_j\mathbf{v}_j$ for every column".
When is it possible? Exactly when $A$ has $n$ linearly independent eigenvectors (so that $P$ is invertible). In practice:
- $n$ distinct eigenvalues ⇒ always diagonalizable.
- Symmetric matrices ⇒ always diagonalizable (next sections).
- Defective matrices (like the shear) ⇒ not diagonalizable, they lack eigenvectors.
- A rotation is not diagonalizable with real numbers (its eigenvectors are complex), though it is with complex numbers.
Powers. The inner $P^{-1}P$ pairs cancel: $A^2 = PDP^{-1}PDP^{-1} = PD^2P^{-1}$, and in general
$$A^k = P D^k P^{-1}, \qquad D^k = \operatorname{diag}(\lambda_1^k, \dots, \lambda_n^k).$$Why do we need it?
Multiplying a matrix by itself many times is slow and messy. On the eigenvector axes the matrix is only a stretch, so powers become easy and the structure of the matrix becomes clear.
Where is it used?
Closed-form solutions of linear dynamical systems, fast $A^k$ (Fibonacci numbers, Markov chains, walks on graphs), PCA, and analysing how signals grow or shrink over many layers or time steps.
How is it used?
Find the eigenvalues and eigenvectors, put the eigenvectors in the columns of $P$ and the eigenvalues on the diagonal of $D$. Then $A^k = PD^kP^{-1}$. In NumPy: vals, P = np.linalg.eig(A), then P @ np.diag(vals**k) @ np.linalg.inv(P).
Order matters. The columns of $P$ and the diagonal entries of $D$ must be in the same order: if $\mathbf{v}_1$ is the first column of $P$, then $\lambda_1$ is the first entry of $D$. Swapping one without the other gives a wrong matrix.
Diagonalizable is not the same as invertible. The matrix $\operatorname{diag}(3,0)$ is diagonalizable but not invertible. And the shear is invertible but not diagonalizable. These are two separate questions.
Eigenvectors can be scaled freely (any non-zero multiple works in $P$), so $P$ is not unique. $D$ is unique up to the order of its entries.
Quick check: $A$ has eigenvalues $3$ and $-1$ with eigenvectors $[1,0]$ and $[1,1]$. What is $D$, and what are the eigenvalues of $A^{10}$?
$D = \operatorname{diag}(3, -1)$ (matching $P = \begin{bmatrix}1&1\\0&1\end{bmatrix}$, eigenvectors as columns). The eigenvalues of $A^{10}$ are $3^{10} = 59049$ and $(-1)^{10} = 1$.
Long-term behaviour: the dominant eigenvalue wins core
Start with any arrow $\mathbf{x}$ and keep applying $A$: $\mathbf{x},\ A\mathbf{x},\ A^2\mathbf{x},\ \dots$. Split $\mathbf{x}$ into its eigen-parts. Each time you apply $A$, every part is multiplied by its own eigenvalue.
It is like two runners, one at 2 metres per second and one at 1. After a few seconds the faster one is far ahead, and the total distance is almost entirely the faster one's. Likewise the part with the largest $|\lambda|$ (the dominant eigenvalue) grows fastest and soon swamps the rest. So whatever you start with, after many steps the arrow lines up with the dominant eigenvector.
Fibonacci numbers. The rule "next = this + previous" can be written as a matrix:
$$\begin{bmatrix} F_{k+1} \\ F_k \end{bmatrix} = \begin{bmatrix} 1 & 1 \\ 1 & 0 \end{bmatrix}\begin{bmatrix} F_k \\ F_{k-1} \end{bmatrix}.$$Starting from $[1, 0]$ and applying the matrix again and again gives $[1,1],\ [2,1],\ [3,2],\ [5,3],\ [8,5],\ \dots$. Look at the ratio of the two entries: $1,\ 2,\ 1.5,\ 1.667,\ 1.6,\ 1.625,\ \dots$ It settles near $1.618$.
Why? The characteristic equation is $\lambda^2 - \lambda - 1 = 0$, so $\lambda = \tfrac{1\pm\sqrt5}{2}$, i.e. $\lambda_1 \approx 1.618$ and $\lambda_2 \approx -0.618$. The dominant eigenvalue $1.618$ is the famous golden ratio. The other part is multiplied by a number smaller than 1 in size, so it dies out, and the arrow lines up with the dominant eigenvector $[\lambda_1, 1]$, whose entries have ratio $\lambda_1$.
Write $\mathbf{x} = c_1\mathbf{v}_1 + c_2\mathbf{v}_2 + \dots + c_n\mathbf{v}_n$ in the eigenvectors of a diagonalizable $A$. Then
$$A^k\mathbf{x} = c_1\lambda_1^k\mathbf{v}_1 + c_2\lambda_2^k\mathbf{v}_2 + \dots + c_n\lambda_n^k\mathbf{v}_n.$$If $|\lambda_1| \gt |\lambda_2| \ge \dots$ and $c_1 \ne 0$, then for large $k$: $A^k\mathbf{x} \approx c_1\lambda_1^k\mathbf{v}_1$. The other parts fade at the rate $|\lambda_2/\lambda_1|^k$.
What the sizes of the eigenvalues mean for repeated application:
| Eigenvalues | What happens to $A^k\mathbf{x}$ |
|---|---|
| all $|\lambda| \lt 1$ | shrinks to $\mathbf{0}$ (stable) |
| some $|\lambda| \gt 1$ | grows without bound, along the dominant eigenvector (unstable) |
| dominant $\lambda = 1$, others smaller | settles to a fixed arrow along $\mathbf{v}_1$ |
| complex pair with $|\lambda| \lt 1$ or $\gt 1$ | spirals in, or out |
Why do we need it?
When something is repeated again and again, we want to know where it ends up: does it blow up, die out or settle? The dominant eigenvalue answers that without simulating every step.
Where is it used?
PageRank and Markov chains (long-run probabilities), population growth models, stability of numerical algorithms, and exploding or vanishing signals in recurrent networks.
How is it used?
Compute the eigenvalues and compare their sizes. If every $|\lambda| \lt 1$ the result shrinks to zero. If the biggest $|\lambda| \gt 1$ it grows along the dominant eigenvector. If it equals 1 it settles. The ratio $|\lambda_2/\lambda_1|$ says how fast.
The starting vector matters a little. If you start with exactly none of the dominant eigenvector ($c_1 = 0$), you stay in the weaker direction. In practice rounding error or noise always adds a little of the dominant part, so it takes over eventually.
Equal-size eigenvalues. If $|\lambda_1| = |\lambda_2|$ (for example $+2$ and $-2$, or a complex pair) there is no single winner, and the arrow may flip or spin forever.
Quick check: $A$ has eigenvalues $0.9$ and $0.5$. What happens to $A^k\mathbf{x}$ for large $k$? And if the eigenvalues were $1.1$ and $0.5$?
With $0.9$ and $0.5$ both smaller than 1 in size, $A^k\mathbf{x}\to\mathbf{0}$ (slowly, along the $0.9$ direction). With $1.1$ and $0.5$ the part along the $1.1$ direction grows like $1.1^k$ without limit, so the vector blows up and points along that eigenvector.
Symmetric matrices and the spectral theorem core
A matrix is symmetric if it equals its own transpose, $A = A^\top$ (the entries mirror across the diagonal). Covariance matrices, Hessians and Gram matrices are all symmetric, so these are the matrices that matter most in ML. (A covariance matrix records how data features vary together, and a Hessian holds the second derivatives of a loss. You will meet both later in this chapter and in Chapter 1.14.)
Symmetric matrices are the best-behaved of all. Their eigen-directions are always at right angles, and their eigenvalues are always real (no rotation hiding inside). So a symmetric matrix does one clean thing: it stretches space along a set of perpendicular axes. The unit circle becomes an ellipse, and the axes of that ellipse are the eigenvectors.
$A = \begin{bmatrix} 2 & 1 \\ 1 & 2 \end{bmatrix}$ is symmetric. We found $\lambda_1 = 3$ with $[1,1]$ and $\lambda_2 = 1$ with $[1,-1]$.
- They are perpendicular: $[1,1]\cdot[1,-1] = 1 - 1 = 0$ ✓.
- Make them length 1 (divide by $\sqrt2$): $\mathbf{q}_1 = \tfrac{1}{\sqrt2}[1,1]$, $\mathbf{q}_2 = \tfrac{1}{\sqrt2}[1,-1]$.
- Each $\mathbf{q}_i\mathbf{q}_i^\top$ is a matrix that projects onto the line of $\mathbf{q}_i$: $\mathbf{q}_1\mathbf{q}_1^\top = \begin{bmatrix} \tfrac12 & \tfrac12 \\ \tfrac12 & \tfrac12 \end{bmatrix}$ and $\mathbf{q}_2\mathbf{q}_2^\top = \begin{bmatrix} \tfrac12 & -\tfrac12 \\ -\tfrac12 & \tfrac12 \end{bmatrix}$.
- Rebuild $A$ from its pieces: $3\begin{bmatrix} \tfrac12 & \tfrac12 \\ \tfrac12 & \tfrac12 \end{bmatrix} + 1\begin{bmatrix} \tfrac12 & -\tfrac12 \\ -\tfrac12 & \tfrac12 \end{bmatrix} = \begin{bmatrix} 2 & 1 \\ 1 & 2 \end{bmatrix}$ ✓.
Spectral theorem. If $A$ is a real symmetric $n\times n$ matrix, then:
- All eigenvalues are real.
- Eigenvectors belonging to different eigenvalues are orthogonal.
- There is always a full set of $n$ eigenvectors that are orthonormal (perpendicular, length 1). So $A$ is never defective.
Putting those orthonormal eigenvectors in the columns of $Q$ (an orthogonal matrix, so $Q^{-1} = Q^\top$, see Chapter 1.9) and the eigenvalues in $\Lambda$ (capital lambda, a diagonal matrix):
$$A = Q\Lambda Q^\top = \lambda_1\mathbf{q}_1\mathbf{q}_1^\top + \lambda_2\mathbf{q}_2\mathbf{q}_2^\top + \dots + \lambda_n\mathbf{q}_n\mathbf{q}_n^\top.$$This is the diagonalization $A = PDP^{-1}$ with the extra gift that $P^{-1}$ is just a transpose. The second form says: a symmetric matrix is a sum of rank-1 projection pieces, each weighted by its eigenvalue. Applying $A$ to $\mathbf{x}$ means: measure how much of $\mathbf{x}$ lies along each $\mathbf{q}_i$ (that is $\mathbf{q}_i^\top\mathbf{x}$), stretch it by $\lambda_i$, and add the pieces back up.
Why do we need it?
General matrices can have complex eigenvalues and slanted, overlapping eigen-directions. Symmetric matrices are clean: real eigenvalues and perpendicular directions, so both theory and computation are safe and simple.
Where is it used?
Covariance matrices (PCA), Hessians of loss functions, Gram and kernel matrices, graph Laplacians (spectral clustering), and the quadratic forms of the next chapter.
How is it used?
Use np.linalg.eigh, made for symmetric matrices: it returns real eigenvalues and orthonormal eigenvectors $Q$. Then $A = Q\Lambda Q^\top$, a sum of rank-1 pieces $\lambda_i\mathbf{q}_i\mathbf{q}_i^\top$. Dropping pieces with tiny $\lambda_i$ simplifies $A$.
Why are the eigenvectors of a symmetric matrix perpendicular? (short proof)
Let $A\mathbf{v}_1 = \lambda_1\mathbf{v}_1$ and $A\mathbf{v}_2 = \lambda_2\mathbf{v}_2$ with $\lambda_1 \ne \lambda_2$, and $A = A^\top$. Compute $\mathbf{v}_1\cdot(A\mathbf{v}_2)$ in two ways.
First, $\mathbf{v}_1\cdot(A\mathbf{v}_2) = \lambda_2\,(\mathbf{v}_1\cdot\mathbf{v}_2)$. Second, because $A$ is symmetric we can move it across the dot product: $\mathbf{v}_1\cdot(A\mathbf{v}_2) = (A\mathbf{v}_1)\cdot\mathbf{v}_2 = \lambda_1\,(\mathbf{v}_1\cdot\mathbf{v}_2)$.
So $(\lambda_1 - \lambda_2)(\mathbf{v}_1\cdot\mathbf{v}_2) = 0$. Since $\lambda_1 \ne \lambda_2$, the dot product must be $0$. ∎ (Real eigenvalues have a similar proof using complex conjugates; we skip it.)
Orthogonal eigenvectors need symmetry. The matrix $\begin{bmatrix}4&1\\2&3\end{bmatrix}$ has eigenvectors $[1,1]$ and $[1,-2]$, whose dot product is $1-2 = -1 \ne 0$. Not perpendicular, because the matrix is not symmetric.
Why this matters for PCA. A covariance matrix is symmetric, so by the proof above its eigenvectors (the principal components) are perpendicular: the new axes are clean, uncorrelated directions. This is the answer to "why do symmetric matrices have orthogonal eigenvectors, and why does it matter?".
If an eigenvalue is repeated, you get a whole eigenspace (a plane or more) and you choose an orthonormal basis inside it. Eigenvectors for different eigenvalues are automatically perpendicular.
Quick check: a symmetric $3\times3$ matrix has eigenvalues $4$, $1$, $0$ with unit eigenvectors $\mathbf{q}_1,\mathbf{q}_2,\mathbf{q}_3$. Write $A$ as a sum, and say what its rank is.
$A = 4\mathbf{q}_1\mathbf{q}_1^\top + 1\,\mathbf{q}_2\mathbf{q}_2^\top + 0\cdot\mathbf{q}_3\mathbf{q}_3^\top = 4\mathbf{q}_1\mathbf{q}_1^\top + \mathbf{q}_2\mathbf{q}_2^\top$. Only two pieces survive, so the rank is $2$ (the number of non-zero eigenvalues).
Eigenvectors in 3D: a ball becomes an ellipsoid
In 2D a symmetric matrix turns the unit circle into an ellipse. In 3D it turns the unit ball (a sphere) into an ellipsoid: a squashed, stretched egg. The picture is the same, with one extra axis.
The three eigenvectors are the three axes of the egg, perpendicular to each other. Each eigenvalue tells you how far the egg sticks out along its axis (a negative eigenvalue means the axis is flipped). If you find an arrow $\mathbf{x}$ such that $A\mathbf{x}$ points the same way, you have found one of those axes.
$A = \begin{bmatrix}2&1&0\\1&2&0\\0&0&1\end{bmatrix}$ (symmetric). Check three arrows:
- $A[1,1,0] = [3,3,0] = 3\,[1,1,0]$: eigenvalue $3$.
- $A[1,-1,0] = [1,-1,0] = 1\,[1,-1,0]$: eigenvalue $1$.
- $A[0,0,1] = [0,0,1] = 1\,[0,0,1]$: eigenvalue $1$.
The three directions are perpendicular. The ellipsoid is stretched 3 times along $[1,1,0]$ and left alone along the other two (the eigenvalue $1$ appears twice, so it has a whole plane of eigenvectors, as in the eigenspace section). The volume grows by $\det A = 3\cdot1\cdot1 = 3$.
For a symmetric $3\times3$ matrix $A = Q\Lambda Q^\top$, the image of the unit sphere $\{\|\mathbf{x}\|=1\}$ is an ellipsoid whose axes point along the orthonormal eigenvectors $\mathbf{q}_1,\mathbf{q}_2,\mathbf{q}_3$, with half-lengths $|\lambda_1|,|\lambda_2|,|\lambda_3|$. The volume is multiplied by $|\det A| = |\lambda_1\lambda_2\lambda_3|$. In $\mathbb{R}^n$ the same holds with $n$ axes.
Why do we need it?
Two dimensions can hide what happens when there are more directions. Watching a sphere turn into an ellipsoid shows what eigenvalues and eigenvectors mean in the real situation: data and models with many directions.
Where is it used?
PCA on 3D (and higher) data clouds, inertia and stress tensors in physics and engineering, covariance ellipsoids of Gaussian models, and the curvature of a loss that depends on three weights.
How is it used?
Call np.linalg.eigh(A). The eigenvectors give the directions of the ellipsoid's axes, and the absolute eigenvalues give how far each axis sticks out. The product of the eigenvalues is the volume scale.
Repeated eigenvalues make a round cross-section. If two eigenvalues are equal, the egg is circular when sliced in that plane, and any two perpendicular directions in the plane are valid eigenvectors.
A zero eigenvalue flattens the egg into a pancake or a needle: the matrix loses a dimension (it is singular).
Quick check: a symmetric $3\times3$ matrix has eigenvalues $4$, $1$ and $0.5$. How much does it scale volumes, and what shape is the image of the sphere?
Volumes are scaled by $4\cdot1\cdot0.5 = 2$. The sphere becomes an ellipsoid with half-lengths $4$, $1$ and $0.5$ along the three perpendicular eigenvector directions: long in one direction, thin in another.
Finding eigenvalues numerically: power iteration and the Rayleigh quotient
The characteristic polynomial is fine for $2\times2$, but useless for a matrix with a million rows (and for degree five and above there is no general formula for the roots at all). Computers need another way.
We already have one. The last section said that repeated multiplication turns any arrow toward the dominant eigenvector. So: pick any arrow, multiply by $A$, shrink it back to length 1 (so numbers don't explode), and repeat. That is power iteration. It needs only matrix-times-vector products, which is cheap even for giant sparse matrices (matrices that are mostly zeros).
Once the arrow has settled, how do we read off the eigenvalue? Ask: "how much does $A$ stretch $\mathbf{x}$, measured along $\mathbf{x}$ itself?" That is the Rayleigh quotient.
$A = \begin{bmatrix} 2 & 1 \\ 1 & 2 \end{bmatrix}$ (true eigenvalues $3$ and $1$). Start with $\mathbf{x}_0 = [1, 0]$ and skip the shrinking for now:
- $A\mathbf{x}_0 = [2, 1]$.
- $A[2,1] = [5, 4]$.
- $A[5,4] = [14, 13]$.
- $A[14,13] = [41, 40]$. The two entries are nearly equal: the arrow is turning toward $[1,1]$, the dominant eigenvector.
Rayleigh quotient of $\mathbf{x} = [2,1]$: $A\mathbf{x} = [5,4]$, so $\mathbf{x}^\top A\mathbf{x} = 2\cdot5 + 1\cdot4 = 14$ and $\mathbf{x}^\top\mathbf{x} = 4 + 1 = 5$. The estimate is $14/5 = 2.8$.
For $\mathbf{x} = [5,4]$: $A\mathbf{x} = [14,13]$, $\mathbf{x}^\top A\mathbf{x} = 70 + 52 = 122$, $\mathbf{x}^\top\mathbf{x} = 41$, estimate $= 2.976$. Closing in on $3$ ✓.
Power iteration. Start with any non-zero $\mathbf{x}_0$ and repeat
$$\mathbf{x}_{k+1} = \frac{A\mathbf{x}_k}{\|A\mathbf{x}_k\|}.$$If $|\lambda_1| \gt |\lambda_2|$ and $\mathbf{x}_0$ has some component along $\mathbf{v}_1$, then $\mathbf{x}_k \to \pm\mathbf{v}_1$. The error shrinks like $|\lambda_2/\lambda_1|^k$, so it is fast when the eigenvalues are well separated and slow when they are close. It finds only the largest eigenvalue (in size).
Rayleigh quotient. For a non-zero $\mathbf{x}$:
$$R(\mathbf{x}) = \frac{\mathbf{x}^\top A\mathbf{x}}{\mathbf{x}^\top\mathbf{x}}.$$- If $\mathbf{x}$ is an eigenvector, $R(\mathbf{x}) = \lambda$ exactly (because $\mathbf{x}^\top A\mathbf{x} = \lambda\,\mathbf{x}^\top\mathbf{x}$).
- For a symmetric $A$, $R(\mathbf{x})$ always lies between the smallest and largest eigenvalue, and it is a very accurate estimate: if $\mathbf{x}$ is off by a small error, $R(\mathbf{x})$ is off by only the square of that error.
Useful variants (awareness): inverse iteration applies $A^{-1}$ instead, and finds the smallest eigenvalue (the eigenvalues of $A^{-1}$ are $1/\lambda$). Shifted versions apply $(A - \sigma I)^{-1}$ and find the eigenvalue closest to $\sigma$.
Why do we need it?
Real matrices can have millions of rows, where solving a characteristic polynomial is impossible. We need a method that only multiplies the matrix by a vector, which is cheap.
Where is it used?
Google's PageRank, finding the first principal component, estimating the sharpest curvature (largest Hessian eigenvalue) of a neural network, and estimating spectral norms.
How is it used?
Start with any vector. Repeat: multiply by $A$, divide by the length. After a few dozen steps you have the dominant eigenvector, and the Rayleigh quotient $\mathbf{x}^\top A\mathbf{x}/\mathbf{x}^\top\mathbf{x}$ gives its eigenvalue. Stop when the vector stops changing.
Power iteration only finds the dominant eigenpair. To get others you need extra ideas (deflation: subtract the part you found; or shifts; or the QR algorithm below).
If the Rayleigh quotient of a non-symmetric matrix is used, it is still a reasonable estimate, but the "squared error" accuracy is a symmetric-matrix gift.
Quick check: what does the Rayleigh quotient of $\mathbf{x} = [1,1]$ give for $A = \begin{bmatrix}2&1\\1&2\end{bmatrix}$?
$A\mathbf{x} = [3,3]$, so $\mathbf{x}^\top A\mathbf{x} = 6$ and $\mathbf{x}^\top\mathbf{x} = 2$. $R = 3$, exactly the eigenvalue, because $[1,1]$ is an eigenvector.
The QR algorithm: how libraries really do it (awareness)
Power iteration finds one eigenvalue. Functions like np.linalg.eig find all of them. The classic method is the QR algorithm. It uses a trick with two ingredients you already know.
- Factor the matrix as $A = QR$ ($Q$ orthogonal, $R$ upper triangular; Gram–Schmidt, Chapter 1.9).
- Multiply the factors back in the opposite order: $A' = RQ$.
Surprisingly, $A'$ has the same eigenvalues as $A$, and it is a little "more triangular". Repeat many times and the numbers below the diagonal melt away. When the matrix is (nearly) triangular, the eigenvalues sit on the diagonal, and we already know how to read those.
$A = \begin{bmatrix} 2 & 1 \\ 1 & 2 \end{bmatrix}$. Gram–Schmidt gives $Q = \begin{bmatrix} 0.894 & -0.447 \\ 0.447 & 0.894 \end{bmatrix}$ and $R = \begin{bmatrix} 2.236 & 1.789 \\ 0 & 1.342 \end{bmatrix}$. Swap the order:
$$RQ \approx \begin{bmatrix} 2.8 & 0.6 \\ 0.6 & 1.2 \end{bmatrix}.$$The diagonal moved toward the true eigenvalues $3$ and $1$, and the off-diagonal fell from $1$ to $0.6$. Another round gives about $\begin{bmatrix} 2.976 & 0.22 \\ 0.22 & 1.024 \end{bmatrix}$, then $\begin{bmatrix} 2.997 & 0.07 \\ 0.07 & 1.003 \end{bmatrix}$, and so on, closing in on $\operatorname{diag}(3, 1)$.
Unshifted QR algorithm. Set $A_0 = A$. For $k = 0, 1, 2, \dots$: factor $A_k = Q_kR_k$, then set $A_{k+1} = R_kQ_k$.
Why the eigenvalues do not change: $R_k = Q_k^{-1}A_k$, so $A_{k+1} = Q_k^{-1}A_kQ_k$, which is similar to $A_k$ (same eigenvalues, see the key properties). Under mild conditions $A_k$ converges to a triangular matrix (diagonal for symmetric $A$) with the eigenvalues on the diagonal, largest first.
Real libraries speed this up a lot: they first reduce $A$ to a nearly-triangular "Hessenberg" form and use shifts. For you the takeaway is: eigenvalue software works by a stable sequence of orthogonal transformations, never by the characteristic polynomial.
Why do we need it?
Power iteration finds only one eigenvalue. We need a reliable method that finds all of them, without ever solving a high-degree polynomial.
Where is it used?
Inside np.linalg.eig, MATLAB's eig and LAPACK (the library under NumPy and SciPy): whenever a program computes all eigenvalues of a dense matrix.
How is it used?
You rarely code it yourself: call np.linalg.eig or eigh and trust it. Conceptually you factor $A = QR$, form $RQ$ and repeat, until the diagonal holds the eigenvalues. It is stable because it only uses orthogonal steps.
Quick check: why does $A_{k+1} = R_kQ_k$ have the same eigenvalues as $A_k = Q_kR_k$?
Because $R_k = Q_k^{-1}A_k$, so $A_{k+1} = Q_k^{-1}A_kQ_k$. That is a similarity transformation, and similar matrices have the same eigenvalues.
PageRank: the dominant eigenvector of the web
Imagine someone surfing the web at random: on each page they click a random link. After a very long time, what fraction of the time have they spent on each page? Pages that many (important) pages link to get visited more. That visit fraction is the page's rank.
One step of surfing turns "where the surfer probably is" into "where they probably are next": a matrix times a vector. The long-run answer is the vector that does not change when you apply the matrix: $\mathbf{r} = M\mathbf{r}$. That is an eigenvector with eigenvalue $1$, and it is the dominant one. So PageRank is power iteration.
Four pages A, B, C, D. Links: A → B, C | B → C | C → A | D → A, C. Column $j$ of $M$ says where a surfer on page $j$ goes next, splitting evenly over its links:
$$M = \begin{bmatrix} 0 & 0 & 1 & \tfrac12 \\ \tfrac12 & 0 & 0 & 0 \\ \tfrac12 & 1 & 0 & \tfrac12 \\ 0 & 0 & 0 & 0 \end{bmatrix}$$- Start with the surfer equally likely to be anywhere: $\mathbf{r}_0 = [\tfrac14,\tfrac14,\tfrac14,\tfrac14]$.
- One step: $\mathbf{r}_1 = M\mathbf{r}_0$. Row A: $0 + 0 + \tfrac14 + \tfrac18 = 0.375$. Row B: $\tfrac18 = 0.125$. Row C: $\tfrac18 + \tfrac14 + 0 + \tfrac18 = 0.5$. Row D: $0$. So $\mathbf{r}_1 = [0.375, 0.125, 0.5, 0]$, and it still adds up to 1 ✓.
- Keep going. Nobody links to D, so D's share drains to zero while A and C gain.
A Markov (column-stochastic) matrix has non-negative entries and each column adds to 1. It always has the eigenvalue $\lambda = 1$, and no eigenvalue has size above 1. A stationary distribution is an eigenvector with $M\mathbf{r} = \mathbf{r}$, scaled to add to 1.
To make the answer unique and the iteration reliable, PageRank adds a small chance of jumping to a random page (the damping factor $d$, about $0.85$):
$$G = d\,M + \frac{1-d}{n}\,\mathbf{1}\mathbf{1}^\top, \qquad \mathbf{r} = \lim_{k\to\infty} G^k\mathbf{r}_0,$$where $\mathbf{1}\mathbf{1}^\top$ is the $n\times n$ matrix of all ones.
The second eigenvalue of $G$ has size at most $d$, so the error shrinks like $d^k$: power iteration converges quickly.
Why do we need it?
A search engine must decide which of billions of pages matter most. "Important pages are those linked from important pages" is circular, and the eigenvector equation resolves the circle.
Where is it used?
Google search ranking, ranking people in social networks, citation analysis for research papers, recommender systems, and the long-run state of any Markov chain (including in reinforcement learning).
How is it used?
Build the link (transition) matrix and add damping to get $G$. Start with equal ranks and multiply by $G$ again and again until the vector stops changing. The final vector, scaled to add up to 1, is the rank of each page.
Quick check: why must the stationary vector have eigenvalue exactly 1?
"Stationary" means the surfer's distribution does not change after one step: $M\mathbf{r} = \mathbf{r}$. Compare with $A\mathbf{v} = \lambda\mathbf{v}$: this is the eigenvalue equation with $\lambda = 1$.
PCA: eigenvectors of the covariance matrix core
A cloud of data points is rarely round. It is usually stretched along some direction, like a cigar. The direction where the cloud is longest is where the data varies the most, so it carries the most information. The direction at a right angle to it is the next best.
PCA (principal component analysis) finds these directions. They are exactly the eigenvectors of the covariance matrix (a symmetric matrix that records how the features vary together). The matching eigenvalues say how much variance lies along each direction. Because the covariance matrix is symmetric, the spectral theorem promises the directions are perpendicular.
Five points, already centred (their mean is $[0,0]$): $[2,1],\ [-2,-1],\ [1,2],\ [-1,-2],\ [0,0]$. The covariance matrix is $C = \tfrac{1}{n-1}\sum \mathbf{x}\mathbf{x}^\top$ (here $n=5$).
- Sum of squares of the first coordinate: $4 + 4 + 1 + 1 + 0 = 10$. Same for the second: $1 + 1 + 4 + 4 + 0 = 10$.
- Sum of products: $2\cdot1 + (-2)(-1) + 1\cdot2 + (-1)(-2) + 0 = 8$.
- Divide by $n - 1 = 4$: $C = \begin{bmatrix} 2.5 & 2 \\ 2 & 2.5 \end{bmatrix}$.
- Eigenvalues: trace $5$, determinant $6.25 - 4 = 2.25$, so $\lambda^2 - 5\lambda + 2.25 = 0$ gives $\lambda = 4.5$ and $0.5$. Eigenvectors: $[1,1]/\sqrt2$ and $[1,-1]/\sqrt2$.
- So the first principal direction is the diagonal $[1,1]$, and it holds $4.5 / (4.5 + 0.5) = 90\%$ of the variance.
Given centred data with covariance matrix $C$ (symmetric), compute $C = Q\Lambda Q^\top$. Then:
- The columns of $Q$ are the principal components (directions), sorted by eigenvalue, largest first.
- The eigenvalue $\lambda_i$ is the variance of the data along $\mathbf{q}_i$.
- The fraction of variance explained by the first $k$ components is $\dfrac{\lambda_1 + \dots + \lambda_k}{\lambda_1 + \dots + \lambda_n}$.
- To compress, keep only the first $k$ components: project each point onto them (Chapter 1.9, projection). The lost information is the variance in the discarded directions.
Why do we need it?
Data has many features, but most of its variation often lies along a few directions. To compress, visualise or denoise it we need those directions, and the covariance matrix's eigenvectors are exactly them.
Where is it used?
Compressing images and embeddings, plotting high-dimensional data in 2D, removing noise, speeding up other models, face recognition ("eigenfaces"), and finding the main risk factors in finance.
How is it used?
Centre the data, compute the covariance matrix, take its eigenvectors with np.linalg.eigh, sort by eigenvalue and project the data onto the first few. The eigenvalue ratios tell you how much information you keep.
Centre the data first (subtract the mean of each feature), otherwise the first "direction" just points at the mean. Scale matters too: a feature measured in large units dominates the variance, so features are often standardised first.
Eigenvectors come with an arbitrary sign: a PCA library may flip PC1 between runs. That is harmless.
Quick check: a covariance matrix has eigenvalues $9, 1, 0$. What fraction of the variance does the first principal component explain, and what does the zero tell you?
The total is $9+1+0 = 10$, so PC1 explains $9/10 = 90\%$. A zero eigenvalue means there is no variation at all along that direction: the data lie flat in a 2-dimensional plane, and that direction can be dropped with no loss.
More places eigenvalues show up in ML
The same idea keeps returning: eigenvalues tell you which directions are strong and which are weak.
- Recurrent neural networks (RNNs) multiply the hidden state by the same matrix at every time step. After $T$ steps, a direction with eigenvalue $\lambda$ has been multiplied by $\lambda^T$. Tiny change per step, enormous change after many steps.
- Loss surfaces. Near a minimum, the loss is a bowl, and the eigenvalues of the Hessian (the matrix of second derivatives) say how steep the bowl is along each principal direction: big $\lambda$ = steep wall, small $\lambda$ = flat valley.
- Graphs. Eigenvectors of a graph's Laplacian reveal its natural clusters.
Exploding and vanishing. Suppose the dominant eigenvalue of an RNN's weight matrix is $0.9$. After 50 steps the signal along that direction is multiplied by $0.9^{50} \approx 0.005$, and after 100 steps by $0.9^{100} \approx 0.00003$. It has vanished. If instead $\lambda = 1.1$: $1.1^{50} \approx 117$ and $1.1^{100} \approx 13{,}781$. It has exploded.
Going backwards (backpropagation through time) the gradient is multiplied by the transpose of the matrix over and over, so gradients explode or vanish for the same reason. Exploding gradients are tamed by gradient clipping. Vanishing ones motivated LSTMs, GRUs and residual connections.
If a (diagonalizable) matrix $W$ is applied $T$ times, then as in the long-term section, $W^T\mathbf{h} \approx c\,\lambda_1^T\mathbf{v}_1$ for large $T$. So the size of the dominant eigenvalue (the spectral radius $\rho(W) = \max_i|\lambda_i|$) decides:
- $\rho(W) \lt 1$: signals and gradients vanish exponentially.
- $\rho(W) \gt 1$: they explode exponentially.
- $\rho(W) \approx 1$: balanced, which is why careful initialisation (for example, orthogonal matrices, whose eigenvalues all have size 1) helps.
Why do we need it?
Many ML problems boil down to "which directions are strong and which are weak?". Eigenvalues answer that for repeated products, for curvature and for graphs.
Where is it used?
Training RNNs (exploding or vanishing gradients, gradient clipping, orthogonal initialisation), studying loss surfaces (sharp versus flat minima), spectral clustering, and judging how hard an optimisation problem is.
How is it used?
Compute the spectral radius (largest $|\lambda|$) of the weight matrix: below 1 gradients vanish, above 1 they explode. For curvature, compute Hessian eigenvalues. For clustering, take the Laplacian eigenvectors with the smallest eigenvalues.
- Hessian curvature and sharp vs flat minima. Near a minimum, $L(\mathbf{w}) \approx L_0 + \tfrac12\boldsymbol\delta^\top H\boldsymbol\delta$ with $H$ symmetric. Its eigenvectors are perpendicular directions in weight space, and its eigenvalues are the curvatures along them. A large top eigenvalue is a sharp minimum (small changes in the weights hurt a lot, and gradient descent needs a small step size). All eigenvalues positive means a true minimum (Chapter 1.12). A mix of positive and negative eigenvalues means a saddle point.
- Spectral clustering. Build a graph whose nodes are data points and whose edges join similar points. Its Laplacian matrix $L = D - W$ is symmetric, so it has orthogonal eigenvectors. The eigenvectors with the smallest eigenvalues act like "cluster indicators": for two clusters, the sign of the second-smallest eigenvector tells you which side each point is on. Then run k-means on those eigenvector coordinates. This finds curved clusters that plain k-means cannot.
- PCA and whitening, Gaussian models, optimisation conditioning all reduce to the eigenvalues of a symmetric matrix. Their spread (largest ÷ smallest) is the condition number, which controls how slowly gradient descent crawls.
Quick check: $W$ has eigenvalues $0.5$ and $1.05$. After about 100 steps, which direction is basically gone and which has exploded?
$0.5^{100}\approx 10^{-30}$: that direction has vanished. $1.05^{100}\approx 131$: that direction has grown more than a hundredfold. The state ends up pointing along the $1.05$ eigenvector, and growing.
Recap, cheat sheet and practice
- An eigenvector of a square matrix is a direction it does not turn: $A\mathbf{v} = \lambda\mathbf{v}$. The eigenvalue $\lambda$ is the stretch (negative = flip, zero = squash). Eigenvectors are defined up to scale.
- By hand: solve $\det(A - \lambda I) = 0$ (the characteristic equation), then solve $(A - \lambda I)\mathbf{v} = \mathbf{0}$ for each $\lambda$. Triangular and diagonal matrices: the eigenvalues are the diagonal entries.
- Rotations have complex eigenvalues $r(\cos\theta \pm i\sin\theta)$: size $r$ = stretch per turn, angle $\theta$ = turn.
- The eigenspace is $N(A - \lambda I)$. Geometric multiplicity $\le$ algebraic multiplicity; if smaller, the matrix is defective.
- $\operatorname{tr}A = \sum\lambda_i$, $\det A = \prod\lambda_i$. Eigenvalues of $A^k$, $A^{-1}$, $A + cI$ are $\lambda^k$, $1/\lambda$, $\lambda + c$. Similar matrices share eigenvalues.
- Diagonalization $A = PDP^{-1}$ needs $n$ independent eigenvectors. Then $A^k = PD^kP^{-1}$, and for large $k$ the dominant eigenvalue takes over.
- Symmetric matrices: real eigenvalues, perpendicular eigenvectors, never defective: $A = Q\Lambda Q^\top = \sum\lambda_i\mathbf{q}_i\mathbf{q}_i^\top$.
- Computers use power iteration (plus the Rayleigh quotient) or the QR algorithm, not the characteristic polynomial.
- ML: PCA (covariance eigenvectors), PageRank (dominant eigenvector), Hessian curvature, exploding/vanishing gradients ($\lambda^T$), spectral clustering.
Cheat sheet
| Idea | Formula | Picture / note |
|---|---|---|
| Eigen equation | $A\mathbf{v} = \lambda\mathbf{v}$, $\mathbf{v}\ne\mathbf{0}$ | direction that is only stretched |
| Characteristic equation | $\det(A - \lambda I) = 0$ | $A - \lambda I$ must be singular |
| 2×2 shortcut | $\lambda^2 - \operatorname{tr}(A)\lambda + \det(A) = 0$ | quadratic formula |
| Eigenspace | $E_\lambda = N(A - \lambda I)$ | dimension = geometric multiplicity |
| Trace and determinant | $\sum\lambda_i$ and $\prod\lambda_i$ | total stretch, area scale |
| Functions of $A$ | $\lambda^k,\ 1/\lambda,\ \lambda + c$ | same eigenvectors |
| Diagonalization | $A = PDP^{-1}$, $A^k = PD^kP^{-1}$ | translate, stretch, translate back |
| Long-term | $A^k\mathbf{x} \approx c_1\lambda_1^k\mathbf{v}_1$ | dominant $|\lambda|$ wins |
| Spectral theorem | $A = Q\Lambda Q^\top = \sum\lambda_i\mathbf{q}_i\mathbf{q}_i^\top$ | circle becomes ellipse, axes = eigenvectors |
| Rayleigh quotient | $\mathbf{x}^\top A\mathbf{x} / \mathbf{x}^\top\mathbf{x}$ | $= \lambda$ when $\mathbf{x}$ is an eigenvector |
| Power iteration | $\mathbf{x} \leftarrow A\mathbf{x}/\|A\mathbf{x}\|$ | finds the dominant eigenpair |
| QR algorithm | $A_k = QR \Rightarrow A_{k+1} = RQ$ | similar matrices, becomes triangular |
import numpy as np
A = np.array([[4., 1.],
[2., 3.]])
# 1. Eigenvalues and eigenvectors (the columns of V are the eigenvectors)
vals, V = np.linalg.eig(A)
vals, V = vals.real, V.real # these eigenvalues are real, so drop any +0j
print(vals) # [5. 2.] (the order may differ on your machine)
print(np.allclose(A @ V, V * vals)) # True: A v = lambda v for every column
print(np.trace(A), vals.sum()) # 7.0 7.0 trace = sum of eigenvalues
print(np.linalg.det(A), vals.prod()) # about 10.0 and 10.0 (tiny rounding noise is normal)
# 2. Diagonalization A = P D P^-1 and a matrix power
P, D = V, np.diag(vals)
print(np.allclose(A, P @ D @ np.linalg.inv(P))) # True
A100 = P @ np.diag(vals**100) @ np.linalg.inv(P) # only 2 scalar powers needed
print(np.allclose(A100, np.linalg.matrix_power(A, 100))) # True
# 3. Symmetric matrices: use eigh (real eigenvalues, orthonormal eigenvectors)
S = np.array([[2., 1.],
[1., 2.]])
lam, Q = np.linalg.eigh(S) # eigenvalues in increasing order: [1. 3.]
print(np.allclose(Q.T @ Q, np.eye(2))) # True: Q is orthogonal
print(np.allclose(S, Q @ np.diag(lam) @ Q.T)) # True: S = Q Lambda Q^T
rank1_sum = sum(l * np.outer(q, q) for l, q in zip(lam, Q.T))
print(np.allclose(S, rank1_sum)) # True: S = sum of lambda_i q_i q_i^T
# 4. Complex eigenvalues of a rotation by 90 degrees
R = np.array([[0., -1.],
[1., 0.]])
print(np.linalg.eigvals(R)) # [0.+1.j 0.-1.j]
# 5. Power iteration + Rayleigh quotient
def power_iteration(M, steps=50, seed=0):
x = np.random.default_rng(seed).normal(size=M.shape[0])
for _ in range(steps):
x = M @ x
x = x / np.linalg.norm(x) # shrink back to length 1
return (x @ M @ x) / (x @ x), x # Rayleigh quotient, eigenvector
lam1, v1 = power_iteration(A)
print(lam1) # ~5.0 (matches np.linalg.eig)
# 6. Circle to ellipse, with the eigenvectors drawn (needs matplotlib)
import matplotlib.pyplot as plt
t = np.linspace(0, 2 * np.pi, 200)
circle = np.vstack([np.cos(t), np.sin(t)])
ellipse = S @ circle
plt.plot(*circle, "--", label="unit circle")
plt.plot(*ellipse, label="S times circle")
for l, q in zip(lam, Q.T):
plt.arrow(0, 0, *(l * q), head_width=0.08, color="green") # axis = lambda * q
plt.gca().set_aspect("equal"); plt.legend(); plt.show()
# 7. Tiny PageRank on a 5-node graph (column j = where page j links to)
links = {0: [1, 2], 1: [2], 2: [0], 3: [0, 2], 4: [3, 2]}
n, d = 5, 0.85
M = np.zeros((n, n))
for j, outs in links.items():
for i in outs:
M[i, j] = 1 / len(outs)
G = d * M + (1 - d) / n # damping: adds a small chance of jumping anywhere
r = np.ones(n) / n
for _ in range(100):
r = G @ r # power iteration
print(r.round(3), r.sum()) # ranks add up to 1
w, U = np.linalg.eig(G) # same answer from the eigenvector with eigenvalue 1
top = np.real(U[:, np.argmax(np.real(w))]); print((top / top.sum()).round(3))
1. What are the eigenvalues of $\begin{bmatrix}2&7\\0&5\end{bmatrix}$?
2. $[1, 2]$ is an eigenvector of $A$ with eigenvalue $3$. What is $A\,[2, 4]$?
3. A $2\times2$ matrix has trace $5$ and determinant $6$. Its eigenvalues are…
4. $A$ has eigenvalues $4$ and $0.5$. What are the eigenvalues of $A^{-1}$?
5. What are the eigenvalues of a $90^\circ$ rotation of the plane?
6. Which statement is guaranteed for a real symmetric matrix?
Practice problems
A. Find the eigenvalues and eigenvectors of $\begin{bmatrix}5&2\\2&2\end{bmatrix}$.
Trace $7$, determinant $10 - 4 = 6$, so $\lambda^2 - 7\lambda + 6 = (\lambda-6)(\lambda-1)$: $\lambda = 6$ and $1$.
$\lambda = 6$: $A - 6I = \begin{bmatrix}-1&2\\2&-4\end{bmatrix}$, so $x = 2y$ and $\mathbf{v} = [2,1]$. $\lambda = 1$: $A - I = \begin{bmatrix}4&2\\2&1\end{bmatrix}$, so $y = -2x$ and $\mathbf{v} = [1,-2]$. The matrix is symmetric, and indeed $[2,1]\cdot[1,-2] = 0$ ✓.
B. Diagonalise $A = \begin{bmatrix}3&1\\0&2\end{bmatrix}$ and find a formula for $A^k$.
Triangular, so $\lambda = 3, 2$. For $3$: $[1,0]$. For $2$: $A - 2I = \begin{bmatrix}1&1\\0&0\end{bmatrix}$ gives $x = -y$, so $[1,-1]$. Then $P = \begin{bmatrix}1&1\\0&-1\end{bmatrix}$, $D = \operatorname{diag}(3,2)$, and $P^{-1} = P$.
$A^k = PD^kP^{-1} = \begin{bmatrix}3^k & 3^k - 2^k\\ 0 & 2^k\end{bmatrix}$. Check $k=2$: $\begin{bmatrix}9&5\\0&4\end{bmatrix}$, which equals $A\cdot A$ ✓. So $A^5 = \begin{bmatrix}243&211\\0&32\end{bmatrix}$.
C. Is $\begin{bmatrix}1&3\\0&1\end{bmatrix}$ diagonalizable?
The only eigenvalue is $1$ (algebraic multiplicity 2). $A - I = \begin{bmatrix}0&3\\0&0\end{bmatrix}$ has rank 1, so its null space has dimension 1: geometric multiplicity 1. That is less than 2, so the matrix is defective and not diagonalizable.
D. A $3\times3$ matrix has eigenvalues $2$, $-1$, $3$. Give its trace, determinant, whether it is invertible, and the eigenvalues of $A^2$, $A^{-1}$ and $A + 5I$.
Trace $= 4$. Determinant $= 2\cdot(-1)\cdot3 = -6$, which is non-zero, so $A$ is invertible. $A^2$: $4, 1, 9$. $A^{-1}$: $\tfrac12, -1, \tfrac13$. $A + 5I$: $7, 4, 8$.
E. Do one step of power iteration (without shrinking) on $A = \begin{bmatrix}3&1\\1&3\end{bmatrix}$ from $\mathbf{x} = [1,0]$, then compute the Rayleigh quotient of the result.
$A\mathbf{x} = [3, 1]$. For $\mathbf{y} = [3,1]$: $A\mathbf{y} = [10, 6]$, $\mathbf{y}^\top A\mathbf{y} = 30 + 6 = 36$, $\mathbf{y}^\top\mathbf{y} = 10$, so $R = 3.6$. The true eigenvalues are $4$ and $2$ (trace 6, determinant 8), so $3.6$ is already close to the dominant one.
F. Write $\begin{bmatrix}3&1\\1&3\end{bmatrix}$ as $\sum\lambda_i\mathbf{q}_i\mathbf{q}_i^\top$.
Eigenvalue $4$ with $\mathbf{q}_1 = [1,1]/\sqrt2$, and eigenvalue $2$ with $\mathbf{q}_2 = [1,-1]/\sqrt2$. Then $4\begin{bmatrix}\tfrac12&\tfrac12\\\tfrac12&\tfrac12\end{bmatrix} + 2\begin{bmatrix}\tfrac12&-\tfrac12\\-\tfrac12&\tfrac12\end{bmatrix} = \begin{bmatrix}3&1\\1&3\end{bmatrix}$ ✓.
Quadratic Forms & Definiteness
Some functions look like a bowl, some like a hill, some like a horse's saddle. A single symmetric matrix can describe all of them, and its eigenvalues tell you which shape you have. This is the bridge between linear algebra and optimisation: it is how we know whether a point is a minimum.
- Write a quadratic form $f(\mathbf{x}) = \mathbf{x}^\top A\mathbf{x}$, and see why only the symmetric part of $A$ matters
- Picture its level sets (ellipses and hyperbolas) and find the principal axes with eigenvectors
- Name the five shapes: positive/negative definite, positive/negative semidefinite, indefinite
- Test definiteness with eigenvalues, with leading minors (Sylvester), and with Cholesky
- Know and prove the key facts: $A^\top A$ is always PSD, covariance matrices are PSD, PD means invertible, and adding $\lambda I$ (ridge) makes PSD into PD
- Connect it to convex losses, Mahalanobis distance, multivariate Gaussians and kernel matrices
Quadratic forms: turning a vector into one number core
With one variable, the simplest curved function is $f(x) = a\,x^2$: a parabola. If $a \gt 0$ it opens upward like a bowl, if $a \lt 0$ it opens downward like a hill.
A quadratic form is the same idea for a vector with many entries. Every term is a square ($x_1^2$) or a product of two entries ($x_1x_2$). Nothing else: no plain $x_1$ terms, no constants. The result is a single number for each vector, so the function draws a surface above the plane: a bowl, a hill, a trough or a saddle.
A matrix is the neat way to store all the coefficients. Multiply $\mathbf{x}$ by $A$, then take the dot product with $\mathbf{x}$ itself. In words: "how much does $A$ push $\mathbf{x}$ along $\mathbf{x}$?"
Let $A = \begin{bmatrix} 1 & 3 \\ 1 & 2 \end{bmatrix}$ and $\mathbf{x} = [2, 1]$.
- Multiply: $A\mathbf{x} = [1\cdot2 + 3\cdot1,\; 1\cdot2 + 2\cdot1] = [5, 4]$.
- Dot with $\mathbf{x}$: $\mathbf{x}^\top A\mathbf{x} = 2\cdot5 + 1\cdot4 = 14$.
- Check with the expanded formula. For general $[x, y]$: $\mathbf{x}^\top A\mathbf{x} = 1\cdot x^2 + 3\cdot xy + 1\cdot yx + 2\cdot y^2 = x^2 + 4xy + 2y^2$. At $[2,1]$: $4 + 8 + 2 = 14$ ✓.
The symmetric partner. Take $S = \begin{bmatrix} 1 & 2 \\ 2 & 2 \end{bmatrix}$ (the off-diagonal entries 3 and 1 were averaged to 2 and 2). Then $\mathbf{x}^\top S\mathbf{x} = x^2 + 2\cdot2\,xy + 2y^2 = x^2 + 4xy + 2y^2$. Exactly the same function. The matrices $A$ and $S$ look different but give the same quadratic form.
For a square matrix $A$ ($n\times n$), the quadratic form is the function
$$f(\mathbf{x}) = \mathbf{x}^\top A\mathbf{x} = \sum_{i=1}^n\sum_{j=1}^n a_{ij}\,x_ix_j.$$For a $2\times2$ matrix with entries $a,b,c,d$ this is $ax^2 + (b+c)\,xy + dy^2$.
Why only the symmetric part matters. The number $\mathbf{x}^\top A\mathbf{x}$ is a single number, and a single number equals its own transpose: $\mathbf{x}^\top A\mathbf{x} = (\mathbf{x}^\top A\mathbf{x})^\top = \mathbf{x}^\top A^\top\mathbf{x}$. So we can average the two:
$$\mathbf{x}^\top A\mathbf{x} = \mathbf{x}^\top\!\left(\frac{A + A^\top}{2}\right)\!\mathbf{x}.$$Split $A = S + K$ into its symmetric part $S = \tfrac{A+A^\top}{2}$ and its skew part $K = \tfrac{A-A^\top}{2}$. The skew part contributes exactly zero ($K\mathbf{x}$ is always perpendicular to $\mathbf{x}$). Two different matrices with the same symmetric part give the same function.
So from now on we always take $A$ symmetric. For a symmetric $\begin{bmatrix} a & b \\ b & c \end{bmatrix}$ the form is $ax^2 + 2b\,xy + cy^2$ (note the factor 2 on the cross term).
Why do we need it?
Many important quantities (squared error, energy, variance, curvature) have the shape "squares and cross-products of the entries". A quadratic form writes all of them as one compact matrix expression that we can analyse.
Where is it used?
The least-squares loss $\|X\mathbf{w}-\mathbf{y}\|^2$, the variance of data along a direction, ridge and weight-decay penalties, the curvature term of any smooth loss, and Gaussian models.
How is it used?
Write the function as $\mathbf{x}^\top A\mathbf{x}$ with $A$ symmetric (average $A$ and $A^\top$ if it is not). In NumPy it is x @ A @ x. Then ask questions about $A$ (eigenvalues, definiteness) instead of about the messy polynomial.
It is $\mathbf{x}^\top A\mathbf{x}$, not $A\mathbf{x}$. $A\mathbf{x}$ is a vector. The form is a number: a vector sandwiched by $\mathbf{x}^\top$ on the left and $\mathbf{x}$ on the right.
The cross term is doubled. For symmetric $\begin{bmatrix} a & b \\ b & c \end{bmatrix}$ the middle term is $2b\,xy$, not $b\,xy$, because the entry $b$ appears twice.
Definiteness (next sections) is about symmetric matrices. For a non-symmetric $A$, replace it by its symmetric part first.
Quick check: write $\mathbf{x}^\top A\mathbf{x}$ as a polynomial for $A = \begin{bmatrix}3&-1\\-1&2\end{bmatrix}$.
$3x^2 + 2(-1)\,xy + 2y^2 = 3x^2 - 2xy + 2y^2$. (The off-diagonal $-1$ appears twice, giving $-2xy$.)
Level sets: ellipses, hyperbolas and principal axes core
Think of a hiking map. A contour line joins all the places with the same height. For $f(\mathbf{x}) = \mathbf{x}^\top A\mathbf{x}$, the contour line at height $k$ is the level set $\{\mathbf{x} : f(\mathbf{x}) = k\}$.
- For a bowl, the contours are nested ellipses (like the rings seen from above).
- For a saddle, they are hyperbolas (two curved pieces that open away from each other).
The ellipses are not tilted randomly. Their axes point along the eigenvectors of $A$, and these are the principal axes. Along an eigenvector with a big eigenvalue the surface climbs steeply, so the contours crowd close together there. Along one with a small eigenvalue it climbs slowly and the ellipse stretches out.
$A = \begin{bmatrix} 2 & 1 \\ 1 & 2 \end{bmatrix}$ has eigenvalues $3$ and $1$, with unit eigenvectors $\mathbf{q}_1 = \tfrac{1}{\sqrt2}[1,1]$ and $\mathbf{q}_2 = \tfrac{1}{\sqrt2}[1,-1]$ (see the spectral theorem).
- Use the eigenvectors as new axes. A point $\mathbf{x}$ has new coordinates $y_1 = \mathbf{q}_1\cdot\mathbf{x} = \tfrac{x+y}{\sqrt2}$ and $y_2 = \mathbf{q}_2\cdot\mathbf{x} = \tfrac{x-y}{\sqrt2}$.
- In these coordinates the form is simple: $f = 3y_1^2 + 1\,y_2^2$. (Check: $3\tfrac{(x+y)^2}{2} + \tfrac{(x-y)^2}{2} = \tfrac{4x^2 + 4xy + 4y^2}{2} = 2x^2 + 2xy + 2y^2$ ✓.)
- The level set $f = 1$ is $3y_1^2 + y_2^2 = 1$, an ellipse. Its half-width along $\mathbf{q}_1$ is $1/\sqrt3 \approx 0.577$ and along $\mathbf{q}_2$ is $1/\sqrt1 = 1$.
- Check a point: $0.577\,\mathbf{q}_1 = [0.408, 0.408]$, and $2(0.1667) + 2(0.1667) + 2(0.1667) = 1$ ✓.
A saddle: $A = \begin{bmatrix} 1 & 2 \\ 2 & 1 \end{bmatrix}$ has eigenvalues $3$ and $-1$, so $f = 3y_1^2 - y_2^2$. The level set $f = 1$ is a hyperbola, and $f = 0$ is the pair of crossing straight lines $y_2 = \pm\sqrt3\,y_1$.
Let $A = Q\Lambda Q^\top$ be symmetric (spectral theorem). With new coordinates $\mathbf{y} = Q^\top\mathbf{x}$ (the coordinates along the eigenvectors),
$$f(\mathbf{x}) = \mathbf{x}^\top A\mathbf{x} = \mathbf{y}^\top\Lambda\mathbf{y} = \lambda_1y_1^2 + \lambda_2y_2^2 + \dots + \lambda_ny_n^2.$$The eigenvectors remove all the cross terms ("completing the square" done once and for all). So the signs of the eigenvalues decide the shape of the level sets $f = k$ (for $k \gt 0$, in 2D):
| Eigenvalues | Level set $f = k$ | Surface |
|---|---|---|
| both $\gt 0$ | ellipse, half-widths $\sqrt{k/\lambda_i}$ along $\mathbf{q}_i$ | bowl |
| one $\gt 0$, one $= 0$ | two parallel straight lines | trough |
| one $\gt 0$, one $\lt 0$ | hyperbola | saddle |
| both $\lt 0$ | empty (use $k \lt 0$: ellipses) | hill |
The principal axes are the lines along $\mathbf{q}_1,\dots,\mathbf{q}_n$. They are perpendicular, because $A$ is symmetric.
Why do we need it?
A formula alone does not show shape. Level sets (contour lines) and their axes show at a glance where a function is steep or flat, and in which directions.
Where is it used?
Understanding loss contours (why gradient descent zig-zags in a narrow valley), covariance ellipses in PCA and Gaussians, and error ellipses in statistics and sensor fusion.
How is it used?
Compute the eigen-decomposition of $A$. The eigenvectors are the axes of the ellipses, and $1/\sqrt{\lambda}$ is the half-width of the level set $f=1$. A big eigenvalue means a short axis (steep). Eigenvalues of mixed signs mean hyperbolas.
And in 3D? For three variables, the level set $\mathbf{x}^\top A\mathbf{x} = 1$ is a surface. If all three eigenvalues are positive it is a closed ellipsoid (an egg), with half-axes $1/\sqrt{\lambda_i}$ along the eigenvectors. If an eigenvalue is zero it becomes a tube that never closes. If the eigenvalues have mixed signs, it opens out into a hyperboloid. If all are negative, $\mathbf{x}^\top A\mathbf{x}$ is never $1$ and the level set is empty. So a closed ellipsoid is the 3D picture of "positive definite".
The axes follow the eigenvectors, not the matrix entries. If $b = 0$ (a diagonal matrix) the axes are the plain $x$ and $y$ axes. As soon as $b \ne 0$ the ellipse tilts, and the tilt is exactly the direction of the eigenvectors.
The bigger eigenvalue gives the shorter ellipse axis (half-width $\sqrt{k/\lambda}$): steep direction, tight contours.
Quick check: $f(x,y) = 4x^2 + y^2$. What are the eigenvalues, the principal axes, and the half-widths of the level set $f = 1$?
The matrix is $\operatorname{diag}(4, 1)$: eigenvalues $4$ and $1$, principal axes the $x$- and $y$-axes. The ellipse $4x^2 + y^2 = 1$ has half-width $1/\sqrt4 = 0.5$ along $x$ and $1/\sqrt1 = 1$ along $y$.
Definiteness: bowl, trough, hill, ridge and saddle core
Now ask one simple question about the surface $f(\mathbf{x}) = \mathbf{x}^\top A\mathbf{x}$: what signs does it take? Every such surface passes through height $0$ at the origin. Does it then go up in all directions, down in all directions, or a mixture?
- Bowl: up in every direction. The origin is a lowest point.
- Trough: never below zero, but flat along one line (like a half-pipe).
- Hill: down in every direction. The origin is a highest point.
- Ridge: never above zero, flat along one line.
- Saddle: up in some directions and down in others (like a horse's saddle, or a mountain pass).
These five shapes have official names. They are the vocabulary for "is this a minimum?" in optimisation.
- $A = \begin{bmatrix}2&1\\1&2\end{bmatrix}$: $f = 2x^2 + 2xy + 2y^2 = x^2 + y^2 + (x+y)^2$. A sum of squares, so $f \ge 0$, and $f = 0$ forces $x = y = 0$. Bowl.
- $A = \begin{bmatrix}1&1\\1&1\end{bmatrix}$: $f = x^2 + 2xy + y^2 = (x+y)^2$. Never negative, but $f = 0$ along the whole line $y = -x$ (try $[1,-1]$). Trough.
- $A = \begin{bmatrix}-1&0\\0&-2\end{bmatrix}$: $f = -x^2 - 2y^2 \le 0$. Hill.
- $A = \begin{bmatrix}1&2\\2&1\end{bmatrix}$: $f(1,0) = 1 \gt 0$ but $f(1,-1) = 1 - 4 + 1 = -2 \lt 0$. Both signs, so saddle.
Let $A$ be symmetric.
| Name | Condition on $f(\mathbf{x}) = \mathbf{x}^\top A\mathbf{x}$ | Eigenvalues of $A$ | Shape |
|---|---|---|---|
| Positive definite (PD) | $f \gt 0$ for all $\mathbf{x}\ne\mathbf{0}$ | all $\gt 0$ | bowl |
| Positive semidefinite (PSD) | $f \ge 0$ for all $\mathbf{x}$ | all $\ge 0$ | trough |
| Negative definite (ND) | $f \lt 0$ for all $\mathbf{x}\ne\mathbf{0}$ | all $\lt 0$ | hill |
| Negative semidefinite (NSD) | $f \le 0$ for all $\mathbf{x}$ | all $\le 0$ | ridge |
| Indefinite | $f$ takes both signs | some $\gt 0$, some $\lt 0$ | saddle |
Why the eigenvalue column is the same as the sign condition: from the last section, $f = \lambda_1y_1^2 + \dots + \lambda_ny_n^2$ with squares $y_i^2 \ge 0$. If all $\lambda_i \gt 0$, then $f \gt 0$ unless every $y_i = 0$. If some $\lambda_i \lt 0$, put all the weight on that $y_i$ and $f$ goes negative. And so on.
Notation: $A \succ 0$ means PD, $A \succeq 0$ means PSD. PD implies PSD (but not the other way round). Flipping the sign, $-A$ is PD exactly when $A$ is ND.
Why do we need it?
Optimisation asks "is this point a bottom?". Definiteness answers that for quadratic shapes. Without these five names we could not say whether a surface is a bowl, a hill or a saddle.
Where is it used?
Second-derivative tests in optimisation, convexity checks for losses, checking that covariance and kernel matrices are valid, and Newton's method (which needs a bowl-shaped Hessian).
How is it used?
Find the eigenvalues of the symmetric matrix and look at their signs: all positive is a bowl (PD), zeros allowed is PSD, all negative is a hill, mixed is a saddle. np.linalg.eigvalsh(A) gives the list quickly and reliably.
Definite is not the same as "all entries positive". $\begin{bmatrix}1&2\\2&1\end{bmatrix}$ has all positive entries but is indefinite. And $\begin{bmatrix}2&-1\\-1&2\end{bmatrix}$ has negative entries but is positive definite (eigenvalues $3$ and $1$). Look at eigenvalues, not entries.
Positive diagonal is necessary but not enough. A positive definite matrix always has a positive diagonal ($f(\mathbf{e}_i) = a_{ii}$), but a positive diagonal alone proves nothing.
Most symmetric matrices are none of PD, PSD, ND or NSD. Indefinite is the common case for a random symmetric matrix.
Quick check: classify $\begin{bmatrix}3&0\\0&-1\end{bmatrix}$ and $\begin{bmatrix}0&0\\0&5\end{bmatrix}$.
The first is diagonal with eigenvalues $3$ and $-1$ (mixed signs): indefinite (saddle). The second has eigenvalues $0$ and $5$: positive semidefinite but not definite ($f = 5y^2$ is zero along the whole $x$-axis).
How to test definiteness core
The definition says "$f \gt 0$ for every non-zero $\mathbf{x}$". We cannot try infinitely many vectors, so we need a finite test. There are three common ones.
- Eigenvalues. We showed $f = \sum\lambda_iy_i^2$. So just look at the signs of the eigenvalues. This is the most direct test.
- Leading principal minors (Sylvester's criterion, awareness). Take the determinants of the top-left $1\times1$, $2\times2$, $3\times3$, … corners. If every one is positive, the matrix is positive definite. This works by hand for small matrices.
- Try Cholesky. A positive number has a real square root. A positive definite matrix has a "matrix square root" $L$ with $A = LL^\top$. Run the Cholesky algorithm: if it finishes, the matrix is PD; if it needs the square root of a negative number (or zero), the matrix is not PD. It is the fastest test on a computer.
Is $A = \begin{bmatrix}2&1\\1&3\end{bmatrix}$ positive definite?
- Eigenvalues: trace $5$, determinant $6 - 1 = 5$, so $\lambda^2 - 5\lambda + 5 = 0$ and $\lambda = \tfrac{5\pm\sqrt5}{2} \approx 3.618,\ 1.382$. Both positive.
- Leading minors: top-left $1\times1$ is $2 \gt 0$. The full determinant is $5 \gt 0$. Both positive.
- Cholesky: $L_{11} = \sqrt2 \approx 1.414$. $L_{21} = 1/L_{11} \approx 0.707$. $L_{22} = \sqrt{3 - 0.707^2} = \sqrt{2.5} \approx 1.581$. It works. Check: $LL^\top$ has entries $1.414^2 = 2$, $1.414\cdot0.707 = 1$, and $0.707^2 + 1.581^2 = 0.5 + 2.5 = 3$ ✓.
All three agree: positive definite.
Now $B = \begin{bmatrix}1&2\\2&1\end{bmatrix}$. Leading minors: $1 \gt 0$ but $\det B = 1 - 4 = -3 \lt 0$. Fails. Cholesky: $L_{11} = 1$, $L_{21} = 2$, and $L_{22} = \sqrt{1 - 4} = \sqrt{-3}$: impossible. Fails. (Its eigenvalues are $3$ and $-1$: indefinite.)
For a symmetric $n\times n$ matrix $A$:
- Eigenvalue test. PD $\iff$ all $\lambda_i \gt 0$. PSD $\iff$ all $\lambda_i \ge 0$. ND/NSD: all $\lt 0$ / all $\le 0$. Mixed signs: indefinite.
- Sylvester's criterion. Let $\Delta_k$ be the determinant of the top-left $k\times k$ block. Then PD $\iff \Delta_1, \Delta_2, \dots, \Delta_n$ are all $\gt 0$. For ND the signs must alternate: $\Delta_1 \lt 0,\ \Delta_2 \gt 0,\ \Delta_3 \lt 0,\dots$. Careful: for PSD it is not enough that the leading minors are $\ge 0$; you would need every principal minor.
- Cholesky. PD $\iff$ $A = LL^\top$ for a lower-triangular $L$ with a positive diagonal. If you try the algorithm and every pivot (the number under a square root) is positive, you have also proved PD. (Details in Chapter 1.13.)
Why Cholesky proves PD: if $A = LL^\top$ with $L$ invertible, then $\mathbf{x}^\top A\mathbf{x} = \mathbf{x}^\top LL^\top\mathbf{x} = \|L^\top\mathbf{x}\|^2 \gt 0$ for $\mathbf{x}\ne\mathbf{0}$.
Why do we need it?
We cannot try infinitely many vectors $\mathbf{x}$ to check the definition. A quick finite test lets us decide definiteness from a few numbers.
Where is it used?
Libraries use Cholesky to validate covariance and kernel matrices (np.linalg.cholesky), optimisers check Hessian eigenvalues, and Sylvester's minors appear in hand calculations and economics.
How is it used?
In code, try np.linalg.cholesky(A): if it works $A$ is PD, if it raises an error it is not. For a clearer verdict, or for small matrices, check the signs of the eigenvalues or that all leading minors are positive.
Check symmetry first. All three tests assume a symmetric matrix. For a non-symmetric $A$, test $\tfrac{A+A^\top}{2}$.
Sylvester has two traps. It uses the leading (top-left) blocks, not any blocks. And a zero minor means "not PD", but it does not by itself tell you PSD versus indefinite.
Floating point. A matrix that is PSD in theory can show an eigenvalue like $-10^{-17}$ on a computer. Libraries use a small tolerance (or add a tiny "jitter" $\epsilon I$ before Cholesky).
Quick check: use Sylvester's criterion on $\begin{bmatrix}2&-1\\-1&1\end{bmatrix}$.
$\Delta_1 = 2 \gt 0$ and $\Delta_2 = 2\cdot1 - (-1)(-1) = 1 \gt 0$. Both positive, so it is positive definite (despite the negative off-diagonal entries).
Four key facts (and the ridge trick) core
One principle explains most of them: anything built as a "square" is bowl-shaped. A number squared is never negative. A vector's squared length is never negative. So if a quadratic form is secretly a squared length, it can never dip below zero.
And a bowl-shaped (PD) matrix is safe to invert: no direction is flattened to nothing. A matrix that is only PSD may have a flat valley (a zero eigenvalue), which is where inversion breaks. The cure is to add a tiny bit of bowl in every direction: that is what adding $\lambda I$ does, and that is ridge regularisation.
Let $B = \begin{bmatrix}1&2\\0&1\\1&0\end{bmatrix}$ (a $3\times2$ matrix; think of 3 data rows, 2 features). Then
$$B^\top B = \begin{bmatrix}1&0&1\\2&1&0\end{bmatrix}\begin{bmatrix}1&2\\0&1\\1&0\end{bmatrix} = \begin{bmatrix}2&2\\2&5\end{bmatrix}.$$- It is symmetric. Its eigenvalues: trace $7$, determinant $10 - 4 = 6$, so $\lambda^2 - 7\lambda + 6 = 0$ gives $6$ and $1$. Both positive: PD.
- Pick $\mathbf{x} = [1,-1]$. Then $B\mathbf{x} = [1-2,\ 0-1,\ 1-0] = [-1,-1,1]$, and $\|B\mathbf{x}\|^2 = 1 + 1 + 1 = 3$.
- And $\mathbf{x}^\top(B^\top B)\mathbf{x}$: $(B^\top B)\mathbf{x} = [2-2,\ 2-5] = [0,-3]$, and $\mathbf{x}\cdot[0,-3] = 0 + 3 = 3$ ✓. The same number, as the proof below says.
Now make the columns dependent: $B = \begin{bmatrix}1&2\\2&4\\3&6\end{bmatrix}$ (column 2 = 2 × column 1). Then $B^\top B = \begin{bmatrix}14&28\\28&56\end{bmatrix}$, with eigenvalues $70$ and $0$: PSD but singular. Add $\lambda = 1$: $B^\top B + I = \begin{bmatrix}15&28\\28&57\end{bmatrix}$ has eigenvalues $71$ and $1$ (each lifted by 1), determinant $855 - 784 = 71 \gt 0$. Now it is PD and invertible.
- $A^\top A$ is always PSD (for any matrix $A$, any shape). One-line proof: $$\mathbf{x}^\top(A^\top A)\mathbf{x} = (A\mathbf{x})^\top(A\mathbf{x}) = \|A\mathbf{x}\|^2 \ge 0.$$ It is also symmetric: $(A^\top A)^\top = A^\top A$. It is PD exactly when $A$ has independent columns (because $\|A\mathbf{x}\| = 0$ only if $A\mathbf{x} = \mathbf{0}$). The same goes for $AA^\top$.
- Covariance matrices are PSD. For centred data (rows $\mathbf{x}_i$, stacked in a matrix $X$), $C = \tfrac{1}{n-1}X^\top X$, which is PSD by fact 1. Meaning: $\mathbf{u}^\top C\mathbf{u}$ is the variance of the data projected on direction $\mathbf{u}$, and a variance can never be negative.
- PD $\Rightarrow$ invertible. If $A\mathbf{x} = \mathbf{0}$ then $\mathbf{x}^\top A\mathbf{x} = 0$, which for PD forces $\mathbf{x} = \mathbf{0}$. So $A$ has only the trivial null space. (Also: no zero eigenvalue means $\det A = \prod\lambda_i \ne 0$.) Moreover $A^{-1}$ is PD too, with eigenvalues $1/\lambda_i$. A PSD matrix with a zero eigenvalue is singular.
- Adding $\lambda I$ (with $\lambda \gt 0$) turns PSD into PD. The eigenvalues of $A + \lambda I$ are $\lambda_i + \lambda$ (same eigenvectors), and $\lambda_i \ge 0$ gives $\lambda_i + \lambda \ge \lambda \gt 0$. So $A^\top A + \lambda I$ is always invertible: that is the ridge regression system $(X^\top X + \lambda I)\mathbf{w} = X^\top\mathbf{y}$ (Chapter 1.10).
Why do we need it?
We often need a matrix to be safely invertible and bowl-shaped, yet all we have is something like $X^\top X$. These facts say what is guaranteed for free and how to repair the rest.
Where is it used?
The normal equations of linear regression, ridge regression and weight decay, covariance estimation (shrinkage), and Gram matrices in kernel methods.
How is it used?
Matrices built as $X^\top X$ are automatically PSD. If you must invert one that may be singular, add a small multiple of the identity, $X^\top X + \lambda I$, which is always PD, then solve with np.linalg.solve. Use the smallest $\lambda$ that works.
PSD can be singular, PD cannot. $A^\top A$ is PSD always, but it is invertible only when $A$ has independent columns. That is exactly why the plain normal equations can fail and the ridge version never does.
Adding $\lambda I$ changes the answer. Ridge trades a little bias for stability: it lifts every eigenvalue, including the ones you wanted to keep. Use the smallest $\lambda$ that works.
The converse of fact 1 is also true: every PSD matrix can be written as $B^\top B$ for some $B$ (for example $B = \Lambda^{1/2}Q^\top$ from its spectral decomposition).
Quick check: prove that $XX^\top$ is PSD, and say when it is PD.
$\mathbf{x}^\top(XX^\top)\mathbf{x} = (X^\top\mathbf{x})^\top(X^\top\mathbf{x}) = \|X^\top\mathbf{x}\|^2 \ge 0$. It is PD exactly when $\|X^\top\mathbf{x}\| = 0$ only for $\mathbf{x} = \mathbf{0}$, that is, when $X^\top$ has a trivial null space (the rows of $X$ are independent).
A PD Hessian means a minimum (and convexity) core
Zoom in on any smooth surface near a point where the ground is level (the gradient is zero; such a point is called a critical point). Close up, the surface looks like a tilt-free patch plus a gentle curve: a bowl, a hill or a saddle. That curve is a quadratic form.
The matrix inside that quadratic form is the Hessian $H$, the matrix of second derivatives (how the slope itself changes; Chapter 1.14 shows how to compute it, here we are simply given it). Near a critical point $\mathbf{x}_0$:
$f(\mathbf{x}_0 + \boldsymbol\delta) \approx f(\mathbf{x}_0) + \tfrac12\,\boldsymbol\delta^\top H\,\boldsymbol\delta.$
So the definiteness of $H$ tells you what kind of critical point you are standing on. That is the second-derivative test, in matrix form.
Take $f(x, y) = (x^2 - 1)^2 + y^2$ (two valleys with a ridge between them). The gradient is $[4x(x^2-1),\ 2y]$, which is zero at $(0,0)$, $(1,0)$ and $(-1,0)$. The Hessian is $H = \begin{bmatrix} 12x^2 - 4 & 0 \\ 0 & 2 \end{bmatrix}$.
- At $(1, 0)$: $H = \begin{bmatrix}8&0\\0&2\end{bmatrix}$. Eigenvalues $8, 2$, both positive: PD, so a local minimum. Same at $(-1, 0)$.
- At $(0, 0)$: $H = \begin{bmatrix}-4&0\\0&2\end{bmatrix}$. Eigenvalues $-4$ and $2$: indefinite, so a saddle point. The surface curves down along $x$ (the pass between the valleys) and up along $y$.
Least squares. For $L(\mathbf{w}) = \|X\mathbf{w} - \mathbf{y}\|^2$ the Hessian is $2X^\top X$, which is PSD everywhere (key fact 1). A loss with a PSD Hessian everywhere is a convex bowl, so any minimum you find is the global one.
Let $\nabla f(\mathbf{x}_0) = \mathbf{0}$ and let $H$ be the (symmetric) Hessian at $\mathbf{x}_0$.
| Hessian $H$ at the critical point | Conclusion |
|---|---|
| positive definite (all $\lambda \gt 0$) | local minimum |
| negative definite (all $\lambda \lt 0$) | local maximum |
| indefinite (both signs) | saddle point |
| only semidefinite (some $\lambda = 0$) | inconclusive: the test cannot decide |
Convexity. If $H(\mathbf{x})$ is PSD at every point, $f$ is convex: bowl-shaped everywhere, with no traps. Every local minimum is then a global minimum. If $H$ is PD everywhere, $f$ is strictly convex, with at most one minimum.
The eigenvalues are the curvatures along the principal directions (the eigenvectors). Large eigenvalue: steep. Small: flat.
Why do we need it?
A zero gradient does not say whether you found a bottom, a top or a mountain pass. The definiteness of the Hessian separates them, and convexity tells us when training cannot get trapped.
Where is it used?
Second-order optimisers (Newton, L-BFGS), proving that linear regression, logistic regression and SVMs have one global minimum, and studying saddle points and sharp or flat minima in deep networks.
How is it used?
At a point where the gradient is zero, compute the Hessian and its eigenvalues. All positive: local minimum. All negative: local maximum. Mixed: saddle. Any zero: look further. If the Hessian is PSD everywhere, the function is convex.
A zero gradient is not enough. Minimum, maximum and saddle all have zero gradient. Only the Hessian separates them.
Semidefinite Hessian = inconclusive. Both $x^2 + y^4$ (a minimum) and $x^2 - y^4$ (a saddle) have Hessian $\operatorname{diag}(2, 0)$ at the origin. You must look at higher-order terms.
"Local" is the keyword. A PD Hessian proves a minimum only nearby. Global statements need PSD Hessians everywhere (convexity).
Quick check: at a critical point the Hessian is $\begin{bmatrix}2&0\\0&-3\end{bmatrix}$. What is it?
Eigenvalues $2$ and $-3$ have opposite signs: indefinite, so a saddle point. The surface curves up along one axis and down along the other.
Mahalanobis distance and multivariate Gaussians
How far is a point from the centre of a data cloud? The ordinary (Euclidean) distance ignores the shape of the cloud. Picture a long thin cigar of points. A point 3 units out along the cigar is perfectly normal. A point 3 units out across the cigar is very unusual. Same distance, very different surprise.
Mahalanobis distance measures distance in units of spread: it stretches space so the cloud becomes round, then measures. Along directions where the data varies a lot, a step counts for little; along directions with little variation, the same step counts for a lot.
The formula is a quadratic form, with the inverse covariance $\Sigma^{-1}$ in the middle. For a Gaussian ("bell curve") cloud, the contour lines of the density are exactly the level sets of this quadratic form: tilted ellipses along the principal axes of $\Sigma$.
Take the covariance $\Sigma = \begin{bmatrix}2.5&2\\2&2.5\end{bmatrix}$ (eigenvalues $4.5$ along $[1,1]$ and $0.5$ along $[1,-1]$), centred at $\boldsymbol\mu = \mathbf{0}$. Its inverse is $\Sigma^{-1} = \tfrac{1}{2.25}\begin{bmatrix}2.5&-2\\-2&2.5\end{bmatrix}$.
- Point $\mathbf{p} = [1, 1]$ (along the long axis). $\Sigma^{-1}\mathbf{p} = \tfrac{1}{2.25}[0.5, 0.5] = [0.222, 0.222]$. So $d_M^2 = \mathbf{p}\cdot\Sigma^{-1}\mathbf{p} = 0.444$, and $d_M = 0.667$.
- Point $\mathbf{q} = [1,-1]$ (along the short axis). $\Sigma^{-1}\mathbf{q} = \tfrac{1}{2.25}[4.5, -4.5] = [2, -2]$. So $d_M^2 = 1\cdot2 + (-1)(-2) = 4$, and $d_M = 2$.
- Both points are $\sqrt2 \approx 1.414$ from the centre in the ordinary sense. But $\mathbf{q}$ is 3 times as "far" in spread units. It is the unusual one.
(Shortcut: along an eigenvector with eigenvalue $\lambda$ at Euclidean distance $r$, $d_M = r/\sqrt\lambda$. Check: $\sqrt2/\sqrt{4.5} = 0.667$ ✓ and $\sqrt2/\sqrt{0.5} = 2$ ✓.)
For a centre $\boldsymbol\mu$ and a covariance matrix $\Sigma$ (symmetric, positive definite), the Mahalanobis distance of $\mathbf{x}$ is
$$d_M(\mathbf{x}) = \sqrt{(\mathbf{x} - \boldsymbol\mu)^\top\,\Sigma^{-1}\,(\mathbf{x} - \boldsymbol\mu)}.$$With $\Sigma = I$ it is the ordinary distance. Since $\Sigma$ is PD, so is $\Sigma^{-1}$ (eigenvalues $1/\lambda_i$), so $d_M^2 \gt 0$ for any $\mathbf{x}\ne\boldsymbol\mu$: it really is a distance.
The multivariate Gaussian density on $\mathbb{R}^n$ is
$$p(\mathbf{x}) = \frac{1}{(2\pi)^{n/2}\sqrt{\det\Sigma}}\exp\!\Big(-\tfrac12\,d_M(\mathbf{x})^2\Big).$$It needs $\Sigma$ to be PD: we need $\Sigma^{-1}$ (so $\Sigma$ must be invertible) and $\det\Sigma = \prod\lambda_i \gt 0$. If $\Sigma$ is only PSD (a zero eigenvalue), all the data lie in a flat slice of space and no ordinary density exists. The level sets $d_M = k$ are ellipses with axes along the eigenvectors of $\Sigma$ and half-widths $k\sqrt{\lambda_i}$.
Why do we need it?
Plain distance ignores the shape of a data cloud. To judge how unusual a point is, we must measure in units of the data's own spread. That is a quadratic form with the inverse covariance in the middle.
Where is it used?
Anomaly and outlier detection, Gaussian mixture models, Gaussian discriminant analysis, Kalman filters, Gaussian processes, and whitening of features.
How is it used?
Estimate the mean and the covariance $\Sigma$, invert it (np.linalg.inv or a Cholesky solve), then compute $\sqrt{(\mathbf{x}-\boldsymbol\mu)^\top\Sigma^{-1}(\mathbf{x}-\boldsymbol\mu)}$. Large values mean unusual points. If $\Sigma$ is singular, add a small $\epsilon I$ first.
You must invert $\Sigma$. With fewer data points than features, the sample covariance is only PSD and singular, so Mahalanobis distance fails unless you add a small $\epsilon I$ (the ridge trick again).
It is the square root of the form. $d_M^2$ is the quadratic form; $d_M$ is its square root. Many formulas use the squared version.
Quick check: $\Sigma = \operatorname{diag}(4, 1)$ and $\boldsymbol\mu = \mathbf{0}$. Find $d_M$ of $[2, 0]$ and of $[0, 2]$.
$\Sigma^{-1} = \operatorname{diag}(\tfrac14, 1)$. For $[2,0]$: $d_M^2 = 4/4 = 1$, so $d_M = 1$. For $[0,2]$: $d_M^2 = 4\cdot1 = 4$, so $d_M = 2$. Same Euclidean distance (2), but $[0,2]$ is twice as unusual because the data barely vary in the $y$ direction.
Kernel (Gram) matrices are PSD
Imagine a table that says, for every pair of data points, how similar they are. That table is a kernel matrix $K$, with $K_{ij}$ = similarity of point $i$ and point $j$.
Not every table of numbers is a sensible similarity. A sensible one behaves like dot products of some feature vectors (possibly in a space with enormous or infinite dimension). Dot products between vectors form a Gram matrix $XX^\top$, and we proved such matrices are PSD. So a valid kernel matrix must be PSD. If a table has a negative eigenvalue, no feature space exists for it, and algorithms that assume one (SVMs, Gaussian processes, kernel ridge regression) can break.
Three points on a line: $x = 0, 1, 3$.
- Linear kernel $K_{ij} = x_ix_j$: $K = \begin{bmatrix}0&0&0\\0&1&3\\0&3&9\end{bmatrix} = \mathbf{x}\mathbf{x}^\top$. Eigenvalues $10, 0, 0$: PSD ✓ (it is a Gram matrix).
- RBF (Gaussian) kernel $K_{ij} = e^{-(x_i-x_j)^2/2}$: $K \approx \begin{bmatrix}1&0.607&0.011\\0.607&1&0.135\\0.011&0.135&1\end{bmatrix}$. All eigenvalues are positive: PD ✓.
- Distance table $D_{ij} = |x_i - x_j|$: $D = \begin{bmatrix}0&1&3\\1&0&2\\3&2&0\end{bmatrix}$. Its trace is $0$, so the eigenvalues add up to $0$. They cannot all be $\ge 0$ (the matrix is not zero), so some are negative: not PSD, not a kernel. Distances measure the opposite of similarity.
A symmetric function $k(\mathbf{x}, \mathbf{x}')$ is a valid kernel (Mercer's condition) exactly when, for every finite set of points, the Gram matrix $K_{ij} = k(\mathbf{x}_i, \mathbf{x}_j)$ is positive semidefinite. Common kernels: linear $\mathbf{x}\cdot\mathbf{x}'$, polynomial $(\mathbf{x}\cdot\mathbf{x}' + 1)^d$, and RBF $e^{-\|\mathbf{x}-\mathbf{x}'\|^2/(2\sigma^2)}$.
PSD means all eigenvalues $\ge 0$, i.e. $\mathbf{c}^\top K\mathbf{c} = \sum_{ij}c_ic_jk(\mathbf{x}_i,\mathbf{x}_j) \ge 0$ for any weights $\mathbf{c}$. The "kernel trick": any method that uses data only through dot products can swap $\mathbf{x}_i\cdot\mathbf{x}_j$ for $K_{ij}$.
Why do we need it?
Many methods need only a table of pairwise similarities. For that table to behave like real dot products, with all their nice guarantees, it must be positive semidefinite.
Where is it used?
Support vector machines, Gaussian processes, kernel ridge regression, kernel PCA and spectral clustering.
How is it used?
Build $K$ from a valid kernel such as the RBF, check that its smallest eigenvalue is not negative (np.linalg.eigvalsh(K)), and add a tiny jitter $\epsilon I$ before Cholesky. A clearly negative eigenvalue means your "similarity" is not a kernel.
PSD, not PD. Kernel matrices are often only PSD (if two points coincide, or with the linear kernel on few features). That is why Gaussian-process and kernel-ridge code adds a small jitter $\epsilon I$ before using Cholesky.
Similarity is not distance. Turning a distance into a kernel needs a transformation such as $e^{-d^2/2\sigma^2}$.
Quick check: a $2\times2$ "similarity" table is $\begin{bmatrix}1&2\\2&1\end{bmatrix}$. Could it be a kernel matrix?
No. Its eigenvalues are $3$ and $-1$, so it is indefinite, not PSD. (Another way: $\mathbf{c} = [1,-1]$ gives $\mathbf{c}^\top K\mathbf{c} = 1 - 2 - 2 + 1 = -2 \lt 0$.)
Recap, cheat sheet and practice
- A quadratic form is $f(\mathbf{x}) = \mathbf{x}^\top A\mathbf{x}$: one number per vector, made of squares and cross-products. Only the symmetric part $\tfrac{A+A^\top}{2}$ matters, so we take $A$ symmetric.
- Its level sets are ellipses (same-sign eigenvalues) or hyperbolas (mixed signs). The principal axes are the eigenvectors, and $f = \sum\lambda_iy_i^2$ in those axes.
- Definiteness: PD (all $\lambda \gt 0$, bowl), PSD ($\ge 0$, trough), ND ($\lt 0$, hill), NSD ($\le 0$, ridge), indefinite (mixed, saddle).
- Tests: eigenvalue signs; Sylvester (all leading principal minors positive means PD); Cholesky $A = LL^\top$ succeeds exactly for PD.
- Facts: $A^\top A$ is always PSD ($\mathbf{x}^\top A^\top A\mathbf{x} = \|A\mathbf{x}\|^2$); covariance matrices are PSD; PD implies invertible; $A + \lambda I$ lifts all eigenvalues by $\lambda$, so PSD becomes PD (ridge).
- ML: a PD Hessian at a critical point means a local minimum, PSD Hessians everywhere mean convexity; Mahalanobis distance uses $\Sigma^{-1}$; Gaussians need PD $\Sigma$; kernel matrices must be PSD.
Cheat sheet
| Idea | Formula | Picture / note |
|---|---|---|
| Quadratic form | $f(\mathbf{x}) = \mathbf{x}^\top A\mathbf{x} = \sum a_{ij}x_ix_j$ | height of a surface |
| Symmetric part | $\tfrac{A + A^\top}{2}$ gives the same $f$ | skew part adds 0 |
| 2×2 symmetric | $ax^2 + 2bxy + cy^2$ | cross term doubled |
| Principal axes | $f = \sum\lambda_iy_i^2$, $\mathbf{y} = Q^\top\mathbf{x}$ | eigenvectors = ellipse axes |
| PD / PSD | $\lambda_i \gt 0$ / $\lambda_i \ge 0$ | bowl / trough |
| ND / NSD | $\lambda_i \lt 0$ / $\lambda_i \le 0$ | hill / ridge |
| Indefinite | both signs | saddle |
| Sylvester | PD $\iff \Delta_1,\dots,\Delta_n \gt 0$ | leading minors only |
| Cholesky | PD $\iff A = LL^\top$ exists | fastest numerical test |
| Gram matrix | $\mathbf{x}^\top A^\top A\mathbf{x} = \|A\mathbf{x}\|^2 \ge 0$ | always PSD |
| Ridge | eigenvalues of $A + \lambda I$ are $\lambda_i + \lambda$ | PSD becomes PD, invertible |
| Hessian test | PD: min, ND: max, indefinite: saddle | semidefinite: inconclusive |
| Mahalanobis | $\sqrt{(\mathbf{x}-\boldsymbol\mu)^\top\Sigma^{-1}(\mathbf{x}-\boldsymbol\mu)}$ | distance in units of spread |
import numpy as np
# 1. A quadratic form, and why only the symmetric part matters
B = np.array([[1., 3.],
[1., 2.]])
x = np.array([2., 1.])
S = (B + B.T) / 2 # symmetric part
print(x @ B @ x, x @ S @ x) # 14.0 14.0 (same number)
# 2. Classify a symmetric matrix by its eigenvalues
def classify(M, tol=1e-9):
lam = np.linalg.eigvalsh((M + M.T) / 2) # eigvalsh: real eigenvalues, ascending
if np.all(lam > tol): return "positive definite (bowl)"
if np.all(lam >= -tol): return "positive semidefinite (trough)"
if np.all(lam < -tol): return "negative definite (hill)"
if np.all(lam <= tol): return "negative semidefinite (ridge)"
return "indefinite (saddle)"
shapes = {"bowl": [[2, 0.5], [0.5, 1]],
"trough": [[1, 1], [1, 1]],
"hill": [[-2, -0.5], [-0.5, -1]],
"saddle": [[1, 2], [2, 1]]}
for name, M in shapes.items():
print(name, "->", classify(np.array(M, dtype=float)))
# 3. Other tests: leading principal minors (Sylvester) and Cholesky
A = np.array([[2., -1., 0.],
[-1., 2., -1.],
[0., -1., 2.]])
minors = [np.linalg.det(A[:k, :k]) for k in (1, 2, 3)]
print(np.round(minors, 3)) # [2. 3. 4.] all positive => PD
L = np.linalg.cholesky(A) # works only if A is positive definite
print(np.allclose(L @ L.T, A)) # True
try:
np.linalg.cholesky(np.array([[1., 2.], [2., 1.]]))
except np.linalg.LinAlgError:
print("Cholesky failed: not positive definite")
# 4. A^T A is always PSD; adding lambda*I (ridge) makes it PD
X = np.array([[1., 2.], [2., 4.], [3., 6.]]) # dependent columns
G = X.T @ X
print(np.linalg.eigvalsh(G).round(6) + 0.0) # [ 0. 70.] PSD, singular
ridge = G + 1.0 * np.eye(2)
print(np.linalg.eigvalsh(ridge).round(6)) # [ 1. 71.] every eigenvalue lifted by 1
w = np.linalg.solve(ridge, X.T @ np.array([1., 2., 3.])) # ridge solution exists
# 5. Mahalanobis distance
Sigma = np.array([[2.5, 2.0], [2.0, 2.5]])
p = np.array([1., -1.])
d_m = np.sqrt(p @ np.linalg.inv(Sigma) @ p)
print(np.linalg.norm(p), d_m) # 1.414 2.0 (same Euclidean distance, very different Mahalanobis)
# 6. Hessian test at the critical points of f(x,y) = (x^2-1)^2 + y^2
def hessian(x, y):
return np.array([[12 * x**2 - 4, 0.], [0., 2.]])
for pt in [(0, 0), (1, 0), (-1, 0)]:
print(pt, classify(hessian(*pt))) # (0,0) saddle, (+-1,0) bowl => local minima
# 7. Contours and a 3D surface of x^T A x (needs matplotlib)
import matplotlib.pyplot as plt
g = np.linspace(-2, 2, 200)
XX, YY = np.meshgrid(g, g)
fig = plt.figure(figsize=(12, 3))
for i, (name, M) in enumerate(shapes.items()):
M = np.array(M, dtype=float)
Z = M[0, 0] * XX**2 + 2 * M[0, 1] * XX * YY + M[1, 1] * YY**2 # x^T M x
ax = fig.add_subplot(1, 4, i + 1)
ax.contour(XX, YY, Z, levels=15); ax.set_title(name); ax.set_aspect("equal")
fig3 = plt.figure(); ax3 = fig3.add_subplot(projection="3d")
ax3.plot_surface(XX, YY, Z, cmap="coolwarm") # the last shape (the saddle) in 3D
plt.show()
1. What is $\mathbf{x}^\top A\mathbf{x}$ for $A = \begin{bmatrix}1&2\\2&1\end{bmatrix}$ and $\mathbf{x} = [1,-1]$?
2. Which matrix is positive definite?
3. For any real matrix $A$ (any shape), the matrix $A^\top A$ is always…
4. At a point with zero gradient, the Hessian has eigenvalues $3$ and $-2$. This point is…
5. $X^\top X$ has eigenvalues $5$ and $0$. After adding $0.1\,I$ (ridge), the eigenvalues are…
6. With $\Sigma = \operatorname{diag}(4, 1)$ and mean $\mathbf{0}$, what is the Mahalanobis distance of $[0, 2]$?
Practice problems
A. Write $\mathbf{x}^\top A\mathbf{x}$ for $A = \begin{bmatrix}3&2\\2&1\end{bmatrix}$ as a polynomial, and classify $A$.
$f = 3x^2 + 2\cdot2\,xy + 1\,y^2 = 3x^2 + 4xy + y^2$. Determinant $3 - 4 = -1 \lt 0$, so the eigenvalues have opposite signs: indefinite (a saddle). Check: $f(1,-1) = 3 - 4 + 1 = 0$ and $f(1,-2) = 3 - 8 + 4 = -1 \lt 0$, while $f(1,0) = 3 \gt 0$.
B. Is $A = \begin{bmatrix}2&5\\-1&3\end{bmatrix}$ positive definite as a quadratic form?
Only the symmetric part counts: $S = \begin{bmatrix}2&2\\2&3\end{bmatrix}$ (average of $5$ and $-1$ is $2$). $f = 2x^2 + 4xy + 3y^2$. Leading minors: $2 \gt 0$ and $6 - 4 = 2 \gt 0$. Both positive, so yes, positive definite.
C. Use Sylvester's criterion on $\begin{bmatrix}4&2&0\\2&3&1\\0&1&2\end{bmatrix}$.
$\Delta_1 = 4$. $\Delta_2 = 12 - 4 = 8$. $\Delta_3 = 4(3\cdot2 - 1\cdot1) - 2(2\cdot2 - 1\cdot0) + 0 = 20 - 8 = 12$. All positive, so it is positive definite.
D. Show $A = \begin{bmatrix}1&2\\2&4\end{bmatrix}$ is PSD but not PD, and find a direction where $f = 0$.
$f = x^2 + 4xy + 4y^2 = (x + 2y)^2 \ge 0$, so PSD. It is $0$ whenever $x = -2y$, for example $\mathbf{x} = [2,-1]$ (a non-zero vector), so it is not PD. Eigenvalues: trace $5$, determinant $0$, so $5$ and $0$. The null direction $[2,-1]$ is the eigenvector for $0$.
E. $X^\top X$ has eigenvalues $9$, $4$ and $0$. What are the eigenvalues and the ratio $\lambda_{\max}/\lambda_{\min}$ after adding $0.5\,I$?
Eigenvalues become $9.5$, $4.5$, $0.5$. The ratio is $9.5/0.5 = 19$. Before ridge the smallest eigenvalue was $0$ (infinite ratio, not invertible). Now it is PD and well behaved.
F. Classify the critical point at the origin of $f(x,y) = x^2 + 4xy + y^2$.
The gradient is $[2x + 4y,\ 4x + 2y]$, zero at $(0,0)$. The Hessian is $\begin{bmatrix}2&4\\4&2\end{bmatrix}$, with trace $4$, determinant $4 - 16 = -12 \lt 0$. Eigenvalues are $6$ and $-2$: indefinite, so a saddle point.
Matrix Decompositions
Every number can be broken into smaller factors, like $12 = 3 \times 4$. Every matrix can be too. A decomposition splits a hard matrix into simple pieces, and each piece makes one job easy: solving equations, finding structure, or squeezing data. The star of this chapter is the SVD, which works for every matrix and shows up all over machine learning.
- Explain why we factorise a matrix: solve faster, reveal structure, compress
- Know LU, Cholesky, QR and eigendecomposition: what each one is and when to reach for it
- Understand the SVD as rotate, stretch, rotate, and read rank and all four subspaces from it
- Build the best low-rank approximation with a truncated SVD (image compression, PCA, LoRA)
- See how the six decompositions connect, and pick the right one for a job
Why factorise a matrix? core
Think about the number $12$. On its own it tells you little. Written as $2 \times 2 \times 3$, you suddenly see that it is even, that it is divisible by 6, and that it is not prime. Breaking something into simple pieces reveals its structure.
Matrices work the same way. Some matrices are very easy to work with:
- Triangular matrices (zeros on one side of the diagonal). A system with a triangular matrix is solved by simple substitution, one unknown at a time.
- Orthogonal matrices (rotations and flips). Their inverse is just the transpose, and they never stretch anything.
- Diagonal matrices. They only scale each axis separately.
A decomposition rewrites a general matrix as a product of these easy pieces. We pay once to factorise, and then every later job is cheap.
Why triangular is easy. Solve $\begin{bmatrix} 2 & 1 \\ 0 & 3 \end{bmatrix}\mathbf{x} = \begin{bmatrix} 5 \\ 6 \end{bmatrix}$.
- The bottom row says $3x_2 = 6$, so $x_2 = 2$.
- The top row says $2x_1 + 1\cdot x_2 = 5$. Put in $x_2 = 2$: $2x_1 = 3$, so $x_1 = 1.5$.
No elimination was needed. That is why most decompositions try to produce triangular, orthogonal or diagonal pieces.
A matrix decomposition (or factorisation) writes a matrix as a product of simpler matrices, for example $A = LU$ or $A = U\Sigma V^\top$. There are three big reasons to do it:
- Solve faster. Factor once, then reuse the factors for many right-hand sides (many different vectors $\mathbf{b}$ in $A\mathbf{x} = \mathbf{b}$). LU, Cholesky and QR all do this.
- Reveal structure. The factors show rank, directions, stretch factors, and the four fundamental subspaces (eigen, SVD).
- Compress. Keep only the biggest pieces and throw the rest away (truncated SVD).
The map of this chapter: LU (for solving), Cholesky (LU's fast cousin for symmetric positive definite matrices), QR (stable least squares), eigendecomposition (square matrices with enough eigenvectors), and the SVD (any matrix at all).
Why do we need it?
Big matrix jobs (solving equations, fitting a model, compressing data) are slow or messy on the raw matrix. Writing the matrix as simple pieces makes each job easy, safe and fast.
Where is it used?
np.linalg.solve (LU), least-squares fitting (QR), PCA and image compression (SVD), Gaussian sampling (Cholesky), and almost every routine in NumPy, SciPy, PyTorch and scikit-learn.
How is it used?
Decide the job first: solve, fit, sample, find structure, or compress. Pick the factorisation that fits (the picture at the end of this chapter helps), call the library once, and reuse the factors.
In practice, never compute $A^{-1}$ just to solve $A\mathbf{x} = \mathbf{b}$. Forming the inverse costs about three times a factorisation and is usually less accurate. Factorise, then do two cheap triangular solves. Libraries do exactly this when you call np.linalg.solve.
Quick check: after you have the factors, what does one more solve cost compared with starting over?
It costs about $n^2$-scale work (two triangular solves) instead of $n^3$-scale work (a fresh elimination). For $n = 1000$ that is a few hundred times less: about $2n^2 = 2$ million operations instead of about $\tfrac23 n^3 \approx 670$ million.
LU decomposition
In Chapter 1.6 you solved equations with Gaussian elimination: subtract multiples of one row from the rows below it until everything under the diagonal is zero. The result is an upper triangular matrix.
Normally you forget how much of each row you subtracted. LU decomposition simply writes those amounts down. Picture a cooking diary. $U$ is the finished dish (what elimination produced). $L$ is the diary of the steps you took. Together they let you redo the whole job at any time.
Let $A = \begin{bmatrix} 2 & 1 & 1 \\ 4 & 3 & 3 \\ 8 & 7 & 9 \end{bmatrix}$. Eliminate below the first pivot (the $2$):
- Row 2 has $4$ under the pivot. The multiplier is $4 / 2 = 2$. Row 2 $-$ $2\times$ Row 1 $= [0, 1, 1]$.
- Row 3 has $8$ under the pivot. The multiplier is $8 / 2 = 4$. Row 3 $-$ $4\times$ Row 1 $= [0, 3, 5]$.
- Now the second pivot is $1$. Row 3 has $3$ under it, so the multiplier is $3 / 1 = 3$. Row 3 $-$ $3\times$ Row 2 $= [0, 0, 2]$.
The finished triangular matrix is $U$. The multipliers $2, 4, 3$ go into $L$, each in the place where it cleared a zero:
$$\underbrace{\begin{bmatrix} 2 & 1 & 1 \\ 4 & 3 & 3 \\ 8 & 7 & 9 \end{bmatrix}}_{A} = \underbrace{\begin{bmatrix} 1 & 0 & 0 \\ 2 & 1 & 0 \\ 4 & 3 & 1 \end{bmatrix}}_{L}\underbrace{\begin{bmatrix} 2 & 1 & 1 \\ 0 & 1 & 1 \\ 0 & 0 & 2 \end{bmatrix}}_{U}$$Check one entry. Row 3 of $LU$ is $4\cdot[2,1,1] + 3\cdot[0,1,1] + 1\cdot[0,0,2] = [8, 7, 9]$, which is row 3 of $A$. ✓
An LU decomposition of a square matrix $A$ is
$$A = LU$$where $L$ is lower triangular with 1s on its diagonal (its entries below the diagonal are the elimination multipliers) and $U$ is upper triangular (the result of elimination).
Elimination sometimes meets a zero (or a tiny) pivot and must swap rows. Keeping a record of the swaps in a permutation matrix $P$ gives the version that always works for an invertible matrix:
$$PA = LU$$Solving $A\mathbf{x} = \mathbf{b}$ with it. Since $LU\mathbf{x} = P\mathbf{b}$, do two cheap triangular solves:
- Forward substitution: solve $L\mathbf{y} = P\mathbf{b}$ from the top row down.
- Back substitution: solve $U\mathbf{x} = \mathbf{y}$ from the bottom row up.
Factorising costs about $\tfrac23 n^3$ operations, but each new $\mathbf{b}$ costs only about $2n^2$. So one LU serves many right-hand sides. As a bonus, $\det A = \pm$ the product of the diagonal of $U$ (the sign comes from the number of swaps).
Why do we need it?
Gaussian elimination solves one system. Real work often has many right-hand sides with the same matrix. LU saves the elimination so every later solve is cheap, and it also gives determinants.
Where is it used?
np.linalg.solve and scipy.linalg.lu_factor, determinants (slogdet), engineering and circuit simulations, and fitting the same model matrix to many target columns at once.
How is it used?
Call lu, piv = scipy.linalg.lu_factor(A) once, then x = lu_solve((lu, piv), b) for each new b. Keep the pivoting (the P) so zero or tiny pivots cannot break it. Never invert A.
- A zero pivot breaks plain $A = LU$. The matrix $\begin{bmatrix} 0 & 1 \\ 1 & 1 \end{bmatrix}$ is perfectly invertible, but elimination has to divide by $0$ in the first step. Swap the two rows and all is well: that is $PA = LU$.
- A tiny pivot is nearly as bad as a zero one. It gives huge multipliers and amplifies rounding errors. So real software always uses partial pivoting (pick the largest entry in the column as pivot). You will see why in Chapter 1.15.
- $L$ always has 1s on its diagonal. The pivots live in $U$.
Quick check: in $A = LU$, what do the entries below the diagonal of $L$ mean?
They are the multipliers used during elimination. Entry $l_{ij}$ says "I subtracted $l_{ij}$ times row $j$ from row $i$ to clear the entry in position $(i,j)$".
Cholesky decomposition: the square root of a matrix core
Start with a number. $9$ can be split into two equal halves: $9 = 3 \times 3$. We call $3$ the square root of $9$.
Can we do the same with a matrix? Can we split a matrix $A$ into two equal halves, $A = L \times L^\top$, where the second half is just the mirror image of the first? For one special family of matrices, yes. The family is the symmetric positive definite matrices (we will explain "positive definite" in a moment). The splitting is called the Cholesky decomposition (said "Ko-LESS-kee").
There is one more detail. We ask that $L$ is lower triangular: all the numbers above its diagonal are zero. That is a gift, because a triangular matrix is very easy to work with (you solve equations with it one unknown at a time, as in the LU section). So "half of a matrix" is not only half the size of the problem. It is also the easiest kind of half.
Why "positive"? You can only take the square root of a positive number (in real numbers, $\sqrt{-4}$ does not exist). In the same way, you can only take the square root of a positive matrix. "Positive definite" is the matrix version of "positive number". Picture the surface $\mathbf{x}^\top A\mathbf{x}$: for a positive definite matrix it is a bowl, curving up in every direction. If any direction curves down or goes flat, there is no real square root.
A tiny case first. A $1\times1$ matrix $[9]$ has "square root" $[3]$, because $[3][3] = [9]$. This is just $\sqrt9 = 3$.
A $2\times2$ case. Take $A = \begin{bmatrix} 4 & 2 \\ 2 & 5 \end{bmatrix}$ and look for $L = \begin{bmatrix} l_{11} & 0 \\ l_{21} & l_{22} \end{bmatrix}$. Multiplying out, $LL^\top = \begin{bmatrix} l_{11}^2 & l_{11}l_{21} \\ l_{11}l_{21} & l_{21}^2 + l_{22}^2 \end{bmatrix}$. Match each entry of $A$:
- Top-left: $l_{11}^2 = 4$, so $l_{11} = 2$.
- Bottom-left: $l_{11}\,l_{21} = 2$, so $l_{21} = 2 / 2 = 1$.
- Bottom-right: $l_{21}^2 + l_{22}^2 = 5$, so $l_{22}^2 = 5 - 1 = 4$ and $l_{22} = 2$.
A $3\times3$ case. Take $A = \begin{bmatrix} 4 & 2 & 2 \\ 2 & 5 & 3 \\ 2 & 3 & 6 \end{bmatrix}$. We fill $L$ one column at a time. In each step we use only numbers we already know.
- $l_{11} = \sqrt{a_{11}} = \sqrt4 = 2$.
- $l_{21} = a_{21} / l_{11} = 2 / 2 = 1$.
- $l_{31} = a_{31} / l_{11} = 2 / 2 = 1$.
- $l_{22} = \sqrt{a_{22} - l_{21}^2} = \sqrt{5 - 1} = \sqrt4 = 2$.
- $l_{32} = (a_{32} - l_{31}l_{21}) / l_{22} = (3 - 1\cdot1) / 2 = 1$.
- $l_{33} = \sqrt{a_{33} - l_{31}^2 - l_{32}^2} = \sqrt{6 - 1 - 1} = \sqrt4 = 2$.
Check one entry of $LL^\top$: row 3 times row 3 of $L$ is $1\cdot1 + 1\cdot1 + 2\cdot2 = 6 = a_{33}$ ✓. And row 3 times row 2: $1\cdot1 + 1\cdot2 = 3 = a_{32}$ ✓.
A failing case. $A = \begin{bmatrix} 1 & 2 \\ 2 & 1 \end{bmatrix}$: $l_{11} = 1$, $l_{21} = 2$, and then $l_{22} = \sqrt{1 - 2^2} = \sqrt{-3}$. There is no real answer, so $A$ is not positive definite.
A matrix $A$ is symmetric positive definite (SPD) if $A = A^\top$ and $\mathbf{x}^\top A\mathbf{x} \gt 0$ for every non-zero vector $\mathbf{x}$ (Chapter 1.12). For such a matrix the Cholesky decomposition is
$$A = L L^\top,$$where $L$ is lower triangular with positive diagonal entries. It exists and is unique exactly when $A$ is SPD. Filling $L$ column by column ($j$ is the column, $i$ the row), the entries are
$$l_{jj} = \sqrt{\,a_{jj} - \sum_{k=1}^{j-1} l_{jk}^2\,}, \qquad\qquad l_{ij} = \frac{1}{l_{jj}}\Big(a_{ij} - \sum_{k=1}^{j-1} l_{ik}\,l_{jk}\Big) \quad (i \gt j).$$In words: the diagonal entry is the square root of "$a_{jj}$ minus the squares of the entries already placed to its left". An entry below the diagonal is "$a_{ij}$ minus the dot product of the already-known left parts of rows $i$ and $j$", divided by the diagonal entry above it. The sums have $j - 1$ terms, so the first column has no sums at all (as in the examples).
Cost: about $\tfrac13 n^3$ operations, which is half of LU. Because the algorithm only works when each number under a square root is positive, it is also a free test for positive definiteness.
Why do we need it?
Many problems hand us a "squared" object: a covariance matrix (spread of data), a kernel matrix, or $X^\top X$. To work with it we want its "un-squared" half. Cholesky gives that half in a shape that is cheap to use: triangular.
Where is it used?
Gaussian processes, Kalman filters, ridge and linear regression, Bayesian optimisation, VAEs and diffusion models with full covariance, and Newton-type optimisers. The next sections go through the main jobs.
How is it used?
Call L = np.linalg.cholesky(A) (it returns the lower-triangular $L$). If it raises LinAlgError, the matrix is not positive definite. Then use $L$ for triangular solves, sampling, or log-determinants.
- It must be both symmetric and positive definite. A symmetric matrix that is not positive definite (like $\begin{bmatrix} 1 & 2 \\ 2 & 1 \end{bmatrix}$) fails. A positive-definite-looking matrix that is not symmetric is not allowed either.
- $L$ is lower triangular and $L^\top$ is upper triangular. The two halves are mirror images, not the same matrix.
Quick check: what is the Cholesky factor of the diagonal matrix $\begin{bmatrix} 9 & 0 \\ 0 & 16 \end{bmatrix}$?
$l_{11} = \sqrt9 = 3$, $l_{21} = 0/3 = 0$, $l_{22} = \sqrt{16 - 0} = 4$. So $L = \begin{bmatrix} 3 & 0 \\ 0 & 4 \end{bmatrix}$. For a diagonal matrix, Cholesky is just "take the square root of each diagonal entry".
Build the Cholesky factor by hand, one entry at a time
Think of $L$ as a crossword puzzle that you fill in from the top-left, column by column. Every new box has a clue that only uses boxes you have already filled in: the matching number from $A$, and the numbers to the left in $L$. You never have to guess and you never go back.
At each box you do the same small job. For a box on the diagonal: take the matching number of $A$, subtract the squares of the filled boxes on its left, and take the square root. For a box below the diagonal: subtract the products of the filled boxes to the left, then divide by the diagonal box above.
Use the $3\times3$ matrix from before. Look at how the fifth box, $l_{32}$ (row 3, column 2), is built.
- Take $a_{32} = 3$ from $A$.
- The filled boxes to the left in rows 3 and 2 are $l_{31} = 1$ and $l_{21} = 1$. Multiply them: $1\cdot1 = 1$.
- Subtract: $3 - 1 = 2$.
- Divide by the diagonal box above, $l_{22} = 2$: $2/2 = 1$. So $l_{32} = 1$.
Every one of the six entries is built this way. The widget below shows all of them, with the numbers highlighted.
The algorithm (for an $n\times n$ SPD matrix), for $j = 1, 2, \dots, n$:
- Diagonal: $l_{jj} = \sqrt{a_{jj} - l_{j1}^2 - \dots - l_{j,j-1}^2}$. If the number under the root is $\le 0$, stop: $A$ is not positive definite.
- Below the diagonal (for each row $i \gt j$): $l_{ij} = \big(a_{ij} - l_{i1}l_{j1} - \dots - l_{i,j-1}l_{j,j-1}\big) / l_{jj}$.
Entry $l_{ij}$ in column $j$ needs about $j$ multiplications. There are about $n^2/2$ entries, and adding up all the work gives about $\tfrac16 n^3$ multiplications and as many additions, so roughly $\tfrac13n^3$ operations in total. We only ever read the lower triangle of $A$, because $A$ is symmetric.
Why do we need it?
Seeing the steps removes the mystery. It also shows where and why a matrix fails: the failure is always one specific square root of a number that is zero or negative.
Where is it used?
This exact loop runs inside np.linalg.cholesky, scipy.linalg.cho_factor, Cholesky solvers in every Gaussian-process library, and in the "jitter" safeguard code of probabilistic programming tools.
How is it used?
You almost never code it yourself. Call the library. Do code it once by hand (as here) so that error messages like "matrix is not positive definite" make sense: they mean one of these root steps went wrong.
- The order matters: you need $l_{11}$ before $l_{21}$, and the whole of column 1 before column 2. Do not try to fill boxes in a random order.
- A square root gives $\pm$ two answers, but Cholesky always takes the positive one. That is what makes $L$ unique.
- Adding jitter changes the matrix a little. It is a practical safeguard, not a true fix: if the matrix really is not positive definite, you should ask why.
Quick check: in the $3\times3$ example, which already-known entries of $L$ are used to compute $l_{33}$?
Those on the left of row 3: $l_{31} = 1$ and $l_{32} = 1$. So $l_{33} = \sqrt{a_{33} - l_{31}^2 - l_{32}^2} = \sqrt{6 - 1 - 1} = 2$.
Why Cholesky (1): solving $A\mathbf{x} = \mathbf{b}$ twice as fast
Suppose you have a locked box with two locks, one after the other. Opening one lock at a time is easy. Opening both at once is hard.
The system $A\mathbf{x} = \mathbf{b}$ with $A = LL^\top$ is like that. We write it as $L(L^\top\mathbf{x}) = \mathbf{b}$, give the middle part a name, $\mathbf{y} = L^\top\mathbf{x}$, and open the two locks separately:
- Lock 1: $L\mathbf{y} = \mathbf{b}$ is triangular, so find $\mathbf{y}$ from the top row down (forward substitution).
- Lock 2: $L^\top\mathbf{x} = \mathbf{y}$ is triangular too, so find $\mathbf{x}$ from the bottom row up (backward substitution).
Each lock needs only a short pass over the numbers. And because $A$ is symmetric, the factoring step itself costs half as much as LU.
$A = \begin{bmatrix} 4 & 2 \\ 2 & 5 \end{bmatrix}$ with $L = \begin{bmatrix} 2 & 0 \\ 1 & 2 \end{bmatrix}$, and $\mathbf{b} = [6, 7]$.
- Forward ($L\mathbf{y} = \mathbf{b}$): $2y_1 = 6$, so $y_1 = 3$. Then $1\cdot y_1 + 2y_2 = 7$, so $2y_2 = 7 - 3 = 4$ and $y_2 = 2$.
- Backward ($L^\top\mathbf{x} = \mathbf{y}$, with $L^\top = \begin{bmatrix} 2 & 1 \\ 0 & 2 \end{bmatrix}$): $2x_2 = 2$, so $x_2 = 1$. Then $2x_1 + 1\cdot x_2 = 3$, so $2x_1 = 2$ and $x_1 = 1$.
Answer: $\mathbf{x} = [1, 1]$. Check in the original: $4\cdot1 + 2\cdot1 = 6$ ✓ and $2\cdot1 + 5\cdot1 = 7$ ✓. The middle vector $\mathbf{y} = [3, 2]$ is only a stepping stone.
For SPD $A = LL^\top$:
$$A\mathbf{x} = \mathbf{b} \;\Longleftrightarrow\; L\mathbf{y} = \mathbf{b} \ \text{ (forward)}, \quad L^\top\mathbf{x} = \mathbf{y} \ \text{ (backward)}.$$- Cost to factor: $\approx \tfrac13 n^3$ operations (LU: $\approx \tfrac23 n^3$). So Cholesky is about 2× faster than LU, and it needs about half the memory (only one triangle).
- Cost per solve: $\approx 2n^2$ operations for the two triangular passes. With many right-hand sides, factor once and reuse $L$.
- No pivoting needed: for SPD matrices Cholesky is numerically stable as it is.
Why do we need it?
Solving $A\mathbf{x} = \mathbf{b}$ is the most common job in numerical computing, and it can be slow for big $A$. When $A$ is symmetric positive definite we can do it in half the time and half the memory.
Where is it used?
Ridge and linear regression (the normal equations), Gaussian-process prediction, Kalman filters, Newton steps in optimisation, finite-element simulations, and anywhere $A$ is a covariance or $X^\top X$.
How is it used?
Factor once: c, low = scipy.linalg.cho_factor(A). Then for every right-hand side: x = scipy.linalg.cho_solve((c, low), b). Never write np.linalg.inv(A) @ b.
- Do not compute $A^{-1}$ and multiply. It costs about three times more and is less accurate. "Solve" and "invert" are different jobs.
- Remember the order of the passes: forward with $L$ first, then backward with $L^\top$. Swapping them gives a wrong answer.
Quick check: with $L = \begin{bmatrix} 3 & 0 \\ 0 & 2 \end{bmatrix}$ and $\mathbf{b} = [6, 4]$, what are $\mathbf{y}$ and $\mathbf{x}$?
Forward: $3y_1 = 6$, $y_1 = 2$; $2y_2 = 4$, $y_2 = 2$. Backward ($L^\top = L$ here): $3x_1 = 2$, $x_1 = 2/3$; $2x_2 = 2$, $x_2 = 1$. So $\mathbf{x} = [2/3, 1]$.
Why Cholesky (2): making correlated random numbers, $\mathbf{x} = \boldsymbol\mu + L\mathbf{z}$
Computers are great at making independent random numbers: each one ignores all the others. Plot two of them against each other and you get a round, shapeless cloud.
Real data is not like that. A person's height and weight go together: tall people tend to weigh more. Daily returns of two shares move together. Pixels next to each other have similar brightness. We want random numbers with a chosen pattern of togetherness, which is exactly what a covariance matrix $\Sigma$ describes. Its diagonal holds how much each quantity varies on its own (the variance), and the other entries say how much each pair varies together.
Here is the trick. Start from the round cloud $\mathbf{z}$ and multiply it by $L$, the Cholesky factor of $\Sigma$. $L$ stretches and tilts the round cloud into exactly the right oval.
There is a one-dimensional version you already know. To get a normal number with spread $\sigma$ you compute $x = \mu + \sigma z$. The standard deviation $\sigma$ is the square root of the variance $\sigma^2$. In many dimensions, $L$ is the "standard deviation matrix": the square root of the covariance.
$\Sigma = \begin{bmatrix} 4 & 2 \\ 2 & 5 \end{bmatrix}$ has $L = \begin{bmatrix} 2 & 0 \\ 1 & 2 \end{bmatrix}$. Let the mean be $\boldsymbol\mu = [10, 20]$ and let the independent noise be $\mathbf{z} = [1, -1]$.
- Multiply: $L\mathbf{z} = [2\cdot1 + 0\cdot(-1),\; 1\cdot1 + 2\cdot(-1)] = [2, -1]$.
- Shift by the mean: $\mathbf{x} = \boldsymbol\mu + L\mathbf{z} = [12, 19]$.
Do this for thousands of fresh $\mathbf{z}$ vectors and the cloud of $\mathbf{x}$ points has mean close to $[10,20]$ and covariance close to $\Sigma$. (With more and more points, "close" becomes "equal".)
Why the covariance comes out right. $\mathbf{z}$ has covariance $I$ (independent, variance 1). Multiplying by $L$ changes a covariance $C$ into $LCL^\top$. So $\operatorname{cov}(L\mathbf{z}) = L\,I\,L^\top = LL^\top = \Sigma$. ✓
To draw $\mathbf{x} \sim \mathcal{N}(\boldsymbol\mu, \Sigma)$:
- Factor $\Sigma = LL^\top$ once (Cholesky).
- Draw $\mathbf{z}$ with independent standard normal entries, $\mathbf{z} \sim \mathcal{N}(\mathbf{0}, I)$.
- Set $\mathbf{x} = \boldsymbol\mu + L\mathbf{z}$.
This is called the reparameterisation trick when used in a neural network: all the randomness sits in $\mathbf{z}$, and $\boldsymbol\mu$ and $L$ enter through a smooth formula, so gradients can flow back through them during training.
Why do we need it?
We often need random data with a specific pattern of correlation (to simulate, test, or train), but the computer's random-number generator only gives independent numbers. $L$ converts one into the other.
Where is it used?
Variational autoencoders (the reparameterisation trick), sampling from Gaussian processes, Monte Carlo risk simulation of correlated assets, Bayesian posteriors, and making synthetic correlated features for tests.
How is it used?
Compute L = np.linalg.cholesky(Sigma) once. Draw Z = rng.standard_normal((n, N)). Then X = mu[:, None] + L @ Z gives $N$ samples as columns.
- Use the standard normal $\mathbf{z}$ (mean 0, variance 1, independent). If $\mathbf{z}$ is not standard, the covariance comes out wrong.
- The mean $\boldsymbol\mu$ is added after multiplying by $L$.
- Covariances you pick by hand must be consistent. Choosing correlations that cannot coexist gives a matrix that is not positive definite, and Cholesky fails (as in the 3D widget).
- For a singular covariance (a perfectly flat cloud), plain Cholesky fails. Use the eigendecomposition or the SVD, or add a tiny jitter.
Quick check: if $\Sigma = \begin{bmatrix} 9 & 0 \\ 0 & 4 \end{bmatrix}$ (no correlation), what does $\boldsymbol\mu + L\mathbf{z}$ do to $\mathbf{z} = [1, 1]$ with $\boldsymbol\mu = \mathbf{0}$?
$L = \begin{bmatrix} 3 & 0 \\ 0 & 2 \end{bmatrix}$ (the standard deviations), so $L\mathbf{z} = [3, 2]$. Each coordinate is just scaled by its own standard deviation.
Why Cholesky (3): the log-determinant for free
The determinant of $A$ tells you how much $A$ scales areas and volumes (Chapter 1.7). Gaussian formulas need it: the bell curve in many dimensions has $\det\Sigma$ inside it.
For a triangular matrix the determinant is easy: multiply the diagonal. Our $A = LL^\top$ is made of two triangular matrices, and they have the same diagonal. So $\det A$ is the product of the diagonal of $L$, squared.
One more idea. For big matrices a determinant is a gigantic or a tiny number, like $10^{-400}$, and the computer rounds that to $0$ (or to infinity). Taking the logarithm turns the long product into a short, safe sum. So we always work with the log-determinant.
$A = \begin{bmatrix} 4 & 2 \\ 2 & 5 \end{bmatrix}$, $L = \begin{bmatrix} 2 & 0 \\ 1 & 2 \end{bmatrix}$.
- Directly: $\det A = 4\cdot5 - 2\cdot2 = 16$, so $\log\det A = \ln 16 \approx 2.77$.
- From $L$: diagonal $2, 2$. $\det A = (2\cdot2)^2 = 16$ ✓. And $2(\ln2 + \ln2) = 4\ln2 \approx 2.77$ ✓.
For the $3\times3$ matrix from before, $L$ has diagonal $2, 2, 2$, so $\det A = (2\cdot2\cdot2)^2 = 64$ and $\log\det A = 2\cdot3\ln2 = 6\ln2 \approx 4.16$. (Direct check: $4(30-9) - 2(12-6) + 2(6-10) = 84 - 12 - 8 = 64$ ✓.)
Because $\det(LL^\top) = \det L\cdot\det L^\top = (\det L)^2$ and the determinant of a triangular matrix is the product of its diagonal,
$$\det A = \Big(\prod_{i=1}^n l_{ii}\Big)^2, \qquad \log\det A = 2\sum_{i=1}^n \log l_{ii}.$$The log-determinant costs nothing extra once you have $L$. It is also safe: each $\log l_{ii}$ is an ordinary-sized number, even when $\det A$ would overflow or underflow.
Why do we need it?
The Gaussian bell curve in many dimensions contains $\det\Sigma$. Computing a huge determinant directly overflows or underflows, but its log is a normal-sized number, and Cholesky delivers it for free.
Where is it used?
Gaussian log-likelihoods, the marginal likelihood of a Gaussian process (used to learn kernel settings), Gaussian mixture models, Kalman-filter likelihoods, and Bayesian model comparison.
How is it used?
L = np.linalg.cholesky(A), then logdet = 2 * np.sum(np.log(np.diag(L))). Compare: np.linalg.slogdet(A) gives sign and log-determinant for any matrix.
- $\log\det A = 2\sum\log l_{ii}$ has the 2. Forgetting it is the classic slip. You need it because $\det A = (\det L)^2$.
- The determinant is positive for a positive definite matrix, so the logarithm is defined.
Quick check: $L$ has diagonal $1, 3, 3$. What is $\det A$ for $A = LL^\top$?
$\det A = (1\cdot3\cdot3)^2 = 81$. (And $\log\det A = 2(\ln1 + \ln3 + \ln3) = 4\ln3 \approx 4.39$.)
Why Cholesky (4): the positive-definiteness test
Sometimes the question is not "what is $L$?" but simply "is this matrix positive definite at all?" You could compute all its eigenvalues and check that each one is positive. That works, but it is slow.
There is a cheaper way: just try to take the matrix's square root. If every square root along the way is of a positive number, it worked, and the matrix is positive definite. The first time you need the square root of zero or a negative number, you stop: the answer is "no". Cholesky is a built-in lie detector.
- $\begin{bmatrix} 4 & 2 \\ 2 & 5 \end{bmatrix}$: roots of $4$ and then of $5 - 1 = 4$. Both positive. ✓ Positive definite.
- $\begin{bmatrix} 1 & 2 \\ 2 & 1 \end{bmatrix}$: root of $1$, then of $1 - 2^2 = -3$. ✗ Not positive definite. (Its eigenvalues are $3$ and $-1$.)
- $\begin{bmatrix} 1 & 1 \\ 1 & 1 \end{bmatrix}$: root of $1$, then of $1 - 1 = 0$. ✗ Not strictly positive: it is only positive semidefinite (eigenvalues $2$ and $0$).
PD test by Cholesky. For a symmetric matrix $A$, attempt the Cholesky algorithm. Then:
- Every number under a square root is $\gt 0$ for all steps $\Longleftrightarrow$ $A$ is positive definite.
- Some number is $\le 0$ $\Longleftrightarrow$ $A$ is not positive definite. (A $0$ means "borderline": positive semidefinite at best.)
The cost is only $\tfrac13 n^3$, and you get the factor $L$ as a bonus if it works. Check symmetry first (it costs almost nothing): the test is only valid for symmetric matrices.
Why do we need it?
Many methods are only valid if a matrix is positive definite: a covariance must be, a kernel must be, and a Hessian must be for a safe minimum. We need a fast yes-or-no check before we trust the result.
Where is it used?
Validating covariance and kernel matrices in Gaussian processes, checking that a Hessian describes a true minimum, trust-region and Levenberg–Marquardt optimisers, and convex optimisation solvers.
How is it used?
Wrap np.linalg.cholesky(A) in try / except np.linalg.LinAlgError. If it raises, the matrix is not positive definite. Decide what to do: add jitter, fix the data, or use a different method.
- Check symmetry first. Cholesky only reads one triangle, so it will happily "pass" a non-symmetric matrix and give you the answer for a different, symmetrised one.
- Because of rounding, a matrix that is positive definite on paper but nearly singular can fail on a computer (a tiny positive number becomes slightly negative). See the "jitter" trick below.
- Checking only the entries (all positive) or only the determinant is not a valid test.
Quick check: does $\begin{bmatrix} 2 & 3 \\ 3 & 2 \end{bmatrix}$ pass the Cholesky test?
$l_{11} = \sqrt2$, $l_{21} = 3/\sqrt2$, and then $l_{22}^2 = 2 - 9/2 = -2.5 \lt 0$. It fails: not positive definite (eigenvalues $5$ and $-1$).
Why Cholesky (5): whitening and distance that respects correlation
Sampling used $L$ to turn a round cloud into a correlated one. Run it backwards and you get the reverse: $L^{-1}$ turns a correlated cloud back into a round one. This is called whitening (the data becomes "white noise": uncorrelated, equal spread).
Why would you want that? Imagine a long, thin oval cloud of data. A point $2$ units away from the centre along the long direction is perfectly normal. A point $2$ units away across the thin direction is a shocking outlier. The ordinary (Euclidean) distance says both are equally far. After whitening, the oval is a circle, and plain distance now tells the truth: "how many standard deviations away". That distance is called the Mahalanobis distance.
Take $\Sigma = \begin{bmatrix} 4 & 2 \\ 2 & 5 \end{bmatrix}$, $L = \begin{bmatrix} 2 & 0 \\ 1 & 2 \end{bmatrix}$, and mean $\boldsymbol\mu = \mathbf{0}$. Consider the point $\mathbf{x} = [2, 1]$.
- Solve $L\mathbf{z} = \mathbf{x}$ by forward substitution: $2z_1 = 2$ gives $z_1 = 1$; then $1\cdot1 + 2z_2 = 1$ gives $z_2 = 0$. So $\mathbf{z} = [1, 0]$.
- The Mahalanobis distance is $\|\mathbf{z}\| = \sqrt{1^2 + 0^2} = 1$.
- The ordinary distance from the centre is $\|\mathbf{x}\| = \sqrt{4+1} \approx 2.24$.
So this point is only 1 standard deviation from the centre in this data's own units, even though it looks far in plain units. (We used forward substitution, so we never had to compute $L^{-1}$ at all.)
For data with mean $\boldsymbol\mu$ and covariance $\Sigma = LL^\top$, the whitened vector and the Mahalanobis distance are
$$\mathbf{z} = L^{-1}(\mathbf{x} - \boldsymbol\mu), \qquad d_M(\mathbf{x}) = \sqrt{(\mathbf{x}-\boldsymbol\mu)^\top\Sigma^{-1}(\mathbf{x}-\boldsymbol\mu)} = \|\mathbf{z}\|.$$The two forms agree because $\Sigma^{-1} = (L^\top)^{-1}L^{-1}$, so $(\mathbf{x}-\boldsymbol\mu)^\top\Sigma^{-1}(\mathbf{x}-\boldsymbol\mu) = \|L^{-1}(\mathbf{x}-\boldsymbol\mu)\|^2$. In code, get $\mathbf{z}$ by solving $L\mathbf{z} = \mathbf{x}-\boldsymbol\mu$ with a triangular solver. If $d_M = 1$, the point is one standard deviation away. The set of points with $d_M = c$ is an ellipse (or ellipsoid), and after whitening it becomes a circle (sphere).
Why do we need it?
Plain distance ignores that features have different scales and move together. We need a distance in the data's own units, to tell real outliers from ordinary points and to compare features fairly.
Where is it used?
Outlier and anomaly detection, Gaussian discriminant analysis and Gaussian mixture models, whitening as a preprocessing step (ICA, speech and image models), and gating in Kalman-filter tracking.
How is it used?
z = scipy.linalg.solve_triangular(L, x - mu, lower=True), then d = np.linalg.norm(z). For a whole dataset, solve once with all points as columns.
- The Mahalanobis distance needs a covariance that is positive definite (so $L$ exists). With a singular covariance, use a pseudoinverse or add jitter.
- Never form $\Sigma^{-1}$ explicitly for this. Solve with $L$. It is cheaper and more accurate.
- "Whitening" can mean slightly different things (Cholesky, or via eigenvectors). They all make the data uncorrelated with unit variance, but they rotate the result differently.
Quick check: in 1D, what are $L$ and the Mahalanobis distance of a point $x$ from mean $\mu$ when the variance is $\sigma^2$?
$L = \sigma$ (the standard deviation) and $z = (x-\mu)/\sigma$, so $d_M = |x-\mu|/\sigma$: the familiar "z-score".
Where Cholesky is used, the jitter trick, and common mistakes
Every use of Cholesky has the same shape: there is a covariance-like matrix, and we need to treat it as a "square" of something simpler. Solve with it, sample from it, take its determinant, check it, measure distance with it.
One practical story is very common. In theory a kernel matrix or covariance matrix is positive definite. On a computer, when the data points are very close together, the matrix becomes almost singular: its smallest eigenvalue is a tiny positive number, rounding makes it slightly negative, and Cholesky crashes. The standard cure is to add a very small number to the diagonal: the jitter trick, $A + \varepsilon I$.
Ridge regression. Fit weights $\mathbf{w}$ to data $X$ and targets $\mathbf{y}$ by solving $(X^\top X + \lambda I)\mathbf{w} = X^\top\mathbf{y}$. Take $X = \begin{bmatrix} 1 & 0 \\ 0 & 1 \\ 1 & 1 \end{bmatrix}$, $\mathbf{y} = [1,2,3]$, $\lambda = 1$.
- $X^\top X = \begin{bmatrix} 2 & 1 \\ 1 & 2 \end{bmatrix}$, so $A = X^\top X + I = \begin{bmatrix} 3 & 1 \\ 1 & 3 \end{bmatrix}$. This is symmetric, and positive definite for any $\lambda \gt 0$.
- $X^\top\mathbf{y} = [1 + 3,\; 2 + 3] = [4, 5]$.
- Cholesky: $l_{11} = \sqrt3 \approx 1.732$, $l_{21} = 1/\sqrt3 \approx 0.577$, $l_{22} = \sqrt{3 - 1/3} \approx 1.633$. Then solve the two triangular systems.
- The answer is $\mathbf{w} = [0.875,\, 1.375]$. (Check: $3\cdot0.875 + 1.375 = 4$ ✓ and $0.875 + 3\cdot1.375 = 5$ ✓.)
Adding $\lambda I$ is itself a "jitter": it guarantees the matrix is positive definite. That is one more reason ridge regression is so well behaved.
The jitter trick. If Cholesky fails on a matrix that should be positive definite, replace $A$ by $A + \varepsilon I$ with a tiny $\varepsilon$ (often $10^{-6}$ to $10^{-10}$ times the typical size of the diagonal). This lifts every eigenvalue by $\varepsilon$, which makes them all positive, while changing the answer by almost nothing.
Cost summary. Factoring an $n\times n$ SPD matrix: Cholesky $\approx \tfrac13n^3$, LU $\approx \tfrac23n^3$, QR (Householder) $\approx \tfrac43n^3$, SVD more again. Each later solve: $\approx 2n^2$.
Which decomposition for which job?
| Job | Best tool | Why |
|---|---|---|
| Solve $A\mathbf{x} = \mathbf{b}$, $A$ symmetric positive definite | Cholesky | Fastest, stable, half the memory |
| Solve $A\mathbf{x} = \mathbf{b}$, general square $A$ | LU with pivoting | Works for any invertible matrix |
| Least squares, tall $A$ | QR (or SVD) | Does not square the condition number |
| Sample $\mathcal{N}(\boldsymbol\mu,\Sigma)$, log-determinant, whitening | Cholesky | $L$ is exactly the "square root" needed |
| Singular or nearly singular covariance | Eigendecomposition or SVD (or Cholesky with jitter) | Plain Cholesky needs strictly positive definite |
| Rank, pseudoinverse, compression | SVD | Works for every matrix |
Why do we need it?
Knowing the main uses lets you recognise the pattern "SPD matrix plus a solve, a sample, a determinant or a distance" and reach for Cholesky at once. The jitter trick keeps it from crashing on nearly singular matrices.
Where is it used?
Ridge regression and the normal equations, Gaussian processes and Bayesian optimisation, Kalman filters (and "square-root" filters), VAEs and diffusion models with full covariance, second-order optimisers such as Newton's method and K-FAC, and Monte Carlo sampling.
How is it used?
Factor once (np.linalg.cholesky or scipy.linalg.cho_factor). Reuse $L$ for solves, samples, $\log\det$ and whitening. If it raises an error, add jitter such as A + 1e-6 * np.eye(n) and try again.
Common mistakes
- The matrix must be symmetric AND positive definite. Symmetric alone is not enough; positive alone is not enough. Check symmetry first, and treat a failure as information ("not PD"), not as a bug.
- Never invert, always solve. $A^{-1}\mathbf{b}$ should be written as two triangular solves (
cho_solve), and $(\mathbf{x}-\boldsymbol\mu)^\top\Sigma^{-1}(\mathbf{x}-\boldsymbol\mu)$ as $\|L^{-1}(\mathbf{x}-\boldsymbol\mu)\|^2$ with a triangular solve. - $L$ versus $L^\top$ conventions. Maths writes $A = LL^\top$ (lower). NumPy's
np.linalg.choleskyreturns the lower $L$. Butscipy.linalg.choleskyreturns the upper factor $U$ with $A = U^\top U$ unless you passlower=True. Mixing them up gives wrong samples and wrong distances. - Do not use a huge jitter to "make it work": it changes the model. Start with $10^{-10}$ and increase only as far as needed.
Quick check: a Gaussian-process code crashes with "matrix is not positive definite". What is the first thing to try?
Add a small jitter to the diagonal of the kernel matrix, for example K + 1e-6 * np.eye(n). The matrix is nearly singular (points very close together), and rounding pushed a tiny eigenvalue below zero. If it still fails, also check that the kernel is symmetric and that duplicate data points are not present.
QR decomposition core
Imagine the columns of $A$ as a set of crooked, overlapping arrows. Gram–Schmidt (Chapter 1.9) straightens them: keep the first arrow's direction, then take the second arrow and remove the part that lies along the first, then the third arrow minus what lies along the first two, and so on. Scale each result to length 1.
You end up with tidy, perpendicular, unit-length arrows: the columns of $Q$. The matrix $R$ is the recipe that says how to rebuild each original column from the tidy ones. Because column 1 only uses tidy arrow 1, column 2 only uses tidy arrows 1 and 2, and so on, $R$ is upper triangular.
Let $A = \begin{bmatrix} 3 & 1 \\ 4 & 3 \end{bmatrix}$, with columns $\mathbf{a}_1 = [3,4]$ and $\mathbf{a}_2 = [1,3]$.
- Length of $\mathbf{a}_1$: $r_{11} = \sqrt{9 + 16} = 5$. First tidy arrow: $\mathbf{q}_1 = \mathbf{a}_1 / 5 = [0.6,\, 0.8]$.
- How much of $\mathbf{a}_2$ lies along $\mathbf{q}_1$? $r_{12} = \mathbf{q}_1\cdot\mathbf{a}_2 = 0.6\cdot1 + 0.8\cdot3 = 3$.
- Remove that part: $\mathbf{w} = \mathbf{a}_2 - 3\mathbf{q}_1 = [1,3] - [1.8,\,2.4] = [-0.8,\, 0.6]$.
- Its length is $r_{22} = \sqrt{0.64 + 0.36} = 1$, so $\mathbf{q}_2 = [-0.8,\, 0.6]$.
Check the first row: $0.6\cdot5 = 3$ and $0.6\cdot3 + (-0.8)\cdot1 = 1$. ✓ Notice $\mathbf{q}_1\cdot\mathbf{q}_2 = -0.48 + 0.48 = 0$.
Every $m\times n$ matrix $A$ with independent columns ($m \ge n$) can be written as
$$A = QR,$$where $Q$ has orthonormal columns ($Q^\top Q = I$) and $R$ is upper triangular. There are two versions:
- Reduced (thin) QR: $Q$ is $m\times n$ and $R$ is $n\times n$. This is the one you usually want.
- Full QR: $Q$ is a square $m\times m$ orthogonal matrix (we add $m-n$ extra perpendicular columns) and $R$ is $m \times n$ with rows of zeros at the bottom: $$A = \begin{bmatrix} Q_1 & Q_2 \end{bmatrix}\begin{bmatrix} R_1 \\ 0 \end{bmatrix} = Q_1R_1.$$ The extra columns $Q_2$ span the space perpendicular to the columns of $A$.
Use 1: stable least squares. To minimise $\|A\mathbf{x} - \mathbf{b}\|$, note that multiplying by an orthogonal matrix does not change lengths, so we may multiply both sides of the problem by $Q^\top$. The best $\mathbf{x}$ then solves $R\mathbf{x} = Q^\top\mathbf{b}$ (reduced $Q$ and $R$), which back substitution does quickly. This avoids forming $A^\top A$, which squares the condition number (the number that says how much small errors in the data get magnified in the answer; see Chapter 1.7 and the widget below).
Use 2: the QR eigenvalue algorithm. Repeat "factor $A_k = Q_kR_k$, then multiply in the reverse order $A_{k+1} = R_kQ_k$". The matrices $A_k$ slowly turn into a triangular matrix with the eigenvalues on its diagonal. This is the basis of how eigenvalues are computed numerically (Chapter 1.11).
Why do we need it?
Least squares through the normal equations squares the condition number and can lose all accuracy. QR solves the same problem safely, and it also builds orthonormal bases.
Where is it used?
Regression solvers (numpy.linalg.lstsq style), orthogonal weight initialisation in neural networks, Gram–Schmidt orthogonalisation, and the QR algorithm that computes eigenvalues.
How is it used?
Q, R = np.linalg.qr(A) gives the reduced factors. For least squares solve R x = Q.T @ b by back substitution. Use mode="complete" only if you need the extra perpendicular columns.
Three ways to build $Q$ (awareness). You do not need to code these by heart. Just know what they are and which one libraries pick.
| Method | Idea | Good for |
|---|---|---|
| Gram–Schmidt (classical / modified) | Subtract projections, one column at a time | Teaching; modified GS is more accurate than classical |
| Householder reflections | Reflect a whole column onto an axis in one step, repeat for each column | The default in libraries (np.linalg.qr). Very accurate |
| Givens rotations | Rotate to zero out one entry at a time | Sparse or nearly-triangular matrices, updating a factorisation |
- Reduced or full? Reduced $Q$ is tall and thin, with only $Q^\top Q = I$. Full $Q$ is square and orthogonal, so both $Q^\top Q = I$ and $QQ^\top = I$. For reduced $Q$, $QQ^\top$ is a projection matrix, not the identity.
- Signs are a choice. You can flip a column of $Q$ and the matching row of $R$ and still have a valid factorisation. Libraries often return negative diagonal entries in $R$. If you insist on a positive diagonal, the QR factorisation is unique.
- If the columns of $A$ are dependent, $R$ has a zero on its diagonal and the plain "solve $R\mathbf{x} = Q^\top\mathbf{b}$" breaks. Use the SVD or the pseudoinverse (Chapter 1.10) instead.
Quick check: why does $Q^\top Q = I$ turn $\|QR\mathbf{x} - \mathbf{b}\|$ into a simple triangular problem?
Multiplying by $Q^\top$ keeps lengths, so the distance is the same. The normal equations $A^\top A\mathbf{x} = A^\top\mathbf{b}$ become $R^\top Q^\top Q R\mathbf{x} = R^\top Q^\top\mathbf{b}$, which is $R^\top R\mathbf{x} = R^\top Q^\top\mathbf{b}$. Cancel $R^\top$ (it is invertible) to get $R\mathbf{x} = Q^\top\mathbf{b}$, and $R$ is triangular.
Eigendecomposition core
In Chapter 1.11 you met eigenvectors: special directions that a matrix only stretches (or flips) and never turns. If you describe space using these special directions as your axes, the matrix becomes very simple: it just scales each axis by its eigenvalue. A diagonal matrix.
Eigendecomposition says this in three moves. (1) Switch to the eigen-axes. (2) Stretch each axis by its eigenvalue. (3) Switch back to the normal axes. It is like translating a sentence into a language where the grammar is easy, doing the easy job there, and translating back.
$A = \begin{bmatrix} 4 & 1 \\ 2 & 3 \end{bmatrix}$ has eigenvalue $5$ with eigenvector $[1,1]$ (check: $A[1,1] = [5,5]$) and eigenvalue $2$ with eigenvector $[1,-2]$ (check: $A[1,-2] = [2,-4]$). Put the eigenvectors in the columns of $P$ and the eigenvalues on the diagonal of $D$:
$$A = PDP^{-1}: \quad \begin{bmatrix} 4 & 1 \\ 2 & 3 \end{bmatrix} = \begin{bmatrix} 1 & 1 \\ 1 & -2 \end{bmatrix}\begin{bmatrix} 5 & 0 \\ 0 & 2 \end{bmatrix}\begin{bmatrix} 2/3 & 1/3 \\ 1/3 & -1/3 \end{bmatrix}$$Powers become easy. The middle factors cancel in pairs, $A^2 = PD\,P^{-1}P\,DP^{-1} = PD^2P^{-1}$, and $D^2$ just squares the diagonal:
$$A^2 = \begin{bmatrix} 1 & 1 \\ 1 & -2 \end{bmatrix}\begin{bmatrix} 25 & 0 \\ 0 & 4 \end{bmatrix}\begin{bmatrix} 2/3 & 1/3 \\ 1/3 & -1/3 \end{bmatrix} = \begin{bmatrix} 18 & 7 \\ 14 & 11 \end{bmatrix}.$$(Direct check: $A\cdot A$ has first row $[4\cdot4 + 1\cdot2,\; 4\cdot1 + 1\cdot3] = [18, 7]$. ✓) For $A^{100}$ you only compute $5^{100}$ and $2^{100}$.
A square $n\times n$ matrix $A$ is diagonalisable if it has $n$ linearly independent eigenvectors. Then
$$A = PDP^{-1}, \qquad P = \begin{bmatrix} \mathbf{p}_1 & \cdots & \mathbf{p}_n \end{bmatrix}, \quad D = \begin{bmatrix} \lambda_1 & & \\ & \ddots & \\ & & \lambda_n \end{bmatrix},$$where $A\mathbf{p}_i = \lambda_i\mathbf{p}_i$. Consequently $A^k = PD^kP^{-1}$.
Symmetric matrices are the best case (spectral theorem). If $A = A^\top$, then the eigenvalues are real, the eigenvectors can be chosen perpendicular with length 1, and so $P$ can be an orthogonal matrix $Q$ (with $Q^{-1} = Q^\top$):
$$A = Q\Lambda Q^\top.$$This always exists for a real symmetric matrix, with no "if". No matrix inverse is needed, only a transpose.
Limitations (these are exactly why we need the SVD):
- Only square matrices have eigenvalues. A $3\times2$ matrix has none.
- Even a square matrix may not be diagonalisable. A shear such as $\begin{bmatrix} 1 & 1 \\ 0 & 1 \end{bmatrix}$ has only one eigen-direction. A rotation has no real eigenvectors at all.
- For non-symmetric matrices, the eigenvectors are usually not perpendicular, so $P$ can be badly behaved.
Why do we need it?
Some questions ask what a square matrix does again and again (its powers) or which directions it only stretches. In eigen-axes the matrix is just a diagonal scaling, so those questions become easy.
Where is it used?
PCA on a covariance matrix, Markov chains and PageRank (long-run behaviour), stability of recurrent networks and dynamical systems, and spectral clustering with a graph Laplacian.
How is it used?
For a symmetric matrix call np.linalg.eigh(A) (real eigenvalues, orthogonal Q). For a general square matrix call np.linalg.eig(A). Then A^k = P D^k P⁻¹. First check P is invertible and the eigenvalues are real.
- $A = PDP^{-1}$ is not the same as $A = U\Sigma V^\top$ (SVD) even though both have a diagonal middle. In eigendecomposition the two outer factors are inverses of each other ($P$ and $P^{-1}$). In the SVD they are two different orthogonal matrices.
- "Diagonalisable" is about having enough independent eigenvectors, not about whether the eigenvalues are distinct. Distinct eigenvalues guarantee it, but repeated ones may or may not break it (the identity matrix repeats $1$ and is perfectly diagonalisable).
Quick check: can you eigendecompose a $3\times2$ matrix? Why or why not?
No. An eigenvector must satisfy $A\mathbf{v} = \lambda\mathbf{v}$, which compares $A\mathbf{v}$ (a vector in $\mathbb{R}^3$) with $\mathbf{v}$ (in $\mathbb{R}^2$). They live in different spaces, so "same direction" has no meaning. This is the gap the SVD fills.
The SVD: rotate, stretch, rotate core
Eigendecomposition is picky: it needs a square matrix with enough eigenvectors. The Singular Value Decomposition (SVD) is the opposite. It works for every matrix: tall, wide, square, singular, anything.
Here is the picture. Take a perfectly round ball of dough (a circle in 2D, a sphere in 3D) and push it through the matrix. It always comes out as a squashed ball: an ellipse (an ellipsoid in 3D). No matter how wild the matrix looks, the output is just a stretched, tilted oval.
An oval has a long axis and a short axis. So any matrix does three simple things in a row:
- Rotate the plane so that the special input directions line up with the coordinate axes.
- Stretch along each coordinate axis by a different amount (the singular values).
- Rotate again into the final tilted position.
Rotate, stretch, rotate. That is the whole SVD. Every matrix is "two rotations around a stretch".
Take $A = \begin{bmatrix} 1 & -1 \\ 2 & 2 \end{bmatrix}$. We find the three pieces in four steps.
- Form $A^\top A = \begin{bmatrix} 1 & 2 \\ -1 & 2 \end{bmatrix}\begin{bmatrix} 1 & -1 \\ 2 & 2 \end{bmatrix} = \begin{bmatrix} 5 & 3 \\ 3 & 5 \end{bmatrix}$.
- Its eigenvalues are $8$ and $2$ (with eigenvectors $[1,1]$ and $[-1,1]$). The singular values are the square roots: $\sigma_1 = \sqrt8 = 2\sqrt2 \approx 2.83$ and $\sigma_2 = \sqrt2 \approx 1.41$.
- The unit eigenvectors are the right singular vectors: $\mathbf{v}_1 = \tfrac{1}{\sqrt2}[1,1]$ and $\mathbf{v}_2 = \tfrac{1}{\sqrt2}[-1,1]$.
- The left singular vectors come from $\mathbf{u}_i = A\mathbf{v}_i/\sigma_i$. First, $A\mathbf{v}_1 = \tfrac{1}{\sqrt2}[1-1,\; 2+2] = \tfrac{1}{\sqrt2}[0, 4]$, and dividing by $2\sqrt2$ gives $\mathbf{u}_1 = [0, 1]$. Next, $A\mathbf{v}_2 = \tfrac{1}{\sqrt2}[-1-1,\; -2+2] = \tfrac{1}{\sqrt2}[-2, 0]$, and dividing by $\sqrt2$ gives $\mathbf{u}_2 = [-1, 0]$.
Read it right to left. $V^\top$ rotates the plane by $45^\circ$ clockwise. $\Sigma$ stretches $x$ by $2.83$ and $y$ by $1.41$. $U$ rotates by $90^\circ$ counter-clockwise. Check: $\Sigma V^\top = \begin{bmatrix} 2 & 2 \\ -1 & 1 \end{bmatrix}$, and $U$ times that gives $\begin{bmatrix} 1 & -1 \\ 2 & 2 \end{bmatrix} = A$. ✓
(Signs are a choice: you can flip $\mathbf{u}_i$ and $\mathbf{v}_i$ together. Software may give a different but equally correct set of signs.)
Every real $m\times n$ matrix $A$ can be written as
$$A = U\,\Sigma\,V^\top$$- $V$ is an $n\times n$ orthogonal matrix. Its columns $\mathbf{v}_1,\dots,\mathbf{v}_n$ are the right singular vectors (directions in the input space).
- $U$ is an $m\times m$ orthogonal matrix. Its columns $\mathbf{u}_1,\dots,\mathbf{u}_m$ are the left singular vectors (directions in the output space).
- $\Sigma$ is an $m\times n$ diagonal matrix (non-zero only on the main diagonal) holding the singular values $\sigma_1 \ge \sigma_2 \ge \dots \ge \sigma_p \ge 0$, where $p = \min(m,n)$. They are never negative, and by convention they are sorted from biggest to smallest.
The key link is $A\mathbf{v}_i = \sigma_i\mathbf{u}_i$: $A$ sends the perpendicular input direction $\mathbf{v}_i$ to the perpendicular output direction $\mathbf{u}_i$, stretched by $\sigma_i$.
Geometry. $A$ maps the unit sphere (all $\|\mathbf{x}\| = 1$) onto an ellipsoid whose semi-axes are the vectors $\sigma_i\mathbf{u}_i$. The largest singular value $\sigma_1$ is the most that $A$ can stretch any unit vector, and the smallest tells you the least it stretches (or that it flattens a direction completely when it is $0$).
Orthogonal matrices are rotations, possibly combined with a mirror flip. For a square matrix, a flip appears when $\det A < 0$. Neither rotation nor flip changes lengths, so all the stretching lives in $\Sigma$.
Why do we need it?
We need one tool that works for every matrix, even rectangular or singular ones, and that shows what the matrix really does: which directions it stretches a lot and which it flattens.
Where is it used?
PCA, low-rank compression of images and models, latent semantic analysis, recommender systems, pseudoinverses for least squares, the spectral norm (spectral normalisation of layers) and the condition number.
How is it used?
U, s, Vt = np.linalg.svd(A, full_matrices=False). Read the singular values s: large means important, near zero means redundant. To apply A: rotate with Vt, scale by s, rotate with U. Keep only the first k to simplify.
- Singular values are not eigenvalues. They are always $\ge 0$ and real, for any matrix, even a rotation. (A pure rotation has complex eigenvalues but singular values all equal to $1$, because it does not stretch anything.)
- The directions are not unique when singular values repeat. For a pure rotation, every direction is as good as any other. The singular values themselves are always unique.
- The two rotations $U$ and $V$ are different in general. Do not mix them up: $V$ lives in the input space, $U$ in the output space. For a $3\times2$ matrix they even have different sizes.
Quick check: a matrix has singular values $5$ and $0$. What does the unit circle become?
A line segment, of half-length $5$. The second singular value is $0$, so the ellipse is completely flattened in one direction. The matrix has rank 1.
SVD for any shape: full versus thin core
When $A$ is not square, $U$ and $V$ have different sizes. Take a tall $3\times2$ matrix. It eats vectors with 2 numbers and produces vectors with 3 numbers. The input world has only 2 directions to rotate, but the output world has 3.
The stretch step $\Sigma$ can only use two of the three output directions, because it only has two inputs to stretch. The third output direction gets multiplied by zero: it is never used. The thin (or economy) SVD simply leaves out the unused directions. The full SVD keeps them so that $U$ stays a nice square orthogonal matrix.
Let $A = \begin{bmatrix} 3 & 0 \\ 0 & 2 \\ 0 & 0 \end{bmatrix}$ ($3\times 2$). A valid full SVD is
$$A = \underbrace{\begin{bmatrix} 1 & 0 & 0 \\ 0 & 1 & 0 \\ 0 & 0 & 1 \end{bmatrix}}_{U\ (3\times3)}\underbrace{\begin{bmatrix} 3 & 0 \\ 0 & 2 \\ 0 & 0 \end{bmatrix}}_{\Sigma\ (3\times2)}\underbrace{\begin{bmatrix} 1 & 0 \\ 0 & 1 \end{bmatrix}}_{V^\top\ (2\times2)}.$$The third row of $\Sigma$ is all zeros, so the third column of $U$ is multiplied by zero in every product. Delete that column and that row of $\Sigma$ and you get the thin SVD, with $U$ of size $3\times2$ and $\Sigma$ of size $2\times2$. The product is unchanged.
Let $A$ be $m\times n$ and $p = \min(m, n)$.
| $U$ | $\Sigma$ | $V$ | Property | |
|---|---|---|---|---|
| Full | $m\times m$ | $m\times n$ | $n\times n$ | $U$ and $V$ are square orthogonal |
| Thin / economy | $m\times p$ | $p\times p$ | $n\times p$ | $U^\top U = I$, $V^\top V = I$ (orthonormal columns) |
Both give the same product $A = U\Sigma V^\top$. The thin form saves memory, and it is what np.linalg.svd(A, full_matrices=False) returns. If the rank is $r \lt p$ you can go further and keep only $r$ columns: the compact SVD.
The extra columns of the full $U$ (or $V$) are a free choice of orthonormal vectors that fill out the space. They are important in one place only: they form bases for the subspaces that $A$ does not reach, as the next section shows.
Why do we need it?
Data matrices are rarely square, and a full square U for a tall matrix can be gigantic. The thin SVD keeps only what is needed, so memory stays small.
Where is it used?
np.linalg.svd(..., full_matrices=False), scikit-learn's PCA and TruncatedSVD, and any code that decomposes a tall data matrix with many rows.
How is it used?
Use full_matrices=False for tall or wide data. Use the full version only if you need the extra columns (they span the spaces A never reaches). Check the shapes: U is m×p, s has p numbers, Vt is p×n.
- Do not expect to "see" the pretty numbers from a textbook in your own output. Singular vectors are only unique up to a sign (and up to rotation within a pair of equal singular values), so software may print $-\mathbf{u}_1$ and $-\mathbf{v}_1$ instead. The product is the same.
- For a wide matrix, the roles flip: $V$ gets the extra columns. A practical trick: the SVD of $A^\top$ is $V\Sigma^\top U^\top$. Same singular values.
- NumPy returns $V^\top$ (already transposed), called
Vt, not $V$.
Quick check: for a $5\times3$ matrix, what are the sizes of $U$, $\Sigma$, $V$ in the thin SVD?
$p = \min(5,3) = 3$, so $U$ is $5\times3$, $\Sigma$ is $3\times3$ and $V$ is $3\times3$. (The full version would use $5\times5$, $5\times3$ and $3\times3$.)
How the SVD is related to eigendecomposition core
The SVD looks new, but it hides two familiar friends. The trick is that for any matrix $A$ (even a rectangular one), the matrices $A^\top A$ and $AA^\top$ are square and symmetric. And symmetric matrices have the nicest eigendecomposition there is (perpendicular eigenvectors, real eigenvalues).
It turns out the singular vectors are those eigenvectors, and the singular values are the square roots of those eigenvalues. So the SVD is "eigen-analysis done on the right matrix".
For $A = \begin{bmatrix} 1 & -1 \\ 2 & 2 \end{bmatrix}$ (the example before):
- $A^\top A = \begin{bmatrix} 5 & 3 \\ 3 & 5 \end{bmatrix}$ has eigenvalues $8$ and $2$, with unit eigenvectors $\tfrac1{\sqrt2}[1,1]$ and $\tfrac1{\sqrt2}[-1,1]$. These are $\mathbf{v}_1, \mathbf{v}_2$.
- $AA^\top = \begin{bmatrix} 1 & -1 \\ 2 & 2 \end{bmatrix}\begin{bmatrix} 1 & 2 \\ -1 & 2 \end{bmatrix} = \begin{bmatrix} 2 & 0 \\ 0 & 8 \end{bmatrix}$ has eigenvalues $8$ (eigenvector $[0,1]$) and $2$ (eigenvector $[1,0]$). These are $\mathbf{u}_1 = [0,1]$ and $\mathbf{u}_2 = \pm[1,0]$.
- The same eigenvalues $8$ and $2$ appear both times, and $\sigma_i = \sqrt{\lambda_i}$: $\sqrt8 \approx 2.83$ and $\sqrt2 \approx 1.41$. ✓
Substitute $A = U\Sigma V^\top$ and use $U^\top U = I$:
$$A^\top A = (U\Sigma V^\top)^\top(U\Sigma V^\top) = V\Sigma^\top\underbrace{U^\top U}_{I}\Sigma V^\top = V\,\Sigma^2\,V^\top,$$ $$AA^\top = U\Sigma V^\top V\Sigma^\top U^\top = U\,\Sigma^2\,U^\top.$$Each right-hand side is an eigendecomposition $Q\Lambda Q^\top$ with $\Lambda = \Sigma^2$. Therefore:
- the right singular vectors $\mathbf{v}_i$ are eigenvectors of $A^\top A$;
- the left singular vectors $\mathbf{u}_i$ are eigenvectors of $AA^\top$;
- the singular values are $\sigma_i = \sqrt{\lambda_i(A^\top A)} = \sqrt{\lambda_i(AA^\top)}$ (the two matrices share their non-zero eigenvalues).
Singular values vs eigenvalues.
| Eigenvalues $\lambda$ | Singular values $\sigma$ | |
|---|---|---|
| Which matrices | Square only | Any $m\times n$ |
| Can be negative or complex? | Yes | Never: always real and $\ge 0$ |
| Special directions | $P$: not perpendicular in general | $U$ and $V$: always orthonormal |
| Meaning | Stretch along directions that stay put | Stretch from the input direction $\mathbf{v}_i$ to the (different) output direction $\mathbf{u}_i$ |
| When equal | For a symmetric positive semidefinite matrix they coincide ($U = V = Q$, $\sigma_i = \lambda_i$). For symmetric with a negative eigenvalue, $\sigma_i = |\lambda_i|$. | |
Why do we need it?
It explains where the singular vectors and values come from, and ties the new SVD to the eigenvalue tools you already know. It also shows why PCA can be done either way.
Where is it used?
PCA (eigenvectors of the covariance are the right singular vectors of the centred data), kernel methods, spectral clustering, and checking SVD output with eigenvalue code.
How is it used?
By hand: form AᵀA, find its eigenvalues λ, set σ = √λ, take v from the eigenvectors, then u = Av/σ. In real code call svd directly. Do not form AᵀA, because that squares the condition number.
- Do not compute the SVD by forming $A^\top A$ in practice. It squares the condition number (exactly like the QR widget showed), so small singular values lose their accuracy. Real SVD routines work on $A$ directly. The relationship is for understanding and for small examples.
- The signs of $\mathbf{u}_i$ and $\mathbf{v}_i$ must be chosen together ($\mathbf{u}_i = A\mathbf{v}_i/\sigma_i$). Taking eigenvectors of $A^\top A$ and $AA^\top$ separately can give mismatched signs.
Quick check: a matrix has singular values $3$ and $2$. What are the eigenvalues of $A^\top A$?
$\sigma_i^2$: they are $9$ and $4$. (And $AA^\top$ has the same non-zero eigenvalues.)
The SVD gives the four fundamental subspaces, and the rank core
Remember the four subspaces from Chapter 1.8. The SVD sorts every input direction and every output direction into two piles:
- Directions where $\sigma_i \gt 0$: the matrix really uses them. Input $\mathbf{v}_i$ becomes output $\mathbf{u}_i$, stretched.
- Directions where $\sigma_i = 0$ (or where there is no $\sigma$ at all): the matrix ignores or never reaches them. An input $\mathbf{v}_i$ is crushed to zero. An output $\mathbf{u}_i$ is never produced.
The first pile gives the column space and the row space. The second pile gives the two null spaces. And the number of directions in the first pile is the rank.
$A = \begin{bmatrix} 1 & 2 \\ 2 & 4 \\ 3 & 6 \end{bmatrix}$. The second column is twice the first, so the rank is $1$. Its SVD has $\sigma_1 = \sqrt{70} \approx 8.37$ and $\sigma_2 = 0$, with
- $\mathbf{u}_1 = \tfrac{1}{\sqrt{14}}[1,2,3] \approx [0.27, 0.53, 0.80]$: the direction of every column. This spans the column space.
- $\mathbf{v}_1 = \tfrac{1}{\sqrt5}[1,2]$: every row is a multiple of this. It spans the row space.
- $\mathbf{v}_2 = \tfrac{1}{\sqrt5}[2,-1]$: $A\mathbf{v}_2 = \tfrac{1}{\sqrt5}[2-2,\,4-4,\,6-6] = \mathbf{0}$. It spans the null space.
- $\mathbf{u}_2, \mathbf{u}_3$: any two perpendicular unit vectors that are also perpendicular to $\mathbf{u}_1$. They span the left null space $N(A^\top)$.
Counting dimensions: $1 + 1 = 2 = n$ on the input side, and $1 + 2 = 3 = m$ on the output side. ✓
Let $A = U\Sigma V^\top$ have rank $r$ (so $\sigma_1,\dots,\sigma_r \gt 0$ and the rest are $0$). Then:
| Subspace | Lives in | Orthonormal basis | Dimension |
|---|---|---|---|
| Column space $C(A)$ | $\mathbb{R}^m$ | $\mathbf{u}_1,\dots,\mathbf{u}_r$ | $r$ |
| Left null space $N(A^\top)$ | $\mathbb{R}^m$ | $\mathbf{u}_{r+1},\dots,\mathbf{u}_m$ | $m - r$ |
| Row space $C(A^\top)$ | $\mathbb{R}^n$ | $\mathbf{v}_1,\dots,\mathbf{v}_r$ | $r$ |
| Null space $N(A)$ | $\mathbb{R}^n$ | $\mathbf{v}_{r+1},\dots,\mathbf{v}_n$ | $n - r$ |
Why: $A\mathbf{v}_i = \sigma_i\mathbf{u}_i$. For $i \le r$ this makes $\mathbf{u}_i$ reachable. For $i \gt r$ it gives $A\mathbf{v}_i = \mathbf{0}$, so $\mathbf{v}_i$ is in the null space. The same argument with $A^\top\mathbf{u}_i = \sigma_i\mathbf{v}_i$ handles the other two spaces.
Rank = the number of non-zero singular values. In floating-point arithmetic, "zero" never appears exactly. The numerical rank counts the singular values bigger than a small tolerance, for example $\text{tol} = \max(m,n)\cdot\varepsilon_{\text{machine}}\cdot\sigma_1$. This is how np.linalg.matrix_rank works. It is far more trustworthy than counting pivots in elimination.
Why do we need it?
We need to know how much information a matrix really holds (its rank) and which directions it uses or ignores. Counting pivots is fragile on real data; the SVD is reliable.
Where is it used?
np.linalg.matrix_rank, spotting redundant or collinear features, pseudoinverses and minimum-norm solutions, null spaces in model analysis, and checking whether an embedding matrix has collapsed.
How is it used?
Compute s from the SVD, choose a tolerance (for example 1e-10 * s[0]), and count how many s exceed it: that is the rank r. Columns of U beyond r span the left null space; rows of Vt beyond r span the null space.
- Exact "zero" does not exist on a computer. Always compare $\sigma_i$ to a tolerance. A tiny non-zero singular value means "numerically rank-deficient": the matrix is almost singular even though it is technically invertible.
- The null-space columns of $V$ (and left-null columns of $U$) are only unique as a subspace. Any other orthonormal basis of the same subspace is just as valid.
Quick check: a $5\times4$ matrix has singular values $7, 3, 0, 0$. What is its rank, and what are the dimensions of its null space and left null space?
Rank $r = 2$. The null space (in $\mathbb{R}^4$) has dimension $n - r = 2$. The left null space (in $\mathbb{R}^5$) has dimension $m - r = 3$.
The outer-product form: a matrix as a stack of layers core
A rank-1 matrix is the simplest non-zero matrix: one column pattern times one row pattern. Think of a multiplication table: row $i$ times column $j$ is $r_i \cdot c_j$. Every row is a copy of the same pattern, just scaled.
The SVD says every matrix is a sum of rank-1 layers, like a photo made of transparent sheets laid on top of each other. The sheets come sorted by importance: the first sheet (with the largest $\sigma$) holds the main structure, and each later sheet adds finer detail.
For $A = \begin{bmatrix} 1 & -1 \\ 2 & 2 \end{bmatrix}$ with $\sigma_1 = 2\sqrt2$, $\mathbf{u}_1 = [0,1]$, $\mathbf{v}_1 = \tfrac1{\sqrt2}[1,1]$ and $\sigma_2 = \sqrt2$, $\mathbf{u}_2 = [-1,0]$, $\mathbf{v}_2 = \tfrac1{\sqrt2}[-1,1]$:
- Layer 1: $\sigma_1\mathbf{u}_1\mathbf{v}_1^\top = 2\sqrt2\begin{bmatrix} 0 \\ 1 \end{bmatrix}\tfrac1{\sqrt2}\begin{bmatrix} 1 & 1 \end{bmatrix} = 2\begin{bmatrix} 0 & 0 \\ 1 & 1 \end{bmatrix} = \begin{bmatrix} 0 & 0 \\ 2 & 2 \end{bmatrix}$.
- Layer 2: $\sigma_2\mathbf{u}_2\mathbf{v}_2^\top = \sqrt2\begin{bmatrix} -1 \\ 0 \end{bmatrix}\tfrac1{\sqrt2}\begin{bmatrix} -1 & 1 \end{bmatrix} = \begin{bmatrix} 1 & -1 \\ 0 & 0 \end{bmatrix}$.
- Add: $\begin{bmatrix} 0 & 0 \\ 2 & 2 \end{bmatrix} + \begin{bmatrix} 1 & -1 \\ 0 & 0 \end{bmatrix} = \begin{bmatrix} 1 & -1 \\ 2 & 2 \end{bmatrix} = A$. ✓
The product $U\Sigma V^\top$ can be read column-times-row, which gives the outer-product form of the SVD:
$$A = \sum_{i=1}^{r}\sigma_i\,\mathbf{u}_i\mathbf{v}_i^\top = \sigma_1\mathbf{u}_1\mathbf{v}_1^\top + \sigma_2\mathbf{u}_2\mathbf{v}_2^\top + \dots + \sigma_r\mathbf{u}_r\mathbf{v}_r^\top.$$Each term $\mathbf{u}_i\mathbf{v}_i^\top$ is an outer product: an $m\times n$ matrix with entries $(u_i)_a(v_i)_b$, of rank 1. The $\sigma_i$ says how much that layer counts.
The layers do not overlap in a harmful way: because the $\mathbf{u}_i$ are perpendicular and so are the $\mathbf{v}_i$, the "size" simply adds up: $\|A\|_F^2 = \sigma_1^2 + \sigma_2^2 + \dots + \sigma_r^2$ (the total energy).
Why do we need it?
A matrix with millions of numbers is hard to understand or store. Writing it as a sum of simple rank-1 layers, sorted by importance, lets us see and keep only the part that matters.
Where is it used?
Image compression ideas, latent semantic analysis (topics as layers), recommender systems (taste factors), and the weight-matrix compression and low-rank adapters used in large neural networks.
How is it used?
Compute the SVD, then add layers one by one with (U[:, :k] * s[:k]) @ Vt[:k]. Each layer helps less than the one before. Stop when the error or the kept energy is good enough.
- An outer product $\mathbf{u}\mathbf{v}^\top$ (column times row, an $m\times n$ matrix) is a different thing from the dot product $\mathbf{u}^\top\mathbf{v}$ (row times column, a number).
- The sum stops at $r$, the rank. Layers with $\sigma_i = 0$ contribute nothing.
Quick check: how many numbers does a rank-1 layer $\mathbf{u}\mathbf{v}^\top$ of a $100\times50$ matrix need, compared with the full matrix?
$100 + 50 = 150$ numbers (the vectors $\mathbf{u}$ and $\mathbf{v}$), against $100\cdot50 = 5000$ for the full matrix. That is about 33 times smaller.
Low-rank approximation and truncated SVD core
If a matrix is a stack of layers sorted by importance, we can keep the first few layers and throw the rest away. The result is a simpler matrix (lower rank) that is still very close to the original, as long as the discarded layers were small.
This is how a blurry-but-recognisable thumbnail relates to a full photo. The first layers carry the big shapes and smooth gradients. The last layers carry fine detail and noise. A rank-$k$ copy stores only the $k$ most important sheets.
And here is the beautiful part: this is not just a good way to simplify a matrix. It is provably the best way.
Again $A = \begin{bmatrix} 1 & -1 \\ 2 & 2 \end{bmatrix}$, with $\sigma_1 = 2\sqrt2$ and $\sigma_2 = \sqrt2$. Keep only the first layer ($k = 1$):
$$A_1 = \sigma_1\mathbf{u}_1\mathbf{v}_1^\top = \begin{bmatrix} 0 & 0 \\ 2 & 2 \end{bmatrix}, \qquad A - A_1 = \begin{bmatrix} 1 & -1 \\ 0 & 0 \end{bmatrix}.$$The size of the leftover (Frobenius norm) is $\sqrt{1 + 1} = \sqrt2$, which is exactly $\sigma_2$. The kept energy is $\sigma_1^2/(\sigma_1^2 + \sigma_2^2) = 8/10 = 80\%$.
A bigger one. Suppose a matrix has singular values $10, 5, 1, 0.5$. Their squares are $100, 25, 1, 0.25$ and the total is $126.25$. Cumulative energy: $k=1$: $100/126.25 = 79.2\%$; $k=2$: $125/126.25 = 99.0\%$. Two layers out of four already hold 99% of the energy, so $k = 2$ is a good choice.
The truncated SVD of rank $k$ keeps the $k$ largest singular values:
$$A_k = \sum_{i=1}^{k}\sigma_i\mathbf{u}_i\mathbf{v}_i^\top = U_k\Sigma_kV_k^\top.$$Eckart–Young theorem. Among all matrices $B$ with rank at most $k$, $A_k$ is the closest to $A$:
$$\|A - A_k\|_F = \sqrt{\sigma_{k+1}^2 + \dots + \sigma_r^2} \;=\; \min_{\operatorname{rank}B\le k}\|A - B\|_F, \qquad \|A - A_k\|_2 = \sigma_{k+1}.$$So no other rank-$k$ matrix does better, and the error is exactly the sum of what you threw away.
- Storage: $A_k$ needs $U_k$ ($mk$ numbers), $V_k$ ($nk$) and $k$ singular values: $k(m+n+1)$ numbers, against $mn$ for $A$. It saves memory when $k(m+n+1) \lt mn$.
- Choosing $k$ by explained energy: $\text{energy}(k) = \dfrac{\sigma_1^2 + \dots + \sigma_k^2}{\sigma_1^2 + \dots + \sigma_r^2}$. Pick the smallest $k$ with energy above a target (90%, 95%, 99%), or look for an elbow where the singular values suddenly flatten. For centred data this is the "explained variance" of PCA.
- Randomized SVD (awareness). For a huge matrix you may only want the top $k$ layers, and a full SVD is too slow. The randomized method multiplies $A$ by a thin random matrix $\Omega$ to capture where $A$ "points", orthonormalises the result with QR to get $Q$, then takes the SVD of the small matrix $Q^\top A$. It gets an excellent approximation of $A_k$ at a fraction of the cost (
sklearn.utils.extmath.randomized_svd, orTruncatedSVD).
Why do we need it?
Real matrices are often close to low-rank: a few directions carry most of the information. Keeping only those gives a smaller, faster, denoised matrix, and Eckart–Young promises it is the best possible one of that size.
Where is it used?
Image compression, PCA, latent semantic analysis and topic models, recommender systems, denoising, compressing neural-network weights, and fast approximate matrix products.
How is it used?
Take the SVD, pick k from the cumulative energy (np.cumsum(s**2) / np.sum(s**2) reaching 0.9 or 0.99) or from an elbow in the singular-value plot, then form A_k = (U[:, :k] * s[:k]) @ Vt[:k]. For huge matrices use randomized SVD (scikit-learn TruncatedSVD).
- "Best" depends on the measure. Eckart–Young is about total squared error (Frobenius) and about the worst-case stretch (spectral norm). A rank-$k$ copy may still look wrong in a way you care about, such as a thin bright line in an image.
- Truncating is not the same as deleting columns. A rank-$k$ matrix still has all $m\times n$ entries. What got smaller is the description ($U_k$, $\Sigma_k$, $V_k$), not the grid.
- For PCA, subtract the mean of each column first. The SVD of uncentred data mostly finds the direction towards the mean.
- Do not pick $k$ blindly. Plot the singular values or the cumulative energy, and decide with the task in mind.
Quick check: a $1000\times500$ matrix is approximated at rank $k = 20$. How many numbers are stored, and what fraction of the original is that?
$k(m+n+1) = 20\cdot(1000 + 500 + 1) = 30\,020$ numbers. The original has $500\,000$, so the fraction is $30\,020/500\,000 \approx 6\%$.
Matrix efficiency from low rank: storing and using $BC$ instead of $A$
Imagine a huge table of ratings: $1000$ people by $1000$ films, one million numbers. But people's tastes follow only a few "themes" (likes comedies, likes action, likes slow films...). If there are just $10$ themes, then every rating is a mix of $10$ theme scores.
So instead of storing the whole million-number table, store two thin tables: a $1000\times10$ table (how much each person likes each theme) and a $10\times1000$ table (how much each film has of each theme). Multiplying the two thin tables rebuilds the big one. Two thin matrices can hold the same information as one fat matrix, in a far smaller space.
That is low-rank factorisation (also called low-rank approximation): write a big matrix as a product of two thin ones, $A \approx BC$.
Let $A$ be $1000\times1000$ with rank (or approximate rank) $r = 10$. Write $A \approx BC$ with $B$ of size $1000\times10$ and $C$ of size $10\times1000$.
- Storage: $A$ needs $1000\cdot1000 = 1\,000\,000$ numbers. $B$ and $C$ together need $1000\cdot10 + 10\cdot1000 = 20\,000$. That is 50 times smaller.
- Speed: multiplying $A$ by a vector $\mathbf{x}$ costs $1\,000\,000$ multiplications. With the factors, do $C\mathbf{x}$ first ($10\cdot1000 = 10\,000$ multiplications, giving 10 numbers), then $B(C\mathbf{x})$ ($1000\cdot10 = 10\,000$ more). Total $20\,000$: again 50 times faster.
The bracket order matters: $B(C\mathbf{x})$ is cheap. Building the big matrix $BC$ first and then multiplying would throw the saving away.
For an $m\times n$ matrix, a rank-$r$ factorisation is $A \approx BC$ with $B$ of size $m\times r$ and $C$ of size $r\times n$. The best choice (Eckart–Young) comes from the truncated SVD: $B = U_r\Sigma_r$ and $C = V_r^\top$.
- Numbers to store: $(m + n)\,r$, instead of $mn$.
- Multiplications for $A\mathbf{x}$: $(m+n)\,r$, instead of $mn$.
- It saves effort only when $r \lt \dfrac{mn}{m+n}$. For a square $n\times n$ matrix that means $r \lt n/2$.
- The ratio $(m+n)r / (mn)$ is the fraction you keep.
Why do we need it?
Modern neural networks have weight matrices with millions of entries. Training, storing and running them is expensive. If the useful information is low-rank, two thin matrices can stand in for the fat one at a fraction of the cost.
Where is it used?
Compressing trained networks, speeding up large embedding and projection layers, recommender systems, and LoRA (Low-Rank Adaptation), a way of fine-tuning large language models by training only thin matrices. LoRA is explained in Chapter 1.17, where the two thin matrices are called $B$ ($d\times r$) and $A$ ($r\times d$) and the new weights are $W + BA$. Same idea, different letters.
How is it used?
Take the SVD of the weight matrix, keep the top $r$ values, and replace $W$ by two thin layers $B$ and $C$. Or create two thin matrices from the start and learn them directly. Always apply them as $B(C\mathbf{x})$, never by first forming $BC$.
- A low-rank copy is only good if the matrix really is close to low-rank. The saving is guaranteed; the accuracy is not. Check the singular values or the error first.
- Rank $r$ cannot exceed $\min(m, n)$, and the saving disappears as $r$ approaches $mn/(m+n)$.
- Order of multiplication: $B(C\mathbf{x})$ is cheap, $(BC)\mathbf{x}$ is not.
Quick check: a $4096\times4096$ layer is replaced by two factors of rank $16$. How many numbers do the factors have, and what fraction of the original is that?
$(m+n)r = (4096 + 4096)\cdot16 = 131\,072$ numbers. The original has $4096^2 = 16\,777\,216$. The fraction is $131\,072/16\,777\,216 = 1/128 \approx 0.78\%$.
How the decompositions connect, and where they appear in ML
The six factorisations are not six unrelated tricks. They form a family tree, and each member is "the right tool for one situation".
- LU is Gaussian elimination, written down. Cholesky is LU for the special case of a symmetric positive definite matrix: the same idea, half the work.
- QR is Gram–Schmidt, written down. It works for tall matrices and is the careful way to do least squares.
- Eigendecomposition finds the directions a square matrix only stretches. For a symmetric matrix it is perfectly tidy ($Q\Lambda Q^\top$).
- SVD is what you get when you do that tidy symmetric analysis on $A^\top A$ and $AA^\top$. It is the only one that works for every matrix and the one that shows the most: rank, subspaces, pseudoinverse, best low-rank copies.
A good rule of thumb: if the matrix has structure (triangular-friendly, symmetric positive definite), use the faster specialised factorisation. If you need insight, or the matrix is awkward, use the SVD.
The same job, different tools. Solving $A\mathbf{x} = \mathbf{b}$ for a symmetric positive definite $A$ (such as a covariance matrix):
- With Cholesky: $A = LL^\top$, then two triangular solves. Cost $\approx \tfrac13 n^3$.
- With LU: $PA = LU$, then two triangular solves. Cost $\approx \tfrac23 n^3$.
- With the SVD: $\mathbf{x} = V\Sigma^{-1}U^\top\mathbf{b}$. Cost several times higher, but it also tells you whether $A$ is nearly singular and still works if $A$ is exactly singular (via the pseudoinverse).
All three give the same $\mathbf{x}$. They differ in speed, safety and how much they reveal.
| Decomposition | Form | Applies to | Pieces | Main use |
|---|---|---|---|---|
| LU | $PA = LU$ | Square | triangular × triangular | Solving linear systems, determinants |
| Cholesky | $A = LL^\top$ | Symmetric positive definite | triangular × its transpose | Fast solves, Gaussian sampling, log-determinant, whitening, PD test |
| QR | $A = QR$ | Any $m\times n$ (usually $m \ge n$) | orthonormal × triangular | Least squares, eigenvalue algorithms |
| Eigen | $A = PDP^{-1}$ | Square, diagonalisable | eigenvectors × diagonal × inverse | Powers, dynamics, long-run behaviour |
| Spectral | $A = Q\Lambda Q^\top$ | Symmetric | orthogonal × diagonal × transpose | PCA, quadratic forms, symmetric analysis |
| SVD | $A = U\Sigma V^\top$ | Any matrix | orthogonal × diagonal × orthogonal | Everything: rank, pseudoinverse, compression, PCA |
Links worth remembering: Cholesky follows from LU for symmetric PD matrices (the pivots are positive, and $U = DL^\top$). For symmetric positive semidefinite $A$, the SVD and the spectral decomposition are the same ($U = V = Q$, $\sigma_i = \lambda_i$). For any $A$, the SVD is the spectral decomposition of $A^\top A$ and $AA^\top$ glued together by $A\mathbf{v}_i = \sigma_i\mathbf{u}_i$.
Why do we need it?
With six tools it is easy to choose the wrong one, which wastes time or loses accuracy. Seeing how they connect gives a simple rule for which to use.
Where is it used?
Every numerical library (NumPy, SciPy, PyTorch, scikit-learn) chooses among these inside. Knowing them helps you read error messages and set options such as solver="cholesky" or "svd" in scikit-learn's Ridge.
How is it used?
Ask two questions: what shape and structure does my matrix have (square, symmetric, positive definite, tall)? What do I need (a solution, a sample, structure, compression)? Then use the picker below.
Where each idea shows up in machine learning
- PCA via SVD. Centre the data, take the thin SVD, keep the top $k$ right singular vectors. It is more accurate than eigen-decomposing the covariance.
- LSA and topic models. Truncated SVD of a word-by-document matrix groups words into topics.
- Image compression. A rank-$k$ copy stores $k(m+n+1)$ numbers (the widget above).
- Recommender systems. Rating matrices are approximately low-rank. Users and items become short "taste" vectors from $U_k\Sigma_k^{1/2}$ and $V_k\Sigma_k^{1/2}$.
- LoRA (low-rank adaptation of large language models) trains only thin low-rank matrices instead of the whole weight matrix. See the section above for the arithmetic and Chapter 1.17 for LoRA itself.
- Model compression. Factor a trained weight matrix with a truncated SVD to make the network smaller and faster.
- Gaussian processes use the Cholesky factor of the kernel matrix to solve systems, sample functions and compute $\log\det$.
- Numerical rank of embeddings. Count the singular values of an embedding matrix above a tolerance to see how many dimensions are really used. A collapsed embedding space shows up as a few large singular values and many tiny ones.
Quick check: which factorisation would you use to (a) solve $A\mathbf{x} = \mathbf{b}$ for a covariance matrix many times, (b) fit a regression with many features, (c) compress a photo?
(a) Cholesky: covariance matrices are symmetric positive definite, and the factor is reused. (b) QR (or the SVD if some features may be dependent). (c) Truncated SVD.
Recap, cheat sheet and practice
- We factorise to solve faster, to reveal structure, and to compress. Good pieces are triangular, orthogonal or diagonal.
- LU ($PA = LU$) is Gaussian elimination with the multipliers saved in $L$. Factor once, then solve for many $\mathbf{b}$ with forward and back substitution.
- Cholesky ($A = LL^\top$) is the "square root" of a symmetric positive definite matrix, and the fast LU for such matrices (half the work). It gives: solves via two triangular passes, samples $\mathbf{x} = \boldsymbol\mu + L\mathbf{z}$, $\log\det A = 2\sum\log l_{ii}$, a positive-definiteness test (it fails otherwise), and whitening $L^{-1}(\mathbf{x}-\boldsymbol\mu)$. Add a tiny jitter $\varepsilon I$ if rounding makes it fail.
- QR ($A = QR$) is Gram–Schmidt written down. It solves least squares stably through $R\mathbf{x} = Q^\top\mathbf{b}$. Libraries build it with Householder reflections.
- Eigendecomposition ($A = PDP^{-1}$, or $Q\Lambda Q^\top$ if symmetric) needs a square matrix and is not always possible.
- SVD ($A = U\Sigma V^\top$) works for every matrix: rotate, stretch, rotate. The unit sphere becomes an ellipsoid with semi-axes $\sigma_i\mathbf{u}_i$. It gives the rank (number of non-zero $\sigma$), bases for all four subspaces, and $A = \sum\sigma_i\mathbf{u}_i\mathbf{v}_i^\top$. Also $A^\top A = V\Sigma^2V^\top$ and $AA^\top = U\Sigma^2U^\top$.
- Low-rank factorisation $A \approx BC$ stores and applies $(m+n)r$ numbers instead of $mn$ (apply it as $B(C\mathbf{x})$): the idea behind compression and LoRA (Chapter 1.17).
- Truncated SVD $A_k$ is the best rank-$k$ approximation (Eckart–Young). Error $= \sqrt{\sigma_{k+1}^2+\dots}$. Pick $k$ by explained energy. Storage is $k(m+n+1)$.
Cheat sheet
| Decomposition | Form | Needs | Remember |
|---|---|---|---|
| LU | $PA = LU$ | square | elimination diary; reuse for many $\mathbf{b}$; cost $\tfrac23n^3$ |
| Cholesky | $A = LL^\top$ | symmetric PD | matrix square root; $\tfrac13n^3$; solve, $\boldsymbol\mu + L\mathbf{z}$, $2\sum\log l_{ii}$, PD test, jitter |
| QR | $A = QR$ | any (tall) | $R\mathbf{x} = Q^\top\mathbf{b}$; reduced vs full |
| Eigen | $PDP^{-1}$ | square, diagonalisable | $A^k = PD^kP^{-1}$ |
| Spectral | $Q\Lambda Q^\top$ | symmetric | always exists; real $\lambda$, perpendicular $\mathbf{q}$ |
| SVD | $U\Sigma V^\top$ | any | $\sigma_i = \sqrt{\lambda_i(A^\top A)} \ge 0$; $A\mathbf{v}_i = \sigma_i\mathbf{u}_i$ |
| Rank from SVD | count $\sigma_i \gt \text{tol}$ | any | $\mathbf{u}_{1..r}$: column space; $\mathbf{v}_{r+1..n}$: null space |
| Truncated SVD | $A_k = U_k\Sigma_kV_k^\top$ | any | best rank-$k$; error $\sqrt{\sum_{i\gt k}\sigma_i^2}$ |
import numpy as np
# ---------- 1. LU without pivoting, from scratch ----------
def lu_nopivot(A):
n = A.shape[0]
L, U = np.eye(n), A.astype(float).copy()
for c in range(n - 1):
for r in range(c + 1, n):
L[r, c] = U[r, c] / U[c, c] # the multiplier goes into L
U[r, c:] -= L[r, c] * U[c, c:] # row r minus multiplier times row c
return L, np.triu(U)
A = np.array([[2., 1., 1.], [4., 3., 3.], [8., 7., 9.]])
L, U = lu_nopivot(A)
print(np.allclose(L @ U, A)) # True
# with pivoting and reuse for many right-hand sides (SciPy)
from scipy.linalg import lu_factor, lu_solve
lu, piv = lu_factor(A) # factor once
for b in (np.array([4., 10., 24.]), np.array([4., 10., 22.])):
print(lu_solve((lu, piv), b)) # cheap solves: [1 1 1] and [1 2 0]
# ---------- 2. Cholesky from scratch + PD test ----------
def cholesky_scratch(A):
n = A.shape[0]
L = np.zeros((n, n))
for j in range(n):
s = A[j, j] - L[j, :j] @ L[j, :j]
if s <= 0:
raise ValueError("not positive definite")
L[j, j] = np.sqrt(s)
for i in range(j + 1, n):
L[i, j] = (A[i, j] - L[i, :j] @ L[j, :j]) / L[j, j]
return L
S = np.array([[4., 2.], [2., 5.]])
print(cholesky_scratch(S)) # [[2 0] [1 2]]
print(np.allclose(cholesky_scratch(S), np.linalg.cholesky(S))) # True
# ---------- 3. Sample a 2-D Gaussian with Cholesky: x = mu + L z ----------
mu = np.array([1., 0.5])
Lc = np.linalg.cholesky(S)
Z = np.random.default_rng(0).standard_normal((2, 5000))
X = mu[:, None] + Lc @ Z # each column is one sample
print(np.cov(X)) # close to S
# plot with: import matplotlib.pyplot as plt; plt.scatter(X[0], X[1], s=3)
# ---------- 3b. Cholesky for solving, log-det, whitening, PD test, jitter ----------
from scipy.linalg import cho_factor, cho_solve, solve_triangular
c, low = cho_factor(S) # factor once
print(cho_solve((c, low), np.array([6., 7.]))) # [1. 1.] two triangular solves, no inverse
logdet = 2 * np.sum(np.log(np.diag(Lc))) # log det S = 2 * sum(log(diag(L)))
print(np.isclose(logdet, np.linalg.slogdet(S)[1])) # True (log 16)
z = solve_triangular(Lc, np.array([2., 1.]), lower=True) # whitening: solve L z = x - mu (mu = 0)
print(z, np.linalg.norm(z)) # [1. 0.] Mahalanobis distance = 1.0
def is_pd(A):
try:
np.linalg.cholesky(A) # raises LinAlgError if A is not PD
return True
except np.linalg.LinAlgError:
return False
print(is_pd(S), is_pd(np.array([[1., 2.], [2., 1.]]))) # True False
# the jitter trick: a nearly singular kernel matrix
xs = np.linspace(0, 1, 30)
K = np.exp(-(xs[:, None] - xs[None, :])**2 / (2 * 0.5**2))
print(is_pd(K)) # typically False (rounding errors)
print(is_pd(K + 1e-6 * np.eye(30))) # True
# careful: scipy.linalg.cholesky(A) returns the UPPER factor unless you pass lower=True
# ---------- 4. QR and stable least squares ----------
x = np.array([-2., -1., 0., 1., 2.])
y = np.array([-1.5, -0.5, 1., 1., 3.])
M = np.column_stack([np.ones_like(x), x])
Q, R = np.linalg.qr(M) # reduced QR (default)
beta = np.linalg.solve(R, Q.T @ y) # R beta = Q^T y
print(beta) # [0.6 1.05] (intercept, slope)
# ---------- 5. SVD and its link with eigenvalues ----------
B = np.array([[1., -1.], [2., 2.]])
U, s, Vt = np.linalg.svd(B) # note: returns V transposed
print(s) # [2.828 1.414]
print(np.sqrt(np.linalg.eigvalsh(B.T @ B))[::-1]) # same numbers
print(np.allclose(B, U @ np.diag(s) @ Vt)) # True
# thin vs full on a tall matrix
T = np.random.default_rng(1).standard_normal((6, 3))
print([a.shape for a in np.linalg.svd(T, full_matrices=True)]) # (6,6) (3,) (3,3)
print([a.shape for a in np.linalg.svd(T, full_matrices=False)]) # (6,3) (3,) (3,3)
# rank and the four subspaces of a rank-1 matrix
C = np.array([[1., 2.], [2., 4.], [3., 6.]])
U, s, Vt = np.linalg.svd(C)
r = int(np.sum(s > 1e-10 * s[0])) # numerical rank = 1
col_space, left_null = U[:, :r], U[:, r:]
row_space, null_space = Vt[:r].T, Vt[r:].T
print(r, np.allclose(C @ null_space, 0)) # 1 True
# ---------- 6. Image compression with a truncated SVD ----------
# any 2-D grey image in [0, 1]; here a tiny synthetic one so the code runs anywhere
i, j = np.mgrid[0:128, 0:128]
img = 0.2 + 0.3 * (i / 127) + 0.5 * (np.hypot(i - 50, j - 60) < 25)
U, s, Vt = np.linalg.svd(img, full_matrices=False)
m, n = img.shape
for k in (5, 20, 50, 100):
Ak = (U[:, :k] * s[:k]) @ Vt[:k] # sum of the first k layers
err = np.linalg.norm(img - Ak) / np.linalg.norm(img)
ratio = k * (m + n + 1) / (m * n) # storage compared with the original
print(k, round(err, 4), round(ratio, 3))
energy = np.cumsum(s**2) / np.sum(s**2)
k90 = int(np.searchsorted(energy, 0.90)) + 1 # smallest k keeping 90% energy
# ---------- 7. Low-rank factors: store and apply B @ C instead of A ----------
Bf, Cf = U[:, :20] * s[:20], Vt[:20] # from the SVD above (rank 20)
x = np.random.default_rng(2).standard_normal(Cf.shape[1])
y = Bf @ (Cf @ x) # cheap: (m + n) * r multiplications
print(Bf.size + Cf.size, 'numbers instead of', img.size)
# ---------- 8. Randomized SVD (sketch) ----------
def randomized_svd(A, k, extra=5, seed=0):
rng = np.random.default_rng(seed)
Omega = rng.standard_normal((A.shape[1], k + extra)) # thin random matrix
Qr, _ = np.linalg.qr(A @ Omega) # basis for the range of A
Ub, sv, Vt = np.linalg.svd(Qr.T @ A, full_matrices=False) # SVD of a small matrix
return Qr @ Ub[:, :k], sv[:k], Vt[:k]
Ur, sr, Vtr = randomized_svd(img, 2)
print(np.round(sr, 1), np.round(s[:2], 1)) # [55.2 11.7] [55.2 11.8]: close to the exact top two values
1. Which statement about the SVD $A = U\Sigma V^\top$ is true?
2. You want to test whether a symmetric matrix is positive definite, and the matrix is large. A good way is…
3. A $4\times3$ matrix has singular values $6, 2, 0$. What is its rank, and what is the dimension of its null space?
4. The eigenvalues of $A^\top A$ are $16$ and $4$. What are the singular values of $A$?
5. A $200\times100$ image is approximated with a rank-10 truncated SVD. How many numbers are stored, roughly?
6. What does the Eckart–Young theorem tell you?
7. $A = LL^\top$ and $L$ has diagonal entries $2, 2, 2$. What is $\det A$?
8. You need $(\mathbf{x}-\boldsymbol\mu)^\top\Sigma^{-1}(\mathbf{x}-\boldsymbol\mu)$ for a covariance $\Sigma$. What is the best way to compute it?
Practice problems
A. Find the LU decomposition of $\begin{bmatrix} 1 & 2 \\ 3 & 4 \end{bmatrix}$ and use it to get the determinant.
The multiplier is $3/1 = 3$. Row 2 $-$ $3\times$ Row 1 $= [0,\; 4 - 6] = [0, -2]$. So $L = \begin{bmatrix} 1 & 0 \\ 3 & 1 \end{bmatrix}$ and $U = \begin{bmatrix} 1 & 2 \\ 0 & -2 \end{bmatrix}$. Check: $LU = \begin{bmatrix} 1 & 2 \\ 3 & 6 - 2 \end{bmatrix} = \begin{bmatrix} 1 & 2 \\ 3 & 4 \end{bmatrix}$ ✓. No swaps, so $\det A = 1\cdot(-2) = -2$ (and directly: $1\cdot4 - 2\cdot3 = -2$ ✓).
B. Compute the Cholesky factor of $\begin{bmatrix} 9 & 3 \\ 3 & 5 \end{bmatrix}$, then turn $\mathbf{z} = [1, -1]$ into a sample with mean $\mathbf{0}$.
$l_{11} = \sqrt9 = 3$. $l_{21} = 3/3 = 1$. $l_{22} = \sqrt{5 - 1^2} = 2$. So $L = \begin{bmatrix} 3 & 0 \\ 1 & 2 \end{bmatrix}$. Check: $LL^\top = \begin{bmatrix} 9 & 3 \\ 3 & 1 + 4 \end{bmatrix}$ ✓. The sample is $L\mathbf{z} = [3\cdot1,\; 1\cdot1 + 2\cdot(-1)] = [3, -1]$.
C. Find an SVD of $A = \begin{bmatrix} 3 & 0 \\ 0 & -2 \end{bmatrix}$. Why is a minus sign allowed in $U$ but not in $\Sigma$?
The singular values must be non-negative: $\sigma_1 = 3$, $\sigma_2 = 2$. Move the sign into $U$: $A = \begin{bmatrix} 1 & 0 \\ 0 & -1 \end{bmatrix}\begin{bmatrix} 3 & 0 \\ 0 & 2 \end{bmatrix}\begin{bmatrix} 1 & 0 \\ 0 & 1 \end{bmatrix}$. Here $\mathbf{u}_2 = [0,-1]$ and $\mathbf{v}_2 = [0,1]$, so $A\mathbf{v}_2 = [0,-2] = 2\mathbf{u}_2$ ✓. The matrix flips the $y$ direction, and a flip is an orthogonal operation, so it belongs in $U$. Note $\det A = -6$ and $\sigma_1\sigma_2 = 6$: the sign of the determinant records the flip.
D. A matrix has singular values $10, 4, 1, 0.5$. Which $k$ keeps at least 95% of the energy, and what is the Frobenius error of that rank-$k$ approximation?
Squares: $100, 16, 1, 0.25$, total $117.25$. $k = 1$: $100/117.25 = 85.3\%$ (too low). $k = 2$: $116/117.25 = 98.9\%$ ✓. So $k = 2$. The error is $\sqrt{1 + 0.25} = \sqrt{1.25} \approx 1.118$.
E. For $A = \begin{bmatrix} 1 & 2 \\ 2 & 4 \end{bmatrix}$ find $\sigma_1$, $\sigma_2$, a basis for the null space, and write $A$ in outer-product form.
$A = \begin{bmatrix} 1 \\ 2 \end{bmatrix}\begin{bmatrix} 1 & 2 \end{bmatrix}$ has rank 1. $A^\top A = \begin{bmatrix} 5 & 10 \\ 10 & 20 \end{bmatrix}$ has eigenvalues $25$ and $0$, so $\sigma_1 = 5$, $\sigma_2 = 0$. With $\mathbf{u}_1 = \mathbf{v}_1 = \tfrac1{\sqrt5}[1,2]$ we get $A = 5\,\mathbf{u}_1\mathbf{v}_1^\top$ (check: $5\cdot\tfrac15\begin{bmatrix}1\\2\end{bmatrix}[1\ 2]$ ✓). The null space is spanned by $\mathbf{v}_2 = \tfrac1{\sqrt5}[2,-1]$, since $A\mathbf{v}_2 = \tfrac1{\sqrt5}[2-2,\; 4-4] = \mathbf{0}$.
F. Explain in two or three sentences why QR gives a more accurate least-squares answer than the normal equations.
The normal equations multiply $A$ by $A^\top$, which squares the condition number: if $\kappa(A) = 10^8$ then $\kappa(A^\top A) = 10^{16}$, and all accuracy is lost in double precision. QR works with $A$ itself. Multiplying by the orthogonal $Q^\top$ does not change lengths or amplify errors, so only the triangular system $R\mathbf{x} = Q^\top\mathbf{b}$ remains, with $\kappa(R) = \kappa(A)$.
G. Find the Cholesky factor of $\begin{bmatrix} 9 & 3 & 3 \\ 3 & 5 & 3 \\ 3 & 3 & 3 \end{bmatrix}$ and use it to get $\det A$ and $\log\det A$.
$l_{11} = \sqrt9 = 3$; $l_{21} = 3/3 = 1$; $l_{31} = 3/3 = 1$; $l_{22} = \sqrt{5 - 1^2} = 2$; $l_{32} = (3 - 1\cdot1)/2 = 1$; $l_{33} = \sqrt{3 - 1^2 - 1^2} = 1$. So $L = \begin{bmatrix} 3 & 0 & 0 \\ 1 & 2 & 0 \\ 1 & 1 & 1 \end{bmatrix}$. The diagonal is $3, 2, 1$, so $\det A = (3\cdot2\cdot1)^2 = 36$ and $\log\det A = 2(\ln3 + \ln2 + \ln1) = 2\ln6 \approx 3.58$.
H. With $\Sigma = \begin{bmatrix} 4 & 2 \\ 2 & 5 \end{bmatrix}$ and $\boldsymbol\mu = \mathbf{0}$, compare the Euclidean and Mahalanobis distances of $\mathbf{x} = [0, 2]$.
$L = \begin{bmatrix} 2 & 0 \\ 1 & 2 \end{bmatrix}$. Solve $L\mathbf{z} = [0, 2]$: $2z_1 = 0$ gives $z_1 = 0$; then $1\cdot0 + 2z_2 = 2$ gives $z_2 = 1$. So $d_M = \|[0, 1]\| = 1$, while the Euclidean distance is $2$. Compare with $[2, 1]$, which has $d_M = 1$ too but Euclidean distance $2.24$: both are one standard deviation from the centre.
I. Why does adding $\lambda I$ in ridge regression guarantee that Cholesky works?
$X^\top X$ is symmetric positive semidefinite, so all its eigenvalues are $\ge 0$. Adding $\lambda I$ with $\lambda \gt 0$ lifts every eigenvalue by $\lambda$, so they are all $\ge \lambda \gt 0$. A symmetric matrix with all eigenvalues positive is positive definite, so every square root in the Cholesky algorithm is of a positive number.
Matrix Calculus
How does the output change when you nudge the input? When the input is a whole vector, the answer is an arrow (the gradient) or a table (the Jacobian). This is the language of training: backpropagation is just the chain rule written with matrices.
- Refresh derivatives, partial derivatives and the chain rule
- Understand the gradient: the direction of steepest ascent, always perpendicular to the level curves
- Understand the Jacobian as the best local linear approximation of a vector-valued function
- Understand the Hessian: curvature, local minima, and second-order Taylor expansion
- Know and use the essential identities, and check any gradient numerically
- Follow the chain rule through a computational graph: forward versus reverse mode, and backprop through a linear layer
- Derive the gradient of linear regression and connect it to Chapter 1.10
We use ideas from Chapter 1.4 (matrix multiplication and transpose), 1.5 (a matrix as a linear map), 1.10 (least squares), 1.11 (eigenvalues) and 1.12 (positive definite matrices). You do not need to have studied calculus deeply: the first section is a refresher.
Derivatives, partial derivatives and the chain rule
A derivative answers one question: if I nudge the input a tiny bit, how much does the output move? Think of a road. The derivative at a point is how steeply the road climbs right there. A derivative of 3 means "for every tiny step forward, the road rises about 3 times as much".
On a graph, it is the slope of the tangent line: the straight line that just touches the curve. Zoom in far enough on any smooth curve and it looks like that straight line. This "zoom in until it looks straight" idea will come back for the Jacobian and the Hessian.
When a function has several inputs, like a mixing desk with many knobs, you can turn one knob at a time while leaving the others frozen. How much the output moves per unit turn of one knob is a partial derivative.
The chain rule is about gears. If the first gear turns 3 times for every turn of the handle, and the second gear turns 2 times for every turn of the first, then the second gear turns $3\times 2 = 6$ times per turn of the handle. When functions are chained, slopes multiply.
- Derivative. For $f(x)=x^2$ the derivative is $2x$. At $x=3$ the slope is $6$. Check by nudging: $f(3.01) = 9.0601$, a rise of $0.0601$ for a step of $0.01$, and $6\times0.01 = 0.06$. ✓
- Partial derivatives. For $f(x,y)=x^2y$: treat $y$ as a constant to get $\dfrac{\partial f}{\partial x}=2xy$, and treat $x$ as a constant to get $\dfrac{\partial f}{\partial y}=x^2$. At $(3,2)$ these are $12$ and $9$. Check: $f(3,2)=18$ and $f(3.01,2)=18.1202$, a rise of $0.1202\approx12\times0.01$. ✓
- Chain rule. $y=(3x+1)^2$. Break it into two steps: $u=3x+1$, then $y=u^2$. Then $\dfrac{dy}{du}=2u$ and $\dfrac{du}{dx}=3$, so $\dfrac{dy}{dx}=2u\cdot3$. At $x=1$: $u=4$, so the slope is $2\cdot4\cdot3=24$. Check: at $x=1.01$, $u=4.03$ and $y=16.2409$, a rise of $0.2409\approx24\times0.01$. ✓
The derivative of $f$ at $x$ is the limit of the "rise over run" as the step shrinks:
$$f'(x) = \frac{df}{dx} = \lim_{h\to0}\frac{f(x+h)-f(x)}{h}.$$Equivalent view, the linear approximation: for small $h$, $\;f(x+h)\approx f(x) + f'(x)\,h$.
For $f(x_1,\dots,x_n)$, the partial derivative $\dfrac{\partial f}{\partial x_i}$ is the ordinary derivative with respect to $x_i$, holding all the other inputs fixed.
Chain rule (one variable). If $y=f(u)$ and $u=g(x)$, then
$$\frac{dy}{dx}=\frac{dy}{du}\cdot\frac{du}{dx} = f'(g(x))\,g'(x).$$A few derivatives to know: $(x^n)'=nx^{n-1}$, $(e^x)'=e^x$, $(\ln x)'=1/x$, $(\sin x)'=\cos x$, and the sum rule $(f+g)'=f'+g'$. Product rule: $(fg)'=f'g+fg'$.
Why do we need it?
To improve a model we must know which way to change a number and how much the result moves. The derivative is exactly that: the response of the output to a tiny nudge of the input.
Where is it used?
Training every machine-learning model (the derivative of the loss), optimisation in physics and economics, sensitivity analysis ("how much does the answer depend on this input?"), and reasoning about learning rates.
How is it used?
Write the function as small steps, differentiate each step, and multiply the slopes (chain rule). Verify with a tiny nudge: (f(x + h) − f(x)) / h should be close to f′(x).
- The derivative is local. It describes behaviour for tiny nudges only. A large step can behave quite differently (try a big $h$ above).
- Partial does not mean "part of". $\partial f/\partial x$ is a full derivative, just with the other variables held fixed. The curly $\partial$ reminds you there are other inputs.
- In the chain rule, evaluate the outer slope at the inner value: $f'(g(x))$, not $f'(x)$. (In the example, $2u$ used $u=4$, not $x=1$.)
Quick check: use the chain rule to differentiate $y=e^{3x}$, and give the slope at $x=0$.
Let $u=3x$ and $y=e^u$. Then $dy/du=e^u$ and $du/dx=3$, so $dy/dx = 3e^{3x}$. At $x=0$: $3e^0 = 3$.
The gradient core
You are standing on a hillside holding a map with contour lines (lines joining places of equal height). Which way is straight uphill? It is the direction that crosses the contour lines at a right angle, because that is the quickest way to change height. If you walk along a contour line, you stay at the same height: no climbing at all.
The gradient is exactly that uphill arrow. It points toward steepest ascent, and its length says how steep the hill is. Walk the opposite way and you go down as fast as possible. That is why training a model (making the loss smaller) uses $-\nabla f$.
How do we build the arrow? Measure the slope in the $x$-direction (a partial derivative), measure the slope in the $y$-direction, and stack the two numbers into a vector.
Let $f(x,y) = x^2 + 3y^2$. A bowl, steeper in the $y$-direction.
- Partial derivatives: $\partial f/\partial x = 2x$ and $\partial f/\partial y = 6y$. So $\nabla f = [2x,\;6y]$.
- At the point $(1,1)$: $\nabla f = [2, 6]$. The bowl climbs fastest mostly in the $y$ direction. Its steepness is $\|\nabla f\| = \sqrt{4+36}=\sqrt{40}\approx 6.32$.
- The contour through $(1,1)$ is the ellipse $x^2+3y^2=4$. A direction along it is $[-6, 2]$ (swap the entries of the gradient and flip one sign).
- Check the right angle: $[2,6]\cdot[-6,2] = -12+12 = 0$. ✓ The gradient is perpendicular to the contour.
- Climbing speed in the direction $\mathbf{u}=[1,0]$ (along $x$): $\nabla f\cdot\mathbf{u} = 2$. In the gradient's own direction it is the full $\sqrt{40}\approx6.32$. Along the contour it is $0$.
For a function $f:\mathbb{R}^n\to\mathbb{R}$ (many inputs, one output), the gradient is the vector of all partial derivatives:
$$\nabla f(\mathbf{x}) = \begin{bmatrix} \partial f/\partial x_1 \\ \partial f/\partial x_2 \\ \vdots \\ \partial f/\partial x_n \end{bmatrix}\in\mathbb{R}^n.$$It is a column vector with the same shape as $\mathbf{x}$. It gives the linear approximation $f(\mathbf{x}+\boldsymbol{\delta}) \approx f(\mathbf{x}) + \nabla f(\mathbf{x})^\top\boldsymbol{\delta}$.
Steepest ascent. The rate of change of $f$ in the direction of a unit vector $\mathbf{u}$ (the directional derivative) is
$$D_{\mathbf{u}}f = \nabla f\cdot\mathbf{u} = \|\nabla f\|\cos\theta,$$where $\theta$ is the angle between $\mathbf{u}$ and $\nabla f$. This is largest ($=\|\nabla f\|$) when $\theta=0$, i.e. when $\mathbf{u}$ points along $\nabla f$. It is smallest ($=-\|\nabla f\|$) along $-\nabla f$, and $0$ at right angles.
Gradient $\perp$ level sets. A level set is a set where $f$ is constant, $f(\mathbf{x})=c$ (a contour line in 2D, a surface in 3D). Take a path $\mathbf{x}(t)$ that stays on it. Then $f(\mathbf{x}(t))=c$ never changes, so its derivative is zero. By the chain rule that derivative is $\nabla f\cdot\mathbf{x}'(t)=0$. The path's velocity $\mathbf{x}'(t)$ points along the level set, so the gradient is perpendicular to it.
Where $\nabla f=\mathbf{0}$ the ground is flat: a critical point (a minimum, maximum or saddle). Gradient descent updates $\mathbf{x}\leftarrow\mathbf{x}-\eta\nabla f(\mathbf{x})$.
Why do we need it?
A model has many weights, not just one. We need a single arrow that says, for all weights at once, which way makes the loss rise fastest, so that we can walk the other way.
Where is it used?
Gradient descent, SGD and Adam when training neural networks, logistic and linear regression, gradient clipping, and saliency maps that show which pixels of an image mattered for a prediction.
How is it used?
Compute every partial derivative and stack them into one vector, then update x ← x − η∇f with a small learning rate η. Watch ‖∇f‖: when it is close to zero you are at a flat spot, so stop.
Gradient descent is a local rule. The gradient points to the steepest ascent at this point, for tiny steps only. A big step along $-\nabla f$ can overshoot the valley and make things worse (you will see this with a learning-rate slider at the end of the chapter).
Layout conventions (awareness). Books disagree about whether the derivative of a scalar $y$ with respect to a vector $\mathbf{x}$ is a column or a row. This causes endless sign and transpose confusion:
- Denominator layout: the answer is laid out like the thing you differentiate by ($\mathbf{x}$). So $\partial y/\partial\mathbf{x}$ is a column ($n\times1$), the gradient.
- Numerator layout: the answer is laid out like the thing being differentiated ($y$, the numerator). For a scalar $y$ this gives a row ($1\times n$). For a vector-valued function it gives one row per output ($m\times n$), which is the Jacobian.
This guide uses the most common choice in machine learning: the gradient of a scalar is a column with the same shape as $\mathbf{x}$, and the Jacobian is $m\times n$ (one row per output). So the Jacobian of a scalar function is the row $\nabla f^\top$.
Practical advice: whichever book you read, find its convention first, and always check that matrix shapes multiply. Libraries (PyTorch, JAX) return a gradient with the same shape as the parameter.
Quick check: $f(x,y)=x^2y$. What is $\nabla f$ at $(1,3)$, and which direction should you step to decrease $f$ fastest?
$\nabla f = [2xy,\,x^2] = [6, 1]$ at $(1,3)$. To decrease $f$ fastest, step along $-\nabla f=[-6,-1]$ (normalised if you like: divide by $\sqrt{37}$).
The Jacobian core
The gradient handles functions with many inputs and one output. But a layer of a neural network takes a vector in and gives a vector out. A map of the Earth is another example: it takes (latitude, longitude) and gives (x, y) on the paper. Now every output has its own gradient.
Stack those gradients as the rows of a table: that table is the Jacobian matrix. Row $i$ answers "how does output $i$ respond to each input?".
Here is the beautiful part. Zoom in on a smooth map around a point, and the curved, warped grid starts to look like a perfectly straight, evenly spaced grid. That straight picture is a matrix acting on the plane (Chapter 1.5). The Jacobian is that matrix: the best local linear approximation of the function.
Take $F(x,y) = \big(x^2y,\;\; 5x + \sin y\big)$, two inputs and two outputs.
- Gradient of output 1: $[\,2xy,\; x^2\,]$. Gradient of output 2: $[\,5,\; \cos y\,]$.
- Stack them as rows: $J = \begin{bmatrix} 2xy & x^2 \\ 5 & \cos y \end{bmatrix}$.
- At the point $(1,0)$: $J = \begin{bmatrix} 0 & 1 \\ 5 & 1 \end{bmatrix}$ and $F(1,0) = (0, 5)$.
- Nudge the input by $\boldsymbol{\delta}=[0.01,\;0.02]$. The Jacobian predicts the change $J\boldsymbol{\delta} = [0\cdot0.01+1\cdot0.02,\;5\cdot0.01+1\cdot0.02] = [0.02,\;0.07]$.
- The true new output is $F(1.01,0.02) = (0.020402,\; 5.069999)$. The true change is $[0.0204,\;0.0700]$. ✓ Very close.
For $F:\mathbb{R}^n\to\mathbb{R}^m$ with outputs $F_1,\dots,F_m$, the Jacobian matrix is the $m\times n$ matrix of all first partial derivatives:
$$J_F(\mathbf{x}) = \begin{bmatrix} \partial F_1/\partial x_1 & \cdots & \partial F_1/\partial x_n\\ \vdots & & \vdots\\ \partial F_m/\partial x_1 & \cdots & \partial F_m/\partial x_n \end{bmatrix},\qquad (J_F)_{ij}=\frac{\partial F_i}{\partial x_j}.$$Row $i$ is $\nabla F_i^\top$ (the gradient of output $i$, laid down as a row). Column $j$ says how all outputs respond to input $j$. For small changes,
$$F(\mathbf{x}+\boldsymbol{\delta}) \;\approx\; F(\mathbf{x}) + J_F(\mathbf{x})\,\boldsymbol{\delta}.$$- If $F(\mathbf{x}) = A\mathbf{x}$ is already linear, then $J_F = A$ everywhere: the best linear approximation of a linear map is itself.
- If $m=1$ (one output), the Jacobian is the single row $\nabla f^\top$.
- If $m=n$, $\det J$ is the factor by which the map locally stretches areas (Chapter 1.7). For polar coordinates $(r,\theta)\mapsto(r\cos\theta,\,r\sin\theta)$, $\det J = r$, the familiar "$r\,dr\,d\theta$".
Why do we need it?
Layers of a network output vectors, not single numbers. We need one object that says how every output changes with every input, so that sensitivity can be passed through a layer.
Where is it used?
Backpropagation through every layer, normalising flows (they use the determinant of the Jacobian), Gauss–Newton fitting of nonlinear curves, robot-arm kinematics, and measuring how sensitive a network is to small input changes.
How is it used?
Stack the gradients of the outputs as rows (an m×n matrix). To predict a small change use F(x + δ) ≈ F(x) + Jδ. In code, use autograd (for example torch.autograd.functional.jacobian) and check the shape.
- Shape: the Jacobian of a map from $\mathbb{R}^n$ to $\mathbb{R}^m$ is $m\times n$: rows = outputs, columns = inputs. A map from 3 inputs to 2 outputs has a $2\times3$ Jacobian. If you get the shape wrong, the matrix products later will not fit.
- The Jacobian depends on the point. It is a different matrix at every $\mathbf{x}$ (unless $F$ is linear).
- "Best linear approximation" only holds for small steps, as the widget shows at large $h$.
Quick check: $F(x,y)=(x+y,\;xy,\;x^2)$. What is the shape of $J_F$, and what is $J_F(2,3)$?
3 outputs and 2 inputs, so $J_F$ is $3\times2$. Rows are the gradients: $[1,1]$, $[y,x]$, $[2x,0]$. At $(2,3)$: $J=\begin{bmatrix}1&1\\3&2\\4&0\end{bmatrix}$.
The Hessian: curvature core
The gradient tells you the slope. The Hessian tells you how the slope itself is changing: the curvature.
In one variable: if $f''\gt0$ the curve is shaped like a smile (a valley floor), if $f''\lt0$ it is a frown (a hilltop). With several variables, the ground can curve up in one direction and down in another, like a mountain pass (a saddle). So curvature depends on direction, and we need a whole table of second derivatives to describe it.
Why care? At a flat spot (gradient zero), curvature decides what you are standing on. If the ground curves up in every direction, you are at the bottom of a bowl: a local minimum.
- $f(x,y)=x^2y$. First derivatives: $f_x=2xy$, $f_y=x^2$. Second derivatives: $f_{xx}=2y$, $f_{xy}=2x$ (differentiate $f_x$ by $y$), $f_{yx}=2x$ (differentiate $f_y$ by $x$), $f_{yy}=0$. At $(1,2)$: $H=\begin{bmatrix}4&2\\2&0\end{bmatrix}$. Note it is symmetric: $f_{xy}=f_{yx}$.
- $f(x,y)=x^2+xy+2y^2$ has $\nabla f=[2x+y,\;x+4y]$, which is $\mathbf{0}$ only at the origin. The Hessian is $H=\begin{bmatrix}2&1\\1&4\end{bmatrix}$ everywhere. Its determinant is $8-1=7>0$ and $2>0$, so it is positive definite (Chapter 1.12). So the origin is a local minimum (the lowest point of the bowl).
- $f(x,y)=x^2-y^2$ has $H=\begin{bmatrix}2&0\\0&-2\end{bmatrix}$: up in $x$, down in $y$. A saddle.
For $f:\mathbb{R}^n\to\mathbb{R}$, the Hessian is the $n\times n$ matrix of second partial derivatives:
$$H_f(\mathbf{x}) = \nabla^2 f(\mathbf{x}),\qquad H_{ij} = \frac{\partial^2 f}{\partial x_i\,\partial x_j}.$$It is the Jacobian of the gradient. If the second derivatives are continuous (almost always true), then $H_{ij}=H_{ji}$: the Hessian is symmetric. The curvature along a unit direction $\mathbf{u}$ is $\mathbf{u}^\top H\mathbf{u}$.
Second-order Taylor expansion (keep the slope and the curvature):
$$f(\mathbf{x}+\boldsymbol{\delta}) \;\approx\; f(\mathbf{x}) + \nabla f(\mathbf{x})^\top\boldsymbol{\delta} + \tfrac12\,\boldsymbol{\delta}^\top H(\mathbf{x})\,\boldsymbol{\delta}.$$Second-derivative test. At a critical point ($\nabla f=\mathbf{0}$), look at the eigenvalues of $H$ (Chapter 1.11):
- All eigenvalues $\gt0$ ($H$ positive definite) $\Rightarrow$ local minimum.
- All eigenvalues $\lt0$ (negative definite) $\Rightarrow$ local maximum.
- Mixed signs $\Rightarrow$ saddle point.
- Some eigenvalue $=0$ $\Rightarrow$ the test is inconclusive (flat in some direction).
The eigenvectors are the directions of greatest and least curvature, and the eigenvalues are how strongly the ground curves along them. For a quadratic $f(\mathbf{x})=\tfrac12\mathbf{x}^\top A\mathbf{x}-\mathbf{b}^\top\mathbf{x}$ with $A$ symmetric, the Hessian is the constant matrix $A$, and the Taylor expansion is exact.
Why do we need it?
The gradient only gives the slope. To know whether a flat spot is a valley, a hilltop or a saddle, and how far we can safely step, we need the curvature.
Where is it used?
Newton and quasi-Newton optimisers (such as L-BFGS), the study of saddle points in deep learning, Laplace approximations that give uncertainty in Bayesian models, and learning-rate limits.
How is it used?
At a point with zero gradient, compute H and its eigenvalues: all positive means a minimum, mixed signs a saddle. The ratio of the largest to the smallest eigenvalue tells how hard the bowl is for gradient descent.
- A zero gradient is not enough to say "minimum". It could be a maximum or a saddle. The Hessian decides.
- "Positive definite" is about all directions. One negative eigenvalue (a downhill direction) is enough to destroy a minimum.
- The Hessian of $f:\mathbb{R}^n\to\mathbb{R}$ has $n^2$ entries. With millions of weights it is far too big to store, so deep learning mostly avoids it and uses first-order methods.
Quick check: at a critical point the Hessian is $\begin{bmatrix}3&0\\0&-1\end{bmatrix}$. What is it?
The eigenvalues are $3$ and $-1$: mixed signs. It is a saddle point: the ground curves up along the first axis and down along the second. It is not a minimum.
Essential identities core
Nobody re-derives every gradient from scratch. In one variable you remember a few rules: $(ax)'=a$ and $(x^2)'=2x$. Matrix calculus has the same rules in vector costume.
- A linear function $\mathbf{a}^\top\mathbf{x}$ (a weighted sum) has the same slope $\mathbf{a}$ everywhere, like $ax$ has slope $a$.
- A quadratic $\mathbf{x}^\top A\mathbf{x}$ is the vector version of $ax^2$. Its slope is "$2A\mathbf{x}$", like $2ax$, with a small twist when $A$ is not symmetric.
- A squared distance $\|A\mathbf{x}-\mathbf{b}\|^2$ is "(inside)$^2$", so by the chain rule: $2\times$(inside)$\times$(slope of the inside).
These three cover almost all of classical machine learning (linear and ridge regression, quadratic forms in PCA and Gaussians). A few matrix identities (trace, log-determinant, inverse) are worth knowing exist.
Everything with two unknowns, so you can check by hand.
- Linear. $\mathbf{a}=[3,-1]$: $\mathbf{a}^\top\mathbf{x} = 3x_1 - x_2$. Its partial derivatives are $3$ and $-1$, which is $\mathbf{a}$. ✓
- Quadratic. $A=\begin{bmatrix}1&2\\0&3\end{bmatrix}$: $\mathbf{x}^\top A\mathbf{x} = x_1(x_1+2x_2) + x_2(3x_2) = x_1^2+2x_1x_2+3x_2^2$. Partials: $2x_1+2x_2$ and $2x_1+6x_2$. And $(A+A^\top)\mathbf{x} = \begin{bmatrix}2&2\\2&6\end{bmatrix}\mathbf{x} = [2x_1+2x_2,\;2x_1+6x_2]$. ✓ (The naive "$2A\mathbf{x}$" would give $[2x_1+4x_2,\;6x_2]$, which is wrong because $A$ is not symmetric.)
- Squared distance. $A=\begin{bmatrix}1&0\\0&2\end{bmatrix}$, $\mathbf{b}=[1,1]$, at $\mathbf{x}=[1,1]$. The inside is $A\mathbf{x}-\mathbf{b}=[0,1]$, so $2A^\top(A\mathbf{x}-\mathbf{b}) = 2\,[0,\;2] = [0,4]$. Direct check: $f=(x_1-1)^2+(2x_2-1)^2$ gives $\partial f/\partial x_1 = 2(x_1-1)=0$ and $\partial f/\partial x_2 = 2(2x_2-1)\cdot2=4$. ✓
Here $\mathbf{x}\in\mathbb{R}^n$, $\mathbf{a},\mathbf{b}$ are constant vectors and $A$ is a constant matrix.
| Function | Gradient | One-variable cousin |
|---|---|---|
| $\mathbf{a}^\top\mathbf{x}$ | $\mathbf{a}$ | $(ax)' = a$ |
| $\mathbf{x}^\top A\mathbf{x}$ | $(A+A^\top)\mathbf{x}$, which is $2A\mathbf{x}$ if $A=A^\top$ | $(ax^2)'=2ax$ |
| $\|\mathbf{x}\|^2=\mathbf{x}^\top\mathbf{x}$ | $2\mathbf{x}$ | $(x^2)'=2x$ |
| $\|A\mathbf{x}-\mathbf{b}\|^2$ | $2A^\top(A\mathbf{x}-\mathbf{b})$ | $((ax-b)^2)'=2a(ax-b)$ |
Why they hold (index notation). (1) $\mathbf{a}^\top\mathbf{x}=\sum_i a_ix_i$, so $\partial/\partial x_k$ leaves only $a_k$. (2) $\mathbf{x}^\top A\mathbf{x}=\sum_{i,j}A_{ij}x_ix_j$. The variable $x_k$ appears as $x_i$ (when $i=k$) and as $x_j$ (when $j=k$), giving $\sum_jA_{kj}x_j+\sum_iA_{ik}x_i=(A\mathbf{x})_k+(A^\top\mathbf{x})_k$. (3) Let $\mathbf{r}=A\mathbf{x}-\mathbf{b}$, so $f=\mathbf{r}^\top\mathbf{r}$. The gradient of $f$ with respect to $\mathbf{r}$ is $2\mathbf{r}$, and the Jacobian of $\mathbf{r}$ with respect to $\mathbf{x}$ is $A$. The chain rule (in its vector form, below) gives $\nabla_{\mathbf{x}}f = A^\top(2\mathbf{r})$.
Awareness: matrix-variable identities. For a matrix variable $W$, treat each entry as an input:
- $\nabla_W\operatorname{tr}(W^\top A)=A$, because $\operatorname{tr}(W^\top A)=\sum_{i,j}W_{ij}A_{ij}$ is a weighted sum of the entries.
- $\nabla_X\log\det X = X^{-\top}$ (the inverse transposed), for $\det X>0$. Appears in the Gaussian log-likelihood.
- $d(X^{-1}) = -X^{-1}\,(dX)\,X^{-1}$. In one variable this is $(1/x)'=-1/x^2$.
Why do we need it?
Deriving every gradient from scratch is slow and easy to get wrong. A small toolbox of rules gives the gradient of most classic losses in one line.
Where is it used?
Linear and ridge regression, PCA and Gaussian models (quadratic forms and log-determinants), weight decay, and writing your own loss function in PyTorch or JAX.
How is it used?
Match your loss to a rule, write the gradient, check that the shapes fit, then run a finite-difference check (nudge each entry by about 1e-6). If both agree to many digits, your formula is right.
- $\nabla(\mathbf{x}^\top A\mathbf{x})=2A\mathbf{x}$ only if $A$ is symmetric. In general it is $(A+A^\top)\mathbf{x}$. (A quadratic form only "sees" the symmetric part of $A$, as you saw in Chapter 1.12.)
- Always check shapes. $\nabla_{\mathbf{x}}$ of a scalar has the shape of $\mathbf{x}$. In $2A^\top(A\mathbf{x}-\mathbf{b})$: $A^\top$ is $n\times m$ and the residual is $m\times1$, giving $n\times1$. ✓
- Numerical gradients are a test, not a method for training: they cost two function evaluations per input. Use them to catch mistakes in a hand-derived gradient ("gradient checking").
Quick check: what is $\nabla_{\mathbf{x}}\,\|\mathbf{x}-\mathbf{c}\|^2$?
Write it as $\|A\mathbf{x}-\mathbf{b}\|^2$ with $A=I$ and $\mathbf{b}=\mathbf{c}$. Then the gradient is $2I^\top(\mathbf{x}-\mathbf{c}) = 2(\mathbf{x}-\mathbf{c})$: it points away from $\mathbf{c}$, and its length is twice the distance to $\mathbf{c}$.
The chain rule in vector form, and computational graphs core
A neural network is a long chain of simple steps: multiply by a matrix, add a bias, apply a ReLU, multiply again, ... and finally measure a loss. Each step is a smooth map, and (from the Jacobian section) each smooth map looks like a matrix when you zoom in. Chain the steps and you chain the matrices. So the derivative of the whole chain is the product of the step Jacobians. The gears from the first section, now with matrices.
To organise the bookkeeping, draw the computation as a computational graph: every small operation (a multiply, an add, a ReLU) is a node, and arrows show which results feed which. To get derivatives, start at the loss and walk backwards through the graph, multiplying by each node's tiny local derivative. If one node feeds into two places, the gradients flowing back add up.
A tiny "neuron": two inputs, two weights and a bias $b$ (a constant that is added), a ReLU, then a squared error against a target $t$. The ReLU is $\max(0,s)$: it keeps a positive number and turns a negative one into $0$.
$$p_1=w_1x_1,\quad p_2=w_2x_2,\quad s=p_1+p_2+b,\quad a=\max(0,s),\quad r=a-t,\quad L=r^2.$$- Forward with $x_1=2$, $w_1=0.5$, $x_2=-1$, $w_2=-1$, $b=0.5$, $t=1$: $p_1=1$, $p_2=1$, $s=2.5$, $a=2.5$, $r=1.5$, $L=2.25$.
- Local derivatives: $\partial L/\partial r=2r=3$, $\partial r/\partial a=1$, $\partial a/\partial s=1$ (because $s>0$), $\partial s/\partial p_1=1$, $\partial p_1/\partial w_1=x_1=2$.
- Multiply along the path from $L$ back to $w_1$: $\dfrac{\partial L}{\partial w_1}=3\cdot1\cdot1\cdot1\cdot2=6$.
- Similarly $\partial L/\partial w_2 = 3\cdot x_2=-3$, $\;\partial L/\partial b=3$, $\;\partial L/\partial x_1=3\cdot w_1=1.5$.
- If $s$ had been negative, the ReLU slope would be $0$ and every gradient behind it would be $0$: a "dead" ReLU passes no learning signal.
The widget below does exactly this, one step at a time, with every number visible.
Vector chain rule. If $\mathbf{y}=F(\mathbf{x})$ and $\mathbf{z}=G(\mathbf{y})$, the Jacobian of the composition is the product of the Jacobians, in order:
$$J_{G\circ F}(\mathbf{x}) = J_G\big(F(\mathbf{x})\big)\;J_F(\mathbf{x}).$$Shapes: with $\mathbf{x}\in\mathbb{R}^n$, $\mathbf{y}\in\mathbb{R}^m$, $\mathbf{z}\in\mathbb{R}^p$ we multiply $(p\times m)(m\times n)=p\times n$. ✓ For a chain of many steps, $J=J_k\cdots J_2J_1$.
For a scalar loss $L(\mathbf{y})$ with $\mathbf{y}=F(\mathbf{x})$, the Jacobian of $L$ is the row $\nabla_{\mathbf{y}}L^\top$, so as a column gradient
$$\boxed{\;\nabla_{\mathbf{x}}L = J_F(\mathbf{x})^\top\,\nabla_{\mathbf{y}}L\;}$$This is backpropagation in one line: to move a gradient back through a step, multiply it by the transpose of that step's Jacobian. The scalar chain rule is the special case $m=n=1$.
Computational graph rules. (i) Local rule: each node knows only its own derivative, e.g. a "multiply" node passes back the other factor, an "add" node passes the gradient unchanged to both inputs, a ReLU passes it unchanged where the input was positive and blocks it elsewhere. (ii) Fan-out rule: if a value is used in several places, the gradient arriving at it is the sum over all the places (multivariable chain rule).
Why do we need it?
A deep network is a long chain of steps. We need a systematic way to get the derivative of the whole chain from the simple derivative of each step.
Where is it used?
Backpropagation in every deep-learning framework (PyTorch loss.backward(), JAX grad, TensorFlow GradientTape), and automatic differentiation in scientific computing and physics simulation.
How is it used?
Run the forward pass and keep the values. Then go backwards, multiplying by each step's local derivative (the transposed Jacobian) and adding where paths merge. Check one weight with a tiny nudge.
- Order matters. The Jacobian product is $J_k\cdots J_1$: the last step is on the left. Matrix multiplication is not commutative.
- Moving a gradient backwards uses the transpose: $\nabla_{\mathbf{x}}L=J^\top\nabla_{\mathbf{y}}L$. Forgetting the transpose is the number-one bug in hand-written backprop. Shape checks catch it.
- When a value is used twice (fan-out), gradients are added, not multiplied.
- Backprop needs the numbers from the forward pass (here $x_1$, $w_1$, $s$, $r$, ...), so frameworks store the intermediate values. That is where training memory goes.
Quick check: in the example, suppose $b=-3$ so that $s=-1$. What is $\partial L/\partial w_1$?
With $s=-1\lt0$, the ReLU output is $a=0$ and its slope is $0$. So $\partial L/\partial s=0$, and by the chain rule every gradient behind it ($\partial L/\partial w_1$, $\partial L/\partial b$, ...) is $0$ as well. Changing $w_1$ a little changes nothing, because the ReLU is switched off.
Forward mode versus reverse mode
The Jacobian of a chain is a product like $J_3J_2J_1$. Matrix multiplication is associative: you may group the product any way you like, $(J_3J_2)J_1$ or $J_3(J_2J_1)$, and get the same answer. But the cost can be wildly different.
- Forward mode starts at the input end and works rightwards: $J_3(J_2J_1)$. One pass pushes one input direction through the whole chain. You need $n$ passes (one per input) to get the full Jacobian.
- Reverse mode starts at the output end and works leftwards: $(J_3J_2)J_1$. One pass pulls one output back through the chain. You need $m$ passes (one per output).
Think of a river delta: to find how every source contributes to one sea, trace backwards from the sea (reverse). To find where one source ends up, trace forwards. A network has millions of inputs (weights) and one output (the loss), so reverse mode needs just one backward pass for all gradients. That is why backpropagation is reverse mode.
A chain with $n=1000$ inputs, hidden size $h=1000$ and a single output ($m=1$). The Jacobians are $J_1$: $h\times n$, $J_2$: $h\times h$, $J_3$: $m\times h$. Count multiplications (an $a\times b$ times $b\times c$ product costs $a\,b\,c$):
- Forward grouping $J_3(J_2J_1)$: first $J_2J_1$ costs $h\cdot h\cdot n=10^9$, then $J_3(\cdots)$ costs $m\cdot h\cdot n=10^6$. Total $\approx1.001\times10^9$.
- Reverse grouping $(J_3J_2)J_1$: first $J_3J_2$ costs $m\cdot h\cdot h=10^6$, then $(\cdots)J_1$ costs $m\cdot h\cdot n=10^6$. Total $2\times10^6$.
- Reverse is about 500 times cheaper, because the left-most matrix is a single row ($m=1$), so every product stays a thin row.
For a function $F:\mathbb{R}^n\to\mathbb{R}^m$ built from simple steps:
- Forward mode computes Jacobian–vector products $J\mathbf{v}$ (a directional derivative: how outputs change if inputs move along $\mathbf{v}$). Cost per product: about the cost of evaluating $F$. Full Jacobian: $n$ passes. Best when $n\ll m$.
- Reverse mode computes vector–Jacobian products $\mathbf{u}^\top J$ (how a chosen combination of outputs responds to each input). Cost per product: a small multiple of the cost of $F$. Full Jacobian: $m$ passes. Best when $m\ll n$, e.g. a scalar loss ($m=1$), where one pass gives the entire gradient. The price is memory: it must remember the forward values.
Both give exact derivatives (up to rounding). That distinguishes automatic differentiation from numerical differentiation (which approximates) and from symbolic differentiation (which manipulates formulas).
Why do we need it?
The same derivative can be computed in two orders with very different cost. Picking the right order is what makes training large models possible.
Where is it used?
Reverse mode in all deep-learning training (one loss, many weights); forward mode in sensitivity analysis, Hessian-vector products and JAX's jvp; gradient checkpointing to save memory.
How is it used?
Count the inputs n and outputs m. If n is much bigger than m (a single loss), use reverse mode (vjp, backward). If n is small and m is large, use forward mode (jvp).
- Reverse mode is not "free": it needs the forward pass first, plus the memory to keep intermediate values. Tricks like gradient checkpointing trade extra computing for less memory.
- Reverse mode gives a gradient for one scalar output per backward pass. To get a full Jacobian of a vector output you need one pass per output.
- "Automatic differentiation" is not the same as numerical differentiation (finite differences). It is exact, and costs a small constant times one function evaluation, not $2n$ evaluations.
Quick check: a model has 5 inputs and 2,000 outputs. Which mode needs fewer passes for the full Jacobian?
Forward mode needs one pass per input: 5. Reverse mode needs one per output: 2,000. So forward mode is the better choice here.
Backprop through a linear layer core
A linear layer computes $\mathbf{y}=W\mathbf{x}$ (plus a bias). Suppose a gradient $\mathbf{g}=\partial L/\partial\mathbf{y}$ arrives from the layers after it. Entry $g_i$ says: "the loss would change by $g_i$ per unit of output $i$".
Now ask how the loss depends on one weight $W_{ij}$. That weight connects input $j$ to output $i$. It scales input $x_j$ and feeds the result into output $i$. So its influence on the loss is (how much the loss cares about output $i$) times (how big input $j$ was): $g_i\cdot x_j$. Do this for every pair $(i,j)$ and you get a table of products of two lists of numbers: an outer product, $\mathbf{g}\mathbf{x}^\top$.
And for the input: $x_j$ pushes on every output $i$ with strength $W_{ij}$, so it receives $\sum_i W_{ij}g_i$ in total. That is $W^\top\mathbf{g}$. The transpose carries the gradient backwards along the same wires.
$W=\begin{bmatrix}1&0&2\\0&1&-1\end{bmatrix}$, input $\mathbf{x}=[1,2,3]$, target $\mathbf{t}=[3,0]$, loss $L=\tfrac12\|\mathbf{y}-\mathbf{t}\|^2$.
- Forward: $\mathbf{y}=W\mathbf{x}=[1\cdot1+0+2\cdot3,\;0+2-3]=[7,-1]$. So $L=\tfrac12(4^2+(-1)^2)=8.5$.
- The upstream gradient is $\mathbf{g}=\partial L/\partial\mathbf{y}=\mathbf{y}-\mathbf{t}=[4,-1]$.
- Weight gradient: $\dfrac{\partial L}{\partial W}=\mathbf{g}\mathbf{x}^\top=\begin{bmatrix}4\\-1\end{bmatrix}\begin{bmatrix}1&2&3\end{bmatrix}=\begin{bmatrix}4&8&12\\-1&-2&-3\end{bmatrix}$. (Same shape as $W$, $2\times3$.)
- Input gradient: $W^\top\mathbf{g}=[1\cdot4+0,\;0+1\cdot(-1),\;2\cdot4+(-1)(-1)]=[4,-1,9]$.
- Check one entry: raise $W_{11}$ by $0.01$. Then $y_1=7.01$, so $L=\tfrac12(4.01^2+1)=8.54005$, a rise of $0.04005\approx4\times0.01$. ✓
For a linear layer $\mathbf{y}=W\mathbf{x}+\mathbf{b}$ with $W\in\mathbb{R}^{m\times n}$, and any downstream scalar loss $L$, write $\mathbf{g}=\partial L/\partial\mathbf{y}\in\mathbb{R}^m$. Then
$$\boxed{\;\frac{\partial L}{\partial W}=\mathbf{g}\,\mathbf{x}^\top\;}\in\mathbb{R}^{m\times n},\qquad \frac{\partial L}{\partial\mathbf{x}}=W^\top\mathbf{g}\in\mathbb{R}^n,\qquad \frac{\partial L}{\partial\mathbf{b}}=\mathbf{g}\in\mathbb{R}^m.$$Derivation. $L$ depends on $W_{ij}$ only through $y_i=\sum_jW_{ij}x_j+b_i$, and $\partial y_i/\partial W_{ij}=x_j$. By the chain rule $\partial L/\partial W_{ij}=g_i\,x_j$. Similarly $\partial y_i/\partial x_j=W_{ij}$, and $x_j$ affects all outputs, so $\partial L/\partial x_j=\sum_ig_iW_{ij}=(W^\top\mathbf{g})_j$.
Batches. If the examples are the columns of $X\in\mathbb{R}^{n\times B}$ and the upstream gradients the columns of $G$, then $\partial L/\partial W=G\,X^\top$, the sum of one outer product per example. And if an activation follows, e.g. $\mathbf{a}=\operatorname{ReLU}(\mathbf{z})$, the gradient passes through entry by entry: $\partial L/\partial\mathbf{z}=\partial L/\partial\mathbf{a}\odot\mathbf{1}[\mathbf{z}\gt0]$. Here $\odot$ means multiply entry by entry, and $\mathbf{1}[\mathbf{z}\gt0]$ is a list with 1 where $z_i$ is positive and 0 elsewhere.
Three ready-made facts: the gradient of a weight matrix always has the same shape as the matrix; the layer's backward pass is a multiplication by $W^\top$ (compare the forward pass, a multiplication by $W$); and the weight gradient is rank 1 per example (an outer product).
Why do we need it?
Linear layers make up most of a network. Their backward formulas tell us how to update the weights and what gradient to send further back.
Where is it used?
Dense layers, the Q, K and V projections of attention, convolutions, embeddings, and low-rank LoRA updates, in every Transformer and CNN.
How is it used?
Keep the layer input x, receive the upstream gradient g, then compute dW = g xᵀ (or G Xᵀ for a batch) and dx = Wᵀ g. Check that dW has the same shape as W.
- Shapes tell you the formula. $W$ is $m\times n$, so $\partial L/\partial W$ must be $m\times n$. The only way to build that from $\mathbf{g}$ ($m$) and $\mathbf{x}$ ($n$) is $\mathbf{g}\mathbf{x}^\top$. And $\partial L/\partial\mathbf{x}$ must have $n$ entries: $W^\top\mathbf{g}$ ($n\times m$ times $m$). Mixing these up produces shape errors immediately.
- With batches, sum (or average) over examples: $GX^\top$ already does that.
- Some books write the weights as $\mathbf{x}^\top W$ (a row-vector convention). Then the formulas appear transposed. Same maths, different layout.
Quick check: $W=\begin{bmatrix}1&2\\3&4\end{bmatrix}$, $\mathbf{x}=[1,1]$ and the upstream gradient is $\mathbf{g}=[1,-1]$. Find $\partial L/\partial W$ and $\partial L/\partial\mathbf{x}$.
$\partial L/\partial W=\mathbf{g}\mathbf{x}^\top=\begin{bmatrix}1&1\\-1&-1\end{bmatrix}$. $\partial L/\partial\mathbf{x}=W^\top\mathbf{g}=\begin{bmatrix}1&3\\2&4\end{bmatrix}\begin{bmatrix}1\\-1\end{bmatrix}=[1-3,\;2-4]=[-2,-2]$.
Closing the loop: the gradient of linear regression core
In Chapter 1.10 we found the bottom of the "error bowl" in one shot, by asking the gradient to be zero and solving the normal equations. There is another way to reach the bottom of a bowl: walk downhill. Measure the gradient where you stand, take a small step the opposite way, and repeat. This is gradient descent. Both roads lead to the same $\hat{\mathbf{x}}$.
The Hessian tells you how the bowl is shaped. A round bowl is easy: any step heads straight to the bottom. A long, narrow valley is hard: the gradient points mostly across the valley rather than along it, so you zigzag. And a third option, Newton's method, uses the Hessian to jump to the bottom of the bowl in one go.
The running example from Chapter 1.10: $A=\begin{bmatrix}1&1\\1&2\\1&3\end{bmatrix}$, $\mathbf{b}=[1,2,2]$, error $E(\mathbf{x})=\|A\mathbf{x}-\mathbf{b}\|^2$. Derive the gradient with the tools of this chapter.
- Name the inside: $\mathbf{r}=A\mathbf{x}-\mathbf{b}$, so $E=\mathbf{r}^\top\mathbf{r}$. Then $\partial E/\partial\mathbf{r}=2\mathbf{r}$ and the Jacobian of $\mathbf{r}$ with respect to $\mathbf{x}$ is $A$.
- Chain rule (transposed Jacobian times upstream gradient): $\nabla E = A^\top(2\mathbf{r}) = 2A^\top(A\mathbf{x}-\mathbf{b})$.
- Start at $\mathbf{x}=[0,0]$: $\mathbf{r}=-\mathbf{b}=[-1,-2,-2]$, $A^\top\mathbf{r}=[-5,-11]$, so $\nabla E=[-10,-22]$ and $E=9$.
- One gradient-descent step with $\eta=0.03$: $\mathbf{x}\leftarrow[0,0]-0.03\,[-10,-22]=[0.3,\,0.66]$. The new error: residual $[0.3+0.66-1,\;0.3+1.32-2,\;0.3+1.98-2]=[-0.04,-0.38,0.28]$, so $E=0.0016+0.1444+0.0784=0.2244$. Big drop from 9.
- The target $\hat{\mathbf{x}}=[\tfrac23,\tfrac12]\approx[0.667,0.5]$. We are near it, but not there yet.
For least squares $E(\mathbf{x})=\|A\mathbf{x}-\mathbf{b}\|^2$ (for the mean squared error, divide everything by the number of examples):
$$\nabla E(\mathbf{x}) = 2A^\top(A\mathbf{x}-\mathbf{b}),\qquad \nabla^2E = 2A^\top A\quad(\text{constant}).$$Setting $\nabla E=\mathbf{0}$ gives $A^\top A\mathbf{x}=A^\top\mathbf{b}$: the normal equations (Chapter 1.10, derivation 2, now fully justified). The Hessian $2A^\top A$ is positive semi-definite, and positive definite when $A$ has independent columns (Chapter 1.12), so the flat spot is the unique minimum.
Gradient descent: $\mathbf{x}_{k+1}=\mathbf{x}_k-\eta\,\nabla E(\mathbf{x}_k)=\mathbf{x}_k-2\eta A^\top(A\mathbf{x}_k-\mathbf{b})$. The error vector $\mathbf{e}_k=\mathbf{x}_k-\hat{\mathbf{x}}$ obeys $\mathbf{e}_{k+1}=(I-2\eta A^\top A)\,\mathbf{e}_k$. Along an eigen-direction of $A^\top A$ with eigenvalue $\mu$ it shrinks by the factor $|1-2\eta\mu|$. So:
- It converges if and only if $\eta<\dfrac{1}{\mu_{\max}}$ (where $\mu_{\max}$ is the largest eigenvalue of $A^\top A$, i.e. $\eta<2/\lambda_{\max}(H)$). Larger steps overshoot and diverge.
- The slowest direction (smallest $\mu$) takes about $\kappa(A^\top A)=\kappa(A)^2$ steps. Remember the squared condition number from Chapter 1.10? It returns as the speed limit.
Newton's method: $\mathbf{x}_{k+1}=\mathbf{x}_k-H^{-1}\nabla E=\mathbf{x}_k-(A^\top A)^{-1}A^\top(A\mathbf{x}_k-\mathbf{b})=(A^\top A)^{-1}A^\top\mathbf{b}=\hat{\mathbf{x}}$. For a quadratic error it reaches the answer in one step, from anywhere.
Why do we need it?
We need to connect two big ideas: the closed-form least-squares answer, and the step-by-step training used for models that have no closed form.
Where is it used?
Training linear and logistic regression, softmax cross-entropy (gradient p − y), Newton's method and IRLS, and choosing the learning rate in every optimiser.
How is it used?
Write the loss, derive ∇ = 2Aᵀ(Ax − b), check it numerically, then loop x ← x − η∇. Pick η below 2/λmax(H), and scale the features so the bowl is round.
- Learning rate: too small means slow, too large means divergence. The safe limit depends on the largest curvature, while the speed depends on the smallest. Badly scaled features stretch the bowl and make this worse, which is why people normalise features.
- The "$2$" in $2A^\top(A\mathbf{x}-\mathbf{b})$ is often absorbed: for $L=\tfrac12\|A\mathbf{x}-\mathbf{b}\|^2$ the gradient is $A^\top(A\mathbf{x}-\mathbf{b})$, and for the mean squared error $\tfrac1m\|\cdot\|^2$ it is $\tfrac2mA^\top(\cdot)$. Check which loss a formula refers to.
- Newton's method needs $H^{-1}$: for a model with $n$ weights that costs about $n^3$ operations, which is why it is not used for big neural networks.
Checkpoint: derive the gradient of the linear-regression loss $L(\mathbf{w})=\|X\mathbf{w}-\mathbf{y}\|^2$ without looking anything up.
Let $\mathbf{r}=X\mathbf{w}-\mathbf{y}$, so $L=\mathbf{r}^\top\mathbf{r}=\sum_ir_i^2$. Then $\partial L/\partial r_i=2r_i$, so $\partial L/\partial\mathbf{r}=2\mathbf{r}$. Each residual is $r_i=\sum_jX_{ij}w_j-y_i$, so $\partial r_i/\partial w_j=X_{ij}$ (the Jacobian of $\mathbf{r}$ is $X$). Chain rule: $\partial L/\partial w_j=\sum_i2r_iX_{ij}=2(X^\top\mathbf{r})_j$. Hence $\nabla L=2X^\top(X\mathbf{w}-\mathbf{y})$. Setting it to zero gives $X^\top X\mathbf{w}=X^\top\mathbf{y}$, the normal equations.
Recap, cheat sheet and practice
- Derivative = slope = response to a tiny nudge; partial derivative = nudge one input and freeze the rest; the chain rule multiplies slopes of chained steps.
- Gradient $\nabla f\in\mathbb{R}^n$ (for $f:\mathbb{R}^n\to\mathbb{R}$): points to steepest ascent, has length equal to the steepness, and is perpendicular to the level curves. Descent goes along $-\nabla f$.
- Jacobian $J\in\mathbb{R}^{m\times n}$ (rows = outputs): the best local linear approximation, $F(\mathbf{x}+\boldsymbol{\delta})\approx F(\mathbf{x})+J\boldsymbol{\delta}$.
- Hessian $H=\nabla^2f$: symmetric matrix of curvature. At a critical point, $H$ positive definite gives a local minimum; mixed signs give a saddle. Second-order Taylor: $f+\nabla f^\top\boldsymbol{\delta}+\tfrac12\boldsymbol{\delta}^\top H\boldsymbol{\delta}$.
- Identities: $\nabla\mathbf{a}^\top\mathbf{x}=\mathbf{a}$, $\nabla\mathbf{x}^\top A\mathbf{x}=(A+A^\top)\mathbf{x}$, $\nabla\|A\mathbf{x}-\mathbf{b}\|^2=2A^\top(A\mathbf{x}-\mathbf{b})$. Check any gradient with finite differences.
- Vector chain rule: Jacobians multiply, and backwards gradients use the transpose: $\nabla_{\mathbf{x}}L=J^\top\nabla_{\mathbf{y}}L$. Reverse mode (backprop) gets a whole gradient of a scalar loss in one backward pass.
- Linear layer: $\partial L/\partial W=\mathbf{g}\mathbf{x}^\top$, $\partial L/\partial\mathbf{x}=W^\top\mathbf{g}$, $\partial L/\partial\mathbf{b}=\mathbf{g}$.
- Linear regression: $\nabla E=2A^\top(A\mathbf{x}-\mathbf{b})$ gives the normal equations at zero; gradient descent walks to the same point, at a speed set by $\kappa(A)^2$; Newton gets there in one step.
Cheat sheet
| Object | Definition | Shape | Remember |
|---|---|---|---|
| Gradient | $\nabla f=[\partial f/\partial x_i]$ | $n\times1$ | steepest ascent, $\perp$ level sets |
| Directional derivative | $\nabla f\cdot\mathbf{u}$ | scalar | $\|\nabla f\|\cos\theta$ |
| Jacobian | $J_{ij}=\partial F_i/\partial x_j$ | $m\times n$ | local linear map |
| Hessian | $H_{ij}=\partial^2f/\partial x_i\partial x_j$ | $n\times n$, symmetric | curvature |
| Taylor (order 2) | $f+\nabla f^\top\boldsymbol{\delta}+\tfrac12\boldsymbol{\delta}^\top H\boldsymbol{\delta}$ | scalar | slope + curvature |
| Chain rule | $J_{G\circ F}=J_GJ_F$; $\nabla_xL=J_F^\top\nabla_yL$ | — | backprop = transpose |
| Linear layer | $\partial L/\partial W=\mathbf{g}\mathbf{x}^\top$ | $m\times n$ | outer product |
| Least squares | $\nabla\|A\mathbf{x}-\mathbf{b}\|^2=2A^\top(A\mathbf{x}-\mathbf{b})$ | $n\times1$ | zero gives normal equations |
| Gradient descent | $\mathbf{x}\leftarrow\mathbf{x}-\eta\nabla f$ | — | need $\eta<2/\lambda_{\max}(H)$ |
| Newton | $\mathbf{x}\leftarrow\mathbf{x}-H^{-1}\nabla f$ | — | one step on a quadratic |
import numpy as np
# A generic finite-difference gradient checker (central differences)
def num_grad(f, x, h=1e-6):
g = np.zeros(x.shape)
for i in range(x.size):
e = np.zeros(x.shape); e.flat[i] = h
g.flat[i] = (f(x + e) - f(x - e)) / (2 * h)
return g
rng = np.random.default_rng(0)
X = rng.normal(size=(20, 3))
y = rng.normal(size=20)
w = rng.normal(size=3)
# 1) Mean squared error: L = mean((Xw - y)^2), grad = (2/m) X^T (Xw - y)
mse = lambda w: np.mean((X @ w - y) ** 2)
grad_mse = lambda w: 2 / len(y) * X.T @ (X @ w - y)
print(np.max(np.abs(grad_mse(w) - num_grad(mse, w)))) # tiny, about 1e-10
# 2) Logistic loss (labels 0/1): grad = (1/m) X^T (sigmoid(Xw) - y)
sigmoid = lambda z: 1 / (1 + np.exp(-z))
yb = (y > 0).astype(float)
def logloss(w):
p = sigmoid(X @ w)
return -np.mean(yb * np.log(p) + (1 - yb) * np.log(1 - p))
grad_log = lambda w: X.T @ (sigmoid(X @ w) - yb) / len(yb)
print(np.max(np.abs(grad_log(w) - num_grad(logloss, w))))
# 3) Softmax cross-entropy for one example with logits z and true class k: grad = softmax(z) - onehot(k)
def softmax(z):
e = np.exp(z - z.max())
return e / e.sum()
z, k = rng.normal(size=4), 2
ce = lambda z: -np.log(softmax(z)[k])
grad_ce = lambda z: softmax(z) - np.eye(4)[k]
print(np.max(np.abs(grad_ce(z) - num_grad(ce, z))))
# 4) Backprop for a 2-layer MLP, by hand. NumPy convention here: examples are the ROWS of X, so formulas look transposed.
# Z1 = X W1 + b1, A1 = relu(Z1), Z2 = A1 W2 + b2, loss = mean((Z2 - y)^2)
W1, b1 = rng.normal(size=(3, 4)), np.zeros(4)
W2, b2 = rng.normal(size=(4, 1)), np.zeros(1)
def forward(W1, b1, W2, b2):
Z1 = X @ W1 + b1
A1 = np.maximum(0, Z1)
Z2 = A1 @ W2 + b2
return Z1, A1, Z2, np.mean((Z2[:, 0] - y) ** 2)
Z1, A1, Z2, loss = forward(W1, b1, W2, b2)
B = len(y)
dZ2 = (2 / B) * (Z2[:, 0] - y)[:, None] # dL/dZ2, shape (B, 1)
dW2 = A1.T @ dZ2 # (4, 1): same shape as W2 (outer products summed over the batch)
db2 = dZ2.sum(axis=0)
dA1 = dZ2 @ W2.T # (B, 4): the gradient goes back through W2 (transpose)
dZ1 = dA1 * (Z1 > 0) # ReLU passes the gradient only where Z1 > 0
dW1 = X.T @ dZ1 # (3, 4)
db1 = dZ1.sum(axis=0)
# check dW1 against finite differences
num_dW1 = num_grad(lambda M: forward(M, b1, W2, b2)[3], W1)
print(np.max(np.abs(dW1 - num_dW1))) # tiny
# (Optional) compare with PyTorch: build the same network with requires_grad=True,
# call loss.backward(), and compare W1.grad with dW1.
# 5) Gradient descent versus Newton on least squares
A = np.array([[1., 1], [1, 2], [1, 3]]); b = np.array([1., 2, 2])
grad = lambda x: 2 * A.T @ (A @ x - b)
H = 2 * A.T @ A
eta = 1 / np.linalg.eigvalsh(H).max() # a safe step size (below 2 / lambda_max)
x = np.array([-0.8, 1.4])
for _ in range(500):
x = x - eta * grad(x)
print(x) # close to [0.6667 0.5] after 500 small steps (about 0.6666 0.5000)
x0 = np.array([-0.8, 1.4])
print(x0 - np.linalg.solve(H, grad(x0))) # Newton: exactly [0.6667 0.5] in one step
1. The gradient $\nabla f$ at a point…
2. For a non-symmetric matrix $A$, what is $\nabla_{\mathbf{x}}(\mathbf{x}^\top A\mathbf{x})$?
3. $F:\mathbb{R}^3\to\mathbb{R}^2$. What is the shape of its Jacobian?
4. At a point where $\nabla f=\mathbf{0}$, the Hessian is positive definite. What is the point?
5. In a linear layer $\mathbf{y}=W\mathbf{x}$ with upstream gradient $\mathbf{g}=\partial L/\partial\mathbf{y}$, the gradient with respect to $W$ is…
6. Why do deep-learning frameworks use reverse-mode differentiation?
Practice problems
A. Find $\nabla f$ for $f(x,y)=x^2y+3y$ at $(1,2)$.
$f_x=2xy=4$ and $f_y=x^2+3=4$, so $\nabla f(1,2)=[4,4]$.
B. Linear regression: $A=\begin{bmatrix}1\\2\end{bmatrix}$, $\mathbf{b}=[1,3]$. Compute the gradient of $\|A\mathbf{x}-\mathbf{b}\|^2$ at $x=1$.
$A\mathbf{x}-\mathbf{b}=[1-1,\;2-3]=[0,-1]$. Then $2A^\top(A\mathbf{x}-\mathbf{b})=2(1\cdot0+2\cdot(-1))=-4$. (Direct check: $E=(x-1)^2+(2x-3)^2$, $E'=2(x-1)+4(2x-3)=0+(-4)=-4$ at $x=1$. ✓)
C. Compute the Jacobian of $F(x,y)=(x+y,\;xy,\;x^2)$ at $(2,3)$.
Rows are the gradients: $[1,1]$, $[y,x]=[3,2]$, $[2x,0]=[4,0]$. So $J=\begin{bmatrix}1&1\\3&2\\4&0\end{bmatrix}$ ($3\times2$).
D. Find the Hessian of $f=x^2+xy+2y^2$ and classify its critical point.
$f_x=2x+y$, $f_y=x+4y$, so the only critical point is $(0,0)$. $H=\begin{bmatrix}2&1\\1&4\end{bmatrix}$ has trace $6>0$ and determinant $7>0$, so both eigenvalues are positive (they are $3\pm\sqrt2$). $H$ is positive definite, so $(0,0)$ is a local (and global) minimum.
E. For $W=\begin{bmatrix}1&2\\3&4\end{bmatrix}$, $\mathbf{x}=[2,1]$, $\mathbf{g}=[1,0]$: compute $\partial L/\partial W$ and $\partial L/\partial\mathbf{x}$.
$\partial L/\partial W=\mathbf{g}\mathbf{x}^\top=\begin{bmatrix}1\\0\end{bmatrix}\begin{bmatrix}2&1\end{bmatrix}=\begin{bmatrix}2&1\\0&0\end{bmatrix}$. $\partial L/\partial\mathbf{x}=W^\top\mathbf{g}=\begin{bmatrix}1&3\\2&4\end{bmatrix}\begin{bmatrix}1\\0\end{bmatrix}=[1,2]$.
F. For $E=\|A\mathbf{x}-\mathbf{b}\|^2$, the largest eigenvalue of $A^\top A$ is 16. What is the largest learning rate for which gradient descent converges?
The Hessian is $2A^\top A$ with largest eigenvalue $32$, so we need $\eta<2/32=1/16=0.0625$. (Check: the error along that direction is multiplied by $|1-2\eta\cdot16|=|1-32\eta|$ each step, which is below 1 exactly when $0<\eta<1/16$.)
Numerical Linear Algebra
On paper, maths is exact. In a computer, every number is rounded to fit in a few bits. This chapter shows how that rounding can quietly wreck a correct formula, and the standard tricks that keep ML code honest, fast and stable.
- Understand how computers store real numbers (floating point) and where their limits are
- See rounding errors, cancellation and order effects, and know how big they can get
- Separate a hard problem (conditioning) from a careless method (stability)
- Use the standard ML tricks: log-sum-exp, stable softmax, log space, jitter
- Know when to solve with iterative methods and sparse matrices
- Estimate the cost of linear algebra operations, including attention
How computers store real numbers core
A computer has a fixed number of boxes (bits) for each number. It cannot store "all the digits of π". So it uses the same trick scientists use: scientific notation.
Think of $6.02 \times 10^{23}$. It has three parts: a sign (plus or minus), a few significant digits ($6.02$), and an exponent ($23$) that says where the decimal point floats to. That is why it is called floating point.
Computers do the same, but in base 2. The number of significant digits is fixed. Anything beyond them is rounded away. Imagine a ruler that only has certain marks: a number between two marks is moved to the nearest mark.
Let us store $6.5$ as a 32-bit float (called float32).
- Write 6.5 in binary: $6.5 = 4 + 2 + 0.5 = 110.1_2$.
- Move the point so that one digit is before it: $110.1_2 = 1.101_2 \times 2^{2}$.
- Sign: positive, so the sign bit is $0$.
- Exponent: it is $2$. Stored exponents have a "bias" of 127 added, so we store $2 + 127 = 129 = 10000001_2$.
- Mantissa: the digits after the point: $101$. The leading $1$ is always there, so it is not stored. Pad with zeros: $10100000000000000000000$.
All 32 bits: 0 10000001 10100000000000000000000. Reading it back: $+1 \times 2^{129-127} \times (1 + 0.625) = 4 \times 1.625 = 6.5$ ✓.
The IEEE 754 standard stores a number as three fields: a sign bit $s$, an exponent field $e$ ($E$ bits) and a mantissa (or fraction) field $m$ ($M$ bits). For ordinary ("normal") numbers:
$$x = (-1)^{s} \times 2^{\,e - \text{bias}} \times \left(1 + \frac{m}{2^{M}}\right), \qquad \text{bias} = 2^{E-1} - 1.$$Two special exponent values are reserved: all zeros means zero or a "subnormal" tiny number (no hidden leading 1), and all ones means $\pm\infty$ (mantissa 0) or NaN (mantissa not 0).
| Format | Bits | Sign | Exponent | Mantissa | Machine epsilon | Largest value | Decimal digits |
|---|---|---|---|---|---|---|---|
float64 (double) | 64 | 1 | 11 | 52 | $2.2\times10^{-16}$ | $1.8\times10^{308}$ | about 16 |
float32 (single) | 32 | 1 | 8 | 23 | $1.2\times10^{-7}$ | $3.4\times10^{38}$ | about 7 |
float16 (half) | 16 | 1 | 5 | 10 | $9.8\times10^{-4}$ | $65\,504$ | about 3 |
bfloat16 (brain float) | 16 | 1 | 8 | 7 | $7.8\times10^{-3}$ | $3.4\times10^{38}$ | about 2 |
More exponent bits mean a bigger range (how huge or tiny). More mantissa bits mean more precision (how many correct digits). Notice that bfloat16 keeps float32's range but throws away most of its precision.
Why do we need it?
A computer has only a fixed number of bits per number, yet we need both huge values (like 10^30) and tiny ones (like 10^-30). Floating point is the compromise that makes this possible.
Where is it used?
Every number in NumPy, PyTorch and TensorFlow. Choosing float32, float16 or bfloat16 decides the memory use and speed of mixed-precision training on GPUs and TPUs, and the size of saved model weights.
How is it used?
Pick a format by asking: how many digits do I need (mantissa) and how large or small can values get (exponent)? Use float32 by default, float64 for delicate maths, bfloat16 or float16 for fast training. Check np.finfo(dtype).
Not every number fits. The number $0.1$ has an endless pattern in binary ($0.000110011001100\ldots_2$), just like $1/3 = 0.3333\ldots$ in decimal. So the computer stores a nearby number instead. You can see this in the widget by typing 0.1 and reading the long decimal below it.
Integers are safe, up to a point. Whole numbers are stored exactly as long as they fit in the mantissa ($2^{24}$ for float32, $2^{53}$ for float64). Above that, even whole numbers start to be skipped.
Quick check: which format has more correct digits, float16 or bfloat16? Which has the larger range?
float16 has 10 mantissa bits, so more digits (about 3 against about 2). bfloat16 has 8 exponent bits against 5, so it has the much larger range ($3.4\times10^{38}$ against $65\,504$).
Machine epsilon: the gaps between floats core
Floating-point numbers are like marks on a ruler, but a strange ruler: the marks are packed tightly near zero and drift further apart as you go to bigger numbers. Between $1$ and $2$ there are the same number of marks as between $1024$ and $2048$, so the marks near $1024$ are $1024$ times further apart.
What stays the same is the relative gap: the gap divided by the number. Around every number, about the first 7 digits (in float32) are right, and the rest is noise.
In float32, the number after $1$ is $1 + 2^{-23} \approx 1.0000001192$. So the gap at $1$ is $2^{-23} \approx 1.2\times10^{-7}$.
- Try $1 + 2^{-24}$. It lies exactly halfway between two marks, and ties go to the even one, which is $1$. Result: 1. The addition did nothing.
- Try $1 + 2^{-23}$. This is exactly a mark. Result: $1.0000001192$. It worked.
- At $16\,777\,216 = 2^{24}$ the gap in
float32is already $2$. So $16\,777\,216 + 1 = 16\,777\,216$ again.
In float64 the same thing happens at $1 + 10^{-16}$ and at $2^{53} \approx 9\times10^{15}$.
Machine epsilon $\varepsilon$ is the gap between $1$ and the next number the format can store: $\varepsilon = 2^{-M}$ (so $2.2\times10^{-16}$ for float64, $1.2\times10^{-7}$ for float32).
Rounding to the nearest mark gives a tiny relative error. Writing $\text{fl}(x)$ for "the stored version of $x$":
$$\text{fl}(x) = x\,(1 + \delta), \qquad |\delta| \le u = \tfrac{1}{2}\varepsilon.$$The number $u$ is called the unit roundoff. Near a number $x$, the gap between neighbouring floats is between $\tfrac12 \varepsilon |x|$ and $\varepsilon |x|$. In NumPy: np.finfo(np.float32).eps.
Why do we need it?
We need a way to say how accurate a computed answer can be at best. Machine epsilon is that yardstick: the smallest relative step the format can notice.
Where is it used?
Test tolerances such as np.isclose, convergence checks in optimisers, deciding whether a matrix is numerically singular, and explaining why tiny weight updates vanish in float16 training.
How is it used?
Read epsilon with np.finfo(np.float32).eps. Compare floats with a tolerance a few times epsilon (times the size of the numbers), never with ==. Expect about 7 correct digits in float32 and 16 in float64.
Epsilon is not "the smallest number". $\varepsilon \approx 10^{-16}$ is the gap at 1. The smallest positive float64 is about $10^{-308}$, hugely smaller. Epsilon tells you about relative precision. The smallest number tells you about range.
Never test floats with ==. Compare with a tolerance: $|a - b| \le \text{tol}\cdot\max(|a|, |b|)$. np.isclose does something very similar (it also adds a tiny absolute tolerance, so that numbers near $0$ can still compare equal).
Quick check: roughly how many correct decimal digits does float32 give you for a number near 1000?
About 7. The relative precision is the same at every size ($\approx 10^{-7}$), so $1000.0001$ is stored as a nearby mark ($1000.0000610\ldots$), but $1000.000001$ is stored as plain $1000$: the gap at 1000 is $6.1\times10^{-5}$.
Overflow, underflow, Inf and NaN core
Every format has a biggest and a smallest number it can hold, like a measuring cup with a rim.
- Overflow: the true answer is too big, so the cup overflows. The computer writes Inf ("infinity").
- Underflow: the true answer is too small to tell apart from zero. The computer writes 0. (Just before that, numbers go "subnormal": they lose digits gradually.)
- NaN ("not a number"): the question has no sensible answer, like $0/0$. And NaN spreads: any calculation that touches a NaN becomes NaN.
In float16 the largest number is $65\,504$.
- $300 \times 300 = 90\,000$ is bigger than $65\,504$, so the result is Inf.
- Then $\text{Inf} - \text{Inf}$ has no meaning, so it is NaN.
- Then $\text{NaN} + 5 = \text{NaN}$ and even $\text{NaN} = \text{NaN}$ is false.
In float32 the same story starts at $e^{89} \approx 4.5\times10^{38}$, which is above $3.4\times10^{38}$. In float64 it starts at $e^{710}$.
Special values defined by IEEE 754:
| Expression | Result | Why |
|---|---|---|
| $1/0$, $\log 0$ | $+\infty$, $-\infty$ | the limit as the input shrinks to 0 |
| $0/0$, $\infty - \infty$, $0\times\infty$ | NaN | no single sensible answer |
| $\sqrt{-1}$, $\log(-1)$ | NaN | not a real number |
| NaN with anything | NaN | "not a number" is contagious |
| NaN $=$ NaN | false | test with np.isnan(x) instead |
| too big / too small | Inf / 0 | overflow / underflow |
The normal range of a format is from its smallest normal number up to its largest. Below the smallest normal number come the subnormals, which still work but have fewer and fewer correct digits, until they reach $0$.
Why do we need it?
Real programs hit impossible or too-big results. IEEE 754 defines Inf, 0 and NaN, so the program keeps running with a clear marker instead of crashing.
Where is it used?
NaN losses in neural network training, overflowing exp in softmax, loss scaling in float16 mixed-precision training, and gradient underflow. Debuggers and tools like torch.autograd.set_detect_anomaly look for them.
How is it used?
When a loss turns NaN, check inputs with np.isnan and np.isinf, then look for log(0), 0/0, a huge learning rate or an overflowing exp. Prefer bfloat16 or loss scaling to avoid float16 range problems.
One NaN ruins everything. Once a NaN enters a sum or a matrix product, every result it touches becomes NaN. In training, a single NaN in the loss means every weight turns into NaN after one update. Find the first place it appears.
Silent underflow. Tiny values flush to $0$ without any warning. A probability of $10^{-50}$ is fine in float64 and exactly $0$ in float16.
Quick check: in float32, what is np.exp(100)? And np.exp(-110)?
$e^{100} \approx 2.7\times10^{43}$ is above $3.4\times10^{38}$, so it overflows to Inf. $e^{-110} \approx 1.7\times10^{-48}$ is below the smallest subnormal ($1.4\times10^{-45}$), so it underflows to 0.
Rounding error: why 0.1 + 0.2 is not 0.3 core
Suppose you only have a ruler with marks every millimetre, and you must record every length to the nearest mark. Each measurement is slightly off. Add two measurements and the little errors add as well.
A computer does exactly this after every single operation: it computes the true answer, then moves it to the nearest representable number. Usually the damage is tiny (about the 16th digit in float64). But "tiny" is not "zero", and the famous example is $0.1 + 0.2$.
Why can't we store $0.1$? Convert it to binary by repeatedly doubling the fraction. The bit you take off each time is the next binary digit:
- $0.1 \times 2 = 0.2$, bit 0.
- $0.2 \times 2 = 0.4$, bit 0.
- $0.4 \times 2 = 0.8$, bit 0.
- $0.8 \times 2 = 1.6$, bit 1, keep $0.6$.
- $0.6 \times 2 = 1.2$, bit 1, keep $0.2$.
- We are back at $0.2$, so the bits $0011$ repeat forever: $0.1 = 0.0\,\overline{0011}_2$.
The computer cuts this off after 53 significant bits. The stored value of $0.1$ is slightly bigger than $0.1$ ($0.1000000000000000055\ldots$). The stored $0.2$ is also a bit too big. Their sum rounds to $0.3000000000000000444\ldots$, which is not the same number as the stored version of $0.3$ ($0.2999999999999999889\ldots$). So 0.1 + 0.2 == 0.3 is False.
The standard model of floating-point arithmetic: each basic operation $\circ \in \{+,-,\times,\div\}$ returns the exactly-rounded true result,
$$\text{fl}(a \circ b) = (a \circ b)(1 + \delta), \qquad |\delta| \le u.$$Two consequences matter. First, one operation is almost perfect (relative error at most $u \approx 10^{-16}$). Second, millions of operations can pile up many small errors (see the next sections). The errors are not random noise you can ignore: they follow rules, and good algorithms keep them small.
Why do we need it?
Most decimals such as 0.1 cannot be stored exactly in binary. Knowing this stops you from being surprised, and from writing code that depends on exact equality.
Where is it used?
Unit tests of numerical code, comparing a loop with a vectorised version, financial code (which avoids floats for money), and checking that two implementations of the same model agree.
How is it used?
Never test floats with ==. Use np.isclose(a, b) or np.allclose with rtol and atol. Expect differences near 1e-16 (float64) or 1e-7 (float32), and treat bigger ones as a real bug.
"Wrong" does not mean "random". Run the same code twice and you get the identical error. The difference is tiny: about $5.6\times10^{-17}$ here. What you must not do is if x == 0.3, or build a loop that stops when a float "equals" a target exactly.
Quick check: is 0.5 + 0.25 == 0.75 true in floating point? Why?
True. $0.5 = 2^{-1}$, $0.25 = 2^{-2}$ and $0.75 = 2^{-1} + 2^{-2}$ all have short exact binary forms, so nothing is rounded. Rounding trouble only starts with numbers like $0.1$ that need infinitely many binary digits.
Catastrophic cancellation core
You want the weight of a ship's captain. You weigh the ship with the captain: $50\,000.2$ tonnes. Then without him: $50\,000.1$ tonnes. The difference is $0.1$. But each weighing could be off by $\pm 0.05$. The leftover answer $0.1$ is the same size as the error. Almost all of the information was in the digits that cancelled.
Subtracting two numbers that are almost equal cancels their leading digits. What is left was formed from the last, least-reliable digits. That is a catastrophic cancellation.
Suppose we only keep 7 significant digits.
- $a = 1.234568$ (really $1.2345678\ldots$) and $b = 1.234567$ (really $1.2345671\ldots$).
- Subtract: $a - b = 0.000001 = 1\times10^{-6}$.
- The true difference is $1.2345678 - 1.2345671 = 0.0000007 = 7\times10^{-7}$.
Seven good digits went in. One bad digit came out: the answer is $1$ when it should be $7$. The subtraction itself was exact, but it exposed the earlier rounding.
A real case. $\sqrt{10^8 + 1} - \sqrt{10^8}$ in float64 gives $5.00000005559\times10^{-5}$. The true value is $4.99999998750\times10^{-5}$. Only 8 digits are right out of 16. Rewriting it as $\dfrac{1}{\sqrt{10^8+1} + \sqrt{10^8}}$ (no subtraction) gives the exact digits.
For $x - y$ with $x \approx y$, each input carries a relative error up to $u$. The result's relative error can be as large as
$$\frac{|x - y|_{\text{error}}}{|x-y|} \approx u\,\frac{|x| + |y|}{|x - y|}.$$The closer $x$ and $y$ are, the bigger the blow-up. The subtraction is harmless only when the inputs are exact (such as small integers).
The cure is algebra: rewrite the formula so the subtraction disappears. Standard rewrites: $1 - \cos x = 2\sin^2(x/2)$; $\sqrt{x+1}-\sqrt{x} = \dfrac{1}{\sqrt{x+1}+\sqrt{x}}$; use np.log1p(x) for $\log(1+x)$ and np.expm1(x) for $e^x - 1$ when $x$ is tiny.
Why do we need it?
Subtracting two nearly equal numbers throws away the leading digits and leaves mostly noise. We must spot this trap and rewrite the formula.
Where is it used?
Computing variance as E[x²] minus (E[x])², log(1+x) for tiny x (use np.log1p), the quadratic formula, and the pairwise-distance formula for points that are close together but far from zero.
How is it used?
When you see a subtraction of close values, look for an algebra trick (a conjugate, a half-angle identity) or a special function (log1p, expm1). For variance use the two-pass formula or Welford's method.
Cancellation hides in innocent code. Computing a variance as $E[x^2] - (E[x])^2$ subtracts two big, nearly equal numbers when the mean is large. Use the two-pass formula $\frac1n\sum (x_i-\bar x)^2$ (first find the mean, then sum the squared distances from it), or Welford's method (a one-pass update that avoids the subtraction), instead.
The quadratic formula has the same trap: when $b^2 \gg 4ac$, one of the roots subtracts nearly equal numbers. Compute the other root first, then get this one from the product of roots.
Quick check: how would you compute $\sqrt{x+1} - \sqrt{x}$ for a very large $x$ without losing digits?
Multiply top and bottom by the "conjugate": $\sqrt{x+1} - \sqrt{x} = \dfrac{(x+1) - x}{\sqrt{x+1}+\sqrt{x}} = \dfrac{1}{\sqrt{x+1}+\sqrt{x}}$. Now there is only an addition of positive numbers, which loses nothing.
Order matters: non-associativity and long sums core
In school, $(a + b) + c = a + (b + c)$ always. In floating point, it is not true, because every addition rounds. Where the brackets go changes which roundings happen.
Imagine a huge boulder and a pile of tiny pebbles. The scale is so coarse that one pebble on top of the boulder does not move the needle at all. If you drop the pebbles on the boulder one by one, the scale never changes. But if you first collect all the pebbles into one bag and then put the bag on the boulder, the scale does move, and you get the right answer.
- $(0.1 + 0.2) + 0.3 = 0.6000000000000001$, but $0.1 + (0.2 + 0.3) = 0.6$.
- $(10^{16} + 1) - 10^{16} = 0$, but $10^{16} - 10^{16} + 1 = 1$. The gap at $10^{16}$ is $2$, so $10^{16} + 1$ rounds back to $10^{16}$.
A long sum. Add $0.1$ ten times in float64. You get $0.9999999999999999$, not $1$. Each of the ten additions has a tiny error, and they do not cancel.
In general, a naive sum of $n$ numbers can lose up to about $n\,u$ in relative terms (worst case). For a sum of $10^7$ numbers in float32 this is large: $10^7 \times 6\times10^{-8} = 0.6$.
Floating-point addition is commutative ($a+b = b+a$) but not associative: $(a+b)+c \ne a+(b+c)$ in general. So the order of a long sum changes the answer. Ways to sum more accurately:
- Sort by size and add the small ones first.
- Pairwise (tree) summation: add in pairs, then pairs of pairs. The error grows like $\log n$ instead of $n$. NumPy's
np.sumdoes this. - Kahan (compensated) summation ○: keep a running "lost change" $c$ and feed it back in:
s = 0.0; c = 0.0 # c holds the part that was rounded away
for x in values:
y = x - c # add back what we lost last time
t = s + y # big + small: low digits of y get lost here
c = (t - s) - y # (t - s) is what was actually added; minus y = the loss
s = t
Kahan's method gives an error that does not depend on $n$ (to first order). You do not need to write it yourself, but you should know why library sums beat a plain loop.
Why do we need it?
Float addition rounds every time, so the order of a long sum changes the answer. We need sums that stay accurate, and we need to understand why two runs can differ.
Where is it used?
np.sum (pairwise summation), float32 accumulators in mixed-precision matmul, loss sums over millions of samples, parallel GPU reductions, and Kahan summation in numerical libraries.
How is it used?
Use the library sum, not a hand-written loop. Accumulate in higher precision (float32 or float64) when adding many small numbers. If one number is huge, add the small ones first. Do not expect bit-identical results across GPU runs.
Parallel sums are not reproducible. A GPU adds numbers in whatever order the threads finish. Run the same training twice and the sums can differ in their last digits. After many steps this grows into visibly different models, even with the same random seed. It is not a bug: float addition really is not associative.
Quick check: you must add a million positive numbers of very different sizes. Which order is best: smallest first or largest first?
Smallest first. The running total stays small for as long as possible, so small numbers are added to totals of a similar size and keep their digits. Adding them after the total has become huge would lose them (as in the widget).
Condition number and ill-conditioned problems core
Two roads cross at a junction. If the roads meet at a wide angle, then moving one road by a metre moves the junction by about a metre. But if the two roads run almost parallel, then moving one road by a metre can slide the junction hundreds of metres away.
Solving $A\mathbf{x} = \mathbf{b}$ is finding where lines (or planes) cross. A system whose lines are almost parallel is ill-conditioned: a tiny change in the input $\mathbf{b}$ (or in $A$) makes a huge change in the answer $\mathbf{x}$. And the computer always does make tiny changes, because of rounding.
The key point: this is a property of the question, not of the method. Nothing you do afterwards can repair it, because the input's own uncertainty is already larger than the answer's precision.
Take $A = \begin{bmatrix} 1 & 1 \\ 1 & 1.0001 \end{bmatrix}$ and $\mathbf{b} = \begin{bmatrix} 2 \\ 2.0001 \end{bmatrix}$.
- Solve $x + y = 2$ and $x + 1.0001\,y = 2.0001$. Subtract: $0.0001\,y = 0.0001$, so $y = 1$ and $x = 1$. Solution: $[1, 1]$.
- Now change one entry of $\mathbf{b}$ by only $0.0001$: $\mathbf{b}' = [2,\ 2.0002]$.
- Subtract again: $0.0001\,y = 0.0002$, so $y = 2$ and $x = 0$. New solution: $[0, 2]$.
A change of 0.005% in the input moved the answer by 100%. The singular values of $A$ are about $2.00005$ and $0.00005$, so $\kappa(A) \approx 40\,000$.
The condition number of an invertible matrix (you met it in Chapter 1.7) is the ratio of its largest to its smallest singular value (singular values are explained in Chapter 1.13):
$$\kappa(A) = \frac{\sigma_{\max}}{\sigma_{\min}} = \|A\|\,\|A^{-1}\|.$$It always satisfies $\kappa(A) \ge 1$. For the system $A\mathbf{x}=\mathbf{b}$, a small change $\delta\mathbf{b}$ in the input can change the answer by at most
$$\frac{\|\delta\mathbf{x}\|}{\|\mathbf{x}\|} \;\le\; \kappa(A)\,\frac{\|\delta\mathbf{b}\|}{\|\mathbf{b}\|}.$$Rule of thumb: you lose about $\log_{10}\kappa(A)$ digits. With 16 digits in float64 and $\kappa = 10^{10}$, you keep about 6.
- $\kappa \approx 1$: well-conditioned (orthogonal matrices have exactly $\kappa = 1$, the best possible).
- $\kappa$ large: ill-conditioned. Once $\kappa \approx 1/u \approx 10^{16}$, the matrix is numerically singular: the computer cannot tell it from a singular one.
Why do we need it?
Some problems are touchy: a tiny change in the input changes the answer a lot. The condition number measures that touchiness so we know how much to trust a result.
Where is it used?
Linear regression with collinear features, solving Ax = b, Gaussian processes, deciding if a matrix is numerically singular, understanding why ridge regularisation helps and why optimisers struggle in long narrow valleys.
How is it used?
Compute np.linalg.cond(A). Subtract its log10 from your 16 digits (float64) to guess how many digits of x you can trust. If it is huge, rescale features, regularise, or use a more careful method. Never judge a matrix by its determinant.
Small residual does not mean small error. A good algorithm always finds an $\hat{\mathbf{x}}$ with $A\hat{\mathbf{x}} \approx \mathbf{b}$ (tiny residual). But for an ill-conditioned $A$, many quite different $\mathbf{x}$ give nearly the same $\mathbf{b}$, so a tiny residual tells you little about how close $\hat{\mathbf{x}}$ is to the truth.
The determinant is the wrong test. $\det(A)$ can be tiny for a perfectly healthy matrix (for example $0.1\,I$ in 20 dimensions has determinant $10^{-20}$) and large for an awful one. Use $\kappa(A)$ instead: np.linalg.cond(A).
Quick check: if $\kappa(A) = 10^{6}$ and your data has 16 correct digits, about how many digits of $\mathbf{x}$ can you trust?
About $16 - 6 = 10$ digits, if you use a good algorithm. Conditioning sets the best you can hope for.
Conditioning vs stability: normal equations against QR core
Compare two things that sound alike but are not:
- Conditioning is how bumpy the road is. It belongs to the problem. A mountain pass is dangerous for every driver.
- Stability is how carefully you drive. It belongs to the algorithm. A careful driver on a bumpy road arrives with small errors. A reckless driver on a smooth road can still crash.
A stable algorithm gives an answer as good as the problem allows: its error is about $\kappa\cdot u$. An unstable one is much worse than $\kappa\cdot u$, even though the exact maths is correct.
Least squares (Chapter 1.10) can be solved two ways. Suppose $\kappa(A) = 10^{8}$ and we use float64 ($u \approx 10^{-16}$).
- Normal equations: solve $A^\top A\,\mathbf{x} = A^\top\mathbf{b}$. The matrix $A^\top A$ has $\kappa(A^\top A) = \kappa(A)^2 = 10^{16}$. Error about $10^{16}\times10^{-16} = 1$: no correct digits.
- QR: factor $A = QR$ with $Q$ orthogonal (condition number 1) and solve $R\mathbf{x} = Q^\top\mathbf{b}$. The condition number involved is just $\kappa(A) = 10^8$. Error about $10^{8}\times10^{-16} = 10^{-8}$: 8 correct digits.
Both are correct on paper. The normal equations square the condition number, so they lose twice as many digits.
An algorithm is backward stable if the answer $\hat{\mathbf{x}}$ it returns is the exact answer to a slightly different problem:
$$(A + \Delta A)\,\hat{\mathbf{x}} = \mathbf{b} + \Delta\mathbf{b}, \qquad \frac{\|\Delta A\|}{\|A\|},\ \frac{\|\Delta \mathbf{b}\|}{\|\mathbf{b}\|} = O(u).$$Combine this with conditioning and you get the master rule:
$$\text{error in } \hat{\mathbf{x}} \;\lesssim\; \kappa(A) \times (\text{backward error}) \;\approx\; \kappa(A)\cdot u.$$So: stable algorithm + well-conditioned problem = accurate answer. Stable algorithm + ill-conditioned problem = answer limited by $\kappa$ (not the algorithm's fault). Unstable algorithm = error beyond $\kappa u$ (the algorithm's fault).
Householder QR, LU with pivoting and the SVD are backward stable. Forming $A^\top A$ is an example of a step that makes things worse: it turns a $\kappa$ problem into a $\kappa^2$ one.
Why do we need it?
A correct formula can still give a wrong number if the algorithm magnifies rounding errors. We need to tell a hard problem (conditioning) from a careless method (stability).
Where is it used?
Least squares in np.linalg.lstsq, scikit-learn and statsmodels (QR or SVD, not the normal equations), LAPACK solvers, and any library routine that claims to be backward stable.
How is it used?
For least squares call np.linalg.lstsq or use QR, not inv(A.T @ A). Use np.linalg.solve, not inv(A) @ b. Remember: error is about condition number times epsilon for a stable method, and condition number squared for the normal equations.
"Stable" and "well-conditioned" are not the same word. People say "the matrix is unstable" when they mean "ill-conditioned". Say matrix or problem for conditioning and algorithm for stability.
The normal equations are not always wrong: when $\kappa(A)$ is small (say below $10^3$) they are fast and fine. They are risky when $A$ is nearly rank-deficient.
Quick check: $\kappa(A) = 10^{9}$. Roughly how many digits does each method give in float64?
QR: about $16 - 9 = 7$ digits. Normal equations: $\kappa^2 = 10^{18} > 10^{16}$, so no correct digits at all.
Classical vs modified Gram–Schmidt core
Gram–Schmidt (Chapter 1.9) builds perpendicular vectors one at a time: take the next vector and remove its shadow on every earlier one. Two orderings give the same answer on paper:
- Classical (CGS): measure all the shadows on the original vector, then subtract them all.
- Modified (MGS): subtract the shadow on the first vector. Then measure the next shadow on the leftover, subtract that, and so on.
MGS re-measures after each cleaning step, so rounding errors from earlier steps get caught. CGS never looks back, so errors pile up and the "perpendicular" vectors may end up not perpendicular at all.
The vectors $\mathbf{a}_1 = [1, \varepsilon, 0, 0]$, $\mathbf{a}_2 = [1, 0, \varepsilon, 0]$, $\mathbf{a}_3 = [1, 0, 0, \varepsilon]$ (columns of the "Läuchli matrix") are nearly the same vector when $\varepsilon$ is tiny. Take $\varepsilon = 10^{-8}$.
- Step 1: $\mathbf{q}_1 \approx \mathbf{a}_1$ (its length is $\sqrt{1 + 10^{-16}}$, which the computer rounds to exactly $1$).
- Step 2: $\mathbf{a}_2 - (\mathbf{q}_1\cdot\mathbf{a}_2)\,\mathbf{q}_1 = [0, -\varepsilon, \varepsilon, 0]$. All the "1" parts cancelled. The result is only $10^{-8}$ long, so it is correct only to about 8 digits. (Cancellation!)
- Step 3 (CGS) measures the shadow of the original $\mathbf{a}_3$ on $\mathbf{q}_2$. That shadow is $0$, so nothing is removed. But after removing the part along $\mathbf{q}_1$, the leftover $[0, -\varepsilon, 0, \varepsilon]$ still has a component along $\mathbf{q}_2$ (half as long as the leftover itself), because $\mathbf{q}_2$ is not perfectly perpendicular to $\mathbf{q}_1$. That component is never removed, so $\mathbf{q}_3$ is not perpendicular to $\mathbf{q}_2$: their dot product is about $0.5$.
- MGS measures the second shadow on the already cleaned vector. The dot product stays near $10^{-8}$.
Let $Q$ hold the computed orthonormal vectors as columns. The loss of orthogonality is $\|Q^\top Q - I\|$: it should be about $u$.
- CGS: $\|Q^\top Q - I\| \approx u\,\kappa(A)^2$ (can be $O(1)$: no orthogonality at all).
- MGS: $\|Q^\top Q - I\| \approx u\,\kappa(A)$ (much better, still not perfect).
- Householder QR (what
np.linalg.qruses): $\|Q^\top Q - I\| \approx u$, independent of $\kappa$. This is why libraries use it.
The same pattern as before: reordering the same maths changes the power of $\kappa$ that shows up.
Why do we need it?
Gram–Schmidt builds perpendicular vectors, but rounding can quietly destroy the perpendicularity. We need the ordering and the method that keeps it.
Where is it used?
QR factorisation, least squares, orthogonal weight initialisation, the Arnoldi and Lanczos methods for large eigenvalue problems, and building orthonormal bases for subspaces.
How is it used?
In code, call np.linalg.qr (Householder) rather than writing Gram–Schmidt yourself. If you must write it, use the modified version and re-orthogonalise. Check quality with the loss of orthogonality np.abs(Q.T @ Q - I).max().
Even MGS is not safe for very ill-conditioned matrices, and both Gram–Schmidt methods break if a column is (nearly) a combination of earlier ones: the "leftover" is $\approx 0$ and you divide by a tiny length. Householder reflections avoid both problems.
Quick check: which orthogonalisation method keeps $Q^\top Q \approx I$ no matter how ill-conditioned $A$ is?
Householder QR. Its orthogonality error is about $u$, while MGS is about $u\kappa$ and CGS about $u\kappa^2$.
Pivoting: never divide by a tiny number core
Gaussian elimination (Chapter 1.6) divides by the pivot, the number on the diagonal. If the pivot is very small, the multiplier is huge, and a huge number swallows the small ones in its row (just like the boulder and pebbles earlier). Information is lost.
The fix is simple and cheap: swap rows so that the biggest available number in the column becomes the pivot. Then every multiplier has size at most 1.
Solve $\begin{bmatrix} \delta & 1 \\ 1 & 1\end{bmatrix}\mathbf{x} = \begin{bmatrix} 1 \\ 2\end{bmatrix}$ with $\delta = 10^{-20}$. The true answer is $\mathbf{x} \approx [1, 1]$.
Without swapping.
- Multiplier $= 1/\delta = 10^{20}$.
- Row 2 becomes $[0,\ 1 - 10^{20}]$ with right side $2 - 10^{20}$. In floating point, $1 - 10^{20} = -10^{20}$ and $2 - 10^{20} = -10^{20}$: the 1 and the 2 are swallowed.
- So $x_2 = (-10^{20})/(-10^{20}) = 1$. Fine so far.
- Back-substitute: $x_1 = (1 - x_2)/\delta = (1 - 1)/10^{-20} = 0$. Wrong: the answer should be $1$.
With a row swap (pivot $=1$): multiplier $= \delta = 10^{-20}$, row 2 becomes $[0,\ 1 - 10^{-20}] \to [0, 1]$, $x_2 = 1$, $x_1 = 2 - 1 = 1$. Correct.
Partial pivoting: at column $k$, find the entry with the largest absolute value in that column (at or below the diagonal) and swap its row into the pivot position. This gives the factorisation
$$PA = LU,$$where $P$ is the permutation matrix recording the swaps and every entry of $L$ has $|l_{ij}| \le 1$. With pivoting, the algorithm is backward stable in practice. Complete pivoting ○ searches rows and columns for the biggest entry: a little safer, a lot slower, rarely needed.
Pivoting is not about whether an exact answer exists (the matrix was never singular). It is about stability, a property of the algorithm.
Why do we need it?
Gaussian elimination divides by the pivot. A tiny pivot gives a giant multiplier that swamps the other numbers. Swapping rows keeps every multiplier small.
Where is it used?
np.linalg.solve, np.linalg.det, scipy.linalg.lu and every LAPACK solver (LU with partial pivoting). Cholesky for symmetric positive definite matrices needs no pivoting.
How is it used?
You normally do nothing: library routines pivot for you and return the permutation as PA = LU. If you ever write elimination by hand, swap in the row with the largest entry in the column first, at every step.
Pivoting cannot rescue a truly ill-conditioned matrix. It only stops the algorithm from making things worse than $\kappa u$. If the problem is ill-conditioned, some digits are lost whatever you do.
Quick check: why must the multipliers in Gaussian elimination be at most 1 in size?
A multiplier bigger than 1 makes the new rows bigger than the old ones, so small numbers get added to large ones and lose digits (as with $1/\delta$ above). Choosing the largest entry as pivot keeps every multiplier $\le 1$ and the numbers from growing.
Stable softmax and the log-sum-exp trick core
Softmax turns a list of scores into probabilities: make every score positive with $e^{z}$, then divide by the total. The trouble is that $e^{z}$ grows very fast. A score of $100$ gives $e^{100} \approx 10^{43}$, which already overflows in float32.
Here is the rescue. If you subtract the same number from every score, the probabilities do not change: the matching factor cancels between the top and the bottom of the fraction (like measuring everybody's height from a different floor: the differences between people do not change). So subtract the biggest score. Then the biggest becomes $0$ and $e^0 = 1$, and every other term is between $0$ and $1$. Nothing can overflow, and the sum is at least $1$, so we never divide by zero.
Scores $\mathbf{z} = [1000, 1001, 1002]$. Naively, $e^{1000}$ is Inf in float64, and Inf / Inf is NaN.
- The maximum is $m = 1002$. Subtract it: $\mathbf{z} - m = [-2, -1, 0]$.
- Exponentiate: $[e^{-2}, e^{-1}, e^{0}] = [0.1353,\ 0.3679,\ 1]$.
- Sum: $0.1353 + 0.3679 + 1 = 1.5032$.
- Divide: $[0.0900,\ 0.2447,\ 0.6652]$. These are the right probabilities (the same as for $[0,1,2]$).
For scores $\mathbf{z}$ with maximum $m = \max_j z_j$:
$$\text{softmax}(\mathbf{z})_i = \frac{e^{z_i}}{\sum_j e^{z_j}} = \frac{e^{z_i - m}}{\sum_j e^{z_j - m}}.$$The log-sum-exp of the scores is the log of the denominator, computed safely as
$$\text{LSE}(\mathbf{z}) = \log\sum_j e^{z_j} = m + \log\sum_j e^{z_j - m}.$$Then $\log\text{softmax}(\mathbf{z})_i = z_i - \text{LSE}(\mathbf{z})$. The cross-entropy loss for the true class $y$ is $\text{LSE}(\mathbf{z}) - z_y$, and it is computed directly, without ever forming a probability.
Why do we need it?
The exponential in softmax overflows for big scores, which breaks attention and classifiers. A one-line shift by the maximum fixes it without changing the result.
Where is it used?
The softmax layer and cross-entropy loss of every classifier, attention weights in Transformers, log-likelihoods, mixture models, and torch.logsumexp, scipy.special.logsumexp and log_softmax.
How is it used?
Subtract the maximum score before exp. Use log_softmax or logsumexp instead of log(softmax(z)). Feed raw scores (logits) to cross_entropy losses, which do this safely inside.
Never write np.log(softmax(z)). A tiny probability underflows to $0$, and $\log 0 = -\infty$. Use log_softmax (or the LSE form).
Never use the "add a small number inside the log" patch (log(p + 1e-8)) when a stable form exists. It silently changes the answer.
Quick check: why can't the shifted softmax overflow, and why can't it divide by zero?
After subtracting the maximum, every exponent is $\le 0$, so every term is $\le 1$: no overflow. The largest entry becomes $e^0 = 1$, so the sum is at least $1$: no division by zero.
Working in log space core
Probabilities of independent events multiply. A sentence with 400 words, each of probability $0.1$, has probability $0.1^{400} = 10^{-400}$. That is smaller than the smallest float64 number ($\approx 10^{-324}$), so the computer writes 0. Every sentence now looks impossible.
Logarithms turn products into sums: $\log(ab) = \log a + \log b$. And sums of logs do not underflow: $400\times\log(0.1) = -921$ is an ordinary number. So we keep the logarithm of every probability and never go back, unless we have to.
- Ten words, each with probability $0.1$. Product: $0.1^{10} = 10^{-10}$. Fine.
- Log version: $10 \times \ln 0.1 = 10 \times (-2.3026) = -23.026$. Check: $e^{-23.026} = 10^{-10}$ ✓.
- Now 400 words. Product: $10^{-400}$, which underflows to $0$. Log version: $400 \times (-2.3026) = -921.03$. Still perfectly fine.
- To compare two sentences, compare their log-probabilities: bigger (less negative) wins. The order is the same, because $\log$ is increasing.
Log-likelihood. For independent items with probabilities $p_1,\dots,p_n$:
$$\log\prod_i p_i = \sum_i \log p_i.$$To add two probabilities that are stored as logs, use log-sum-exp (the previous section):
$$\log(e^{a} + e^{b}) = \max(a,b) + \log\left(1 + e^{-|a-b|}\right).$$This is what np.logaddexp(a, b) computes.
Why do we need it?
Products of many probabilities shrink below the smallest number the computer can hold and become exactly 0. Adding logarithms instead of multiplying avoids that.
Where is it used?
Naive Bayes classifiers, hidden Markov models, language-model scoring and perplexity, log-likelihood losses, and any probabilistic model with many independent factors.
How is it used?
Store log-probabilities. Turn products into sums. Add two probabilities stored as logs with np.logaddexp. Compare candidates by their log scores, and only exponentiate at the very end, if at all.
Staying in log space is only free for products. A sum of probabilities needs the logsumexp trick, never "exponentiate, add, take the log" (which brings the underflow straight back).
Only exponentiate at the very end, and only if you really need a probability.
Quick check: you have $\log p_1 = -1000$ and $\log p_2 = -1001$. How do you compute $\log(p_1 + p_2)$ safely?
Use $\max + \log(1 + e^{-|a-b|}) = -1000 + \log(1 + e^{-1}) = -1000 + 0.3133 = -999.687$. Never compute $e^{-1000}$ directly: it underflows to $0$.
Adding a small ε to the diagonal (jitter) core
Some matrices are almost singular. A kernel matrix built from points that are very close together has rows that are nearly copies of each other. Mathematically it is still fine. For the computer, its smallest eigenvalue is so close to $0$ (or even slightly negative through rounding) that factorising it fails.
The fix is a pinch of salt: add a tiny positive number $\varepsilon$ to every diagonal entry. This lifts all eigenvalues by $\varepsilon$, so none can be zero or negative. The answer changes a little, but the problem becomes solvable.
Two identical points give $K = \begin{bmatrix} 1 & 1 \\ 1 & 1 \end{bmatrix}$, which is singular (its eigenvalues are $2$ and $0$).
- Cholesky (Chapter 1.13) tries $L_{11} = \sqrt{1} = 1$, $L_{21} = 1/1 = 1$, then $L_{22} = \sqrt{1 - 1^2} = \sqrt{0}$. A zero pivot: failure (no inverse exists).
- Add $\varepsilon = 10^{-6}$: $K + \varepsilon I = \begin{bmatrix} 1.000001 & 1 \\ 1 & 1.000001 \end{bmatrix}$. Now $L_{22} = \sqrt{1.000001 - 1/1.000001} \approx \sqrt{2\times10^{-6}} \approx 0.0014 > 0$. It works.
- The eigenvalues are now $2 + \varepsilon$ and $\varepsilon$, so $\kappa = (2+\varepsilon)/\varepsilon \approx 2\times10^{6}$: finite, large but manageable.
For a symmetric positive semi-definite matrix $K$ (from Chapter 1.12), the jittered matrix is
$$K_\varepsilon = K + \varepsilon I, \qquad \lambda_i(K_\varepsilon) = \lambda_i(K) + \varepsilon, \qquad \kappa(K_\varepsilon) = \frac{\lambda_{\max} + \varepsilon}{\lambda_{\min} + \varepsilon}.$$Typical starting choices: $\varepsilon \approx 10^{-6}$ for float64 and $10^{-4}$ for float32 (a bit above the rounding noise, and small compared with the typical eigenvalue). The right value depends on the matrix, so treat these as starting points. The same idea appears under many names: jitter, nugget, damping, ridge or Tikhonov regularisation.
Why do we need it?
Matrices built from similar data points are nearly singular, so a Cholesky factorisation or inverse can fail. A tiny constant added to the diagonal lifts the weakest directions.
Where is it used?
Gaussian process kernel matrices, covariance matrices in LDA and Mahalanobis distance, ridge regression, Levenberg–Marquardt damping, and (same idea: keep things away from zero) the small epsilon in Adam, BatchNorm and LayerNorm.
How is it used?
Replace K by K + eps·I, with eps around 1e-6 in float64 or 1e-4 in float32. Start small and raise it only if the factorisation fails. If you need a very large eps, suspect a bug in how the matrix was built.
Jitter changes the problem. Too much jitter over-smooths the answer. Too little does nothing against rounding noise. Start from about $10^{-6}$ (float64) and increase only if needed.
If a matrix should be positive definite but is far from it (large negative eigenvalues), the bug is elsewhere: jitter hides real mistakes.
Quick check: a matrix has eigenvalues $[5, 1, 10^{-12}]$. What is $\kappa$ after adding $\varepsilon = 10^{-6}$ to the diagonal?
Before: $5/10^{-12} = 5\times10^{12}$. After: eigenvalues $5.000001$, $1.000001$, $\approx 10^{-6}$, so $\kappa \approx 5/10^{-6} = 5\times10^{6}$. The condition number dropped by a factor of a million.
Why iterate? Jacobi and Gauss–Seidel (awareness)
A direct method (like LU) is a long recipe that gives the exact answer at the end, but it costs about $n^3$ steps and needs the whole matrix in memory. For $n = 10^6$ unknowns, $n^3 = 10^{18}$: far too slow. And the matrices in real problems (graphs, images, physics, recommenders) are usually huge and sparse, where the only cheap thing to do is multiply the matrix by a vector.
An iterative method plays "hot or cold". Make a guess. Check how wrong it is. Use that to make a better guess. Repeat until the error is small enough. Each round uses only a matrix–vector product. You never get the exact answer, but you can stop as soon as it is good enough.
Solve $4x + y = 6$ and $x + 3y = 7$ (answer $x = 1,\ y = 2$). Rewrite each equation to get one unknown alone: $x = (6 - y)/4$ and $y = (7 - x)/3$. Start from $(0, 0)$.
- Jacobi (use only the old values): $x_1 = (6-0)/4 = 1.5$, $y_1 = (7-0)/3 = 2.333$. Next round: $x_2 = (6 - 2.333)/4 = 0.917$, $y_2 = (7 - 1.5)/3 = 1.833$.
- Gauss–Seidel (use the newest value straight away): $x_1 = 1.5$, then $y_1 = (7 - 1.5)/3 = 1.833$ (it already uses the new $x$). Next: $x_2 = (6 - 1.833)/4 = 1.042$, $y_2 = (7 - 1.042)/3 = 1.986$.
After two rounds Gauss–Seidel $(1.042, 1.986)$ is closer to $(1, 2)$ than Jacobi $(0.917, 1.833)$. Both are creeping towards the answer.
Split the matrix as $A = D + L + U$: the diagonal $D$, the strictly lower part $L$ and the strictly upper part $U$. (These $L$ and $U$ are just pieces of $A$ that get added up. They are not the factors of an LU factorisation.) Then
$$\text{Jacobi: } \mathbf{x}^{(k+1)} = D^{-1}\bigl(\mathbf{b} - (L+U)\,\mathbf{x}^{(k)}\bigr), \qquad \text{Gauss–Seidel: } \mathbf{x}^{(k+1)} = (D+L)^{-1}\bigl(\mathbf{b} - U\,\mathbf{x}^{(k)}\bigr).$$Each is a "residual correction": $\mathbf{x} \leftarrow \mathbf{x} + M^{-1}(\mathbf{b} - A\mathbf{x})$ with $M = D$ or $M = D+L$. They converge when the iteration matrix has all eigenvalues of size less than 1, which is guaranteed if $A$ is strictly diagonally dominant (each diagonal entry bigger than the sum of the other entries in its row). If the off-diagonal "coupling" is too strong, the iteration diverges.
These two are mostly of historical and teaching interest as stand-alone solvers. Gauss–Seidel survives as a "smoother" inside multigrid methods.
Why do we need it?
Direct methods cost about n³ steps and fill up memory, which is hopeless for millions of unknowns. Iterative methods improve a guess using only cheap matrix–vector products.
Where is it used?
Huge sparse systems: PageRank, graph Laplacians, physics and finance simulations, Gaussian processes at scale. Gauss–Seidel survives as the smoother inside multigrid solvers.
How is it used?
Start with a guess, compute the residual (b minus Ax), use it to correct the guess, and stop when the residual is small enough. Check convergence first: Jacobi and Gauss–Seidel need a diagonally dominant matrix.
"Iterative" does not mean "approximate and sloppy". With enough steps the answer can be as accurate as the direct one, and you choose the accuracy. Also remember: they only help when one step (a matrix–vector product) is much cheaper than a full factorisation. For a small dense matrix, just call np.linalg.solve.
Quick check: why do iterative solvers store only $O(n)$ extra numbers while LU on a sparse matrix can need far more?
An iterative solver only needs the matrix (kept sparse) and a few vectors of length $n$. LU creates fill-in: zeros of $A$ become non-zeros in $L$ and $U$, so the factors can be much denser than $A$.
Gradient descent as an iterative linear solver
When $A$ is symmetric positive definite (Chapter 1.12), solving $A\mathbf{x} = \mathbf{b}$ is the same problem as finding the bottom of a bowl. The bowl is the quadratic function $f(\mathbf{x}) = \tfrac12\mathbf{x}^\top A\mathbf{x} - \mathbf{b}^\top\mathbf{x}$. Its lowest point is exactly where $A\mathbf{x} = \mathbf{b}$.
The slope of the bowl at $\mathbf{x}$ is $A\mathbf{x} - \mathbf{b}$. That is minus the residual (how wrong the equation currently is). So "roll downhill" means "move in the direction of the residual". Gradient descent is a linear solver, and each step costs a single matrix–vector product.
The smallest case: $A = [2]$ and $b = 4$, so we solve $2x = 4$. The bowl is $f(x) = x^2 - 4x$, with slope $2x - 4$. Take step size $\alpha = 0.25$ and start at $x = 0$.
- Residual $r = b - Ax = 4 - 0 = 4$. New $x = 0 + 0.25\times4 = 1$.
- $r = 4 - 2 = 2$. New $x = 1 + 0.25\times2 = 1.5$.
- $r = 4 - 3 = 1$. New $x = 1.5 + 0.25 = 1.75$.
- $r = 0.5$, $x = 1.875$. Then $1.9375,\ 1.96875,\dots$ The error halves every step, and $x \to 2$ ✓.
For symmetric positive definite $A$, minimising $f(\mathbf{x}) = \tfrac12\mathbf{x}^\top A\mathbf{x} - \mathbf{b}^\top\mathbf{x}$ is equivalent to solving $A\mathbf{x} = \mathbf{b}$, because $\nabla f = A\mathbf{x} - \mathbf{b} = -\mathbf{r}$. Gradient descent is
$$\mathbf{x}_{k+1} = \mathbf{x}_k + \alpha\,\mathbf{r}_k, \qquad \mathbf{r}_k = \mathbf{b} - A\mathbf{x}_k.$$For steepest descent the best step along $\mathbf{r}_k$ is found exactly: $\alpha_k = \dfrac{\mathbf{r}_k^\top\mathbf{r}_k}{\mathbf{r}_k^\top A\,\mathbf{r}_k}$.
The contours of $f$ are ellipses. If $A$ has eigenvalues $\lambda_{\min}$ and $\lambda_{\max}$, the ellipse is $\sqrt{\kappa}$ times longer than it is wide, with $\kappa = \lambda_{\max}/\lambda_{\min}$. Each steepest-descent step multiplies the error by at most the factor
$$\frac{\kappa - 1}{\kappa + 1}.$$(Here the error is measured in the "energy norm" $\sqrt{\mathbf{e}^\top A\,\mathbf{e}}$, a length that fits the bowl's shape.) For $\kappa = 1$ (round bowl) this is $0$: one step. For $\kappa = 100$ it is $0.98$: the error drops by only 2% per step.
Why do we need it?
For symmetric positive definite A, solving Ax = b is the same as finding the bottom of a bowl. This links linear solving to optimisation and explains why training behaves as it does.
Where is it used?
Linear regression by gradient descent (the loss is a quadratic bowl), the analysis of learning rates, understanding zig-zagging in badly scaled problems, and the idea behind momentum and Adam.
How is it used?
Update x with x + α·(b minus Ax), where α is the step size. A smaller condition number means a rounder bowl and fewer steps. Standardise your features to make the bowl rounder before training.
This picture is exactly why plain gradient descent is slow on badly scaled problems. The loss surface of a neural network has long thin valleys too (a large "condition number" of the Hessian, the matrix of second derivatives from Chapter 1.14), and the same zig-zag appears. Rescaling the features (standardising) makes the bowl rounder.
In machine learning the step $\alpha$ is a fixed learning rate, not the exact line-search value. Too big and the iteration diverges, just like Jacobi did when the coupling was too strong.
Quick check: what is the gradient of $f(\mathbf{x}) = \tfrac12\mathbf{x}^\top A\mathbf{x} - \mathbf{b}^\top\mathbf{x}$ for symmetric $A$?
$\nabla f = A\mathbf{x} - \mathbf{b}$, which is minus the residual. It is zero exactly when $A\mathbf{x} = \mathbf{b}$.
Conjugate gradient and Krylov subspaces (awareness)
Steepest descent has a short memory. Each new step only looks at the current slope, so it keeps undoing earlier progress and zig-zags. Conjugate gradient (CG) keeps its memory: every new direction is chosen so that it does not spoil what the earlier directions already achieved.
Think of a stretched bowl. A hiker who walks along one axis of the ellipse until the height stops dropping, and then along the other axis, reaches the bottom in two steps. CG finds directions with the same "do not spoil each other" property, called A-conjugate directions, without knowing the axes in advance. In $n$ dimensions it needs at most $n$ steps (in exact arithmetic), and it usually gets very close in far fewer.
Solve $A\mathbf{x} = \mathbf{b}$ with $A = \begin{bmatrix} 4 & 1 \\ 1 & 3 \end{bmatrix}$, $\mathbf{b} = [1, 2]$, from $\mathbf{x}_0 = \mathbf{0}$.
- $\mathbf{r}_0 = \mathbf{b} = [1, 2]$ and $\mathbf{p}_0 = \mathbf{r}_0$. $A\mathbf{p}_0 = [6, 7]$. $\alpha_0 = \dfrac{\mathbf{r}_0\cdot\mathbf{r}_0}{\mathbf{p}_0\cdot A\mathbf{p}_0} = \dfrac{5}{20} = 0.25$. So $\mathbf{x}_1 = [0.25, 0.5]$.
- $\mathbf{r}_1 = \mathbf{r}_0 - 0.25\,A\mathbf{p}_0 = [-0.5, 0.25]$. $\beta_0 = \dfrac{\mathbf{r}_1\cdot\mathbf{r}_1}{\mathbf{r}_0\cdot\mathbf{r}_0} = \dfrac{0.3125}{5} = 0.0625$. New direction $\mathbf{p}_1 = \mathbf{r}_1 + 0.0625\,\mathbf{p}_0 = [-0.4375, 0.375]$.
- $A\mathbf{p}_1 = [-1.375, 0.6875]$, $\alpha_1 = \dfrac{0.3125}{0.859375} = 0.3636$. So $\mathbf{x}_2 = [0.25 - 0.1591,\ 0.5 + 0.1364] = [0.0909,\ 0.6364]$.
Check: $A\mathbf{x}_2 = [4(0.0909) + 0.6364,\ 0.0909 + 3(0.6364)] = [1.0,\ 2.0]$ ✓. Done in exactly 2 steps (the dimension).
Conjugate gradient for symmetric positive definite $A$: start with $\mathbf{r}_0 = \mathbf{b} - A\mathbf{x}_0$, $\mathbf{p}_0 = \mathbf{r}_0$, and repeat
alpha = (r @ r) / (p @ (A @ p)) # best step along the search direction p
x = x + alpha * p
r_new = r - alpha * (A @ p) # new residual (one matrix-vector product per step)
beta = (r_new @ r_new) / (r @ r)
p = r_new + beta * p # new direction: mostly the residual, plus a bit of the old direction
r = r_new
The directions satisfy $\mathbf{p}_i^\top A\,\mathbf{p}_j = 0$ for $i \ne j$ (A-conjugate). After $k$ steps, $\mathbf{x}_k$ is the best possible point in $\mathbf{x}_0 + \mathcal{K}_k$, where the Krylov subspace is
$$\mathcal{K}_k = \text{span}\{\mathbf{r}_0,\ A\mathbf{r}_0,\ A^2\mathbf{r}_0,\ \ldots,\ A^{k-1}\mathbf{r}_0\}.$$In words: each step reaches one new direction by multiplying by $A$ once more. The error (in the energy norm from the last section) shrinks by about $\dfrac{\sqrt\kappa - 1}{\sqrt\kappa + 1}$ per step. Compare this with steepest descent's $\dfrac{\kappa - 1}{\kappa + 1}$: CG needs about $\sqrt{\kappa}$ steps where steepest descent needs about $\kappa$. Krylov methods for non-symmetric matrices (GMRES, BiCGSTAB) ○ follow the same idea.
Why do we need it?
Gradient descent forgets its past directions and zig-zags. Conjugate gradient keeps directions that do not spoil earlier progress, so it needs far fewer steps (about the square root).
Where is it used?
Large symmetric positive definite systems, Hessian-free optimisation, Gaussian processes on GPUs, and Krylov methods (Lanczos, Arnoldi) for big eigenvalue and SVD problems.
How is it used?
Call scipy.sparse.linalg.cg(A, b), ideally with a preconditioner (a cheap approximate inverse). It needs A symmetric positive definite and only matrix–vector products. Stop on a residual tolerance, not after n steps.
CG needs $A$ to be symmetric positive definite. For other matrices use different Krylov methods (GMRES, BiCGSTAB), or apply CG to $A^\top A$ (which squares $\kappa$ again).
"At most $n$ steps" is only true in exact arithmetic. With rounding errors the directions slowly lose their conjugacy, so in practice we stop on a tolerance, usually long before $n$.
In practice CG is combined with a preconditioner: a cheap approximate inverse $M^{-1}$ that makes the effective $\kappa$ small. A good preconditioner is worth more than any other tuning.
Quick check: roughly how many steps would CG need for a system with $\kappa = 10^4$, compared with steepest descent?
CG: on the order of $\sqrt{10^4} = 100$ steps times a small constant. Steepest descent: on the order of $\kappa = 10^4$ steps times a constant. CG is about 100 times faster.
Convergence depends on the condition number
The condition number appears again, but in a new role. Before, it said "how much will errors be amplified?". Now it says "how many steps will an iterative method need?"
A round bowl ($\kappa = 1$) is solved in one step. A very stretched bowl ($\kappa = 1000$) needs about a thousand steps for gradient descent, and about thirty for conjugate gradient. The square root makes all the difference.
To shrink the error by a factor of $10^{6}$ (about 14 "e-folds", since $\ln 10^6 = 13.8$) with $\kappa = 100$:
- Gradient descent, per-step factor about $\frac{\kappa-1}{\kappa+1} = 0.98$: $\;\dfrac{13.8}{-\ln 0.98} \approx 680$ steps.
- Conjugate gradient, per-step factor about $\frac{\sqrt\kappa-1}{\sqrt\kappa+1} = \frac{9}{11} = 0.82$: $\;\dfrac{13.8}{-\ln 0.82} \approx 70$ steps (and in practice often fewer).
Multiply $\kappa$ by $100$: gradient descent needs about $100$ times more steps. CG needs only about $10$ times more.
To reach a relative accuracy $\text{tol}$ you need roughly
$$k_{\text{GD}} \approx \tfrac{1}{2}\,\kappa\,\ln\frac{1}{\text{tol}}, \qquad k_{\text{CG}} \approx \tfrac{1}{2}\,\sqrt{\kappa}\,\ln\frac{2}{\text{tol}}$$iterations. These are worst-case rules; clustered eigenvalues make CG even faster (it needs about one step per distinct cluster of eigenvalues). The remedies are the ones from earlier: rescale the data, regularise (add $\lambda I$) or use a preconditioner. All of them lower $\kappa$.
Why do we need it?
We need to predict how long an iterative method will take before running it. The condition number tells us: gradient descent needs about κ steps and conjugate gradient about √κ.
Where is it used?
Choosing optimisers and learning rates, feature scaling, batch normalisation, preconditioning in scientific computing, and explaining why second-order methods can need far fewer steps.
How is it used?
Estimate κ (or the ratio of the largest to smallest curvature). If it is large, lower it: standardise features, add regularisation, or precondition. Then plan iterations by rule of thumb: κ for gradient descent, √κ for CG.
The rules of thumb are upper limits. Real counts can be smaller (CG here beats the bound because it is only a 50-dimensional system). The growth rate is the lesson: GD scales like $\kappa$, CG like $\sqrt\kappa$.
Quick check: you reduce $\kappa$ from 10 000 to 100 by rescaling the features. By roughly what factor does plain gradient descent speed up? And CG?
GD scales with $\kappa$: $10\,000/100 = 100$ times faster. CG scales with $\sqrt\kappa$: $\sqrt{10\,000}/\sqrt{100} = 100/10 = 10$ times faster.
Sparse matrices: store only what is there
Many real matrices are almost all zeros. In a table of "who follows whom" on a social network, each person follows a few hundred of millions of people. Storing every zero would waste memory and make every multiplication do millions of pointless "times 0".
A sparse matrix stores only the non-zero entries together with their addresses. It is like a phone book: you list the people who exist, not a row for every possible name.
Take $A = \begin{bmatrix} 10 & 0 & 0 & 2 \\ 0 & 3 & 0 & 0 \\ 0 & 0 & 0 & 7 \\ 1 & 0 & 4 & 0 \end{bmatrix}$. It has $16$ entries but only 6 non-zeros. Reading row by row, the non-zeros are $10, 2, 3, 7, 1, 4$.
- COO (coordinate): three lists, one entry per non-zero.
row = [0,0,1,2,3,3],col = [0,3,1,3,0,2],data = [10,2,3,7,1,4]. - CSR (compressed sparse row): keep
dataandcolas above, but replace the row list byindptr = [0,2,3,4,6]. It says "row $i$ owns the entries from positionindptr[i]up toindptr[i+1]". Row 0 owns positions 0 to 2, so $10$ and $2$. - CSC (compressed sparse column): the same idea by columns.
data = [10,1,3,4,2,7],row = [0,3,1,3,0,2],indptr = [0,2,3,4,6]. - Matrix–vector product with $\mathbf{x} = [1,2,3,4]$ using CSR: $y_0 = 10\cdot1 + 2\cdot4 = 18$, $y_1 = 3\cdot2 = 6$, $y_2 = 7\cdot4 = 28$, $y_3 = 1\cdot1 + 4\cdot3 = 13$. Only 6 multiplications instead of 16.
With $\text{nnz}$ non-zeros:
| Format | Stores | Good at |
|---|---|---|
| COO | (row, col, value) triples | building a matrix; easy to append to |
| CSR | values, column indices, row pointers | fast row slicing and matrix–vector products |
| CSC | values, row indices, column pointers | fast column slicing; many direct solvers |
A sparse matrix–vector product costs $O(\text{nnz})$ instead of $O(mn)$. Memory drops from $mn$ numbers to about $2\,\text{nnz} + m$. The density is $\text{nnz}/(mn)$. In Python: scipy.sparse.csr_matrix(A), then A_sp @ x.
Why do we need it?
Many big matrices are almost all zeros. Storing and multiplying the zeros wastes memory and time, so we store only the non-zero entries and their positions.
Where is it used?
Bag-of-words and TF-IDF text matrices, graph adjacency matrices and graph neural networks, recommender systems (user by item), one-hot encodings, and finite-element simulations. Tools: scipy.sparse, torch.sparse.
How is it used?
Build with COO (easy to assemble), convert to CSR for fast row access and products, or CSC for columns. Multiply with A @ x: the cost follows the number of non-zeros. Stay sparse; the inverse and the LU factors may become dense.
Sparse is not always better. For a matrix that is more than a few percent full, a well-tuned dense multiply on optimised hardware beats the sparse version, because sparse access jumps around in memory. The sparse formats also need extra index storage.
The inverse of a sparse matrix is usually dense (and so are the factors of an LU, through "fill-in"). That is another reason to avoid explicit inverses and to prefer iterative methods for huge sparse systems.
Quick check: a $1000\times1000$ matrix has 5000 non-zeros. How many multiplications does a sparse matrix–vector product need, against a dense one?
Sparse: $\text{nnz} = 5000$. Dense: $1000\times1000 = 10^6$. That is $200$ times fewer.
How much work is it? Counting operations core
Before you run something big, ask: if the input gets 10 times larger, how much longer will it take? You find out by counting multiplications and additions, and seeing how the count grows with the size $n$.
- A dot product pairs up $n$ numbers: work grows in proportion to $n$.
- A matrix times a vector is $m$ dot products of length $n$: work grows like $mn$.
- A matrix times a matrix is $n^2$ dot products of length $n$: work grows like $n^3$. Double the size and it takes 8 times longer.
- Dot product of two vectors of length $n = 1000$: $1000$ multiplications and $999$ additions, about $2n = 2000$ flops ("flop" = one floating-point operation).
- Matrix–vector with a $1000\times1000$ matrix: $1000$ dot products, about $2n^2 = 2\times10^6$ flops.
- Matrix–matrix with two $1000\times1000$ matrices: $10^6$ output entries, each a dot product of length $1000$, about $2n^3 = 2\times10^9$ flops.
- On a machine doing $10^{10}$ flops per second, these take $0.2$ microseconds, $0.2$ milliseconds and $0.2$ seconds.
- Now make $n = 10\,000$: the matrix–matrix product is $1000$ times more work: 200 seconds.
"$O(\cdot)$" ("big-O") describes how the work grows, ignoring constant factors:
| Operation | Cost | Flops (approx.) |
|---|---|---|
| dot product of length $n$ | $O(n)$ | $2n$ |
| matrix ($m\times n$) times vector | $O(mn)$ | $2mn$ |
| matrix ($n\times n$) times matrix | $O(n^3)$ | $2n^3$ |
| solve $A\mathbf{x}=\mathbf{b}$ by LU | $O(n^3)$ | $\tfrac23 n^3$ (then $2n^2$ per extra $\mathbf{b}$) |
| Cholesky (symmetric positive definite only) | $O(n^3)$ | $\tfrac13 n^3$ (half of LU) |
| QR (Householder) | $O(n^3)$ | $\tfrac43 n^3$ |
| eigenvalues / SVD | $O(n^3)$ | roughly $10$ to $25\,n^3$ (iterative, bigger constant) |
Strassen's algorithm ○ multiplies two $n\times n$ matrices with $O(n^{2.81})$ operations by cleverly using 7 block multiplications instead of 8 (and recursing). Research algorithms go as low as about $n^{2.37}$. In practice libraries still use the plain $n^3$ method with careful memory use, because the clever ones have large constants, are less stable and rarely pay off at normal sizes.
Why do we need it?
Before running a big computation we must know if it will take a second or a year. Counting operations as a function of size n answers that.
Where is it used?
Planning training budgets, choosing between Cholesky and SVD, understanding why Gaussian processes struggle past about 10,000 points, and picking truncated or randomised SVD for big data.
How is it used?
Count the multiply–adds: dot product 2n, matrix–vector 2mn, matrix–matrix 2n³. Time is roughly the count divided by the machine speed. Double the size and expect 2×, 4× or 8×. Never form an inverse just to solve a system.
Cholesky vs LU: when to use which. LU (with pivoting) works for any square invertible matrix and costs about $\tfrac23 n^3$ flops: it is what np.linalg.solve does. Cholesky ($A = LL^\top$) works only for symmetric positive definite matrices, costs about $\tfrac13 n^3$ (half), needs no pivoting and is very stable. Use it for covariance and kernel matrices, Gaussian processes and the system $(X^\top X + \lambda I)\mathbf{w} = X^\top\mathbf{y}$. A failing Cholesky is also a free test for "not positive definite" (see the jitter section). The factorisations themselves are explained in Chapter 1.13.
Big-O hides constants. Cholesky and LU are both $O(n^3)$, but Cholesky does half the work. An $O(n^2)$ method with a huge constant can lose to an $O(n^3)$ one at normal sizes.
Flops are not the whole story. A dot product and a matrix product do very different amounts of work per number read from memory. That is the topic of the next section.
Never form an inverse to solve a system. Computing $A^{-1}$ costs about $2n^3$ flops, three times an LU solve, and then you still need a matrix–vector product. Use np.linalg.solve.
Quick check: a matrix–matrix product with $n = 500$ takes 1 second. About how long for $n = 2000$?
Four times larger gives $4^3 = 64$ times more work: about 64 seconds.
Memory vs compute: BLAS, LAPACK and why GPUs love matmul core
Picture a chef (the processor) and a pantry (the memory). The chef can chop very fast, but each trip to the pantry is slow. If every ingredient is used once (as in a dot product: multiply a pair, add, done), the chef spends most of the time waiting at the pantry door. If every ingredient is used many times (as in matrix multiplication, where every number is used $n$ times), the chef is always busy.
A GPU is a kitchen with thousands of chefs and a very wide pantry door. It is amazing when the work is "lots of cooking per ingredient fetched", which is exactly what matrix–matrix multiplication offers. That is why deep learning is built from matmuls.
Count "flops per number moved" (called arithmetic intensity) for vectors and matrices of size $n$:
- Dot product: $2n$ flops, reads $2n$ numbers. Intensity $= 2n / 2n = 1$.
- Matrix–vector: $2n^2$ flops, reads about $n^2$ numbers (the matrix). Intensity $\approx 2$.
- Matrix–matrix: $2n^3$ flops, reads (and writes) $3n^2$ numbers. Intensity $= 2n^3/3n^2 = \tfrac23 n$. For $n = 3000$ that is $2000$ flops per number!
On a GPU that can do $10^{13}$ flops/s but only move $1.25\times10^{11}$ numbers/s, a dot product reaches $1 \times 1.25\times10^{11}$ flops/s (about 1% of peak), while a big matmul reaches the full $10^{13}$.
BLAS (Basic Linear Algebra Subprograms) is a standard list of fast routines in three levels: level 1 vector–vector (dot, axpy), level 2 matrix–vector, level 3 matrix–matrix (gemm). LAPACK builds LU, Cholesky, QR, eigenvalue and SVD routines on top of BLAS, arranged so that most of the work is level 3. Implementations (OpenBLAS, Intel MKL, Apple Accelerate, and cuBLAS on NVIDIA GPUs) split matrices into small tiles that fit in fast cache, so each loaded number is reused many times.
NumPy's @, np.linalg.solve, np.linalg.svd and PyTorch's matmul all call these libraries. Writing your own triple loop in Python is easily 100 to 1000 times slower.
Why do we need it?
Fast arithmetic is useless if the numbers arrive slowly. Matrix products reuse each number many times, so they keep the hardware busy, and that shapes how we write ML code.
Where is it used?
BLAS and LAPACK under NumPy and SciPy, cuBLAS and tensor cores on NVIDIA GPUs, batched inference, quantisation of weights, and kernel fusion in PyTorch and JAX compilers.
How is it used?
Hand whole arrays to library calls instead of looping. Make matmuls big by batching. Fuse element-wise steps. Remember that batch size 1 is a matrix–vector product and is limited by memory speed, not by arithmetic.
The same arithmetic can be slow or fast depending on how it is arranged. Many small matmuls, or big element-wise operations (adding, exp, normalising), are memory-bound: they run at the speed of the pantry, not the chefs. This is why GPU code fuses element-wise steps into one pass.
Python loops over matrix entries pay an extra price: each iteration runs slow interpreter code. Always hand whole arrays to BLAS (more on this in Chapter 1.16).
Quick check: why is multiplying a $4096\times4096$ matrix by a single vector so much less efficient per flop than multiplying two $4096\times4096$ matrices?
The matrix–vector product reads every one of the $4096^2$ matrix numbers and uses each only once (intensity about 2), so it waits on memory. In the matrix–matrix product every number is reused about $4096$ times, so the arithmetic units stay busy.
The cost of attention: $O(n^2 d)$ core
In a Transformer, every token looks at every other token to decide what is relevant. For a sequence of $n$ tokens that is $n \times n$ pairs. Each pair needs a dot product of length $d$ (the size of the query and key vectors). So the cost is $n \times n \times d$.
The key fact: the cost grows with the square of the sequence length. A document twice as long costs four times as much. A document 100 times longer costs 10 000 times as much. Everything else in the network (feed-forward layers) only grows in proportion to $n$.
One attention head with $n = 1000$ tokens and $d = 64$.
- Scores $S = QK^\top$: a $(1000\times64)$ times $(64\times1000)$ product. Result: $1000\times1000$ scores, each a dot product of length 64. Work: $1000^2\times64 = 6.4\times10^{7}$ multiply–adds.
- Softmax over each row: about $10^6$ operations, small.
- Output $= \text{softmax}(S)V$: a $(1000\times1000)$ times $(1000\times64)$ product: another $6.4\times10^{7}$.
- Total about $1.3\times10^{8}$ multiply–adds. At $n = 100\,000$: $10^4$ times more, $1.3\times10^{12}$.
- Memory: the score matrix holds $n^2$ numbers. For $n = 100\,000$ in
float16that is $2\times10^{10}$ bytes = 20 GB, for one head in one layer.
For sequence length $n$ and head size $d$, scaled dot-product attention $\text{softmax}\!\left(\dfrac{QK^\top}{\sqrt d}\right)V$ costs
$$\underbrace{n^2 d}_{QK^\top} + \underbrace{n^2}_{\text{softmax}} + \underbrace{n^2 d}_{\text{(weights)}\,V} \;=\; O(n^2 d)\ \text{time}, \qquad O(n^2)\ \text{memory for the scores}.$$With model width $D$ (all heads together), one layer costs about $2n^2D$ multiply–adds in attention and about $12\,nD^2$ in the projections and feed-forward network. They are equal at $n \approx 6D$: beyond that, attention dominates.
The division by $\sqrt d$ is a numerical trick in the same spirit as the stable softmax: dot products of random length-$d$ vectors have size about $\sqrt d$, which without the division would push softmax into saturation (one entry $\approx1$, the rest $\approx 0$, near-zero gradients).
Why do we need it?
Every token compares itself with every other token, so the cost grows with the square of the sequence length. We must see this to understand context limits.
Where is it used?
Every Transformer: language models, translation, vision transformers, speech models. Tricks such as FlashAttention, sparse or sliding-window attention, linear attention and the KV cache all exist because of this cost.
How is it used?
Estimate the cost as n²d multiply–adds and n² numbers of memory per head. Doubling the context makes attention four times as expensive. Use FlashAttention to avoid storing the n by n table, and a KV cache when generating text.
The $n^2$ memory is often the real limit, long before the time. FlashAttention computes exactly the same result in small tiles and never writes the full $n\times n$ matrix to memory (memory $O(n)$), though the number of flops is still $O(n^2 d)$. Other approaches (sparse, sliding-window, low-rank or linear attention) really reduce the flops, usually at some cost in quality.
Quick check: a model reads 4096 tokens. You move to 16 384 tokens. How much more does the attention part cost?
The sequence is 4 times longer, so attention costs $4^2 = 16$ times as much (and the score matrix needs 16 times the memory). The rest of the layer costs only 4 times as much.
Recap, cheat sheet and practice
- Floats store a sign, an exponent and a mantissa: relative precision $\varepsilon = 2^{-M}$ ($10^{-16}$ in
float64, $10^{-7}$ infloat32). Range and precision are traded off infloat16/bfloat16. Overflow gives Inf, underflow gives 0, undefined gives NaN. - Errors: every operation rounds. Subtracting nearly equal numbers (cancellation) exposes old errors, and addition is not associative, so long sums depend on order. Rewrite formulas, add small numbers first, use pairwise or compensated sums.
- Conditioning ($\kappa = \sigma_{\max}/\sigma_{\min}$) belongs to the problem: you lose about $\log_{10}\kappa$ digits. Stability belongs to the algorithm. Normal equations square $\kappa$; QR does not. MGS beats CGS; Householder beats both. Always pivot.
- ML tricks: subtract the max (stable softmax), use log-sum-exp and log-softmax, work in log space for products of probabilities, add jitter or $\varepsilon$ on the diagonal.
- Iterative methods (Jacobi, Gauss–Seidel, gradient descent, conjugate gradient) need only matrix–vector products. GD needs about $\kappa$ steps; CG about $\sqrt\kappa$. Sparse formats (COO, CSR, CSC) cost $O(\text{nnz})$.
- Cost: dot $O(n)$, mat-vec $O(mn)$, mat-mat and factorisations $O(n^3)$. Matmul reuses data, so GPUs love it. Attention costs $O(n^2d)$ time and $O(n^2)$ memory.
Cheat sheet
| Idea | Formula / rule | Remember |
|---|---|---|
| Float value | $(-1)^s\,2^{e-\text{bias}}\,(1 + m/2^M)$ | sign, exponent, mantissa |
| Machine epsilon | $2^{-M}$: $2.2\times10^{-16}$ / $1.2\times10^{-7}$ | relative gap at 1 |
| Compare floats | $|a-b| \le \text{tol}\cdot\max(|a|,|b|)$ | np.isclose, never == |
| Cancellation cure | rewrite algebraically; log1p, expm1 | avoid subtracting near-equals |
| Condition number | $\kappa = \sigma_{\max}/\sigma_{\min}$; lose $\log_{10}\kappa$ digits | property of the problem |
| Least squares | QR or SVD, not $A^\top A$ | normal equations: $\kappa^2$ |
| Stable softmax | $e^{z_i-m}/\sum_j e^{z_j-m}$, $m=\max z$ | LSE $= m + \log\sum e^{z-m}$ |
| Jitter | $K + \varepsilon I$, $\varepsilon\approx10^{-6}$ | lifts eigenvalues by $\varepsilon$ |
| Iterations | GD $\sim\kappa$, CG $\sim\sqrt\kappa$ | lower $\kappa$: scale, regularise, precondition |
| Sparse | CSR: data, indices, indptr | mat-vec $O(\text{nnz})$ |
| Costs | dot $n$; mat-vec $mn$; mat-mat $n^3$; Cholesky $n^3/3$ | double $n$: ×2, ×4, ×8 |
| Attention | $O(n^2d)$ time, $O(n^2)$ memory | double context: ×4 |
import time
import math
import numpy as np
import scipy.sparse as sp
from scipy.linalg import hilbert, solve_triangular
# 1. floating point basics ---------------------------------------------
print(0.1 + 0.2 == 0.3) # False
print(np.isclose(0.1 + 0.2, 0.3)) # True
print(np.finfo(np.float32).eps, np.finfo(np.float64).eps) # 1.19e-07 2.22e-16
x = np.float32(2**24)
print(x + 1 == x) # True: the gap at 2**24 is 2
# sums depend on order
v = np.array([1e16] + [1.0] * 1000)
print(sum(v)) # 1e16 (every 1.0 was lost)
print(sum(v[::-1])) # 1.0000000000001e16 (ones first)
print(math.fsum(v), np.sum(v)) # 1.0000000000001e16 (exact) and about 1.0000000000001e16 (numpy's pairwise sum keeps most of the ones)
# 2. naive vs stable softmax -------------------------------------------
def softmax_naive(z):
e = np.exp(z)
return e / e.sum()
def softmax_stable(z):
e = np.exp(z - z.max()) # largest exponent becomes 0
return e / e.sum()
def logsumexp(z):
m = z.max()
return m + np.log(np.exp(z - m).sum())
z = np.array([1000.0, 1001.0, 1002.0])
print(softmax_naive(z)) # [nan nan nan] (overflow warning)
print(softmax_stable(z)) # [0.0900 0.2447 0.6652]
print(z - logsumexp(z)) # log-softmax, never takes log(0)
# 3. normal equations vs QR on the Hilbert matrix -----------------------
for n in (4, 8, 10, 12):
A = hilbert(n)
x_true = np.ones(n)
b = A @ x_true
x_ne = np.linalg.solve(A.T @ A, A.T @ b) # normal equations
Q, R = np.linalg.qr(A)
x_qr = solve_triangular(R, Q.T @ b) # QR
print(n, f"cond={np.linalg.cond(A):.1e}",
f"NE error={np.abs(x_ne - x_true).max():.1e}",
f"QR error={np.abs(x_qr - x_true).max():.1e}") # NE is far worse
# 4. conjugate gradient, iterations vs condition number ----------------
def cg(A, b, tol=1e-8, maxit=100000):
x = np.zeros_like(b); r = b.copy(); p = r.copy()
rs = r @ r; bn = np.linalg.norm(b)
for k in range(maxit):
if np.sqrt(rs) <= tol * bn:
return x, k
Ap = A @ p
alpha = rs / (p @ Ap) # best step along p
x += alpha * p
r -= alpha * Ap # one matrix-vector product per step
rs_new = r @ r
p = r + (rs_new / rs) * p # new conjugate direction
rs = rs_new
return x, maxit
for kappa in (1, 10, 100, 1000):
A = np.diag(np.linspace(1, kappa, 200))
print(kappa, cg(A, np.ones(200))[1], "iterations") # 1, 29, 70, 89: far fewer than kappa; bounded by about sqrt(kappa)
# 5. sparse matrices ----------------------------------------------------
A = sp.random(10_000, 10_000, density=0.001, format="csr", random_state=0)
x = np.ones(10_000)
y = A @ x # about 1e5 multiplications, not 1e8
print(A.nnz, A.indptr[:5], A.indices[:5], A.data[:3])
A_coo, A_csc = A.tocoo(), A.tocsc() # other formats
# 6. benchmark matmul and fit the O(n^3) curve --------------------------------
sizes = [100, 200, 400, 800, 1600]
times = []
for n in sizes:
M = np.random.rand(n, n)
t0 = time.perf_counter()
M @ M
times.append(time.perf_counter() - t0)
slope = np.polyfit(np.log(sizes), np.log(times), 1)[0]
print("measured exponent:", slope) # between 2 and 3 (about 2.5 on one test machine): BLAS is cache-friendly and multi-threaded,
# so small sizes are dominated by overhead and the curve is flatter than n^3
1. In float64, why is 0.1 + 0.2 == 0.3 false?
2. A least-squares problem has $\kappa(A) = 10^{9}$ and you use float64. Which statement is right?
3. Why does the stable softmax subtract the maximum score?
4. A system has $\kappa = 10^{4}$. Roughly how do gradient descent and conjugate gradient compare?
5. You double $n$ in a dense $n\times n$ matrix–matrix multiplication. The work grows by a factor of…
6. A language model's context grows from 8 000 to 32 000 tokens. By what factor does the attention score computation grow?
Practice problems
A. What number does the 32-bit pattern 0 10000000 10000000000000000000000 represent?
Sign $0$ means positive. The exponent field is $10000000_2 = 128$, so the true exponent is $128 - 127 = 1$. The mantissa field $1000\ldots_2$ means $0.5$, so $1 + 0.5 = 1.5$. Value $= 1.5 \times 2^{1} = 3$.
B. Rewrite $\sqrt{x^2 + 1} - 1$ so it does not lose digits for tiny $x$.
Multiply by the conjugate: $\sqrt{x^2+1} - 1 = \dfrac{(x^2+1) - 1}{\sqrt{x^2+1} + 1} = \dfrac{x^2}{\sqrt{x^2+1}+1}$. For $x = 10^{-8}$ the original gives $0$ in float64 (because $1 + 10^{-16}$ rounds to $1$), while the new form gives the correct $5\times10^{-17}$.
C. Compute the stable softmax and the log-sum-exp of $\mathbf{z} = [2000, 2000]$.
The maximum is $2000$, so the shifted scores are $[0, 0]$, giving $e^0 = 1$ twice. Softmax $= [0.5, 0.5]$. $\text{LSE} = 2000 + \ln(1 + 1) = 2000 + 0.6931 = 2000.6931$.
D. $\kappa(A) = 10^{5}$ in float32 (about 7 digits). How many digits do QR and the normal equations deliver?
QR: about $7 - 5 = 2$ digits. Normal equations: $\kappa^2 = 10^{10}$, far above $1/u \approx 10^{7}$: no correct digits. This is why float32 least squares needs particular care.
E. Write the CSR arrays for $\begin{bmatrix} 0 & 5 & 0 \\ 0 & 0 & 0 \\ 2 & 0 & 3 \end{bmatrix}$.
Reading row by row, the non-zeros are $5$ (column 1), then $2$ (column 0) and $3$ (column 2). So data = [5, 2, 3], indices = [1, 0, 2]. Row 0 has one entry, row 1 has none, row 2 has two, so indptr = [0, 1, 1, 3].
F. Estimate the multiply–adds for one attention head with $n = 4096$ and $d = 64$, and what happens at $n = 8192$.
$QK^\top$ costs $n^2 d = 4096^2\times64 \approx 1.07\times10^{9}$, and the weights-times-$V$ product costs the same, so about $2.1\times10^{9}$ in total. At $n = 8192$ the cost is $4\times$ larger, about $8.6\times10^{9}$.
Tensors & Array Programming
Every ML model is written as operations on arrays with several axes. This chapter teaches you to read and predict their shapes, to let NumPy or PyTorch do the looping for you, and to describe complicated products in one line.
- See scalars, vectors, matrices and higher-dimensional tensors as one idea: arrays with axes
- Reshape, transpose, permute, squeeze and unsqueeze, and know what really happens in memory
- Predict the result of broadcasting, and spot the bugs it can hide
- Replace loops with matrix operations, including the pairwise distance matrix
- Use batched matrix products and
einsumnotation with confidence
Tensors: arrays with axes
You already know three kinds of arrays. A single number (a scalar). A list of numbers (a vector). A table of numbers (a matrix). A tensor is the same idea with as many directions as you like.
Think of a stack of tables, like a pile of spreadsheets: to find a number you need three addresses (which sheet, which row, which column). Add a shelf of piles and you need four. Each address is called an axis. In machine learning, a "tensor" just means an n-dimensional array of numbers.
- A temperature: $23.5$. A scalar, no axes.
- The heights of 5 people: $[1.6, 1.7, \ldots]$. A vector, one axis,
shape (5,). - A grey-scale photo, 28 rows and 28 columns: a matrix,
shape (28, 28). - A batch of 32 sentences, each with 10 words, each word an embedding of 64 numbers: a 3-axis tensor
(batch, seq, features) = (32, 10, 64). - A batch of 32 colour photos, 3 colour channels, 224 rows, 224 columns:
(batch, channels, H, W) = (32, 3, 224, 224). It holds $32\times3\times224\times224 = 4\,816\,896$ numbers, which is about $19.3$ MB infloat32.
A tensor (or array) is a block of numbers addressed by $n$ indices. Words to know:
- ndim (number of axes): $0$ for a scalar, $1$ vector, $2$ matrix, and so on. (This is not the matrix rank from Chapter 1.8.)
- shape: the list of the axis lengths, for example
(32, 10, 64). Axes are numbered from $0$. Negative numbers count from the end: axis $-1$ is the last one. - size: the total number of elements, the product of the shape.
- dtype: the number format of every element (
float32,int64, ...), from Chapter 1.15.
Indexing: x[i, j, k] is one number. x[i] is the whole sub-tensor with the first index fixed, so it has one axis fewer. By convention the first axis is the batch (many examples processed together), and the last axis is the feature axis.
Layout conventions differ: PyTorch images are usually (N, C, H, W) ("channels first"), TensorFlow often uses (N, H, W, C) ("channels last").
Why do we need it?
Data in ML has many axes: examples, positions, features, channels. We need one clear idea, the tensor, to describe all of it and to know what every number means.
Where is it used?
Images (batch, channels, height, width), text (batch, sequence, features), audio, video, and every weight and activation in PyTorch, TensorFlow and JAX. x.shape is the most printed line in deep learning.
How is it used?
Write the shape of every tensor next to the code. Count the axes with x.ndim, the elements with x.size. Keep the batch first. Print x.shape whenever something looks wrong; shape errors are the most common bug.
A vector of shape (5,) is not the same as a column (5, 1) or a row (1, 5). All three hold 5 numbers, but they behave differently in operations (this causes real bugs, see the broadcasting sections).
"Rank" has two meanings. In NumPy and PyTorch, the number of axes is called ndim (or "rank" in TensorFlow). The rank of a matrix in linear algebra is a different idea (the number of independent columns).
Quick check: what is the shape and the number of elements of a batch of 8 sentences, each of 20 tokens, with 128 features per token?
Shape (8, 20, 128), ndim $=3$, size $8\times20\times128 = 20\,480$.
Reshape, view, transpose and permute
In memory, every tensor is just one long row of numbers. The shape is only a way of reading that row. Imagine a long ribbon of beads. Reshape keeps the same ribbon, in the same order, and just chops it into rows of a different length: no bead moves. This is cheap, because nothing is copied.
Transpose and permute are different: they change which direction counts as "next". The beads stay where they are, but you now read them in a new order. The numbers in the result are not just re-cut, they are re-ordered.
Take the numbers $0, 1, \ldots, 11$ stored in order, and read them as a $3\times4$ matrix (rows are filled first, "row-major"):
$$X = \begin{bmatrix} 0 & 1 & 2 & 3 \\ 4 & 5 & 6 & 7 \\ 8 & 9 & 10 & 11 \end{bmatrix}.$$- Reshape to $4\times3$: keep the order and cut every 3: $\begin{bmatrix} 0 & 1 & 2 \\ 3 & 4 & 5 \\ 6 & 7 & 8 \\ 9 & 10 & 11 \end{bmatrix}$.
- Transpose $X^\top$ (also $4\times3$): rows become columns: $\begin{bmatrix} 0 & 4 & 8 \\ 1 & 5 & 9 \\ 2 & 6 & 10 \\ 3 & 7 & 11 \end{bmatrix}$.
- Both have shape $(4, 3)$, but they are different arrays. Reshape is not transpose.
x.reshape(new_shape): same elements in the same row-major order, new shape. The sizes must multiply to the same total. One axis may be written as-1("work it out for me"). Usually this returns a view (shares memory withx: changing one changes the other).x.Torx.transpose(1, 0)swaps the two axes of a matrix.x.permute(p)(PyTorch) orx.transpose(p)/np.transpose(NumPy) reorders any axes:y.shape[k] = x.shape[p[k]].- Strides are how a tensor knows its layout: the number of elements to skip in memory to move one step along each axis. For a fresh row-major tensor of shape $(a, b, c)$, the strides are $(bc,\ c,\ 1)$. Permuting just permutes the strides, so no data moves, but the tensor is no longer contiguous.
- PyTorch's
.view()only works on contiguous tensors. After a permute, call.contiguous()(it copies the data into a fresh row-major layout) or use.reshape()(which copies only if it must).
Why do we need it?
We often need the same numbers arranged differently: flatten an image for a linear layer, or split a feature axis into heads. We must know which operations move data and which only re-read it.
Where is it used?
Flattening before a dense layer, multi-head attention (reshape then permute), switching between NCHW and NHWC image layouts, and writing X.T @ X for a covariance matrix.
How is it used?
Use reshape to re-cut the same numbers and transpose or permute to reorder axes. Never use reshape when you mean transpose. Call .contiguous() after a permute if .view() complains. Copy with .clone() if you need independence.
Reshape is not transpose. To turn $(N, T, F)$ into $(N, F, T)$ you need permute / transpose, never reshape. Reshaping would scramble the numbers silently: the shape would be right and the answer wrong.
Views share memory. After y = x.reshape(...), writing into y may also change x. If you need an independent copy, use .copy() (NumPy) or .clone() (PyTorch).
Flatten before a linear layer: x.reshape(batch, -1) keeps the batch axis and merges the rest into one feature axis.
Quick check: x has shape (2, 3, 4). What is the shape of x.permute(2, 0, 1)? And of x.reshape(4, 6)?
permute(2, 0, 1) puts old axis 2 first, then old axis 0, then old axis 1: shape (4, 2, 3). reshape(4, 6) gives shape (4, 6) (and $2\cdot3\cdot4 = 24 = 4\cdot6$).
Adding and removing axes: unsqueeze and squeeze
An axis of length 1 costs no memory and holds no extra numbers, it is only a label that says "this direction exists, but has one slot". Why would you want one? Because a column $(5, 1)$ and a row $(1, 5)$ behave differently from a plain list $(5,)$. Inserting a length-1 axis tells the computer which direction a list should extend in. Unsqueeze adds such an axis, squeeze removes the length-1 axes.
A vector v has shape $(3,)$ and holds $[1, 2, 3]$.
v.unsqueeze(0)(orv[None, :]) gives shape $(1, 3)$: a row, $[[1, 2, 3]]$.v.unsqueeze(1)(orv[:, None]) gives shape $(3, 1)$: a column, $[[1],[2],[3]]$.x.squeeze()on a tensor of shape $(3, 1, 4, 1)$ removes both length-1 axes: shape $(3, 4)$.
x.unsqueeze(k)/np.expand_dims(x, k)/x[..., None]inserts a new axis of length 1 at position $k$. The number of elements does not change.x.squeeze()removes all axes of length 1.x.squeeze(k)removes only axis $k$ (and only if its length is 1).- Both return views: no data is copied.
Why do we need it?
A plain list has no row or column direction. A length-1 axis tells the computer which way a list points, which is essential for broadcasting.
Where is it used?
Adding a batch axis to one example before a model, shaping a mask as (batch, 1, 1, length) for attention, and fixing the (N, 1) outputs of a regression head before a loss.
How is it used?
Use x.unsqueeze(k) or x[:, None] to add a length-1 axis, and x.squeeze(k) to remove one. Prefer naming the axis to squeeze, so you never remove a batch axis of size 1 by accident.
Forgetting squeeze/unsqueeze is the classic cause of silent bugs. A model that outputs (N, 1) compared against labels of shape (N,) will broadcast to $(N, N)$ (see the "accidental broadcasting" section). Always make the shapes match deliberately.
squeeze() with no argument can remove an axis you wanted to keep (for example a batch of size 1). Prefer squeeze(k) with an explicit axis.
Quick check: x has shape (1, 5, 1). What do x.squeeze() and x.squeeze(0) give?
x.squeeze() removes both length-1 axes: shape (5,). x.squeeze(0) removes only the first: shape (5, 1).
Broadcasting core
What should [[1,2,3],[4,5,6]] + [10,20,30] mean? The shapes differ, but the sensible reading is: add the list $[10,20,30]$ to every row. The small array is stretched (copied virtually) to match the big one.
That is broadcasting. It lets you add a bias to every example, subtract a mean from every column, or scale every row, without writing a loop and without actually copying the small array. The computer just re-uses it.
Add $\mathbf{b} = [10, 20, 30]$ (shape $(3,)$) to $A$ (shape $(2,3)$):
$$\begin{bmatrix} 1 & 2 & 3 \\ 4 & 5 & 6 \end{bmatrix} + [10,\ 20,\ 30] = \begin{bmatrix} 11 & 22 & 33 \\ 14 & 25 & 36 \end{bmatrix}.$$Now the "outer" case. A column $\mathbf{a} = [[1],[2],[3]]$ (shape $(3,1)$) plus a row $\mathbf{b} = [[10, 20, 30, 40]]$ (shape $(1,4)$) makes a full $3\times4$ table:
$$\begin{bmatrix} 1 \\ 2 \\ 3 \end{bmatrix} + [10,\ 20,\ 30,\ 40] = \begin{bmatrix} 11 & 21 & 31 & 41 \\ 12 & 22 & 32 & 42 \\ 13 & 23 & 33 & 43 \end{bmatrix}.$$The column is copied across the columns, the row is copied down the rows. Both are stretched.
The broadcasting rules. To combine two shapes:
- Line them up from the right (the last axes first). If one shape has fewer axes, imagine extra axes of length $1$ added on its left.
- For each pair of axis lengths, they are compatible if they are equal, or one of them is 1.
- The result length on that axis is the larger of the two. An axis of length 1 is stretched to match.
- If any pair is compatible in neither way, the operation fails with an error.
Example: $(8,1,6,1)$ and $(7,1,5)$ line up as $(8,1,6,1)$ and $(1,7,1,5)$, giving $(8,7,6,5)$. And $(3,4)$ with $(3,)$ lines up as $(3,4)$ with $(1,3)$: the last pair $4$ against $3$ is incompatible, an error.
Broadcasting applies to all element-wise operations ($+,-,\times,\div,$ comparisons, np.maximum, ...). It does not change how the matrix product @ works (that has its own shape rule).
Why do we need it?
We want to add a bias to every example or subtract a mean from every column without writing loops or copying data. Broadcasting stretches the smaller array virtually.
Where is it used?
Bias addition in every layer, standardising features, attention masks and scaling, outer sums and products, and distance formulas. It works the same in NumPy, PyTorch, TensorFlow and JAX.
How is it used?
Line the two shapes up from the right. Each pair of lengths must be equal or one must be 1. Use keepdims=True or None to put length-1 axes where you want the stretching. Predict the result shape before running the code.
Broadcasting is silent. If the shapes happen to be compatible, NumPy and PyTorch do it without a warning, even if you did not mean it. The next section shows what that can do.
Broadcasting is not copying. A broadcast array uses no extra memory. But the result is a real array of the full broadcast shape, so a $(N,1)$ with an $(1,N)$ makes $N^2$ numbers.
Quick check: do shapes $(4, 1, 3)$ and $(5, 3)$ broadcast? To what shape?
Align from the right: $(4,1,3)$ with $(1,5,3)$. Axis pairs: $3$ and $3$ equal; $1$ and $5$ (stretch to 5); $4$ and $1$ (stretch to 4). Result: (4, 5, 3).
Common bugs from accidental broadcasting core
Broadcasting is helpful when you mean it and dangerous when you don't. An error message is a gift: it tells you at once that your shapes are wrong. Broadcasting that quietly succeeds is worse: your code runs, produces numbers, and the numbers are wrong.
The classic trap: a list of $N$ predictions $(N,)$ against a column of $N$ targets $(N,1)$ (or the other way round). Broadcasting pairs every prediction with every target and returns an $N\times N$ table instead of $N$ differences.
Targets $y = [1, 2, 3, 4]$, shape $(4,)$. A perfect model predicts exactly $[[1],[2],[3],[4]]$, shape $(4, 1)$. The mean squared error should be $0$.
- The code computes
pred - y: shapes $(4,1)$ and $(4,)$, which broadcast to $(4,4)$. Entry $(i,j)$ is $\text{pred}_i - y_j$. - The table has non-zero numbers, for example $\text{pred}_1 - y_4 = 1 - 4 = -3$.
- The mean of the 16 squared entries is $\dfrac{(0+1+4+9) + (1+0+1+4) + (4+1+0+1) + (9+4+1+0)}{16} = \dfrac{40}{16} = 2.5$.
The code says the error is $2.5$ for a model that is perfect. No error message was shown.
Defences against accidental broadcasting:
- Check shapes at the boundaries:
assert pred.shape == y.shape. - Make a vector's role explicit:
pred.squeeze(-1)ory[:, None]. - Use
keepdims=Truewhen reducing, so the result still has the axis (and broadcasts the way you expect):x.mean(axis=1, keepdims=True)has shape(N, 1). - Beware of square arrays: a $(N,N)$ matrix minus a $(N,)$ vector of row means always "works", but the vector lines up with the columns, so entry $j$ is subtracted from column $j$ and not from row $j$. A shape error would have warned you. Test with $N \ne D$.
- Do a quick sanity check: the result of an operation on $(N,)$ should have $N$ elements, not $N^2$.
Why do we need it?
Broadcasting never complains when shapes happen to fit, so a mistake can produce wrong numbers with no error message. We need habits that catch it.
Where is it used?
Mean-squared-error losses with (N, 1) against (N,), mean subtraction over the wrong axis, binary classifier outputs against labels, and square matrices that hide axis mix-ups.
How is it used?
Assert shapes at the edges, for example assert pred.shape == y.shape. Use keepdims=True when reducing. Test with non-square sizes (N different from D). Check that an operation on N items returns N numbers, not N².
Silent bugs are hard to find because training still runs, the loss is just strange (often it plateaus at a suspicious value). When a model "does not learn", print every shape in the loss computation first.
The rule of thumb: never rely on broadcasting between two arrays that both came from the network. Use it for deliberate stretching (a bias, a mean, a mask) and nothing else.
Quick check: X has shape (5, 3). You compute X - X.mean(axis=1). What happens, and how do you fix it?
X.mean(axis=1) has shape (5,), which lines up against the last axis (length 3): $5$ vs $3$ is an error here. If $X$ were $(5,5)$ it would silently subtract the wrong thing. Fix: X - X.mean(axis=1, keepdims=True), with mean of shape (5, 1), so each row has its own mean subtracted.
Vectorisation: replacing loops with matrix operations core
Imagine moving 1000 boxes. You can carry them one at a time, walking back and forth 1000 times (a Python for loop). Or you can load them onto a truck and drive once (an array operation). The work on each box is the same, but the truck skips the walking.
A Python loop makes the interpreter handle one number at a time, with lots of overhead. A vectorised operation hands the whole array to fast compiled code (often BLAS, a very well tuned maths library) that streams through the numbers. It is typically 10 to 1000 times faster. It also reads like the maths: y = X @ w is the formula $\mathbf{y} = X\mathbf{w}$.
The predictions of a linear model for $n$ examples (the rows of $X$) are $\hat y_i = \mathbf{x}_i\cdot\mathbf{w} + b$.
Vectorised, the same thing is one line: y = X @ w + b. The matrix–vector product $X\mathbf{w}$ computes all $n$ dot products at once, and + b broadcasts the single number $b$ over all $n$ results.
Standardising the columns of $X$:
Recipe for vectorising:
- Find what the loop repeats: "for each row", "for each pair", "for each column".
- Put that repeated index into an axis of an array.
- Write the loop body as an operation on whole arrays: element-wise operations, broadcasting, reductions along an axis (
sum(axis=...)) and matrix products (@).
Useful tools: np.sum(x, axis=k) (collapses axis $k$), keepdims=True, np.where(cond, a, b) (vector "if"), X @ Y, np.einsum (later in this chapter). Vectorising is not about hiding a loop: the loop is still done, but in fast compiled code, and in an order the hardware likes.
Why do we need it?
Python loops handle one number at a time and are very slow. Whole-array operations run in fast compiled code and read like the maths.
Where is it used?
Predictions for a whole dataset (X @ w), standardising columns, row-wise softmax, batched forward passes on GPUs, and everything that automatic differentiation has to handle.
How is it used?
Find what the loop repeats, turn that index into an axis, then use element-wise operations, reductions along an axis, broadcasting and matrix products. Compare with the loop on a small example using np.allclose before trusting it.
Vectorised code can use a lot of memory. The "all pairs" trick builds an $N\times N$ table. For huge $N$, process in chunks (mini-batches of rows) instead of everything at once.
Not every loop vectorises. A loop where step $t$ needs the result of step $t-1$ (a recurrence, like an RNN over time) cannot be turned into one array operation along that axis. Vectorise over the batch and keep the time loop.
Do not loop over rows to fill a matrix in NumPy or PyTorch unless you truly have to. If you see for i in range(n) around array code, ask: which axis could this index become?
Quick check: how would you compute, without a loop, the mean of every row of a matrix X of shape $(n, d)$ so that you can subtract it from X?
X - X.mean(axis=1, keepdims=True). The mean has shape (n, 1), which broadcasts across the $d$ columns, so each row loses its own mean.
The pairwise distance matrix without loops core
A very common job: you have $n$ points and $m$ other points, and you want the distance from every one to every other, an $n\times m$ table. For k-nearest neighbours, k-means, kernels and clustering, this table is the first step.
The slow way is a double loop. The fast way uses an old school identity: the squared distance is $\|\mathbf{x}-\mathbf{y}\|^2 = \|\mathbf{x}\|^2 + \|\mathbf{y}\|^2 - 2\,\mathbf{x}\cdot\mathbf{y}$. The first two terms are one number per point, and the last term for all pairs at once is just one matrix product $XY^\top$.
$\mathbf{x} = [1, 2]$ and $\mathbf{y} = [4, 6]$.
- $\|\mathbf{x}\|^2 = 1 + 4 = 5$ and $\|\mathbf{y}\|^2 = 16 + 36 = 52$.
- $\mathbf{x}\cdot\mathbf{y} = 4 + 12 = 16$, so $2\,\mathbf{x}\cdot\mathbf{y} = 32$.
- Squared distance $= 5 + 52 - 32 = 25$, so the distance is $5$ ✓. (Direct check: $(4-1)^2 + (6-2)^2 = 9 + 16 = 25$.)
For $X \in \mathbb{R}^{n\times d}$ (rows are points) and $Y\in\mathbb{R}^{m\times d}$, the matrix of squared distances is
$$D^2_{ij} = \|\mathbf{x}_i\|^2 + \|\mathbf{y}_j\|^2 - 2\,\mathbf{x}_i\cdot\mathbf{y}_j, \qquad D^2 = \underbrace{\mathbf{p}\,\mathbf{1}_m^\top}_{n\times m} + \underbrace{\mathbf{1}_n\,\mathbf{q}^\top}_{n\times m} - 2\,XY^\top,$$where $\mathbf{p}_i = \|\mathbf{x}_i\|^2$ and $\mathbf{q}_j = \|\mathbf{y}_j\|^2$. In NumPy, broadcasting builds the first two terms for free:
The heavy part is one matrix product (fast BLAS). The clip matters: because of rounding the formula can return a tiny negative squared distance, and the square root of a negative number is NaN.
Why do we need it?
Many algorithms need the distance from every point to every other point. A double loop is far too slow; one matrix product computes the whole table.
Where is it used?
k-nearest neighbours, k-means clustering, RBF kernels in SVMs and Gaussian processes, contrastive learning on embeddings, and retrieval and semantic search.
How is it used?
Compute the squared norms of both sets, then x2 + y2 - 2·X @ Y.T with broadcasting, clip at 0, then take the square root. Centre the data first, process large sets in chunks, and use float64 if the points are close but far from the origin.
Cancellation again. The formula subtracts $2\,\mathbf{x}\cdot\mathbf{y}$ from $\|\mathbf{x}\|^2+\|\mathbf{y}\|^2$. When two points are close but both are far from the origin, those numbers are huge and nearly equal, and a catastrophic cancellation (Chapter 1.15) destroys the small difference. The fix: centre the data first (subtract the mean), compute in float64, or use the direct formula on small chunks.
Always clip at zero before the square root: np.maximum(d2, 0). The diagonal $D_{ii}$ of a self-distance matrix should be exactly 0 but often comes out as $\pm10^{-7}$.
Quick check: X has shape (n, d) and Y has shape (m, d). What are the shapes of x2, y2 and X @ Y.T in the code above?
x2: (n, 1). y2: (1, m). X @ Y.T: $(n,d)\cdot(d,m) = $ (n, m). Adding them broadcasts to (n, m).
Batched matrix multiplication core
One matrix product handles one pair of matrices. But in ML you have a whole batch: 32 sentences, each needing the same kind of product. A batched matrix multiplication does all 32 products in one call. The leading axes are the batch, the last two axes are the matrices, and the product is done for every batch item separately and at the same time.
A useful trick: if one side has no batch axes, it is shared: the same weight matrix is applied to every item (broadcasting again).
- Plain: $(3,4)\ @\ (4,2) \to (3,2)$. The inner $4$ must match and disappears.
- Batched: $(8,3,4)\ @\ (8,4,5) \to (8,3,5)$. Eight independent products of a $3\times4$ with a $4\times5$ matrix.
- Shared weights: a linear layer on a batch of sentences: $(32,10,64)\ @\ (64,128) \to (32,10,128)$. The one $64\times128$ matrix is applied to all $32\times10$ token vectors.
- Attention scores: queries $(B,h,T,d)$ times transposed keys $(B,h,d,T)$ give $(B,h,T,T)$: for every example and every head, a $T\times T$ table of scores. (Queries and keys are the two sets of vectors a Transformer compares; $h$ is the number of heads, $T$ the number of tokens. See Chapter 1.15.)
For A @ B (NumPy, torch.matmul, or torch.bmm when both are exactly 3-D), with shapes $(\ldots, n, k)$ and $(\ldots, k, m)$:
- The last two axes follow the matrix rule: $(n,k)\cdot(k,m)\to(n,m)$, and the two $k$'s must be equal.
- All the leading axes are batch axes. They are broadcast against each other (the rules from before), and the result has the broadcast batch shape.
- A 1-D argument is promoted to a row or column for the product and the extra axis is removed again.
- Cost: $2\cdot(\text{batch size})\cdot n\,k\,m$ flops (a flop is one floating-point add or multiply; each multiply-add counts as 2).
The element-wise operators ($*$) are not matrix products. A * B multiplies matching entries and broadcasts; A @ B does the matrix product.
Why do we need it?
We often run the same matrix product on a whole batch at once. A batched matmul does all of them in one call, and broadcasting lets a weight matrix be shared.
Where is it used?
Every linear layer on a sequence, attention scores QKᵀ and the weighted sum with V, batched covariance and Gram matrices, and torch.bmm, torch.matmul, np.matmul.
How is it used?
Write the shapes: (…, n, k) @ (…, k, m) gives (…, n, m). The two k must match, and the leading batch axes must broadcast. To transpose only the last two axes of a tensor use transpose(-2, -1), never .T.
* is not @. Writing A * B for a matrix product is a very common slip. With compatible shapes it runs and returns an element-wise product.
Transpose the right axes. For a 4-D tensor, K.T reverses all four axes and gives nonsense. You want to swap only the last two: K.transpose(-2, -1) or K.swapaxes(-1, -2).
torch.bmm needs exactly 3-D tensors with equal batch sizes (no broadcasting). torch.matmul and @ are more flexible.
Quick check: what is the shape of (16, 5, 8) @ (16, 8, 3), and what error do you get from (16, 5, 8) @ (16, 5, 8)?
The first is (16, 5, 3): 16 products of $5\times8$ by $8\times3$. The second fails: the last axis of A is $8$ but the second-to-last axis of B is $5$. You would need to transpose B's last two axes first.
Einstein summation: einsum core
Matrix products, transposes, traces, outer products and attention scores all have the same shape of recipe: "multiply some entries together, and add up over some of the indices". einsum lets you write that recipe directly, by naming the axes with letters.
You give each input a string of letters, one per axis. Then you say which letters you keep in the output. Any letter that is not kept is summed over. That is the whole idea. The recipe "ij,jk->ik" reads: "take $A$ with axes $(i,j)$ and $B$ with axes $(j,k)$; the output has axes $(i,k)$; and since $j$ is not in the output, multiply and sum over $j$." That is exactly matrix multiplication.
"ij,jk->ik": $C_{ik} = \sum_j A_{ij}B_{jk}$. Matrix product."ij->ji": $C_{ji} = A_{ij}$. Transpose (no letter is summed)."ii->": $\sum_i A_{ii}$. The trace (a repeated letter walks down the diagonal; the output has no letters, so it is a single number)."i,i->": $\sum_i a_ib_i$. Dot product."i,j->ij": $a_ib_j$. Outer product (nothing summed)."bij,bjk->bik": batched matmul. The batch letter $b$ is kept in the output, so every batch item is handled separately."bhqd,bhkd->bhqk": attention scores. For each batch $b$ and head $h$: $S_{qk} = \sum_d Q_{qd}K_{kd}$. Only $d$ disappears.
np.einsum("subscripts", A, B, ...). Rules:
- Each input has one lower-case letter per axis, separated by commas. Equal letters (in different inputs, or twice in one) refer to the same index: the sizes must agree, and the entries are multiplied at equal index values.
- After
->comes the output's letters, in the order of the output axes. - Letters missing from the output are summed over. Letters present in the output are kept as axes.
- Without
->("implicit mode"), the output is every letter that appears exactly once, in alphabetical order. Writing->explicitly is clearer and safer.
The total work is the product of the sizes of all the letters (kept and summed), one multiply–add per combination.
Why do we need it?
Products, transposes, traces and attention scores are all 'multiply and add over some indices'. Einsum states that recipe directly by naming each axis with a letter.
Where is it used?
Attention scores bhqd,bhkd->bhqk, bilinear forms, covariance matrices, tensor contractions in physics, and np.einsum, torch.einsum, JAX and einops.
How is it used?
Give each input one letter per axis, write the output letters after the arrow, and let every missing letter be summed. Predict the output shape from the output letters first. Always write the arrow explicitly rather than using implicit mode.
Read the output letters first. The output shape is the sizes of the letters after ->, nothing else. If a letter you expected is summed away, the axis disappears.
Same letter = same size. If you use j on two inputs, the two axes must have equal length, or you get an error. This is einsum's built-in shape check, and it is why many people prefer it to a chain of transposes.
Do not rely on implicit mode. "ij,jk" works because of the alphabetical rule, but "ji,jk" silently computes $A^\top B$ (the letter $i$ is now the second axis of $A$), which is not what you may expect. Always write ->.
Quick check: what is the output shape of np.einsum("bhqd,bhkd->bhqk", Q, K) if Q and K are both (4, 8, 10, 64)?
The output letters are $b,h,q,k$ with sizes $4, 8, 10, 10$. Output shape (4, 8, 10, 10). The letter $d$ (size 64) is summed over.
The Kronecker product (awareness)
Take a small matrix $A$. Replace every entry $a_{ij}$ by a whole copy of another matrix $B$, scaled by $a_{ij}$. You get a bigger, block matrix: the Kronecker product $A\otimes B$. Picture a photo made of tiny photos: $A$ says how bright each tile is and $B$ is the picture inside every tile.
Let $A = \begin{bmatrix} 1 & 2 \\ 3 & 4 \end{bmatrix}$ and $B = \begin{bmatrix} 0 & 1 \\ 1 & 0 \end{bmatrix}$.
$$A \otimes B = \begin{bmatrix} 1\cdot B & 2\cdot B \\ 3\cdot B & 4\cdot B \end{bmatrix} = \begin{bmatrix} 0 & 1 & 0 & 2 \\ 1 & 0 & 2 & 0 \\ 0 & 3 & 0 & 4 \\ 3 & 0 & 4 & 0 \end{bmatrix}.$$A $2\times2$ with a $2\times2$ gives $4\times4$. In general a $(m\times n)$ matrix and a $(p\times q)$ matrix give an $(mp\times nq)$ matrix.
Useful facts (awareness level): $(A\otimes B)(C\otimes D) = (AC)\otimes(BD)$; $(A\otimes B)^\top = A^\top\otimes B^\top$; $(A\otimes B)^{-1} = A^{-1}\otimes B^{-1}$. The famous vec trick (with $\text{vec}$ stacking the columns of a matrix into one long column):
$$\text{vec}(A\,X\,B) = (B^\top \otimes A)\,\text{vec}(X).$$It says a product of small matrices on the left and right of $X$ equals a single huge matrix times a vector. You use it in the opposite direction: never build the huge Kronecker matrix, do the cheap small products instead. In NumPy: np.kron(A, B).
Why do we need it?
Sometimes a huge structured matrix is really two small ones combined. The Kronecker product describes that structure so we never have to build the huge matrix.
Where is it used?
K-FAC and Shampoo optimisers (cheap curvature approximations), 2D transforms such as the 2D Fourier transform, grid-structured Gaussian processes, and quantum computing.
How is it used?
Build a block matrix with np.kron(A, B) only for small sizes. For large ones use identities such as vec(AXB) = (Bᵀ ⊗ A) vec(X), which need only the small factors. Check the size: (m·p) by (n·q).
The Kronecker product grows fast: two $1000\times1000$ matrices give a $10^6\times10^6$ matrix with $10^{12}$ entries. Never form it explicitly for large sizes. Use the identities to work with the small factors.
Do not confuse it with the ordinary product $AB$, the element-wise product $A\odot B$, or the outer product $\mathbf{a}\mathbf{b}^\top$ (which is a Kronecker product of a column and a row).
Quick check: what is the shape of $A\otimes B$ when $A$ is $2\times3$ and $B$ is $4\times5$?
$(2\cdot4)\times(3\cdot5) = 8\times15$, with $120$ entries.
Recap, cheat sheet and practice
- A tensor is an n-dimensional array. Its shape lists the axis lengths, ndim counts the axes, size is the product. Typical shapes: $(B,T,F)$ for sequences, $(B,C,H,W)$ for images.
- Reshape re-cuts the same row-major memory (no data moves). Transpose/permute reorder the axes by changing strides (and make the tensor non-contiguous). Reshape is not transpose.
- unsqueeze adds a length-1 axis, squeeze removes it. They decide whether a list acts as a row or a column.
- Broadcasting: align shapes from the right; each axis pair must be equal or contain a 1. It is silent, so accidental broadcasting (like $(N,1)-(N,)\to(N,N)$) is a classic bug.
- Vectorise: turn the loop index into an axis, then use element-wise operations, reductions and matmuls. Pairwise distances: $\|x\|^2 + \|y\|^2 - 2x^\top y$ with broadcasting, clipped at 0.
- Batched matmul multiplies the last two axes and broadcasts the batch axes. einsum names the axes: letters not in the output are summed. Kronecker products make block matrices.
Cheat sheet
| Task | NumPy / PyTorch | Shape rule |
|---|---|---|
| Shape, ndim, size | x.shape, x.ndim, x.size | size = product of the shape |
| Reshape / flatten | x.reshape(a, -1) | same total number of elements |
| Swap / reorder axes | x.T, x.transpose(...), x.permute(...) | y.shape[k] = x.shape[p[k]] |
| Add / remove an axis | x[:, None], unsqueeze, squeeze | length-1 axes only |
| Broadcast | a + b, a * b | align right; equal or 1 |
| Reduce | x.sum(axis=k, keepdims=True) | axis $k$ disappears (or becomes 1) |
| Matrix product | A @ B, np.matmul | $(\ldots,n,k)(\ldots,k,m)\to(\ldots,n,m)$ |
| Pairwise distances | x2 + y2 - 2 * X @ Y.T | $(n,1)+(1,m)-(n,m)\to(n,m)$ |
| einsum | np.einsum("bij,bjk->bik", A, B) | output = sizes of the letters after -> |
| Kronecker | np.kron(A, B) | $(m,n)\otimes(p,q)\to(mp,nq)$ |
import numpy as np
# 1. shapes, reshape, transpose, views ---------------------------------
x = np.arange(24).reshape(2, 3, 4) # (batch=2, seq=3, features=4)
print(x.shape, x.ndim, x.size) # (2, 3, 4) 3 24
print(x.reshape(4, 6).shape, x.reshape(2, -1).shape) # (4, 6) (2, 12)
y = x.transpose(0, 2, 1) # swap seq and features
print(y.shape, y.flags["C_CONTIGUOUS"]) # (2, 4, 3) False: only strides changed
print(np.shares_memory(x, y)) # True: no data was copied
print(x[:, None].shape, np.squeeze(x[:, None], axis=1).shape) # (2, 1, 3, 4) (2, 3, 4)
# 2. broadcasting and its classic bug -------------------------------------
a = np.arange(3)[:, None] # (3, 1)
b = np.arange(4)[None, :] # (1, 4)
print((a + b).shape) # (3, 4)
pred = np.array([[1.0], [2.0], [3.0], [4.0]]) # (4, 1) a perfect model...
target = np.array([1.0, 2.0, 3.0, 4.0]) # (4,)
print(((pred - target) ** 2).mean()) # 2.5 WRONG: (4,1)-(4,) -> (4,4)
print(((pred.squeeze(-1) - target) ** 2).mean()) # 0.0 right
# 3. five loops rewritten without loops -------------------------------------
X = np.random.randn(100, 5); w = np.random.randn(5); v = np.random.randn(50)
def sum_squares_loop(v):
t = 0.0
for e in v: t += e * e
return t
assert np.isclose(sum_squares_loop(v), v @ v)
def predict_loop(X, w, b):
return np.array([X[i] @ w + b for i in range(len(X))])
assert np.allclose(predict_loop(X, w, 0.5), X @ w + 0.5)
def standardise_loop(X):
Z = X.copy()
for j in range(X.shape[1]): Z[:, j] = (X[:, j] - X[:, j].mean()) / X[:, j].std()
return Z
assert np.allclose(standardise_loop(X), (X - X.mean(0)) / X.std(0))
def softmax_rows_loop(X):
return np.array([np.exp(r - r.max()) / np.exp(r - r.max()).sum() for r in X])
E = np.exp(X - X.max(axis=1, keepdims=True))
assert np.allclose(softmax_rows_loop(X), E / E.sum(axis=1, keepdims=True))
def dist_loop(X, Y):
return np.array([[np.linalg.norm(a - c) for c in Y] for a in X])
def dist_vec(X, Y):
d2 = (X**2).sum(1)[:, None] + (Y**2).sum(1)[None, :] - 2 * X @ Y.T
return np.sqrt(np.maximum(d2, 0))
# a distance of exactly 0 comes out near 3e-8 (rounding, then sqrt makes it bigger), so allow a small atol
assert np.allclose(dist_loop(X[:20], X[:30]), dist_vec(X[:20], X[:30]), atol=1e-6)
# 4. batched attention scores with einsum --------------------------------
B, h, T, d = 2, 4, 5, 8
Q, K, V = (np.random.randn(B, h, T, d) for _ in range(3))
scores = np.einsum("bhqd,bhkd->bhqk", Q, K) / np.sqrt(d) # (B, h, T, T)
assert np.allclose(scores, Q @ K.transpose(0, 1, 3, 2) / np.sqrt(d))
wts = np.exp(scores - scores.max(-1, keepdims=True)) # stable softmax
wts /= wts.sum(-1, keepdims=True)
out = np.einsum("bhqk,bhkd->bhqd", wts, V) # (B, h, T, d)
print(scores.shape, out.shape)
# 5. pairwise distances for 10,000 points, no loops over pairs ---------------
P = np.random.randn(10_000, 50)
nearest = np.empty(len(P), dtype=int)
for s in range(0, len(P), 1000): # chunks: a 1000 x 10000 table, not 10000 x 10000
D2 = (P[s:s+1000]**2).sum(1)[:, None] + (P**2).sum(1)[None, :] - 2 * P[s:s+1000] @ P.T
D2[np.arange(len(D2)), np.arange(s, s + len(D2))] = np.inf # ignore each point itself
nearest[s:s+1000] = D2.argmin(axis=1)
print(nearest[:5])
# 6. Kronecker product and the vec trick --------------------------------
A, Bm, Xm = np.random.randn(2, 3), np.random.randn(4, 5), np.random.randn(3, 4)
print(np.kron(A, Bm).shape) # (8, 15)
vec = lambda M: M.reshape(-1, order="F") # stack the columns
print(np.allclose(vec(A @ Xm @ Bm), np.kron(Bm.T, A) @ vec(Xm))) # True
1. What is the shape of a + b if a has shape $(3, 1, 4)$ and b has shape $(5, 1)$?
2. x = np.arange(6). Compare x.reshape(2, 3).T with x.reshape(3, 2).
3. A batch of token vectors has shape $(B, T, D)$. It is multiplied by a weight matrix of shape $(D, H)$ with @. The result has shape…
4. In np.einsum("bij,bjk->bik", A, B), which index is summed over?
-> (b, i, k) are kept. The letter missing from the output (j) is multiplied and summed. This is a batched matrix product.5. Why do we clip with np.maximum(d2, 0) in the pairwise-distance formula before taking the square root?
sqrt of a negative number.6. x has shape $(32, 1)$ and y has shape $(32,)$. What is the shape of x - y?
x is paired with every entry of y. A silent bug if you wanted element-wise differences.Practice problems
A. A has shape $(4,3,5)$ and B has shape $(4,5,2)$. What are the output shape and the number of multiply–adds of einsum("bij,bjk->bik", A, B)?
Output letters $b,i,k$ have sizes $4, 3, 2$: shape (4, 3, 2). The summed letter $j$ has size $5$. Total multiply–adds $= 4\cdot3\cdot5\cdot2 = 120$.
B. Broadcast $(8,1,6,1)$ with $(7,1,5)$. Result shape and number of elements?
Align: $(8,1,6,1)$ with $(1,7,1,5)$. Result $(8,7,6,5)$, with $8\cdot7\cdot6\cdot5 = 1680$ elements.
C. x has shape $(2,3,4)$. What are the shapes of x.permute(1,2,0) and x.reshape(6,4)? Is the permuted tensor contiguous?
permute(1,2,0): new axes are old axes 1, 2, 0, so shape (3, 4, 2). reshape(6,4): shape (6, 4). The permuted tensor is not contiguous (its strides $(4,1,12)$ are not the row-major strides $(8,2,1)$ of a $(3,4,2)$ array).
D. $X$ has 5 points and $Y$ has 3 points, both in $\mathbb{R}^2$. Give the shapes in x2 + y2 - 2 * X @ Y.T.
x2: $(5,1)$. y2: $(1,3)$. X @ Y.T: $(5,2)\cdot(2,3) = (5,3)$. Sum: $(5,1)+(1,3)\to(5,3)$.
E. Rewrite without a loop: s = 0; for i in range(n): s += X[i] @ X[i].
It adds the squared length of every row, i.e. the sum of all squared entries: s = (X * X).sum() (or np.sum(X**2), or np.einsum("ij,ij->", X, X)).
F. $A$ is $3\times3$ and $B$ is $4\times4$. How large is $A\otimes B$ and why should you avoid building it when both are $1000\times1000$?
$A\otimes B$ is $12\times12$. For two $1000\times1000$ matrices it would be $10^6\times10^6$ with $10^{12}$ entries (8 terabytes). Use identities like the vec trick, which needs only products of the small matrices.
ML Applications
This is the payoff. Every chapter so far was a tool. Now we open the machine-learning toolbox and find the same few tools in every drawer: dot products, matrix products, projections, eigenvectors and the SVD. By the end, regression, PCA, embeddings, recommenders and the Transformer will all look like old friends.
- Write linear regression as projection, solve it in closed form and by gradient descent, and add ridge and lasso penalties
- See logistic and softmax regression as a linear score plus a squashing function, with a convex loss
- Build covariance matrices, measure distance with Mahalanobis, and run PCA two ways (eigen and SVD)
- Treat words and items as vectors (embeddings), and fill in missing ratings with matrix factorisation
- Read a neural-network layer, initialisation, backprop and normalisation as matrix operations
- Know the name of the matrix-efficiency idea (low-rank factorisation) and follow LoRA from the problem it solves to training, merging and real use
- Compute attention step by step, understand the $\sqrt d$ scaling and the cost $O(n^2 d)$, and see why batched matrix multiplication matters
- Recognise linear algebra in spectral clustering, kernels, Gaussian processes, Markov chains and graph networks
How this chapter works. Each section starts with a line called What we need from earlier chapters, with links back to the exact ideas used. If a link feels shaky, click it, refresh your memory, and come back.
Here is the map of where we are going. Read it as "ML idea → the linear algebra underneath".
| ML idea | The linear algebra inside | Chapters used |
|---|---|---|
| Linear regression | Projection onto the column space; normal equations | 1.9, 1.10 |
| Logistic / softmax regression | Dot product, hyperplane, Hessian is PSD | 1.12, 1.14 |
| Covariance, Mahalanobis, Gaussian | $X^\top X$, symmetric PSD matrices, eigenvectors | 1.11, 1.12 |
| PCA | Eigendecomposition of $\Sigma$, or SVD of the data | 1.11, 1.13 |
| Embeddings | Rows of a matrix; cosine similarity | 1.2, 1.4 |
| Recommenders | Low-rank factorisation, SVD, alternating least squares | 1.10, 1.13 |
| Neural-network layers | Matrix products, Jacobians, low-rank updates | 1.4, 1.14 |
| Attention | $QK^\top$, softmax, batched matmul | 1.15, 1.16 |
Linear regression: the model and the closed form core
What we need from earlier chapters: the dot product, the matrix–vector product and transpose (Chapter 1.4), rank (Chapter 1.7) and the normal equations of least squares (Chapter 1.10).
Guess a house price from its features. A simple recipe: price ≈ (a number) × area + (another number) × bedrooms + a base price. That recipe is a weighted sum, which is a dot product between a weight vector and the house's feature vector.
Learning means choosing the weights so the guesses are close to the real prices of houses we already know. Stack all the houses as rows of one matrix $X$. Then a single matrix–vector product $X\mathbf{w}$ gives every guess at once.
"Close" means a small total of squared mistakes. We square so that positive and negative mistakes cannot cancel, and so that one huge mistake hurts more than several tiny ones.
Three houses. Each feature row is $[1, \text{size}]$ (the leading 1 pays for the "base price"), and the prices are $\mathbf{y}$.
$$X = \begin{bmatrix} 1 & 1 \\ 1 & 2 \\ 1 & 3 \end{bmatrix}, \qquad \mathbf{y} = \begin{bmatrix} 1 \\ 2 \\ 2 \end{bmatrix}.$$- Form $X^\top X = \begin{bmatrix} 3 & 6 \\ 6 & 14 \end{bmatrix}$ (for example, the bottom-right entry is $1+4+9 = 14$).
- Form $X^\top\mathbf{y} = \begin{bmatrix} 1+2+2 \\ 1\cdot1+2\cdot2+3\cdot2 \end{bmatrix} = \begin{bmatrix} 5 \\ 11 \end{bmatrix}$.
- Solve $X^\top X\,\mathbf{w} = X^\top\mathbf{y}$. The determinant is $3\cdot14 - 6\cdot6 = 6$, so $\mathbf{w} = \frac{1}{6}\begin{bmatrix} 14 & -6 \\ -6 & 3 \end{bmatrix}\begin{bmatrix} 5 \\ 11 \end{bmatrix} = \frac16\begin{bmatrix} 4 \\ 3 \end{bmatrix} = \begin{bmatrix} 2/3 \\ 1/2 \end{bmatrix}$.
- Predictions: $X\mathbf{w} = [\tfrac76, \tfrac53, \tfrac{13}{6}]$. Mistakes (residuals): $\mathbf{r} = \mathbf{y} - X\mathbf{w} = [-\tfrac16, \tfrac13, -\tfrac16]$.
- Check: $X^\top\mathbf{r} = [\,-\tfrac16+\tfrac13-\tfrac16,\; -\tfrac16+\tfrac23-\tfrac12\,] = [0, 0]$. The mistakes are perpendicular to every column of $X$. Remember this.
Data: a matrix $X \in \mathbb{R}^{n\times d}$ (one row per example, one column per feature) and targets $\mathbf{y}\in\mathbb{R}^n$. The model predicts
$$\hat{\mathbf{y}} = X\mathbf{w}, \qquad \mathbf{r} = \mathbf{y} - X\mathbf{w}\ \text{(residuals)}, \qquad L(\mathbf{w}) = \tfrac12\|X\mathbf{w}-\mathbf{y}\|^2.$$Setting the gradient to zero (Chapter 1.14) gives the normal equations and the closed form:
$$X^\top X\,\mathbf{w} = X^\top\mathbf{y}\quad\Longrightarrow\quad \mathbf{w} = (X^\top X)^{-1}X^\top\mathbf{y} = X^{+}\mathbf{y}.$$The first form needs the columns of $X$ to be independent (rank $d$). The last form, with the pseudoinverse $X^{+}$ ($=VS^{-1}U^\top$ from the SVD, where $S$ holds the singular values), works even when they are not: it picks the shortest $\mathbf{w}$ among all the best ones.
Bias trick: add a column of ones to $X$, so the intercept becomes one more weight.
Why do we need it?
We often want to predict a number (a price, a temperature, next month's sales) from other numbers. We need a rule that a computer can fit by itself from examples. A weighted sum is the simplest such rule, and it can be fitted exactly in one calculation.
Where is it used?
House-price and sales forecasting, the baseline model in almost every ML project, curve fitting in science, and the 'linear probe' that sits on top of the features learned by a big neural network.
How is it used?
Put the examples in the rows of X (add a column of ones), the answers in y, and call np.linalg.lstsq(X, y). Then look at the residuals y − Xw. If they look like random noise, a straight-line model is good enough.
- Do not literally invert $X^\top X$ in real code. Squaring $X$ squares its condition number (Chapter 1.15). Use
np.linalg.lstsq, QR or the SVD instead. - Dependent features (for example "area in m²" and "area in ft²") make $X^\top X$ singular. The ridge penalty below fixes this.
- The size of $X^\top X$ is $d\times d$. It does not grow with the number of examples $n$.
Quick check: $X$ has 1000 rows and 5 columns. What shape is $X^\top X$, and what shape is $\mathbf{w}$?
$X^\top X$ is $5\times5$ (a $5\times1000$ times a $1000\times5$). The weight vector $\mathbf{w}$ has 5 entries, one per feature. The 1000 examples only affect the numbers inside.
Linear regression by gradient descent core
What we need from earlier chapters: the gradient and Hessian (Chapter 1.14), positive semi-definite matrices (Chapter 1.12), eigenvalues (Chapter 1.11) and the condition number (Chapter 1.15).
The closed form solves everything in one go, but it needs a big matrix solve. With millions of weights that is too slow. So we walk downhill instead: stand somewhere on the loss "bowl", feel which way is steepest uphill (the gradient), and take a small step the opposite way. Repeat.
For a bowl-shaped loss, small enough steps always end at the bottom. The only questions are how big a step (the learning rate) and how round the bowl is. A round bowl is easy. A long thin valley makes the walk zig-zag.
One weight, two data points: $x = [1, 2]$ and $y = [2, 4]$ (so the perfect weight is $w = 2$). With the average loss $L(w) = \frac{1}{2n}\sum (wx_i - y_i)^2$ the gradient is
$$L'(w) = \tfrac12\big[1\cdot(w-2) + 2\cdot(2w-4)\big] = 2.5\,(w-2).$$Take the learning rate $\eta = 0.2$ and start at $w=0$:
- $L'(0) = 2.5(0-2) = -5$, so $w \leftarrow 0 - 0.2\cdot(-5) = 1$.
- $L'(1) = -2.5$, so $w \leftarrow 1 + 0.5 = 1.5$.
- $L'(1.5) = -1.25$, so $w \leftarrow 1.5 + 0.25 = 1.75$.
Each step halves the distance to $2$ (the error is multiplied by $1 - 2.5\eta = 0.5$). If $\eta$ were above $0.8$, the factor $1-2.5\eta$ would be below $-1$: the steps overshoot more each time and the weight explodes.
With the mean loss $L(\mathbf{w}) = \frac{1}{2n}\|X\mathbf{w}-\mathbf{y}\|^2$ (dividing by $n$ only rescales the loss; the minimum does not move):
$$\nabla L(\mathbf{w}) = \tfrac1n X^\top(X\mathbf{w}-\mathbf{y}), \qquad \mathbf{w} \leftarrow \mathbf{w} - \eta\,\nabla L(\mathbf{w}), \qquad H = \tfrac1n X^\top X.$$- $H$ (the Hessian) is symmetric and positive semi-definite (Chapter 1.12), so $L$ is a convex bowl with no false valleys.
- The error shrinks if $\eta < 2/\lambda_{\max}(H)$. Too large and it diverges.
- The speed is set by the condition number $\kappa = \lambda_{\max}/\lambda_{\min}$. A big $\kappa$ means a long thin valley and slow, zig-zag progress.
- Centring or standardising the features makes $H$ rounder (smaller $\kappa$), so GD is faster.
Why do we need it?
The closed form needs one big matrix solve. That is too slow with millions of weights, and for most models it does not exist at all. Walking downhill works for any loss whose slope we can compute.
Where is it used?
Training every neural network (SGD and Adam are gradient descent with extra tricks), logistic regression on large data, matrix factorisation, and almost every large-scale ML system.
How is it used?
Compute the gradient, subtract a small multiple of it (the learning rate), and repeat. Watch the loss. If it explodes, lower the learning rate. If it crawls, standardise the features first. Stop when the loss stops improving.
"Learning rate too high" is not a mystery. For a quadratic loss it is exactly $\eta > 2/\lambda_{\max}$, an eigenvalue fact from Chapter 1.11. For deep networks the picture is messier, but the same eigenvalues of the Hessian still decide what step is safe.
Quick check: in the one-weight example, what is the error multiplier per step when $\eta = 0.4$?
The factor is $1 - 2.5\cdot0.4 = 0$. The first step lands exactly on $w=2$. (It is the "perfect" step for a one-weight bowl.) For two or more weights with different eigenvalues, no single $\eta$ can be perfect in all directions at once.
Ridge and lasso: norms as penalties core
What we need from earlier chapters: the L1 and L2 norms and their unit balls (Chapter 1.2), positive definite matrices (Chapter 1.12), regularised least squares (Chapter 1.10) and the SVD (Chapter 1.13).
With few examples, or with features that copy each other, plain least squares can pick huge weights that cancel one another out. The fit looks perfect but breaks on new data.
The cure is a leash: add a cost for large weights, so the model must earn every large weight. There are two popular leashes, and they come from two of the norms you already know:
- Ridge pays for the squared L2 length of $\mathbf{w}$. The unit ball is a circle, which is smooth.
- Lasso pays for the L1 size of $\mathbf{w}$. The unit ball is a diamond with sharp corners on the axes.
Corners on the axes are the whole story: a solution that touches a corner has some weights exactly zero.
Take the simplest case, where the columns of $X$ are orthonormal ($X^\top X = I$) and the plain least-squares weight is $w = 3$. Use penalty strength $\lambda = 1$.
- Ridge: $w/(1+\lambda) = 3/2 = 1.5$. Everything is shrunk by the same factor.
- Lasso: move toward zero by $\lambda$ and stop at zero: $3 - 1 = 2$.
Now let the plain weight be small, $w = 0.8$. Ridge gives $0.8/2 = 0.4$ (small, not zero). Lasso gives $\max(0.8 - 1, 0) = 0$. Lasso deletes weak features; ridge only shrinks them.
- $X^\top X$ is PSD, so $X^\top X+\lambda I$ has every eigenvalue $\ge\lambda>0$. It is always invertible and well conditioned. Ridge fixes singular $X^\top X$.
- SVD view: ridge multiplies the part of the solution along singular direction $i$ by $\sigma_i^2/(\sigma_i^2+\lambda)$. Directions with small $\sigma_i$ (the noisy, weak ones) are damped hardest.
- Equivalent picture: minimise the loss subject to $\|\mathbf{w}\|\le t$. The answer sits where a loss contour first touches the ball.
- Do not penalise the intercept (the weight on the column of ones).
Why do we need it?
With few examples, or with features that copy each other, plain least squares picks huge weights that cancel out and then fail on new data. A cost on the size of the weights keeps the model calm. Lasso also switches useless features off.
Where is it used?
Weight decay in neural networks (the same idea as ridge), feature selection in medicine and finance (lasso), the Ridge and Lasso classes in scikit-learn, and any regression whose XᵀX is nearly singular.
How is it used?
Standardise the features, add the penalty with strength λ, and choose λ by cross-validation: try a ladder of values and keep the one that does best on held-out data. For lasso, read off which weights became exactly 0.
Scale your features before using a penalty. The leash is on the size of the weights, and a feature measured in tiny units needs a huge weight, so it gets unfairly punished.
Quick check: why is $X^\top X + \lambda I$ invertible even when $X^\top X$ is not?
$X^\top X$ is positive semi-definite, so its eigenvalues are $\ge 0$. Adding $\lambda I$ adds $\lambda$ to every eigenvalue, so they are all $\ge\lambda>0$. A matrix with no zero eigenvalue is invertible.
The geometry: regression is projection core
What we need from earlier chapters: the projection of a vector (Chapter 1.2), the column space and its orthogonal complement (Chapter 1.8) and orthogonal projection matrices (Chapter 1.9).
Every prediction $X\mathbf{w}$ is a linear combination of the columns of $X$. So the set of predictions the model can ever make is the column space of $X$: a flat sheet (a plane, in this picture) through the origin.
The real answer $\mathbf{y}$ usually floats off the sheet, so we cannot hit it exactly. The best we can do is the closest point on the sheet. That point is the shadow of $\mathbf{y}$ on the sheet, and the leftover arrow from the shadow to $\mathbf{y}$ is perpendicular to the sheet.
That is all least-squares regression is: drop a perpendicular.
Back to our three houses. The columns of $X$ are $[1,1,1]$ and $[1,2,3]$, and $\mathbf{y} = [1,2,2]$. We found the prediction $\hat{\mathbf{y}} = [\tfrac76,\tfrac53,\tfrac{13}6]$ and the residual $\mathbf{r} = [-\tfrac16,\tfrac13,-\tfrac16]$. Test perpendicularity with a dot product:
$$\hat{\mathbf{y}}\cdot\mathbf{r} = \tfrac76\cdot(-\tfrac16) + \tfrac53\cdot\tfrac13 + \tfrac{13}6\cdot(-\tfrac16) = \tfrac{-7+20-13}{36} = 0\ \checkmark$$So $\hat{\mathbf{y}}$ lies in the sheet and $\mathbf{r}$ sticks straight out of it.
The prediction is an orthogonal projection of $\mathbf{y}$ onto the column space $C(X)$:
$$\hat{\mathbf{y}} = P\mathbf{y}, \qquad P = X(X^\top X)^{-1}X^\top, \qquad \mathbf{r} = (I-P)\mathbf{y}.$$- $P$ is called the hat matrix (it puts the hat on $y$). It is symmetric and $P^2 = P$ (projecting twice changes nothing).
- The residual is perpendicular to every column: $X^\top\mathbf{r}=\mathbf{0}$. That sentence is the normal equations. So $\mathbf{r}$ lives in the left null space $N(X^\top)$.
- Pythagoras holds: $\|\mathbf{y}\|^2 = \|\hat{\mathbf{y}}\|^2 + \|\mathbf{r}\|^2$.
Why do we need it?
It tells us what regression really does, so we can trust it and debug it. The answer is the closest point the model can reach, and the leftover is perpendicular to everything the model can use.
Where is it used?
Explaining least squares and R², regression diagnostics such as leverage and outlier checks, and later PCA, which is the same projection with a cleverly chosen sheet.
How is it used?
To check a fit, compute the residual r = y − ŷ and verify that Xᵀr is almost zero. Large diagonal entries of the hat matrix P flag single points that pull the whole fit toward themselves.
Quick check: what is $P\hat{\mathbf{y}}$? Why?
It is $\hat{\mathbf{y}}$ itself. $\hat{\mathbf{y}}$ is already on the sheet, and projecting a point that is already on the sheet changes nothing. This is $P^2=P$.
Logistic regression: a score and a squashing function core
What we need from earlier chapters: the dot product and projection (Chapter 1.2, signed distance is a projection), hyperplanes and affine spaces (Chapter 1.3) and the unit vector.
Now the answer is a yes/no: is this email spam? We still start with a weighted sum, called the score. Big positive score means "looks like spam"; big negative means "looks fine".
But a score can be any number, and we want a probability between 0 and 1. So we pass the score through an S-shaped curve, the sigmoid. Score $0$ becomes probability $\tfrac12$ (unsure). Large scores flatten out near $1$; very negative scores flatten out near $0$.
The place where the score is exactly $0$ is the decision boundary. In 2D it is a line, in 3D a plane, in general a hyperplane. The weight vector $\mathbf{w}$ points straight across it, toward the "yes" side.
Two features per email: the number of times the word "free" appears and the number of links. Weights $\mathbf{w} = [1.5, 0.5]$, bias $b = -2$.
- Email A has $\mathbf{x} = [2, 1]$. Score: $z = 1.5\cdot2 + 0.5\cdot1 - 2 = 1.5$.
- Probability: $\sigma(1.5) = 1/(1+e^{-1.5}) = 1/(1+0.223) \approx 0.82$. Likely spam.
- Email B has $\mathbf{x} = [0, 1]$. Score: $0 + 0.5 - 2 = -1.5$, so $\sigma(-1.5) \approx 0.18$. Likely fine.
- Notice the symmetry: $\sigma(-z) = 1-\sigma(z)$, and $0.82 + 0.18 = 1$.
- The decision boundary is $\{\mathbf{x} : \mathbf{w}^\top\mathbf{x} + b = 0\}$, a hyperplane.
- $\mathbf{w}$ is the normal vector: it is perpendicular to the boundary, because moving along the boundary does not change $z$, so $\mathbf{w}\cdot(\text{step}) = 0$ for every step along it.
- The boundary is $-b/\|\mathbf{w}\|$ away from the origin, measured along $\mathbf{w}$.
- The signed distance of a point $\mathbf{x}$ from the boundary is $z/\|\mathbf{w}\|$ (positive on the "yes" side).
The loss is the cross-entropy: $-\log p$ if the true label is $1$, and $-\log(1-p)$ if it is $0$. It is small when the model is confidently right and huge when confidently wrong.
Why do we need it?
Many questions have yes/no answers: spam or not, sick or healthy. We need a model that turns evidence into a probability between 0 and 1, and a boundary we can draw to separate the classes.
Where is it used?
Spam filters, medical risk scores, click prediction in online ads, credit scoring, and the final layer of many classifiers.
How is it used?
Compute the score w·x + b, squash it with the sigmoid, and predict class 1 if p ≥ 0.5. Use scikit-learn's LogisticRegression, or train w yourself with gradient descent. Check accuracy and loss on held-out data.
- Logistic regression is "regression" in name only; it outputs a probability and is used for classification.
- The boundary is flat. If the classes are tangled (for example one ring inside another), a flat boundary cannot separate them. You then need feature maps, kernels or a neural network.
- Scaling $(\mathbf{w}, b)$ by 2 gives the same boundary but a sharper sigmoid (more confident). The direction of $\mathbf{w}$ decides the boundary; the length decides the confidence.
Quick check: $\mathbf{w}=[3,4]$, $b=-10$. How far is the boundary from the origin?
$\|\mathbf{w}\| = 5$, so the distance is $-b/\|\mathbf{w}\| = 10/5 = 2$, measured along $\mathbf{w}$. (Check: the point $2\cdot[0.6,0.8] = [1.2,1.6]$ has score $3.6+6.4-10=0$ ✓.)
Fitting logistic regression: gradient, Hessian and convexity core
What we need from earlier chapters: the gradient, Jacobian and Hessian and the chain rule (Chapter 1.14), positive semi-definite matrices (Chapter 1.12) and solving $A\mathbf{x}=\mathbf{b}$ (Chapter 1.6).
There is no closed-form answer this time, so we need to walk downhill as before. To walk, we need the slope (the gradient). To walk smartly, we can also use the curvature (the Hessian), which tells us how the slope itself changes.
The nice surprise: the gradient looks almost exactly like the one for linear regression. It is still "$X^\top$ times the mistakes". And the curvature is never negative in any direction, so the loss is a single bowl: no false valleys, one best answer.
One training example $\mathbf{x} = [2,1]$ with label $y=1$. Use the bias trick: $\mathbf{a} = [2, 1, 1]$ (the last 1 is for $b$). Start at $\mathbf{w}=\mathbf{0}$.
- Score $z = 0$, probability $p = \sigma(0) = 0.5$, mistake $p - y = -0.5$.
- Gradient $= (p-y)\,\mathbf{a} = -0.5\cdot[2,1,1] = [-1, -0.5, -0.5]$.
- Step with $\eta = 1$: $\mathbf{w} \leftarrow [1, 0.5, 0.5]$. The new score is $2+0.5+0.5 = 3$, and $p = \sigma(3) \approx 0.95$. Much better.
- Curvature at the start: $p(1-p) = 0.25$, so $H = 0.25\,\mathbf{a}\mathbf{a}^\top = 0.25\begin{bmatrix}4&2&2\\2&1&1\\2&1&1\end{bmatrix}$. Its eigenvalues are $0.25\|\mathbf{a}\|^2 = 1.5$, $0$ and $0$. All are $\ge 0$ ✓.
With augmented rows $\mathbf{a}_i = [\mathbf{x}_i, 1]$, probabilities $\mathbf{p} = \sigma(A\mathbf{w})$ and the average cross-entropy loss $L$:
$$\nabla L = \tfrac1n A^\top(\mathbf{p}-\mathbf{y}), \qquad H = \tfrac1n A^\top S A,\quad S = \operatorname{diag}\big(p_i(1-p_i)\big).$$- The first formula comes from the chain rule: $\partial L/\partial z_i = p_i - y_i$, and $z = A\mathbf{w}$ has Jacobian $A$. Compare with linear regression: $\nabla L = \frac1n X^\top(\hat{\mathbf{y}}-\mathbf{y})$. Same shape!
- $H$ is positive semi-definite. For any direction $\mathbf{v}$: $\mathbf{v}^\top H\mathbf{v} = \frac1n\sum_i p_i(1-p_i)\,(\mathbf{a}_i\cdot\mathbf{v})^2 \ge 0$, because every term is non-negative.
- PSD Hessian everywhere $\Rightarrow$ $L$ is convex $\Rightarrow$ every local minimum is the global minimum.
- Newton's method uses the curvature: $\mathbf{w}\leftarrow\mathbf{w} - H^{-1}\nabla L$. This is a linear solve (Chapter 1.6). It is also called iteratively reweighted least squares, because each step is a weighted linear regression.
Why do we need it?
To fit the weights we need the slope of the loss. The Hessian also tells us the loss is a single bowl, so training cannot get stuck in a false valley, and it lets us take smarter steps than plain gradient descent.
Where is it used?
Training logistic regression (Newton or IRLS in statsmodels, L-BFGS in scikit-learn) and the habit 'Xᵀ times (prediction − target)' that gives the weight gradient of every neural-network layer.
How is it used?
Compute p = σ(Aw), then the gradient Aᵀ(p − y)/n. For a Newton step build H = AᵀSA/n and solve H·step = gradient. Check that the eigenvalues of H are all ≥ 0 if you want proof that the loss is convex.
- If the two classes can be perfectly separated, the best $\mathbf{w}$ runs off to infinity (more confidence always lowers the loss). Regularisation (ridge, above) fixes this.
- "PSD" gives convexity, but not a unique minimiser. Duplicate features make $H$ singular, and then many $\mathbf{w}$ are equally good.
- Newton needs a $d\times d$ solve per step. That is great for 100 weights and hopeless for a billion, which is why deep learning uses gradient methods.
Quick check: why is a sum of matrices of the form $s_i\,\mathbf{a}_i\mathbf{a}_i^\top$ with $s_i\ge0$ always PSD?
For any $\mathbf{v}$: $\mathbf{v}^\top(s_i\mathbf{a}_i\mathbf{a}_i^\top)\mathbf{v} = s_i(\mathbf{a}_i\cdot\mathbf{v})^2\ge0$. A sum of non-negative numbers is non-negative, so $\mathbf{v}^\top H\mathbf{v}\ge0$ for every $\mathbf{v}$. That is the definition of PSD.
Softmax regression: more than two classes core
What we need from earlier chapters: the matrix–vector product (Chapter 1.4, one row of $W$ per class), the dot product, and numerical stability (Chapter 1.15, subtract the max before exponentiating).
With $K$ classes, give each class its own weight vector and its own score. The class whose score is biggest wins. To turn the $K$ scores into probabilities that add up to 1, make every score positive with $e^{z}$ and divide by the total. That recipe is the softmax.
Geometrically, each class owns a wedge-shaped region of the plane: the region where its own score $\mathbf{w}_k\cdot\mathbf{x}$ is the biggest of all.
Three classes and scores $\mathbf{z} = [2, 1, 0]$.
- Exponentiate: $e^2 = 7.39$, $e^1 = 2.72$, $e^0 = 1$. Total: $11.11$.
- Divide: $\mathbf{p} \approx [0.665, 0.245, 0.090]$. They add to $1$ ✓.
- If the true class is the first one, the loss is $-\log 0.665 = 0.408$, and the gradient with respect to the scores is $\mathbf{p} - \mathbf{y} = [0.665-1, 0.245, 0.090] = [-0.335, 0.245, 0.090]$.
"Push up the true class, push down the others, in proportion to how much probability they stole."
For $K$ classes keep a weight matrix $W\in\mathbb{R}^{K\times d}$ (row $k$ is $\mathbf{w}_k$) and a bias vector $\mathbf{b}$. For one input $\mathbf{x}$:
$$\mathbf{z} = W\mathbf{x} + \mathbf{b}, \qquad p_k = \frac{e^{z_k}}{\sum_j e^{z_j}}, \qquad L = -\log p_y.$$- Gradient: $\partial L/\partial\mathbf{z} = \mathbf{p} - \mathbf{e}_y$ (where $\mathbf{e}_y$ is the one-hot vector of the true class), and $\partial L/\partial W = (\mathbf{p}-\mathbf{e}_y)\,\mathbf{x}^\top$, an outer product (Chapter 1.4).
- With a whole batch $X\in\mathbb{R}^{n\times d}$: $Z = XW^\top + \mathbf{1}\mathbf{b}^\top$, and $\nabla_W L = \frac1n (P-Y)^\top X$.
- The loss is still convex in $W$ (its Hessian is PSD for the same reason as before).
- Adding the same number to every score changes nothing, so real code subtracts $\max_j z_j$ first. With $K=2$ softmax regression is equivalent to logistic regression (the single weight vector is $\mathbf{w}_1-\mathbf{w}_2$).
Why do we need it?
With more than two classes we need one score per class, and probabilities that are all positive and add up to 1. Softmax does exactly that, and its gradient is the simple difference p − y.
Where is it used?
The last layer of image classifiers such as ResNet, language models choosing the next word out of tens of thousands, and multi-class logistic regression.
How is it used?
Compute the scores z = Wx + b, subtract the largest score for stability, exponentiate, and divide by the sum. Train with cross-entropy. The gradient on the scores is the probabilities minus the one-hot label.
- Never compute $e^{z}$ naively for big scores: $e^{1000}$ overflows to infinity and you get NaN. Subtract the largest score first; the answer is identical.
- The softmax is not "just a normalised score". Because of the exponential, it exaggerates: a score gap of 3 already gives a 20-to-1 ratio. We will meet this again in attention.
Quick check: scores $[5,5,5]$. What are the probabilities?
All equal: each is $1/3$. Softmax only cares about differences between scores. Adding the same number to all of them (here 5) changes nothing.
Covariance matrices core
What we need from earlier chapters: the matrix product and transpose (Chapter 1.4, $X^\top X$ is a table of dot products between columns), PSD matrices and quadratic forms (Chapter 1.12) and eigenvalues of symmetric matrices (Chapter 1.11).
Take a cloud of data points, say people's heights and weights. Two questions matter: how spread out is each feature? and do the features rise and fall together? Taller people tend to be heavier, so when height is above average, weight tends to be above average too.
First centre the cloud: subtract the average of each feature, so the middle of the cloud sits at the origin. Then every number measures "how far from typical". The covariance matrix is a table that holds all these answers at once: the diagonal says how spread out each feature is, and the off-diagonal entries say how much each pair moves together.
Four people, two features: $X = \begin{bmatrix}2&1\\4&3\\6&2\\8&6\end{bmatrix}$.
- Column means: $\mu_1 = (2+4+6+8)/4 = 5$ and $\mu_2 = (1+3+2+6)/4 = 3$.
- Centre (subtract the means): $X_c = \begin{bmatrix}-3&-2\\-1&0\\1&-1\\3&3\end{bmatrix}$. Each column now adds up to $0$.
- $X_c^\top X_c$: top-left $9+1+1+9 = 20$; off-diagonal $6+0-1+9 = 14$; bottom-right $4+0+1+9 = 14$.
- Divide by $n = 4$: $\Sigma = \begin{bmatrix}5 & 3.5\\3.5&3.5\end{bmatrix}$.
- Correlation: $\rho = 3.5/\sqrt{5\cdot3.5} = 3.5/4.18 \approx 0.84$. Strongly together.
For data $X\in\mathbb{R}^{n\times d}$ with column means $\boldsymbol\mu$, the centred data is $X_c = X - \mathbf{1}\boldsymbol\mu^\top$ and the covariance matrix is
$$\Sigma = \frac1n X_c^\top X_c, \qquad \Sigma_{jk} = \frac1n\sum_{i=1}^n (x_{ij}-\mu_j)(x_{ik}-\mu_k).$$(NumPy's np.cov divides by $n-1$ instead; for large $n$ it makes no difference.)
- $\Sigma_{jj}$ is the variance of feature $j$. $\Sigma_{jk}$ is the covariance of features $j,k$.
- Correlation rescales to unit-free numbers: $\rho_{jk} = \Sigma_{jk}/(\sigma_j\sigma_k)\in[-1,1]$. The correlation matrix is the covariance of the standardised features.
- Symmetric: $\Sigma_{jk}=\Sigma_{kj}$, so $\Sigma^\top=\Sigma$ and the spectral theorem applies (real eigenvalues, orthogonal eigenvectors).
- Positive semi-definite: for any direction $\mathbf{v}$, $\ \mathbf{v}^\top\Sigma\mathbf{v} = \frac1n\|X_c\mathbf{v}\|^2 \ge 0$. This number is the variance of the data projected onto $\mathbf{v}$. A variance can never be negative.
Why do we need it?
We need one object that records how spread out each feature is and which features rise and fall together. Without it we cannot describe the shape of a data cloud with numbers.
Where is it used?
PCA, Gaussian models, linear discriminant analysis, risk of a portfolio in finance, Kalman filters in tracking and robotics, and checks for features that are almost copies of each other.
How is it used?
Subtract the column means, then compute X_cᵀX_c/n (np.cov uses n−1). Read the diagonal (the variances), the correlation matrix (numbers from −1 to 1) and the eigenvalues (all ≥ 0).
- Forgetting to centre gives $\frac1nX^\top X$, which is a different matrix (it mixes the mean into the spread). Always subtract the mean first.
- Correlation is not causation, and it is only linear. A perfect circle of points has covariance $0$ but the features are completely dependent.
- With $n$ examples and $d$ features, $\Sigma$ has rank at most $n-1$. If $d \ge n$, it is singular.
Quick check: why can't $\Sigma$ have a negative eigenvalue?
If $\Sigma\mathbf{v}=\lambda\mathbf{v}$ with $\|\mathbf{v}\|=1$, then $\lambda = \mathbf{v}^\top\Sigma\mathbf{v} = \frac1n\|X_c\mathbf{v}\|^2\ge0$. The eigenvalue is exactly the variance of the data along that eigenvector, and variances are never negative.
Mahalanobis distance and the multivariate Gaussian core
What we need from earlier chapters: Euclidean distance (Chapter 1.2), the inverse and determinant (Chapter 1.7), quadratic forms $\mathbf{d}^\top A\mathbf{d}$ (Chapter 1.12) and the eigen-decomposition (Chapter 1.11).
Stand in a long, thin field. Walking 2 metres across the field takes you to the fence. Walking 2 metres along it is nothing. Ordinary (Euclidean) distance treats both walks as equal, which is silly for a cloud that is long in one direction and thin in the other.
Mahalanobis distance measures distance in units of "how much the data usually varies in that direction". Stretch space until the cloud is round, then use an ordinary ruler. A point that is only 2 metres away but off the thin side is a genuine outlier; a point 2 metres away along the long side is perfectly ordinary.
Cloud with $\boldsymbol\mu=\mathbf{0}$ and $\Sigma = \begin{bmatrix}4&0\\0&1\end{bmatrix}$ (spread 2 along $x$, spread 1 along $y$). Two points, both Euclidean distance $2$ from the centre: $\mathbf{a}=[2,0]$ and $\mathbf{b}=[0,2]$. Here $\Sigma^{-1}=\begin{bmatrix}1/4&0\\0&1\end{bmatrix}$.
- $d_M(\mathbf{a})^2 = \mathbf{a}^\top\Sigma^{-1}\mathbf{a} = \tfrac14\cdot 2^2 = 1$, so $d_M(\mathbf{a}) = 1$ (one standard deviation: ordinary).
- $d_M(\mathbf{b})^2 = \mathbf{b}^\top\Sigma^{-1}\mathbf{b} = 1\cdot 2^2 = 4$, so $d_M(\mathbf{b}) = 2$ (two standard deviations: unusual).
- With $\Sigma = I$ this is plain Euclidean distance. In general it equals the Euclidean length of the whitened vector $\Sigma^{-1/2}(\mathbf{x}-\boldsymbol\mu)$. Here $\Sigma^{-1/2}$ is the matrix that "un-stretches" the cloud into a round one (you will build it in the whitening section below).
- Every set $\{d_M = k\}$ is an ellipse whose axes point along the eigenvectors of $\Sigma$ and have half-lengths $k\sqrt{\lambda_i}$.
The multivariate Gaussian (bell curve in many dimensions) is built from it:
$$p(\mathbf{x}) = \frac{1}{(2\pi)^{d/2}\sqrt{\det\Sigma}}\exp\!\Big(-\tfrac12\,d_M(\mathbf{x})^2\Big).$$- Its contours of equal density are exactly the Mahalanobis ellipses. $\det\Sigma=\prod\lambda_i$ measures the volume of the cloud.
- $\Sigma$ must be positive definite (invertible). Fitting the best Gaussian to data just means taking $\boldsymbol\mu$ = the sample mean and $\Sigma$ = the sample covariance.
Why do we need it?
Plain distance ignores the shape of the data, so it calls a point 'ordinary' even when it is far off the thin side of the cloud. We need a distance measured in units of normal variation.
Where is it used?
Outlier and anomaly detection, Gaussian mixture models, classification with LDA and QDA, quality control, and checking whether new data looks like the training data.
How is it used?
Fit the mean and covariance on normal data. For a new point solve Σy = (x − μ) and take √((x − μ)·y). Flag points beyond about 3. For the Gaussian density, also divide by the square root of det Σ.
- $\Sigma$ must be invertible. With strongly correlated or duplicated features it is (nearly) singular and $\Sigma^{-1}$ explodes. Add a small ridge $\Sigma+\epsilon I$ (as in ridge regression), or work in the PCA basis.
- Do not compute $\Sigma^{-1}$ explicitly; solve $\Sigma\mathbf{y}=\mathbf{d}$ with Cholesky (Chapter 1.13) and then take $\sqrt{\mathbf{d}\cdot\mathbf{y}}$.
Quick check: for $\Sigma=\operatorname{diag}(9,1)$, what is $d_M$ of $[3,0]$ and of $[0,3]$?
$d_M([3,0]) = \sqrt{3^2/9} = 1$. $d_M([0,3]) = \sqrt{3^2/1} = 3$. Same Euclidean distance (3), but the second point is three standard deviations out.
PCA: finding the direction of most variance core
What we need from earlier chapters: change of basis (Chapter 1.3), orthonormal bases (Chapter 1.9), the spectral theorem (Chapter 1.11), the SVD (Chapter 1.13) and the covariance matrix just above.
Imagine photographing a cigar-shaped cloud of points. Photograph it end-on and it looks like a small circle; you have lost almost everything. Photograph it from the side and the cigar looks as long as possible; you have kept the most information.
Principal Component Analysis (PCA) finds the best viewing direction: the one along which the data is most spread out. That is the first principal component. The second is the most spread-out direction among those perpendicular to the first, and so on.
Put together, those directions are a new set of axes lined up with the cloud: a rotation (a change of basis), nothing more.
Four centred points: $(2,2),\ (-2,-2),\ (1,-1),\ (-1,1)$. They stretch along the diagonal.
- Covariance: $\Sigma = \frac14X_c^\top X_c = \begin{bmatrix}2.5&1.5\\1.5&2.5\end{bmatrix}$ (for example $\Sigma_{12} = (4+4-1-1)/4 = 1.5$).
- Eigenvalues: $\det(\Sigma-\lambda I) = (2.5-\lambda)^2-1.5^2 = 0$, so $\lambda = 2.5\pm1.5 = 4$ and $1$.
- Eigenvectors: $\mathbf{v}_1 = [1,1]/\sqrt2$ for $\lambda=4$ and $\mathbf{v}_2 = [1,-1]/\sqrt2$ for $\lambda=1$. They are perpendicular, as the spectral theorem promises.
- Check: variance along $\mathbf{v}_1$ is $\mathbf{v}_1^\top\Sigma\mathbf{v}_1 = (2.5+2.5+1.5+1.5)/2 = 4$ ✓. The scores of the four points on $\mathbf{v}_1$ are $2.83, -2.83, 0, 0$; their mean square is $(8+8)/4 = 4$ ✓.
PCA on centred data $X_c\in\mathbb{R}^{n\times d}$. First principal direction:
$$\mathbf{v}_1 = \arg\max_{\|\mathbf{u}\|=1}\ \mathbf{u}^\top\Sigma\mathbf{u}\quad(\text{the variance along }\mathbf{u}).$$Its answer is the top eigenvector of $\Sigma$, and the maximum value is $\lambda_1$ (this is the Rayleigh quotient of Chapter 1.11). The next directions are the other eigenvectors, in order of eigenvalue.
Two equivalent recipes:
- Eigen route: $\Sigma = V\Lambda V^\top$. The columns of $V$ are the principal directions, $\lambda_i$ is the variance along $\mathbf{v}_i$.
- SVD route (preferred): $X_c = U S V^\top$. The same $V$ gives the directions, and $\lambda_i = s_i^2/n$. The scores are $X_cV = US$.
The SVD route never forms $X_c^\top X_c$ (which would square the condition number, Chapter 1.15) and also works when $d>n$. The new coordinates ("scores") of the data are $Z = X_cV$: since $V$ is orthogonal, this is just a rotation, so nothing is lost until we drop columns.
Why do we need it?
Data often has far more numbers than ideas. PCA finds the few directions that carry most of the variation, so that we can plot, compress or de-noise the data.
Where is it used?
Plotting high-dimensional data in 2D, image compression ('eigenfaces'), speeding up other models, removing noise, and exploring genetics, finance and sensor data.
How is it used?
Centre the data, run np.linalg.svd on it, take the rows of Vᵀ as the directions and s²/n as their variances. Look at the explained-variance ratios to decide how many directions are worth keeping.
- Centre first. PCA on uncentred data finds the direction to the mean, not the direction of spread.
- Scale matters. A feature measured in millimetres has a huge variance and will hog the first component. Standardise the features (or use the correlation matrix) when units differ.
- Signs are arbitrary: $\mathbf{v}_1$ and $-\mathbf{v}_1$ are equally good. Do not be alarmed if NumPy and your hand calculation differ by a sign.
- PCA finds directions of variance, not directions that are best for a given task. The most spread-out direction is not always the most useful one.
Quick check: the singular values of $X_c$ are $4$ and $2$ with $n=4$ rows. What are the PCA variances?
$\lambda_i = s_i^2/n$: $\lambda_1 = 16/4 = 4$ and $\lambda_2 = 4/4 = 1$. These match the eigenvalues of $\Sigma$ in the worked example, which is the whole point: the two routes agree.
PCA: explained variance, reduction and reconstruction core
What we need from earlier chapters: orthogonal projection matrices (Chapter 1.9), low-rank approximation and the SVD (Chapter 1.13) and the principal directions from the section above.
Once the axes line up with the cloud, the last axes (the ones with tiny variance) carry almost no information. Throw them away. Keep only the first $k$ coordinates of each point. That is dimensionality reduction.
To get back to the original space, put the kept coordinates back on their axes and set the dropped ones to zero. The result is the point's shadow on the flat sheet spanned by the first $k$ directions. How much did we lose? Exactly the variance of the directions we dropped.
The four points from the last section, kept to $k=1$ dimension. The variances were $\lambda_1 = 4$ and $\lambda_2 = 1$.
- Explained variance ratio: $\lambda_1/(\lambda_1+\lambda_2) = 4/5 = 80\%$.
- Score of $(2,2)$: $z = (2,2)\cdot[1,1]/\sqrt2 = 2.83$. Reconstruct: $z\,\mathbf{v}_1 = 2.83\cdot[0.707, 0.707] = (2, 2)$. Perfect, because this point lies on the first axis.
- Score of $(1,-1)$: $z = 0$. Reconstruct: $(0,0)$. The error is $\|(1,-1)\|^2 = 2$.
- Total squared error: $0+0+2+2 = 4$, average $4/4 = 1 = \lambda_2$ ✓. The average error is exactly the variance we dropped.
Keep the first $k$ columns $V_k$ of $V$ (a $d\times k$ matrix with orthonormal columns).
$$\mathbf{z} = V_k^\top(\mathbf{x}-\boldsymbol\mu)\ \ (\text{reduce: } d\to k \text{ numbers}), \qquad \hat{\mathbf{x}} = \boldsymbol\mu + V_k\mathbf{z} = \boldsymbol\mu + V_kV_k^\top(\mathbf{x}-\boldsymbol\mu)\ \ (\text{reconstruct}).$$- $V_kV_k^\top$ is the orthogonal projection matrix onto the span of the first $k$ directions.
- Explained variance ratio: $\text{EVR}_k = \dfrac{\lambda_1+\dots+\lambda_k}{\lambda_1+\dots+\lambda_d}$. A common rule: pick the smallest $k$ with $\text{EVR}_k\ge 0.90$ or $0.95$.
- Average squared reconstruction error $=\lambda_{k+1}+\dots+\lambda_d$ (the dropped variances).
- In SVD terms, $\hat{X}_c = U_kS_kV_k^\top$ is the best rank-$k$ approximation of $X_c$ (Eckart–Young, Chapter 1.13).
Why do we need it?
Knowing the directions is not enough. We must decide how many to keep, and know how much information we throw away when we drop the rest.
Where is it used?
Compressing and de-noising data, shrinking 1000 features to 50 before feeding another model, image compression, and the 'scree plot' that data analysts draw.
How is it used?
Plot the cumulative explained variance and keep the smallest k that reaches about 90–95%. Reduce with Z = X_c V_k, reconstruct with Z V_kᵀ + mean, and check the reconstruction error. Fit on training data only.
- Fit the centring and the $V$ on the training data only, then reuse the same $\boldsymbol\mu$ and $V$ for test data. Refitting PCA on the test set leaks information.
- A low-variance direction is not always unimportant. A tiny but telling signal can hide there. Check performance, not only EVR.
Quick check: eigenvalues of $\Sigma$ are $5, 3, 1, 1$. How many components give at least 80% explained variance?
The total is $10$. One component gives $5/10 = 50\%$. Two give $8/10 = 80\%$. So $k=2$ is enough.
Whitening: making the cloud round
What we need from earlier chapters: change of basis (Chapter 1.3), scaling transformations (Chapter 1.5) and the PCA rotation.
PCA rotates the cloud so that it lines up with the axes. But it is still a cigar: long in one direction, thin in the other. Whitening takes one more step: divide each new coordinate by its own spread. The cigar becomes a round ball. No direction is bigger than any other, like white noise.
Two steps: rotate (PCA), then rescale each axis. Both are matrix multiplications.
Take the four points again. After the PCA rotation, the first coordinate has variance $4$ (spread $2$) and the second has variance $1$ (spread $1$). Dividing by $2$ and $1$ gives each coordinate variance $1$, and since PCA made them uncorrelated, the covariance is now the identity $I$.
For the point $(2,2)$: PCA scores $[2.83, 0]$, whitened $[2.83/2, 0/1] = [1.41, 0]$.
- $V^\top$ rotates, $\Lambda^{-1/2}=\operatorname{diag}(1/\sqrt{\lambda_i})$ rescales. This is PCA whitening.
- Rotating back with $V$ afterwards gives ZCA whitening, which stays as close as possible to the original data.
- The Euclidean length of a whitened vector is the Mahalanobis distance of the original.
- Directions with tiny $\lambda_i$ get hugely amplified (mostly noise). In practice use $1/\sqrt{\lambda_i+\epsilon}$.
Why do we need it?
Many algorithms work best when features are uncorrelated and on an equal footing. Whitening turns the cigar-shaped cloud into a round ball, so no direction dominates.
Where is it used?
Pre-processing for independent component analysis (ICA) and some SVMs, 'preconditioning' in optimisation (Adam rescales each weight's step, a cheap cousin of whitening), natural-gradient methods, and the computation of Mahalanobis distance.
How is it used?
Compute the PCA rotation V and the eigenvalues λ. Multiply the centred data by V, then divide each coordinate by √(λ + ε). Always add a small ε, or tiny directions will blow up the noise.
Quick check: why does whitening need $\lambda_i>0$?
We divide by $\sqrt{\lambda_i}$. A zero eigenvalue means the data has no spread at all in that direction (it lies flat), and dividing by zero is impossible. A very tiny one amplifies noise enormously. That is why we add a small $\epsilon$.
Embeddings: words and items as vectors core
What we need from earlier chapters: vectors (Chapter 1.2), and the matrix–vector product (Chapter 1.4: a product with a one-hot vector picks out one row).
A computer cannot do arithmetic on the word "cat". So we give every word its own list of numbers, a vector, chosen so that words with similar meaning get similar vectors. The list is called an embedding because the word is "embedded" as a point in a space of numbers.
All the vectors live in one big table with one row per word: the embedding matrix. To look up a word, you read its row.
Reading one row can also be described as a matrix product. Multiply the matrix by a one-hot vector (all zeros except a single 1 at the word's position) and the single 1 selects exactly that row. Real systems skip the multiplication and just index the row, because the answer is the same and far cheaper.
A tiny vocabulary of 5 tokens, each embedded in 3 numbers. The embedding matrix $E$ has 5 rows and 3 columns; the row for "dog" is $[0.8, 0.5, -0.2]$ (row 2).
- The one-hot vector for "dog" is $\mathbf{e} = [0, 1, 0, 0, 0]$ (a 1 in position 2).
- Compute $\mathbf{e}^\top E = 0\cdot\text{row}_1 + 1\cdot\text{row}_2 + 0\cdot\text{row}_3 + 0\cdot\text{row}_4 + 0\cdot\text{row}_5$.
- Every term but one is zero, so the result is row 2: $[0.8, 0.5, -0.2]$. That is the embedding of "dog".
It is a linear combination of the rows with weights $[0,1,0,0,0]$.
A vocabulary of $V$ tokens, each embedded in $d$ dimensions, is stored as an embedding matrix $E\in\mathbb{R}^{V\times d}$. Token number $i$ has the vector
$$\mathbf{x}_i = E_{i,:} = \mathbf{e}_i^\top E, \qquad \mathbf{e}_i\in\mathbb{R}^V \text{ one-hot}.$$- For a sentence of $n$ tokens, stack the one-hot rows into $H\in\mathbb{R}^{n\times V}$. Then $X = HE\in\mathbb{R}^{n\times d}$ is the sentence as a matrix, one row per token. Real code does
E[token_ids]. - The numbers in $E$ are learned during training: they start random and are nudged by gradient descent. Only the rows of the tokens that appeared get a gradient.
- Anything can be embedded: words, sub-words, users, products, songs, images, graph nodes.
Why do we need it?
A model cannot do maths on a word, a song or a user. An embedding gives each one a learned list of numbers, so that similar things end up close together.
Where is it used?
The first layer of every language model, word2vec, recommender systems (users and items), search and retrieval, and image-and-text models such as CLIP.
How is it used?
Keep a matrix E with one row per token. Turn text into token numbers and read E[ids]. Training adjusts E by gradient descent. Afterwards the rows are the vectors you compare.
- The one-hot vector has length $V$ (often 50,000 or more). Building it and doing the full product would waste time and memory, so libraries index instead. The matrix view is for understanding and for deriving gradients.
- A single embedding vector means nothing alone. Only relationships (distances, angles) between vectors carry meaning.
Quick check: $E$ is $50{,}000\times768$. How many numbers does one token's embedding have, and how many does $E$ store?
One embedding has $d=768$ numbers (a row). $E$ stores $50{,}000\times768 = 38{,}400{,}000$ numbers.
Nearest neighbours and vector arithmetic core
What we need from earlier chapters: cosine similarity and normalising a vector (Chapter 1.2), adding and subtracting vectors (Chapter 1.2) and matrix–vector products (Chapter 1.4).
If similar words have similar vectors, then "which word is most like this one?" becomes a geometry question: which arrow points in nearly the same direction? We compare directions with the cosine, ignoring length.
Even better, directions can capture relationships. The step from "man" to "king" is "add royalty". The step from "woman" to "queen" is the same step. So start at "king", undo "man", add "woman", and you should land near "queen":
king − man + woman ≈ queen
Similarity. From the last section: cat $=[0.9, 0.4, -0.3]$, dog $=[0.8, 0.5, -0.2]$, car $=[-0.6, 0.1, 0.9]$.
- cat·dog $= 0.72+0.20+0.06 = 0.98$. Lengths: $\sqrt{1.06}=1.030$ and $\sqrt{0.93}=0.964$. Cosine $= 0.98/(1.030\cdot0.964)\approx 0.99$. Almost the same direction.
- cat·car $= -0.54+0.04-0.27 = -0.77$. Length of car: $\sqrt{1.18}=1.086$. Cosine $\approx -0.77/(1.030\cdot1.086) \approx -0.69$. Pointing away from each other.
Analogy with a toy 2D map that we made up by hand so the pattern is clean (horizontal = male to female, vertical = how royal): man $=(-3,1)$, woman $=(3,1)$, king $=(-3,4)$, queen $=(3,4)$.
king − man $= (0, 3)$ ("add royalty"). Add woman: $(0,3)+(3,1) = (3, 4) = $ queen ✓.
Normalise every row of $E$ to length 1 to get $\bar E$. For a query vector $\mathbf{q}$ (normalised to $\bar{\mathbf{q}}$), all the cosine similarities come out of one matrix–vector product:
$$\mathbf{s} = \bar E\,\bar{\mathbf{q}}, \qquad s_i = \cos(\text{angle between word } i \text{ and the query}).$$The $k$ nearest neighbours are the $k$ largest entries of $\mathbf{s}$ (excluding the query itself).
Analogy "a is to b as c is to ?": form $\mathbf{t} = \mathbf{x}_a - \mathbf{x}_b + \mathbf{x}_c$ and return the word nearest to $\mathbf{t}$, leaving out $a$, $b$ and $c$.
Why do we need it?
Once things are vectors, 'which one is most like this?' becomes a calculation. We need a similarity that ignores vector length, and a way to see a relationship as a direction.
Where is it used?
Semantic search, retrieval-augmented generation (RAG), 'users like you' recommendations, duplicate detection, and word-analogy tests.
How is it used?
Normalise the rows of E, multiply by the normalised query to get every cosine in one product, and take the top k with argsort. For an analogy form a − b + c and leave out the three input words.
- This is a hand-made 2D toy. Real embeddings have 100 to 4000 dimensions and are learned from text. In them, analogies work only roughly, and they also copy biases found in the training text.
- Always exclude the input words when answering an analogy, otherwise "king" or "woman" is usually the closest point.
- For millions of vectors, exact search is too slow. Approximate nearest-neighbour indexes (HNSW, FAISS) are used instead; they still rely on dot products and cosines.
Quick check: why compare embeddings with cosine rather than plain distance?
Cosine compares direction only, so a frequent word whose vector happens to be long is not unfairly treated as "far" from a rarer word pointing the same way. (After normalising to length 1, the two ways of comparing give the same ranking.)
Recommendation systems and matrix factorisation core
What we need from earlier chapters: rank (Chapter 1.7), least squares and ridge (Chapter 1.10, every ALS step is one), and the SVD and best low-rank approximation (Chapter 1.13).
Picture a table: users down the side, films across the top, and a rating in each cell. Most cells are empty, because nobody watches everything. The job is to guess the empty cells, so we can recommend what a user would probably enjoy.
The key assumption: tastes have only a few themes (say action and romance). Every user is described by how much they like each theme. Every film is described by how much of each theme it contains. A rating is then roughly "taste · content", a dot product of two short vectors.
If that is true, the huge table is really the product of two thin matrices. We say the table has low rank. Learn the two thin matrices from the cells we can see, multiply them, and every empty cell gets a prediction.
Two themes: [action, romance]. Dee's taste is $\mathbf{u} = [1, 2]$ (a little action, a lot of romance). Film B has content $\mathbf{v} = [0.5, 2]$.
- Predicted rating $= \mathbf{u}\cdot\mathbf{v} = 1\cdot0.5 + 2\cdot2 = 4.5$.
- For an action film $[2, 0.2]$: $1\cdot2 + 2\cdot0.2 = 2.4$. Lower, as expected.
With 6 users and 5 films stored as $U$ ($6\times2$) and $V$ ($5\times2$), the whole $6\times5$ table is $UV^\top$. That is $6\cdot2+5\cdot2 = 22$ numbers instead of $30$, and the saving grows quickly with size: for a million users and ten thousand films at rank 20 it is about 20 million numbers instead of ten billion.
Let $R\in\mathbb{R}^{m\times n}$ be the ratings, with $\Omega$ the set of observed cells. A rank-$k$ factorisation is
$$R \approx UV^\top,\quad U\in\mathbb{R}^{m\times k},\ V\in\mathbb{R}^{n\times k},\qquad \hat r_{ij} = \mathbf{u}_i\cdot\mathbf{v}_j.$$Choose $U,V$ to minimise the error on observed cells only, with a ridge penalty:
$$\min_{U,V}\ \sum_{(i,j)\in\Omega}\big(r_{ij}-\mathbf{u}_i\cdot\mathbf{v}_j\big)^2 + \lambda\big(\|U\|_F^2+\|V\|_F^2\big).$$- SVD recommender: fill the blanks with something (0, or each film's average), take the SVD, keep the top $k$ singular values. This is the best rank-$k$ approximation of the filled table (Chapter 1.13), but it treats the fill-ins as real ratings.
- Alternating least squares (ALS): freeze $V$. Now each user's vector $\mathbf{u}_i$ is an independent ridge least-squares problem, using only the films that user rated: $\mathbf{u}_i = (V_\Omega^\top V_\Omega+\lambda I)^{-1}V_\Omega^\top\mathbf{r}_\Omega$. Then freeze $U$ and solve for every film the same way. Repeat. Each half-step can only lower the loss.
- Missing values never enter the loss. That is the big advantage of ALS over filled-in SVD.
Why do we need it?
Most cells of a ratings table are empty, but we still want to guess them. If tastes have only a few themes, two thin matrices can store the whole table and predict the blanks.
Where is it used?
Movie, music and shopping recommendations (the Netflix Prize winners used this), two-tower recommenders at large companies, and filling gaps in survey or sensor tables.
How is it used?
Hide some known ratings as a test. Fit U and V on the visible cells with ALS or gradient descent (with a ridge penalty). Pick the rank k by the error on the hidden cells. Predict a blank cell as u_i · v_j.
- Overfitting: too large a rank (or too small a $\lambda$) lets the model memorise the known cells. Always judge on ratings you hid, never on the ones you trained on.
- The factors are not unique: rotating $U$ by $Q$ and $V$ by the same $Q$ gives the same $UV^\top$. So the "taste dimensions" are not automatically labelled "action" and "romance".
- Cold start: a brand-new user or film has no ratings, so no vector can be learned yet.
- The joint problem in $U$ and $V$ is not convex (two unknown matrices multiplied), but each ALS half-step is a convex least-squares problem. Different starting points can land in different answers.
Quick check: why can ALS solve each user's vector separately?
With $V$ frozen, the loss is a sum over users, and user $i$'s vector appears only in user $i$'s own terms. So the big problem splits into $m$ small ridge least-squares problems (each of size $k\times k$) that do not affect each other.
Neural-network layers: the dense layer and why we need non-linearities core
What we need from earlier chapters: matrix–vector and matrix–matrix products (Chapter 1.4), composition of linear maps and affine maps (Chapter 1.5) and the linear combination (Chapter 1.2).
A dense layer (or "fully connected" or "linear" layer) takes a vector of numbers in and produces a new vector of numbers out. Each output number is a neuron: it takes its own weighted sum of all the inputs and adds its own bias. A row of the weight matrix $W$ holds one neuron's weights. So the whole layer is one matrix–vector product, plus a bias vector.
There is a catch. Stack two such layers with nothing in between and you get one layer in disguise, because a linear map followed by a linear map is a linear map ($W_2W_1$ is just another matrix). A hundred layers would be no more powerful than one.
The fix is a small bend after each layer: a simple function, such as ReLU ("replace negatives by 0"), applied to each number separately. The bends let layers build curved, complicated functions out of flat pieces.
$W = \begin{bmatrix}1&2\\3&-1\end{bmatrix}$, $\mathbf{b} = [0, 1]$, input $\mathbf{x} = [1, 2]$.
- $W\mathbf{x} = [1\cdot1+2\cdot2,\ 3\cdot1+(-1)\cdot2] = [5, 1]$.
- Add the bias: $\mathbf{z} = [5, 1+1] = [5, 2]$.
- ReLU: $\max(0,5) = 5$ and $\max(0,2)=2$, so $\mathbf{h}=[5,2]$ (here nothing was negative; $[-3,2]$ would become $[0,2]$).
Collapse test (ignore biases). With a second matrix $W_2 = \begin{bmatrix}1&1\\0&1\end{bmatrix}$: $W_2(W\mathbf{x}) = W_2[5,1] = [6, 1]$. And $W_2W = \begin{bmatrix}4&1\\3&-1\end{bmatrix}$, so $(W_2W)\mathbf{x} = [4+2,\ 3-2] = [6,1]$. Same answer: two layers were really one.
For a batch $X\in\mathbb{R}^{n\times d_{\text{in}}}$ (one example per row), a dense layer with $W\in\mathbb{R}^{d_{\text{out}}\times d_{\text{in}}}$ and $\mathbf{b}\in\mathbb{R}^{d_{\text{out}}}$ computes
$$Y = XW^\top + \mathbf{1}\mathbf{b}^\top\in\mathbb{R}^{n\times d_{\text{out}}}, \qquad H = \varphi(Y)\ \ (\varphi\text{ applied to every entry}).$$- For one example this is $\mathbf{y} = W\mathbf{x}+\mathbf{b}$, an affine map (Chapter 1.5). The transpose appears because the examples are rows.
- Number of learned parameters: $d_{\text{out}}(d_{\text{in}}+1)$.
- Collapse: $W_2(W_1\mathbf{x}+\mathbf{b}_1)+\mathbf{b}_2 = (W_2W_1)\mathbf{x} + (W_2\mathbf{b}_1+\mathbf{b}_2)$. Always a single affine map.
- Common non-linearities: $\text{ReLU}(z)=\max(0,z)$, $\tanh$, the sigmoid $\sigma$, and GELU (a smooth ReLU).
Why do we need it?
A network needs a way to mix its inputs into new features, and a bend in between the mixes. Without the bend, any number of layers is secretly one matrix.
Where is it used?
Every multi-layer perceptron, the feed-forward blocks and the attention projections inside Transformers, and the classifier heads of image models.
How is it used?
Store W (outputs × inputs) and b, compute XWᵀ + b for a whole batch in one matrix product, apply ReLU or GELU, and repeat. Check the shapes first: most bugs here are shape mismatches.
- The bias is not optional decoration: without it every layer maps $\mathbf{0}$ to $\mathbf{0}$, so the shape can never shift away from the origin.
- Check shapes. If $W$ is $d_{\text{out}}\times d_{\text{in}}$ then $W\mathbf{x}$ needs $\mathbf{x}\in\mathbb{R}^{d_{\text{in}}}$. Libraries differ on whether they store $W$ or $W^\top$; the maths is the same.
- ReLU throws away information (all negatives become 0), and that is a feature, not a bug: it is what makes the map non-linear.
Quick check: a dense layer maps 100 inputs to 50 outputs. How many parameters?
$W$ is $50\times100 = 5000$ numbers and $\mathbf{b}$ has 50, so $5050$ parameters in total, which matches $d_{\text{out}}(d_{\text{in}}+1) = 50\cdot101$.
Weight initialisation: keeping signals the right size core
What we need from earlier chapters: the L2 norm (Chapter 1.2), orthogonal matrices preserve length (Chapter 1.9), and the matrix–vector product (Chapter 1.4). We also use one fact from statistics: the variance of a sum of independent numbers is the sum of their variances.
Training starts from random weights. Think of passing a signal through many layers as photocopying a photocopy. If each copy is slightly too faint, after 50 copies the page is blank. If each copy is slightly too dark, after 50 copies it is solid black. The same happens to the gradients on the way back.
We want each layer to keep the signal's size roughly the same. That means choosing the size of the random weights carefully, based on how many numbers are being added up.
A square layer with $n = 100$ inputs and weights drawn independently with standard deviation $\sigma$. Each output is a sum of 100 terms, so the squared length of the signal gets multiplied by about $n\sigma^2$ per layer.
- $\sigma = 0.1$: factor $100\cdot0.01 = 1$. Length stays the same, which is what we want. (This is $\sigma = 1/\sqrt n$.)
- $\sigma = 0.05$: factor $0.25$, so lengths halve each layer. After 10 layers: $0.5^{10}\approx0.001$. Vanishing.
- $\sigma = 0.2$: factor $4$, so lengths double each layer. After 10 layers: $2^{10} = 1024$. Exploding.
With ReLU, about half of the entries are zeroed, so only half the squared length survives. We must double the variance to compensate: $\sigma^2 = 2/n$.
| Scheme | Variance of each weight | Designed for |
|---|---|---|
| Xavier / Glorot | $1/n_{\text{in}}$, or $2/(n_{\text{in}}+n_{\text{out}})$ | linear, tanh, sigmoid |
| He / Kaiming | $2/n_{\text{in}}$ | ReLU |
| Orthogonal | $W = Q$ from the QR of a random matrix (times a gain) | any depth; $\|Q\mathbf{x}\| = \|\mathbf{x}\|$ exactly |
The common idea is norm preservation: the length of the activations should neither shrink nor grow from layer to layer. An orthogonal matrix achieves this perfectly (Chapter 1.9), with no randomness in the singular values: all of them equal 1.
Why do we need it?
A deep network multiplies the signal by a matrix at every layer. If each matrix shrinks or grows the signal a little, it vanishes or explodes before training even starts.
Where is it used?
Keras's default (Xavier, also called Glorot), the Xavier and He helpers in torch.nn.init, very deep networks, recurrent networks (orthogonal initialisation), and any model that fails to train from the first step.
How is it used?
Pick the rule to match the activation: He for ReLU, Xavier for tanh or linear, orthogonal for recurrent nets. Draw random weights with that variance, run one batch and check that the activation sizes stay about the same from layer to layer.
- All of this is about the start of training. After some steps the weights are no longer random, but a good start decides whether training can begin at all.
- The "$n$" is the number of inputs being summed, called the fan-in. For convolutions it is the kernel size times the input channels.
- Normalisation layers (below) and residual connections reduce, but do not remove, the need for sensible initialisation.
Quick check: why does ReLU need $\sigma^2 = 2/n$ instead of $1/n$?
ReLU sets about half of the numbers to zero, so only about half of the squared length passes through. Doubling the variance of the weights doubles the squared length produced by each layer, which exactly cancels the loss.
Backpropagation as a chain of Jacobian products core
What we need from earlier chapters: the Jacobian, the chain rule and gradients (Chapter 1.14), the transpose and matrix products (Chapter 1.4) and the outer product (Chapters 1.2 and 1.4).
A network is a chain of small steps, and the final loss depends on every weight through that chain. The chain rule says: to know how a weight at the start affects the loss, multiply the local slopes along the way. For vectors, each local slope is a matrix: the Jacobian.
The clever part is the order. Start at the loss and work backwards. At each step you carry one row of numbers (the gradient so far) and multiply it by the next Jacobian. That is a cheap vector-times-matrix, never a huge matrix-times-matrix. Backprop is just this order of multiplication, done once for the whole network.
A tiny network: $\mathbf{z} = W_1\mathbf{x}$, $\mathbf{h} = \text{ReLU}(\mathbf{z})$, $y = \mathbf{w}_2\cdot\mathbf{h}$, loss $L = \tfrac12(y-t)^2$. Take $\mathbf{x}=[1,2]$, $W_1 = \begin{bmatrix}1&-1\\2&0\end{bmatrix}$, $\mathbf{w}_2=[1,3]$, target $t=5$.
Forward:
- $\mathbf{z} = [1-2,\ 2+0] = [-1, 2]$. $\mathbf{h} = [0, 2]$. $y = 1\cdot0+3\cdot2 = 6$. $L = \frac12(6-5)^2 = 0.5$.
Backward (each step multiplies by a Jacobian):
- $\partial L/\partial y = y - t = 1$.
- $\partial L/\partial\mathbf{h} = \mathbf{w}_2\cdot1 = [1, 3]$ (the Jacobian of $y=\mathbf{w}_2\cdot\mathbf{h}$ is $\mathbf{w}_2^\top$).
- ReLU's Jacobian is $\operatorname{diag}(1[z>0]) = \operatorname{diag}(0,1)$, so $\partial L/\partial\mathbf{z} = [0\cdot1,\ 1\cdot3] = [0, 3]$. Unit 1 was off, so it passes back nothing.
- $\partial L/\partial W_1 = (\partial L/\partial\mathbf{z})\,\mathbf{x}^\top = \begin{bmatrix}0\\3\end{bmatrix}[1,2] = \begin{bmatrix}0&0\\3&6\end{bmatrix}$, an outer product. Also $\partial L/\partial\mathbf{w}_2 = (y-t)\,\mathbf{h} = [0, 2]$.
If a step computes $\mathbf{y} = f(\mathbf{x})$ with Jacobian $J = \partial\mathbf{y}/\partial\mathbf{x}$, then the gradient flows backwards by the vector–Jacobian product
$$\frac{\partial L}{\partial\mathbf{x}} = J^\top\frac{\partial L}{\partial\mathbf{y}}.$$- Dense layer $\mathbf{y} = W\mathbf{x}$: $J = W$, so $g_{\mathbf{x}} = W^\top g_{\mathbf{y}}$, and the weight gradient is the outer product $g_W = g_{\mathbf{y}}\mathbf{x}^\top$. For a batch: $g_W = G_Y^\top X$ and $G_X = G_Y W$.
- Elementwise function $\varphi$: $J = \operatorname{diag}(\varphi'(\mathbf{z}))$, so just multiply entry by entry.
- The backward pass costs about twice the forward pass, and it reuses the forward numbers ($\mathbf{x}$, $\mathbf{z}$) that were saved.
- Going backwards (reverse mode) is efficient because the loss is a single number, so the carried object is one row vector, not a full Jacobian matrix.
Why do we need it?
A network has millions of weights, and we need every weight's gradient cheaply. Multiplying Jacobians from the loss backwards gets all of them for about twice the cost of one forward pass.
Where is it used?
Training every neural network. PyTorch's loss.backward() and TensorFlow's GradientTape are automated versions of exactly this chain of products.
How is it used?
Run forward and save the intermediate values. Then go backwards: at each layer multiply the incoming gradient by Wᵀ (and by ReLU's 0/1 mask) and form the weight gradient as an outer product. Compare with a numerical gradient to debug.
- We never build the full Jacobian of a big layer: it would have (outputs × inputs) entries. We only ever compute its product with a vector.
- The forward pass must save its intermediate values ($\mathbf{x}$, $\mathbf{z}$). That is why training needs far more memory than inference.
- Repeated multiplication by $W^\top$ is why initial scale matters: if $\|W\|$ is too big or too small, gradients explode or vanish on the way back.
Quick check: in the worked example, why is the first row of $\partial L/\partial W_1$ all zeros?
The first hidden unit had $z_1=-1\le0$, so ReLU output $0$ and its slope is $0$. Changing the weights feeding that unit (row 1 of $W_1$) does not change $y$, so the loss does not care about them (right now).
Convolution is a structured, sparse matrix multiplication
What we need from earlier chapters: the matrix–vector product, where each output is a row dotted with the input (Chapter 1.4), the dot product (Chapter 1.2) and special matrices (banded, sparse) (Chapter 1.4).
A convolution slides one small window of weights (the kernel) along the input. At every position it takes a weighted sum of the few numbers under the window. A weighted sum is a dot product, and a dot product of a row with $\mathbf{x}$ is exactly what one row of a matrix does.
So a convolution is a matrix multiplication, by a very special matrix: mostly zeros, with each row a shifted copy of the same few numbers. The same kernel is reused everywhere ("weight sharing"). That is why a convolution layer has so few parameters compared with a dense layer.
Input $\mathbf{x} = [1,2,3,4,5]$ and kernel $\mathbf{k} = [1, 0, -1]$ (a "slope detector"). Slide it with no padding:
- Position 1: $1\cdot1 + 0\cdot2 + (-1)\cdot3 = -2$.
- Position 2: $1\cdot2 + 0\cdot3 + (-1)\cdot4 = -2$.
- Position 3: $1\cdot3 + 0\cdot4 + (-1)\cdot5 = -2$.
As a matrix ($3\times5$, each row the kernel shifted one step right):
$$\begin{bmatrix}1&0&-1&0&0\\0&1&0&-1&0\\0&0&1&0&-1\end{bmatrix}\begin{bmatrix}1\\2\\3\\4\\5\end{bmatrix} = \begin{bmatrix}-2\\-2\\-2\end{bmatrix}.$$Same answer. A dense $3\times5$ layer would have 15 free numbers; this one has only 3.
For a kernel $\mathbf{k}$ of length $m$ (deep-learning "convolution", which is technically cross-correlation):
$$y_i = \sum_{j=0}^{m-1} k_j\,x_{i+j} \qquad\Longleftrightarrow\qquad \mathbf{y} = T\mathbf{x},\quad T_{i,\,i+j} = k_j,\ \text{zero elsewhere}.$$- $T$ is banded and Toeplitz (constant along each diagonal). Only the $m$ numbers in $\mathbf{k}$ are free.
- The backward pass multiplies by $T^\top$, which is again a convolution (with the kernel reversed).
- For images the same idea holds with a 2D kernel: the matrix is bigger and block-structured, but still sparse with shared entries.
- In practice libraries never build $T$ (mostly zeros!); they slide the kernel, or use the FFT. The matrix view explains what is computed.
Why do we need it?
A dense layer on an image needs a weight for every pair of pixels, which is far too many. A convolution reuses one small kernel everywhere. Seen as a matrix, it is mostly zeros with repeating rows.
Where is it used?
Image classifiers such as ResNet, speech and audio models, 1D convolutions on text and sensor streams, and U-Nets for image segmentation and generation.
How is it used?
Slide a small kernel over the input (nn.Conv1d or nn.Conv2d). Each output is a dot product with the window. Libraries never build the big matrix. This view just helps you count parameters and understand the backward pass.
- Strictly, the maths definition of convolution flips the kernel first. Deep-learning libraries skip the flip (it only relabels the learned numbers). The matrix picture is the same.
- Padding and stride change the shape of $T$ (extra columns, skipped rows) but not the idea.
Quick check: a 1D kernel of length 3 is applied to an input of length 10 with no padding. What is the shape of $T$?
There are $10-3+1 = 8$ output positions, so $T$ is $8\times10$. It still has just 3 free numbers.
Normalisation layers: centring plus scaling
What we need from earlier chapters: the L2 norm and normalising a vector (Chapter 1.2), projection (Chapter 1.2) and centring the data (earlier in this chapter).
As numbers flow through many layers, their typical size and average can drift. A normalisation layer puts them back on a standard footing: subtract the average (centre) and divide by the spread (scale). Afterwards the numbers have average 0 and spread 1, so the next layer always sees well-behaved input.
The only question is which numbers do you average over?
- LayerNorm: over the features of one example (across a row).
- BatchNorm: over the examples in the batch, for one feature (down a column).
LayerNorm on one example $\mathbf{x} = [2, 4, 6]$:
- Mean: $\mu = (2+4+6)/3 = 4$. Centred: $[-2, 0, 2]$.
- Variance: $\sigma^2 = (4+0+4)/3 = 8/3$, so $\sigma\approx1.633$.
- Divide: $[-2, 0, 2]/1.633 \approx [-1.225, 0, 1.225]$. The mean is $0$ and the variance is $1$ ✓.
Geometry: subtracting the mean is a projection that removes the component along the all-ones vector $\mathbf{1}$: $\mathbf{x}_c = (I - \tfrac1d\mathbf{1}\mathbf{1}^\top)\mathbf{x}$. Dividing by $\sigma$ then rescales the result to length exactly $\sqrt d$. So LayerNorm pushes every vector onto a sphere inside the "mean-zero" subspace.
- $\epsilon$ (about $10^{-5}$) prevents division by zero when all entries are equal.
- $\boldsymbol\gamma$ and $\boldsymbol\beta$ are learned: they let the network undo the normalisation if that helps.
- BatchNorm is the same formula, but $\mu$ and $\sigma^2$ are computed down each column of the batch matrix (and, at test time, replaced by running averages).
- RMSNorm, used in many modern language models, skips the centring and divides by the root-mean-square only.
Why do we need it?
Between layers the numbers can drift to odd sizes, which makes training slow or unstable. Normalising resets them to average 0 and spread 1 before the next layer sees them.
Where is it used?
LayerNorm inside every Transformer block, RMSNorm in many modern language models, and BatchNorm in convolutional networks.
How is it used?
Choose the axis: LayerNorm works across the features of one example, BatchNorm across the examples of a batch. Subtract the mean, divide by √(variance + ε), then apply a learned scale γ and shift β.
- BatchNorm depends on the other examples in the batch, so it behaves differently in training (batch statistics) and testing (running averages). LayerNorm looks only at one example, which is why Transformers prefer it.
- With a batch of 1, BatchNorm statistics are meaningless.
- Normalisation removes information about the original mean and scale of each vector. $\boldsymbol\gamma,\boldsymbol\beta$ give some of it back.
Quick check: LayerNorm of $[5,5,5,5]$ (ignoring $\epsilon$ trouble). What happens?
The mean is 5, so the centred vector is all zeros, and the variance is 0. The result is $0/\sqrt{0+\epsilon}=0$. A constant vector has no "shape" left after centring, so it maps to the zero vector (plus $\boldsymbol\beta$).
Matrix efficiency: the big idea has a name core
What we need from earlier chapters: rank (Chapter 1.7), the outer-product view of matrix multiplication (Chapter 1.4) and the SVD and low-rank approximation (Chapter 1.13). This is the first of eight short sections about one idea that makes modern AI affordable.
Big models are built from huge tables of numbers (matrices). Every number costs something: memory to store it, time to multiply by it and energy to move it around. "Matrix efficiency" means doing the same job with fewer or smaller numbers.
The trick that works best is based on a simple fact: a big table is often far simpler than it looks. Think of a multiplication table. It has 100 cells, but you can rebuild every cell from just two short lists, the row labels and the column labels. If every row of a big table is the same pattern, only stretched, you can store one pattern and one list of stretch factors.
Writing a big table as a product of two skinny tables is called low-rank factorisation (or low-rank approximation). Remember this name. Everything in this part of the chapter is built on it.
Two short lists, $\mathbf{u} = [1,2,3]$ and $\mathbf{v}=[4,5,6]$. Their outer product $\mathbf{u}\mathbf{v}^\top$ builds a full $3\times3$ table:
$$\mathbf{u}\mathbf{v}^\top = \begin{bmatrix}1\cdot4&1\cdot5&1\cdot6\\2\cdot4&2\cdot5&2\cdot6\\3\cdot4&3\cdot5&3\cdot6\end{bmatrix} = \begin{bmatrix}4&5&6\\8&10&12\\12&15&18\end{bmatrix}.$$The table has 9 numbers but we only needed $3+3=6$ to describe it. Every row is a multiple of the first row, so the table has rank 1. For a $1000\times1000$ table of this kind: $2000$ numbers instead of $1{,}000{,}000$, which is 500 times fewer.
Low-rank factorisation. Approximate a $d\times d$ matrix $M$ by a product of two skinny matrices:
$$M \approx B\,A,\qquad B\in\mathbb{R}^{d\times r},\quad A\in\mathbb{R}^{r\times d},\quad r\ll d.$$- It stores $2dr$ numbers instead of $d^2$. The matrix $BA$ has rank at most $r$.
- The best possible rank-$r$ approximation comes from the SVD: keep the $r$ largest singular values (Eckart–Young, Chapter 1.13).
- LoRA = Low-Rank Adaptation. It uses this idea to fine-tune a big model: freeze the original weights $W$ and learn only a low-rank change $\Delta W = BA$.
- LoRA belongs to a family called parameter-efficient fine-tuning (PEFT): methods that adapt a model by training only a tiny fraction of its numbers (others in the family: adapters, prompt tuning, BitFit).
Why do we need it?
Modern models have billions of numbers. Storing, training and shipping all of them for every new task is too slow, too big and too expensive for most people. We need ways to get the same quality from far fewer numbers.
Where is it used?
LoRA fine-tuning of language models, compressing a trained layer with a truncated SVD, recommenders (R ≈ UVᵀ, earlier in this chapter), PCA, and quantised models on phones and small GPUs.
How is it used?
Ask: "is this matrix (or this change to a matrix) close to low rank?" If yes, replace it by two skinny matrices B and A, with r chosen by checking how much error you can accept.
The family of matrix-efficiency tricks. Low-rank factorisation is one of several. Here is the whole toolbox in one table, so you can place each idea:
| Trick | What it saves | Used when | The catch |
|---|---|---|---|
| Low-rank factorisation (LoRA, truncated SVD) | Number of stored and trained numbers: $2dr$ instead of $d^2$ | Compressing a trained matrix; fine-tuning a big model cheaply | The matrix (or the change) must really be close to low rank |
| Quantisation (8-bit, 4-bit, QLoRA) | Bytes per number: 16 bits down to 4 bits is 4 times smaller | Running huge models on small GPUs; storing the frozen base in QLoRA | Some precision is lost; large outlier values need care |
| Sparsity / pruning | Numbers: set many weights to exactly 0 and skip them | Removing weights that matter little | Only faster if the software and hardware can skip zeros; positions must be stored |
| Distillation | The whole model: train a small "student" to copy a big "teacher" | Deploying fast, small models | Needs a training run; the student can be a little worse |
| Structured matrices (Kronecker, convolution, butterfly) | Numbers and time: the big matrix is built from a few small pieces | Convolutions, some efficient Transformer layers | The structure limits what the matrix can represent |
- These tricks combine. QLoRA stores the frozen base model in 4-bit (quantisation) and trains a low-rank update (LoRA).
- "Low-rank" means few independent directions, not "small numbers". A matrix of huge numbers can have rank 1.
- Almost nothing is free. Each trick trades a little accuracy or flexibility for a lot of savings. The art is to check how little you lose.
Quick check: a $2000\times2000$ matrix is replaced by $B$ ($2000\times10$) and $A$ ($10\times2000$). How many numbers are stored now?
$2\cdot2000\cdot10 = 40{,}000$, compared with $2000^2 = 4{,}000{,}000$: 100 times fewer.
LoRA, part 1: the problem with fine-tuning a big model core
What we need from earlier chapters: the dense layer and its weight matrix (earlier in this chapter), gradient descent (earlier in this chapter) and how matrices are counted (Chapter 1.4).
A pre-trained model has learned general skills from a huge pile of text or images. Fine-tuning teaches it one more thing: answer like a doctor, follow instructions, draw in a certain style.
The obvious way is to let every weight change a little. That is costly in three ways:
- You must store the weights.
- You must store a gradient for every weight (one more number each).
- The optimiser (Adam, a popular version of gradient descent) keeps two more numbers per weight, its running averages, and often a full-precision copy too.
That is roughly 16 bytes for every weight (a rule of thumb: the exact number depends on the software and the settings), instead of 2 bytes just to run the model. And if you have ten tasks you need ten full copies of the model.
One weight matrix in a large model is $4096\times4096$.
- Numbers in one matrix: $4096\times4096 = 16{,}777{,}216$ (about 16.8 million).
- A 7-billion-parameter model has about 32 layers, each with several such matrices, giving about $7{,}000{,}000{,}000$ numbers in total.
- Fine-tuning every one with Adam costs about $16\times7\text{ billion bytes} = 112$ GB of GPU memory, before even counting the text being processed. A good consumer GPU has 24 GB.
Full fine-tuning trains all $P$ parameters. Rough memory with Adam in "mixed precision" (most numbers kept in 16 bits, plus one 32-bit master copy of the weights):
$$\underbrace{2P}_{\text{weights}} + \underbrace{2P}_{\text{gradients}} + \underbrace{8P}_{\text{Adam's two averages}} + \underbrace{4P}_{\text{full-precision copy}} \approx 16P\ \text{bytes}.$$- Plus the activations (saved forward-pass values, from Backprop earlier), which depend on batch size and text length.
- Plus one complete copy of the model on disk for every task ($2P$ bytes in 16-bit).
- Parameter-efficient fine-tuning (PEFT) keeps the original weights frozen and trains only a small number $P_{\text{train}}\ll P$ of new or changed numbers. Gradients and optimiser memory are then only needed for those.
Why do we need it?
Understanding the cost explains why most people could not adapt big models at all. The memory is dominated by gradients and optimiser states, not by the weights themselves, so freezing most weights removes most of the cost.
Where is it used?
Planning GPU memory for any fine-tuning job (LLM chat tuning, speech models, image generators) and deciding between full fine-tuning, LoRA and QLoRA.
How is it used?
Estimate 16 bytes per trainable parameter plus 2 bytes per frozen one, add room for activations, and compare with your GPU's memory. If it does not fit, freeze the base model and train a small adapter.
- These are rough numbers to build intuition. Real memory also depends on batch size, text length, gradient checkpointing and the library.
- LoRA saves the memory for gradients and optimiser states. It does not make the forward pass cheaper: the full base model still has to run.
Quick check: a model has 3 billion parameters. Roughly how much GPU memory do the weights, gradients and Adam states of a full fine-tune need?
About $16\times3\text{ billion} = 48$ GB, before activations.
LoRA, part 2: the key idea, a simple change is a low-rank change core
What we need from earlier chapters: rank (Chapter 1.7), the SVD, singular values and the Eckart–Young theorem (Chapter 1.13) and PCA reconstruction (earlier in this chapter, the same truncation idea).
Do not think of fine-tuning as learning a new weight matrix. Think of it as learning a change: $W_{\text{new}} = W + \Delta W$. Here is the bet behind LoRA:
teaching an already-clever model one new thing needs only a few directions of change, so $\Delta W$ is "simple".
"Simple" has an exact meaning: low rank. A matrix of rank $r$ is a sum of just $r$ layers of the form (a column) times (a row). The SVD lists these layers in order of importance: the biggest singular values come first, and the rest are often small. Keeping only the top $r$ layers loses little.
So instead of storing all $d^2$ numbers of $\Delta W$, store the $r$ layers: $\Delta W = BA$ with $B$ having $r$ columns and $A$ having $r$ rows. This is a bet, not a law: it was found to work well in practice for many tasks.
Suppose the change we need is $\Delta W = \begin{bmatrix}2&4&6\\1&2&3\\3&6&9\end{bmatrix}$. Every row is a multiple of $[1,2,3]$, so the rank is 1:
$$\Delta W = \begin{bmatrix}2\\1\\3\end{bmatrix}\begin{bmatrix}1&2&3\end{bmatrix} = BA,\qquad B = \begin{bmatrix}2\\1\\3\end{bmatrix}\ (3\times1),\ A = [1,2,3]\ (1\times3).$$Nine numbers become $3+3=6$. For a real $4096\times4096$ matrix with $r=8$: $2\cdot4096\cdot8 = 65{,}536$ numbers instead of $16{,}777{,}216$, which is 256 times fewer (0.39%).
- Parameters: $2dr$ instead of $d^2$ (for a $d_{\text{out}}\times d_{\text{in}}$ matrix: $r(d_{\text{out}}+d_{\text{in}})$).
- $\operatorname{rank}(BA)\le r$, because $BA = \sum_{k=1}^r \mathbf{b}_k\mathbf{a}_k^\top$ is a sum of $r$ outer products (Chapter 1.4).
- If $\Delta W$ has SVD $\sum_i \sigma_i\mathbf{u}_i\mathbf{v}_i^\top$, the best rank-$r$ approximation keeps the top $r$ terms, and the error is $\sqrt{\sigma_{r+1}^2+\sigma_{r+2}^2+\cdots}$ (Eckart–Young). LoRA does not compute the SVD; it learns $B$ and $A$ directly by gradient descent.
- Note: $W$ itself is usually not low rank. Only the change is assumed to be.
Why do we need it?
We need the update to be small enough to train and store cheaply. The low-rank assumption is exactly what lets two skinny matrices stand in for a huge one, and the SVD tells us how much accuracy we give up.
Where is it used?
LoRA for language models and image generators, truncated-SVD compression of trained layers, and the low-rank structure behind recommenders and PCA earlier in this chapter.
How is it used?
Pick a small rank r (4, 8 or 16 are common starting points). If the fine-tuned model underfits, raise r. To test the idea on a known update, take its SVD and see how fast the singular values fall.
- Low rank is an assumption about the update, tested by experiments, not a theorem. If the task needs a big rewrite of the model, a small $r$ underfits.
- In real LoRA you never see $\Delta W$ or its SVD during training. You only learn $B$ and $A$. The SVD is how we reason about what $BA$ can represent.
- Rank is not the same as size. A rank-1 update can still change every entry of $W$.
Quick check: the singular values of an update are $10, 5, 0.1, 0.05$. Which rank would you pick, and why?
$r=2$. The first two values are big and the rest are tiny: the relative error is $\sqrt{0.1^2+0.05^2}/\sqrt{100+25+0.01+0.0025}\approx0.010$, about 1%.
LoRA, part 3: how the forward pass works core
What we need from earlier chapters: the matrix–vector product and why the order of multiplication matters (Chapter 1.4), the dense layer (earlier in this chapter) and counting FLOPs (Chapter 1.15).
A LoRA layer has two paths that start at the same input $\mathbf{x}$ and are added at the end:
- Main road: the original frozen weights, $W\mathbf{x}$. Unchanged.
- Side road: squeeze $\mathbf{x}$ down to only $r$ numbers with $A$, then expand it back to $d$ numbers with $B$. This is the learned correction.
The side road is cheap because of the order. Compute $A\mathbf{x}$ first: that is only $r$ numbers. Then $B$ times those $r$ numbers. We never build the big $d\times d$ matrix $BA$ during the forward pass.
Tiny sizes: $d=4$, $r=1$, $\alpha=2$ (so $\alpha/r=2$). $\mathbf{x} = [1,2,0,1]$, $A = [1, 0, 1, 1]$ ($1\times4$), $B = [1,2,0,1]^\top$ ($4\times1$).
- Squeeze: $A\mathbf{x} = 1\cdot1 + 0\cdot2 + 1\cdot0 + 1\cdot1 = 2$. Just one number.
- Expand: $B(A\mathbf{x}) = 2\cdot[1,2,0,1] = [2,4,0,2]$.
- Scale by $\alpha/r=2$: $[4,8,0,4]$. Add this to $W\mathbf{x}$.
Building $BA$ first would have made a $4\times4$ matrix (16 numbers) before multiplying. The squeeze-then-expand order used 4 + 4 multiplications instead of 16 + 16.
For a column input $\mathbf{x}\in\mathbb{R}^{d}$:
$$\mathbf{h} = W\mathbf{x} + \frac{\alpha}{r}\,B\,(A\mathbf{x}),\qquad W\ \text{frozen},\quad A\in\mathbb{R}^{r\times d},\ B\in\mathbb{R}^{d\times r}\ \text{trained}.$$- Cost per token (multiply and add = 2 FLOPs): main road $2d^2$; side road $2dr + 2dr = 4dr$. The side road adds only a fraction $2r/d$ of the main road.
- Only $A$ and $B$ get gradients. $W$ gets none, so no gradient or optimiser memory is spent on it. (The gradient still flows through $W$ to earlier layers.)
- $\alpha$ ("alpha") is a fixed scale knob. The factor $\alpha/r$ multiplies the correction.
- For a batch of row vectors $X\in\mathbb{R}^{n\times d}$ (how libraries store data): $H = XW^\top + \frac\alpha r (XA^\top)B^\top$. Some blogs write this as $xW + xAB$ with $A$ and $B$ named for the transposes; the maths is identical.
Why do we need it?
We want the model to behave as if its weights were changed, without ever storing or multiplying a changed d × d matrix during training. Squeezing through r numbers makes the correction cheap in both memory and time.
Where is it used?
The LoRA layers in Hugging Face PEFT, the attention projections of fine-tuned language models, and image-generator LoRAs applied to the attention layers of Stable Diffusion.
How is it used?
Wrap an existing linear layer: keep W frozen, add A and B, and in the forward pass return W x + (α/r)·B(A x). Pass A x through first. Only A and B go to the optimiser.
- Order matters: $B(A\mathbf{x})$ and $(BA)\mathbf{x}$ give the same answer, but the first costs about $4dr$ and the second about $2d^2r$ plus $2d^2$. Always squeeze first.
- LoRA does not make the forward pass of the base model faster. It makes training memory small. After merging (part 6) there is no extra inference cost either.
Quick check: $d=1000$, $r=10$. How large is $A\mathbf{x}$, and what is the side-road cost per token?
$A\mathbf{x}$ has $r=10$ numbers. The side road costs $4dr = 40{,}000$ FLOPs, against $2d^2=2{,}000{,}000$ for the main road: 2%.
LoRA, part 4: starting safely (B = 0), the $\alpha/r$ scale and choosing $r$
What we need from earlier chapters: weight initialisation (earlier in this chapter), the L2 norm (Chapter 1.2) and the Frobenius norm of a matrix (Chapter 1.7).
A pre-trained model is already good. Fine-tuning should start from exactly that model and move away gradually, not begin with a damaged one.
LoRA makes sure of this with a simple trick. Start with $B = 0$ (all zeros) and $A$ small random numbers. Then $BA$ is the zero matrix, so $\Delta W = 0$ and the model's outputs are identical to the original. Gradient descent then grows $B$ step by step.
Why not make both zero? If both were zero, neither would get a gradient and nothing could ever move ($A$'s gradient contains $B$, and $B$'s contains $A$). One random factor, one zero factor, is the safe start.
$d=3$, $r=1$. Start: $B = [0,0,0]^\top$, $A=[0.5,-1,2]$. Then $BA$ is a $3\times3$ matrix of zeros, and $W\mathbf{x} + BA\mathbf{x} = W\mathbf{x}$: the model is unchanged.
After a few gradient steps, $B$ becomes $[0.1, 0, -0.2]^\top$. Now $BA = \begin{bmatrix}0.05&-0.1&0.2\\0&0&0\\-0.1&0.2&-0.4\end{bmatrix}$, a small rank-1 change. It grew smoothly from zero.
The scale. With $\alpha=8$ and $r=4$ the correction is multiplied by $\alpha/r=2$. If you later change to $r=8$ and keep $\alpha=8$, the factor becomes $1$, which keeps the size of the update roughly comparable, so you do not have to re-tune the learning rate from scratch.
- Initialisation: $A\sim$ small random numbers (a normal or uniform distribution), $B=0$. So $\Delta W=\frac\alpha r BA=0$ at step 0.
- Scaling: the correction is $\frac\alpha r BA$. $\alpha$ is a fixed constant you choose (common: $\alpha = r$ or $\alpha=2r$).
- Which matrices get LoRA? The original paper used the attention query and value projections ($W_Q$, $W_V$). Today people often add the other attention matrices and the MLP matrices too, because more places give better quality for more adapter size.
- Choosing the rank: start with $r=8$ or $16$. If the model underfits, raise it (32, 64). If $r$ stops helping, the task does not need more. Cost grows linearly with $r$.
Why do we need it?
If the adapter started out random, it would immediately corrupt the pre-trained model, which would then have to un-learn the damage. Starting from ΔW = 0 keeps everything the model already knows.
Where is it used?
Every standard LoRA library, including Hugging Face PEFT, Microsoft's original LoRA code and the training scripts for Stable Diffusion LoRAs. The defaults are B = 0 and a scale of α/r.
How is it used?
Choose r (8 or 16), choose α (often r or 2r), pick which layers to adapt, and train. Before training, check that the output of the model with the adapter equals the output without it.
- It is a common bug to initialise both factors to zero. Then no gradient ever flows and the adapter stays at exactly zero forever.
- Raising $r$ does not automatically help. Beyond the rank the task really needs, you only pay more memory.
- The scale $\alpha/r$ is a convention that helps keep learning rates comparable across ranks. Treat $\alpha$ as one more setting to tune if results look off.
Quick check: why not initialise both $A$ and $B$ to zero?
The gradient with respect to $B$ is proportional to $A$ and the gradient with respect to $A$ is proportional to $B$. If both are zero, both gradients are zero, so training can never move them.
LoRA, part 5: watch it learn, full fine-tune versus LoRA core
What we need from earlier chapters: gradient descent (earlier in this chapter), backprop gradients as products (earlier in this chapter) and the Eckart–Young limit on low-rank approximation (Chapter 1.13).
Let us test the bet in a tiny world. Imagine the model needs a change $T$ (a $16\times16$ table). We train in two ways: full fine-tuning moves all 256 numbers of $\Delta W$ directly; LoRA moves only the entries of $B$ and $A$, with $\Delta W=BA$.
If $T$ really is simple (low rank), a small LoRA should reach nearly the same quality using a fraction of the numbers. If the rank is too small, LoRA cannot do better than the best low-rank approximation of $T$, and will plateau there.
The target change here is a rank-2 matrix: $T = 3\,\mathbf{u}_1\mathbf{v}_1^\top + 2\,\mathbf{u}_2\mathbf{v}_2^\top$, with unit vectors. Its squared size is $\|T\|_F^2 = 3^2+2^2=13$.
- LoRA with $r=1$ can capture only the biggest layer ($3$). The leftover is the layer with $\sigma=2$: relative error$^2 = 4/13 = 0.31$. It will stop there, whatever you do.
- LoRA with $r\ge2$ can represent $T$ exactly, so its error should reach 0.
- Parameter counts: full 256; LoRA $r=1$: $2\cdot16\cdot1 = 32$; $r=2$: 64; $r=4$: 128.
Toy loss: $L = \tfrac12\|\Delta W - T\|_F^2$ (how far the learned change is from the wanted change). Gradient descent with step size $\eta$:
$$\text{full: }\ \Delta W \leftarrow \Delta W - \eta(\Delta W - T);\qquad \text{LoRA: }\ E = BA - T,\ \ B \leftarrow B - \eta\,EA^\top,\ \ A\leftarrow A - \eta\,B^\top E.$$- For LoRA, the gradients are the matrix-calculus results from Chapter 1.14: $\partial L/\partial B = EA^\top$ and $\partial L/\partial A = B^\top E$.
- We start with $B=0$ and a small random $A$, as in part 4. At first nothing seems to happen: $A$'s gradient is zero while $B=0$. Then $B$ grows and learning accelerates.
- The plotted number is the relative squared error $\|\Delta W - T\|_F^2/\|T\|_F^2$: 1 at the start, 0 when perfect.
Why do we need it?
To see with our own eyes that training only a few numbers can work, and also when it cannot: the rank limits what LoRA can express. It turns the abstract claim into a loss curve you can watch.
Where is it used?
The same experiment is how practitioners choose r for real LoRA runs: train with several ranks, compare validation loss, and keep the smallest rank that is close to the best.
How is it used?
Train several ranks on the same data, plot the loss curves, and look for the rank where the curve stops improving. Pick that rank, since a larger one only adds memory.
- This is a toy: a clean least-squares target with a known rank. Real fine-tuning losses are noisy and not convex, but the same pattern appears: a modest rank usually gets close to full fine-tuning.
- The flat start of LoRA (before $B$ grows) is real. With a large learning rate it is short; with a tiny one it can look like nothing is happening.
- Full fine-tuning here is just a straight line toward $T$ because the toy loss is a perfect bowl. In a real model it can overfit more easily, and LoRA's low rank acts as a built-in restraint.
Quick check: the target has singular values $3$ and $2$. What is the best relative squared error a rank-1 update can reach?
$\sigma_2^2/(\sigma_1^2+\sigma_2^2) = 4/13\approx0.31$. That is the Eckart–Young bound: the part of the target that a single layer cannot capture.
LoRA, part 6: merging and swapping adapters core
What we need from earlier chapters: distributivity: $(W+M)\mathbf{x}=W\mathbf{x}+M\mathbf{x}$ (Chapter 1.4) and the LoRA forward pass (part 3).
After training, the correction is fixed. Then we can fold the side road into the main road: add the small change straight into the weights. The model then has the same shape and speed as the original, with no extra work at all.
Or we can keep the adapter separate. One big base model sits on disk once, and each task has its own tiny adapter file. To switch tasks, load a different adapter. That is why image-generation sites can offer thousands of "style" files for one base model.
Take $W = \begin{bmatrix}1&0\\0&1\end{bmatrix}$, $B = \begin{bmatrix}1\\2\end{bmatrix}$, $A = [3, 1]$, $\alpha/r = 1$, and $\mathbf{x} = [1, 1]$.
- Unmerged: $A\mathbf{x} = 3+1 = 4$; $B(A\mathbf{x}) = [4, 8]$; $W\mathbf{x} = [1,1]$; total $[5, 9]$.
- Merge: $BA = \begin{bmatrix}3&1\\6&2\end{bmatrix}$, so $W' = W + BA = \begin{bmatrix}4&1\\6&3\end{bmatrix}$.
- Merged: $W'\mathbf{x} = [4+1, 6+3] = [5, 9]$. Same answer.
- Merging is one matrix addition per layer, done once. Afterwards inference costs exactly as much as the original model.
- Un-merging: subtract $\frac\alpha rBA$ to get the base model back.
- Swapping: keep $W$ fixed and load a different pair $(A_k, B_k)$ for each task or style. Storage: one base model plus a small file for each adapter. Adding several adapters is also possible: $W + \sum_k s_k B_kA_k$ (the strength $s_k$ is the "weight" slider in image tools).
- If the base is stored in low precision (4-bit), merging requires care: the merged matrix must be re-quantised, which loses a little accuracy. Many systems keep the adapter separate instead.
Why do we need it?
We want the adapted model to be as fast as the original when we serve it, and we want to store many tasks without keeping many copies of the big model. Linearity makes both possible.
Where is it used?
Serving a fine-tuned chat model after merge_and_unload() in Hugging Face PEFT, style and character LoRAs for Stable Diffusion, and hosting one base model with a different adapter for every customer.
How is it used?
For one fixed task, merge once and ship the merged weights. To support many tasks, keep the base model loaded and switch the (A, B) pair per request. A strength slider simply scales the correction.
- Merging is exact in full precision. If the base weights are quantised, merging introduces a small rounding error.
- Merged weights cannot be "un-merged" unless you kept $A$ and $B$. Keep the adapter files.
- Mixing several adapters only works well if they were trained on the same base model.
Quick check: why is $(W + BA)\mathbf{x}$ equal to $W\mathbf{x} + B(A\mathbf{x})$?
Matrix multiplication distributes over addition: $(W+BA)\mathbf{x} = W\mathbf{x}+(BA)\mathbf{x}$. And matrix multiplication is associative: $(BA)\mathbf{x} = B(A\mathbf{x})$.
LoRA, part 7: where it is used, QLoRA and the limits
What we need from earlier chapters: everything from LoRA parts 1 to 6, plus the table of matrix-efficiency tricks from part 0.
Once adapting a model costs megabytes instead of a whole new model, many things become possible: a hobbyist can adapt a language model on one GPU; an app can serve a thousand customers from one base model; an artist can share a "style" as a small file that anyone can plug in.
And because the base weights are frozen, we can store them with fewer bits. We shrink the base model to 4-bit numbers (quantisation), lose a little precision, and still train 16-bit adapters on top. That combination is called QLoRA.
A 7-billion-parameter model.
- Full fine-tune: about $16\times7 = 112$ GB. Too big for one consumer GPU.
- LoRA (16-bit base, $r=16$ on $q,v$): $2\times7 = 14$ GB for the frozen base plus about $0.13$ GB of trainable memory: fits on a 24 GB GPU.
- QLoRA (4-bit base): $0.5\times7 = 3.5$ GB for the base. It fits on a 12 GB card.
- Ten tasks: ten full copies are 140 GB. One base plus ten adapters of about 17 MB each is about 14.2 GB.
- LoRA fine-tuning of language models: instruction tuning and chat tuning, domain adaptation (legal, medical, code), on a single GPU.
- Image generators (Stable Diffusion and similar): style, character and concept LoRAs, typically tens to a few hundred MB, applied to the attention layers of the base model with a strength setting.
- Speech and vision adapters: adapt a speech model to a new accent or a vision transformer to a new image domain.
- QLoRA = a frozen base in 4-bit (e.g. NF4) + trainable 16-bit LoRA matrices. Gradients flow through the dequantised base to the adapters.
- Limits: too low a rank underfits hard tasks; the method assumes the update is low rank; it is not a replacement for pre-training from scratch (a model learning everything from nothing needs full-rank changes); results can lag full fine-tuning on large shifts in domain.
Why do we need it?
To make adapting big models affordable and shareable. Without PEFT methods like LoRA, only organisations with large GPU clusters could customise a modern model, and every customisation would be a full copy.
Where is it used?
Chat and instruction tuning of open language models, character and style files for Stable Diffusion, per-customer adapters in AI products, and research labs trying many fine-tunes cheaply.
How is it used?
Load the base model (optionally in 4-bit), attach LoRA layers with a library such as Hugging Face PEFT, train for a few epochs, then save the small adapter file. Merge it for deployment, or load it on demand.
- LoRA adapters are tied to the base model they were trained on. An adapter for one model will not work on another.
- "Fine-tuning on one GPU" still needs enough memory for the base model's weights plus the activations. QLoRA's 4-bit base is what makes the largest models fit.
- If your adapted model is much worse than expected, test a higher rank, adapt more layers, or check the learning rate before concluding that LoRA cannot do the task.
Quick check: which two efficiency tricks does QLoRA combine?
Quantisation (the frozen base model is stored in 4-bit numbers) and low-rank factorisation (the trainable update is $BA$ with a small rank $r$).
Attention: queries, keys and values core
What we need from earlier chapters: the dot product as a similarity score (Chapter 1.2), matrix products (Chapter 1.4), softmax (earlier in this chapter), linear combinations (Chapter 1.2) and tensor shapes (Chapter 1.16).
In the sentence "The cat sat on the mat because it was tired", the word "it" needs to look back at "cat" to be understood. Attention lets every word look at every other word and decide whom to listen to.
Think of a library. Each word makes three things from its vector:
- a query: "what am I looking for?"
- a key: "what do I offer, as a label?"
- a value: "what is my actual content?"
A word compares its query with every key (a dot product: big means "relevant"), turns the scores into percentages (softmax), and then takes that percentage-mix of everyone's values. The result is a new vector for the word that is blended with its context.
Three tokens with 2-number vectors $X = \begin{bmatrix}1&0\\0&1\\1&1\end{bmatrix}$. To keep the arithmetic short, let $W_Q=W_K=W_V=I$, so $Q=K=V=X$. (Real models have learned matrices, and divide by $\sqrt d$; we add that in the next section.) Look at token 1.
- Scores: $\mathbf{q}_1\cdot\mathbf{k}_j = [1,0]\cdot[1,0],\ [0,1],\ [1,1] = [1, 0, 1]$.
- Softmax: $e^1=2.718$, $e^0=1$, $e^1=2.718$; total $6.437$; weights $= [0.422, 0.155, 0.422]$.
- Output: $0.422\cdot[1,0] + 0.155\cdot[0,1] + 0.422\cdot[1,1] = [0.844, 0.577]$.
Token 1 listened mostly to itself and to token 3 (they share its first feature), and little to token 2.
Stack the $n$ token vectors as rows of $X\in\mathbb{R}^{n\times d_{\text{model}}}$. With learned matrices $W_Q,W_K\in\mathbb{R}^{d_{\text{model}}\times d_k}$ and $W_V\in\mathbb{R}^{d_{\text{model}}\times d_v}$:
$$Q = XW_Q,\quad K = XW_K,\quad V = XW_V,$$ $$S = \frac{QK^\top}{\sqrt{d_k}},\qquad A = \operatorname{softmax}_{\text{rows}}(S),\qquad \text{Attention}(Q,K,V) = A\,V.$$- $S_{ij} = \mathbf{q}_i\cdot\mathbf{k}_j/\sqrt{d_k}$: all pairwise dot products in one matrix product, an $n\times n$ table.
- Each row of $A$ is non-negative and adds up to 1: a set of "how much do I listen to each token" percentages.
- $AV$: row $i$ of the output is a weighted average of the rows of $V$ (a linear combination with weights $A_{i,:}$).
- For text generation, a causal mask sets $S_{ij}=-\infty$ for $j>i$ so a token cannot look at the future; after softmax those weights are exactly 0.
Why do we need it?
The meaning of a word depends on the words around it. Attention lets every token look at all the others and blend in the ones that are relevant.
Where is it used?
Transformers: GPT-style chat models, BERT, machine translation, vision transformers, and protein-structure and speech models.
How is it used?
Make Q, K and V with three matrix products, compute the scores QKᵀ/√d, softmax each row, and multiply by V. Look at the weight matrix A to see which token is listening to which.
- Attention has no built-in sense of order: swap two tokens and the same numbers come out in swapped rows. Models add position information to $X$ (positional encodings).
- Query and key must have the same length $d_k$ (to take their dot product). The value can have a different length $d_v$.
- Weights add to 1 along each row (over the tokens being looked at), not down columns.
Quick check: $A$ is $4\times4$ and $V$ is $4\times8$. What shape is $AV$, and what does its first row mean?
$4\times8$: one 8-number vector per token. The first row is a weighted average of the 4 value vectors, using the first row of $A$ as the weights.
Why attention divides by $\sqrt{d}$ core
What we need from earlier chapters: the dot product as a sum of products (Chapter 1.2), softmax (earlier in this chapter) and Jacobians (Chapter 1.14). We use again that the variance of a sum of independent numbers is the sum of their variances.
A dot product adds up $d$ small products. Add more terms, and the total wanders further from zero. So with long vectors, the scores $\mathbf{q}\cdot\mathbf{k}$ become big numbers, positive or negative, just because $d$ is large, not because anything is more relevant.
Softmax then behaves badly. It exponentiates, so a score gap of 8 means a ratio of about 3000 to 1. Nearly all the weight goes to one token (almost a hard choice) and the others get almost nothing. In that state the softmax is flat in every direction, so gradients vanish and learning stalls.
The cure is a simple rescale: divide the scores by $\sqrt{d}$, which cancels the growth exactly.
Suppose the entries of $\mathbf{q}$ and $\mathbf{k}$ are independent with mean 0 and variance 1.
- Each product $q_ik_i$ has mean 0 and variance $1\cdot1 = 1$.
- The dot product is a sum of $d$ independent such terms, so its variance is $d$ and its standard deviation is $\sqrt d$.
- For $d = 64$ a typical score is about $\pm8$. Dividing by $\sqrt{64}=8$ brings it back to about $\pm1$.
Effect on softmax with scores $[4, 0, -4]$ (as if $d=16$): without scaling the weights are $[0.982, 0.018, 0.0003]$, nearly one-hot. With scaling by $\sqrt{16}=4$ the scores are $[1,0,-1]$ and the weights are $[0.665, 0.245, 0.090]$, still soft.
If the entries of $\mathbf{q},\mathbf{k}\in\mathbb{R}^{d}$ are independent, mean 0, variance 1, then
$$\operatorname{Var}(\mathbf{q}\cdot\mathbf{k}) = d, \qquad \operatorname{Var}\!\Big(\frac{\mathbf{q}\cdot\mathbf{k}}{\sqrt d}\Big) = \frac{d}{d} = 1.$$The softmax Jacobian is $\operatorname{diag}(\mathbf{p}) - \mathbf{p}\mathbf{p}^\top$. When $\mathbf{p}$ is nearly one-hot, every entry of it is close to 0: no gradient flows. Scaling keeps the scores in the range where the softmax still has a useful slope.
Why do we need it?
Dot products of long vectors are large, which makes the softmax almost one-hot and kills its gradients. Dividing by √d keeps the scores a steady size, whatever the vector length.
Where is it used?
Scaled dot-product attention in every Transformer. The same variance-keeping idea is behind the weight-initialisation rules.
How is it used?
Divide QKᵀ by √d_k before the softmax, where d_k is the key length. If training stalls with a very large head size, check that this line is present.
- The argument assumes independent, unit-variance entries. Trained vectors are not exactly like that, but the $\sqrt d$ rule is a good, cheap default and it works in practice.
- The scaling uses $d_k$ (the key length), not the model width, and not the number of tokens.
- Scaling is not "just a constant". Without it, a model with large $d_k$ often fails to train.
Quick check: $d_k = 256$. By what do we divide the scores?
$\sqrt{256} = 16$. Without it the raw scores would have a typical size around 16, so the softmax would be very sharp.
Multi-head attention and the $O(n^2d)$ cost core
What we need from earlier chapters: reshape, transpose and tensor shapes (Chapter 1.16), matrix multiplication and its cost (Chapter 1.4) and counting FLOPs (Chapter 1.15).
One attention pattern can only look for one kind of relationship at a time. But language has many at once: which noun a pronoun refers to, which word comes just before, which words share a topic. So we run several attentions in parallel, called heads, each with its own small queries, keys and values. Then we glue their answers together.
No loop over heads is needed. We reshape the long vectors into $h$ shorter ones and use one batched matrix multiplication (next section) that handles all heads at once.
The price: every token looks at every token, so the score table has $n\times n$ entries. Double the text length and the cost goes up four times.
$n=4$ tokens, model width $d=8$, $h=2$ heads, so each head has $d_h = d/h = 4$.
- $Q$ has shape $(4, 8)$. Reshape to $(4, 2, 4)$ (split the 8 numbers into 2 groups of 4), then swap axes to $(2, 4, 4)$: head, token, feature.
- $QK^\top$ per head: $(2,4,4)\times(2,4,4)^\top\to(2,4,4)$. That is two $4\times4$ score tables, from one batched call.
- Cost in multiply–adds: per head $4\cdot4\cdot4 = 64$, times 2 heads $=128$. And $n^2d = 16\cdot8 = 128$. Same, however many heads we use.
- After attention, swap back to $(4,2,4)$, reshape to $(4,8)$, and multiply by an output matrix $W_O$ to mix the heads.
- Shapes: $X:(n,d)\to Q,K,V:(h,n,d_h)\to\text{scores}:(h,n,n)\to\text{output}:(n,d)$.
- Cost (a multiply–add counted as 2 FLOPs): scores $QK^\top$ and weighted sum $AV$ each take $2n^2d$, the four projections ($W_Q,W_K,W_V,W_O$) take $8nd^2$. Attention part: $O(n^2d)$; projections: $O(nd^2)$.
- Memory: the score tables hold $h\cdot n^2$ numbers. For long inputs this, not the arithmetic, is often the limit. (FlashAttention avoids storing it.)
- The share of time in the $n^2$ part is $\dfrac{n}{n+2d}$: short texts are dominated by the projections, long texts by attention.
Why do we need it?
One attention pattern can find only one kind of relation. Several heads look for different patterns at once, and a reshape lets us compute them all in one batched step. We also need the cost, because it grows with n².
Where is it used?
Every Transformer. The n² memory is why long-context models need special kernels such as FlashAttention.
How is it used?
Reshape Q, K and V from (n, d) to (h, n, d/h), run one batched attention, reshape back and multiply by W_O. To plan hardware, estimate 4n²d FLOPs and h·n² score numbers per layer.
- More heads does not mean more compute: the total width $d$ is split between them, so each head is thinner.
- The $n\times n$ table is the real bottleneck for long documents. That is why efficient variants (FlashAttention, sparse and linear attention) exist. They change how $QK^\top$ and the softmax are computed, but not the mathematical idea.
- "$O(n^2d)$" hides constants: the FLOP counts above are the ones used for planning real hardware.
Quick check: $n$ goes from 1,000 to 4,000. By what factor does the $QK^\top$ cost grow?
It scales with $n^2$: $(4000/1000)^2 = 16$ times more.
Batch matrix multiplication and GPU efficiency core
What we need from earlier chapters: tensors, shapes, broadcasting and einsum (Chapter 1.16), cost counting and memory versus compute (Chapter 1.15) and the matrix product (Chapter 1.4).
A GPU is a factory with thousands of workers. One big job keeps them all busy. A thousand tiny jobs, each started separately, mostly leaves the workers standing around, waiting for the next order to be sent. Each start-up has a small fixed cost, like a trip to the shop.
So instead of multiplying many small matrix pairs one after the other, we stack them into a 3D array (a "batch", with pages like a book) and ask for all the products in one go. This is batch matrix multiplication. One trip, a full factory.
Batching examples also helps for a second reason: every example uses the same weight matrix, so a matrix–vector product per example becomes one matrix–matrix product, and each weight loaded from memory gets reused many times.
Three pairs of $2\times2$ matrices, stacked into arrays of shape $(3,2,2)$. Page by page:
- Page 0: $\begin{bmatrix}1&2\\3&4\end{bmatrix}\begin{bmatrix}1&0\\0&1\end{bmatrix} = \begin{bmatrix}1&2\\3&4\end{bmatrix}$ (times the identity changes nothing).
- Page 1: $\begin{bmatrix}0&1\\1&0\end{bmatrix}\begin{bmatrix}5&6\\7&8\end{bmatrix} = \begin{bmatrix}7&8\\5&6\end{bmatrix}$ (the left matrix swaps the rows).
- Page 2: $\begin{bmatrix}2&0\\0&2\end{bmatrix}\begin{bmatrix}1&1\\1&1\end{bmatrix} = \begin{bmatrix}2&2\\2&2\end{bmatrix}$ (doubling).
The result is another $(3,2,2)$ array. Each page is its own independent product.
For arrays $A$ of shape $(B, m, k)$ and $C$ of shape $(B, k, n)$:
$$(A\,@\,C)[b] = A[b]\,C[b]\quad\Longrightarrow\quad \text{shape }(B,m,n),\qquad \texttt{einsum('bij,bjk->bik')}.$$- Any extra leading dimensions work the same way: $(B,h,n,d_h)\,@\,(B,h,d_h,n)\to(B,h,n,n)$ is exactly the attention score computation for a batch of $B$ sequences and $h$ heads.
- One matrix can be broadcast over a batch: a dense layer maps $(B,n,d_{\text{in}})\,@\,W^\top\,(d_{\text{in}},d_{\text{out}})\to(B,n,d_{\text{out}})$ with the same $W$ for every page.
- Names:
np.matmul/@,torch.bmm,torch.matmul. - Why it is fast: (1) one launch instead of $B$; (2) big enough work to fill the thousands of cores; (3) data reuse: a matrix–matrix product does about $n$ multiplications for each number it reads, far more than a matrix–vector product; (4) GPUs have special matrix hardware (tensor cores) that likes large, regular shapes.
Why do we need it?
A GPU is fast only when it is given big jobs. Many small matrix products, one after another, waste it. Stacking them into a 3D array lets a single call do them all.
Where is it used?
Attention heads, batches of sequences during training, convolutions turned into matrix products, and any code that calls torch.bmm, torch.matmul or einsum.
How is it used?
Stack the inputs into shapes (B, m, k) and (B, k, n), call A @ C or torch.bmm, and get (B, m, n). Choose the biggest batch that fits in memory, and pad sequences to the same length.
- All pages in a batch must have the same shape. Sequences of different lengths are padded to a common length (and masked), which wastes some work.
- Batching does not change the answer, only the speed. Page $b$ of the result depends only on page $b$ of the inputs.
- Bigger batches use more memory (all activations are kept for backprop). You pick the largest size that fits.
Quick check: $(32, 12, 128, 64)\ @\ (32, 12, 64, 128)$ has what shape?
The leading dimensions $(32,12)$ are batch dimensions (32 sequences, 12 heads). Each page multiplies $128\times64$ by $64\times128$, giving $128\times128$. Result: $(32, 12, 128, 128)$, the attention score tables.
Other connections (1/5): spectral clustering
What we need from earlier chapters: eigenvectors of symmetric matrices (Chapter 1.11), PSD matrices and the quadratic form (Chapter 1.12) and adjacency matrices (Chapter 1.4).
A graph is a set of dots (people) joined by lines (friendships). Suppose there are two friend groups joined by only a few weak links. Cutting those weak links separates the groups. How can a computer find the cut?
Give every person one number so that friends get similar numbers. Then the two groups get two different numbers, and the sign of the number tells you the group. The matrix whose eigenvectors give "the smoothest possible numbers" is the graph Laplacian.
Four people in a line: 1–2 (strength 1), 2–3 (strength $0.1$, weak), 3–4 (strength 1). Try the labelling $\mathbf{x} = [1, 1, -1, -1]$. The "roughness" is the sum over links of strength × (difference)$^2$:
$$\mathbf{x}^\top L\mathbf{x} = 1\cdot(1-1)^2 + 0.1\cdot(1-(-1))^2 + 1\cdot(-1-(-1))^2 = 0.4.$$Divided by $\|\mathbf{x}\|^2=4$ it is $0.1$: very smooth. The only rough link is the weak one we want to cut. The all-ones labelling has roughness $0$ (it ignores the groups), so we ask for the next smoothest labelling, which is orthogonal to it.
For a graph with weights $A_{ij}\ge0$ and degrees $D_{ii} = \sum_j A_{ij}$, the graph Laplacian is $L = D - A$.
$$\mathbf{x}^\top L\mathbf{x} = \sum_{\text{links }(i,j)} A_{ij}\,(x_i-x_j)^2 \ \ge 0.$$- $L$ is symmetric and PSD, so it has real eigenvalues $0=\lambda_1\le\lambda_2\le\dots$ and orthogonal eigenvectors.
- The eigenvector for $\lambda_1=0$ is all ones. The eigenvector for $\lambda_2$ (the Fiedler vector) is the smoothest non-trivial labelling. Splitting nodes by its sign gives a good 2-cluster cut.
- For $k$ clusters: take the $k$ smallest eigenvectors, treat each node's row of them as a short embedding, and run $k$-means on those rows.
- The number of zero eigenvalues equals the number of disconnected pieces.
Why do we need it?
Finding the weakest links in a graph by trying every possible cut is impossibly slow. The eigenvectors of the Laplacian give a good cut from one matrix calculation.
Where is it used?
Community detection in social networks, image segmentation, clustering data that is not round, and graph embeddings.
How is it used?
Build a similarity graph, form L = D − A, compute its smallest few eigenvectors (np.linalg.eigh), and run k-means on their rows. For two groups, the sign of the Fiedler vector is enough.
Quick check: a graph has 3 separate connected pieces. How many zero eigenvalues does $L$ have?
Three: one all-ones-on-that-piece labelling per piece has zero roughness.
Other connections (2/5): kernel methods and Gram matrices
What we need from earlier chapters: the dot product (Chapter 1.2), PSD matrices (Chapter 1.12) and $XX^\top$ as a table of dot products (Chapter 1.4).
A flat boundary cannot separate one ring of points sitting inside another. One fix: send every point into a bigger space where the classes can be separated by a flat boundary. The trick of kernel methods is that you never need to go there. All the algorithm ever uses is the dot products between pairs of points, and a kernel function $k(\mathbf{x},\mathbf{x}')$ computes that dot product in the big space directly, from the original points.
Collect all pairwise similarities in a table, the Gram matrix. Close points score near 1, far points near 0.
Three points on a line: $0$, $1$ and $3$. Use the RBF kernel $k(x,x') = e^{-\gamma(x-x')^2}$ with $\gamma=1$.
- $k(0,1) = e^{-1} = 0.368$, $\ k(1,3) = e^{-4}=0.018$, $\ k(0,3) = e^{-9} \approx 0.0001$, and $k(x,x) = e^0 = 1$.
Symmetric, ones on the diagonal, and the near neighbours (0 and 1) are the most similar pair.
Given points $\mathbf{x}_1,\dots,\mathbf{x}_n$ and a kernel $k$, the Gram matrix is $G_{ij} = k(\mathbf{x}_i,\mathbf{x}_j)$. For the plain dot product $k(\mathbf{x},\mathbf{x}')=\mathbf{x}\cdot\mathbf{x}'$ it is $G = XX^\top$.
- A function $k$ is a valid kernel exactly when every Gram matrix it produces is symmetric and positive semi-definite (Mercer's condition). Then $G=\Phi\Phi^\top$ for some (possibly infinite) feature matrix $\Phi$.
- Popular kernels: linear $\mathbf{x}\cdot\mathbf{x}'$; polynomial $(\mathbf{x}\cdot\mathbf{x}'+1)^p$; RBF $e^{-\gamma\|\mathbf{x}-\mathbf{x}'\|^2}$.
- Kernel ridge regression: instead of solving for $d$ weights, solve $(G+\lambda I)\boldsymbol\alpha=\mathbf{y}$ for $n$ coefficients; predict $f(\mathbf{x}) = \sum_i\alpha_i\,k(\mathbf{x}_i,\mathbf{x})$. The same linear-system solve as ridge regression.
Why do we need it?
A flat boundary cannot separate tangled classes. Kernels let an algorithm act as if the points lived in a much bigger space, while it only ever computes similarities between pairs of points.
Where is it used?
Support vector machines, kernel ridge regression, Gaussian processes, and kernel PCA.
How is it used?
Pick a kernel (RBF is the usual start), build the Gram matrix G of all pairs, add λI and solve for the coefficients. Tune γ and λ on held-out data. The cost grows with the square of the number of points.
Quick check: why must the diagonal of an RBF Gram matrix be all ones?
Each point is at distance $0$ from itself, and $e^{-\gamma\cdot0}=1$. Every point is perfectly similar to itself.
Other connections (3/5): Gaussian processes and the Cholesky factor
What we need from earlier chapters: the Cholesky decomposition (Chapter 1.13), the multivariate Gaussian (earlier in this chapter) and matrix–vector products (Chapter 1.4).
A random function is just a long list of numbers: its height at many points along the $x$-axis. A Gaussian process is a recipe for random functions in which neighbouring heights are strongly related (a smooth curve does not jump), and distant ones are almost unrelated. The relatedness comes from a kernel (previous section), collected in a covariance matrix $K$.
How do you draw a random list whose covariance is $K$? Draw independent random numbers, then mix them with a matrix. The right mixing matrix is the Cholesky factor $L$ with $K=LL^\top$.
Two points whose heights have variance 1 and correlation $0.8$: $K = \begin{bmatrix}1&0.8\\0.8&1\end{bmatrix}$.
- Cholesky: $L = \begin{bmatrix}1&0\\0.8&0.6\end{bmatrix}$, because $0.6 = \sqrt{1-0.8^2}$. Check: $LL^\top = \begin{bmatrix}1&0.8\\0.8&0.64+0.36\end{bmatrix} = K$ ✓.
- Draw independent numbers $\mathbf{z} = [1, -1]$.
- Mix: $\mathbf{f} = L\mathbf{z} = [1,\ 0.8\cdot1+0.6\cdot(-1)] = [1, 0.2]$.
The second value is pulled toward the first, as the correlation demands.
Choose points $x_1,\dots,x_n$ and a kernel, for example $K_{ij} = \exp\!\big(-(x_i-x_j)^2/(2\ell^2)\big)$ with length-scale $\ell$. Then:
$$K = LL^\top,\qquad \mathbf{z}\sim\mathcal{N}(\mathbf{0}, I),\qquad \mathbf{f} = L\mathbf{z}\ \Longrightarrow\ \operatorname{Cov}(\mathbf{f}) = L\,I\,L^\top = K.$$- $K$ must be symmetric and positive definite (Chapter 1.12) for the Cholesky factorisation to exist. Smooth kernels give nearly singular $K$, so a tiny "jitter" $\epsilon I$ is added.
- Prediction (regression with a GP) conditions on observed data: the mean is $K_*^\top(K+\sigma^2I)^{-1}\mathbf{y}$, computed by Cholesky solves, never an explicit inverse. This costs $O(n^3)$.
Why do we need it?
Sometimes we want predictions with honest error bars from very little data. A Gaussian process says how related nearby points are, and the Cholesky factor lets us sample from it and predict with it.
Where is it used?
Bayesian optimisation (tuning hyper-parameters), small-data science and engineering models, geostatistics, and the reparameterisation trick in variational autoencoders.
How is it used?
Build K from a kernel, add a small jitter, compute L = chol(K), sample with f = Lz, and predict with Cholesky solves rather than inverses. The cost grows with the cube of the number of data points.
Quick check: why do we add a tiny jitter $\epsilon I$ before the Cholesky factorisation?
Smooth kernels make $K$ almost singular (some eigenvalues are essentially $0$), and rounding can push one slightly negative, which makes Cholesky fail. Adding $\epsilon I$ lifts every eigenvalue by $\epsilon$.
Other connections (4/5): Markov chains, PageRank and linear dynamics
What we need from earlier chapters: matrix powers (Chapter 1.4) and eigenvalues and eigenvectors, and what repeated multiplication does (Chapter 1.11).
A web surfer clicks a random link, over and over. At any moment we only know the chance of being on each page. One click turns today's chances into tomorrow's with one matrix–vector product. After many clicks the chances settle down and stop changing. That settled state is the steady state, and it is an eigenvector with eigenvalue 1.
The same idea works for any system $\mathbf{x}_{t+1} = A\mathbf{x}_t$: repeated multiplication is dominated by the biggest eigenvalue. Eigenvalues below 1 in size fade, above 1 explode.
Weather: after a sunny day, 90% sunny and 10% rainy. After a rainy day, 50% and 50%. Each column of $P$ lists the chances for tomorrow:
$$P = \begin{bmatrix}0.9&0.5\\0.1&0.5\end{bmatrix},\qquad \mathbf{p}_0 = \begin{bmatrix}1\\0\end{bmatrix},\quad \mathbf{p}_1 = P\mathbf{p}_0 = \begin{bmatrix}0.9\\0.1\end{bmatrix},\quad \mathbf{p}_2 = \begin{bmatrix}0.81+0.05\\0.09+0.05\end{bmatrix} = \begin{bmatrix}0.86\\0.14\end{bmatrix}.$$The steady state is $\boldsymbol\pi=[5/6, 1/6]$: check $P\boldsymbol\pi = [0.9\cdot\tfrac56+0.5\cdot\tfrac16,\ \dots] = [\tfrac{4.5+0.5}{6}, \dots] = [\tfrac56,\tfrac16]$ ✓. The eigenvalues of $P$ are $1$ (trace $1.4$ minus $0.4$) and $0.4$, so the distance to the steady state shrinks by about a factor of $0.4$ per day.
A column-stochastic matrix $P$ has non-negative entries and every column adds to 1. Then $\mathbf{p}_{t+1} = P\mathbf{p}_t$ keeps the probabilities adding to 1.
- $1$ is always an eigenvalue and no eigenvalue is bigger in size. The steady state solves $P\boldsymbol\pi=\boldsymbol\pi$ (the eigenvector for $\lambda=1$, scaled to add to 1).
- Convergence speed is set by the second-largest eigenvalue $|\lambda_2|$: the error shrinks like $|\lambda_2|^t$. The power method is exactly this repeated multiplication.
- PageRank adds "teleporting": with probability $1-d$ the surfer jumps to a random page, so $G = dP + \frac{1-d}{n}\mathbf{1}\mathbf{1}^\top$. A page's importance is its entry in the steady state of $G$ (typical $d=0.85$).
- Linear dynamics $\mathbf{x}_{t+1}=A\mathbf{x}_t$: $\mathbf{x}_t = A^t\mathbf{x}_0$ and the long-run behaviour follows the eigenvalues (stable if all $|\lambda|<1$).
Why do we need it?
Many systems move step by step with fixed chances, and we want to know where they end up. The answer is an eigenvector, and repeated multiplication finds it.
Where is it used?
Google's PageRank, reinforcement learning, MCMC sampling, hidden Markov models, queues and population models, and the stability of recurrent networks.
How is it used?
Write down the transition matrix. Multiply a probability vector by it again and again (the power method) until it stops changing, or solve Pπ = π. The second-largest eigenvalue tells you how fast it settles.
Quick check: a chain has $|\lambda_2|=0.9$. Does it settle faster or slower than one with $|\lambda_2|=0.4$?
Slower. The error shrinks by a factor of $0.9$ per step instead of $0.4$. After 10 steps: $0.9^{10}\approx0.35$ versus $0.4^{10}\approx0.0001$.
Other connections (5/5): graph neural networks
What we need from earlier chapters: the adjacency matrix and matrix products as sums over neighbours (Chapter 1.4), the dense layer (earlier in this chapter) and the repeated multiplication story from the previous section.
Many things are graphs: molecules (atoms and bonds), social networks, road maps. Each node carries a vector of features. A graph neural network layer lets every node collect its neighbours' features and average them with its own. After one layer a node knows about its neighbours; after two layers, about neighbours of neighbours; and so on.
Doing that for all nodes at once is a single matrix product: multiply the feature matrix by the (normalised) adjacency matrix. Then apply a weight matrix and a non-linearity, just like a dense layer.
Three nodes in a line, 1–2–3, with a single number each, $H = [1, 0, 0]^\top$ (only node 1 "lit"). Add self-loops: $A+I = \begin{bmatrix}1&1&0\\1&1&1\\0&1&1\end{bmatrix}$, row sums $2,3,2$. Dividing each row by its sum gives
$$\hat A = \begin{bmatrix}\tfrac12&\tfrac12&0\\\tfrac13&\tfrac13&\tfrac13\\0&\tfrac12&\tfrac12\end{bmatrix},\qquad \hat AH = \begin{bmatrix}0.5\\0.333\\0\end{bmatrix},\qquad \hat A(\hat AH) = \begin{bmatrix}0.417\\0.278\\0.167\end{bmatrix}.$$The light has spread out from node 1; after two steps it has even reached node 3.
One graph-convolution layer (Kipf and Welling):
$$H^{(l+1)} = \varphi\big(\hat A\,H^{(l)}\,W^{(l)}\big),\qquad \hat A = \tilde D^{-1/2}(A+I)\tilde D^{-1/2},$$where $H^{(l)}\in\mathbb{R}^{n\times d}$ has one row of features per node, $\tilde D$ is the degree matrix of $A+I$, and $W^{(l)}$ is a learned weight matrix shared by all nodes. (The simpler $D^{-1}(A+I)$, as in the example, is "mean of neighbours".)
- $\hat AH$ replaces each node's row by a weighted average of its neighbours' rows.
- $k$ layers $\Rightarrow$ information from $k$ hops away. $\hat A^k$ is the $k$-step spreading matrix, the same repeated multiplication as a Markov chain.
- Over-smoothing: after too many layers, $\hat A^k$ is dominated by its top eigenvector and all nodes end up with nearly the same features.
- $A$ is usually sparse, so real libraries use sparse matrix products.
Why do we need it?
Molecules, social networks and road maps are not grids. We need a layer in which each node gathers information from its neighbours, for a graph of any shape.
Where is it used?
Predicting properties of molecules and drugs, traffic forecasting, fraud detection, recommendation on graphs, and knowledge graphs.
How is it used?
Build the normalised adjacency matrix  (with self-loops), compute ÂHW for the node-feature matrix H, apply a non-linearity, and stack two or three layers. Libraries such as PyTorch Geometric use sparse products.
Quick check: a GNN has 3 layers. How many hops away can a node's output "see"?
Three hops: each layer multiplies by $\hat A$ once, extending the reach by one step along the links.
Recap, cheat sheet and practice
- Linear regression is $\hat{\mathbf{y}} = X\mathbf{w}$. The best $\mathbf{w}$ solves the normal equations $X^\top X\mathbf{w}=X^\top\mathbf{y}$, which says "the residual is perpendicular to the column space": a projection. Gradient descent reaches the same answer; ridge adds $\lambda I$, lasso adds an L1 penalty that creates zeros.
- Logistic and softmax regression are a linear score plus a squashing function. The decision boundary is a hyperplane with normal $\mathbf{w}$, and the loss has a PSD Hessian, so it is convex.
- The covariance matrix $\Sigma=\frac1nX_c^\top X_c$ is symmetric PSD. Mahalanobis distance and the Gaussian use $\Sigma^{-1}$.
- PCA rotates onto the eigenvectors of $\Sigma$ (or the right singular vectors of $X_c$); $\lambda_i=s_i^2/n$; keep the top $k$ to reduce, project back to reconstruct. Whitening also rescales.
- Embeddings are rows of a matrix (lookup = one-hot product); compare with cosine. Recommenders assume a low-rank table $R\approx UV^\top$ and fit it with SVD or ALS.
- Network layers are $XW^\top+\mathbf{b}$ plus a non-linearity; initialisation keeps norms stable; backprop multiplies by Jacobian transposes; LayerNorm centres and scales.
- Matrix efficiency has a name: low-rank factorisation ($M\approx BA$, $2dr$ numbers instead of $d^2$). Applied to fine-tuning it is LoRA (Low-Rank Adaptation, a kind of parameter-efficient fine-tuning, PEFT): freeze $W$, learn $\Delta W=BA$ with $B=0$ at the start, merge afterwards with $W'=W+\frac\alpha rBA$. The other tricks are quantisation (QLoRA), sparsity, distillation and structured matrices.
- Attention is $\operatorname{softmax}(QK^\top/\sqrt{d})V$; heads are a reshape plus a batched matmul; the cost is $O(n^2d)$.
- Spectral clustering, kernels, Gaussian processes, Markov chains and graph networks all reuse eigenvectors, Gram matrices, Cholesky factors and matrix powers.
Cheat sheet
| Topic | Key formulas | Linear algebra inside |
|---|---|---|
| Linear regression | $\hat{\mathbf{y}}=X\mathbf{w}$, $\ \mathbf{w}=(X^\top X)^{-1}X^\top\mathbf{y}$, $\ \nabla L=\frac1nX^\top(X\mathbf{w}-\mathbf{y})$ | projection onto $C(X)$, pseudoinverse |
| Ridge / lasso | $(X^\top X+\lambda I)^{-1}X^\top\mathbf{y}$ / $\lambda\|\mathbf{w}\|_1$ | PSD $+\lambda I$ is invertible; L1 ball has corners |
| Logistic regression | $p=\sigma(\mathbf{w}^\top\mathbf{x}+b)$, $\ \nabla L=\frac1nA^\top(\mathbf{p}-\mathbf{y})$, $\ H=\frac1nA^\top SA$ | hyperplane, PSD Hessian |
| Softmax | $p_k=e^{z_k}/\sum_je^{z_j}$, $\ \partial L/\partial\mathbf{z}=\mathbf{p}-\mathbf{y}$ | $W\mathbf{x}$, outer-product gradient |
| Covariance | $\Sigma=\frac1nX_c^\top X_c$, $\ \rho_{jk}=\Sigma_{jk}/\sigma_j\sigma_k$ | symmetric PSD |
| Mahalanobis / Gaussian | $d_M^2=(\mathbf{x}-\boldsymbol\mu)^\top\Sigma^{-1}(\mathbf{x}-\boldsymbol\mu)$ | quadratic form, whitening |
| PCA | $X_c=USV^\top$, $\ \lambda_i=s_i^2/n$, $\ Z=X_cV_k$, $\ \hat X=ZV_k^\top$ | SVD, orthogonal projection |
| Embeddings | $\mathbf{x}_i=\mathbf{e}_i^\top E$, $\ \cos=\bar E\bar{\mathbf{q}}$ | matrix rows, normalised dot products |
| Matrix factorisation | $R\approx UV^\top$; ALS: $\mathbf{u}_i=(V_\Omega^\top V_\Omega+\lambda I)^{-1}V_\Omega^\top\mathbf{r}_\Omega$ | low rank, ridge least squares |
| Dense layer | $Y=XW^\top+\mathbf{b}$, $\ H=\varphi(Y)$ | matmul; composition of linear maps is linear |
| Initialisation | Var $=1/n$ (Xavier), $2/n$ (He); orthogonal $Q$ | norm preservation |
| Backprop | $g_{\mathbf{x}}=J^\top g_{\mathbf{y}}$, $\ g_W=g_{\mathbf{y}}\mathbf{x}^\top$ | Jacobian transposes |
| Low-rank factorisation / LoRA / PEFT | $M\approx BA$, params $2dr$ not $d^2$; $\ h=W\mathbf{x}+\frac\alpha rB(A\mathbf{x})$; merge $W'=W+\frac\alpha rBA$; $B=0$ at the start | rank, outer products, SVD and Eckart–Young |
| Other efficiency tricks | quantisation (fewer bits), sparsity (zeros), distillation (small student), structured matrices (Kronecker, convolution) | each trades a little accuracy for big savings |
| Attention | $\operatorname{softmax}(QK^\top/\sqrt{d_k})V$, cost $4n^2d$ | pairwise dot products, weighted average |
| Batch matmul | $(B,m,k)@(B,k,n)\to(B,m,n)$ | tensors; one big call |
| Spectral / kernels / GP / Markov / GNN | $L=D-A$; $G_{ij}=k(\mathbf{x}_i,\mathbf{x}_j)$; $\mathbf{f}=L\mathbf{z}$; $P\boldsymbol\pi=\boldsymbol\pi$; $\hat AHW$ | eigenvectors, PSD, Cholesky, matrix powers |
import numpy as np
rng = np.random.default_rng(0)
# ---------- 1. Linear regression: closed form, gradient descent, ridge ----------
X = np.c_[np.ones(100), rng.normal(size=(100, 3))] # bias column + 3 features
w_true = np.array([1.0, 2.0, -1.0, 0.5])
y = X @ w_true + 0.1 * rng.normal(size=100)
w_cf = np.linalg.lstsq(X, y, rcond=None)[0] # closed form (SVD inside, no inverse)
w_gd = np.zeros(4)
for _ in range(500): # gradient descent, mean loss
w_gd -= 0.1 * X.T @ (X @ w_gd - y) / len(y)
lam = 1.0
w_ridge = np.linalg.solve(X.T @ X + lam * np.eye(4), X.T @ y) # (X^T X + lambda I)^-1 X^T y
print(np.allclose(w_cf, w_gd, atol=1e-3), w_ridge.round(2))
# ---------- 2. Logistic regression: vectorised gradient and Newton ----------
Xc = rng.normal(size=(200, 2))
yc = (Xc @ np.array([2.0, -1.0]) + 0.3 + rng.normal(size=200) > 0).astype(float)
A = np.c_[Xc, np.ones(200)]
sigmoid = lambda z: 1 / (1 + np.exp(-z))
th = np.zeros(3)
for _ in range(10): # Newton: solve H step = gradient
p = sigmoid(A @ th)
g = A.T @ (p - yc) / len(yc)
H = A.T @ (A * (p * (1 - p))[:, None]) / len(yc) # A^T S A / n (PSD)
th -= np.linalg.solve(H + 1e-8 * np.eye(3), g)
print(th.round(2), np.linalg.eigvalsh(H).min() >= -1e-12)
# ---------- 3. PCA from scratch via SVD ----------
Z = rng.normal(size=(300, 3)) @ np.diag([3.0, 1.0, 0.3]) @ rng.normal(size=(3, 3))
Zc = Z - Z.mean(axis=0) # centre first!
U, s, Vt = np.linalg.svd(Zc, full_matrices=False)
var = s**2 / len(Zc) # eigenvalues of the covariance
evr = var / var.sum() # explained variance ratio
k = 2
scores = Zc @ Vt[:k].T # reduce: 3 numbers -> 2
recon = scores @ Vt[:k] + Z.mean(axis=0) # reconstruct
print(evr.round(3), np.mean(np.sum((Z - recon) ** 2, axis=1)), var[k:].sum())
Sigma = Zc.T @ Zc / len(Zc)
print(np.allclose(np.sort(np.linalg.eigvalsh(Sigma))[::-1], var))
# from sklearn.decomposition import PCA
# PCA(2).fit(Z).explained_variance_ratio_ -> same ratios (directions may differ by a sign)
# ---------- 4. Mahalanobis distance ----------
d = Z[0] - Z.mean(axis=0)
m = np.sqrt(d @ np.linalg.solve(Sigma, d)) # solve, do not invert
print(m)
# ---------- 5. Embeddings: lookup and cosine neighbours ----------
E = rng.normal(size=(5, 3)) # 5 tokens, 3 dimensions
ids = np.array([1, 3, 1])
one_hot = np.eye(5)[ids]
print(np.allclose(one_hot @ E, E[ids])) # one-hot product == row lookup
En = E / np.linalg.norm(E, axis=1, keepdims=True)
sims = En @ En[1] # all cosines in one product
print(np.argsort(-sims))
# ---------- 6. Single-head and multi-head self-attention ----------
def softmax(S):
S = S - S.max(axis=-1, keepdims=True) # stable softmax
e = np.exp(S)
return e / e.sum(axis=-1, keepdims=True)
def attention(Q, K, V):
return softmax(Q @ K.swapaxes(-1, -2) / np.sqrt(Q.shape[-1])) @ V
n, d_model, h = 5, 8, 2
X = rng.normal(size=(n, d_model))
WQ, WK, WV, WO = (rng.normal(size=(d_model, d_model)) / np.sqrt(d_model) for _ in range(4))
single = attention(X @ WQ, X @ WK, X @ WV) # (n, d_model)
def split(M): # (n, d) -> (h, n, d/h)
return M.reshape(n, h, d_model // h).transpose(1, 0, 2)
heads = attention(split(X @ WQ), split(X @ WK), split(X @ WV)) # one batched matmul
multi = heads.transpose(1, 0, 2).reshape(n, d_model) @ WO # concat heads, mix with W_O
print(single.shape, heads.shape, multi.shape)
# ---------- 7. A LoRA layer in plain NumPy ----------
class LoRALinear:
def __init__(self, W, r=8, alpha=16, rng=None):
rng = rng or np.random.default_rng(0)
self.W = W # frozen, shape (d_out, d_in)
self.A = rng.normal(scale=0.01, size=(r, W.shape[1])) # trainable, small random numbers
self.B = np.zeros((W.shape[0], r)) # trainable, zeros: delta W = 0 at the start
self.s = alpha / r # the scale alpha / r
def forward(self, X): # X: (n, d_in), one example per row
return X @ self.W.T + self.s * (X @ self.A.T) @ self.B.T # squeeze with A first, then expand with B
def backward(self, X, G): # G = dLoss/dOutput, shape (n, d_out)
gB = self.s * G.T @ (X @ self.A.T) # (d_out, r)
gA = self.s * (G @ self.B).T @ X # (r, d_in)
return gA, gB # W gets no gradient: it is frozen
def merge(self):
return self.W + self.s * self.B @ self.A # W' = W + (alpha/r) B A
lay = LoRALinear(rng.normal(size=(64, 64)) / 8)
X = rng.normal(size=(5, 64))
print(np.allclose(lay.forward(X), X @ lay.W.T)) # True: B = 0, so same as the original model
lay.B = 0.1 * rng.normal(size=lay.B.shape) # pretend training has changed B
print(np.allclose(lay.forward(X), X @ lay.merge().T)) # True: merged weights give the same output
print(lay.W.size, lay.A.size + lay.B.size) # 4096 versus 1024 trainable numbers (r = 8)
# check the gradient of B numerically (loss = sum of outputs squared / 2)
G = lay.forward(X)
gA, gB = lay.backward(X, G)
eps = 1e-6
lay.B[2, 3] += eps
num = (0.5 * np.sum(lay.forward(X) ** 2) - 0.5 * np.sum(G ** 2)) / eps
print(abs(num - gB[2, 3]) < 1e-3 * max(1, abs(num))) # True
# truncated SVD: the best rank-r approximation of any matrix (Eckart-Young)
M = rng.normal(size=(64, 3)) @ rng.normal(size=(3, 64)) + 0.01 * rng.normal(size=(64, 64))
U, s, Vt = np.linalg.svd(M)
r = 3
M_r = (U[:, :r] * s[:r]) @ Vt[:r] # keep only the top r layers
print(np.linalg.norm(M - M_r) / np.linalg.norm(M)) # small: M is close to rank 3
# --- the same idea in PyTorch (a sketch; needs torch installed) ---
# import torch, torch.nn as nn
# class LoRALinear(nn.Module):
# def __init__(self, base: nn.Linear, r=8, alpha=16):
# super().__init__()
# self.base = base
# for p in self.base.parameters():
# p.requires_grad = False # freeze W
# self.A = nn.Parameter(torch.randn(r, base.in_features) * 0.01)
# self.B = nn.Parameter(torch.zeros(base.out_features, r))
# self.s = alpha / r
# def forward(self, x):
# return self.base(x) + self.s * (x @ self.A.T) @ self.B.T
# With Hugging Face PEFT you do not write this yourself:
# from peft import LoraConfig, get_peft_model
# cfg = LoraConfig(r=8, lora_alpha=16, target_modules=["q_proj", "v_proj"])
# model = get_peft_model(model, cfg) # base frozen, only A and B train
1. In least-squares regression, the residual $\mathbf{r}=\mathbf{y}-X\hat{\mathbf{w}}$ is…
2. A logistic regression has $\mathbf{w}=[3,4]$ and $b=-10$. How far is the decision boundary from the origin?
3. Why can't a covariance matrix have a negative eigenvalue?
4. The singular values of a centred data matrix with $n=100$ rows are $20$ and $10$. What fraction of the variance does the first principal component explain?
5. Why does scaled dot-product attention divide the scores by $\sqrt{d_k}$?
6. A ratings table $R$ is $1000\times500$ and is approximated as $UV^\top$ with rank $20$. How many numbers do $U$ and $V$ store together?
7. In LoRA, why is $B$ started at all zeros while $A$ is random?
8. A $4096\times4096$ weight matrix is adapted with LoRA of rank 8. How many numbers are trained?
Practice problems
A. Fit $\hat y = w_0 + w_1x$ to the points $(0,1),(1,2),(2,4)$ with the normal equations. Verify $X^\top\mathbf{r}=\mathbf{0}$.
$X=\begin{bmatrix}1&0\\1&1\\1&2\end{bmatrix}$, $\mathbf{y}=[1,2,4]$. $X^\top X=\begin{bmatrix}3&3\\3&5\end{bmatrix}$, $X^\top\mathbf{y}=[7, 10]$. The determinant is $15-9=6$, so $\mathbf{w}=\frac16\begin{bmatrix}5&-3\\-3&3\end{bmatrix}[7,10] = \frac16[35-30,\ -21+30]=[\tfrac56,\tfrac32]$.
Predictions $[\tfrac56,\tfrac73,\tfrac{23}6]$, residuals $\mathbf{r}=[\tfrac16,-\tfrac13,\tfrac16]$. Then $X^\top\mathbf{r} = [\tfrac16-\tfrac13+\tfrac16,\ 0-\tfrac13+\tfrac13] = [0,0]$ ✓.
B. A logistic model has $\mathbf{w}=[1,-2]$, $b=1$. For $\mathbf{x}=[3,1]$ with true label $0$, find $p$, the loss, and the gradient with respect to $\mathbf{w}$.
$z = 3-2+1 = 2$, so $p=\sigma(2)=1/(1+e^{-2})\approx0.881$. The true label is 0, so the loss is $-\log(1-p) = -\log0.119\approx2.13$ (a confident mistake). The gradient is $(p-y)\mathbf{x} = 0.881\cdot[3,1]\approx[2.64,\,0.88]$. A gradient step will lower the score on this example.
C. Data rows $[1,2]$ and $[3,6]$. Compute $\Sigma$ and its eigenvalues. What does PCA say?
Mean $[2,4]$; centred rows $[-1,-2]$ and $[1,2]$. $\Sigma=\frac12\begin{bmatrix}1+1&2+2\\2+2&4+4\end{bmatrix} = \begin{bmatrix}1&2\\2&4\end{bmatrix}$. Trace $5$, determinant $0$, so $\lambda=5$ and $0$. The first principal direction is $[1,2]/\sqrt5$ and it explains $100\%$ of the variance: the data lies exactly on a line. The correlation is $2/\sqrt{1\cdot4}=1$.
D. With $\Sigma=\operatorname{diag}(4,9)$ and $\boldsymbol\mu=\mathbf{0}$, find the Mahalanobis distance of $\mathbf{x}=[2,3]$.
$d_M^2=\frac{2^2}{4}+\frac{3^2}{9}=1+1=2$, so $d_M=\sqrt2\approx1.41$. In each direction the point is exactly one standard deviation out; the Euclidean distance $\sqrt{13}\approx3.6$ hides that.
E. One query $\mathbf{q}=[1,0]$, keys $\mathbf{k}_1=[1,0]$, $\mathbf{k}_2=[0,1]$, values $\mathbf{v}_1=[10,0]$, $\mathbf{v}_2=[0,10]$, $d_k=2$. Find the attention output.
Scores: $[1,0]/\sqrt2 = [0.707, 0]$. Softmax: $e^{0.707}=2.03$ and $e^0=1$, total $3.03$, weights $[0.670, 0.330]$. Output: $0.670\cdot[10,0]+0.330\cdot[0,10]=[6.70,\,3.30]$. The query matches key 1 better, so the output leans toward value 1.
F. A weight matrix is $1024\times1024$. Compare a full update with a rank-16 LoRA update.
Full: $1024^2=1{,}048{,}576$ numbers. LoRA: $2\cdot1024\cdot16 = 32{,}768$, about $3.1\%$ (32 times fewer). The update $BA$ has rank at most 16.
G. A $1000\times1000$ weight matrix gets a rank-4 LoRA update. Count the trainable numbers and the extra work per token.
Trainable: $2\cdot1000\cdot4 = 8000$ instead of $1{,}000{,}000$ (0.8%, or 125 times fewer). Extra work: squeeze $A\mathbf{x}$ costs $2\cdot4\cdot1000 = 8000$ FLOPs and expand $B(A\mathbf{x})$ another $8000$, so $16{,}000$, against $2\cdot1000^2 = 2{,}000{,}000$ for the main road: 0.8%. Building $BA$ first would cost far more, which is why we always squeeze first.
H. $W = \begin{bmatrix}1&2\\0&1\end{bmatrix}$, $B = [1,-1]^\top$, $A=[2,3]$, $\alpha/r=1$, $\mathbf{x}=[1,1]$. Compute the output with the adapter kept separate, then merged.
Separate: $A\mathbf{x} = 5$, $B(A\mathbf{x}) = [5,-5]$, $W\mathbf{x} = [3, 1]$, total $[8, -4]$. Merged: $BA = \begin{bmatrix}2&3\\-2&-3\end{bmatrix}$, so $W' = \begin{bmatrix}3&5\\-2&-2\end{bmatrix}$ and $W'\mathbf{x} = [8, -4]$. Identical.
Glossary
Every important word in this guide, in one place, explained in plain English. Type in the box to filter. Each entry links to the chapter that teaches it.
- Absolute value
- How far a number is from zero. Always 0 or positive: $|-3| = 3$. 1.1
- Affine transformation
- A linear transformation followed by a shift: $\mathbf{y} = A\mathbf{x} + \mathbf{b}$. A neural-network layer is affine. 1.5
- Algebraic multiplicity
- How many times an eigenvalue appears as a root of the characteristic polynomial. 1.11
- Angle between vectors
- The angle $\theta$ with $\cos\theta = \dfrac{\mathbf{a}\cdot\mathbf{b}}{\|\mathbf{a}\|\|\mathbf{b}\|}$. 1.2
- Attention
- The Transformer step that mixes value vectors using weights from softmax of scaled dot products between queries and keys. 1.17
- Augmented matrix
- The matrix $A$ with the right-hand side $\mathbf{b}$ attached as an extra column, written $[A \mid \mathbf{b}]$. 1.6
- Basis
- A set of vectors that are independent and that span the whole space. Every vector is built from them in exactly one way. 1.3
- Batch matrix multiplication
- Doing many matrix products at once by stacking matrices into a 3D tensor. 1.16
- Broadcasting
- The rule that lets arrays of different shapes be combined by automatically stretching the smaller one. 1.16
- Cauchy–Schwarz inequality
- $|\mathbf{a}\cdot\mathbf{b}| \le \|\mathbf{a}\|\|\mathbf{b}\|$. The shadow is never longer than the arrow. 1.2
- Change of basis
- Describing the same vector using the coordinates of a different basis. 1.3
- Characteristic polynomial
- $\det(A - \lambda I)$. Its roots are the eigenvalues of $A$. 1.11
- Cholesky decomposition
- The "square root" of a symmetric positive definite matrix: $A = LL^\top$ with $L$ lower triangular. Used to solve systems about twice as fast, to sample correlated random numbers, to get the log-determinant cheaply, and to test positive definiteness. 1.13
- Column space
- All the outputs $A\mathbf{x}$ can produce. It is the span of the columns of $A$. 1.8
- Complex number
- A number $a + bi$ with $i^2 = -1$. Think of it as the point $(a, b)$ on a flat page. Needed because some matrices have eigenvalues that are not real. 1.1 1.11
- Condition number
- $\kappa(A) = \sigma_{\max}/\sigma_{\min}$. Says how much small errors in the input can be magnified in the answer. Large means "ill-conditioned". 1.7 1.15
- Consistent system
- A system of equations that has at least one solution. 1.6
- Cosine similarity
- The cosine of the angle between two vectors. $+1$ same direction, $0$ unrelated, $-1$ opposite. 1.2
- Covariance matrix
- A table showing how every pair of features varies together. Always symmetric and positive semidefinite. 1.17
- Determinant
- The signed factor by which a matrix scales area (2D) or volume (3D). Zero means the matrix squashes space and cannot be inverted. 1.7
- Diagonal matrix
- A matrix whose only non-zero entries are on the main diagonal. 1.4
- Diagonalization
- Writing $A = PDP^{-1}$ with $D$ diagonal. Makes powers of $A$ easy. 1.11
- Dimension
- For a vector: the number of entries. For a space: the number of vectors in any basis. 1.3
- Dot product
- Multiply matching entries and add: $\sum a_ib_i$. Measures how much two vectors agree. 1.2
- Eckart–Young theorem
- Keeping the $k$ biggest singular values gives the best possible rank-$k$ approximation of a matrix. 1.13
- Eigenvalue / eigenvector
- A direction $\mathbf{v}$ that a matrix only stretches, $A\mathbf{v} = \lambda\mathbf{v}$. The stretch factor is $\lambda$. 1.11
- einsum
- A compact notation that describes tensor products and sums using index letters, such as
ij,jk->ik. 1.16 - Embedding
- A vector that represents a word, image, user or item so that similar things get nearby vectors. 1.17
- Euclidean norm (L2)
- The straight-line length $\sqrt{v_1^2 + \dots + v_n^2}$. 1.2
- Floating-point number
- The way computers store real numbers approximately, using a sign, an exponent and a mantissa. 1.15
- Forward / back substitution
- Solving a triangular system one unknown at a time, starting from the first (or last) row. It is why triangular matrices are so easy to work with. 1.13
- Frobenius norm
- The square root of the sum of the squares of all matrix entries. 1.7
- Full rank
- A matrix whose rank is as large as it can be ($\min(m, n)$). 1.7
- Function
- A rule that gives exactly one output for each input. 1.1
- Gaussian elimination
- Using row operations to turn a system into a staircase (echelon) shape that is easy to solve. 1.6
- Gradient
- The vector of all partial derivatives. It points in the direction of steepest increase. 1.14
- Gram matrix
- $A^\top A$: the table of dot products between the columns of $A$. Always symmetric and positive semidefinite. 1.4
- Gram–Schmidt process
- A recipe that turns independent vectors into orthonormal ones by subtracting projections. 1.9
- Hadamard product
- Entry-by-entry multiplication of two same-shape arrays, written $A \odot B$. 1.4
- Hessian
- The matrix of all second derivatives. Describes curvature. 1.14
- Identity matrix
- $I$: ones on the diagonal, zeros elsewhere. Multiplying by it changes nothing. 1.4
- Ill-conditioned
- A problem where tiny changes in the input cause big changes in the output. 1.15
- Injective / surjective / bijective
- Injective (one-to-one): different inputs always give different outputs. Surjective (onto): every possible output is produced. Bijective: both, so the function can be undone perfectly. 1.1
- Inverse
- The matrix $A^{-1}$ that undoes $A$: $AA^{-1} = I$. Exists only for non-singular square matrices. 1.7
- Jacobian
- The matrix of all first partial derivatives of a vector-valued function. It is the best local linear approximation. 1.14
- Jitter
- A tiny number (such as $10^{-6}$) added to the diagonal of a matrix so that a nearly positive definite matrix becomes safely positive definite and Cholesky does not fail. 1.13 1.15
- L1 norm
- The sum of absolute values of the entries (taxi distance). 1.2
- Least squares
- Finding the $\mathbf{x}$ that makes $\|A\mathbf{x} - \mathbf{b}\|^2$ as small as possible when no exact solution exists. 1.10
- Left null space
- All $\mathbf{y}$ with $A^\top\mathbf{y} = \mathbf{0}$. 1.8
- Linear combination
- A recipe $c_1\mathbf{v}_1 + \dots + c_k\mathbf{v}_k$: some scoops of each vector, added up. 1.2
- Linear independence
- No vector in the set can be built from the others. 1.3
- Linear transformation
- A function that keeps adding and scaling intact: $T(a\mathbf{x} + b\mathbf{y}) = aT(\mathbf{x}) + bT(\mathbf{y})$. Every one (from one list-of-numbers space to another) can be written as a matrix. 1.5
- L∞ norm
- The largest absolute entry (chess-king distance). 1.2
- Log-determinant
- $\log\det A$. With Cholesky it is cheap and safe: $\log\det A = 2\sum_i \log L_{ii}$. Appears in Gaussian likelihoods and Gaussian processes. 1.13
- Logarithm
- The inverse of an exponent: $\log_b x = y$ means $b^y = x$. It turns products into sums, $\log(xy) = \log x + \log y$. $\ln$ is the log to base $e \approx 2.718$. 1.1
- Log-likelihood
- The sum of the logs of the probabilities a model gives to the data, $\sum_i \log p_i$. It is used instead of the product of probabilities, which would underflow to 0 on a computer. 1.1
- LoRA (Low-Rank Adaptation)
- A way to fine-tune a huge model cheaply: freeze the original weights $W$ and learn only a small low-rank update $\Delta W = BA$ with a tiny rank $r$. 1.17
- Low-rank factorisation
- Writing a big $m \times n$ matrix as the product of two skinny matrices, $A \approx BC$, so you store and multiply $(m+n)r$ numbers instead of $mn$. This is the matrix-efficiency idea behind LoRA, compression and recommenders. 1.13 1.17
- LU decomposition
- Factoring a matrix as lower-triangular × upper-triangular (with row swaps: $PA = LU$). 1.13
- Mahalanobis distance
- A distance that accounts for the shape of the data cloud: $\sqrt{(\mathbf{x}-\boldsymbol{\mu})^\top\Sigma^{-1}(\mathbf{x}-\boldsymbol{\mu})}$. 1.12 1.17
- Matrix
- A rectangular table of numbers. Also a machine that transforms vectors. 1.4
- Moore–Penrose pseudoinverse
- $A^{+}$: a generalised inverse that works for any matrix and gives the least-squares / minimum-norm answer. 1.10
- Norm
- A way of measuring the size of a vector. It is never negative, it is zero only for the zero vector, it scales properly ($\|c\mathbf{v}\| = |c|\,\|\mathbf{v}\|$), and it obeys the triangle inequality. 1.2
- Normal equations
- $A^\top A\,\hat{\mathbf{x}} = A^\top\mathbf{b}$. Solving them gives the least-squares answer. 1.10
- Null space
- All inputs $\mathbf{x}$ that a matrix sends to zero: $A\mathbf{x} = \mathbf{0}$. 1.8
- Nullity
- The dimension of the null space. 1.8
- Orthogonal
- Perpendicular. Two vectors are orthogonal when their dot product is 0. 1.2
- Orthogonal matrix
- A square matrix with $Q^\top Q = I$. It rotates or reflects without changing lengths. 1.9
- Orthonormal
- Orthogonal and each of length 1. 1.9
- Outer product
- $\mathbf{x}\mathbf{y}^\top$: a column times a row gives a whole matrix of rank 1. 1.4
- PCA (principal component analysis)
- Finding the directions along which data varies most, then describing the data with just a few of them. Uses the eigenvectors of the covariance matrix, or the SVD. 1.17
- PEFT (parameter-efficient fine-tuning)
- The family of methods (LoRA is the best known) that adapt a big pre-trained model by training only a small number of extra parameters. 1.17
- Pivot
- The first non-zero entry in a row during elimination. 1.6
- Positive definite
- $\mathbf{x}^\top A\mathbf{x} \gt 0$ for every non-zero $\mathbf{x}$ (for a symmetric $A$, this means all eigenvalues are positive). A perfect bowl. 1.12
- Positive semidefinite
- $\mathbf{x}^\top A\mathbf{x} \ge 0$ for every $\mathbf{x}$ (for a symmetric $A$, this means all eigenvalues are $\ge 0$). 1.12
- QLoRA
- LoRA applied on top of a base model stored in 4-bit compressed numbers, so even larger models fit on one GPU. 1.17
- Rank
- The number of independent columns (equals the number of independent rows). 1.7
- Rank–nullity theorem
- rank + nullity = number of columns. 1.8
- Residual
- What is left over: the true value minus the predicted value (or $\mathbf{b} - A\hat{\mathbf{x}}$). In least squares it is perpendicular to the column space. 1.10
- Row echelon form
- The staircase shape you get after Gaussian elimination. 1.6
- Row space
- The span of the rows of a matrix. 1.8
- Scalar
- An ordinary single number (as opposed to a list or a table of numbers). Used, for example, to scale vectors. 1.1 1.2
- Set
- A collection of things where only membership matters (not order or repeats), such as $\{1, 2, 3\}$. Written with $\in$, $\cup$, $\cap$. 1.1
- Singular matrix
- A square matrix with no inverse (determinant 0). 1.7
- Singular value
- One of the stretch factors $\sigma_i \ge 0$ in the SVD. 1.13
- Span
- The set of all linear combinations of some vectors. 1.3
- Sparse matrix
- A matrix that is mostly zeros, stored in a special way to save memory and time. 1.15
- Spectral theorem
- A symmetric matrix has real eigenvalues and orthogonal eigenvectors: $A = Q\Lambda Q^\top$. 1.11
- Summation ($\Sigma$)
- Shorthand for adding a list of terms, like a for-loop that adds: $\sum_{i=1}^{n} a_i = a_1 + \dots + a_n$. 1.1
- SVD (singular value decomposition)
- $A = U\Sigma V^\top$ for any matrix: rotate, stretch, rotate. 1.13
- Symmetric matrix
- A matrix equal to its own transpose. 1.4
- Tensor
- An array with any number of axes: scalar, vector, matrix, and beyond. 1.16
- Trace
- The sum of the diagonal entries of a square matrix. 1.7
- Transpose
- Flip a matrix over its diagonal so rows become columns: $A^\top$. 1.4
- Triangle inequality
- $\|\mathbf{a} + \mathbf{b}\| \le \|\mathbf{a}\| + \|\mathbf{b}\|$. A detour is never shorter than the straight path. 1.2
- Triangular matrix
- A square matrix with zeros on one side of the diagonal (lower or upper). Systems with triangular matrices are solved by simple substitution. 1.4
- Underflow
- When a number is so close to zero that the computer cannot store it and rounds it to exactly 0. Multiplying many probabilities causes this. 1.1 1.15
- Unit vector
- A vector of length 1. 1.2
- Vector
- An ordered list of numbers; also an arrow. 1.2
- Vector space
- A collection of objects you can add and scale without ever leaving the collection. 1.3
- Whitening
- Transforming data so its covariance becomes the identity matrix (uncorrelated, equal variance). Can be done with Cholesky or PCA. 1.17
- Zero vector
- The vector with every entry 0. 1.2
Formula cheat sheet
The formulas you will use most, on one page. Use it to refresh your memory. If a formula does not make sense, the chapter link tells you where to relearn the idea.
Vectors and matrices
| Idea | Formula | Chapter |
|---|---|---|
| Dot product | $\mathbf{a}\cdot\mathbf{b} = \sum a_ib_i = \|\mathbf{a}\|\|\mathbf{b}\|\cos\theta$ | 1.2 |
| L2 / L1 / L∞ norms | $\sqrt{\sum v_i^2}$, $\sum |v_i|$, $\max |v_i|$ | 1.2 |
| Cosine similarity | $\dfrac{\mathbf{a}\cdot\mathbf{b}}{\|\mathbf{a}\|\|\mathbf{b}\|}$ | 1.2 |
| Projection of $\mathbf{b}$ on $\mathbf{a}$ | $\dfrac{\mathbf{a}\cdot\mathbf{b}}{\mathbf{a}\cdot\mathbf{a}}\,\mathbf{a}$ | 1.2 |
| Matrix product entry | $(AB)_{ij} = \sum_k A_{ik}B_{kj}$ | 1.4 |
| Transpose of a product | $(AB)^\top = B^\top A^\top$ | 1.4 |
| Inverse of a product | $(AB)^{-1} = B^{-1}A^{-1}$ | 1.7 |
| 2×2 inverse | $\dfrac{1}{ad-bc}\begin{bmatrix} d & -b \\ -c & a \end{bmatrix}$ | 1.7 |
| Rank–nullity | $\operatorname{rank}(A) + \operatorname{nullity}(A) = n$ | 1.8 |
Solving, fitting and decomposing
| Idea | Formula | Chapter |
|---|---|---|
| Orthogonal matrix | $Q^\top Q = I \;\Rightarrow\; Q^{-1} = Q^\top$ | 1.9 |
| Projection matrix onto $C(A)$ | $P = A(A^\top A)^{-1}A^\top$ (with orthonormal $Q$: $P = QQ^\top$) | 1.9 |
| Normal equations | $A^\top A\,\hat{\mathbf{x}} = A^\top\mathbf{b}$ | 1.10 |
| Ridge regression | $(A^\top A + \lambda I)\,\mathbf{x} = A^\top\mathbf{b}$ | 1.10 |
| Pseudoinverse via SVD | $A^{+} = V\Sigma^{+}U^\top$ | 1.10 |
| Eigen equation | $A\mathbf{v} = \lambda\mathbf{v}, \quad \det(A - \lambda I) = 0$ | 1.11 |
| Trace and determinant | $\operatorname{tr}(A) = \sum\lambda_i, \quad \det(A) = \prod\lambda_i$ | 1.11 |
| Diagonalization / spectral | $A = PDP^{-1}$; symmetric: $A = Q\Lambda Q^\top$ | 1.11 |
| Quadratic form | $f(\mathbf{x}) = \mathbf{x}^\top A\mathbf{x}$ | 1.12 |
| SVD and low-rank approximation | $A = U\Sigma V^\top = \sum\sigma_i\mathbf{u}_i\mathbf{v}_i^\top; \quad A_k = \sum_{i\le k}\sigma_i\mathbf{u}_i\mathbf{v}_i^\top$ | 1.13 |
| Cholesky (SPD matrices) | $A = LL^\top$; solve $A\mathbf{x}=\mathbf{b}$: $L\mathbf{y}=\mathbf{b}$ then $L^\top\mathbf{x}=\mathbf{y}$; $\log\det A = 2\sum\log L_{ii}$ | 1.13 |
| Condition number | $\kappa(A) = \sigma_{\max}/\sigma_{\min}$ | 1.15 |
Calculus and machine learning
| Idea | Formula | Chapter |
|---|---|---|
| Gradient identities | $\nabla_{\mathbf{x}}\,\mathbf{a}^\top\mathbf{x} = \mathbf{a}, \quad \nabla_{\mathbf{x}}\,\mathbf{x}^\top A\mathbf{x} = (A + A^\top)\mathbf{x}$ | 1.14 |
| Least-squares gradient | $\nabla_{\mathbf{x}}\|A\mathbf{x} - \mathbf{b}\|^2 = 2A^\top(A\mathbf{x} - \mathbf{b})$ | 1.14 |
| Gradient descent step | $\mathbf{w} \leftarrow \mathbf{w} - \eta\,\nabla L(\mathbf{w})$ | 1.14 |
| Linear layer | $\mathbf{y} = W\mathbf{x} + \mathbf{b}$, batch: $Y = XW^\top + \mathbf{b}$ | 1.17 |
| Covariance (centered data) | $\Sigma = \dfrac{1}{n}X^\top X$ (many books use $\frac{1}{n-1}$) | 1.17 |
| PCA | principal directions = eigenvectors of $\Sigma$ = right singular vectors of centered $X$ | 1.17 |
| Sampling correlated Gaussians | $\mathbf{x} = \boldsymbol{\mu} + L\mathbf{z}$, $\Sigma = LL^\top$, $\mathbf{z}\sim\mathcal{N}(0, I)$ | 1.13 |
| LoRA layer | $\mathbf{h} = W\mathbf{x} + \dfrac{\alpha}{r}\,B A\,\mathbf{x}$ with $B\in\mathbb{R}^{d\times r},\, A\in\mathbb{R}^{r\times d}$; parameters $2dr$ instead of $d^2$ | 1.17 |
| Attention | $\operatorname{softmax}\!\left(\dfrac{QK^\top}{\sqrt{d_k}}\right)V$ | 1.17 |
Where to go next
You finished the tour. Here is how to make it stick, and where to dig deeper.
How to make it stick
- Code every idea. Each chapter ends with a NumPy block. Type it out (do not paste) and change the numbers.
- Re-derive, don't memorise. If you can explain why the normal equations come from "the error is perpendicular", you will never forget them.
- Teach it. Explain SVD to a friend using only "rotate, stretch, rotate".
- Build something. Compress an image with SVD, write PCA from scratch, or code a tiny attention layer.
Next in the series
Linear algebra tells you what vectors and matrices are. Calculus tells you how a model's error changes when those numbers change, which is how learning works. Continue with the Calculus for Machine Learning guide: it picks up exactly where the matrix-calculus chapter (1.14) leaves off and goes much deeper into gradients, the chain rule and backpropagation.
Free resources
- 3Blue1Brown, "Essence of Linear Algebra" (YouTube): the best visual intuition.
- Gilbert Strang, MIT 18.06 (OCW): the classic full course. Then MIT 18.065 for matrix methods in data and ML.
- Mathematics for Machine Learning (Deisenroth, Faisal, Ong): free PDF, matches this guide closely.
- Introduction to Applied Linear Algebra (Boyd & Vandenberghe): free, very practical.
- The Matrix Cookbook (Petersen & Pedersen): free reference for matrix calculus identities.