USAAIO Lesson 22, from Week 8, fully worked. It defines the affine map T(x)=Wx+b and separates it into its linear and translation parts, tests the linearity axioms one at a time, and reads the geometric transforms - rotation, reflection, scaling, and shear - off W. It proves the affine invariants for lines, parallelism, and ratios, then derives the composition T2 of T1 move by move to (W2W1)x + (W2b1 + b2), and gives the homogeneous-coordinate trick that turns Wx+b into a single matrix. It then shows a deep linear stack collapsing into one matrix, three layers at a time, and why a ReLU breaks that collapse, since additivity fails on real numbers. It settles XOR once and for all: the best linear fit is stuck at 0.5 while a tanh MLP reaches 1.0, alongside a hand-built two-ReLU XOR. Every snippet runs standalone, and every number came from real execution with numpy 2.2.6 and torch 2.7.1, seeded. The lesson runs to 61 slides.
Subject: Machine Learning · 114 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 22 · Week 8
Why a deep stack of linear layers is secretly just one layer — and why that single fact makes nonlinearity non-negotiable. We build Wx + b from its two halves, prove composition stays affine one move at a time, watch three layers collapse into one matrix, and settle XOR with real numbers.
Objectives
T(x) = Wx + b and split it into its linear part Wx and translation + b+ b breaks them (an affine map is linear only when b = 0)W, and name what affine maps preserve (lines, parallelism, ratios)T₂∘T₁ step by step to (W₂W₁)x + (W₂b₁ + b₂) — and repackage it as one matrix in homogeneous coordinates0.5, a tanh MLP at 1.0Warm-up
Discussion prompt
Before we open Lesson 22: Affine Transforms: without looking back, what was the main idea of Second-Order Methods, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
the Hessian and what its definiteness says about minima vs saddles, Newton's method and its quadratic convergence, why second-order methods cost O(d³), and the L-BFGS / natural-gradient approximations. Implement Newton for logistic regression and race it against gradient descent.
Section
Part 1 of 6 — what an affine map is
Concept
An affine map takes a vector x, applies a matrix W, then adds a fixed vector b. The Wx rotates, scales, or shears; the + b slides the whole result.
\[ T(x) = Wx + b \]
This is exactly one layer of a neural network before its activation — the nn.Linear step. Master this object and you understand every layer's skeleton.
Counterexample
Discussion prompt
An affine map takes a vector x, applies a matrix W, then adds a fixed vector b. The Wx rotates, scales, or shears; the + b slides the whole result.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Concept
We fix one concrete 2-D map and carry it through the whole lesson. Here is W, b, and the input x we will keep reusing:
\[ W = \begin{bmatrix} 2 & 0 \\ 1 & 3 \end{bmatrix}, \qquad b = \begin{bmatrix} 1 \\ -1 \end{bmatrix}, \qquad x = \begin{bmatrix} 3 \\ 4 \end{bmatrix} \]
Two pieces: the (2×2) matrix W and the shift vector b
Why: W acts on the 2-vector x to give a 2-vector; then b (also a 2-vector) is added. Shapes must line up — that is the first thing to check on any affine map.
Ranking
Put in order
Put the moves of Evaluate T(x), entry by entry into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. First row [2, 0] dotted with [3, 4]: the second feature is zeroed out, so only 2·3 survives.
Worked example
Do Wx first: each row of W dotted with x = [3, 4]. Then add b component-wise. No steps skipped.
Row 0 of Wx: 2·3 + 0·4 = 6
Why: First row [2, 0] dotted with [3, 4]: the second feature is zeroed out, so only 2·3 survives.
Row 1 of Wx: 1·3 + 3·4 = 3 + 12 = 15
Why: Second row [1, 3] dotted with [3, 4]. So Wx = [6, 15].
\[ Wx = \begin{bmatrix} 2 & 0 \\ 1 & 3 \end{bmatrix}\!\begin{bmatrix} 3 \\ 4 \end{bmatrix} = \begin{bmatrix} 6 \\ 15 \end{bmatrix} \]
Add b = [1, −1]: T(x) = [6+1, 15−1] = [7, 14]
Why: The translation slides Wx by b component-wise. This [7, 14] is T(x), and we will confirm it in code next.
\[ T(x) = Wx + b = \begin{bmatrix} 6 \\ 15 \end{bmatrix} + \begin{bmatrix} 1 \\ -1 \end{bmatrix} = \begin{bmatrix} 7 \\ 14 \end{bmatrix} \]
Notation
Annotate
From Evaluate T(x), entry by entry — read this one piece at a time. What is each part doing?
On: \( Wx = \begin{bmatrix} 2 & 0 \\ 1 & 3 \end{bmatrix}\!\begin{bmatrix} 3 \\ 4 \end{bmatrix} = \begin{bmatrix} 6 \\ 15 \end{bmatrix} \)
Estimation
Predict first
Type the map out and print Wx, then T(x). Predict [7, 14] before you run it. This block is complete and runnable on its own:
Commit before you compute: what does Confirm T(x) in code come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Wx = [6, 15], then + b gives T(x) = [7, 14]
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Matches the hand computation exactly.
Worked example
Type the map out and print Wx, then T(x). Predict [7, 14] before you run it. This block is complete and runnable on its own:
import numpy as np
W = np.array([[2., 0.], [1., 3.]])
b = np.array([1., -1.])
x = np.array([3., 4.])
print('Wx =', (W @ x).tolist())
print('T(x) =', (W @ x + b).tolist())Wx = [6, 15], then + b gives T(x) = [7, 14]
Why: Matches the hand computation exactly. The '@' operator is matrix-vector multiply; adding b broadcasts component-wise.
| expression | value (verified) |
|---|---|
| W @ x | [6.0, 15.0] |
| W @ x + b | [7.0, 14.0] |
| shapes | W (2×2), x (2,), b (2,) → T(x) (2,) |
Comparison
Comparison matrix
From Confirm T(x) in code: refill the value (verified) column from what you know. The rest of the table is as it appeared.
| expression | value (verified) |
|---|---|
| W @ x | [6.0, 15.0] |
| W @ x + b | [7.0, 14.0] |
| shapes | W (2×2), x (2,), b (2,) → T(x) (2,) |
Missing information
Discussion prompt
To see what T does, push the four corners of the unit square through it. The image is a parallelogram — never a curved blob, because affine maps keep lines straight. Runnable:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
T(0) = W·0 + b = b. The affine map moves the origin to the shift vector — the visible signature that b ≠ 0.
Worked example
To see what T does, push the four corners of the unit square through it. The image is a parallelogram — never a curved blob, because affine maps keep lines straight. Runnable:
import numpy as np
W = np.array([[2., 0.], [1., 3.]]); b = np.array([1., -1.])
corners = np.array([[0., 0.], [1., 0.], [0., 1.], [1., 1.]])
for p in corners:
print(p.tolist(), '->', (W @ p + b).tolist())The origin [0,0] maps to b = [1,−1]
Why: T(0) = W·0 + b = b. The affine map moves the origin to the shift vector — the visible signature that b ≠ 0.
Opposite edges stay parallel in the image
Why: Edge [0,0]→[1,0] maps to [1,−1]→[3,0] (direction [2,1]); the opposite edge [0,1]→[1,1] maps to [1,2]→[3,3] (also [2,1]). Parallelism preserved — the square becomes a parallelogram.
| corner | T(corner) |
|---|---|
| [0, 0] | [1, −1] |
| [1, 0] | [3, 0] |
| [0, 1] | [1, 2] |
| [1, 1] | [3, 3] |
Trade off
Comparison matrix
From Map the whole unit square: every row here is a choice with a cost. Fill the T(corner) column, then say which row you would actually pick and what you give up for it.
| corner | T(corner) |
|---|---|
| [0, 0] | [1, −1] |
| [1, 0] | [3, 0] |
| [0, 1] | [1, 2] |
| [1, 1] | [3, 3] |
Concept
A map T is linear when it satisfies two axioms for all vectors u, v and scalars c:
\[ \text{additivity: } T(u+v) = T(u)+T(v), \qquad \text{homogeneity: } T(cu) = c\,T(u) \]
The pure matrix map x ↦ Wx obeys both. The shift + b is what an affine map adds on top — and that shift is exactly what can break linearity.
Analogy
Discussion prompt
Explain When is an affine map linear? by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
A map T is linear when it satisfies two axioms for all vectors u, v and scalars c:
Step zero
Discussion prompt
The shift breaks homogeneity — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Left side: T(2x) = W(2x) + b = 2Wx + b
Answer:
Worked example
Test homogeneity T(cx) = c·T(x) on our map with c = 2. Watch the bias fail to scale.
Left side: T(2x) = W(2x) + b = 2Wx + b
Why: Doubling x doubles Wx (matrix mult is linear), but b is added once, unchanged.
\[ T(2x) = 2\,Wx + b \]
Right side: 2·T(x) = 2(Wx + b) = 2Wx + 2b
Why: Scaling the whole output doubles b too.
\[ 2\,T(x) = 2\,Wx + 2b \]
They differ by exactly b (unless b = 0)
Why: T(2x) − 2T(x) = b − 2b = −b. So the map is linear only when b = 0; with a nonzero shift it is affine but NOT linear.
\[ T(2x) - 2\,T(x) = b - 2b = -b \;\neq\; 0 \]
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
They differ by exactly b (unless b = 0)
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Test homogeneity T(cx) = c·T(x) on our map with c = 2. Watch the bias fail to scale.
Intuition
A linear map must send 0 to 0 — plug in x = 0 and W·0 = 0. An affine map sends 0 to b: the origin is free to move.
Picture a rubber sheet. The linear part W stretches and rotates the sheet around a pinned origin; the + b then picks the whole sheet up and slides it. Nothing bends — that is the key limitation we exploit later.
Explain it
Discussion prompt
Explain Affine = linear, but the origin can move to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
A linear map must send 0 to 0 — plug in x = 0 and W·0 = 0. An affine map sends 0 to b: the origin is free to move.
Anomaly
Predict first
A student writes this, and it looks reasonable:
T(x) = Wx + b is built from a matrix, so it must be a linear map.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Then T(u+v) = W(u+v) + b = Wu + Wv + b, but T(u)+T(v) = Wu + Wv + 2b.
Call it affine. It is linear only in the special case b = 0.
Why: Then T(u+v) = W(u+v) + b = Wu + Wv + b, but T(u)+T(v) = Wu + Wv + 2b. They differ by b. With b ≠ 0 additivity FAILS — the map is affine, not linear.
Trap
T(x) = Wx + b is built from a matrix, so it must be a linear map.
Assume T(u + v) = T(u) + T(v)
Why: Then T(u+v) = W(u+v) + b = Wu + Wv + b, but T(u)+T(v) = Wu + Wv + 2b. They differ by b. With b ≠ 0 additivity FAILS — the map is affine, not linear.
Call it affine. It is linear only in the special case b = 0.
Affine = (linear part Wx) + (translation b)
Why: Linearity requires T(0) = 0; here T(0) = b. Reserve 'linear' for b = 0. In ML the bias is almost always nonzero, so a 'linear layer' nn.Linear(with bias) is really an affine layer.
Section
Part 2 of 6 — transforms & invariants
Concept
Rotation, reflection, scaling, and shear are not separate machinery — each is just a particular matrix W (with b = 0, or a translation added):
| transform | what W does | signature |
|---|---|---|
| rotation | turns space about the origin | orthogonal W, det = +1 |
| reflection | flips across a line | orthogonal W, det = −1 |
| scaling | stretches along axes | diagonal W |
| shear | slides one axis along another | off-diagonal entries ≠ 0 |
Our running W = [[2,0],[1,3]] is a scale-and-shear: the diagonal scales, the lower-left 1 shears.
Pattern
Predict first
The table runs: W (rounded) | [[0, −1], [1, 0]] · det(W) | 1.0 (rotation)
In A rotation matrix, checked, given the rows so far: what is the next one — the row where quantity is W @ [1, 0]?
Correct: W @ [1, 0] | [0.0, 1.0]
| quantity | value (verified) |
|---|---|
| W (rounded) | [[0, −1], [1, 0]] |
| det(W) | 1.0 (rotation) |
| W @ [1, 0] | [0.0, 1.0] |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. det = +1 confirms a pure rotation (no reflection, no scaling).
Worked example
A 90° rotation should send [1, 0] to [0, 1]. Build the rotation W for θ = 90° and apply it. Runnable standalone:
import numpy as np
theta = np.pi / 2 # 90 degrees
W = np.array([[np.cos(theta), -np.sin(theta)],
[np.sin(theta), np.cos(theta)]])
e1 = np.array([1., 0.])
print('W =', np.round(W, 3).tolist())
print('det(W) =', round(np.linalg.det(W), 6))
print('W @ [1,0] =', np.round(W @ e1, 6).tolist())det(W) = 1.0, and W·[1,0] = [0, 1]
Why: det = +1 confirms a pure rotation (no reflection, no scaling). The x-axis unit vector lands on the y-axis — a 90° turn, exactly as expected.
| quantity | value (verified) |
|---|---|
| W (rounded) | [[0, −1], [1, 0]] |
| det(W) | 1.0 (rotation) |
| W @ [1, 0] | [0.0, 1.0] |
Concept
For an orthogonal W, the sign of the determinant distinguishes a rotation from a reflection: +1 keeps orientation (rotation), −1 flips it (reflection). A reflection across the x-axis negates the second coordinate:
\[ W_{\text{reflect}} = \begin{bmatrix} 1 & 0 \\ 0 & -1 \end{bmatrix}, \qquad \det W_{\text{reflect}} = (1)(-1) - 0 = -1 \]
Apply it: W_reflect · [2, 3] = [2, −3]
Why: The x-coordinate is untouched, the y-coordinate flips sign — a mirror across the x-axis. det = −1 flags exactly this orientation reversal.
Concept
Because an affine map only stretches and slides — never bends — it keeps three things intact:
What it cannot do: bend a straight line into a curve, or make two parallel lines cross. That inability is precisely why one affine map — however large W — can never carve XOR's boundary.
Fill the middle
Fill in the blanks
From Verify: midpoints and parallels survive — one line has had its right-hand side removed. Put it back.
import numpy as np
W = np.array([[2., 1.], [0., 1.]]); b = np.array([1., -2.])
def T(p): return W @ p + b
P = np.array([0., 0.]); Q = np.array([1., 0.]); R = np.array([0., 1.])
mid = (P + Q) / 2
print('T(mid PQ) =', T(mid).tolist())
print('mid of T(P),T(Q) =', ((T(P) + T(Q)) / 2).tolist())
e1 = T(Q) - T(P)
e2 = T(R + (Q - P)) - T(R)
print('edge1 image =', e1.tolist(), ' edge2 image =', e2.tolist())
Why: e1 is what everything below it consumes, so the wrong expression here fails later and somewhere else. Both print [2, -2]: the affine map takes the midpoint of PQ to the midpoint of T(P),T(Q).
Worked example
Take a shear-scale-shift map. Send the midpoint of an edge through it, and separately send two parallel edges through it, and check both invariants. Runnable:
import numpy as np
W = np.array([[2., 1.], [0., 1.]]); b = np.array([1., -2.])
def T(p): return W @ p + b
P = np.array([0., 0.]); Q = np.array([1., 0.]); R = np.array([0., 1.])
mid = (P + Q) / 2
print('T(mid PQ) =', T(mid).tolist())
print('mid of T(P),T(Q) =', ((T(P) + T(Q)) / 2).tolist())
e1 = T(Q) - T(P)
e2 = T(R + (Q - P)) - T(R)
print('edge1 image =', e1.tolist(), ' edge2 image =', e2.tolist())T(midpoint) equals the midpoint of the images
Why: Both print [2, -2]: the affine map takes the midpoint of PQ to the midpoint of T(P),T(Q). Ratios along a line are preserved.
Two parallel edges map to identical direction vectors
Why: edge1 and edge2 both map to [2, 0] — still parallel. An affine map cannot make parallel lines converge.
| invariant | before → after | verified |
|---|---|---|
| T(P) | [0,0] → [1, −2] | — |
| T(Q) | [1,0] → [3, −2] | — |
| midpoint of PQ | → [2, −2] | = mid of images ✓ |
| two parallel edges | → [2, 0] and [2, 0] | still parallel ✓ |
Discrimination
Sort into buckets
Sort these by verified, from memory, without looking back at Verify: midpoints and parallels survive. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Section
Part 3 of 6 — every step
Intuition
A feed-forward network is a chain: the output of layer 1 is the input to layer 2, whose output feeds layer 3. Mathematically that is function composition — layer₃(layer₂(layer₁(x))).
So the question 'what does a whole linear network compute?' is really 'what is the composition of these affine maps?' Answer that once, for two maps, and induction handles any depth. That is why the next few slides matter far beyond a single example.
Concept
A network applies maps one after another. Suppose T₁(x) = W₁x + b₁ runs first, then T₂(u) = W₂u + b₂ runs on its output. What is the combined map T₂(T₁(x))?
\[ T_1(x) = W_1 x + b_1, \qquad T_2(u) = W_2 u + b_2 \]
We will substitute and simplify — and the punchline is that the result is again a single Wx + b.
Ranking
Put in order
Put the moves of Derive the composition, one move per beat into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. The output of the first map becomes the input of the second: replace u by W₁x + b₁.
Worked example
Substitute u = T₁(x) into T₂
Why: The output of the first map becomes the input of the second: replace u by W₁x + b₁.
\[ T_2(T_1(x)) = W_2\,(W_1 x + b_1) + b_2 \]
Distribute W₂ across the parentheses
Why: Matrix multiplication distributes over addition: W₂(W₁x + b₁) = W₂W₁x + W₂b₁.
\[ = W_2 W_1 x + W_2 b_1 + b_2 \]
Group the x-term and the constant terms
Why: Everything multiplying x is (W₂W₁); everything without x is (W₂b₁ + b₂). That is the Wx + b shape.
\[ = \underbrace{(W_2 W_1)}_{W_{\text{eff}}}\,x + \underbrace{(W_2 b_1 + b_2)}_{b_{\text{eff}}} \]
The composition IS affine, with new W and b
Why: Combined weight W₂W₁, combined bias W₂b₁ + b₂. Notice b₁ is first transformed by W₂ — the biases do NOT simply add.
\[ \boxed{\,T_2\!\circ\! T_1 \,:\; x \mapsto (W_2 W_1)\,x + (W_2 b_1 + b_2)\,} \]
Translation
\( = \underbrace{(W_2 W_1)}_{W_{\text{eff}}}\,x + \underbrace{(W_2 b_1 + b_2)}_{b_{\text{eff}}} \)
Draw it
Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.
Concept
To make the formula bite, fix a second map T₂ and reuse T₁ from Part 1. We will run the two-step computation and the single-map formula and check they agree.
\[ W_1 = \begin{bmatrix} 2 & 0 \\ 1 & 3 \end{bmatrix},\; b_1 = \begin{bmatrix} 1 \\ -1 \end{bmatrix}, \qquad W_2 = \begin{bmatrix} 1 & -1 \\ 0 & 2 \end{bmatrix},\; b_2 = \begin{bmatrix} 2 \\ 0 \end{bmatrix} \]
Same input x = [3, 4]. From Part 1 we already have T₁(x) = [7, 14].
Step zero
Discussion prompt
Two-step path: T₂(T₁(x)) by hand — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: W₂·[7,14]: row 0 = 1·7 + (−1)·14 = 7 − 14 = −7
Answer:
Worked example
Run the maps in sequence. T₁(x) = [7, 14]; now push that through T₂.
W₂·[7,14]: row 0 = 1·7 + (−1)·14 = 7 − 14 = −7
Why: First row [1, −1] of W₂ dotted with u = [7, 14].
W₂·[7,14]: row 1 = 0·7 + 2·14 = 28
Why: Second row [0, 2] dotted with [7, 14]. So W₂u = [−7, 28].
\[ W_2\,T_1(x) = \begin{bmatrix} -7 \\ 28 \end{bmatrix} \]
Add b₂ = [2, 0]: T₂(T₁(x)) = [−7+2, 28+0] = [−5, 28]
Why: The final output of running both maps in sequence.
\[ T_2(T_1(x)) = \begin{bmatrix} -5 \\ 28 \end{bmatrix} \]
Blank canvas
Draw it
Draw what Two-step path: T₂(T₁(x)) by hand just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Ranking
Put in order
Put the moves of One-step path: the combined map by hand into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Row 0: [1,−1]·cols of W₁ = [2−1, 0−3] = [1, −3].
Worked example
Now build W_eff = W₂W₁ and b_eff = W₂b₁ + b₂ directly, then apply to x. Must land on [−5, 28].
W₂W₁: compute the (2×2) product
Why: Row 0: [1,−1]·cols of W₁ = [2−1, 0−3] = [1, −3]. Row 1: [0,2]·cols = [0+2, 0+6] = [2, 6].
\[ W_2 W_1 = \begin{bmatrix} 1 & -1 \\ 0 & 2 \end{bmatrix}\!\begin{bmatrix} 2 & 0 \\ 1 & 3 \end{bmatrix} = \begin{bmatrix} 1 & -3 \\ 2 & 6 \end{bmatrix} \]
b_eff = W₂b₁ + b₂ = [2,−2] + [2,0] = [4, −2]
Why: W₂b₁ = [1·1+(−1)(−1), 0·1+2·(−1)] = [2, −2]; add b₂ = [2, 0].
\[ W_2 b_1 + b_2 = \begin{bmatrix} 2 \\ -2 \end{bmatrix} + \begin{bmatrix} 2 \\ 0 \end{bmatrix} = \begin{bmatrix} 4 \\ -2 \end{bmatrix} \]
Apply: (W₂W₁)x + b_eff = [1·3−3·4, 2·3+6·4] + [4,−2]
Why: [3−12, 6+24] + [4,−2] = [−9, 30] + [4, −2] = [−5, 28]. Identical to the two-step path — composition really is one affine map.
\[ \begin{bmatrix} 1 & -3 \\ 2 & 6 \end{bmatrix}\!\begin{bmatrix} 3 \\ 4 \end{bmatrix} + \begin{bmatrix} 4 \\ -2 \end{bmatrix} = \begin{bmatrix} -9 \\ 30 \end{bmatrix} + \begin{bmatrix} 4 \\ -2 \end{bmatrix} = \begin{bmatrix} -5 \\ 28 \end{bmatrix} \]
Notation
Annotate
From One-step path: the combined map by hand — read this one piece at a time. What is each part doing?
On: \( W_2 b_1 + b_2 = \begin{bmatrix} 2 \\ -2 \end{bmatrix} + \begin{bmatrix} 2 \\ 0 \end{bmatrix} = \begin{bmatrix} 4 \\ -2 \end{bmatrix} \)
Faded example
Fill in the blanks
Confirm both paths in code, with the scaffolding fading: two lines are gone now — fill both.
import numpy as np
W1 = np.array([[2., 0.], [1., 3.]]); b1 = np.array([1., -1.])
W2 = np.array([[1., -1.], [0., 2.]]); b2 = np.array([2., 0.])
x = np.array([3., 4.])
direct = W2 @ (W1 @ x + b1) + b2
composed = (W2 @ W1) @ x + (W2 @ b1 + b2)
print('W2@W1 =', (W2 @ W1).tolist())
print('W2@b1+b2 =', (W2 @ b1 + b2).tolist())
print('direct =', direct.tolist())
print('composed =', composed.tolist())
print('match:', np.allclose(direct, composed))
Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. The sequential two-step computation and the single combined map agree exactly.
Worked example
Compute the two-step and the one-step results and compare with np.allclose. Predict True and [−5, 28]. Runnable standalone:
import numpy as np
W1 = np.array([[2., 0.], [1., 3.]]); b1 = np.array([1., -1.])
W2 = np.array([[1., -1.], [0., 2.]]); b2 = np.array([2., 0.])
x = np.array([3., 4.])
direct = W2 @ (W1 @ x + b1) + b2
composed = (W2 @ W1) @ x + (W2 @ b1 + b2)
print('W2@W1 =', (W2 @ W1).tolist())
print('W2@b1+b2 =', (W2 @ b1 + b2).tolist())
print('direct =', direct.tolist())
print('composed =', composed.tolist())
print('match:', np.allclose(direct, composed))direct == composed == [−5, 28], match = True
Why: The sequential two-step computation and the single combined map agree exactly. Composition of affine maps is affine — confirmed against the hand math.
| computation | result (verified) |
|---|---|
| W₂W₁ | [[1, −3], [2, 6]] |
| W₂b₁ + b₂ | [4, −2] |
| T₂(T₁(x)) directly | [−5.0, 28.0] |
| (W₂W₁)x + (W₂b₁+b₂) | [−5.0, 28.0] |
| np.allclose(...) | True |
Anomaly
Predict first
A student writes this, and it looks reasonable:
Composing T₁ then T₂, the combined bias is surely b₁ + b₂.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: This ignores that b₁ is produced INSIDE T₁ and then passes through W₂.
b₁ is transformed by W₂ before b₂ is added.
Why: This ignores that b₁ is produced INSIDE T₁ and then passes through W₂. Plugging [3,−1] into the check gives the wrong output — it disagrees with the [−5, 28] both honest paths produced.
Trap
Composing T₁ then T₂, the combined bias is surely b₁ + b₂.
b_eff = b₁ + b₂ = [1,−1] + [2,0] = [3, −1]
Why: This ignores that b₁ is produced INSIDE T₁ and then passes through W₂. Plugging [3,−1] into the check gives the wrong output — it disagrees with the [−5, 28] both honest paths produced.
b₁ is transformed by W₂ before b₂ is added.
b_eff = W₂b₁ + b₂ = [2,−2] + [2,0] = [4, −2]
Why: Trace the algebra: W₂(W₁x + b₁) + b₂ = W₂W₁x + W₂b₁ + b₂. The first bias rides through the second matrix. Order matters — W₂ on the left, and it hits b₁ too.
Break the constraint
Discussion prompt
The rule this trap just fixed:
Trace the algebra: W₂(W₁x + b₁) + b₂ = W₂W₁x + W₂b₁ + b₂. The first bias rides through the second matrix. Order matters — W₂ on the left, and it hits b₁ too.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
This ignores that b₁ is produced INSIDE T₁ and then passes through W₂. Plugging [3,−1] into the check gives the wrong output — it disagrees with the [−5, 28] both honest paths produced.
Section
Part 4 of 6 — the packaging trick
Concept
Carrying W and b separately is a nuisance. The trick: append a 1 to x, and build a bigger matrix M whose last column is b. Then a single matrix multiply does Wx + b.
\[ M = \begin{bmatrix} W & b \\ 0 & 1 \end{bmatrix}, \qquad \tilde{x} = \begin{bmatrix} x \\ 1 \end{bmatrix} \;\Longrightarrow\; M\tilde{x} = \begin{bmatrix} Wx + b \\ 1 \end{bmatrix} \]
The extra 1 rides along untouched (bottom row [0, 1]), so you can chain these M's freely. This is how graphics pipelines and transformer position code pack affine maps.
Worked example
Using our W, b, x, build the 3×3 homogeneous M, append a 1 to x, and confirm the top two outputs equal Wx + b = [7, 14]. Runnable:
import numpy as np
W = np.array([[2., 0.], [1., 3.]]); b = np.array([1., -1.])
x = np.array([3., 4.])
M = np.block([[W, b.reshape(2, 1)], [np.zeros((1, 2)), np.ones((1, 1))]])
xh = np.array([3., 4., 1.]) # append a 1
yh = M @ xh
print('M =', M.tolist())
print('M @ [x; 1] =', yh.tolist())
print('Wx + b =', (W @ x + b).tolist())M is 3×3 with b in the last column, [0,0,1] on the bottom
Why: np.block stacks W (2×2) beside b (2×1), then a bottom row [0, 0, 1]. The last row keeps the trailing 1 intact.
M·[x;1] = [7, 14, 1] → top two match Wx + b = [7, 14]
Why: One matrix multiply reproduced the affine map, and the appended 1 came back as 1 — ready to feed into the next M.
| object | value (verified) |
|---|---|
| M row 0 | [2, 0, 1] |
| M row 1 | [1, 3, −1] |
| M row 2 | [0, 0, 1] |
| M @ [3, 4, 1] | [7.0, 14.0, 1.0] |
| Wx + b | [7.0, 14.0] |
Error analysis
Annotate
Walk the callouts on Build M and check it reproduces Wx + b. Each one is a place this is easy to get subtly wrong.
Estimation
Predict first
The payoff: in homogeneous form, composing two affine maps is just M₂ M₁ — one matrix product, no separate bias bookkeeping. Build both and read off the blocks:
Commit before you compute: what does Composition becomes a plain matrix product come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: M₂M₁ last column (top) = W₂b₁+b₂ = [4,−2]
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. And the combined bias appears in the last column — exactly [4, −2].
Worked example
The payoff: in homogeneous form, composing two affine maps is just M₂ M₁ — one matrix product, no separate bias bookkeeping. Build both and read off the blocks:
import numpy as np
W1 = np.array([[2., 0.], [1., 3.]]); b1 = np.array([1., -1.])
W2 = np.array([[1., -1.], [0., 2.]]); b2 = np.array([2., 0.])
def homo(W, b):
return np.block([[W, b.reshape(2, 1)], [np.zeros((1, 2)), np.ones((1, 1))]])
M = homo(W2, b2) @ homo(W1, b1)
print('M2 @ M1 =', M.tolist())
print('top-left 2x2 (= W2W1) =', M[:2, :2].tolist())
print('last col top (= W2b1+b2)=', M[:2, 2].tolist())M₂M₁ top-left block = W₂W₁ = [[1,−3],[2,6]]
Why: The matrix product automatically produces the combined weight in the upper-left 2×2 — the same [[1,−3],[2,6]] we derived by hand.
M₂M₁ last column (top) = W₂b₁+b₂ = [4,−2]
Why: And the combined bias appears in the last column — exactly [4, −2]. The homogeneous product encodes both the weight and the transformed-bias rule in one step.
| block of M₂M₁ | value (verified) | matches |
|---|---|---|
| top-left 2×2 | [[1, −3], [2, 6]] | W₂W₁ ✓ |
| last column (top 2) | [4, −2] | W₂b₁ + b₂ ✓ |
| bottom row | [0, 0, 1] | preserved ✓ |
Concept
The Wx + b skeleton is not a toy — it is the pre-activation of every dense layer, and it recurs across the whole USAAIO syllabus:
nn.Linear(in, out) is literally x ↦ Wx + bWM₂M₁ trickEvery one of these inherits the collapse problem we are about to prove — which is why every one is paired with a nonlinearity.
Section
Part 5 of 6 — why linear depth is free
Concept
A neural network with no activation functions is just a chain of affine maps. By the composition rule, that chain is itself a single affine map:
\[ A_3\big(A_2(A_1 x)\big) = (A_3 A_2 A_1)\,x \]
So a 3-layer linear network has exactly the expressive power of a 1-layer one. Depth without nonlinearity buys nothing — this is the central fact of the lesson.
Intuition
Each linear layer multiplies its input by a matrix. Chain them and the matrices multiply into one matrix — and any single matrix is reachable by a single layer.
It is like composing the functions f(x) = 3x and g(x) = 5x: g(f(x)) = 15x is still just multiply-by-a-constant. You never escape the family of straight-through-the-origin maps, no matter how many you stack.
Fill the middle
Fill in the blanks
From Three linear layers = one matrix — one line has had its right-hand side removed. Put it back.
import numpy as np
np.random.seed(0)
A1 = np.random.randn(4, 3)
A2 = np.random.randn(5, 4)
A3 = np.random.randn(2, 5)
x = np.random.randn(3)
stacked = A3 @ (A2 @ (A1 @ x))
W_eff = A3 @ A2 @ A1
single = W_eff @ x
print('stacked =', np.round(stacked, 6).tolist())
print('single =', np.round(single, 6).tolist())
print('effective shape:', W_eff.shape)
print('match:', np.allclose(stacked, single))
Why: x is what everything below it consumes, so the wrong expression here fails later and somewhere else. Running the three layers in sequence gives the identical output to multiplying by the single collapsed matrix.
Worked example
Chain three random linear layers (3→4→5→2) and compare running them in sequence against multiplying their matrices into one. Seeded, so your numbers match. Runnable:
import numpy as np
np.random.seed(0)
A1 = np.random.randn(4, 3)
A2 = np.random.randn(5, 4)
A3 = np.random.randn(2, 5)
x = np.random.randn(3)
stacked = A3 @ (A2 @ (A1 @ x))
W_eff = A3 @ A2 @ A1
single = W_eff @ x
print('stacked =', np.round(stacked, 6).tolist())
print('single =', np.round(single, 6).tolist())
print('effective shape:', W_eff.shape)
print('match:', np.allclose(stacked, single))stacked == single == [1.562102, 8.741537]
Why: Running the three layers in sequence gives the identical output to multiplying by the single collapsed matrix. The intermediate widths 4 and 5 vanish.
The effective map is a single (2×3) matrix
Why: 3 inputs, 2 outputs — a 2×3 matrix. The whole 3→4→5→2 network has exactly the capacity of one 2×3 linear layer.
| network | output | effective map |
|---|---|---|
| 3 layers stacked (3→4→5→2) | [1.562102, 8.741537] | — |
| single collapsed matrix | [1.562102, 8.741537] | one 2×3 matrix |
| np.allclose | True | identical |
Faded example
Fill in the blanks
With biases too, the affine stack collapses, with the scaffolding fading: two lines are gone now — fill both.
import numpy as np
np.random.seed(0)
A1 = np.random.randn(4, 3); c1 = np.random.randn(4)
A2 = np.random.randn(5, 4); c2 = np.random.randn(5)
A3 = np.random.randn(2, 5); c3 = np.random.randn(2)
x = np.random.randn(3)
stacked = A3 @ (A2 @ (A1 @ x + c1) + c2) + c3
W_eff = A3 @ A2 @ A1
b_eff = A3 @ (A2 @ c1 + c2) + c3
single = W_eff @ x + b_eff
print('W_eff shape:', W_eff.shape, ' b_eff shape:', b_eff.shape)
print('match:', np.allclose(stacked, single))
Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. The composition formula (W₂b₁ + b₂) generalizes: each earlier bias passes through all later matrices.
Worked example
Adding a bias to each layer changes nothing about the conclusion — the stack is still one affine map W_eff x + b_eff. Verify with three biased layers:
import numpy as np
np.random.seed(0)
A1 = np.random.randn(4, 3); c1 = np.random.randn(4)
A2 = np.random.randn(5, 4); c2 = np.random.randn(5)
A3 = np.random.randn(2, 5); c3 = np.random.randn(2)
x = np.random.randn(3)
stacked = A3 @ (A2 @ (A1 @ x + c1) + c2) + c3
W_eff = A3 @ A2 @ A1
b_eff = A3 @ (A2 @ c1 + c2) + c3
single = W_eff @ x + b_eff
print('W_eff shape:', W_eff.shape, ' b_eff shape:', b_eff.shape)
print('match:', np.allclose(stacked, single))b_eff = A₃(A₂c₁ + c₂) + c₃ — the nested-bias rule
Why: The composition formula (W₂b₁ + b₂) generalizes: each earlier bias passes through all later matrices. c₁ is hit by A₂ then A₃, c₂ by A₃, c₃ raw.
stacked == single → still one affine map
Why: The biased 3-layer stack equals a single W_eff x + b_eff. Biases do not rescue linear depth — the collapse is total.
| quantity | value (verified) |
|---|---|
| W_eff shape | (2, 3) |
| b_eff shape | (2,) |
| stacked == single | True |
Comparison
Comparison matrix
From With biases too, the affine stack collapses: refill the value (verified) column from what you know. The rest of the table is as it appeared.
| quantity | value (verified) |
|---|---|
| W_eff shape | (2, 3) |
| b_eff shape | (2,) |
| stacked == single | True |
Anomaly
Predict first
A student writes this, and it looks reasonable:
The model underfits — stack more linear layers to make it more expressive.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Every extra linear layer just multiplies into the running matrix product.
Insert a nonlinearity between layers — that is what makes depth expressive.
Why: Every extra linear layer just multiplies into the running matrix product. 3 layers collapsed to one 2×3 matrix; 10 layers collapse to one matrix too. Capacity gained: zero.
Trap
The model underfits — stack more linear layers to make it more expressive.
Add depth with no activations to boost capacity
Why: Every extra linear layer just multiplies into the running matrix product. 3 layers collapsed to one 2×3 matrix; 10 layers collapse to one matrix too. Capacity gained: zero.
Insert a nonlinearity between layers — that is what makes depth expressive.
Use σ(W₂ σ(W₁x + b₁) + b₂)
Why: The activation σ sits between the affine maps and blocks the collapse: σ(W₂σ(W₁x+b₁)+b₂) is not any single Wx+b. Now each added layer contributes genuinely new structure.
Section
Part 6 of 6 — XOR settles it
Concept
Put an activation σ (ReLU, tanh, sigmoid, …) between the affine maps, and the composition no longer simplifies to a single Wx + b:
\[ \sigma\!\big(W_2\,\sigma(W_1 x + b_1) + b_2\big) \;\neq\; \text{any single } Wx + b \]
The reason: σ is not linear, so it does not distribute through the matrix products. Now the network can bend space, carve curved decision boundaries, and approximate any function — the universal-approximation property.
Intuition
Think of relu(z) = max(0, z) as a switch: for each hidden unit it either passes the value through or clamps it to zero. Which units switch off depends on the input.
So different inputs travel through different effective matrices — the net is a patchwork of many linear pieces stitched along fold lines. That is why doubling the input does not double the output, and why a whole square of inputs can fold onto itself. One affine map has no folds; a hidden ReLU layer has as many as it needs.
Explain it
Discussion prompt
Explain What a ReLU actually breaks to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Think of relu(z) = max(0, z) as a switch: for each hidden unit it either passes the value through or clamps it to zero. Which units switch off depends on the input.
Fill the middle
Fill in the blanks
From Prove a ReLU net is not linear — one line has had its right-hand side removed. Put it back.
import numpy as np
np.random.seed(0)
A1 = np.random.randn(4, 2)
A2 = np.random.randn(1, 4)
relu = lambda z: np.maximum(0.0, z)
def net(x): return A2 @ relu(A1 @ x) # one hidden ReLU layer
a = np.array([1.0, 1.0]); b = np.array([-1.0, 2.0])
print('net(a) =', round(net(a)[0], 4))
print('net(b) =', round(net(b)[0], 4))
print('net(a)+net(b) =', round(net(a)[0] + net(b)[0], 4))
print('net(a+b) =', round(net(a + b)[0], 4))
print('additive?', np.allclose(net(a + b), net(a) + net(b)))
Why: A2 is what everything below it consumes, so the wrong expression here fails later and somewhere else. Additivity fails by more than a full unit.
Worked example
A one-hidden-layer ReLU net must fail an additivity test net(a+b) = net(a) + net(b) for some inputs. Pick a = [1,1], b = [−1,2] and check. Seeded, runnable:
import numpy as np
np.random.seed(0)
A1 = np.random.randn(4, 2)
A2 = np.random.randn(1, 4)
relu = lambda z: np.maximum(0.0, z)
def net(x): return A2 @ relu(A1 @ x) # one hidden ReLU layer
a = np.array([1.0, 1.0]); b = np.array([-1.0, 2.0])
print('net(a) =', round(net(a)[0], 4))
print('net(b) =', round(net(b)[0], 4))
print('net(a)+net(b) =', round(net(a)[0] + net(b)[0], 4))
print('net(a+b) =', round(net(a + b)[0], 4))
print('additive?', np.allclose(net(a + b), net(a) + net(b)))net(a)+net(b) = 3.8267 but net(a+b) = 2.6364
Why: Additivity fails by more than a full unit. The ReLU zeros out different coordinates for different inputs, so the map cannot be written as one matrix.
additive? → False → the net is genuinely nonlinear
Why: Because it breaks a linearity axiom, no single affine map equals this net. THAT broken linearity is the capacity the network gains from σ.
| quantity | value (verified) |
|---|---|
| net(a), a=[1,1] | 2.3884 |
| net(b), b=[−1,2] | 1.4383 |
| net(a) + net(b) | 3.8267 |
| net(a + b), a+b=[0,3] | 2.6364 |
| additive? | False (nonlinear ✓) |
Blank canvas
Draw it
Draw what Prove a ReLU net is not linear just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Intuition
XOR labels the four corners of a square: [0,0] and [1,1] are class 0; [0,1] and [1,0] are class 1. The two classes sit on opposite diagonals.
No single straight line can put both 0-corners on one side and both 1-corners on the other — the classes are not linearly separable. Since an affine map only draws straight boundaries, one layer is doomed. This is the historical example that stalled neural nets for a decade.
Analogy
Discussion prompt
Explain XOR: the smallest problem depth needs by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
XOR labels the four corners of a square: [0,0] and [1,1] are class 0; [0,1] and [1,0] are class 1. The two classes sit on opposite diagonals.
Pattern
Predict first
The table runs: [0, 0] | 0 | 0.5 | 0 · [0, 1] | 1 | 0.5 | 0 ✗ · [1, 0] | 1 | 0.5 | 0 ✗
In The best possible linear fit still fails XOR, given the rows so far: what is the next one — the row where point is [1, 1]?
Correct: [1, 1] | 0 | 0.5 | 0
| point | target | raw output | predicted |
|---|---|---|---|
| [0, 0] | 0 | 0.5 | 0 |
| [0, 1] | 1 | 0.5 | 0 ✗ |
| [1, 0] | 1 | 0.5 | 0 ✗ |
| [1, 1] | 0 | 0.5 | 0 |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. The best line can do no better than predict the mean (0.5) everywhere: the symmetry of XOR forces both slopes to zero.
Worked example
Do not just try one linear model — find the best one. Least squares gives the optimal linear weights; then threshold at 0.5 and count how many of the four points it gets right. Runnable:
import numpy as np
X = np.array([[0., 0.], [0., 1.], [1., 0.], [1., 1.]])
y = np.array([0., 1., 1., 0.])
Xb = np.c_[np.ones(4), X] # bias + 2 features
w, *_ = np.linalg.lstsq(Xb, y, rcond=None) # best least-squares line
pred = Xb @ w
print('best weights [b, w1, w2] =', np.round(w, 4).tolist())
print('raw outputs =', np.round(pred, 4).tolist())
print('thresholded =', (pred > 0.5).astype(int).tolist())
print('accuracy =', ((pred > 0.5).astype(float) == y).mean())Optimal weights = [0.5, 0.0, 0.0] → every prediction is 0.5
Why: The best line can do no better than predict the mean (0.5) everywhere: the symmetry of XOR forces both slopes to zero. It literally cannot tilt to separate the classes.
Thresholded predictions = [0,0,0,0], accuracy = 0.5
Why: It labels everything class 0, matching 2 of 4 points — pure chance. No linear (or stacked-linear) model beats 0.5 on XOR, ever.
| point | target | raw output | predicted |
|---|---|---|---|
| [0, 0] | 0 | 0.5 | 0 |
| [0, 1] | 1 | 0.5 | 0 ✗ |
| [1, 0] | 1 | 0.5 | 0 ✗ |
| [1, 1] | 0 | 0.5 | 0 |
Discrimination
Sort into buckets
Sort these by target, from memory, without looking back at The best possible linear fit still fails XOR. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Missing information
Discussion prompt
Same data, but now a 2-layer MLP with a tanh hidden layer. Train both a bare linear model and the MLP and compare accuracy. Seeded torch, runnable:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Even trained with Adam for 4000 steps, the single linear layer is stuck at chance — matching the least-squares result. Training cannot fix a representational limit.
Worked example
Same data, but now a 2-layer MLP with a tanh hidden layer. Train both a bare linear model and the MLP and compare accuracy. Seeded torch, runnable:
import torch
torch.manual_seed(0)
X = torch.tensor([[0., 0.], [0., 1.], [1., 0.], [1., 1.]])
y = torch.tensor([[0.], [1.], [1.], [0.]])
def train(model, steps=4000, lr=0.05):
opt = torch.optim.Adam(model.parameters(), lr=lr)
loss_fn = torch.nn.BCEWithLogitsLoss()
for _ in range(steps):
opt.zero_grad(); loss_fn(model(X), y).backward(); opt.step()
with torch.no_grad():
pred = (torch.sigmoid(model(X)) > 0.5).float()
return (pred == y).float().mean().item(), pred.int().ravel().tolist()
lin = torch.nn.Linear(2, 1)
mlp = torch.nn.Sequential(torch.nn.Linear(2, 8), torch.nn.Tanh(), torch.nn.Linear(8, 1))
print('linear:', train(lin))
print('MLP :', train(mlp))linear → accuracy 0.5, predictions [0,0,0,0]
Why: Even trained with Adam for 4000 steps, the single linear layer is stuck at chance — matching the least-squares result. Training cannot fix a representational limit.
MLP → accuracy 1.00, predictions [0,1,1,0]
Why: The tanh hidden layer warps the input space so XOR becomes linearly separable in the hidden representation. One nonlinearity is the whole difference.
| model | accuracy (verified) | predictions |
|---|---|---|
| linear nn.Linear(2,1) | 0.50 (chance) | [0, 0, 0, 0] |
| MLP (tanh hidden, 8 units) | 1.00 | [0, 1, 1, 0] |
| targets | — | [0, 1, 1, 0] |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
MLP → accuracy 1.00, predictions [0,1,1,0]
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Same data, but now a 2-layer MLP with a tanh hidden layer. Train both a bare linear model and the MLP and compare accuracy. Seeded torch, runnable:
Concept
You do not need training to see the mechanism. Two hand-picked ReLU units build XOR exactly: an OR feature and an AND feature, combined as OR − 2·AND.
\[ h_{\text{or}} = \operatorname{relu}(x_1 + x_2 - 0.5), \quad h_{\text{and}} = \operatorname{relu}(x_1 + x_2 - 1.5), \quad \text{xor} = h_{\text{or}} - 2\,h_{\text{and}} \]
h_or fires when at least one input is 1; h_and fires only when both are 1. Subtracting twice the AND cancels the [1,1] corner back down — the trick a hidden layer discovers on its own.
Counterexample
Discussion prompt
You do not need training to see the mechanism. Two hand-picked ReLU units build XOR exactly: an OR feature and an AND feature, combined as OR − 2·AND.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
h_or fires when at least one input is 1; h_and fires only when both are 1. Subtracting twice the AND cancels the [1,1] corner back down — the trick a hidden layer discovers on its own.
Estimation
Predict first
Run the two ReLU features across all four points and read off the XOR column. No training — just the formula. Runnable:
Commit before you compute: what does Hand-built XOR from two ReLU features come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: xor column separates the classes (0.0 vs 0.5)
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Corners [0,0]→0.0 and [1,1]→0.5 vs [0,1],[1,0]→0.5...
Worked example
Run the two ReLU features across all four points and read off the XOR column. No training — just the formula. Runnable:
import numpy as np
X = np.array([[0., 0.], [0., 1.], [1., 0.], [1., 1.]])
relu = lambda z: np.maximum(0.0, z)
print('x1 x2 | h_or h_and | xor')
for x1, x2 in X:
h_or = relu(x1 + x2 - 0.5) # fires when >= one input is 1
h_and = relu(x1 + x2 - 1.5) # fires only when both are 1
xor = h_or - 2 * h_and
print(int(x1), int(x2), ' |', round(h_or, 1), round(h_and, 1), ' |', round(xor, 1))The [1,1] corner: h_or=1.5, h_and=0.5 → 1.5 − 2(0.5) = 0.5
Why: Without the AND correction, [1,1] would score 1.5 and be misread as class 1. Subtracting 2·AND pulls it back down to 0.5 — the low class, correct for XOR.
xor column separates the classes (0.0 vs 0.5)
Why: Corners [0,0]→0.0 and [1,1]→0.5 vs [0,1],[1,0]→0.5... the two 'on' corners score 0.5 while [0,0] scores 0.0 — now a single threshold in the HIDDEN space splits them. The nonlinearity made XOR separable.
| x1 | x2 | h_or | h_and | xor = h_or − 2·h_and |
|---|---|---|---|---|
| 0 | 0 | 0.0 | 0.0 | 0.0 |
| 0 | 1 | 0.5 | 0.0 | 0.5 |
| 1 | 0 | 0.5 | 0.0 | 0.5 |
| 1 | 1 | 1.5 | 0.5 | 0.5 |
Error analysis
Annotate
Walk the callouts on Hand-built XOR from two ReLU features. Each one is a place this is easy to get subtly wrong.
Anomaly
Predict first
A student writes this, and it looks reasonable:
The linear model failed — throw more parameters (or more linear layers) at XOR until it fits.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Stacked linear layers collapse to ONE linear map (Part 5), and one linear map draws one straight boundary.
Add a hidden layer with a nonlinearity — even 2 hidden units suffice.
Why: Stacked linear layers collapse to ONE linear map (Part 5), and one linear map draws one straight boundary. XOR needs a non-convex region, so accuracy stays pinned at 0.50 no matter the size.
Trap
The linear model failed — throw more parameters (or more linear layers) at XOR until it fits.
Scale up / stack a linear classifier on XOR
Why: Stacked linear layers collapse to ONE linear map (Part 5), and one linear map draws one straight boundary. XOR needs a non-convex region, so accuracy stays pinned at 0.50 no matter the size.
Add a hidden layer with a nonlinearity — even 2 hidden units suffice.
One hidden layer + activation → 100% on XOR
Why: The nonlinearity warps the input space (OR − 2·AND) so XOR becomes linearly separable in the hidden representation. Capacity comes from σ, not from parameter count.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
x, applies a matrix W, then adds a fixed vector b. The Wx rotates, scales, or shears; the + b slides the whole result.; A map T is linear when it satisfies two axioms for all vectors u, v and scalars c:; A linear map must send 0 to 0 — plug in x = 0 and W·0 = 0. An affine map sends 0 to b: the origin is free to move.T(x) = Wx + b is built from a matrix, so it must be a linear map.; Composing T₁ then T₂, the combined bias is surely b₁ + b₂.Ranking
Put in order
These are the steps of Why activations exist — the whole argument, scrambled. Put them back in order before the next slide shows you.
Wx + b (linear part + shift)(W₂W₁)x + (W₂b₁+b₂) (verified: [−5, 28] two ways)net(a+b) ≠ net(a)+net(b))0.5, tanh MLP at 1.0) — and, by extension, approximate anythingWhy: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
Wx + b (linear part + shift)(W₂W₁)x + (W₂b₁+b₂) (verified: [−5, 28] two ways)net(a+b) ≠ net(a)+net(b))0.5, tanh MLP at 1.0) — and, by extension, approximate anythingEdge cases
Discussion prompt
Why activations exist — the whole argument works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
Wx + b (linear part + shift)(W₂W₁)x + (W₂b₁+b₂) (verified: [−5, 28] two ways)net(a+b) ≠ net(a)+net(b))0.5, tanh MLP at 1.0) — and, by extension, approximate anythingElimination
Eliminate the wrong options
Composing T₁(x)=W₁x+b₁ then T₂(x)=W₂x+b₂ gives a single affine map with weight W and bias b:
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Substitute: T₂(T₁(x)) = W₂(W₁x+b₁)+b₂ = (W₂W₁)x + (W₂b₁+b₂). Because T₁ runs first, W₁ sits on the right of the product, and b₁ is transformed by W₂ before b₂ is added.
Check
Combine two affine maps in order: T₁ first, then T₂.
Check your understanding
Composing T₁(x)=W₁x+b₁ then T₂(x)=W₂x+b₂ gives a single affine map with weight W and bias b:
Answer: A
Why: Substitute: T₂(T₁(x)) = W₂(W₁x+b₁)+b₂ = (W₂W₁)x + (W₂b₁+b₂). Because T₁ runs first, W₁ sits on the right of the product, and b₁ is transformed by W₂ before b₂ is added.
Prediction
Predict first
For which maps T(x) = Wx + b is T actually a LINEAR map (not merely affine)?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Only when b = 0
Why: A linear map must satisfy T(0) = 0, but T(0) = W·0 + b = b. So T is linear exactly when b = 0. With a nonzero bias, additivity fails: T(u+v) = Wu+Wv+b ≠ Wu+Wv+2b = T(u)+T(v).
Check
Recall the linearity axioms and what T(0) must be.
Check your understanding
For which maps T(x) = Wx + b is T actually a LINEAR map (not merely affine)?
Answer: A
Why: A linear map must satisfy T(0) = 0, but T(0) = W·0 + b = b. So T is linear exactly when b = 0. With a nonzero bias, additivity fails: T(u+v) = Wu+Wv+b ≠ Wu+Wv+2b = T(u)+T(v).
Prediction
Predict first
A 5-layer network with NO activation functions has the same expressive power as:
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: a single linear (affine) layer
Why: Composition of affine maps is affine, so the 5 weight matrices multiply into one (and the biases collapse into one b_eff). The network equals a single affine layer — verified when 3 layers collapsed to one 2×3 matrix.
Check
How expressive is a deep network with no activations?
Check your understanding
A 5-layer network with NO activation functions has the same expressive power as:
Answer: A
Why: Composition of affine maps is affine, so the 5 weight matrices multiply into one (and the biases collapse into one b_eff). The network equals a single affine layer — verified when 3 layers collapsed to one 2×3 matrix.
Elimination
Eliminate the wrong options
A single linear classifier cannot solve XOR because…
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: The classes ([0,0],[1,1] vs [0,1],[1,0]) sit on opposite diagonals, so no straight boundary separates them. An affine map draws only straight boundaries, so accuracy is stuck at 0.50 — even the optimal least-squares line predicts 0.5 everywhere. A hidden nonlinearity fixes it.
Check
Why exactly does one linear layer fail here?
Check your understanding
A single linear classifier cannot solve XOR because…
Answer: A
Why: The classes ([0,0],[1,1] vs [0,1],[1,0]) sit on opposite diagonals, so no straight boundary separates them. An affine map draws only straight boundaries, so accuracy is stuck at 0.50 — even the optimal least-squares line predicts 0.5 everywhere. A hidden nonlinearity fixes it.
Section
The project
Concept
Reproduce the three load-bearing facts of the lesson with your own hands: affine composition is affine, a linear stack collapses, and only a nonlinear net solves XOR. You have derived every piece — now assemble them.
| # | requirement | tool |
|---|---|---|
| 1 | Compose two affines; check the (W₂W₁, W₂b₁+b₂) formula | numpy @, np.allclose |
| 2 | Collapse 3 linear layers into one matrix | matrix product |
| 3 | XOR: linear (0.5) vs MLP (1.0) | torch MLP |
Build rules: type every line yourself, fix the seed (np.random.seed(0), torch.manual_seed(0)), and compare with np.allclose — the collapse must be exact, not approximate. When something errors, read the array shapes in the message.
Worked example
Your turn: verify T₂(T₁(x)) equals (W₂W₁)x + (W₂b₁+b₂). Predict out loud: exact match, or only approximate?
Hint: compute the two-step W2@(W1@x+b1)+b2 and the one-step (W2@W1)@x + (W2@b1+b2), then compare with np.allclose.
import numpy as np
W1 = np.array([[2., 0.], [1., 3.]]); b1 = np.array([1., -1.])
W2 = np.array([[1., -1.], [0., 2.]]); b2 = np.array([2., 0.])
x = np.array([3., 4.])
direct = W2 @ (W1 @ x + b1) + b2
composed = (W2 @ W1) @ x + (W2 @ b1 + b2)
print(direct.tolist(), composed.tolist(), np.allclose(direct, composed))| check | expected value |
|---|---|
| direct output | [−5.0, 28.0] |
| composed output | [−5.0, 28.0] |
| np.allclose | True (exact) |
Worked example
Your turn: chain three linear layers (3→4→5→2) and confirm it equals one matrix product. Predict the effective shape before you print it.
Hint: seed first, then compare A3@(A2@(A1@x)) with (A3@A2@A1)@x; the effective W should be (2, 3).
import numpy as np
np.random.seed(0)
A1 = np.random.randn(4, 3); A2 = np.random.randn(5, 4); A3 = np.random.randn(2, 5)
x = np.random.randn(3)
stacked = A3 @ (A2 @ (A1 @ x))
single = (A3 @ A2 @ A1) @ x
print(np.allclose(stacked, single), (A3 @ A2 @ A1).shape)| check | expected value |
|---|---|
| stacked == single | True |
| stacked output | [1.562102, 8.741537] |
| effective W shape | (2, 3) |
Worked example
Your turn: train a bare linear model and a tanh MLP on XOR. Predict each one's accuracy before running.
Hint: nn.Linear(2,1) stalls at 0.5; nn.Sequential(Linear(2,8), Tanh, Linear(8,1)) reaches 1.0 within a few thousand Adam steps. Seed with torch.manual_seed(0).
import torch
torch.manual_seed(0)
X = torch.tensor([[0., 0.], [0., 1.], [1., 0.], [1., 1.]])
y = torch.tensor([[0.], [1.], [1.], [0.]])
def train(m, steps=4000, lr=0.05):
opt = torch.optim.Adam(m.parameters(), lr=lr)
lf = torch.nn.BCEWithLogitsLoss()
for _ in range(steps):
opt.zero_grad(); lf(m(X), y).backward(); opt.step()
with torch.no_grad():
return ((torch.sigmoid(m(X)) > 0.5).float() == y).float().mean().item()
print(train(torch.nn.Linear(2, 1)))
print(train(torch.nn.Sequential(torch.nn.Linear(2, 8), torch.nn.Tanh(), torch.nn.Linear(8, 1))))| model | expected accuracy |
|---|---|
| linear nn.Linear(2,1) | 0.50 |
| MLP (tanh hidden) | 1.00 |
Trade off
Comparison matrix
From Milestone 3 — XOR, linear vs MLP: every row here is a choice with a cost. Fill the expected accuracy column, then say which row you would actually pick and what you give up for it.
| model | expected accuracy |
|---|---|
| linear nn.Linear(2,1) | 0.50 |
| MLP (tanh hidden) | 1.00 |
Concept
import numpy as np, torch
# 1. composition of affines is affine
W1 = np.array([[2., 0.], [1., 3.]]); b1 = np.array([1., -1.])
W2 = np.array([[1., -1.], [0., 2.]]); b2 = np.array([2., 0.]); x = np.array([3., 4.])
print('composition affine:', np.allclose(W2@(W1@x+b1)+b2, (W2@W1)@x + (W2@b1+b2)))
# 2. linear depth collapses
np.random.seed(0)
A1 = np.random.randn(4, 3); A2 = np.random.randn(5, 4); A3 = np.random.randn(2, 5)
v = np.random.randn(3)
print('linear collapse:', np.allclose(A3@(A2@(A1@v)), (A3@A2@A1)@v))
# 3. XOR: linear fails, MLP solves
torch.manual_seed(0)
Xt = torch.tensor([[0., 0.], [0., 1.], [1., 0.], [1., 1.]])
yt = torch.tensor([[0.], [1.], [1.], [0.]])
def train(m, steps=4000, lr=0.05):
opt = torch.optim.Adam(m.parameters(), lr=lr); lf = torch.nn.BCEWithLogitsLoss()
for _ in range(steps):
opt.zero_grad(); lf(m(Xt), yt).backward(); opt.step()
with torch.no_grad():
return ((torch.sigmoid(m(Xt)) > 0.5).float() == yt).float().mean().item()
la = train(torch.nn.Linear(2, 1))
ma = train(torch.nn.Sequential(torch.nn.Linear(2, 8), torch.nn.Tanh(), torch.nn.Linear(8, 1)))
print('XOR linear / MLP:', round(la, 2), '/', round(ma, 2))| printed line | value (verified) |
|---|---|
| composition affine: | True |
| linear collapse: | True |
| XOR linear / MLP: | 0.5 / 1.0 |
If composition prints True, the three layers collapse to True, and only the MLP reaches 1.0 on XOR — you know exactly why activations exist.
Comparison
Comparison matrix
From The full program: refill the value (verified) column from what you know. The rest of the table is as it appeared.
| printed line | value (verified) |
|---|---|
| composition affine: | True |
| linear collapse: | True |
| XOR linear / MLP: | 0.5 / 1.0 |
Concept
Slides closed, out loud: explain (1) why composing affines keeps the Wx + b form and why b₁ gets multiplied by W₂, (2) why that collapses a linear network to one layer, and (3) what the nonlinearity does for XOR geometrically.
Stretch: prove by induction that any depth-N linear network equals a single layer, and swap the tanh MLP's hidden width from 8 down to 2 — confirm it still hits 1.0. Affine maps return as attention's QKV projections (Week 27) and positional encodings (Week 28), so this Wx + b skeleton is everywhere ahead.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — The map: Wx + b · The geometry of W · Composition stays affine · One matrix: homogeneous coordinates · The collapse · Why nonlinearity. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
Wx + b, split it into linear part and shift, and show it is linear only when b = 0W and name the affine invariants — lines, parallelism, ratios (verified: midpoints and parallels survive)(W₂W₁)x + (W₂b₁+b₂), and pack it as one matrix in homogeneous coordinates0.5, a tanh MLP at 1.0, and a hand-built 2-ReLU XOR that shows how| idea | the one thing to remember |
|---|---|
| affine map | Wx + b: linear part + shift; linear only if b = 0 |
| composition | affine ∘ affine = affine — b₁ rides through W₂ |
| homogeneous | append a 1; composition becomes M₂M₁ |
| linear depth | collapses to one matrix → adds no capacity |
| nonlinearity | breaks the collapse — XOR needs it (0.5 → 1.0) |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.