Lesson 22: Affine Transforms

USAAIO Lesson 22, from Week 8, fully worked. It defines the affine map T(x)=Wx+b and separates it into its linear and translation parts, tests the linearity axioms one at a time, and reads the geometric transforms - rotation, reflection, scaling, and shear - off W. It proves the affine invariants for lines, parallelism, and ratios, then derives the composition T2 of T1 move by move to (W2W1)x + (W2b1 + b2), and gives the homogeneous-coordinate trick that turns Wx+b into a single matrix. It then shows a deep linear stack collapsing into one matrix, three layers at a time, and why a ReLU breaks that collapse, since additivity fails on real numbers. It settles XOR once and for all: the best linear fit is stuck at 0.5 while a tanh MLP reaches 1.0, alongside a hand-built two-ReLU XOR. Every snippet runs standalone, and every number came from real execution with numpy 2.2.6 and torch 2.7.1, seeded. The lesson runs to 61 slides.

Subject: Machine Learning · 114 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Affine Transforms

Title

USAAIO · Lesson 22 · Week 8

Why a deep stack of linear layers is secretly just one layer — and why that single fact makes nonlinearity non-negotiable. We build Wx + b from its two halves, prove composition stays affine one move at a time, watch three layers collapse into one matrix, and settle XOR with real numbers.

2. By the end of this lesson you can

Objectives

  1. Define an affine map T(x) = Wx + b and split it into its linear part Wx and translation + b
  2. State the two linearity axioms and show why + b breaks them (an affine map is linear only when b = 0)
  3. Read the classic geometric transforms — rotation, reflection, scaling, shear — off the matrix W, and name what affine maps preserve (lines, parallelism, ratios)
  4. Derive composition T₂∘T₁ step by step to (W₂W₁)x + (W₂b₁ + b₂) — and repackage it as one matrix in homogeneous coordinates
  5. Prove a stack of linear layers collapses to a single matrix, explain why a nonlinearity breaks that collapse, and demonstrate it on XOR: a best linear fit stuck at 0.5, a tanh MLP at 1.0

3. What survived from Second-Order Methods?

Warm-up

Discussion prompt

Before we open Lesson 22: Affine Transforms: without looking back, what was the main idea of Second-Order Methods, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

the Hessian and what its definiteness says about minima vs saddles, Newton's method and its quadratic convergence, why second-order methods cost O(d³), and the L-BFGS / natural-gradient approximations. Implement Newton for logistic regression and race it against gradient descent.

4. The map: Wx + b

Section

Part 1 of 6 — what an affine map is

5. A linear part plus a shift

Concept

An affine map takes a vector x, applies a matrix W, then adds a fixed vector b. The Wx rotates, scales, or shears; the + b slides the whole result.

\[ T(x) = Wx + b \]

This is exactly one layer of a neural network before its activation — the nn.Linear step. Master this object and you understand every layer's skeleton.

6. Break it if you can: A linear part plus a shift

Counterexample

Discussion prompt

An affine map takes a vector x, applies a matrix W, then adds a fixed vector b. The Wx rotates, scales, or shears; the + b slides the whole result.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

7. Our running map

Concept

We fix one concrete 2-D map and carry it through the whole lesson. Here is W, b, and the input x we will keep reusing:

\[ W = \begin{bmatrix} 2 & 0 \\ 1 & 3 \end{bmatrix}, \qquad b = \begin{bmatrix} 1 \\ -1 \end{bmatrix}, \qquad x = \begin{bmatrix} 3 \\ 4 \end{bmatrix} \]

Two pieces: the (2×2) matrix W and the shift vector b

Why: W acts on the 2-vector x to give a 2-vector; then b (also a 2-vector) is added. Shapes must line up — that is the first thing to check on any affine map.

8. What has to happen first: Evaluate T(x), entry by entry

Ranking

Put in order

Put the moves of Evaluate T(x), entry by entry into the order they have to happen.

  1. Row 0 of Wx: 2·3 + 0·4 = 6
  2. Row 1 of Wx: 1·3 + 3·4 = 3 + 12 = 15
  3. Add b = [1, −1]: T(x) = [6+1, 15−1] = [7, 14]

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. First row [2, 0] dotted with [3, 4]: the second feature is zeroed out, so only 2·3 survives.

9. Evaluate T(x), entry by entry

Worked example

Do Wx first: each row of W dotted with x = [3, 4]. Then add b component-wise. No steps skipped.

Row 0 of Wx: 2·3 + 0·4 = 6

Why: First row [2, 0] dotted with [3, 4]: the second feature is zeroed out, so only 2·3 survives.

Row 1 of Wx: 1·3 + 3·4 = 3 + 12 = 15

Why: Second row [1, 3] dotted with [3, 4]. So Wx = [6, 15].

\[ Wx = \begin{bmatrix} 2 & 0 \\ 1 & 3 \end{bmatrix}\!\begin{bmatrix} 3 \\ 4 \end{bmatrix} = \begin{bmatrix} 6 \\ 15 \end{bmatrix} \]

Add b = [1, −1]: T(x) = [6+1, 15−1] = [7, 14]

Why: The translation slides Wx by b component-wise. This [7, 14] is T(x), and we will confirm it in code next.

\[ T(x) = Wx + b = \begin{bmatrix} 6 \\ 15 \end{bmatrix} + \begin{bmatrix} 1 \\ -1 \end{bmatrix} = \begin{bmatrix} 7 \\ 14 \end{bmatrix} \]

10. Decode the notation: Evaluate T(x), entry by entry

Notation

Annotate

From Evaluate T(x), entry by entry — read this one piece at a time. What is each part doing?

On: \( Wx = \begin{bmatrix} 2 & 0 \\ 1 & 3 \end{bmatrix}\!\begin{bmatrix} 3 \\ 4 \end{bmatrix} = \begin{bmatrix} 6 \\ 15 \end{bmatrix} \)

  • First row [2, 0] dotted with [3, 4]: the second feature is zeroed out, so only 2·3 survives.
  • Second row [1, 3] dotted with [3, 4]. So Wx = [6, 15].
  • The translation slides Wx by b component-wise. This [7, 14] is T(x), and we will confirm it in code next.

11. Guess the shape of the answer: Confirm T(x) in code

Estimation

Predict first

Type the map out and print Wx, then T(x). Predict [7, 14] before you run it. This block is complete and runnable on its own:

Commit before you compute: what does Confirm T(x) in code come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Wx = [6, 15], then + b gives T(x) = [7, 14]

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Matches the hand computation exactly.

12. Confirm T(x) in code

Worked example

Type the map out and print Wx, then T(x). Predict [7, 14] before you run it. This block is complete and runnable on its own:

import numpy as np
W = np.array([[2., 0.], [1., 3.]])
b = np.array([1., -1.])
x = np.array([3., 4.])
print('Wx     =', (W @ x).tolist())
print('T(x)   =', (W @ x + b).tolist())

Wx = [6, 15], then + b gives T(x) = [7, 14]

Why: Matches the hand computation exactly. The '@' operator is matrix-vector multiply; adding b broadcasts component-wise.

expressionvalue (verified)
W @ x[6.0, 15.0]
W @ x + b[7.0, 14.0]
shapesW (2×2), x (2,), b (2,) → T(x) (2,)

13. Fill in: value (verified) for Confirm T(x) in code

Comparison

Comparison matrix

From Confirm T(x) in code: refill the value (verified) column from what you know. The rest of the table is as it appeared.

expressionvalue (verified)
W @ x[6.0, 15.0]
W @ x + b[7.0, 14.0]
shapesW (2×2), x (2,), b (2,) → T(x) (2,)

14. What has to be given first: Map the whole unit square

Missing information

Discussion prompt

To see what T does, push the four corners of the unit square through it. The image is a parallelogram — never a curved blob, because affine maps keep lines straight. Runnable:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

T(0) = W·0 + b = b. The affine map moves the origin to the shift vector — the visible signature that b ≠ 0.

15. Map the whole unit square

Worked example

To see what T does, push the four corners of the unit square through it. The image is a parallelogram — never a curved blob, because affine maps keep lines straight. Runnable:

import numpy as np
W = np.array([[2., 0.], [1., 3.]]); b = np.array([1., -1.])
corners = np.array([[0., 0.], [1., 0.], [0., 1.], [1., 1.]])
for p in corners:
    print(p.tolist(), '->', (W @ p + b).tolist())

The origin [0,0] maps to b = [1,−1]

Why: T(0) = W·0 + b = b. The affine map moves the origin to the shift vector — the visible signature that b ≠ 0.

Opposite edges stay parallel in the image

Why: Edge [0,0]→[1,0] maps to [1,−1]→[3,0] (direction [2,1]); the opposite edge [0,1]→[1,1] maps to [1,2]→[3,3] (also [2,1]). Parallelism preserved — the square becomes a parallelogram.

cornerT(corner)
[0, 0][1, −1]
[1, 0][3, 0]
[0, 1][1, 2]
[1, 1][3, 3]

16. What each one costs: Map the whole unit square

Trade off

Comparison matrix

From Map the whole unit square: every row here is a choice with a cost. Fill the T(corner) column, then say which row you would actually pick and what you give up for it.

cornerT(corner)
[0, 0][1, −1]
[1, 0][3, 0]
[0, 1][1, 2]
[1, 1][3, 3]

17. When is an affine map linear?

Concept

A map T is linear when it satisfies two axioms for all vectors u, v and scalars c:

\[ \text{additivity: } T(u+v) = T(u)+T(v), \qquad \text{homogeneity: } T(cu) = c\,T(u) \]

The pure matrix map x ↦ Wx obeys both. The shift + b is what an affine map adds on top — and that shift is exactly what can break linearity.

18. By analogy: When is an affine map linear?

Analogy

Discussion prompt

Explain When is an affine map linear? by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

A map T is linear when it satisfies two axioms for all vectors u, v and scalars c:

19. Plan first: The shift breaks homogeneity

Step zero

Discussion prompt

The shift breaks homogeneity — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Left side: T(2x) = W(2x) + b = 2Wx + b

Answer:

  1. Left side: T(2x) = W(2x) + b = 2Wx + b
  2. Right side: 2·T(x) = 2(Wx + b) = 2Wx + 2b
  3. They differ by exactly b (unless b = 0)

20. The shift breaks homogeneity

Worked example

Test homogeneity T(cx) = c·T(x) on our map with c = 2. Watch the bias fail to scale.

Left side: T(2x) = W(2x) + b = 2Wx + b

Why: Doubling x doubles Wx (matrix mult is linear), but b is added once, unchanged.

\[ T(2x) = 2\,Wx + b \]

Right side: 2·T(x) = 2(Wx + b) = 2Wx + 2b

Why: Scaling the whole output doubles b too.

\[ 2\,T(x) = 2\,Wx + 2b \]

They differ by exactly b (unless b = 0)

Why: T(2x) − 2T(x) = b − 2b = −b. So the map is linear only when b = 0; with a nonzero shift it is affine but NOT linear.

\[ T(2x) - 2\,T(x) = b - 2b = -b \;\neq\; 0 \]

21. Work backwards from the answer: The shift breaks homogeneity

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

They differ by exactly b (unless b = 0)

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Test homogeneity T(cx) = c·T(x) on our map with c = 2. Watch the bias fail to scale.

22. Affine = linear, but the origin can move

Intuition

A linear map must send 0 to 0 — plug in x = 0 and W·0 = 0. An affine map sends 0 to b: the origin is free to move.

Picture a rubber sheet. The linear part W stretches and rotates the sheet around a pinned origin; the + b then picks the whole sheet up and slides it. Nothing bends — that is the key limitation we exploit later.

23. Teach it back: Affine = linear, but the origin can move

Explain it

Discussion prompt

Explain Affine = linear, but the origin can move to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

A linear map must send 0 to 0 — plug in x = 0 and W·0 = 0. An affine map sends 0 to b: the origin is free to move.

24. Something is wrong here: calling an affine map linear

Anomaly

Predict first

A student writes this, and it looks reasonable:

T(x) = Wx + b is built from a matrix, so it must be a linear map.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Then T(u+v) = W(u+v) + b = Wu + Wv + b, but T(u)+T(v) = Wu + Wv + 2b.

Call it affine. It is linear only in the special case b = 0.

Why: Then T(u+v) = W(u+v) + b = Wu + Wv + b, but T(u)+T(v) = Wu + Wv + 2b. They differ by b. With b ≠ 0 additivity FAILS — the map is affine, not linear.

25. Trap: calling an affine map linear

Trap

The trap

T(x) = Wx + b is built from a matrix, so it must be a linear map.

Assume T(u + v) = T(u) + T(v)

Why: Then T(u+v) = W(u+v) + b = Wu + Wv + b, but T(u)+T(v) = Wu + Wv + 2b. They differ by b. With b ≠ 0 additivity FAILS — the map is affine, not linear.

The fix

Call it affine. It is linear only in the special case b = 0.

Affine = (linear part Wx) + (translation b)

Why: Linearity requires T(0) = 0; here T(0) = b. Reserve 'linear' for b = 0. In ML the bias is almost always nonzero, so a 'linear layer' nn.Linear(with bias) is really an affine layer.

26. The geometry of W

Section

Part 2 of 6 — transforms & invariants

27. Classic transforms are choices of W

Concept

Rotation, reflection, scaling, and shear are not separate machinery — each is just a particular matrix W (with b = 0, or a translation added):

transformwhat W doessignature
rotationturns space about the originorthogonal W, det = +1
reflectionflips across a lineorthogonal W, det = −1
scalingstretches along axesdiagonal W
shearslides one axis along anotheroff-diagonal entries ≠ 0

Our running W = [[2,0],[1,3]] is a scale-and-shear: the diagonal scales, the lower-left 1 shears.

28. Predict the next row: A rotation matrix, checked

Pattern

Predict first

The table runs: W (rounded) | [[0, −1], [1, 0]] · det(W) | 1.0 (rotation)

In A rotation matrix, checked, given the rows so far: what is the next one — the row where quantity is W @ [1, 0]?

Correct: W @ [1, 0] | [0.0, 1.0]

quantityvalue (verified)
W (rounded)[[0, −1], [1, 0]]
det(W)1.0 (rotation)
W @ [1, 0][0.0, 1.0]

Why: The relationship between the columns, not the individual numbers, is what generates the next row. det = +1 confirms a pure rotation (no reflection, no scaling).

29. A rotation matrix, checked

Worked example

A 90° rotation should send [1, 0] to [0, 1]. Build the rotation W for θ = 90° and apply it. Runnable standalone:

import numpy as np
theta = np.pi / 2                       # 90 degrees
W = np.array([[np.cos(theta), -np.sin(theta)],
              [np.sin(theta),  np.cos(theta)]])
e1 = np.array([1., 0.])
print('W =', np.round(W, 3).tolist())
print('det(W) =', round(np.linalg.det(W), 6))
print('W @ [1,0] =', np.round(W @ e1, 6).tolist())

det(W) = 1.0, and W·[1,0] = [0, 1]

Why: det = +1 confirms a pure rotation (no reflection, no scaling). The x-axis unit vector lands on the y-axis — a 90° turn, exactly as expected.

quantityvalue (verified)
W (rounded)[[0, −1], [1, 0]]
det(W)1.0 (rotation)
W @ [1, 0][0.0, 1.0]

30. The determinant's sign tells you the flip

Concept

For an orthogonal W, the sign of the determinant distinguishes a rotation from a reflection: +1 keeps orientation (rotation), −1 flips it (reflection). A reflection across the x-axis negates the second coordinate:

\[ W_{\text{reflect}} = \begin{bmatrix} 1 & 0 \\ 0 & -1 \end{bmatrix}, \qquad \det W_{\text{reflect}} = (1)(-1) - 0 = -1 \]

Apply it: W_reflect · [2, 3] = [2, −3]

Why: The x-coordinate is untouched, the y-coordinate flips sign — a mirror across the x-axis. det = −1 flags exactly this orientation reversal.

31. What affine maps preserve

Concept

Because an affine map only stretches and slides — never bends — it keeps three things intact:

  1. Straight lines stay straight (a line maps to a line)
  2. Parallel lines stay parallel
  3. Ratios along a line are preserved — the midpoint stays the midpoint

What it cannot do: bend a straight line into a curve, or make two parallel lines cross. That inability is precisely why one affine map — however large W — can never carve XOR's boundary.

32. Restore the missing line: Verify: midpoints and parallels survive

Fill the middle

Fill in the blanks

From Verify: midpoints and parallels survive — one line has had its right-hand side removed. Put it back.

import numpy as np
W = np.array([[2., 1.], [0., 1.]]); b = np.array([1., -2.])
def T(p): return W @ p + b
P = np.array([0., 0.]); Q = np.array([1., 0.]); R = np.array([0., 1.])
mid = (P + Q) / 2
print('T(mid PQ) =', T(mid).tolist())
print('mid of T(P),T(Q) =', ((T(P) + T(Q)) / 2).tolist())
e1 = T(Q) - T(P)
e2 = T(R + (Q - P)) - T(R)
print('edge1 image =', e1.tolist(), ' edge2 image =', e2.tolist())

Why: e1 is what everything below it consumes, so the wrong expression here fails later and somewhere else. Both print [2, -2]: the affine map takes the midpoint of PQ to the midpoint of T(P),T(Q).

33. Verify: midpoints and parallels survive

Worked example

Take a shear-scale-shift map. Send the midpoint of an edge through it, and separately send two parallel edges through it, and check both invariants. Runnable:

import numpy as np
W = np.array([[2., 1.], [0., 1.]]); b = np.array([1., -2.])
def T(p): return W @ p + b
P = np.array([0., 0.]); Q = np.array([1., 0.]); R = np.array([0., 1.])
mid = (P + Q) / 2
print('T(mid PQ)         =', T(mid).tolist())
print('mid of T(P),T(Q)  =', ((T(P) + T(Q)) / 2).tolist())
e1 = T(Q) - T(P)
e2 = T(R + (Q - P)) - T(R)
print('edge1 image =', e1.tolist(), ' edge2 image =', e2.tolist())

T(midpoint) equals the midpoint of the images

Why: Both print [2, -2]: the affine map takes the midpoint of PQ to the midpoint of T(P),T(Q). Ratios along a line are preserved.

Two parallel edges map to identical direction vectors

Why: edge1 and edge2 both map to [2, 0] — still parallel. An affine map cannot make parallel lines converge.

invariantbefore → afterverified
T(P)[0,0] → [1, −2]—
T(Q)[1,0] → [3, −2]—
midpoint of PQ→ [2, −2]= mid of images ✓
two parallel edges→ [2, 0] and [2, 0]still parallel ✓

34. Which is which, by verified

Discrimination

Sort into buckets

Sort these by verified, from memory, without looking back at Verify: midpoints and parallels survive. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

—
T(P); T(Q)
= mid of images ✓
midpoint of PQ
still parallel ✓
two parallel edges
g1
verified is "—" for T(P), T(Q) — that is what the table on "Verify: midpoints and parallels survive" records, and it is the single property separating this group from the rest.
g2
verified is "= mid of images ✓" for midpoint of PQ — that is what the table on "Verify: midpoints and parallels survive" records, and it is the single property separating this group from the rest.
g3
verified is "still parallel ✓" for two parallel edges — that is what the table on "Verify: midpoints and parallels survive" records, and it is the single property separating this group from the rest.

35. Composition stays affine

Section

Part 3 of 6 — every step

36. A network is just function composition

Intuition

A feed-forward network is a chain: the output of layer 1 is the input to layer 2, whose output feeds layer 3. Mathematically that is function composition — layer₃(layer₂(layer₁(x))).

So the question 'what does a whole linear network compute?' is really 'what is the composition of these affine maps?' Answer that once, for two maps, and induction handles any depth. That is why the next few slides matter far beyond a single example.

37. Stack two affine maps

Concept

A network applies maps one after another. Suppose T₁(x) = W₁x + b₁ runs first, then T₂(u) = W₂u + b₂ runs on its output. What is the combined map T₂(T₁(x))?

\[ T_1(x) = W_1 x + b_1, \qquad T_2(u) = W_2 u + b_2 \]

We will substitute and simplify — and the punchline is that the result is again a single Wx + b.

38. What has to happen first: Derive the composition, one move per beat

Ranking

Put in order

Put the moves of Derive the composition, one move per beat into the order they have to happen.

  1. Substitute u = T₁(x) into T₂
  2. Distribute W₂ across the parentheses
  3. Group the x-term and the constant terms
  4. The composition IS affine, with new W and b

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. The output of the first map becomes the input of the second: replace u by W₁x + b₁.

39. Derive the composition, one move per beat

Worked example

Substitute u = T₁(x) into T₂

Why: The output of the first map becomes the input of the second: replace u by W₁x + b₁.

\[ T_2(T_1(x)) = W_2\,(W_1 x + b_1) + b_2 \]

Distribute W₂ across the parentheses

Why: Matrix multiplication distributes over addition: W₂(W₁x + b₁) = W₂W₁x + W₂b₁.

\[ = W_2 W_1 x + W_2 b_1 + b_2 \]

Group the x-term and the constant terms

Why: Everything multiplying x is (W₂W₁); everything without x is (W₂b₁ + b₂). That is the Wx + b shape.

\[ = \underbrace{(W_2 W_1)}_{W_{\text{eff}}}\,x + \underbrace{(W_2 b_1 + b_2)}_{b_{\text{eff}}} \]

The composition IS affine, with new W and b

Why: Combined weight W₂W₁, combined bias W₂b₁ + b₂. Notice b₁ is first transformed by W₂ — the biases do NOT simply add.

\[ \boxed{\,T_2\!\circ\! T_1 \,:\; x \mapsto (W_2 W_1)\,x + (W_2 b_1 + b_2)\,} \]

40. Say it in words: Derive the composition, one move per beat

Translation

\( = \underbrace{(W_2 W_1)}_{W_{\text{eff}}}\,x + \underbrace{(W_2 b_1 + b_2)}_{b_{\text{eff}}} \)

Draw it

Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.

41. Our two concrete maps

Concept

To make the formula bite, fix a second map T₂ and reuse T₁ from Part 1. We will run the two-step computation and the single-map formula and check they agree.

\[ W_1 = \begin{bmatrix} 2 & 0 \\ 1 & 3 \end{bmatrix},\; b_1 = \begin{bmatrix} 1 \\ -1 \end{bmatrix}, \qquad W_2 = \begin{bmatrix} 1 & -1 \\ 0 & 2 \end{bmatrix},\; b_2 = \begin{bmatrix} 2 \\ 0 \end{bmatrix} \]

Same input x = [3, 4]. From Part 1 we already have T₁(x) = [7, 14].

42. Plan first: Two-step path: T₂(T₁(x)) by hand

Step zero

Discussion prompt

Two-step path: T₂(T₁(x)) by hand — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: W₂·[7,14]: row 0 = 1·7 + (−1)·14 = 7 − 14 = −7

Answer:

  1. W₂·[7,14]: row 0 = 1·7 + (−1)·14 = 7 − 14 = −7
  2. W₂·[7,14]: row 1 = 0·7 + 2·14 = 28
  3. Add b₂ = [2, 0]: T₂(T₁(x)) = [−7+2, 28+0] = [−5, 28]

43. Two-step path: T₂(T₁(x)) by hand

Worked example

Run the maps in sequence. T₁(x) = [7, 14]; now push that through T₂.

W₂·[7,14]: row 0 = 1·7 + (−1)·14 = 7 − 14 = −7

Why: First row [1, −1] of W₂ dotted with u = [7, 14].

W₂·[7,14]: row 1 = 0·7 + 2·14 = 28

Why: Second row [0, 2] dotted with [7, 14]. So W₂u = [−7, 28].

\[ W_2\,T_1(x) = \begin{bmatrix} -7 \\ 28 \end{bmatrix} \]

Add b₂ = [2, 0]: T₂(T₁(x)) = [−7+2, 28+0] = [−5, 28]

Why: The final output of running both maps in sequence.

\[ T_2(T_1(x)) = \begin{bmatrix} -5 \\ 28 \end{bmatrix} \]

44. Draw the shape of it: Two-step path: T₂(T₁(x)) by hand

Blank canvas

Draw it

Draw what Two-step path: T₂(T₁(x)) by hand just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.

45. What has to happen first: One-step path: the combined map by hand

Ranking

Put in order

Put the moves of One-step path: the combined map by hand into the order they have to happen.

  1. W₂W₁: compute the (2×2) product
  2. b_eff = W₂b₁ + b₂ = [2,−2] + [2,0] = [4, −2]
  3. Apply: (W₂W₁)x + b_eff = [1·3−3·4, 2·3+6·4] + [4,−2]

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Row 0: [1,−1]·cols of W₁ = [2−1, 0−3] = [1, −3].

46. One-step path: the combined map by hand

Worked example

Now build W_eff = W₂W₁ and b_eff = W₂b₁ + b₂ directly, then apply to x. Must land on [−5, 28].

W₂W₁: compute the (2×2) product

Why: Row 0: [1,−1]·cols of W₁ = [2−1, 0−3] = [1, −3]. Row 1: [0,2]·cols = [0+2, 0+6] = [2, 6].

\[ W_2 W_1 = \begin{bmatrix} 1 & -1 \\ 0 & 2 \end{bmatrix}\!\begin{bmatrix} 2 & 0 \\ 1 & 3 \end{bmatrix} = \begin{bmatrix} 1 & -3 \\ 2 & 6 \end{bmatrix} \]

b_eff = W₂b₁ + b₂ = [2,−2] + [2,0] = [4, −2]

Why: W₂b₁ = [1·1+(−1)(−1), 0·1+2·(−1)] = [2, −2]; add b₂ = [2, 0].

\[ W_2 b_1 + b_2 = \begin{bmatrix} 2 \\ -2 \end{bmatrix} + \begin{bmatrix} 2 \\ 0 \end{bmatrix} = \begin{bmatrix} 4 \\ -2 \end{bmatrix} \]

Apply: (W₂W₁)x + b_eff = [1·3−3·4, 2·3+6·4] + [4,−2]

Why: [3−12, 6+24] + [4,−2] = [−9, 30] + [4, −2] = [−5, 28]. Identical to the two-step path — composition really is one affine map.

\[ \begin{bmatrix} 1 & -3 \\ 2 & 6 \end{bmatrix}\!\begin{bmatrix} 3 \\ 4 \end{bmatrix} + \begin{bmatrix} 4 \\ -2 \end{bmatrix} = \begin{bmatrix} -9 \\ 30 \end{bmatrix} + \begin{bmatrix} 4 \\ -2 \end{bmatrix} = \begin{bmatrix} -5 \\ 28 \end{bmatrix} \]

47. Decode the notation: One-step path: the combined map by hand

Notation

Annotate

From One-step path: the combined map by hand — read this one piece at a time. What is each part doing?

On: \( W_2 b_1 + b_2 = \begin{bmatrix} 2 \\ -2 \end{bmatrix} + \begin{bmatrix} 2 \\ 0 \end{bmatrix} = \begin{bmatrix} 4 \\ -2 \end{bmatrix} \)

  • Row 0: [1,−1]·cols of W₁ = [2−1, 0−3] = [1, −3]. Row 1: [0,2]·cols = [0+2, 0+6] = [2, 6].
  • W₂b₁ = [1·1+(−1)(−1), 0·1+2·(−1)] = [2, −2]; add b₂ = [2, 0].
  • [3−12, 6+24] + [4,−2] = [−9, 30] + [4, −2] = [−5, 28]. Identical to the two-step path — composition really is one affine map.

48. Finish it with less help: Confirm both paths in code

Faded example

Fill in the blanks

Confirm both paths in code, with the scaffolding fading: two lines are gone now — fill both.

import numpy as np
W1 = np.array([[2., 0.], [1., 3.]]); b1 = np.array([1., -1.])
W2 = np.array([[1., -1.], [0., 2.]]); b2 = np.array([2., 0.])
x = np.array([3., 4.])
direct = W2 @ (W1 @ x + b1) + b2
composed = (W2 @ W1) @ x + (W2 @ b1 + b2)
print('W2@W1 =', (W2 @ W1).tolist())
print('W2@b1+b2 =', (W2 @ b1 + b2).tolist())
print('direct =', direct.tolist())
print('composed =', composed.tolist())
print('match:', np.allclose(direct, composed))

Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. The sequential two-step computation and the single combined map agree exactly.

49. Confirm both paths in code

Worked example

Compute the two-step and the one-step results and compare with np.allclose. Predict True and [−5, 28]. Runnable standalone:

import numpy as np
W1 = np.array([[2., 0.], [1., 3.]]); b1 = np.array([1., -1.])
W2 = np.array([[1., -1.], [0., 2.]]); b2 = np.array([2., 0.])
x  = np.array([3., 4.])
direct   = W2 @ (W1 @ x + b1) + b2
composed = (W2 @ W1) @ x + (W2 @ b1 + b2)
print('W2@W1    =', (W2 @ W1).tolist())
print('W2@b1+b2 =', (W2 @ b1 + b2).tolist())
print('direct   =', direct.tolist())
print('composed =', composed.tolist())
print('match:', np.allclose(direct, composed))

direct == composed == [−5, 28], match = True

Why: The sequential two-step computation and the single combined map agree exactly. Composition of affine maps is affine — confirmed against the hand math.

computationresult (verified)
W₂W₁[[1, −3], [2, 6]]
W₂b₁ + b₂[4, −2]
T₂(T₁(x)) directly[−5.0, 28.0]
(W₂W₁)x + (W₂b₁+b₂)[−5.0, 28.0]
np.allclose(...)True

50. Something is wrong here: the biases just add

Anomaly

Predict first

A student writes this, and it looks reasonable:

Composing T₁ then T₂, the combined bias is surely b₁ + b₂.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: This ignores that b₁ is produced INSIDE T₁ and then passes through W₂.

b₁ is transformed by W₂ before b₂ is added.

Why: This ignores that b₁ is produced INSIDE T₁ and then passes through W₂. Plugging [3,−1] into the check gives the wrong output — it disagrees with the [−5, 28] both honest paths produced.

51. Trap: the biases just add

Trap

The trap

Composing T₁ then T₂, the combined bias is surely b₁ + b₂.

b_eff = b₁ + b₂ = [1,−1] + [2,0] = [3, −1]

Why: This ignores that b₁ is produced INSIDE T₁ and then passes through W₂. Plugging [3,−1] into the check gives the wrong output — it disagrees with the [−5, 28] both honest paths produced.

The fix

b₁ is transformed by W₂ before b₂ is added.

b_eff = W₂b₁ + b₂ = [2,−2] + [2,0] = [4, −2]

Why: Trace the algebra: W₂(W₁x + b₁) + b₂ = W₂W₁x + W₂b₁ + b₂. The first bias rides through the second matrix. Order matters — W₂ on the left, and it hits b₁ too.

52. Break it on purpose: the biases just add

Break the constraint

Discussion prompt

The rule this trap just fixed:

Trace the algebra: W₂(W₁x + b₁) + b₂ = W₂W₁x + W₂b₁ + b₂. The first bias rides through the second matrix. Order matters — W₂ on the left, and it hits b₁ too.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

This ignores that b₁ is produced INSIDE T₁ and then passes through W₂. Plugging [3,−1] into the check gives the wrong output — it disagrees with the [−5, 28] both honest paths produced.

53. One matrix: homogeneous coordinates

Section

Part 4 of 6 — the packaging trick

54. Absorb the bias into the matrix

Concept

Carrying W and b separately is a nuisance. The trick: append a 1 to x, and build a bigger matrix M whose last column is b. Then a single matrix multiply does Wx + b.

\[ M = \begin{bmatrix} W & b \\ 0 & 1 \end{bmatrix}, \qquad \tilde{x} = \begin{bmatrix} x \\ 1 \end{bmatrix} \;\Longrightarrow\; M\tilde{x} = \begin{bmatrix} Wx + b \\ 1 \end{bmatrix} \]

The extra 1 rides along untouched (bottom row [0, 1]), so you can chain these M's freely. This is how graphics pipelines and transformer position code pack affine maps.

55. Build M and check it reproduces Wx + b

Worked example

Using our W, b, x, build the 3×3 homogeneous M, append a 1 to x, and confirm the top two outputs equal Wx + b = [7, 14]. Runnable:

import numpy as np
W = np.array([[2., 0.], [1., 3.]]); b = np.array([1., -1.])
x = np.array([3., 4.])
M = np.block([[W, b.reshape(2, 1)], [np.zeros((1, 2)), np.ones((1, 1))]])
xh = np.array([3., 4., 1.])          # append a 1
yh = M @ xh
print('M =', M.tolist())
print('M @ [x; 1] =', yh.tolist())
print('Wx + b     =', (W @ x + b).tolist())

M is 3×3 with b in the last column, [0,0,1] on the bottom

Why: np.block stacks W (2×2) beside b (2×1), then a bottom row [0, 0, 1]. The last row keeps the trailing 1 intact.

M·[x;1] = [7, 14, 1] → top two match Wx + b = [7, 14]

Why: One matrix multiply reproduced the affine map, and the appended 1 came back as 1 — ready to feed into the next M.

objectvalue (verified)
M row 0[2, 0, 1]
M row 1[1, 3, −1]
M row 2[0, 0, 1]
M @ [3, 4, 1][7.0, 14.0, 1.0]
Wx + b[7.0, 14.0]

56. Inspect it line by line: Build M and check it reproduces Wx + b

Error analysis

Annotate

Walk the callouts on Build M and check it reproduces Wx + b. Each one is a place this is easy to get subtly wrong.

  • np.block stacks W (2×2) beside b (2×1), then a bottom row [0, 0, 1]. The last row keeps the trailing 1 intact.
  • One matrix multiply reproduced the affine map, and the appended 1 came back as 1 — ready to feed into the next M.

57. Guess the shape of the answer: Composition becomes a plain matrix product

Estimation

Predict first

The payoff: in homogeneous form, composing two affine maps is just M₂ M₁ — one matrix product, no separate bias bookkeeping. Build both and read off the blocks:

Commit before you compute: what does Composition becomes a plain matrix product come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: M₂M₁ last column (top) = W₂b₁+b₂ = [4,−2]

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. And the combined bias appears in the last column — exactly [4, −2].

58. Composition becomes a plain matrix product

Worked example

The payoff: in homogeneous form, composing two affine maps is just M₂ M₁ — one matrix product, no separate bias bookkeeping. Build both and read off the blocks:

import numpy as np
W1 = np.array([[2., 0.], [1., 3.]]); b1 = np.array([1., -1.])
W2 = np.array([[1., -1.], [0., 2.]]); b2 = np.array([2., 0.])
def homo(W, b):
    return np.block([[W, b.reshape(2, 1)], [np.zeros((1, 2)), np.ones((1, 1))]])
M = homo(W2, b2) @ homo(W1, b1)
print('M2 @ M1 =', M.tolist())
print('top-left 2x2 (= W2W1)   =', M[:2, :2].tolist())
print('last col top (= W2b1+b2)=', M[:2, 2].tolist())

M₂M₁ top-left block = W₂W₁ = [[1,−3],[2,6]]

Why: The matrix product automatically produces the combined weight in the upper-left 2×2 — the same [[1,−3],[2,6]] we derived by hand.

M₂M₁ last column (top) = W₂b₁+b₂ = [4,−2]

Why: And the combined bias appears in the last column — exactly [4, −2]. The homogeneous product encodes both the weight and the transformed-bias rule in one step.

block of M₂M₁value (verified)matches
top-left 2×2[[1, −3], [2, 6]]W₂W₁ ✓
last column (top 2)[4, −2]W₂b₁ + b₂ ✓
bottom row[0, 0, 1]preserved ✓

59. Where affine maps show up

Concept

The Wx + b skeleton is not a toy — it is the pre-activation of every dense layer, and it recurs across the whole USAAIO syllabus:

Every one of these inherits the collapse problem we are about to prove — which is why every one is paired with a nonlinearity.

60. The collapse

Section

Part 5 of 6 — why linear depth is free

61. A linear network is affine ∘ affine ∘ …

Concept

A neural network with no activation functions is just a chain of affine maps. By the composition rule, that chain is itself a single affine map:

\[ A_3\big(A_2(A_1 x)\big) = (A_3 A_2 A_1)\,x \]

So a 3-layer linear network has exactly the expressive power of a 1-layer one. Depth without nonlinearity buys nothing — this is the central fact of the lesson.

62. Why stacking matrices adds no power

Intuition

Each linear layer multiplies its input by a matrix. Chain them and the matrices multiply into one matrix — and any single matrix is reachable by a single layer.

It is like composing the functions f(x) = 3x and g(x) = 5x: g(f(x)) = 15x is still just multiply-by-a-constant. You never escape the family of straight-through-the-origin maps, no matter how many you stack.

63. Restore the missing line: Three linear layers = one matrix

Fill the middle

Fill in the blanks

From Three linear layers = one matrix — one line has had its right-hand side removed. Put it back.

import numpy as np
np.random.seed(0)
A1 = np.random.randn(4, 3)
A2 = np.random.randn(5, 4)
A3 = np.random.randn(2, 5)
x = np.random.randn(3)
stacked = A3 @ (A2 @ (A1 @ x))
W_eff = A3 @ A2 @ A1
single = W_eff @ x
print('stacked =', np.round(stacked, 6).tolist())
print('single =', np.round(single, 6).tolist())
print('effective shape:', W_eff.shape)
print('match:', np.allclose(stacked, single))

Why: x is what everything below it consumes, so the wrong expression here fails later and somewhere else. Running the three layers in sequence gives the identical output to multiplying by the single collapsed matrix.

64. Three linear layers = one matrix

Worked example

Chain three random linear layers (3→4→5→2) and compare running them in sequence against multiplying their matrices into one. Seeded, so your numbers match. Runnable:

import numpy as np
np.random.seed(0)
A1 = np.random.randn(4, 3)
A2 = np.random.randn(5, 4)
A3 = np.random.randn(2, 5)
x  = np.random.randn(3)
stacked = A3 @ (A2 @ (A1 @ x))
W_eff   = A3 @ A2 @ A1
single  = W_eff @ x
print('stacked =', np.round(stacked, 6).tolist())
print('single  =', np.round(single, 6).tolist())
print('effective shape:', W_eff.shape)
print('match:', np.allclose(stacked, single))

stacked == single == [1.562102, 8.741537]

Why: Running the three layers in sequence gives the identical output to multiplying by the single collapsed matrix. The intermediate widths 4 and 5 vanish.

The effective map is a single (2×3) matrix

Why: 3 inputs, 2 outputs — a 2×3 matrix. The whole 3→4→5→2 network has exactly the capacity of one 2×3 linear layer.

networkoutputeffective map
3 layers stacked (3→4→5→2)[1.562102, 8.741537]—
single collapsed matrix[1.562102, 8.741537]one 2×3 matrix
np.allcloseTrueidentical

65. Finish it with less help: With biases too, the affine stack collapses

Faded example

Fill in the blanks

With biases too, the affine stack collapses, with the scaffolding fading: two lines are gone now — fill both.

import numpy as np
np.random.seed(0)
A1 = np.random.randn(4, 3); c1 = np.random.randn(4)
A2 = np.random.randn(5, 4); c2 = np.random.randn(5)
A3 = np.random.randn(2, 5); c3 = np.random.randn(2)
x = np.random.randn(3)
stacked = A3 @ (A2 @ (A1 @ x + c1) + c2) + c3
W_eff = A3 @ A2 @ A1
b_eff = A3 @ (A2 @ c1 + c2) + c3
single = W_eff @ x + b_eff
print('W_eff shape:', W_eff.shape, ' b_eff shape:', b_eff.shape)
print('match:', np.allclose(stacked, single))

Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. The composition formula (W₂b₁ + b₂) generalizes: each earlier bias passes through all later matrices.

66. With biases too, the affine stack collapses

Worked example

Adding a bias to each layer changes nothing about the conclusion — the stack is still one affine map W_eff x + b_eff. Verify with three biased layers:

import numpy as np
np.random.seed(0)
A1 = np.random.randn(4, 3); c1 = np.random.randn(4)
A2 = np.random.randn(5, 4); c2 = np.random.randn(5)
A3 = np.random.randn(2, 5); c3 = np.random.randn(2)
x  = np.random.randn(3)
stacked = A3 @ (A2 @ (A1 @ x + c1) + c2) + c3
W_eff = A3 @ A2 @ A1
b_eff = A3 @ (A2 @ c1 + c2) + c3
single = W_eff @ x + b_eff
print('W_eff shape:', W_eff.shape, ' b_eff shape:', b_eff.shape)
print('match:', np.allclose(stacked, single))

b_eff = A₃(A₂c₁ + c₂) + c₃ — the nested-bias rule

Why: The composition formula (W₂b₁ + b₂) generalizes: each earlier bias passes through all later matrices. c₁ is hit by A₂ then A₃, c₂ by A₃, c₃ raw.

stacked == single → still one affine map

Why: The biased 3-layer stack equals a single W_eff x + b_eff. Biases do not rescue linear depth — the collapse is total.

quantityvalue (verified)
W_eff shape(2, 3)
b_eff shape(2,)
stacked == singleTrue

67. Fill in: value (verified) for With biases too, the affine stack collapses

Comparison

Comparison matrix

From With biases too, the affine stack collapses: refill the value (verified) column from what you know. The rest of the table is as it appeared.

quantityvalue (verified)
W_eff shape(2, 3)
b_eff shape(2,)
stacked == singleTrue

68. Something is wrong here: more linear layers = more power

Anomaly

Predict first

A student writes this, and it looks reasonable:

The model underfits — stack more linear layers to make it more expressive.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Every extra linear layer just multiplies into the running matrix product.

Insert a nonlinearity between layers — that is what makes depth expressive.

Why: Every extra linear layer just multiplies into the running matrix product. 3 layers collapsed to one 2×3 matrix; 10 layers collapse to one matrix too. Capacity gained: zero.

69. Trap: more linear layers = more power

Trap

The trap

The model underfits — stack more linear layers to make it more expressive.

Add depth with no activations to boost capacity

Why: Every extra linear layer just multiplies into the running matrix product. 3 layers collapsed to one 2×3 matrix; 10 layers collapse to one matrix too. Capacity gained: zero.

The fix

Insert a nonlinearity between layers — that is what makes depth expressive.

Use σ(W₂ σ(W₁x + b₁) + b₂)

Why: The activation σ sits between the affine maps and blocks the collapse: σ(W₂σ(W₁x+b₁)+b₂) is not any single Wx+b. Now each added layer contributes genuinely new structure.

70. Why nonlinearity

Section

Part 6 of 6 — XOR settles it

71. A nonlinearity breaks the collapse

Concept

Put an activation σ (ReLU, tanh, sigmoid, …) between the affine maps, and the composition no longer simplifies to a single Wx + b:

\[ \sigma\!\big(W_2\,\sigma(W_1 x + b_1) + b_2\big) \;\neq\; \text{any single } Wx + b \]

The reason: σ is not linear, so it does not distribute through the matrix products. Now the network can bend space, carve curved decision boundaries, and approximate any function — the universal-approximation property.

72. What a ReLU actually breaks

Intuition

Think of relu(z) = max(0, z) as a switch: for each hidden unit it either passes the value through or clamps it to zero. Which units switch off depends on the input.

So different inputs travel through different effective matrices — the net is a patchwork of many linear pieces stitched along fold lines. That is why doubling the input does not double the output, and why a whole square of inputs can fold onto itself. One affine map has no folds; a hidden ReLU layer has as many as it needs.

73. Teach it back: What a ReLU actually breaks

Explain it

Discussion prompt

Explain What a ReLU actually breaks to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Think of relu(z) = max(0, z) as a switch: for each hidden unit it either passes the value through or clamps it to zero. Which units switch off depends on the input.

74. Restore the missing line: Prove a ReLU net is not linear

Fill the middle

Fill in the blanks

From Prove a ReLU net is not linear — one line has had its right-hand side removed. Put it back.

import numpy as np
np.random.seed(0)
A1 = np.random.randn(4, 2)
A2 = np.random.randn(1, 4)
relu = lambda z: np.maximum(0.0, z)
def net(x): return A2 @ relu(A1 @ x) # one hidden ReLU layer
a = np.array([1.0, 1.0]); b = np.array([-1.0, 2.0])
print('net(a) =', round(net(a)[0], 4))
print('net(b) =', round(net(b)[0], 4))
print('net(a)+net(b) =', round(net(a)[0] + net(b)[0], 4))
print('net(a+b) =', round(net(a + b)[0], 4))
print('additive?', np.allclose(net(a + b), net(a) + net(b)))

Why: A2 is what everything below it consumes, so the wrong expression here fails later and somewhere else. Additivity fails by more than a full unit.

75. Prove a ReLU net is not linear

Worked example

A one-hidden-layer ReLU net must fail an additivity test net(a+b) = net(a) + net(b) for some inputs. Pick a = [1,1], b = [−1,2] and check. Seeded, runnable:

import numpy as np
np.random.seed(0)
A1 = np.random.randn(4, 2)
A2 = np.random.randn(1, 4)
relu = lambda z: np.maximum(0.0, z)
def net(x): return A2 @ relu(A1 @ x)      # one hidden ReLU layer
a = np.array([1.0, 1.0]); b = np.array([-1.0, 2.0])
print('net(a)        =', round(net(a)[0], 4))
print('net(b)        =', round(net(b)[0], 4))
print('net(a)+net(b) =', round(net(a)[0] + net(b)[0], 4))
print('net(a+b)      =', round(net(a + b)[0], 4))
print('additive?', np.allclose(net(a + b), net(a) + net(b)))

net(a)+net(b) = 3.8267 but net(a+b) = 2.6364

Why: Additivity fails by more than a full unit. The ReLU zeros out different coordinates for different inputs, so the map cannot be written as one matrix.

additive? → False → the net is genuinely nonlinear

Why: Because it breaks a linearity axiom, no single affine map equals this net. THAT broken linearity is the capacity the network gains from σ.

quantityvalue (verified)
net(a), a=[1,1]2.3884
net(b), b=[−1,2]1.4383
net(a) + net(b)3.8267
net(a + b), a+b=[0,3]2.6364
additive?False (nonlinear ✓)

76. Draw the shape of it: Prove a ReLU net is not linear

Blank canvas

Draw it

Draw what Prove a ReLU net is not linear just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.

77. XOR: the smallest problem depth needs

Intuition

XOR labels the four corners of a square: [0,0] and [1,1] are class 0; [0,1] and [1,0] are class 1. The two classes sit on opposite diagonals.

No single straight line can put both 0-corners on one side and both 1-corners on the other — the classes are not linearly separable. Since an affine map only draws straight boundaries, one layer is doomed. This is the historical example that stalled neural nets for a decade.

78. By analogy: XOR: the smallest problem depth needs

Analogy

Discussion prompt

Explain XOR: the smallest problem depth needs by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

XOR labels the four corners of a square: [0,0] and [1,1] are class 0; [0,1] and [1,0] are class 1. The two classes sit on opposite diagonals.

79. Predict the next row: The best possible linear fit still fails XOR

Pattern

Predict first

The table runs: [0, 0] | 0 | 0.5 | 0 · [0, 1] | 1 | 0.5 | 0 ✗ · [1, 0] | 1 | 0.5 | 0 ✗

In The best possible linear fit still fails XOR, given the rows so far: what is the next one — the row where point is [1, 1]?

Correct: [1, 1] | 0 | 0.5 | 0

pointtargetraw outputpredicted
[0, 0]00.50
[0, 1]10.50 ✗
[1, 0]10.50 ✗
[1, 1]00.50

Why: The relationship between the columns, not the individual numbers, is what generates the next row. The best line can do no better than predict the mean (0.5) everywhere: the symmetry of XOR forces both slopes to zero.

80. The best possible linear fit still fails XOR

Worked example

Do not just try one linear model — find the best one. Least squares gives the optimal linear weights; then threshold at 0.5 and count how many of the four points it gets right. Runnable:

import numpy as np
X = np.array([[0., 0.], [0., 1.], [1., 0.], [1., 1.]])
y = np.array([0., 1., 1., 0.])
Xb = np.c_[np.ones(4), X]                    # bias + 2 features
w, *_ = np.linalg.lstsq(Xb, y, rcond=None)   # best least-squares line
pred = Xb @ w
print('best weights [b, w1, w2] =', np.round(w, 4).tolist())
print('raw outputs   =', np.round(pred, 4).tolist())
print('thresholded   =', (pred > 0.5).astype(int).tolist())
print('accuracy      =', ((pred > 0.5).astype(float) == y).mean())

Optimal weights = [0.5, 0.0, 0.0] → every prediction is 0.5

Why: The best line can do no better than predict the mean (0.5) everywhere: the symmetry of XOR forces both slopes to zero. It literally cannot tilt to separate the classes.

Thresholded predictions = [0,0,0,0], accuracy = 0.5

Why: It labels everything class 0, matching 2 of 4 points — pure chance. No linear (or stacked-linear) model beats 0.5 on XOR, ever.

pointtargetraw outputpredicted
[0, 0]00.50
[0, 1]10.50 ✗
[1, 0]10.50 ✗
[1, 1]00.50

81. Which is which, by target

Discrimination

Sort into buckets

Sort these by target, from memory, without looking back at The best possible linear fit still fails XOR. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

0
[0, 0]; [1, 1]
1
[0, 1]; [1, 0]
g1
target is "0" for [0, 0], [1, 1] — that is what the table on "The best possible linear fit still…" records, and it is the single property separating this group from the rest.
g2
target is "1" for [0, 1], [1, 0] — that is what the table on "The best possible linear fit still…" records, and it is the single property separating this group from the rest.

82. What has to be given first: An MLP with one hidden nonlinearity solves it

Missing information

Discussion prompt

Same data, but now a 2-layer MLP with a tanh hidden layer. Train both a bare linear model and the MLP and compare accuracy. Seeded torch, runnable:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

Even trained with Adam for 4000 steps, the single linear layer is stuck at chance — matching the least-squares result. Training cannot fix a representational limit.

83. An MLP with one hidden nonlinearity solves it

Worked example

Same data, but now a 2-layer MLP with a tanh hidden layer. Train both a bare linear model and the MLP and compare accuracy. Seeded torch, runnable:

import torch
torch.manual_seed(0)
X = torch.tensor([[0., 0.], [0., 1.], [1., 0.], [1., 1.]])
y = torch.tensor([[0.], [1.], [1.], [0.]])
def train(model, steps=4000, lr=0.05):
    opt = torch.optim.Adam(model.parameters(), lr=lr)
    loss_fn = torch.nn.BCEWithLogitsLoss()
    for _ in range(steps):
        opt.zero_grad(); loss_fn(model(X), y).backward(); opt.step()
    with torch.no_grad():
        pred = (torch.sigmoid(model(X)) > 0.5).float()
        return (pred == y).float().mean().item(), pred.int().ravel().tolist()
lin = torch.nn.Linear(2, 1)
mlp = torch.nn.Sequential(torch.nn.Linear(2, 8), torch.nn.Tanh(), torch.nn.Linear(8, 1))
print('linear:', train(lin))
print('MLP   :', train(mlp))

linear → accuracy 0.5, predictions [0,0,0,0]

Why: Even trained with Adam for 4000 steps, the single linear layer is stuck at chance — matching the least-squares result. Training cannot fix a representational limit.

MLP → accuracy 1.00, predictions [0,1,1,0]

Why: The tanh hidden layer warps the input space so XOR becomes linearly separable in the hidden representation. One nonlinearity is the whole difference.

modelaccuracy (verified)predictions
linear nn.Linear(2,1)0.50 (chance)[0, 0, 0, 0]
MLP (tanh hidden, 8 units)1.00[0, 1, 1, 0]
targets—[0, 1, 1, 0]

84. Work backwards from the answer: An MLP with one hidden nonlinearity solves it

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

MLP → accuracy 1.00, predictions [0,1,1,0]

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Same data, but now a 2-layer MLP with a tanh hidden layer. Train both a bare linear model and the MLP and compare accuracy. Seeded torch, runnable:

85. What the hidden layer actually did

Concept

You do not need training to see the mechanism. Two hand-picked ReLU units build XOR exactly: an OR feature and an AND feature, combined as OR − 2·AND.

\[ h_{\text{or}} = \operatorname{relu}(x_1 + x_2 - 0.5), \quad h_{\text{and}} = \operatorname{relu}(x_1 + x_2 - 1.5), \quad \text{xor} = h_{\text{or}} - 2\,h_{\text{and}} \]

h_or fires when at least one input is 1; h_and fires only when both are 1. Subtracting twice the AND cancels the [1,1] corner back down — the trick a hidden layer discovers on its own.

86. Break it if you can: What the hidden layer actually did

Counterexample

Discussion prompt

You do not need training to see the mechanism. Two hand-picked ReLU units build XOR exactly: an OR feature and an AND feature, combined as OR − 2·AND.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

h_or fires when at least one input is 1; h_and fires only when both are 1. Subtracting twice the AND cancels the [1,1] corner back down — the trick a hidden layer discovers on its own.

87. Guess the shape of the answer: Hand-built XOR from two ReLU features

Estimation

Predict first

Run the two ReLU features across all four points and read off the XOR column. No training — just the formula. Runnable:

Commit before you compute: what does Hand-built XOR from two ReLU features come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: xor column separates the classes (0.0 vs 0.5)

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Corners [0,0]→0.0 and [1,1]→0.5 vs [0,1],[1,0]→0.5...

88. Hand-built XOR from two ReLU features

Worked example

Run the two ReLU features across all four points and read off the XOR column. No training — just the formula. Runnable:

import numpy as np
X = np.array([[0., 0.], [0., 1.], [1., 0.], [1., 1.]])
relu = lambda z: np.maximum(0.0, z)
print('x1 x2 | h_or h_and | xor')
for x1, x2 in X:
    h_or  = relu(x1 + x2 - 0.5)   # fires when >= one input is 1
    h_and = relu(x1 + x2 - 1.5)   # fires only when both are 1
    xor = h_or - 2 * h_and
    print(int(x1), int(x2), ' |', round(h_or, 1), round(h_and, 1), ' |', round(xor, 1))

The [1,1] corner: h_or=1.5, h_and=0.5 → 1.5 − 2(0.5) = 0.5

Why: Without the AND correction, [1,1] would score 1.5 and be misread as class 1. Subtracting 2·AND pulls it back down to 0.5 — the low class, correct for XOR.

xor column separates the classes (0.0 vs 0.5)

Why: Corners [0,0]→0.0 and [1,1]→0.5 vs [0,1],[1,0]→0.5... the two 'on' corners score 0.5 while [0,0] scores 0.0 — now a single threshold in the HIDDEN space splits them. The nonlinearity made XOR separable.

x1x2h_orh_andxor = h_or − 2·h_and
000.00.00.0
010.50.00.5
100.50.00.5
111.50.50.5

89. Inspect it line by line: Hand-built XOR from two ReLU features

Error analysis

Annotate

Walk the callouts on Hand-built XOR from two ReLU features. Each one is a place this is easy to get subtly wrong.

  • Without the AND correction, [1,1] would score 1.5 and be misread as class 1. Subtracting 2·AND pulls it back down to 0.5 — the low class, correct for XOR.
  • Corners [0,0]→0.0 and [1,1]→0.5 vs [0,1],[1,0]→0.5... the two 'on' corners score 0.5 while [0,0] scores 0.0 — now a single threshold in the HIDDEN space splits them. The nonlinearity made XOR separable.

90. Something is wrong here: a bigger linear model for XOR

Anomaly

Predict first

A student writes this, and it looks reasonable:

The linear model failed — throw more parameters (or more linear layers) at XOR until it fits.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Stacked linear layers collapse to ONE linear map (Part 5), and one linear map draws one straight boundary.

Add a hidden layer with a nonlinearity — even 2 hidden units suffice.

Why: Stacked linear layers collapse to ONE linear map (Part 5), and one linear map draws one straight boundary. XOR needs a non-convex region, so accuracy stays pinned at 0.50 no matter the size.

91. Trap: a bigger linear model for XOR

Trap

The trap

The linear model failed — throw more parameters (or more linear layers) at XOR until it fits.

Scale up / stack a linear classifier on XOR

Why: Stacked linear layers collapse to ONE linear map (Part 5), and one linear map draws one straight boundary. XOR needs a non-convex region, so accuracy stays pinned at 0.50 no matter the size.

The fix

Add a hidden layer with a nonlinearity — even 2 hidden units suffice.

One hidden layer + activation → 100% on XOR

Why: The nonlinearity warps the input space (OR − 2·AND) so XOR becomes linearly separable in the hidden representation. Capacity comes from σ, not from parameter count.

92. Which of these survive contact with Lesson 22: Affine Transforms?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
An affine map takes a vector x, applies a matrix W, then adds a fixed vector b. The Wx rotates, scales, or shears; the + b slides the whole result.; A map T is linear when it satisfies two axioms for all vectors u, v and scalars c:; A linear map must send 0 to 0 — plug in x = 0 and W·0 = 0. An affine map sends 0 to b: the origin is free to move.
Breaks
T(x) = Wx + b is built from a matrix, so it must be a linear map.; Composing T₁ then T₂, the combined bias is surely b₁ + b₂.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 22: Affine Transforms puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

93. Rebuild the recipe: Why activations exist — the whole argument

Ranking

Put in order

These are the steps of Why activations exist — the whole argument, scrambled. Put them back in order before the next slide shows you.

  1. Each layer, pre-activation, is affine: Wx + b (linear part + shift)
  2. Composition of affines is affine → (W₂W₁)x + (W₂b₁+b₂) (verified: [−5, 28] two ways)
  3. So a linear stack collapses to one matrix → linear depth adds zero capacity (verified: 3 layers = one 2×3 matrix)
  4. A nonlinearity between layers breaks the collapse — additivity fails on real inputs (net(a+b) ≠ net(a)+net(b))
  5. That is what lets an MLP fit XOR (linear stuck at 0.5, tanh MLP at 1.0) — and, by extension, approximate anything

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

94. Why activations exist — the whole argument

Pattern

  1. Each layer, pre-activation, is affine: Wx + b (linear part + shift)
  2. Composition of affines is affine → (W₂W₁)x + (W₂b₁+b₂) (verified: [−5, 28] two ways)
  3. So a linear stack collapses to one matrix → linear depth adds zero capacity (verified: 3 layers = one 2×3 matrix)
  4. A nonlinearity between layers breaks the collapse — additivity fails on real inputs (net(a+b) ≠ net(a)+net(b))
  5. That is what lets an MLP fit XOR (linear stuck at 0.5, tanh MLP at 1.0) — and, by extension, approximate anything

95. Where does it stop working: Why activations exist — the whole argument

Edge cases

Discussion prompt

Why activations exist — the whole argument works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Each layer, pre-activation, is affine: Wx + b (linear part + shift)
  2. Composition of affines is affine → (W₂W₁)x + (W₂b₁+b₂) (verified: [−5, 28] two ways)
  3. So a linear stack collapses to one matrix → linear depth adds zero capacity (verified: 3 layers = one 2×3 matrix)
  4. A nonlinearity between layers breaks the collapse — additivity fails on real inputs (net(a+b) ≠ net(a)+net(b))
  5. That is what lets an MLP fit XOR (linear stuck at 0.5, tanh MLP at 1.0) — and, by extension, approximate anything

96. Rule out three: Check yourself — composition

Elimination

Eliminate the wrong options

Composing T₁(x)=W₁x+b₁ then T₂(x)=W₂x+b₂ gives a single affine map with weight W and bias b:

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. W = W₂W₁, b = W₂b₁ + b₂
  • B. W = W₁W₂, b = b₁ + b₂
  • C. W = W₂W₁, b = b₁ + b₂
  • D. It is not affine in general

Survives elimination: A

Why: Substitute: T₂(T₁(x)) = W₂(W₁x+b₁)+b₂ = (W₂W₁)x + (W₂b₁+b₂). Because T₁ runs first, W₁ sits on the right of the product, and b₁ is transformed by W₂ before b₂ is added.

97. Check yourself — composition

Check

Combine two affine maps in order: T₁ first, then T₂.

Check your understanding

Composing T₁(x)=W₁x+b₁ then T₂(x)=W₂x+b₂ gives a single affine map with weight W and bias b:

  • A. W = W₂W₁, b = W₂b₁ + b₂ (correct)
  • B. W = W₁W₂, b = b₁ + b₂
  • C. W = W₂W₁, b = b₁ + b₂
  • D. It is not affine in general

Answer: A

Why: Substitute: T₂(T₁(x)) = W₂(W₁x+b₁)+b₂ = (W₂W₁)x + (W₂b₁+b₂). Because T₁ runs first, W₁ sits on the right of the product, and b₁ is transformed by W₂ before b₂ is added.

Why B tempts people
Order is reversed and the bias is wrong. T₁ acts first, so its matrix W₁ must be applied first — it goes on the RIGHT: W₂W₁, not W₁W₂. And b₁ passes through W₂.
Why C tempts people
Right weight, wrong bias — the classic slip. b₁ does not just add to b₂; it is produced inside T₁ and then multiplied by W₂, giving W₂b₁ + b₂.
Why D tempts people
The composition IS affine — that's exactly what the derivation shows, and it is the reason linear networks collapse to a single layer.

98. Answer it before you see the options: Check yourself — is it linear?

Prediction

Predict first

For which maps T(x) = Wx + b is T actually a LINEAR map (not merely affine)?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: Only when b = 0

Why: A linear map must satisfy T(0) = 0, but T(0) = W·0 + b = b. So T is linear exactly when b = 0. With a nonzero bias, additivity fails: T(u+v) = Wu+Wv+b ≠ Wu+Wv+2b = T(u)+T(v).

99. Check yourself — is it linear?

Check

Recall the linearity axioms and what T(0) must be.

Check your understanding

For which maps T(x) = Wx + b is T actually a LINEAR map (not merely affine)?

  • A. Only when b = 0 (correct)
  • B. Whenever W is invertible
  • C. Always — any Wx + b is linear
  • D. Only when W is the identity

Answer: A

Why: A linear map must satisfy T(0) = 0, but T(0) = W·0 + b = b. So T is linear exactly when b = 0. With a nonzero bias, additivity fails: T(u+v) = Wu+Wv+b ≠ Wu+Wv+2b = T(u)+T(v).

Why B tempts people
Invertibility of W is unrelated to linearity. An invertible W with b ≠ 0 is still affine-not-linear because T(0) = b ≠ 0.
Why C tempts people
False whenever b ≠ 0: then T(0) = b ≠ 0, which no linear map allows. 'Linear layer' with bias is a loose name — it is really affine.
Why D tempts people
W = I with b ≠ 0 gives T(x) = x + b, a pure translation — the least linear map there is (it moves the origin). The condition is on b, not W.

100. Answer it before you see the options: Check yourself — linear collapse

Prediction

Predict first

A 5-layer network with NO activation functions has the same expressive power as:

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: a single linear (affine) layer

Why: Composition of affine maps is affine, so the 5 weight matrices multiply into one (and the biases collapse into one b_eff). The network equals a single affine layer — verified when 3 layers collapsed to one 2×3 matrix.

101. Check yourself — linear collapse

Check

How expressive is a deep network with no activations?

Check your understanding

A 5-layer network with NO activation functions has the same expressive power as:

  • A. a single linear (affine) layer (correct)
  • B. a 5-layer network with ReLU
  • C. a 2-layer ReLU network
  • D. no function at all — it outputs zero

Answer: A

Why: Composition of affine maps is affine, so the 5 weight matrices multiply into one (and the biases collapse into one b_eff). The network equals a single affine layer — verified when 3 layers collapsed to one 2×3 matrix.

Why B tempts people
Adding ReLU breaks the collapse and makes the network strictly more expressive — that is the entire point of activations. A 5-layer ReLU net is far more powerful than any linear stack.
Why C tempts people
A 2-layer ReLU net is already nonlinear, so it is more expressive than ANY purely linear stack, no matter how deep the linear stack is.
Why D tempts people
It still computes a single affine map W_eff x + b_eff — a real (if limited) function, not the zero map. It is restricted, not empty.

102. Rule out three: Check yourself — XOR

Elimination

Eliminate the wrong options

A single linear classifier cannot solve XOR because…

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. XOR is not linearly separable — no single straight line splits the two classes
  • B. XOR has too many features for a linear model
  • C. linear models can't be trained with gradient descent
  • D. XOR just needs more training data

Survives elimination: A

Why: The classes ([0,0],[1,1] vs [0,1],[1,0]) sit on opposite diagonals, so no straight boundary separates them. An affine map draws only straight boundaries, so accuracy is stuck at 0.50 — even the optimal least-squares line predicts 0.5 everywhere. A hidden nonlinearity fixes it.

103. Check yourself — XOR

Check

Why exactly does one linear layer fail here?

Check your understanding

A single linear classifier cannot solve XOR because…

  • A. XOR is not linearly separable — no single straight line splits the two classes (correct)
  • B. XOR has too many features for a linear model
  • C. linear models can't be trained with gradient descent
  • D. XOR just needs more training data

Answer: A

Why: The classes ([0,0],[1,1] vs [0,1],[1,0]) sit on opposite diagonals, so no straight boundary separates them. An affine map draws only straight boundaries, so accuracy is stuck at 0.50 — even the optimal least-squares line predicts 0.5 everywhere. A hidden nonlinearity fixes it.

Why B tempts people
XOR has only 2 features and 4 points. The obstacle is geometry (non-separability), not dimensionality — adding features you don't have is not the issue.
Why C tempts people
Linear models train fine with gradient descent; the trained linear model still scored 0.50. The limit is representational, not optimization — training cannot fix what the model can't represent.
Why D tempts people
All four XOR points are already given exactly. More copies of the same four points don't make a non-separable problem separable for a straight line.

104. Your turn: prove it yourself

Section

The project

105. Project: composition, collapse & XOR

Concept

Reproduce the three load-bearing facts of the lesson with your own hands: affine composition is affine, a linear stack collapses, and only a nonlinear net solves XOR. You have derived every piece — now assemble them.

#requirementtool
1Compose two affines; check the (W₂W₁, W₂b₁+b₂) formulanumpy @, np.allclose
2Collapse 3 linear layers into one matrixmatrix product
3XOR: linear (0.5) vs MLP (1.0)torch MLP

Build rules: type every line yourself, fix the seed (np.random.seed(0), torch.manual_seed(0)), and compare with np.allclose — the collapse must be exact, not approximate. When something errors, read the array shapes in the message.

106. Milestone 1 — compose two affines

Worked example

Your turn: verify T₂(T₁(x)) equals (W₂W₁)x + (W₂b₁+b₂). Predict out loud: exact match, or only approximate?

Hint: compute the two-step W2@(W1@x+b1)+b2 and the one-step (W2@W1)@x + (W2@b1+b2), then compare with np.allclose.

import numpy as np
W1 = np.array([[2., 0.], [1., 3.]]); b1 = np.array([1., -1.])
W2 = np.array([[1., -1.], [0., 2.]]); b2 = np.array([2., 0.])
x  = np.array([3., 4.])
direct   = W2 @ (W1 @ x + b1) + b2
composed = (W2 @ W1) @ x + (W2 @ b1 + b2)
print(direct.tolist(), composed.tolist(), np.allclose(direct, composed))
checkexpected value
direct output[−5.0, 28.0]
composed output[−5.0, 28.0]
np.allcloseTrue (exact)

107. Milestone 2 — collapse 3 linear layers

Worked example

Your turn: chain three linear layers (3→4→5→2) and confirm it equals one matrix product. Predict the effective shape before you print it.

Hint: seed first, then compare A3@(A2@(A1@x)) with (A3@A2@A1)@x; the effective W should be (2, 3).

import numpy as np
np.random.seed(0)
A1 = np.random.randn(4, 3); A2 = np.random.randn(5, 4); A3 = np.random.randn(2, 5)
x = np.random.randn(3)
stacked = A3 @ (A2 @ (A1 @ x))
single  = (A3 @ A2 @ A1) @ x
print(np.allclose(stacked, single), (A3 @ A2 @ A1).shape)
checkexpected value
stacked == singleTrue
stacked output[1.562102, 8.741537]
effective W shape(2, 3)

108. Milestone 3 — XOR, linear vs MLP

Worked example

Your turn: train a bare linear model and a tanh MLP on XOR. Predict each one's accuracy before running.

Hint: nn.Linear(2,1) stalls at 0.5; nn.Sequential(Linear(2,8), Tanh, Linear(8,1)) reaches 1.0 within a few thousand Adam steps. Seed with torch.manual_seed(0).

import torch
torch.manual_seed(0)
X = torch.tensor([[0., 0.], [0., 1.], [1., 0.], [1., 1.]])
y = torch.tensor([[0.], [1.], [1.], [0.]])
def train(m, steps=4000, lr=0.05):
    opt = torch.optim.Adam(m.parameters(), lr=lr)
    lf = torch.nn.BCEWithLogitsLoss()
    for _ in range(steps):
        opt.zero_grad(); lf(m(X), y).backward(); opt.step()
    with torch.no_grad():
        return ((torch.sigmoid(m(X)) > 0.5).float() == y).float().mean().item()
print(train(torch.nn.Linear(2, 1)))
print(train(torch.nn.Sequential(torch.nn.Linear(2, 8), torch.nn.Tanh(), torch.nn.Linear(8, 1))))
modelexpected accuracy
linear nn.Linear(2,1)0.50
MLP (tanh hidden)1.00

109. What each one costs: Milestone 3 — XOR, linear vs MLP

Trade off

Comparison matrix

From Milestone 3 — XOR, linear vs MLP: every row here is a choice with a cost. Fill the expected accuracy column, then say which row you would actually pick and what you give up for it.

modelexpected accuracy
linear nn.Linear(2,1)0.50
MLP (tanh hidden)1.00

110. The full program

Concept

import numpy as np, torch

# 1. composition of affines is affine
W1 = np.array([[2., 0.], [1., 3.]]); b1 = np.array([1., -1.])
W2 = np.array([[1., -1.], [0., 2.]]); b2 = np.array([2., 0.]); x = np.array([3., 4.])
print('composition affine:', np.allclose(W2@(W1@x+b1)+b2, (W2@W1)@x + (W2@b1+b2)))

# 2. linear depth collapses
np.random.seed(0)
A1 = np.random.randn(4, 3); A2 = np.random.randn(5, 4); A3 = np.random.randn(2, 5)
v = np.random.randn(3)
print('linear collapse:', np.allclose(A3@(A2@(A1@v)), (A3@A2@A1)@v))

# 3. XOR: linear fails, MLP solves
torch.manual_seed(0)
Xt = torch.tensor([[0., 0.], [0., 1.], [1., 0.], [1., 1.]])
yt = torch.tensor([[0.], [1.], [1.], [0.]])
def train(m, steps=4000, lr=0.05):
    opt = torch.optim.Adam(m.parameters(), lr=lr); lf = torch.nn.BCEWithLogitsLoss()
    for _ in range(steps):
        opt.zero_grad(); lf(m(Xt), yt).backward(); opt.step()
    with torch.no_grad():
        return ((torch.sigmoid(m(Xt)) > 0.5).float() == yt).float().mean().item()
la = train(torch.nn.Linear(2, 1))
ma = train(torch.nn.Sequential(torch.nn.Linear(2, 8), torch.nn.Tanh(), torch.nn.Linear(8, 1)))
print('XOR linear / MLP:', round(la, 2), '/', round(ma, 2))
printed linevalue (verified)
composition affine:True
linear collapse:True
XOR linear / MLP:0.5 / 1.0

If composition prints True, the three layers collapse to True, and only the MLP reaches 1.0 on XOR — you know exactly why activations exist.

111. Fill in: value (verified) for The full program

Comparison

Comparison matrix

From The full program: refill the value (verified) column from what you know. The rest of the table is as it appeared.

printed linevalue (verified)
composition affine:True
linear collapse:True
XOR linear / MLP:0.5 / 1.0

112. Show it off

Concept

Slides closed, out loud: explain (1) why composing affines keeps the Wx + b form and why b₁ gets multiplied by W₂, (2) why that collapses a linear network to one layer, and (3) what the nonlinearity does for XOR geometrically.

Stretch: prove by induction that any depth-N linear network equals a single layer, and swap the tanh MLP's hidden width from 8 down to 2 — confirm it still hits 1.0. Affine maps return as attention's QKV projections (Week 27) and positional encodings (Week 28), so this Wx + b skeleton is everywhere ahead.

113. Connect it up: Lesson 22: Affine Transforms

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — The map: Wx + b · The geometry of W · Composition stays affine · One matrix: homogeneous coordinates · The collapse · Why nonlinearity. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

114. What you can do now

Recap

ideathe one thing to remember
affine mapWx + b: linear part + shift; linear only if b = 0
compositionaffine ∘ affine = affine — b₁ rides through W₂
homogeneousappend a 1; composition becomes M₂M₁
linear depthcollapses to one matrix → adds no capacity
nonlinearitybreaks the collapse — XOR needs it (0.5 → 1.0)

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 22 (Week 8 - Affine Transforms) — Barron · USAAIO Round 2 Preparation, 2026
  2. Gilbert Strang, Introduction to Linear Algebra, Ch. 7-8 (Linear Transformations) — Wellesley-Cambridge Press, 5th ed.
  3. Minsky & Papert, Perceptrons (the XOR limitation of linear classifiers) — MIT Press, 1969
  4. Goodfellow, Bengio, Courville, Deep Learning, Ch. 6.1 (XOR & the need for nonlinearity)
  5. Every composition, collapse, and XOR accuracy produced by real execution — numpy 2.2.6 + torch 2.7.1, seeded run, July 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108