Lesson 18: Full Backpropagation

USAAIO Lesson 18, from Week 6, fully worked. One tiny 3-2-2 network with fixed weights is carried end to end. The forward pass is built operation by operation - affine, ReLU, affine, softmax, cross-entropy - and then the complete backward pass is derived one gradient per beat, covering the softmax-and-cross-entropy shortcut, the outer-product weight gradient, and the ReLU mask that kills a dead unit. Every number is matched against PyTorch autograd and against a finite-difference gradient check. It then covers vanishing and exploding gradients with clip-by-norm and He initialization, and caps off by training an MLP to about 97% on the sklearn digits dataset, along with a from-scratch NumPy network. Every snippet runs standalone, and every number was produced by real execution. The lesson runs to 60 slides.

Subject: Machine Learning · 109 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Full Backpropagation

Title

USAAIO · Lesson 18 · Week 6

One tiny network, 3 → 2 → 2, carried end to end. We build the forward pass op by op, then derive every gradient of the backward pass one move at a time — and check each number against PyTorch. Then the pathologies, and a net you train to ~97%.

2. By the end of this lesson you can

Objectives

  1. Run a complete forward pass by hand: affine → ReLU → affine → softmax → cross-entropy, caching every intermediate
  2. Derive the complete backward pass one gradient per step — the softmax+CE seed p − e_y, the outer-product dW = δ aᵀ, and the ReLU mask 1[z>0]
  3. Match every hand-computed gradient against PyTorch autograd and a finite-difference gradient check
  4. Read a network as a computational graph and explain why backprop is one reverse sweep
  5. Diagnose vanishing / exploding gradients and fix them with clip-by-norm and He init
  6. Train an MLP to ~97% on real data, in PyTorch and from scratch in NumPy

3. What survived from Loss Functions & Their Gradients?

Warm-up

Discussion prompt

Before we open Lesson 18: Full Backpropagation: without looking back, what was the main idea of Loss Functions & Their Gradients, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

MSE, MAE, binary and multi-class cross-entropy, and hinge loss with their gradients — the clean dL/dz = p − y result for softmax+cross-entropy and sigmoid+BCE, and why MSE is a poor classification loss.

4. The running network

Section

Part 1 of 7 — set up one concrete example

5. One network, fixed numbers, all lesson

Concept

We fix one network and one training example and never change them. Every derivation below produces a number you can check. The net maps a 3-vector input to 2 class scores: affine → ReLU → affine → softmax → cross-entropy.

\[ x \in \mathbb{R}^3 \;\xrightarrow{W_1,b_1}\; z_1 \;\xrightarrow{\text{ReLU}}\; a_1 \;\xrightarrow{W_2,b_2}\; z_2 \;\xrightarrow{\text{softmax}}\; p \;\xrightarrow{\text{CE}}\; L \]

Two layers of weights (W₁, W₂), one hidden ReLU layer with 2 units, 2 output classes. Small enough to do entirely by hand — big enough to show every real subtlety.

6. Break it if you can: One network, fixed numbers, all lesson

Counterexample

Discussion prompt

Two layers of weights (W₁, W₂), one hidden ReLU layer with 2 units, 2 output classes. Small enough to do entirely by hand — big enough to show every real subtlety.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

7. Forward asks, backward blames

Intuition

Two passes, two questions. The forward pass asks what does the network predict? — it runs the data through and produces a loss. The backward pass asks who is to blame for that loss, and by how much? — it hands each weight a number saying which way to nudge it.

That blame-number is a partial derivative ∂L/∂W. Backprop is just the chain rule organized so we compute every one of them in a single sweep from the loss back to the inputs, reusing shared work instead of redoing it.

The whole lesson is those two passes on one tiny net — small enough that every blame-number is a fraction you can check by hand.

8. By analogy: Forward asks, backward blames

Analogy

Discussion prompt

Explain Forward asks, backward blames by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

The whole lesson is those two passes on one tiny net — small enough that every blame-number is a fraction you can check by hand.

9. The weights and the input

Concept

Here are the exact numbers. W₁ is (2×3) — 2 hidden units, each reading the 3 inputs. W₂ is (2×2) — 2 class scores from the 2 hidden units. The true label is class 0.

\[ W_1 = \begin{bmatrix} 1 & 1 & 0 \\ 1 & -1 & 0 \end{bmatrix},\; b_1 = \begin{bmatrix} 0 \\ 0 \end{bmatrix},\quad W_2 = \begin{bmatrix} 1 & 2 \\ -1 & 3 \end{bmatrix},\; b_2 = \begin{bmatrix} -2 \\ 5 \end{bmatrix} \]

\[ x = \begin{bmatrix} 1 \\ 2 \\ 1 \end{bmatrix}, \qquad y = 0 \;\;(\text{true class}) \]

10. Why cache every intermediate

Intuition

The backward pass reuses the forward pass's leftovers. To get the weight gradient at a layer you need that layer's input activation; to cross a ReLU you need its pre-activation z. Throw them away and you'd have to recompute the whole forward pass at every step.

So the rule is: on the way forward, stash z₁, a₁, z₂, p. On the way back, spend them. This forward-cache-then-backward-sweep is exactly what an autograd engine does under the hood.

11. The forward pass

Section

Part 2 of 7 — op by op

12. Guess the shape of the answer: Layer 1 affine: z₁ = W₁x + b₁

Estimation

Predict first

First affine map. Each hidden unit dots its weight row with x, then adds its bias. Two units, so two dot products:

Commit before you compute: what does Layer 1 affine: z₁ = W₁x + b₁ come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Unit 1: (1)(1) + (−1)(2) + (0)(1) + 0 = −1

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Row 1 of W₁ is [1, −1, 0]; dot with x gives 1 − 2 + 0 = −1.

13. Layer 1 affine: z₁ = W₁x + b₁

Worked example

First affine map. Each hidden unit dots its weight row with x, then adds its bias. Two units, so two dot products:

Unit 0: (1)(1) + (1)(2) + (0)(1) + 0 = 3

Why: Row 0 of W₁ is [1, 1, 0]; dot with x = [1, 2, 1] gives 1 + 2 + 0 = 3, plus bias 0.

\[ z_{1,0} = 1\cdot1 + 1\cdot2 + 0\cdot1 + 0 = 3 \]

Unit 1: (1)(1) + (−1)(2) + (0)(1) + 0 = −1

Why: Row 1 of W₁ is [1, −1, 0]; dot with x gives 1 − 2 + 0 = −1. This unit's pre-activation is NEGATIVE — remember that.

\[ z_1 = \begin{bmatrix} 3 \\ -1 \end{bmatrix} \]

import numpy as np
W1 = np.array([[1.,1.,0.],[1.,-1.,0.]])   # (2,3)
b1 = np.array([0., 0.])
x  = np.array([1., 2., 1.])               # one example
z1 = W1 @ x + b1
print(z1)   # [3. -1.]
hidden unitweight row · x+ biasz₁
01+2+0 = 303
11−2+0 = −10−1
shape(2×3)·(3,)(2,)(2,)

14. Fill in: weight row · x for Layer 1 affine: z₁ = W₁x + b₁

Comparison

Comparison matrix

From Layer 1 affine: z₁ = W₁x + b₁: refill the weight row · x column from what you know. The rest of the table is as it appeared.

hidden unitweight row · x+ biasz₁
01+2+0 = 303
11−2+0 = −10−1
shape(2×3)·(3,)(2,)(2,)

15. What has to be given first: ReLU: a₁ = max(0, z₁)

Missing information

Discussion prompt

The nonlinearity. ReLU passes positives through untouched and clamps negatives to 0. Apply it entrywise to z₁ = [3, −1]:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

Positive pre-activation survives — this unit is ACTIVE and will carry gradient on the way back.

16. ReLU: a₁ = max(0, z₁)

Worked example

The nonlinearity. ReLU passes positives through untouched and clamps negatives to 0. Apply it entrywise to z₁ = [3, −1]:

Unit 0: max(0, 3) = 3

Why: Positive pre-activation survives — this unit is ACTIVE and will carry gradient on the way back.

Unit 1: max(0, −1) = 0

Why: Negative pre-activation is clamped to 0 — this unit is DEAD for this input. Its gradient path will be cut in the backward pass.

\[ a_1 = \max(0, z_1) = \begin{bmatrix} \max(0,3) \\ \max(0,-1) \end{bmatrix} = \begin{bmatrix} 3 \\ 0 \end{bmatrix} \]

import numpy as np
z1 = np.array([3., -1.])
a1 = np.maximum(0, z1)     # ReLU, entrywise
print(a1)   # [3. 0.]
unitz₁ReLU(z₁) = a₁status
033active
1−10dead (clamped)

17. What each one costs: ReLU: a₁ = max(0, z₁)

Trade off

Comparison matrix

From ReLU: a₁ = max(0, z₁): every row here is a choice with a cost. Fill the status column, then say which row you would actually pick and what you give up for it.

unitz₁ReLU(z₁) = a₁status
033active
1−10dead (clamped)

18. Predict the next row: Layer 2 affine: z₂ = W₂a₁ + b₂

Pattern

Predict first

The table runs: 0 | 3+0 = 3 | −2 | 1 · 1 | −3+0 = −3 | 5 | 2

In Layer 2 affine: z₂ = W₂a₁ + b₂, given the rows so far: what is the next one — the row where class is note?

Correct: note | a₁[1]=0 kills col 2 | — | —

classrow · a₁+ biasz₂ (logit)
03+0 = 3−21
1−3+0 = −352
notea₁[1]=0 kills col 2——

Why: The relationship between the columns, not the individual numbers, is what generates the next row. Row 0 of W₂ is [1, 2]; dot with a₁ = [3, 0] gives 3 + 0 = 3, plus bias −2 → 1.

19. Layer 2 affine: z₂ = W₂a₁ + b₂

Worked example

Second affine map turns the 2 hidden activations a₁ = [3, 0] into 2 class scores (logits). Same pattern: dot each row of W₂ with a₁, add the bias.

Class 0: (1)(3) + (2)(0) + (−2) = 1

Why: Row 0 of W₂ is [1, 2]; dot with a₁ = [3, 0] gives 3 + 0 = 3, plus bias −2 → 1.

\[ z_{2,0} = 1\cdot3 + 2\cdot0 - 2 = 1 \]

Class 1: (−1)(3) + (3)(0) + 5 = 2

Why: Row 1 of W₂ is [−1, 3]; dot with a₁ gives −3 + 0 = −3, plus bias 5 → 2.

\[ z_2 = \begin{bmatrix} 1 \\ 2 \end{bmatrix} \]

import numpy as np
W2 = np.array([[1.,2.],[-1.,3.]])   # (2,2)
b2 = np.array([-2., 5.])
a1 = np.array([3., 0.])
z2 = W2 @ a1 + b2
print(z2)   # [1. 2.]
classrow · a₁+ biasz₂ (logit)
03+0 = 3−21
1−3+0 = −352
notea₁[1]=0 kills col 2——

20. Decode the notation: Layer 2 affine: z₂ = W₂a₁ + b₂

Notation

Annotate

From Layer 2 affine: z₂ = W₂a₁ + b₂ — read this one piece at a time. What is each part doing?

On: \( z_{2,0} = 1\cdot3 + 2\cdot0 - 2 = 1 \)

  • Row 0 of W₂ is [1, 2]; dot with a₁ = [3, 0] gives 3 + 0 = 3, plus bias −2 → 1.
  • Row 1 of W₂ is [−1, 3]; dot with a₁ gives −3 + 0 = −3, plus bias 5 → 2.

21. Softmax turns logits into probabilities

Concept

Logits z₂ are unbounded real numbers. Softmax exponentiates then normalizes, producing a probability vector p that is positive and sums to 1:

\[ p_k = \frac{e^{z_{2,k}}}{\sum_j e^{z_{2,j}}} \]

In code we subtract max(z₂) inside the exponent first — it cancels in the ratio but stops exp from overflowing on large logits. A must-have habit.

22. Teach it back: Softmax turns logits into probabilities

Explain it

Discussion prompt

Explain Softmax turns logits into probabilities to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Logits z₂ are unbounded real numbers. Softmax exponentiates then normalizes, producing a probability vector p that is positive and sums to 1:

23. Restore the missing line: Softmax on z₂ = [1, 2]

Fill the middle

Fill in the blanks

From Softmax on z₂ = [1, 2] — one line has had its right-hand side removed. Put it back.

import numpy as np
z2 = np.array([1., 2.])
p = np.exp(z2 - z2.max()) # subtract max for numerical safety
p = p / p.sum()
print(np.round(p, 4)) # [0.2689 0.7311]

Why: p is what everything below it consumes, so the wrong expression here fails later and somewhere else. Raise e to each logit. These are the unnormalized weights of each class.

24. Softmax on z₂ = [1, 2]

Worked example

Exponentiate: e¹ ≈ 2.7183, e² ≈ 7.3891

Why: Raise e to each logit. These are the unnormalized weights of each class.

\[ e^{z_2} = \begin{bmatrix} e^1 \\ e^2 \end{bmatrix} \approx \begin{bmatrix} 2.7183 \\ 7.3891 \end{bmatrix}, \quad \textstyle\sum = 10.1073 \]

Normalize by the sum 10.1073

Why: Divide each by the total so they sum to 1. p₀ = 2.7183/10.1073, p₁ = 7.3891/10.1073.

\[ p = \begin{bmatrix} 0.2689 \\ 0.7311 \end{bmatrix} \]

import numpy as np
z2 = np.array([1., 2.])
p = np.exp(z2 - z2.max())   # subtract max for numerical safety
p = p / p.sum()
print(np.round(p, 4))   # [0.2689 0.7311]
class kz₂e^z₂p = e^z₂ / Σ
012.71830.2689
127.38910.7311
Σ—10.10731.0000

25. Inspect it line by line: Softmax on z₂ = [1, 2]

Error analysis

Annotate

Walk the callouts on Softmax on z₂ = [1, 2]. Each one is a place this is easy to get subtly wrong.

  • Raise e to each logit. These are the unnormalized weights of each class.
  • Divide each by the total so they sum to 1. p₀ = 2.7183/10.1073, p₁ = 7.3891/10.1073.

26. Finish it with less help: Cross-entropy: the loss

Faded example

Fill in the blanks

Cross-entropy: the loss, with the scaffolding fading: two lines are gone now — fill both.

import numpy as np
z2 = np.array([1., 2.]); y = 0
p = np.exp(z2 - z2.max()); p = p / p.sum()
L = -np.log(p[y]) # cross-entropy = -log prob of true class
print(round(L, 4)) # 1.3133

Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. The model is confident in the WRONG class (0.7311 on class 1), so the loss is substantial.

27. Cross-entropy: the loss

Worked example

Cross-entropy for a single label is just the negative log-probability the model assigns to the true class. Our true class is 0, and the model gave it only p₀ = 0.2689:

\[ L = -\log p_{y} = -\log p_0 = -\log(0.2689) \]

L = −log(0.2689) ≈ 1.3133

Why: The model is confident in the WRONG class (0.7311 on class 1), so the loss is substantial. A perfect p₀ = 1 would give L = 0.

\[ \boxed{\,L \approx 1.3133\,} \]

import numpy as np
z2 = np.array([1., 2.]); y = 0
p = np.exp(z2 - z2.max()); p = p / p.sum()
L = -np.log(p[y])          # cross-entropy = -log prob of true class
print(round(L, 4))   # 1.3133
quantityvalue
p (softmax)[0.2689, 0.7311]
true class y0
p[y]0.2689
L = −log p[y]1.3133

28. Work backwards from the answer: Cross-entropy: the loss

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

L = −log(0.2689) ≈ 1.3133

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Cross-entropy for a single label is just the negative log-probability the model assigns to the true class. Our true class is 0, and the model gave it only p₀ = 0.2689:

29. Forward pass complete — the cache

Concept

That is the whole forward pass. Everything we stashed is on the table. The backward pass will consume p (to seed the gradient), a₁ and x (for weight gradients), and z₁ (for the ReLU mask).

cachedvalueused backward for
z₁[3, −1]ReLU mask 1[z₁>0]
a₁[3, 0]dW₂ = δ₂ a₁ᵀ
z₂[1, 2](produced p)
p[0.2689, 0.7311]seed δ₂ = p − e_y
x[1, 2, 1]dW₁ = δ₁ xᵀ
import numpy as np
W1 = np.array([[1.,1.,0.],[1.,-1.,0.]]); b1 = np.array([0.,0.])
W2 = np.array([[1.,2.],[-1.,3.]]);       b2 = np.array([-2.,5.])
x  = np.array([1.,2.,1.])
z1 = W1 @ x + b1; a1 = np.maximum(0, z1)
z2 = W2 @ a1 + b2
p  = np.exp(z2 - z2.max()); p = p / p.sum()
print(z1, a1, z2, np.round(p, 4))   # [3 -1] [3 0] [1 2] [0.2689 0.7311]

30. The backward pass

Section

Part 3 of 7 — one gradient per beat

31. The chain rule, applied in reverse

Concept

Backprop wants ∂L/∂W₁ and ∂L/∂W₂. The chain rule says a gradient at any node is the gradient just downstream times that node's local derivative. So we start at the loss and walk backward, multiplying local derivatives as we go.

We write δ_ℓ = ∂L/∂z_ℓ for the gradient at a layer's pre-activation. Each layer's job: turn the downstream δ into (a) its weight gradient and (b) the δ for the layer before it. Two moves, repeated.

32. Why softmax+CE collapses to p − e_y

Concept

The seed of the whole backward pass is ∂L/∂z₂. Softmax and cross-entropy look nasty separately, but composed they simplify dramatically. Cross-entropy for true class y is L = −log p_y, and p_y = e^{z_y} / Σ_j e^{z_j}, so L = −z_y + log Σ_j e^{z_j}.

Differentiate L = −z_y + log Σ_j e^{z_j} with respect to one logit z_k. The first term contributes −1 only when k = y; the log-sum term contributes e^{z_k}/Σ = p_k for every k.

\[ \frac{\partial L}{\partial z_k} = p_k - \mathbb{1}[k = y] \]

Stacked over all k, that indicator vector is exactly the one-hot e_y. So the whole gradient is p − e_y — no Jacobian to multiply, just a subtraction.

33. Seed: δ₂ = p − e_y (softmax+CE shortcut)

Worked example

The gradient of softmax-then-cross-entropy with respect to the logits collapses to a famously clean form (derived last slide): predicted minus one-hot true.

\[ \delta_2 = \frac{\partial L}{\partial z_2} = p - e_y \]

e_y = [1, 0] since true class is 0

Why: The one-hot vector puts a 1 in the true class slot. Subtract it from p = [0.2689, 0.7311].

\[ \delta_2 = \begin{bmatrix} 0.2689 \\ 0.7311 \end{bmatrix} - \begin{bmatrix} 1 \\ 0 \end{bmatrix} = \begin{bmatrix} -0.7311 \\ 0.7311 \end{bmatrix} \]

import numpy as np
z2 = np.array([1., 2.]); y = 0
p = np.exp(z2 - z2.max()); p = p / p.sum()
onehot = np.zeros(2); onehot[y] = 1.0
dz2 = p - onehot           # the softmax+CE shortcut
print(np.round(dz2, 4))   # [-0.7311  0.7311]
classpe_yδ₂ = p − e_y
0 (true)0.26891−0.7311
10.73110+0.7311
readingtoo low → push up—negative = raise score

34. Where does each piece belong: Lesson 18: Full Backpropagation

Sorting

Sort into buckets

These are the pieces of Lesson 18: Full Backpropagation, out of order. Put each one back under the part of the lesson it belongs to.

The running network
One network, fixed numbers, all lesson; Forward asks, backward blames; The weights and the input
The forward pass
Layer 1 affine: z₁ = W₁x + b₁; ReLU: a₁ = max(0, z₁); Layer 2 affine: z₂ = W₂a₁ + b₂
The backward pass
The chain rule, applied in reverse; Why softmax+CE collapses to p − e_y; Seed: δ₂ = p − e_y (softmax+CE shortcut)
s1
The running network is where Lesson 18: Full Backpropagation puts One network, fixed numbers, all lesson, Forward asks, backward blames, The weights and the input. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
The forward pass is where Lesson 18: Full Backpropagation puts Layer 1 affine: z₁ = W₁x + b₁, ReLU: a₁ = max(0, z₁), Layer 2 affine: z₂ = W₂a₁ + b₂. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
The backward pass is where Lesson 18: Full Backpropagation puts The chain rule, applied in reverse, Why softmax+CE collapses to p − e_y, Seed: δ₂ = p − e_y (softmax+CE shortcut). Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

35. The weight gradient is an outer product

Concept

For a linear layer z = Wa + b, entry ∂L/∂W_{jk} picks up δ_j (error at output j) times a_k (the input it multiplied). Collecting all j, k is exactly the outer product:

\[ \frac{\partial L}{\partial W} = \delta\, a^\top, \qquad \frac{\partial L}{\partial b} = \delta \]

Shape check: δ is (out,), a is (in,), so δ aᵀ is (out×in) — exactly the shape of W. If your dW is not the same shape as W, you have the outer product backwards.

36. dW₂ = δ₂ a₁ᵀ

Worked example

Apply the outer product at layer 2 with δ₂ = [−0.7311, 0.7311] and a₁ = [3, 0]. Each entry is δ₂[j] · a₁[k]:

Column for a₁[0]=3: [−0.7311·3, 0.7311·3] = [−2.1932, 2.1932]

Why: The active hidden unit (a₁[0]=3) receives real gradient — these weights get updated.

Column for a₁[1]=0: [0, 0]

Why: The DEAD unit contributed 0 to the forward pass, so the weights feeding OUT of it get zero gradient this step — you can't blame a unit that said nothing.

\[ dW_2 = \begin{bmatrix} -0.7311 \\ 0.7311 \end{bmatrix}\begin{bmatrix} 3 & 0 \end{bmatrix} = \begin{bmatrix} -2.1932 & 0 \\ 2.1932 & 0 \end{bmatrix} \]

import numpy as np
a1  = np.array([3., 0.])           # hidden activations, cached
p   = np.array([0.26894142, 0.73105858])  # softmax output, cached
dz2 = p - np.eye(2)[0]             # exact upstream grad at z2
dW2 = np.outer(dz2, a1)            # (2,) outer (2,) -> (2,2)
db2 = dz2
print(np.round(dW2, 4))   # [[-2.1932 0.] [2.1932 0.]]
print(dW2.shape)   # (2, 2), matches W2
j (out)k=0 (a₁=3)k=1 (a₁=0)row of dW₂
0−0.7311·3 = −2.19320[−2.1932, 0]
10.7311·3 = 2.19320[2.1932, 0]
db₂= δ₂—[−0.7311, 0.7311]

37. Draw the shape of it: dW₂ = δ₂ a₁ᵀ

Blank canvas

Draw it

Draw what dW₂ = δ₂ a₁ᵀ just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.

38. Why the transpose appears in the push-back

Concept

Forward, z = Wa sends each input a_k to every output via column k of W. Backward, the gradient at a_k must collect contributions from every output it fed — which means reading W down its columns, i.e. across the rows of Wᵀ.

\[ z_j = \sum_k W_{jk}\, a_k \;\Longrightarrow\; \frac{\partial L}{\partial a_k} = \sum_j W_{jk}\,\delta_j = (W^\top \delta)_k \]

So the local Jacobian of an affine layer is W going forward and Wᵀ coming back. The transpose is not a trick to memorize — it is the chain rule summing over the right index. Shape check: Wᵀ is (in×out), δ is (out,), product is (in,) — the shape of a.

39. Restore the missing line: Push back to a₁: da₁ = W₂ᵀ δ₂

Fill the middle

Fill in the blanks

From Push back to a₁: da₁ = W₂ᵀ δ₂ — one line has had its right-hand side removed. Put it back.

import numpy as np
W2 = np.array([[1.,2.],[-1.,3.]])
p = np.array([0.26894142, 0.73105858])
dz2 = p - np.eye(2)[0] # exact upstream grad at z2
da1 = W2.T @ dz2 # gradient arriving at a1
print(np.round(da1, 4)) # [-1.4621 0.7311]

Why: dz2 is what everything below it consumes, so the wrong expression here fails later and somewhere else. Row 0: 1·(−0.7311) + (−1)·(0.7311) = −1.4621.

40. Push back to a₁: da₁ = W₂ᵀ δ₂

Worked example

To keep going we need the gradient at the hidden activations a₁. The transpose of the layer's weight matrix carries δ₂ backward across the affine map:

\[ da_1 = \frac{\partial L}{\partial a_1} = W_2^\top \delta_2 \]

W₂ᵀ = [[1, −1], [2, 3]]; multiply by δ₂ = [−0.7311, 0.7311]

Why: Row 0: 1·(−0.7311) + (−1)·(0.7311) = −1.4621. Row 1: 2·(−0.7311) + 3·(0.7311) = 0.7311.

\[ da_1 = \begin{bmatrix} 1 & -1 \\ 2 & 3 \end{bmatrix}\begin{bmatrix} -0.7311 \\ 0.7311 \end{bmatrix} = \begin{bmatrix} -1.4621 \\ 0.7311 \end{bmatrix} \]

Note da₁[1] = 0.7311 is nonzero — gradient wants to flow into the dead unit. The ReLU on the next slide is what stops it.

import numpy as np
W2  = np.array([[1.,2.],[-1.,3.]])
p   = np.array([0.26894142, 0.73105858])
dz2 = p - np.eye(2)[0]     # exact upstream grad at z2
da1 = W2.T @ dz2           # gradient arriving at a1
print(np.round(da1, 4))   # [-1.4621  0.7311]
unitW₂ᵀ row · δ₂da₁
01(−0.7311) + (−1)(0.7311)−1.4621
12(−0.7311) + 3(0.7311)+0.7311
shape(2×2)·(2,)(2,)

41. Say it in words: Push back to a₁: da₁ = W₂ᵀ δ₂

Translation

\( da_1 = \frac{\partial L}{\partial a_1} = W_2^\top \delta_2 \)

Draw it

Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.

42. A dead unit is a closed gate

Intuition

Think of each ReLU as a gate. When its input was positive the gate was open: nudging the input nudges the output one-for-one, so gradient passes straight through (slope 1). When its input was negative the gate was shut: the output was flatly 0, and a tiny nudge still leaves it 0 (slope 0).

A shut gate can't leak. If a unit contributed nothing to the prediction, it deserves no blame for the error — so its gradient is zeroed. That's the entire content of the ReLU mask.

43. Complete the line: Cross the ReLU: δ₁ = da₁ ⊙ 1[z₁>0]

Fill the middle

Fill in the blanks

From Cross the ReLU: δ₁ = da₁ ⊙ 1[z₁>0] — finish the line. Write what belongs on the right of the equals sign before you look.

\delta_1 = da_1 \odot \mathbb{1}[z_1 > 0]

Why: Producing the right-hand side unprompted is the difference between recognising this line and being able to use it. z₁[0]=3>0 → 1 (gradient passes).

44. Cross the ReLU: δ₁ = da₁ ⊙ 1[z₁>0]

Worked example

ReLU's local derivative is 1 where its input was positive and 0 where it was negative. Crossing it means multiplying the incoming gradient by this mask — computed from the cached z₁, not a₁.

\[ \delta_1 = da_1 \odot \mathbb{1}[z_1 > 0] \]

Mask from z₁ = [3, −1] is [1, 0]

Why: z₁[0]=3>0 → 1 (gradient passes). z₁[1]=−1<0 → 0 (gradient blocked). This is the dead unit paying off.

δ₁ = [−1.4621, 0.7311] ⊙ [1, 0] = [−1.4621, 0]

Why: The 0.7311 that wanted to flow into unit 1 is ZEROED. A dead ReLU has zero slope, so no gradient reaches its incoming weights — the single most-forgotten step in manual backprop.

\[ \delta_1 = \begin{bmatrix} -1.4621 \\ 0 \end{bmatrix} \]

import numpy as np
W2  = np.array([[1.,2.],[-1.,3.]])
z1  = np.array([3., -1.])           # unit 1 is negative -> dead
p   = np.array([0.26894142, 0.73105858])
dz2 = p - np.eye(2)[0]             # exact upstream grad at z2
da1 = W2.T @ dz2
dz1 = da1 * (z1 > 0)               # multiply by the ReLU mask
print(np.round(dz1, 4))   # [-1.4621  0.    ]
unitda₁1[z₁>0]δ₁ = da₁ ⊙ mask
0−1.46211−1.4621
10.731100 (killed)

45. Inspect it line by line: Cross the ReLU: δ₁ = da₁ ⊙ 1[z₁>0]

Error analysis

Annotate

Walk the callouts on Cross the ReLU: δ₁ = da₁ ⊙ 1[z₁>0]. Each one is a place this is easy to get subtly wrong.

  • z₁[0]=3>0 → 1 (gradient passes). z₁[1]=−1<0 → 0 (gradient blocked). This is the dead unit paying off.
  • The 0.7311 that wanted to flow into unit 1 is ZEROED. A dead ReLU has zero slope, so no gradient reaches its incoming weights — the single most-forgotten step in manual backprop.

46. Guess the shape of the answer: dW₁ = δ₁ xᵀ

Estimation

Predict first

Same outer-product rule as layer 2, now with δ₁ = [−1.4621, 0] and the input x = [1, 2, 1]. This gives the (2×3) gradient for W₁:

Commit before you compute: what does dW₁ = δ₁ xᵀ come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Row 1 (dead unit): 0 · [1, 2, 1] = [0, 0, 0]

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. δ₁[1] = 0, so every weight feeding the dead unit gets zero gradient — it learns nothing from this example.

47. dW₁ = δ₁ xᵀ

Worked example

Same outer-product rule as layer 2, now with δ₁ = [−1.4621, 0] and the input x = [1, 2, 1]. This gives the (2×3) gradient for W₁:

Row 0 (active unit): −1.4621 · [1, 2, 1] = [−1.4621, −2.9242, −1.4621]

Why: The active hidden unit's incoming weights all get real gradient, scaled by each input feature.

Row 1 (dead unit): 0 · [1, 2, 1] = [0, 0, 0]

Why: δ₁[1] = 0, so every weight feeding the dead unit gets zero gradient — it learns nothing from this example.

\[ dW_1 = \begin{bmatrix} -1.4621 \\ 0 \end{bmatrix}\begin{bmatrix} 1 & 2 & 1 \end{bmatrix} = \begin{bmatrix} -1.4621 & -2.9242 & -1.4621 \\ 0 & 0 & 0 \end{bmatrix} \]

import numpy as np
x   = np.array([1., 2., 1.])        # the input
dz1 = np.array([-1.4621, 0.])       # grad at z1 (dead unit zeroed)
dW1 = np.outer(dz1, x)             # (2,) outer (3,) -> (2,3)
db1 = dz1
print(np.round(dW1, 4))
print(dW1.shape)   # (2, 3), matches W1
j·x[0]=1·x[1]=2·x[2]=1
0−1.4621−2.9242−1.4621
1000
db₁= δ₁→[−1.4621, 0]

48. Fill in: ·x[1]=2 for dW₁ = δ₁ xᵀ

Comparison

Comparison matrix

From dW₁ = δ₁ xᵀ: refill the ·x[1]=2 column from what you know. The rest of the table is as it appeared.

j·x[0]=1·x[1]=2·x[2]=1
0−1.4621−2.9242−1.4621
1000
db₁= δ₁→[−1.4621, 0]

49. The full gradient, one table

Worked example

The backward pass is done. Here is every gradient we produced, in the order the sweep computed them — from the loss back to the first layer's weights:

The sweep: δ₂ → dW₂, db₂ → da₁ → δ₁ → dW₁, db₁

Why: Each arrow is one chain-rule move. Nothing was skipped, and each result fed the next. This ordering is the reverse topological sweep — the same order autograd uses.

gradientvalueshape
δ₂ = p − e_y[−0.7311, 0.7311](2,)
dW₂ = δ₂ a₁ᵀ[[−2.1932, 0], [2.1932, 0]](2×2)
db₂ = δ₂[−0.7311, 0.7311](2,)
δ₁ = (W₂ᵀδ₂)⊙1[z₁>0][−1.4621, 0](2,)
dW₁ = δ₁ xᵀ[[−1.4621, −2.9242, −1.4621], [0,0,0]](2×3)

Every dW has the same shape as its W — the fastest sanity check there is. Two of the five rows are (partly) zero, and both traced to the one dead unit.

50. Which is which, by shape

Discrimination

Sort into buckets

Sort these by shape, from memory, without looking back at The full gradient, one table. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

(2,)
δ₂ = p − e_y; db₂ = δ₂; δ₁ = (W₂ᵀδ₂)⊙1[z₁>0]
(2×2)
dW₂ = δ₂ a₁ᵀ
(2×3)
dW₁ = δ₁ xᵀ
g1
shape is "(2,)" for δ₂ = p − e_y, db₂ = δ₂, δ₁ = (W₂ᵀδ₂)⊙1[z₁>0] — that is what the table on "The full gradient, one table" records, and it is the single property separating this group from the rest.
g2
shape is "(2×2)" for dW₂ = δ₂ a₁ᵀ — that is what the table on "The full gradient, one table" records, and it is the single property separating this group from the rest.
g3
shape is "(2×3)" for dW₁ = δ₁ xᵀ — that is what the table on "The full gradient, one table" records, and it is the single property separating this group from the rest.

51. Something is wrong here: forgetting the ReLU mask

Anomaly

Predict first

A student writes this, and it looks reasonable:

Gradient just flows from a₁ back to z₁ unchanged — an affine map moved it, so set δ₁ = W₂ᵀδ₂ and carry on.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: That leftover 0.7311 flows into the DEAD unit, so dW₁ row 1 becomes [0.7311, 1.4622, 0.7311] instead of zeros.

Multiply by ReLU's local derivative — 1 where z₁ > 0, else 0 — before forming dW₁.

Why: That leftover 0.7311 flows into the DEAD unit, so dW₁ row 1 becomes [0.7311, 1.4622, 0.7311] instead of zeros. The shapes still match, so nothing errors — it silently disagrees with autograd.

52. Trap: forgetting the ReLU mask

Trap

The trap

Gradient just flows from a₁ back to z₁ unchanged — an affine map moved it, so set δ₁ = W₂ᵀδ₂ and carry on.

\[ \delta_1 = W_2^\top \delta_2 = \begin{bmatrix} -1.4621 \\ \mathbf{0.7311} \end{bmatrix} \;\;(\text{no mask}) \]

dW₁ picks up a bogus second row

Why: That leftover 0.7311 flows into the DEAD unit, so dW₁ row 1 becomes [0.7311, 1.4622, 0.7311] instead of zeros. The shapes still match, so nothing errors — it silently disagrees with autograd.

The fix

Multiply by ReLU's local derivative — 1 where z₁ > 0, else 0 — before forming dW₁.

\[ \delta_1 = (W_2^\top \delta_2) \odot \mathbb{1}[z_1>0] = \begin{bmatrix} -1.4621 \\ \mathbf{0} \end{bmatrix} \]

dW₁ row 1 is correctly all zeros

Why: The mask zeros the gradient for the inactive unit, matching autograd exactly. Use the cached z₁ for the mask, never a₁ — a₁ already lost the sign information.

53. Break it on purpose: forgetting the ReLU mask

Break the constraint

Discussion prompt

The rule this trap just fixed:

The mask zeros the gradient for the inactive unit, matching autograd exactly. Use the cached z₁ for the mask, never a₁ — a₁ already lost the sign information.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

That leftover 0.7311 flows into the DEAD unit, so dW₁ row 1 becomes [0.7311, 1.4622, 0.7311] instead of zeros. The shapes still match, so nothing errors — it silently disagrees with autograd.

54. Verify everything

Section

Part 4 of 7 — autograd & gradient check

55. What the gradients are for: one step

Concept

The gradients aren't the goal — they tell us which way to move the weights to shrink the loss. One gradient-descent step subtracts a small multiple of each gradient:

\[ W_\ell \leftarrow W_\ell - \eta\, dW_\ell, \qquad b_\ell \leftarrow b_\ell - \eta\, db_\ell \]

For example W₂[0,0] = 1 with dW₂[0,0] = −2.1932 and step η = 0.1 becomes 1 − 0.1(−2.1932) = 1.219 — nudged up, because raising it raises the true class's score. Repeat forward→backward→step thousands of times and the net learns. Real training swaps plain descent for Adam (Lesson 15).

56. Manual backward == PyTorch autograd

Worked example

Never trust a hand-derived gradient you haven't checked. Build the identical net in PyTorch, call .backward(), and compare with np.allclose. Full standalone script:

import numpy as np, torch
W1 = np.array([[1.,1.,0.],[1.,-1.,0.]]); b1 = np.array([0.,0.])
W2 = np.array([[1.,2.],[-1.,3.]]);       b2 = np.array([-2.,5.])
x  = np.array([1.,2.,1.]); y = 0
# manual forward + backward
z1 = W1@x+b1; a1 = np.maximum(0, z1); z2 = W2@a1+b2
p = np.exp(z2-z2.max()); p = p/p.sum()
dz2 = p - np.eye(2)[y]
dW2 = np.outer(dz2, a1)
dz1 = (W2.T @ dz2) * (z1 > 0)
dW1 = np.outer(dz1, x)
# torch autograd on the same net
tW1 = torch.tensor(W1, requires_grad=True); tW2 = torch.tensor(W2, requires_grad=True)
tb1 = torch.tensor(b1, requires_grad=True); tb2 = torch.tensor(b2, requires_grad=True)
zz = tW2 @ torch.relu(tW1 @ torch.tensor(x) + tb1) + tb2
torch.nn.functional.cross_entropy(zz.unsqueeze(0), torch.tensor([y])).backward()
print(np.allclose(dW1, tW1.grad.numpy()), np.allclose(dW2, tW2.grad.numpy()))

Prints: True True

Why: Both weight gradients match autograd to machine precision. The ReLU mask was the only per-layer subtlety, and we got it right.

gradientshapematches autograd
dW₂(2×2)True
dW₁(2×3)True
db₂(2,)True
db₁(2,)True

57. Finite-difference gradient check

Worked example

A second, independent check with no autograd at all: nudge one weight by ±ε and estimate its gradient from how the loss moves. This is the technique to keep when you write backprop from scratch.

\[ \frac{\partial L}{\partial W_{ij}} \approx \frac{L(W_{ij}+\varepsilon) - L(W_{ij}-\varepsilon)}{2\varepsilon} \]

import numpy as np
W1 = np.array([[1.,1.,0.],[1.,-1.,0.]]); b1 = np.array([0.,0.])
W2 = np.array([[1.,2.],[-1.,3.]]);       b2 = np.array([-2.,5.])
x = np.array([1.,2.,1.]); y = 0
def loss(W1):
    z1=W1@x+b1; a1=np.maximum(0,z1); z2=W2@a1+b2
    p=np.exp(z2-z2.max()); p=p/p.sum(); return -np.log(p[y])
eps = 1e-6
Wp = W1.copy(); Wp[0,1] += eps
Wm = W1.copy(); Wm[0,1] -= eps
print(round((loss(Wp)-loss(Wm))/(2*eps), 5))   # -2.92423

Numeric dW₁[0,1] = −2.92423, analytic = −2.92423

Why: The finite-difference estimate matches the outer-product value to 5 decimals — independent confirmation the backward pass is right, not just internally consistent with autograd.

weightfinite-diffanalytic (δ₁ xᵀ)
dW₁[0,0]−1.46212−1.46212
dW₁[0,1]−2.92423−2.92423
dW₁[1,0] (dead)0.000000.00000

58. The computational graph

Concept

Formally the network is a DAG: nodes are operations, edges carry tensors forward and gradients backward. Our chain x → z₁ → a₁ → z₂ → p → L is that graph, drawn straight.

Figure (svg): A left-to-right chain of boxes: x, then z1 (affine W1,b1), then a1 (ReLU), then z2 (affine W2,b2), then p (softmax), then L (cross-entropy). A lower arrow runs right-to-left labeled 'gradients' from L back to x.

Forward tensors flow right; backprop is one reverse sweep, each node multiplying by its local derivative.

Backprop is one reverse topological sweep — each node multiplies the incoming gradient by its local Jacobian (Lesson 9). PyTorch's autograd is exactly this graph, built and swept automatically.

59. Gradient pathologies

Section

Part 5 of 7 — vanishing, exploding, fixes

60. Why depth makes gradients unstable

Concept

A gradient through n layers is a product of n per-layer Jacobians. Products of numbers compound: repeated factors below 1 collapse toward 0, repeated factors above 1 blow up (Lesson 9).

\[ \frac{\partial L}{\partial z_1} = \underbrace{J_n J_{n-1} \cdots J_2 J_1}_{n \text{ factors}} \,\frac{\partial L}{\partial z_n} \]

Below-1 factors (stacked sigmoids) give vanishing gradients — early layers stop learning. Above-1 factors (large weights) give exploding gradients — NaNs and wild updates.

61. Compounding is why depth is hard

Intuition

Interest compounds: 1% a day is harmless, but over a year it is a different number entirely. Gradients through layers compound the same way — each layer multiplies, and multiplication of many numbers is wildly sensitive to whether they sit above or below 1.

A hair below 1 and the product decays to nothing (the early layers go blind); a hair above 1 and it detonates (the update NaNs out). A healthy network keeps every layer's factor near 1 — which is exactly what good initialization buys you.

62. Vanishing vs exploding, in numbers

Worked example

Model the compounding with a single per-layer factor raised to the depth. The sigmoid's derivative peaks at 0.25, so a deep sigmoid stack multiplies by at most 0.25 per layer:

import numpy as np
for depth in (5, 10, 20):
    print(f"depth {depth:2d}:  0.25^{depth} = {0.25**depth:.2e}   "
          f"1.5^{depth} = {1.5**depth:.2e}")
# why 0.25? the max slope of the sigmoid:
z = np.linspace(-6, 6, 1001); s = 1/(1+np.exp(-z))
print("max sigmoid'(z) =", round((s*(1-s)).max(), 4))

0.25²⁰ ≈ 9e−13 (vanished); 1.5²⁰ ≈ 3325 (exploded)

Why: Twenty sigmoid layers shrink the gradient a trillion-fold — early weights get essentially no signal. A factor of 1.5 instead blows it into the thousands. Only a factor near 1 stays healthy.

depth0.25ⁿ (vanish)1.5ⁿ (explode)
59.77e−047.59e+00
109.54e−075.77e+01
209.09e−133.33e+03

63. Watch it run: Vanishing vs exploding, in numbers

Pattern

Step through it

Step through Vanishing vs exploding, in numbers one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: depth is 5
  2. Step 2: depth is 10
  3. Step 3: depth is 20

64. Fix 1: gradient clipping

Concept

Clipping caps the gradient's norm so one huge step can't derail training. If the norm exceeds a threshold τ, rescale the whole vector down to length τ; otherwise leave it alone:

\[ g \leftarrow g \cdot \min\!\left(1, \frac{\tau}{\lVert g \rVert}\right) \]

The key word is norm: we scale the entire vector by one factor, so its direction is preserved — the step is shorter but still points downhill. This is the standard cure for exploding gradients.

65. Teach it back: Fix 1: gradient clipping

Explain it

Discussion prompt

Explain Fix 1: gradient clipping to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Clipping caps the gradient's norm so one huge step can't derail training. If the norm exceeds a threshold τ, rescale the whole vector down to length τ; otherwise leave it alone:

66. Finish it with less help: Clip by norm keeps the direction

Faded example

Fill in the blanks

Clip by norm keeps the direction, with the scaffolding fading: two lines are gone now — fill both.

import numpy as np
g = np.array([30., 40.]); tau = 5.0
norm = np.linalg.norm(g) # 50.0
g_clip = **g * min(1.0, tau / norm)** # scale the WHOLE vector
print(round(norm, 1), round(np.linalg.norm(g_clip), 1))
print(np.allclose(g_clip/np.linalg.norm(g_clip), g/norm))

Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. 3² + 4² = 25, so ‖[3,4]‖ = 5 = τ.

67. Clip by norm keeps the direction

Worked example

A gradient g = [30, 40] has norm 50. With threshold τ = 5, scale it down by 5/50 = 0.1 and check the direction is unchanged:

g · (5/50) = [3, 4], norm 5

Why: 3² + 4² = 25, so ‖[3,4]‖ = 5 = τ. The classic 3-4-5 triangle: same direction as [30,40], one-tenth the length.

import numpy as np
g = np.array([30., 40.]); tau = 5.0
norm = np.linalg.norm(g)                  # 50.0
g_clip = g * min(1.0, tau / norm)         # scale the WHOLE vector
print(round(norm, 1), round(np.linalg.norm(g_clip), 1))
print(np.allclose(g_clip/np.linalg.norm(g_clip), g/norm))
quantityvalue
‖g‖ before50.0
g_clip[3, 4]
‖g_clip‖5.0
direction preservedTrue

68. Fix 2: careful initialization

Concept

Where the instability starts is the initial weights. He initialization scales each layer's weights so a ReLU layer preserves the variance of its input — keeping activations (and thus gradients) in a healthy range from step 0:

\[ W_{ij} \sim \mathcal{N}\!\left(0, \tfrac{2}{n_{\text{in}}}\right), \quad \text{i.e. std} = \sqrt{\tfrac{2}{n_{\text{in}}}} \]

The 2 is for ReLU (which zeros half its inputs); tanh/sigmoid use Xavier with a 1 instead. Too-large weights explode the forward variance; too-small ones vanish it.

69. By analogy: Fix 2: careful initialization

Analogy

Discussion prompt

Explain Fix 2: careful initialization by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

The 2 is for ReLU (which zeros half its inputs); tanh/sigmoid use Xavier with a 1 instead. Too-large weights explode the forward variance; too-small ones vanish it.

70. Restore the missing line: He init preserves variance

Fill the middle

Fill in the blanks

From He init preserves variance — one line has had its right-hand side removed. Put it back.

import numpy as np
np.random.seed(0)
n_in = 100
x = np.random.randn(n_in) # input variance ~ 1
W_bad = np.random.randn(100, n_in) * 1.0 # too large
W_he = np.random.randn(100, n_in) * np.sqrt(2.0/n_in) # He
print(round(np.sqrt(2.0/n_in), 4))
print(round(np.var(np.maximum(0, W_bad @ x)), 2))
print(round(np.var(np.maximum(0, W_he @ x)), 2))

Why: x is what everything below it consumes, so the wrong expression here fails later and somewhere else. Std-1 weights amplify variance ~26x per layer — after a few layers that is an explosion.

71. He init preserves variance

Worked example

Feed a unit-variance input through a 100-wide ReLU layer with bad (std 1) vs He init, and measure the output variance. He keeps it near 1; std-1 weights explode it:

import numpy as np
np.random.seed(0)
n_in = 100
x = np.random.randn(n_in)                    # input variance ~ 1
W_bad = np.random.randn(100, n_in) * 1.0     # too large
W_he  = np.random.randn(100, n_in) * np.sqrt(2.0/n_in)   # He
print(round(np.sqrt(2.0/n_in), 4))
print(round(np.var(np.maximum(0, W_bad @ x)), 2))
print(round(np.var(np.maximum(0, W_he  @ x)), 2))

Bad init: output var 26.0; He init: 0.96

Why: Std-1 weights amplify variance ~26x per layer — after a few layers that is an explosion. He scaling holds it near 1, so signal neither grows nor dies with depth.

initweight stdoutput var (post-ReLU)
bad1.026.0
He√(2/100) = 0.14140.96
goal—≈ 1 (preserved)

72. Work backwards from the answer: He init preserves variance

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Bad init: output var 26.0; He init: 0.96

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Feed a unit-variance input through a 100-wide ReLU layer with bad (std 1) vs He init, and measure the output variance. He keeps it near 1; std-1 weights explode it:

73. Something is wrong here: clip by value vs by norm

Anomaly

Predict first

A student writes this, and it looks reasonable:

Just clamp every gradient component into [−τ, τ] — that shrinks the big ones, same idea as norm clipping.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Clamping components INDEPENDENTLY changes their ratio, so [30,40] (pointing at 53°) becomes [5,5] (pointing at 45°).

Scale the whole vector by τ/‖g‖ so every component shrinks by the same factor.

Why: Clamping components INDEPENDENTLY changes their ratio, so [30,40] (pointing at 53°) becomes [5,5] (pointing at 45°). You are no longer stepping downhill — you've turned the gradient.

74. Trap: clip by value vs by norm

Trap

The trap

Just clamp every gradient component into [−τ, τ] — that shrinks the big ones, same idea as norm clipping.

\[ g = \text{clip}(g, -\tau, \tau): \quad \begin{bmatrix} 30 \\ 40 \end{bmatrix} \to \begin{bmatrix} 5 \\ 5 \end{bmatrix} \]

Direction bent from ~53° to 45°

Why: Clamping components INDEPENDENTLY changes their ratio, so [30,40] (pointing at 53°) becomes [5,5] (pointing at 45°). You are no longer stepping downhill — you've turned the gradient.

The fix

Scale the whole vector by τ/‖g‖ so every component shrinks by the same factor.

\[ g \cdot \tfrac{\tau}{\lVert g \rVert}: \quad \begin{bmatrix} 30 \\ 40 \end{bmatrix} \to \begin{bmatrix} 3 \\ 4 \end{bmatrix} \]

Direction preserved exactly

Why: [3,4] is [30,40] scaled by 0.1 — identical direction, norm 5. Clip-by-VALUE is a different (occasionally useful) tool, but it is NOT a norm rescale and should never be swapped in for one.

75. Which of these survive contact with Lesson 18: Full Backpropagation?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Two layers of weights (W₁, W₂), one hidden ReLU layer with 2 units, 2 output classes. Small enough to do entirely by hand — big enough to show every real subtlety.; The whole lesson is those two passes on one tiny net — small enough that every blame-number is a fraction you can check by hand.; Logits z₂ are unbounded real numbers. Softmax exponentiates then normalizes, producing a probability vector p that is positive and sums to 1:
Breaks
Gradient just flows from a₁ back to z₁ unchanged — an affine map moved it, so set δ₁ = W₂ᵀδ₂ and carry on.; Just clamp every gradient component into [−τ, τ] — that shrinks the big ones, same idea as norm clipping.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 18: Full Backpropagation puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

76. Pattern & checks

Section

Part 6 of 7

77. Without one step: The full MLP recipe

Constraint

Discussion prompt

Run The full MLP recipe with this step confiscated:

Backward, per layer: weight grad dW = δ aᵀ (outer product), bias grad db = δ, push back with Wᵀδ, and cross each ReLU with the 1[z>0] mask

Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.

Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.

Answer:

  1. Init weights with He (ReLU) / Xavier (tanh) scaling — not all equal, not too large
  2. Forward, caching every z and a on the way
  3. Loss: softmax + cross-entropy → seed δ = p − e_y (the Lesson 17 shortcut)
  4. Backward, per layer: weight grad dW = δ aᵀ (outer product), bias grad db = δ, push back with Wᵀδ, and cross each ReLU with the 1[z>0] mask
  5. Stabilize: clip the gradient norm; then step with Adam (Lesson 15)
  6. Check against autograd or finite differences before trusting a from-scratch net

78. The full MLP recipe

Pattern

  1. Init weights with He (ReLU) / Xavier (tanh) scaling — not all equal, not too large
  2. Forward, caching every z and a on the way
  3. Loss: softmax + cross-entropy → seed δ = p − e_y (the Lesson 17 shortcut)
  4. Backward, per layer: weight grad dW = δ aᵀ (outer product), bias grad db = δ, push back with Wᵀδ, and cross each ReLU with the 1[z>0] mask
  5. Stabilize: clip the gradient norm; then step with Adam (Lesson 15)
  6. Check against autograd or finite differences before trusting a from-scratch net

79. Where does it stop working: The full MLP recipe

Edge cases

Discussion prompt

The full MLP recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Init weights with He (ReLU) / Xavier (tanh) scaling — not all equal, not too large
  2. Forward, caching every z and a on the way
  3. Loss: softmax + cross-entropy → seed δ = p − e_y (the Lesson 17 shortcut)
  4. Backward, per layer: weight grad dW = δ aᵀ (outer product), bias grad db = δ, push back with Wᵀδ, and cross each ReLU with the 1[z>0] mask
  5. Stabilize: clip the gradient norm; then step with Adam (Lesson 15)
  6. Check against autograd or finite differences before trusting a from-scratch net

80. Rule out three: Check yourself — the weight gradient

Elimination

Eliminate the wrong options

In a linear layer z = Wa + b with upstream gradient δ = ∂L/∂z and input a, the weight gradient dW is:

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. the outer product δ aᵀ
  • B. the dot product δ · a
  • C. Wᵀ δ
  • D. δ ⊙ a (elementwise)

Survives elimination: A

Why: dW = δ aᵀ is the outer product, giving the (out×in) shape of W — each weight's gradient is its output-error δ_j times the input a_k it multiplied. In our net, δ₂=[−0.7311,0.7311] outer a₁=[3,0] gave dW₂ with the correct (2×2) shape.

81. Check yourself — the weight gradient

Check

What shape, and built from what?

Check your understanding

In a linear layer z = Wa + b with upstream gradient δ = ∂L/∂z and input a, the weight gradient dW is:

  • A. the outer product δ aᵀ (correct)
  • B. the dot product δ · a
  • C. Wᵀ δ
  • D. δ ⊙ a (elementwise)

Answer: A

Why: dW = δ aᵀ is the outer product, giving the (out×in) shape of W — each weight's gradient is its output-error δ_j times the input a_k it multiplied. In our net, δ₂=[−0.7311,0.7311] outer a₁=[3,0] gave dW₂ with the correct (2×2) shape.

Why B tempts people
A dot product collapses δ and a into a single scalar, but dW must have W's full (out×in) shape. You need the outer product, which keeps both indices.
Why C tempts people
Wᵀδ is how the gradient propagates BACK to the previous layer's activations (da), not the weight gradient. Confusing these two is the most common backprop error.
Why D tempts people
Elementwise δ ⊙ a needs δ and a to be the same length and returns that same shape — but δ is (out,) and a is (in,), and dW must be (out×in). Elementwise is what the ReLU mask does, not the weight gradient.

82. Answer it before you see the options: Check yourself — the ReLU mask

Prediction

Predict first

During the backward pass, what happens to the gradient flowing into a ReLU unit whose forward pre-activation z was negative?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: It is multiplied by 0 — the gradient is blocked

Why: ReLU has slope 0 where z < 0, so its local derivative is 0 there. The mask 1[z>0] zeros that unit's gradient — in our net, da₁[1]=0.7311 was killed to 0, and W₁'s second row got zero gradient.

83. Check yourself — the ReLU mask

Check

A hidden unit had pre-activation z = −1 on this example.

Check your understanding

During the backward pass, what happens to the gradient flowing into a ReLU unit whose forward pre-activation z was negative?

  • A. It is multiplied by 0 — the gradient is blocked (correct)
  • B. It passes through unchanged
  • C. It is multiplied by z
  • D. It is multiplied by −1 (the sign flips)

Answer: A

Why: ReLU has slope 0 where z < 0, so its local derivative is 0 there. The mask 1[z>0] zeros that unit's gradient — in our net, da₁[1]=0.7311 was killed to 0, and W₁'s second row got zero gradient.

Why B tempts people
Unchanged pass-through is what happens where z > 0 (slope 1). For z < 0 the slope is 0, so the gradient must be blocked, not passed.
Why C tempts people
Multiplying by z would be the derivative of z²/2, not ReLU. ReLU's derivative is the step function 1[z>0], which is 0 or 1 — never z itself.
Why D tempts people
ReLU never negates; it either passes a positive input (slope 1) or clamps a negative one to 0 (slope 0). A sign flip would be a different, non-monotonic function.

84. Answer it before you see the options: Check yourself — exploding gradients

Prediction

Predict first

Training diverges: loss → NaN and gradient norms reach the thousands. The most direct fix is:

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: Gradient clipping — cap the gradient norm

Why: NaNs with huge gradient norms are the signature of exploding gradients. Clipping the norm caps the step so a single blow-up can't derail training; He initialization prevents it from starting. We saw 1.5²⁰ ≈ 3325 — clipping to τ would tame exactly that.

85. Check yourself — exploding gradients

Check

Pick the most direct fix.

Check your understanding

Training diverges: loss → NaN and gradient norms reach the thousands. The most direct fix is:

  • A. Gradient clipping — cap the gradient norm (correct)
  • B. Increase the learning rate
  • C. Add more layers
  • D. Switch the loss from cross-entropy to MSE

Answer: A

Why: NaNs with huge gradient norms are the signature of exploding gradients. Clipping the norm caps the step so a single blow-up can't derail training; He initialization prevents it from starting. We saw 1.5²⁰ ≈ 3325 — clipping to τ would tame exactly that.

Why B tempts people
A higher learning rate MULTIPLIES the already-huge gradient, making the overshoot worse and the divergence faster — the opposite of a fix.
Why C tempts people
More layers means more Jacobian factors in the product, compounding explosion (or vanishing) further. Depth is the cause of the pathology, not the cure.
Why D tempts people
The loss function doesn't control gradient magnitude, and MSE on top of softmax is worse for classification (Lesson 17). It addresses the wrong problem entirely.

86. Rule out three: Check yourself — which clipping

Elimination

Eliminate the wrong options

To shrink a large gradient WITHOUT changing its direction, you should:

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Scale the whole vector by τ/‖g‖
  • B. Clamp each component to [−τ, τ]
  • C. Set the largest component to τ
  • D. Zero out the smallest components

Survives elimination: A

Why: Multiplying the whole vector by τ/‖g‖ shrinks every component by the same factor, so magnitude drops but direction is preserved — clip-by-norm. Verified: [30,40] (norm 50) → [3,4] (norm 5), same direction.

87. Check yourself — which clipping

Check

Direction matters.

Check your understanding

To shrink a large gradient WITHOUT changing its direction, you should:

  • A. Scale the whole vector by τ/‖g‖ (correct)
  • B. Clamp each component to [−τ, τ]
  • C. Set the largest component to τ
  • D. Zero out the smallest components

Answer: A

Why: Multiplying the whole vector by τ/‖g‖ shrinks every component by the same factor, so magnitude drops but direction is preserved — clip-by-norm. Verified: [30,40] (norm 50) → [3,4] (norm 5), same direction.

Why B tempts people
Per-component clamping changes the components' ratios, bending the direction — [30,40] → [5,5] turns a 53° vector into a 45° one. That is clip-by-value, a different operation.
Why C tempts people
Editing a single component distorts the vector's direction and leaves the rest untouched — it is neither a uniform rescale nor a norm cap.
Why D tempts people
Zeroing components is sparsification; it changes both magnitude and direction arbitrarily and has nothing to do with taming a large-norm gradient.

88. From one example to a batch

Concept

Everything so far used one example, x ∈ ℝ³. Real training feeds a batch: stack N examples as rows of a matrix X ∈ ℝ^{N×d}. The forward pass becomes Z = XW + b (broadcast bias), one matmul for the whole batch.

The backward pass barely changes. The per-example outer product δ aᵀ becomes a matmul that sums over the batch, dW = AᵀΔ, and we average by dividing by N. The ReLU mask and the p − e_y seed apply row-wise, unchanged.

\[ dW = \tfrac{1}{N}\,A^\top \Delta, \qquad db = \tfrac{1}{N}\sum_{n} \delta^{(n)} \]

So the hand-derived rules scale directly to the project's (1347 × 64) batch — the shapes grow, the moves are identical.

89. Predict the next row: Batched shapes on the digits data

Pattern

Predict first

The table runs: X | (1347, 64) | batch of images · Z1 = X W1 + b1 | (1347, 32) | hidden pre-activations · A1 = ReLU(Z1) | (1347, 32) | hidden activations

In Batched shapes on the digits data, given the rows so far: what is the next one — the row where tensor is Z2 = A1 W2 + b2?

Correct: Z2 = A1 W2 + b2 | (1347, 10) | class logits

tensorshapemeaning
X(1347, 64)batch of images
Z1 = X W1 + b1(1347, 32)hidden pre-activations
A1 = ReLU(Z1)(1347, 32)hidden activations
Z2 = A1 W2 + b2(1347, 10)class logits

Why: The relationship between the columns, not the individual numbers, is what generates the next row. The batch dimension 1347 rides along untouched; the 64 features contract against W1's 64 rows, leaving 32 hidden units per example.

90. Batched shapes on the digits data

Worked example

Trace the shapes through one batched forward pass of the 64 → 32 → 10 net on the full training batch of 1347 images, so nothing surprises you in the project:

X (1347×64) @ W1 (64×32) → Z1 (1347×32)

Why: The batch dimension 1347 rides along untouched; the 64 features contract against W1's 64 rows, leaving 32 hidden units per example.

ReLU keeps (1347×32); then @ W2 (32×10) → Z2 (1347×10)

Why: ReLU is entrywise so shape is unchanged. The second matmul contracts the 32 hidden units to 10 class logits per example.

\[ X \xrightarrow{W_1} Z_1 \xrightarrow{\text{ReLU}} A_1 \xrightarrow{W_2} Z_2, \quad Z_2 \in \mathbb{R}^{1347 \times 10} \]

tensorshapemeaning
X(1347, 64)batch of images
Z1 = X W1 + b1(1347, 32)hidden pre-activations
A1 = ReLU(Z1)(1347, 32)hidden activations
Z2 = A1 W2 + b2(1347, 10)class logits

91. Decode the notation: Batched shapes on the digits data

Notation

Annotate

From Batched shapes on the digits data — read this one piece at a time. What is each part doing?

On: \( X \xrightarrow{W_1} Z_1 \xrightarrow{\text{ReLU}} A_1 \xrightarrow{W_2} Z_2, \quad Z_2 \in \mathbb{R}^{1347 \times 10} \)

  • The batch dimension 1347 rides along untouched; the 64 features contract against W1's 64 rows, leaving 32 hidden units per example.
  • ReLU is entrywise so shape is unchanged. The second matmul contracts the 32 hidden units to 10 class logits per example.

92. Your turn: train an MLP

Section

Part 7 of 7 — the project

93. Why build it yourself

Intuition

You could call nn.Sequential and never think about a single gradient. But the exam probes whether you know what .backward() does — the outer product, the transpose, the mask. Building it once, by hand, makes those reflexes instead of trivia.

So do both: train the PyTorch version to see it work, then re-derive it in raw NumPy to prove the backward pass in your head matches the one in the library. Same accuracy from both is the proof.

94. Project: MLP on the digits dataset

Concept

Train a one-hidden-layer network on load_digits (8×8 handwritten digit images, 64 features, 10 classes) and beat 90% test accuracy — everything you derived, now assembled into a real classifier.

#requirementtool
1Load + split digits, scale pixels to [0,1]load_digits, train_test_split
264 → 32 → 10 MLP with ReLU + Adamnn.Sequential, Adam
3Train with cross-entropy, report test accuracycross_entropy, argmax
4(Stretch) re-implement forward+backward in pure NumPynp.outer, the ReLU mask

Build rules: type every line yourself, divide pixels by 16 to normalize, seed with torch.manual_seed(0) so your numbers match, and judge on the held-out test split, never the training data.

95. Break it if you can: Project: MLP on the digits dataset

Counterexample

Discussion prompt

Train a one-hidden-layer network on load_digits (8×8 handwritten digit images, 64 features, 10 classes) and beat 90% test accuracy — everything you derived, now assembled into a real classifier.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Build rules: type every line yourself, divide pixels by 16 to normalize, seed with torch.manual_seed(0) so your numbers match, and judge on the held-out test split, never the training data.

96. Milestone 1 — the data

Worked example

Your turn: load digits, normalize to [0,1], and split 75/25. Predict the feature dimension out loud before you print it.

Hint: load_digits(), divide .data by 16 (pixels run 0–16), then train_test_split(..., test_size=0.25, random_state=0). Each image is 8×8 = 64 features.

from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
D = load_digits()
X = D.data / 16.0            # 8x8 pixels in [0,16] -> [0,1]
y = D.target
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.25, random_state=0)
print(X.shape, Xtr.shape, Xte.shape, len(set(y)))
objectshape / value
X (all)(1797, 64)
Xtr(1347, 64)
Xte(450, 64)
classes10

97. Watch it run: Milestone 1 — the data

Pattern

Step through it

Step through Milestone 1 — the data one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: object is X (all)
  2. Step 2: object is Xtr
  3. Step 3: object is Xte
  4. Step 4: object is classes

98. Milestone 2 — the network

Worked example

Your turn: build a 64 → 32 → 10 ReLU MLP and an Adam optimizer. Predict layer 1's parameter count before printing.

Hint: nn.Sequential(nn.Linear(64,32), nn.ReLU(), nn.Linear(32,10)), then Adam(net.parameters(), lr=0.01). Layer 1 has 64·32 weights + 32 biases.

import torch, torch.nn as nn
torch.manual_seed(0)
net = nn.Sequential(nn.Linear(64,32), nn.ReLU(), nn.Linear(32,10))
opt = torch.optim.Adam(net.parameters(), lr=0.01)
print(sum(p.numel() for p in net.parameters()))
print(64*32 + 32, 32*10 + 10)
layerweights + biasesparams
Linear 64→3264·32 + 322080
Linear 32→1032·10 + 10330
total—2410

99. Watch it run: Milestone 2 — the network

Pattern

Step through it

Step through Milestone 2 — the network one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: layer is Linear 64→32
  2. Step 2: layer is Linear 32→10
  3. Step 3: layer is total

100. Milestone 3 — train & evaluate

Worked example

Your turn: train 200 epochs of full-batch cross-entropy, then measure held-out accuracy. Predict: above or below 90%?

Hint: each epoch runs opt.zero_grad(), loss = F.cross_entropy(net(Xtr_t), ytr_t), loss.backward(), opt.step(). Accuracy = (argmax(1) == y).float().mean().

import torch, torch.nn as nn, torch.nn.functional as F
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
D = load_digits(); X = D.data/16.0; y = D.target
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.25, random_state=0)
Xtr_t = torch.tensor(Xtr, dtype=torch.float32); ytr_t = torch.tensor(ytr)
Xte_t = torch.tensor(Xte, dtype=torch.float32); yte_t = torch.tensor(yte)
torch.manual_seed(0)
net = nn.Sequential(nn.Linear(64,32), nn.ReLU(), nn.Linear(32,10))
opt = torch.optim.Adam(net.parameters(), lr=0.01)
for ep in range(200):
    opt.zero_grad()
    F.cross_entropy(net(Xtr_t), ytr_t).backward()
    opt.step()
acc = (net(Xte_t).argmax(1) == yte_t).float().mean().item()
print(round(acc, 4))   # 0.9689
epochtrain lossnote
02.3255≈ log(10), untrained
500.1329learning fast
1000.0518nearly fit
1990.0157converged

101. What has to be given first: The result: ~97% on held-out digits

Missing information

Discussion prompt

The full trained program, with gradient-norm clipping added for stability. Test accuracy lands at 0.9689 — comfortably past the 90% bar.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

Clip-by-norm here is a safety net (the net is small and stable), and the result matches the unclipped run to 4 decimals — clipping never hurts a well-behaved run, and saves an exploding one.

102. The result: ~97% on held-out digits

Worked example

The full trained program, with gradient-norm clipping added for stability. Test accuracy lands at 0.9689 — comfortably past the 90% bar.

import torch, torch.nn as nn, torch.nn.functional as F
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
D = load_digits(); X = D.data/16.0; y = D.target
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.25, random_state=0)
Xtr_t = torch.tensor(Xtr, dtype=torch.float32); ytr_t = torch.tensor(ytr)
Xte_t = torch.tensor(Xte, dtype=torch.float32); yte_t = torch.tensor(yte)
torch.manual_seed(0)
net = nn.Sequential(nn.Linear(64,32), nn.ReLU(), nn.Linear(32,10))
opt = torch.optim.Adam(net.parameters(), lr=0.01)
for ep in range(200):
    opt.zero_grad()
    loss = F.cross_entropy(net(Xtr_t), ytr_t)
    loss.backward()
    torch.nn.utils.clip_grad_norm_(net.parameters(), 5.0)   # stabilize
    opt.step()
acc = (net(Xte_t).argmax(1) == yte_t).float().mean().item()
print('final train loss:', round(loss.item(), 4))
print('test accuracy   :', round(acc, 4))

final train loss 0.0157, test accuracy 0.9689

Why: Clip-by-norm here is a safety net (the net is small and stable), and the result matches the unclipped run to 4 decimals — clipping never hurts a well-behaved run, and saves an exploding one.

metricvalue
final train loss0.0157
test accuracy0.9689 (~97%)
target> 0.90 ✓

103. What each one costs: The result: ~97% on held-out digits

Trade off

Comparison matrix

From The result: ~97% on held-out digits: every row here is a choice with a cost. Fill the value column, then say which row you would actually pick and what you give up for it.

metricvalue
final train loss0.0157
test accuracy0.9689 (~97%)
target> 0.90 ✓

104. Guess the shape of the answer: Stretch — the whole net in pure NumPy

Estimation

Predict first

Your turn (stretch): no autograd at all. Batched forward, softmax+CE seed P − Y, then the exact three moves you derived — aᵀδ for weights, Wᵀδ to push back, ⊙ (z>0) for the ReLU. He init, plain SGD.

Commit before you compute: what does Stretch — the whole net in pure NumPy come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: scratch NumPy test acc = 0.9622

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. No PyTorch, no autograd — just np.outer's batched cousin (aᵀδ), the Wᵀδ push-back, and the (z>0) mask, exactly the moves from Part 3.

105. Stretch — the whole net in pure NumPy

Worked example

Your turn (stretch): no autograd at all. Batched forward, softmax+CE seed P − Y, then the exact three moves you derived — aᵀδ for weights, Wᵀδ to push back, ⊙ (z>0) for the ReLU. He init, plain SGD.

import numpy as np
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
D = load_digits(); X = D.data/16.0; y = D.target
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.25, random_state=0)
rng = np.random.default_rng(0)
W1 = rng.standard_normal((64,32))*np.sqrt(2/64); b1 = np.zeros(32)   # He
W2 = rng.standard_normal((32,10))*np.sqrt(2/32); b2 = np.zeros(10)
Y = np.eye(10)[ytr]
for ep in range(300):
    z1 = Xtr@W1 + b1; a1 = np.maximum(0, z1)          # forward
    z2 = a1@W2 + b2
    P = np.exp(z2 - z2.max(1, keepdims=True)); P /= P.sum(1, keepdims=True)
    dz2 = (P - Y) / len(Xtr)                            # softmax+CE seed
    gW2 = a1.T@dz2; dz1 = (dz2@W2.T) * (z1 > 0)        # backward + ReLU mask
    gW1 = Xtr.T@dz1
    W1 -= 0.5*gW1; b1 -= 0.5*dz1.sum(0)                # SGD step
    W2 -= 0.5*gW2; b2 -= 0.5*dz2.sum(0)
pred = (np.maximum(0, Xte@W1 + b1)@W2 + b2).argmax(1)
print('scratch NumPy test acc:', round((pred == yte).mean(), 4))

scratch NumPy test acc = 0.9622

Why: No PyTorch, no autograd — just np.outer's batched cousin (aᵀδ), the Wᵀδ push-back, and the (z>0) mask, exactly the moves from Part 3. Hits ~96%, confirming you understand backprop end to end.

implementationtest accuracy
PyTorch autograd (Milestone 3)0.9689
from-scratch NumPy0.9622
both> 0.90 ✓

106. Fill in: test accuracy for Stretch — the whole net in pure NumPy

Comparison

Comparison matrix

From Stretch — the whole net in pure NumPy: refill the test accuracy column from what you know. The rest of the table is as it appeared.

implementationtest accuracy
PyTorch autograd (Milestone 3)0.9689
from-scratch NumPy0.9622
both> 0.90 ✓

107. Show it off

Concept

Slides closed, out loud: explain (1) why you cache z₁ (not just a₁) on the forward pass, (2) what the ReLU mask did to the dead unit's gradient in our worked example, and (3) why clip-by-norm preserves the descent direction while clip-by-value does not.

Stretch homework: add a gradient check to the NumPy net (compare aᵀδ against finite differences), then extend it to arbitrary depth and add an L2 penalty. This same manual backward becomes BatchNorm backprop (Week 18) and transformer backprop (Week 30).

108. Connect it up: Lesson 18: Full Backpropagation

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — The running network · The forward pass · The backward pass · Verify everything · Gradient pathologies · Pattern & checks. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

109. What you can do now

Recap

stepthe one thing to remember
forwardcache every z AND a
softmax+CE seedδ = p − e_y
weight gradientdW = δ aᵀ (outer product, shape of W)
push backda = Wᵀδ
ReLU backwardmultiply by 1[z>0] — from cached z
explodingclip by NORM (keeps direction), He init

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 18 (Week 6 — Full Backpropagation) — Barron · USAAIO Round 2 Preparation, 2026
  2. PyTorch autograd mechanics
  3. He et al., Delving Deep into Rectifiers (He initialization) — ICCV 2015
  4. Every gradient, trace table, and accuracy produced by real execution — numpy 2.2.6 + torch 2.7.1, verification run July 2026 — manual backward == autograd, MLP digits 96.89% test

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108