USAAIO Lesson 18, from Week 6, fully worked. One tiny 3-2-2 network with fixed weights is carried end to end. The forward pass is built operation by operation - affine, ReLU, affine, softmax, cross-entropy - and then the complete backward pass is derived one gradient per beat, covering the softmax-and-cross-entropy shortcut, the outer-product weight gradient, and the ReLU mask that kills a dead unit. Every number is matched against PyTorch autograd and against a finite-difference gradient check. It then covers vanishing and exploding gradients with clip-by-norm and He initialization, and caps off by training an MLP to about 97% on the sklearn digits dataset, along with a from-scratch NumPy network. Every snippet runs standalone, and every number was produced by real execution. The lesson runs to 60 slides.
Subject: Machine Learning · 109 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 18 · Week 6
One tiny network, 3 → 2 → 2, carried end to end. We build the forward pass op by op, then derive every gradient of the backward pass one move at a time — and check each number against PyTorch. Then the pathologies, and a net you train to ~97%.
Objectives
p − e_y, the outer-product dW = δ aᵀ, and the ReLU mask 1[z>0]Warm-up
Discussion prompt
Before we open Lesson 18: Full Backpropagation: without looking back, what was the main idea of Loss Functions & Their Gradients, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
MSE, MAE, binary and multi-class cross-entropy, and hinge loss with their gradients — the clean dL/dz = p − y result for softmax+cross-entropy and sigmoid+BCE, and why MSE is a poor classification loss.
Section
Part 1 of 7 — set up one concrete example
Concept
We fix one network and one training example and never change them. Every derivation below produces a number you can check. The net maps a 3-vector input to 2 class scores: affine → ReLU → affine → softmax → cross-entropy.
\[ x \in \mathbb{R}^3 \;\xrightarrow{W_1,b_1}\; z_1 \;\xrightarrow{\text{ReLU}}\; a_1 \;\xrightarrow{W_2,b_2}\; z_2 \;\xrightarrow{\text{softmax}}\; p \;\xrightarrow{\text{CE}}\; L \]
Two layers of weights (W₁, W₂), one hidden ReLU layer with 2 units, 2 output classes. Small enough to do entirely by hand — big enough to show every real subtlety.
Counterexample
Discussion prompt
Two layers of weights (W₁, W₂), one hidden ReLU layer with 2 units, 2 output classes. Small enough to do entirely by hand — big enough to show every real subtlety.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Intuition
Two passes, two questions. The forward pass asks what does the network predict? — it runs the data through and produces a loss. The backward pass asks who is to blame for that loss, and by how much? — it hands each weight a number saying which way to nudge it.
That blame-number is a partial derivative ∂L/∂W. Backprop is just the chain rule organized so we compute every one of them in a single sweep from the loss back to the inputs, reusing shared work instead of redoing it.
The whole lesson is those two passes on one tiny net — small enough that every blame-number is a fraction you can check by hand.
Analogy
Discussion prompt
Explain Forward asks, backward blames by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
The whole lesson is those two passes on one tiny net — small enough that every blame-number is a fraction you can check by hand.
Concept
Here are the exact numbers. W₁ is (2×3) — 2 hidden units, each reading the 3 inputs. W₂ is (2×2) — 2 class scores from the 2 hidden units. The true label is class 0.
\[ W_1 = \begin{bmatrix} 1 & 1 & 0 \\ 1 & -1 & 0 \end{bmatrix},\; b_1 = \begin{bmatrix} 0 \\ 0 \end{bmatrix},\quad W_2 = \begin{bmatrix} 1 & 2 \\ -1 & 3 \end{bmatrix},\; b_2 = \begin{bmatrix} -2 \\ 5 \end{bmatrix} \]
\[ x = \begin{bmatrix} 1 \\ 2 \\ 1 \end{bmatrix}, \qquad y = 0 \;\;(\text{true class}) \]
Intuition
The backward pass reuses the forward pass's leftovers. To get the weight gradient at a layer you need that layer's input activation; to cross a ReLU you need its pre-activation z. Throw them away and you'd have to recompute the whole forward pass at every step.
So the rule is: on the way forward, stash z₁, a₁, z₂, p. On the way back, spend them. This forward-cache-then-backward-sweep is exactly what an autograd engine does under the hood.
Section
Part 2 of 7 — op by op
Estimation
Predict first
First affine map. Each hidden unit dots its weight row with x, then adds its bias. Two units, so two dot products:
Commit before you compute: what does Layer 1 affine: z₁ = W₁x + b₁ come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Unit 1: (1)(1) + (−1)(2) + (0)(1) + 0 = −1
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Row 1 of W₁ is [1, −1, 0]; dot with x gives 1 − 2 + 0 = −1.
Worked example
First affine map. Each hidden unit dots its weight row with x, then adds its bias. Two units, so two dot products:
Unit 0: (1)(1) + (1)(2) + (0)(1) + 0 = 3
Why: Row 0 of W₁ is [1, 1, 0]; dot with x = [1, 2, 1] gives 1 + 2 + 0 = 3, plus bias 0.
\[ z_{1,0} = 1\cdot1 + 1\cdot2 + 0\cdot1 + 0 = 3 \]
Unit 1: (1)(1) + (−1)(2) + (0)(1) + 0 = −1
Why: Row 1 of W₁ is [1, −1, 0]; dot with x gives 1 − 2 + 0 = −1. This unit's pre-activation is NEGATIVE — remember that.
\[ z_1 = \begin{bmatrix} 3 \\ -1 \end{bmatrix} \]
import numpy as np
W1 = np.array([[1.,1.,0.],[1.,-1.,0.]]) # (2,3)
b1 = np.array([0., 0.])
x = np.array([1., 2., 1.]) # one example
z1 = W1 @ x + b1
print(z1) # [3. -1.]| hidden unit | weight row · x | + bias | z₁ |
|---|---|---|---|
| 0 | 1+2+0 = 3 | 0 | 3 |
| 1 | 1−2+0 = −1 | 0 | −1 |
| shape | (2×3)·(3,) | (2,) | (2,) |
Comparison
Comparison matrix
From Layer 1 affine: z₁ = W₁x + b₁: refill the weight row · x column from what you know. The rest of the table is as it appeared.
| hidden unit | weight row · x | + bias | z₁ |
|---|---|---|---|
| 0 | 1+2+0 = 3 | 0 | 3 |
| 1 | 1−2+0 = −1 | 0 | −1 |
| shape | (2×3)·(3,) | (2,) | (2,) |
Missing information
Discussion prompt
The nonlinearity. ReLU passes positives through untouched and clamps negatives to 0. Apply it entrywise to z₁ = [3, −1]:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Positive pre-activation survives — this unit is ACTIVE and will carry gradient on the way back.
Worked example
The nonlinearity. ReLU passes positives through untouched and clamps negatives to 0. Apply it entrywise to z₁ = [3, −1]:
Unit 0: max(0, 3) = 3
Why: Positive pre-activation survives — this unit is ACTIVE and will carry gradient on the way back.
Unit 1: max(0, −1) = 0
Why: Negative pre-activation is clamped to 0 — this unit is DEAD for this input. Its gradient path will be cut in the backward pass.
\[ a_1 = \max(0, z_1) = \begin{bmatrix} \max(0,3) \\ \max(0,-1) \end{bmatrix} = \begin{bmatrix} 3 \\ 0 \end{bmatrix} \]
import numpy as np
z1 = np.array([3., -1.])
a1 = np.maximum(0, z1) # ReLU, entrywise
print(a1) # [3. 0.]| unit | z₁ | ReLU(z₁) = a₁ | status |
|---|---|---|---|
| 0 | 3 | 3 | active |
| 1 | −1 | 0 | dead (clamped) |
Trade off
Comparison matrix
From ReLU: a₁ = max(0, z₁): every row here is a choice with a cost. Fill the status column, then say which row you would actually pick and what you give up for it.
| unit | z₁ | ReLU(z₁) = a₁ | status |
|---|---|---|---|
| 0 | 3 | 3 | active |
| 1 | −1 | 0 | dead (clamped) |
Pattern
Predict first
The table runs: 0 | 3+0 = 3 | −2 | 1 · 1 | −3+0 = −3 | 5 | 2
In Layer 2 affine: z₂ = W₂a₁ + b₂, given the rows so far: what is the next one — the row where class is note?
Correct: note | a₁[1]=0 kills col 2 | — | —
| class | row · a₁ | + bias | z₂ (logit) |
|---|---|---|---|
| 0 | 3+0 = 3 | −2 | 1 |
| 1 | −3+0 = −3 | 5 | 2 |
| note | a₁[1]=0 kills col 2 | — | — |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. Row 0 of W₂ is [1, 2]; dot with a₁ = [3, 0] gives 3 + 0 = 3, plus bias −2 → 1.
Worked example
Second affine map turns the 2 hidden activations a₁ = [3, 0] into 2 class scores (logits). Same pattern: dot each row of W₂ with a₁, add the bias.
Class 0: (1)(3) + (2)(0) + (−2) = 1
Why: Row 0 of W₂ is [1, 2]; dot with a₁ = [3, 0] gives 3 + 0 = 3, plus bias −2 → 1.
\[ z_{2,0} = 1\cdot3 + 2\cdot0 - 2 = 1 \]
Class 1: (−1)(3) + (3)(0) + 5 = 2
Why: Row 1 of W₂ is [−1, 3]; dot with a₁ gives −3 + 0 = −3, plus bias 5 → 2.
\[ z_2 = \begin{bmatrix} 1 \\ 2 \end{bmatrix} \]
import numpy as np
W2 = np.array([[1.,2.],[-1.,3.]]) # (2,2)
b2 = np.array([-2., 5.])
a1 = np.array([3., 0.])
z2 = W2 @ a1 + b2
print(z2) # [1. 2.]| class | row · a₁ | + bias | z₂ (logit) |
|---|---|---|---|
| 0 | 3+0 = 3 | −2 | 1 |
| 1 | −3+0 = −3 | 5 | 2 |
| note | a₁[1]=0 kills col 2 | — | — |
Notation
Annotate
From Layer 2 affine: z₂ = W₂a₁ + b₂ — read this one piece at a time. What is each part doing?
On: \( z_{2,0} = 1\cdot3 + 2\cdot0 - 2 = 1 \)
Concept
Logits z₂ are unbounded real numbers. Softmax exponentiates then normalizes, producing a probability vector p that is positive and sums to 1:
\[ p_k = \frac{e^{z_{2,k}}}{\sum_j e^{z_{2,j}}} \]
In code we subtract max(z₂) inside the exponent first — it cancels in the ratio but stops exp from overflowing on large logits. A must-have habit.
Explain it
Discussion prompt
Explain Softmax turns logits into probabilities to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Logits z₂ are unbounded real numbers. Softmax exponentiates then normalizes, producing a probability vector p that is positive and sums to 1:
Fill the middle
Fill in the blanks
From Softmax on z₂ = [1, 2] — one line has had its right-hand side removed. Put it back.
import numpy as np
z2 = np.array([1., 2.])
p = np.exp(z2 - z2.max()) # subtract max for numerical safety
p = p / p.sum()
print(np.round(p, 4)) # [0.2689 0.7311]
Why: p is what everything below it consumes, so the wrong expression here fails later and somewhere else. Raise e to each logit. These are the unnormalized weights of each class.
Worked example
Exponentiate: e¹ ≈ 2.7183, e² ≈ 7.3891
Why: Raise e to each logit. These are the unnormalized weights of each class.
\[ e^{z_2} = \begin{bmatrix} e^1 \\ e^2 \end{bmatrix} \approx \begin{bmatrix} 2.7183 \\ 7.3891 \end{bmatrix}, \quad \textstyle\sum = 10.1073 \]
Normalize by the sum 10.1073
Why: Divide each by the total so they sum to 1. p₀ = 2.7183/10.1073, p₁ = 7.3891/10.1073.
\[ p = \begin{bmatrix} 0.2689 \\ 0.7311 \end{bmatrix} \]
import numpy as np
z2 = np.array([1., 2.])
p = np.exp(z2 - z2.max()) # subtract max for numerical safety
p = p / p.sum()
print(np.round(p, 4)) # [0.2689 0.7311]| class k | z₂ | e^z₂ | p = e^z₂ / Σ |
|---|---|---|---|
| 0 | 1 | 2.7183 | 0.2689 |
| 1 | 2 | 7.3891 | 0.7311 |
| Σ | — | 10.1073 | 1.0000 |
Error analysis
Annotate
Walk the callouts on Softmax on z₂ = [1, 2]. Each one is a place this is easy to get subtly wrong.
Faded example
Fill in the blanks
Cross-entropy: the loss, with the scaffolding fading: two lines are gone now — fill both.
import numpy as np
z2 = np.array([1., 2.]); y = 0
p = np.exp(z2 - z2.max()); p = p / p.sum()
L = -np.log(p[y]) # cross-entropy = -log prob of true class
print(round(L, 4)) # 1.3133
Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. The model is confident in the WRONG class (0.7311 on class 1), so the loss is substantial.
Worked example
Cross-entropy for a single label is just the negative log-probability the model assigns to the true class. Our true class is 0, and the model gave it only p₀ = 0.2689:
\[ L = -\log p_{y} = -\log p_0 = -\log(0.2689) \]
L = −log(0.2689) ≈ 1.3133
Why: The model is confident in the WRONG class (0.7311 on class 1), so the loss is substantial. A perfect p₀ = 1 would give L = 0.
\[ \boxed{\,L \approx 1.3133\,} \]
import numpy as np
z2 = np.array([1., 2.]); y = 0
p = np.exp(z2 - z2.max()); p = p / p.sum()
L = -np.log(p[y]) # cross-entropy = -log prob of true class
print(round(L, 4)) # 1.3133| quantity | value |
|---|---|
| p (softmax) | [0.2689, 0.7311] |
| true class y | 0 |
| p[y] | 0.2689 |
| L = −log p[y] | 1.3133 |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
L = −log(0.2689) ≈ 1.3133
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Cross-entropy for a single label is just the negative log-probability the model assigns to the true class. Our true class is 0, and the model gave it only p₀ = 0.2689:
Concept
That is the whole forward pass. Everything we stashed is on the table. The backward pass will consume p (to seed the gradient), a₁ and x (for weight gradients), and z₁ (for the ReLU mask).
| cached | value | used backward for |
|---|---|---|
| z₁ | [3, −1] | ReLU mask 1[z₁>0] |
| a₁ | [3, 0] | dW₂ = δ₂ a₁ᵀ |
| z₂ | [1, 2] | (produced p) |
| p | [0.2689, 0.7311] | seed δ₂ = p − e_y |
| x | [1, 2, 1] | dW₁ = δ₁ xᵀ |
import numpy as np
W1 = np.array([[1.,1.,0.],[1.,-1.,0.]]); b1 = np.array([0.,0.])
W2 = np.array([[1.,2.],[-1.,3.]]); b2 = np.array([-2.,5.])
x = np.array([1.,2.,1.])
z1 = W1 @ x + b1; a1 = np.maximum(0, z1)
z2 = W2 @ a1 + b2
p = np.exp(z2 - z2.max()); p = p / p.sum()
print(z1, a1, z2, np.round(p, 4)) # [3 -1] [3 0] [1 2] [0.2689 0.7311]Section
Part 3 of 7 — one gradient per beat
Concept
Backprop wants ∂L/∂W₁ and ∂L/∂W₂. The chain rule says a gradient at any node is the gradient just downstream times that node's local derivative. So we start at the loss and walk backward, multiplying local derivatives as we go.
We write δ_ℓ = ∂L/∂z_ℓ for the gradient at a layer's pre-activation. Each layer's job: turn the downstream δ into (a) its weight gradient and (b) the δ for the layer before it. Two moves, repeated.
Concept
The seed of the whole backward pass is ∂L/∂z₂. Softmax and cross-entropy look nasty separately, but composed they simplify dramatically. Cross-entropy for true class y is L = −log p_y, and p_y = e^{z_y} / Σ_j e^{z_j}, so L = −z_y + log Σ_j e^{z_j}.
Differentiate L = −z_y + log Σ_j e^{z_j} with respect to one logit z_k. The first term contributes −1 only when k = y; the log-sum term contributes e^{z_k}/Σ = p_k for every k.
\[ \frac{\partial L}{\partial z_k} = p_k - \mathbb{1}[k = y] \]
Stacked over all k, that indicator vector is exactly the one-hot e_y. So the whole gradient is p − e_y — no Jacobian to multiply, just a subtraction.
Worked example
The gradient of softmax-then-cross-entropy with respect to the logits collapses to a famously clean form (derived last slide): predicted minus one-hot true.
\[ \delta_2 = \frac{\partial L}{\partial z_2} = p - e_y \]
e_y = [1, 0] since true class is 0
Why: The one-hot vector puts a 1 in the true class slot. Subtract it from p = [0.2689, 0.7311].
\[ \delta_2 = \begin{bmatrix} 0.2689 \\ 0.7311 \end{bmatrix} - \begin{bmatrix} 1 \\ 0 \end{bmatrix} = \begin{bmatrix} -0.7311 \\ 0.7311 \end{bmatrix} \]
import numpy as np
z2 = np.array([1., 2.]); y = 0
p = np.exp(z2 - z2.max()); p = p / p.sum()
onehot = np.zeros(2); onehot[y] = 1.0
dz2 = p - onehot # the softmax+CE shortcut
print(np.round(dz2, 4)) # [-0.7311 0.7311]| class | p | e_y | δ₂ = p − e_y |
|---|---|---|---|
| 0 (true) | 0.2689 | 1 | −0.7311 |
| 1 | 0.7311 | 0 | +0.7311 |
| reading | too low → push up | — | negative = raise score |
Sorting
Sort into buckets
These are the pieces of Lesson 18: Full Backpropagation, out of order. Put each one back under the part of the lesson it belongs to.
Concept
For a linear layer z = Wa + b, entry ∂L/∂W_{jk} picks up δ_j (error at output j) times a_k (the input it multiplied). Collecting all j, k is exactly the outer product:
\[ \frac{\partial L}{\partial W} = \delta\, a^\top, \qquad \frac{\partial L}{\partial b} = \delta \]
Shape check: δ is (out,), a is (in,), so δ aᵀ is (out×in) — exactly the shape of W. If your dW is not the same shape as W, you have the outer product backwards.
Worked example
Apply the outer product at layer 2 with δ₂ = [−0.7311, 0.7311] and a₁ = [3, 0]. Each entry is δ₂[j] · a₁[k]:
Column for a₁[0]=3: [−0.7311·3, 0.7311·3] = [−2.1932, 2.1932]
Why: The active hidden unit (a₁[0]=3) receives real gradient — these weights get updated.
Column for a₁[1]=0: [0, 0]
Why: The DEAD unit contributed 0 to the forward pass, so the weights feeding OUT of it get zero gradient this step — you can't blame a unit that said nothing.
\[ dW_2 = \begin{bmatrix} -0.7311 \\ 0.7311 \end{bmatrix}\begin{bmatrix} 3 & 0 \end{bmatrix} = \begin{bmatrix} -2.1932 & 0 \\ 2.1932 & 0 \end{bmatrix} \]
import numpy as np
a1 = np.array([3., 0.]) # hidden activations, cached
p = np.array([0.26894142, 0.73105858]) # softmax output, cached
dz2 = p - np.eye(2)[0] # exact upstream grad at z2
dW2 = np.outer(dz2, a1) # (2,) outer (2,) -> (2,2)
db2 = dz2
print(np.round(dW2, 4)) # [[-2.1932 0.] [2.1932 0.]]
print(dW2.shape) # (2, 2), matches W2| j (out) | k=0 (a₁=3) | k=1 (a₁=0) | row of dW₂ |
|---|---|---|---|
| 0 | −0.7311·3 = −2.1932 | 0 | [−2.1932, 0] |
| 1 | 0.7311·3 = 2.1932 | 0 | [2.1932, 0] |
| db₂ | = δ₂ | — | [−0.7311, 0.7311] |
Blank canvas
Draw it
Draw what dW₂ = δ₂ a₁ᵀ just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Concept
Forward, z = Wa sends each input a_k to every output via column k of W. Backward, the gradient at a_k must collect contributions from every output it fed — which means reading W down its columns, i.e. across the rows of Wᵀ.
\[ z_j = \sum_k W_{jk}\, a_k \;\Longrightarrow\; \frac{\partial L}{\partial a_k} = \sum_j W_{jk}\,\delta_j = (W^\top \delta)_k \]
So the local Jacobian of an affine layer is W going forward and Wᵀ coming back. The transpose is not a trick to memorize — it is the chain rule summing over the right index. Shape check: Wᵀ is (in×out), δ is (out,), product is (in,) — the shape of a.
Fill the middle
Fill in the blanks
From Push back to a₁: da₁ = W₂ᵀ δ₂ — one line has had its right-hand side removed. Put it back.
import numpy as np
W2 = np.array([[1.,2.],[-1.,3.]])
p = np.array([0.26894142, 0.73105858])
dz2 = p - np.eye(2)[0] # exact upstream grad at z2
da1 = W2.T @ dz2 # gradient arriving at a1
print(np.round(da1, 4)) # [-1.4621 0.7311]
Why: dz2 is what everything below it consumes, so the wrong expression here fails later and somewhere else. Row 0: 1·(−0.7311) + (−1)·(0.7311) = −1.4621.
Worked example
To keep going we need the gradient at the hidden activations a₁. The transpose of the layer's weight matrix carries δ₂ backward across the affine map:
\[ da_1 = \frac{\partial L}{\partial a_1} = W_2^\top \delta_2 \]
W₂ᵀ = [[1, −1], [2, 3]]; multiply by δ₂ = [−0.7311, 0.7311]
Why: Row 0: 1·(−0.7311) + (−1)·(0.7311) = −1.4621. Row 1: 2·(−0.7311) + 3·(0.7311) = 0.7311.
\[ da_1 = \begin{bmatrix} 1 & -1 \\ 2 & 3 \end{bmatrix}\begin{bmatrix} -0.7311 \\ 0.7311 \end{bmatrix} = \begin{bmatrix} -1.4621 \\ 0.7311 \end{bmatrix} \]
Note da₁[1] = 0.7311 is nonzero — gradient wants to flow into the dead unit. The ReLU on the next slide is what stops it.
import numpy as np
W2 = np.array([[1.,2.],[-1.,3.]])
p = np.array([0.26894142, 0.73105858])
dz2 = p - np.eye(2)[0] # exact upstream grad at z2
da1 = W2.T @ dz2 # gradient arriving at a1
print(np.round(da1, 4)) # [-1.4621 0.7311]| unit | W₂ᵀ row · δ₂ | da₁ |
|---|---|---|
| 0 | 1(−0.7311) + (−1)(0.7311) | −1.4621 |
| 1 | 2(−0.7311) + 3(0.7311) | +0.7311 |
| shape | (2×2)·(2,) | (2,) |
Translation
\( da_1 = \frac{\partial L}{\partial a_1} = W_2^\top \delta_2 \)
Draw it
Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.
Intuition
Think of each ReLU as a gate. When its input was positive the gate was open: nudging the input nudges the output one-for-one, so gradient passes straight through (slope 1). When its input was negative the gate was shut: the output was flatly 0, and a tiny nudge still leaves it 0 (slope 0).
A shut gate can't leak. If a unit contributed nothing to the prediction, it deserves no blame for the error — so its gradient is zeroed. That's the entire content of the ReLU mask.
Fill the middle
Fill in the blanks
From Cross the ReLU: δ₁ = da₁ ⊙ 1[z₁>0] — finish the line. Write what belongs on the right of the equals sign before you look.
\delta_1 = da_1 \odot \mathbb{1}[z_1 > 0]
Why: Producing the right-hand side unprompted is the difference between recognising this line and being able to use it. z₁[0]=3>0 → 1 (gradient passes).
Worked example
ReLU's local derivative is 1 where its input was positive and 0 where it was negative. Crossing it means multiplying the incoming gradient by this mask — computed from the cached z₁, not a₁.
\[ \delta_1 = da_1 \odot \mathbb{1}[z_1 > 0] \]
Mask from z₁ = [3, −1] is [1, 0]
Why: z₁[0]=3>0 → 1 (gradient passes). z₁[1]=−1<0 → 0 (gradient blocked). This is the dead unit paying off.
δ₁ = [−1.4621, 0.7311] ⊙ [1, 0] = [−1.4621, 0]
Why: The 0.7311 that wanted to flow into unit 1 is ZEROED. A dead ReLU has zero slope, so no gradient reaches its incoming weights — the single most-forgotten step in manual backprop.
\[ \delta_1 = \begin{bmatrix} -1.4621 \\ 0 \end{bmatrix} \]
import numpy as np
W2 = np.array([[1.,2.],[-1.,3.]])
z1 = np.array([3., -1.]) # unit 1 is negative -> dead
p = np.array([0.26894142, 0.73105858])
dz2 = p - np.eye(2)[0] # exact upstream grad at z2
da1 = W2.T @ dz2
dz1 = da1 * (z1 > 0) # multiply by the ReLU mask
print(np.round(dz1, 4)) # [-1.4621 0. ]| unit | da₁ | 1[z₁>0] | δ₁ = da₁ ⊙ mask |
|---|---|---|---|
| 0 | −1.4621 | 1 | −1.4621 |
| 1 | 0.7311 | 0 | 0 (killed) |
Error analysis
Annotate
Walk the callouts on Cross the ReLU: δ₁ = da₁ ⊙ 1[z₁>0]. Each one is a place this is easy to get subtly wrong.
Estimation
Predict first
Same outer-product rule as layer 2, now with δ₁ = [−1.4621, 0] and the input x = [1, 2, 1]. This gives the (2×3) gradient for W₁:
Commit before you compute: what does dW₁ = δ₁ xᵀ come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Row 1 (dead unit): 0 · [1, 2, 1] = [0, 0, 0]
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. δ₁[1] = 0, so every weight feeding the dead unit gets zero gradient — it learns nothing from this example.
Worked example
Same outer-product rule as layer 2, now with δ₁ = [−1.4621, 0] and the input x = [1, 2, 1]. This gives the (2×3) gradient for W₁:
Row 0 (active unit): −1.4621 · [1, 2, 1] = [−1.4621, −2.9242, −1.4621]
Why: The active hidden unit's incoming weights all get real gradient, scaled by each input feature.
Row 1 (dead unit): 0 · [1, 2, 1] = [0, 0, 0]
Why: δ₁[1] = 0, so every weight feeding the dead unit gets zero gradient — it learns nothing from this example.
\[ dW_1 = \begin{bmatrix} -1.4621 \\ 0 \end{bmatrix}\begin{bmatrix} 1 & 2 & 1 \end{bmatrix} = \begin{bmatrix} -1.4621 & -2.9242 & -1.4621 \\ 0 & 0 & 0 \end{bmatrix} \]
import numpy as np
x = np.array([1., 2., 1.]) # the input
dz1 = np.array([-1.4621, 0.]) # grad at z1 (dead unit zeroed)
dW1 = np.outer(dz1, x) # (2,) outer (3,) -> (2,3)
db1 = dz1
print(np.round(dW1, 4))
print(dW1.shape) # (2, 3), matches W1| j | ·x[0]=1 | ·x[1]=2 | ·x[2]=1 |
|---|---|---|---|
| 0 | −1.4621 | −2.9242 | −1.4621 |
| 1 | 0 | 0 | 0 |
| db₁ | = δ₁ | → | [−1.4621, 0] |
Comparison
Comparison matrix
From dW₁ = δ₁ xᵀ: refill the ·x[1]=2 column from what you know. The rest of the table is as it appeared.
| j | ·x[0]=1 | ·x[1]=2 | ·x[2]=1 |
|---|---|---|---|
| 0 | −1.4621 | −2.9242 | −1.4621 |
| 1 | 0 | 0 | 0 |
| db₁ | = δ₁ | → | [−1.4621, 0] |
Worked example
The backward pass is done. Here is every gradient we produced, in the order the sweep computed them — from the loss back to the first layer's weights:
The sweep: δ₂ → dW₂, db₂ → da₁ → δ₁ → dW₁, db₁
Why: Each arrow is one chain-rule move. Nothing was skipped, and each result fed the next. This ordering is the reverse topological sweep — the same order autograd uses.
| gradient | value | shape |
|---|---|---|
| δ₂ = p − e_y | [−0.7311, 0.7311] | (2,) |
| dW₂ = δ₂ a₁ᵀ | [[−2.1932, 0], [2.1932, 0]] | (2×2) |
| db₂ = δ₂ | [−0.7311, 0.7311] | (2,) |
| δ₁ = (W₂ᵀδ₂)⊙1[z₁>0] | [−1.4621, 0] | (2,) |
| dW₁ = δ₁ xᵀ | [[−1.4621, −2.9242, −1.4621], [0,0,0]] | (2×3) |
Every dW has the same shape as its W — the fastest sanity check there is. Two of the five rows are (partly) zero, and both traced to the one dead unit.
Discrimination
Sort into buckets
Sort these by shape, from memory, without looking back at The full gradient, one table. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Gradient just flows from a₁ back to z₁ unchanged — an affine map moved it, so set δ₁ = W₂ᵀδ₂ and carry on.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: That leftover 0.7311 flows into the DEAD unit, so dW₁ row 1 becomes [0.7311, 1.4622, 0.7311] instead of zeros.
Multiply by ReLU's local derivative — 1 where z₁ > 0, else 0 — before forming dW₁.
Why: That leftover 0.7311 flows into the DEAD unit, so dW₁ row 1 becomes [0.7311, 1.4622, 0.7311] instead of zeros. The shapes still match, so nothing errors — it silently disagrees with autograd.
Trap
Gradient just flows from a₁ back to z₁ unchanged — an affine map moved it, so set δ₁ = W₂ᵀδ₂ and carry on.
\[ \delta_1 = W_2^\top \delta_2 = \begin{bmatrix} -1.4621 \\ \mathbf{0.7311} \end{bmatrix} \;\;(\text{no mask}) \]
dW₁ picks up a bogus second row
Why: That leftover 0.7311 flows into the DEAD unit, so dW₁ row 1 becomes [0.7311, 1.4622, 0.7311] instead of zeros. The shapes still match, so nothing errors — it silently disagrees with autograd.
Multiply by ReLU's local derivative — 1 where z₁ > 0, else 0 — before forming dW₁.
\[ \delta_1 = (W_2^\top \delta_2) \odot \mathbb{1}[z_1>0] = \begin{bmatrix} -1.4621 \\ \mathbf{0} \end{bmatrix} \]
dW₁ row 1 is correctly all zeros
Why: The mask zeros the gradient for the inactive unit, matching autograd exactly. Use the cached z₁ for the mask, never a₁ — a₁ already lost the sign information.
Break the constraint
Discussion prompt
The rule this trap just fixed:
The mask zeros the gradient for the inactive unit, matching autograd exactly. Use the cached z₁ for the mask, never a₁ — a₁ already lost the sign information.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
That leftover 0.7311 flows into the DEAD unit, so dW₁ row 1 becomes [0.7311, 1.4622, 0.7311] instead of zeros. The shapes still match, so nothing errors — it silently disagrees with autograd.
Section
Part 4 of 7 — autograd & gradient check
Concept
The gradients aren't the goal — they tell us which way to move the weights to shrink the loss. One gradient-descent step subtracts a small multiple of each gradient:
\[ W_\ell \leftarrow W_\ell - \eta\, dW_\ell, \qquad b_\ell \leftarrow b_\ell - \eta\, db_\ell \]
For example W₂[0,0] = 1 with dW₂[0,0] = −2.1932 and step η = 0.1 becomes 1 − 0.1(−2.1932) = 1.219 — nudged up, because raising it raises the true class's score. Repeat forward→backward→step thousands of times and the net learns. Real training swaps plain descent for Adam (Lesson 15).
Worked example
Never trust a hand-derived gradient you haven't checked. Build the identical net in PyTorch, call .backward(), and compare with np.allclose. Full standalone script:
import numpy as np, torch
W1 = np.array([[1.,1.,0.],[1.,-1.,0.]]); b1 = np.array([0.,0.])
W2 = np.array([[1.,2.],[-1.,3.]]); b2 = np.array([-2.,5.])
x = np.array([1.,2.,1.]); y = 0
# manual forward + backward
z1 = W1@x+b1; a1 = np.maximum(0, z1); z2 = W2@a1+b2
p = np.exp(z2-z2.max()); p = p/p.sum()
dz2 = p - np.eye(2)[y]
dW2 = np.outer(dz2, a1)
dz1 = (W2.T @ dz2) * (z1 > 0)
dW1 = np.outer(dz1, x)
# torch autograd on the same net
tW1 = torch.tensor(W1, requires_grad=True); tW2 = torch.tensor(W2, requires_grad=True)
tb1 = torch.tensor(b1, requires_grad=True); tb2 = torch.tensor(b2, requires_grad=True)
zz = tW2 @ torch.relu(tW1 @ torch.tensor(x) + tb1) + tb2
torch.nn.functional.cross_entropy(zz.unsqueeze(0), torch.tensor([y])).backward()
print(np.allclose(dW1, tW1.grad.numpy()), np.allclose(dW2, tW2.grad.numpy()))Prints: True True
Why: Both weight gradients match autograd to machine precision. The ReLU mask was the only per-layer subtlety, and we got it right.
| gradient | shape | matches autograd |
|---|---|---|
| dW₂ | (2×2) | True |
| dW₁ | (2×3) | True |
| db₂ | (2,) | True |
| db₁ | (2,) | True |
Worked example
A second, independent check with no autograd at all: nudge one weight by ±ε and estimate its gradient from how the loss moves. This is the technique to keep when you write backprop from scratch.
\[ \frac{\partial L}{\partial W_{ij}} \approx \frac{L(W_{ij}+\varepsilon) - L(W_{ij}-\varepsilon)}{2\varepsilon} \]
import numpy as np
W1 = np.array([[1.,1.,0.],[1.,-1.,0.]]); b1 = np.array([0.,0.])
W2 = np.array([[1.,2.],[-1.,3.]]); b2 = np.array([-2.,5.])
x = np.array([1.,2.,1.]); y = 0
def loss(W1):
z1=W1@x+b1; a1=np.maximum(0,z1); z2=W2@a1+b2
p=np.exp(z2-z2.max()); p=p/p.sum(); return -np.log(p[y])
eps = 1e-6
Wp = W1.copy(); Wp[0,1] += eps
Wm = W1.copy(); Wm[0,1] -= eps
print(round((loss(Wp)-loss(Wm))/(2*eps), 5)) # -2.92423Numeric dW₁[0,1] = −2.92423, analytic = −2.92423
Why: The finite-difference estimate matches the outer-product value to 5 decimals — independent confirmation the backward pass is right, not just internally consistent with autograd.
| weight | finite-diff | analytic (δ₁ xᵀ) |
|---|---|---|
| dW₁[0,0] | −1.46212 | −1.46212 |
| dW₁[0,1] | −2.92423 | −2.92423 |
| dW₁[1,0] (dead) | 0.00000 | 0.00000 |
Concept
Formally the network is a DAG: nodes are operations, edges carry tensors forward and gradients backward. Our chain x → z₁ → a₁ → z₂ → p → L is that graph, drawn straight.
Figure (svg): A left-to-right chain of boxes: x, then z1 (affine W1,b1), then a1 (ReLU), then z2 (affine W2,b2), then p (softmax), then L (cross-entropy). A lower arrow runs right-to-left labeled 'gradients' from L back to x.
Backprop is one reverse topological sweep — each node multiplies the incoming gradient by its local Jacobian (Lesson 9). PyTorch's autograd is exactly this graph, built and swept automatically.
Section
Part 5 of 7 — vanishing, exploding, fixes
Concept
A gradient through n layers is a product of n per-layer Jacobians. Products of numbers compound: repeated factors below 1 collapse toward 0, repeated factors above 1 blow up (Lesson 9).
\[ \frac{\partial L}{\partial z_1} = \underbrace{J_n J_{n-1} \cdots J_2 J_1}_{n \text{ factors}} \,\frac{\partial L}{\partial z_n} \]
Below-1 factors (stacked sigmoids) give vanishing gradients — early layers stop learning. Above-1 factors (large weights) give exploding gradients — NaNs and wild updates.
Intuition
Interest compounds: 1% a day is harmless, but over a year it is a different number entirely. Gradients through layers compound the same way — each layer multiplies, and multiplication of many numbers is wildly sensitive to whether they sit above or below 1.
A hair below 1 and the product decays to nothing (the early layers go blind); a hair above 1 and it detonates (the update NaNs out). A healthy network keeps every layer's factor near 1 — which is exactly what good initialization buys you.
Worked example
Model the compounding with a single per-layer factor raised to the depth. The sigmoid's derivative peaks at 0.25, so a deep sigmoid stack multiplies by at most 0.25 per layer:
import numpy as np
for depth in (5, 10, 20):
print(f"depth {depth:2d}: 0.25^{depth} = {0.25**depth:.2e} "
f"1.5^{depth} = {1.5**depth:.2e}")
# why 0.25? the max slope of the sigmoid:
z = np.linspace(-6, 6, 1001); s = 1/(1+np.exp(-z))
print("max sigmoid'(z) =", round((s*(1-s)).max(), 4))0.25²⁰ ≈ 9e−13 (vanished); 1.5²⁰ ≈ 3325 (exploded)
Why: Twenty sigmoid layers shrink the gradient a trillion-fold — early weights get essentially no signal. A factor of 1.5 instead blows it into the thousands. Only a factor near 1 stays healthy.
| depth | 0.25ⁿ (vanish) | 1.5ⁿ (explode) |
|---|---|---|
| 5 | 9.77e−04 | 7.59e+00 |
| 10 | 9.54e−07 | 5.77e+01 |
| 20 | 9.09e−13 | 3.33e+03 |
Pattern
Step through it
Step through Vanishing vs exploding, in numbers one row at a time. What is driving the change, and what would the row after the last one be?
Concept
Clipping caps the gradient's norm so one huge step can't derail training. If the norm exceeds a threshold τ, rescale the whole vector down to length τ; otherwise leave it alone:
\[ g \leftarrow g \cdot \min\!\left(1, \frac{\tau}{\lVert g \rVert}\right) \]
The key word is norm: we scale the entire vector by one factor, so its direction is preserved — the step is shorter but still points downhill. This is the standard cure for exploding gradients.
Explain it
Discussion prompt
Explain Fix 1: gradient clipping to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Clipping caps the gradient's norm so one huge step can't derail training. If the norm exceeds a threshold τ, rescale the whole vector down to length τ; otherwise leave it alone:
Faded example
Fill in the blanks
Clip by norm keeps the direction, with the scaffolding fading: two lines are gone now — fill both.
import numpy as np
g = np.array([30., 40.]); tau = 5.0
norm = np.linalg.norm(g) # 50.0
g_clip = **g * min(1.0, tau / norm)** # scale the WHOLE vector
print(round(norm, 1), round(np.linalg.norm(g_clip), 1))
print(np.allclose(g_clip/np.linalg.norm(g_clip), g/norm))
Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. 3² + 4² = 25, so ‖[3,4]‖ = 5 = τ.
Worked example
A gradient g = [30, 40] has norm 50. With threshold τ = 5, scale it down by 5/50 = 0.1 and check the direction is unchanged:
g · (5/50) = [3, 4], norm 5
Why: 3² + 4² = 25, so ‖[3,4]‖ = 5 = τ. The classic 3-4-5 triangle: same direction as [30,40], one-tenth the length.
import numpy as np
g = np.array([30., 40.]); tau = 5.0
norm = np.linalg.norm(g) # 50.0
g_clip = g * min(1.0, tau / norm) # scale the WHOLE vector
print(round(norm, 1), round(np.linalg.norm(g_clip), 1))
print(np.allclose(g_clip/np.linalg.norm(g_clip), g/norm))| quantity | value |
|---|---|
| ‖g‖ before | 50.0 |
| g_clip | [3, 4] |
| ‖g_clip‖ | 5.0 |
| direction preserved | True |
Concept
Where the instability starts is the initial weights. He initialization scales each layer's weights so a ReLU layer preserves the variance of its input — keeping activations (and thus gradients) in a healthy range from step 0:
\[ W_{ij} \sim \mathcal{N}\!\left(0, \tfrac{2}{n_{\text{in}}}\right), \quad \text{i.e. std} = \sqrt{\tfrac{2}{n_{\text{in}}}} \]
The 2 is for ReLU (which zeros half its inputs); tanh/sigmoid use Xavier with a 1 instead. Too-large weights explode the forward variance; too-small ones vanish it.
Analogy
Discussion prompt
Explain Fix 2: careful initialization by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
The 2 is for ReLU (which zeros half its inputs); tanh/sigmoid use Xavier with a 1 instead. Too-large weights explode the forward variance; too-small ones vanish it.
Fill the middle
Fill in the blanks
From He init preserves variance — one line has had its right-hand side removed. Put it back.
import numpy as np
np.random.seed(0)
n_in = 100
x = np.random.randn(n_in) # input variance ~ 1
W_bad = np.random.randn(100, n_in) * 1.0 # too large
W_he = np.random.randn(100, n_in) * np.sqrt(2.0/n_in) # He
print(round(np.sqrt(2.0/n_in), 4))
print(round(np.var(np.maximum(0, W_bad @ x)), 2))
print(round(np.var(np.maximum(0, W_he @ x)), 2))
Why: x is what everything below it consumes, so the wrong expression here fails later and somewhere else. Std-1 weights amplify variance ~26x per layer — after a few layers that is an explosion.
Worked example
Feed a unit-variance input through a 100-wide ReLU layer with bad (std 1) vs He init, and measure the output variance. He keeps it near 1; std-1 weights explode it:
import numpy as np
np.random.seed(0)
n_in = 100
x = np.random.randn(n_in) # input variance ~ 1
W_bad = np.random.randn(100, n_in) * 1.0 # too large
W_he = np.random.randn(100, n_in) * np.sqrt(2.0/n_in) # He
print(round(np.sqrt(2.0/n_in), 4))
print(round(np.var(np.maximum(0, W_bad @ x)), 2))
print(round(np.var(np.maximum(0, W_he @ x)), 2))Bad init: output var 26.0; He init: 0.96
Why: Std-1 weights amplify variance ~26x per layer — after a few layers that is an explosion. He scaling holds it near 1, so signal neither grows nor dies with depth.
| init | weight std | output var (post-ReLU) |
|---|---|---|
| bad | 1.0 | 26.0 |
| He | √(2/100) = 0.1414 | 0.96 |
| goal | — | ≈ 1 (preserved) |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Bad init: output var 26.0; He init: 0.96
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Feed a unit-variance input through a 100-wide ReLU layer with bad (std 1) vs He init, and measure the output variance. He keeps it near 1; std-1 weights explode it:
Anomaly
Predict first
A student writes this, and it looks reasonable:
Just clamp every gradient component into [−τ, τ] — that shrinks the big ones, same idea as norm clipping.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Clamping components INDEPENDENTLY changes their ratio, so [30,40] (pointing at 53°) becomes [5,5] (pointing at 45°).
Scale the whole vector by τ/‖g‖ so every component shrinks by the same factor.
Why: Clamping components INDEPENDENTLY changes their ratio, so [30,40] (pointing at 53°) becomes [5,5] (pointing at 45°). You are no longer stepping downhill — you've turned the gradient.
Trap
Just clamp every gradient component into [−τ, τ] — that shrinks the big ones, same idea as norm clipping.
\[ g = \text{clip}(g, -\tau, \tau): \quad \begin{bmatrix} 30 \\ 40 \end{bmatrix} \to \begin{bmatrix} 5 \\ 5 \end{bmatrix} \]
Direction bent from ~53° to 45°
Why: Clamping components INDEPENDENTLY changes their ratio, so [30,40] (pointing at 53°) becomes [5,5] (pointing at 45°). You are no longer stepping downhill — you've turned the gradient.
Scale the whole vector by τ/‖g‖ so every component shrinks by the same factor.
\[ g \cdot \tfrac{\tau}{\lVert g \rVert}: \quad \begin{bmatrix} 30 \\ 40 \end{bmatrix} \to \begin{bmatrix} 3 \\ 4 \end{bmatrix} \]
Direction preserved exactly
Why: [3,4] is [30,40] scaled by 0.1 — identical direction, norm 5. Clip-by-VALUE is a different (occasionally useful) tool, but it is NOT a norm rescale and should never be swapped in for one.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
W₁, W₂), one hidden ReLU layer with 2 units, 2 output classes. Small enough to do entirely by hand — big enough to show every real subtlety.; The whole lesson is those two passes on one tiny net — small enough that every blame-number is a fraction you can check by hand.; Logits z₂ are unbounded real numbers. Softmax exponentiates then normalizes, producing a probability vector p that is positive and sums to 1:a₁ back to z₁ unchanged — an affine map moved it, so set δ₁ = W₂ᵀδ₂ and carry on.; Just clamp every gradient component into [−τ, τ] — that shrinks the big ones, same idea as norm clipping.Section
Part 6 of 7
Constraint
Discussion prompt
Run The full MLP recipe with this step confiscated:
Backward, per layer: weight grad dW = δ aᵀ (outer product), bias grad db = δ, push back with Wᵀδ, and cross each ReLU with the 1[z>0] mask
Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.
Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.
Answer:
z and a on the wayδ = p − e_y (the Lesson 17 shortcut)dW = δ aᵀ (outer product), bias grad db = δ, push back with Wᵀδ, and cross each ReLU with the 1[z>0] maskPattern
z and a on the wayδ = p − e_y (the Lesson 17 shortcut)dW = δ aᵀ (outer product), bias grad db = δ, push back with Wᵀδ, and cross each ReLU with the 1[z>0] maskEdge cases
Discussion prompt
The full MLP recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
z and a on the wayδ = p − e_y (the Lesson 17 shortcut)dW = δ aᵀ (outer product), bias grad db = δ, push back with Wᵀδ, and cross each ReLU with the 1[z>0] maskElimination
Eliminate the wrong options
In a linear layer z = Wa + b with upstream gradient δ = ∂L/∂z and input a, the weight gradient dW is:
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: dW = δ aᵀ is the outer product, giving the (out×in) shape of W — each weight's gradient is its output-error δ_j times the input a_k it multiplied. In our net, δ₂=[−0.7311,0.7311] outer a₁=[3,0] gave dW₂ with the correct (2×2) shape.
Check
What shape, and built from what?
Check your understanding
In a linear layer z = Wa + b with upstream gradient δ = ∂L/∂z and input a, the weight gradient dW is:
Answer: A
Why: dW = δ aᵀ is the outer product, giving the (out×in) shape of W — each weight's gradient is its output-error δ_j times the input a_k it multiplied. In our net, δ₂=[−0.7311,0.7311] outer a₁=[3,0] gave dW₂ with the correct (2×2) shape.
Prediction
Predict first
During the backward pass, what happens to the gradient flowing into a ReLU unit whose forward pre-activation z was negative?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: It is multiplied by 0 — the gradient is blocked
Why: ReLU has slope 0 where z < 0, so its local derivative is 0 there. The mask 1[z>0] zeros that unit's gradient — in our net, da₁[1]=0.7311 was killed to 0, and W₁'s second row got zero gradient.
Check
A hidden unit had pre-activation z = −1 on this example.
Check your understanding
During the backward pass, what happens to the gradient flowing into a ReLU unit whose forward pre-activation z was negative?
Answer: A
Why: ReLU has slope 0 where z < 0, so its local derivative is 0 there. The mask 1[z>0] zeros that unit's gradient — in our net, da₁[1]=0.7311 was killed to 0, and W₁'s second row got zero gradient.
Prediction
Predict first
Training diverges: loss → NaN and gradient norms reach the thousands. The most direct fix is:
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Gradient clipping — cap the gradient norm
Why: NaNs with huge gradient norms are the signature of exploding gradients. Clipping the norm caps the step so a single blow-up can't derail training; He initialization prevents it from starting. We saw 1.5²⁰ ≈ 3325 — clipping to τ would tame exactly that.
Check
Pick the most direct fix.
Check your understanding
Training diverges: loss → NaN and gradient norms reach the thousands. The most direct fix is:
Answer: A
Why: NaNs with huge gradient norms are the signature of exploding gradients. Clipping the norm caps the step so a single blow-up can't derail training; He initialization prevents it from starting. We saw 1.5²⁰ ≈ 3325 — clipping to τ would tame exactly that.
Elimination
Eliminate the wrong options
To shrink a large gradient WITHOUT changing its direction, you should:
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Multiplying the whole vector by τ/‖g‖ shrinks every component by the same factor, so magnitude drops but direction is preserved — clip-by-norm. Verified: [30,40] (norm 50) → [3,4] (norm 5), same direction.
Check
Direction matters.
Check your understanding
To shrink a large gradient WITHOUT changing its direction, you should:
Answer: A
Why: Multiplying the whole vector by τ/‖g‖ shrinks every component by the same factor, so magnitude drops but direction is preserved — clip-by-norm. Verified: [30,40] (norm 50) → [3,4] (norm 5), same direction.
Concept
Everything so far used one example, x ∈ ℝ³. Real training feeds a batch: stack N examples as rows of a matrix X ∈ ℝ^{N×d}. The forward pass becomes Z = XW + b (broadcast bias), one matmul for the whole batch.
The backward pass barely changes. The per-example outer product δ aᵀ becomes a matmul that sums over the batch, dW = AᵀΔ, and we average by dividing by N. The ReLU mask and the p − e_y seed apply row-wise, unchanged.
\[ dW = \tfrac{1}{N}\,A^\top \Delta, \qquad db = \tfrac{1}{N}\sum_{n} \delta^{(n)} \]
So the hand-derived rules scale directly to the project's (1347 × 64) batch — the shapes grow, the moves are identical.
Pattern
Predict first
The table runs: X | (1347, 64) | batch of images · Z1 = X W1 + b1 | (1347, 32) | hidden pre-activations · A1 = ReLU(Z1) | (1347, 32) | hidden activations
In Batched shapes on the digits data, given the rows so far: what is the next one — the row where tensor is Z2 = A1 W2 + b2?
Correct: Z2 = A1 W2 + b2 | (1347, 10) | class logits
| tensor | shape | meaning |
|---|---|---|
| X | (1347, 64) | batch of images |
| Z1 = X W1 + b1 | (1347, 32) | hidden pre-activations |
| A1 = ReLU(Z1) | (1347, 32) | hidden activations |
| Z2 = A1 W2 + b2 | (1347, 10) | class logits |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. The batch dimension 1347 rides along untouched; the 64 features contract against W1's 64 rows, leaving 32 hidden units per example.
Worked example
Trace the shapes through one batched forward pass of the 64 → 32 → 10 net on the full training batch of 1347 images, so nothing surprises you in the project:
X (1347×64) @ W1 (64×32) → Z1 (1347×32)
Why: The batch dimension 1347 rides along untouched; the 64 features contract against W1's 64 rows, leaving 32 hidden units per example.
ReLU keeps (1347×32); then @ W2 (32×10) → Z2 (1347×10)
Why: ReLU is entrywise so shape is unchanged. The second matmul contracts the 32 hidden units to 10 class logits per example.
\[ X \xrightarrow{W_1} Z_1 \xrightarrow{\text{ReLU}} A_1 \xrightarrow{W_2} Z_2, \quad Z_2 \in \mathbb{R}^{1347 \times 10} \]
| tensor | shape | meaning |
|---|---|---|
| X | (1347, 64) | batch of images |
| Z1 = X W1 + b1 | (1347, 32) | hidden pre-activations |
| A1 = ReLU(Z1) | (1347, 32) | hidden activations |
| Z2 = A1 W2 + b2 | (1347, 10) | class logits |
Notation
Annotate
From Batched shapes on the digits data — read this one piece at a time. What is each part doing?
On: \( X \xrightarrow{W_1} Z_1 \xrightarrow{\text{ReLU}} A_1 \xrightarrow{W_2} Z_2, \quad Z_2 \in \mathbb{R}^{1347 \times 10} \)
Section
Part 7 of 7 — the project
Intuition
You could call nn.Sequential and never think about a single gradient. But the exam probes whether you know what .backward() does — the outer product, the transpose, the mask. Building it once, by hand, makes those reflexes instead of trivia.
So do both: train the PyTorch version to see it work, then re-derive it in raw NumPy to prove the backward pass in your head matches the one in the library. Same accuracy from both is the proof.
Concept
Train a one-hidden-layer network on load_digits (8×8 handwritten digit images, 64 features, 10 classes) and beat 90% test accuracy — everything you derived, now assembled into a real classifier.
| # | requirement | tool |
|---|---|---|
| 1 | Load + split digits, scale pixels to [0,1] | load_digits, train_test_split |
| 2 | 64 → 32 → 10 MLP with ReLU + Adam | nn.Sequential, Adam |
| 3 | Train with cross-entropy, report test accuracy | cross_entropy, argmax |
| 4 | (Stretch) re-implement forward+backward in pure NumPy | np.outer, the ReLU mask |
Build rules: type every line yourself, divide pixels by 16 to normalize, seed with torch.manual_seed(0) so your numbers match, and judge on the held-out test split, never the training data.
Counterexample
Discussion prompt
Train a one-hidden-layer network on load_digits (8×8 handwritten digit images, 64 features, 10 classes) and beat 90% test accuracy — everything you derived, now assembled into a real classifier.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Build rules: type every line yourself, divide pixels by 16 to normalize, seed with torch.manual_seed(0) so your numbers match, and judge on the held-out test split, never the training data.
Worked example
Your turn: load digits, normalize to [0,1], and split 75/25. Predict the feature dimension out loud before you print it.
Hint: load_digits(), divide .data by 16 (pixels run 0–16), then train_test_split(..., test_size=0.25, random_state=0). Each image is 8×8 = 64 features.
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
D = load_digits()
X = D.data / 16.0 # 8x8 pixels in [0,16] -> [0,1]
y = D.target
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.25, random_state=0)
print(X.shape, Xtr.shape, Xte.shape, len(set(y)))| object | shape / value |
|---|---|
| X (all) | (1797, 64) |
| Xtr | (1347, 64) |
| Xte | (450, 64) |
| classes | 10 |
Pattern
Step through it
Step through Milestone 1 — the data one row at a time. What is driving the change, and what would the row after the last one be?
Worked example
Your turn: build a 64 → 32 → 10 ReLU MLP and an Adam optimizer. Predict layer 1's parameter count before printing.
Hint: nn.Sequential(nn.Linear(64,32), nn.ReLU(), nn.Linear(32,10)), then Adam(net.parameters(), lr=0.01). Layer 1 has 64·32 weights + 32 biases.
import torch, torch.nn as nn
torch.manual_seed(0)
net = nn.Sequential(nn.Linear(64,32), nn.ReLU(), nn.Linear(32,10))
opt = torch.optim.Adam(net.parameters(), lr=0.01)
print(sum(p.numel() for p in net.parameters()))
print(64*32 + 32, 32*10 + 10)| layer | weights + biases | params |
|---|---|---|
| Linear 64→32 | 64·32 + 32 | 2080 |
| Linear 32→10 | 32·10 + 10 | 330 |
| total | — | 2410 |
Pattern
Step through it
Step through Milestone 2 — the network one row at a time. What is driving the change, and what would the row after the last one be?
Worked example
Your turn: train 200 epochs of full-batch cross-entropy, then measure held-out accuracy. Predict: above or below 90%?
Hint: each epoch runs opt.zero_grad(), loss = F.cross_entropy(net(Xtr_t), ytr_t), loss.backward(), opt.step(). Accuracy = (argmax(1) == y).float().mean().
import torch, torch.nn as nn, torch.nn.functional as F
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
D = load_digits(); X = D.data/16.0; y = D.target
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.25, random_state=0)
Xtr_t = torch.tensor(Xtr, dtype=torch.float32); ytr_t = torch.tensor(ytr)
Xte_t = torch.tensor(Xte, dtype=torch.float32); yte_t = torch.tensor(yte)
torch.manual_seed(0)
net = nn.Sequential(nn.Linear(64,32), nn.ReLU(), nn.Linear(32,10))
opt = torch.optim.Adam(net.parameters(), lr=0.01)
for ep in range(200):
opt.zero_grad()
F.cross_entropy(net(Xtr_t), ytr_t).backward()
opt.step()
acc = (net(Xte_t).argmax(1) == yte_t).float().mean().item()
print(round(acc, 4)) # 0.9689| epoch | train loss | note |
|---|---|---|
| 0 | 2.3255 | ≈ log(10), untrained |
| 50 | 0.1329 | learning fast |
| 100 | 0.0518 | nearly fit |
| 199 | 0.0157 | converged |
Missing information
Discussion prompt
The full trained program, with gradient-norm clipping added for stability. Test accuracy lands at 0.9689 — comfortably past the 90% bar.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Clip-by-norm here is a safety net (the net is small and stable), and the result matches the unclipped run to 4 decimals — clipping never hurts a well-behaved run, and saves an exploding one.
Worked example
The full trained program, with gradient-norm clipping added for stability. Test accuracy lands at 0.9689 — comfortably past the 90% bar.
import torch, torch.nn as nn, torch.nn.functional as F
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
D = load_digits(); X = D.data/16.0; y = D.target
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.25, random_state=0)
Xtr_t = torch.tensor(Xtr, dtype=torch.float32); ytr_t = torch.tensor(ytr)
Xte_t = torch.tensor(Xte, dtype=torch.float32); yte_t = torch.tensor(yte)
torch.manual_seed(0)
net = nn.Sequential(nn.Linear(64,32), nn.ReLU(), nn.Linear(32,10))
opt = torch.optim.Adam(net.parameters(), lr=0.01)
for ep in range(200):
opt.zero_grad()
loss = F.cross_entropy(net(Xtr_t), ytr_t)
loss.backward()
torch.nn.utils.clip_grad_norm_(net.parameters(), 5.0) # stabilize
opt.step()
acc = (net(Xte_t).argmax(1) == yte_t).float().mean().item()
print('final train loss:', round(loss.item(), 4))
print('test accuracy :', round(acc, 4))final train loss 0.0157, test accuracy 0.9689
Why: Clip-by-norm here is a safety net (the net is small and stable), and the result matches the unclipped run to 4 decimals — clipping never hurts a well-behaved run, and saves an exploding one.
| metric | value |
|---|---|
| final train loss | 0.0157 |
| test accuracy | 0.9689 (~97%) |
| target | > 0.90 ✓ |
Trade off
Comparison matrix
From The result: ~97% on held-out digits: every row here is a choice with a cost. Fill the value column, then say which row you would actually pick and what you give up for it.
| metric | value |
|---|---|
| final train loss | 0.0157 |
| test accuracy | 0.9689 (~97%) |
| target | > 0.90 ✓ |
Estimation
Predict first
Your turn (stretch): no autograd at all. Batched forward, softmax+CE seed P − Y, then the exact three moves you derived — aᵀδ for weights, Wᵀδ to push back, ⊙ (z>0) for the ReLU. He init, plain SGD.
Commit before you compute: what does Stretch — the whole net in pure NumPy come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: scratch NumPy test acc = 0.9622
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. No PyTorch, no autograd — just np.outer's batched cousin (aᵀδ), the Wᵀδ push-back, and the (z>0) mask, exactly the moves from Part 3.
Worked example
Your turn (stretch): no autograd at all. Batched forward, softmax+CE seed P − Y, then the exact three moves you derived — aᵀδ for weights, Wᵀδ to push back, ⊙ (z>0) for the ReLU. He init, plain SGD.
import numpy as np
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
D = load_digits(); X = D.data/16.0; y = D.target
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.25, random_state=0)
rng = np.random.default_rng(0)
W1 = rng.standard_normal((64,32))*np.sqrt(2/64); b1 = np.zeros(32) # He
W2 = rng.standard_normal((32,10))*np.sqrt(2/32); b2 = np.zeros(10)
Y = np.eye(10)[ytr]
for ep in range(300):
z1 = Xtr@W1 + b1; a1 = np.maximum(0, z1) # forward
z2 = a1@W2 + b2
P = np.exp(z2 - z2.max(1, keepdims=True)); P /= P.sum(1, keepdims=True)
dz2 = (P - Y) / len(Xtr) # softmax+CE seed
gW2 = a1.T@dz2; dz1 = (dz2@W2.T) * (z1 > 0) # backward + ReLU mask
gW1 = Xtr.T@dz1
W1 -= 0.5*gW1; b1 -= 0.5*dz1.sum(0) # SGD step
W2 -= 0.5*gW2; b2 -= 0.5*dz2.sum(0)
pred = (np.maximum(0, Xte@W1 + b1)@W2 + b2).argmax(1)
print('scratch NumPy test acc:', round((pred == yte).mean(), 4))scratch NumPy test acc = 0.9622
Why: No PyTorch, no autograd — just np.outer's batched cousin (aᵀδ), the Wᵀδ push-back, and the (z>0) mask, exactly the moves from Part 3. Hits ~96%, confirming you understand backprop end to end.
| implementation | test accuracy |
|---|---|
| PyTorch autograd (Milestone 3) | 0.9689 |
| from-scratch NumPy | 0.9622 |
| both | > 0.90 ✓ |
Comparison
Comparison matrix
From Stretch — the whole net in pure NumPy: refill the test accuracy column from what you know. The rest of the table is as it appeared.
| implementation | test accuracy |
|---|---|
| PyTorch autograd (Milestone 3) | 0.9689 |
| from-scratch NumPy | 0.9622 |
| both | > 0.90 ✓ |
Concept
Slides closed, out loud: explain (1) why you cache z₁ (not just a₁) on the forward pass, (2) what the ReLU mask did to the dead unit's gradient in our worked example, and (3) why clip-by-norm preserves the descent direction while clip-by-value does not.
Stretch homework: add a gradient check to the NumPy net (compare aᵀδ against finite differences), then extend it to arbitrary depth and add an L2 penalty. This same manual backward becomes BatchNorm backprop (Week 18) and transformer backprop (Week 30).
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — The running network · The forward pass · The backward pass · Verify everything · Gradient pathologies · Pattern & checks. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
z₁=[3,−1], p=[0.2689,0.7311], L=1.3133δ₂ = p − e_y, weight grad dW = δ aᵀ, push back with Wᵀδ, cross ReLU with 1[z>0]| step | the one thing to remember |
|---|---|
| forward | cache every z AND a |
| softmax+CE seed | δ = p − e_y |
| weight gradient | dW = δ aᵀ (outer product, shape of W) |
| push back | da = Wᵀδ |
| ReLU backward | multiply by 1[z>0] — from cached z |
| exploding | clip by NORM (keeps direction), He init |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.