USAAIO Lesson 9, from Week 3 on calculus, fully worked. It builds the single-variable chain rule from local rates, then the multivariable sum-over-paths rule, then derives the Jacobian form and pins down the order in which the multiplications happen. From there it presents backpropagation as reverse-mode autodiff through a concrete two-layer sigmoid network: every forward and backward number is pushed through by hand, one beat at a time - z1, h, y-hat, L, then dy-hat, dW2, dh, sigma', dz1, dW1 - and checked against torch.autograd to machine precision. It quantifies vanishing gradients as (0.25)^depth, covers finite-difference gradient checking, and ends with a your-turn build of the whole backward pass. Every snippet runs as written, and every number came from real execution. The lesson runs to 64 slides.
Subject: Machine Learning · 116 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 9 · Week 3 (Calculus)
How a network learns: multiply local derivatives backward through the graph. We derive the chain rule three ways — scalar, sum-over-paths, and Jacobian — then push one concrete 2-layer net all the way through by hand and match torch.autograd to the last digit.
Objectives
∂L/∂L = 1 and sweep Jacobians from the loss back to every parametertorch.autograd to machine precision(0.25)^depth, explain why ReLU survives, and gradient-check any analytic derivativeWarm-up
Discussion prompt
Before we open Lesson 9: The Chain Rule & Backpropagation: without looking back, what was the main idea of Maximum Likelihood Estimation, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
MLE as an optimization principle, the log-likelihood trick, MLE for Gaussian/Bernoulli, and the big idea that MSE and cross-entropy ARE negative log-likelihoods — plus MAP as regularized MLE. Build maximum_likelihood_gaussian from scratch.
Section
Part 1 of 6 — local rates
Concept
A composition y = f(g(x)) is two machines wired in series: x goes into g to make an intermediate u = g(x), and u goes into f to make y = f(u).
\[ x \;\xrightarrow{\;g\;}\; u = g(x) \;\xrightarrow{\;f\;}\; y = f(u) \]
To ask how y responds to a nudge in x, we ask two smaller questions: how does u respond to x, and how does y respond to u? Then we combine them.
Counterexample
Discussion prompt
A composition y = f(g(x)) is two machines wired in series: x goes into g to make an intermediate u = g(x), and u goes into f to make y = f(u).
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
To ask how y responds to a nudge in x, we ask two smaller questions: how does u respond to x, and how does y respond to u? Then we combine them.
Intuition
Think of two meshed gears. If the first turns 3× as fast as its input and the second turns 2× as fast as the first, the last spins 6× the input — the rates multiply, they don't add.
A tiny bump Δx becomes Δu ≈ g'(x)·Δx after the first machine, and that becomes Δy ≈ f'(u)·Δu after the second. Substitute the first into the second.
So Δy ≈ f'(u)·g'(x)·Δx. Divide by Δx: the overall rate is the product of the two local rates. That single idea drives every gradient in this lesson.
Analogy
Discussion prompt
Explain Rates multiply, like gears by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Think of two meshed gears. If the first turns 3× as fast as its input and the second turns 2× as fast as the first, the last spins 6× the input — the rates multiply, they don't add.
Concept
Written cleanly, with u = g(x):
\[ \frac{dy}{dx} = \frac{dy}{du}\cdot\frac{du}{dx} = f'(g(x))\cdot g'(x) \]
The Leibniz form dy/dx = (dy/du)(du/dx) makes it look like the dus cancel. They don't literally cancel, but the mnemonic is exactly right: outer derivative (at the inner value) times inner derivative.
Ranking
Put in order
Put the moves of Chain rule on y = sin(x²) into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. The inner machine squares its input; its local rate is 2x, which is 2.6 at x = 1.3.
Worked example
Take y = sin(x²), evaluated at x = 1.3. Split it into inner u = g(x) = x² and outer f(u) = sin(u).
Inner: u = x² = 1.3² = 1.69, and du/dx = 2x = 2.6
Why: The inner machine squares its input; its local rate is 2x, which is 2.6 at x = 1.3.
Outer: dy/du = cos(u) = cos(1.69) = −0.118922
Why: The outer machine is sine; its local rate is cosine, evaluated at the INNER value u = 1.69, not at x.
Multiply: dy/dx = cos(1.69)·2.6 = −0.118922·2.6 = −0.309196
Why: Product of the two local rates. Watch the sign flip — the outer slope is negative here.
| piece | expression | value at x = 1.3 |
|---|---|---|
| inner u | x² | 1.69 |
| du/dx | 2x | 2.6 |
| dy/du | cos(u) | −0.118922 |
| dy/dx | cos(u)·2x | −0.309196 |
Comparison
Comparison matrix
From Chain rule on y = sin(x²): refill the expression column from what you know. The rest of the table is as it appeared.
| piece | expression | value at x = 1.3 |
|---|---|---|
| inner u | x² | 1.69 |
| du/dx | 2x | 2.6 |
| dy/du | cos(u) | −0.118922 |
| dy/dx | cos(u)·2x | −0.309196 |
Concept
Every analytic derivative can be sanity-checked numerically. A central finite difference nudges the input by ±ε and measures the output change:
\[ \frac{dy}{dx} \approx \frac{y(x+\epsilon) - y(x-\epsilon)}{2\epsilon} \]
Central (both sides) is far more accurate than one-sided: its error shrinks like ε², not ε. We use this all lesson as the ground truth check.
Explain it
Discussion prompt
Explain Trust, but verify: the finite difference to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Every analytic derivative can be sanity-checked numerically. A central finite difference nudges the input by ±ε and measures the output change:
Estimation
Predict first
Confirm −0.309196 by nudging x = 1.3 by ε = 1e−6. Runnable as written:
Commit before you compute: what does Numeric check of the sin(x²) derivative come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: analytic = −0.309196, numeric = −0.309196 — they match
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Six-digit agreement means the split into inner/outer and the product were correct.
Worked example
Confirm −0.309196 by nudging x = 1.3 by ε = 1e−6. Runnable as written:
import numpy as np
x0 = 1.3
analytic = np.cos(x0**2) * 2*x0 # cos(u)*2x
eps = 1e-6
num = (np.sin((x0+eps)**2) - np.sin((x0-eps)**2)) / (2*eps)
print(round(analytic, 6), round(num, 6))analytic = −0.309196, numeric = −0.309196 — they match
Why: Six-digit agreement means the split into inner/outer and the product were correct. If these disagreed, one of the two local rates would be wrong.
| method | dy/dx at x = 1.3 |
|---|---|
| analytic (chain rule) | −0.309196 |
| central difference (ε = 1e−6) | −0.309196 |
| difference | < 1e−9 |
Trade off
Comparison matrix
From Numeric check of the sin(x²) derivative: every row here is a choice with a cost. Fill the dy/dx at x = 1.3 column, then say which row you would actually pick and what you give up for it.
| method | dy/dx at x = 1.3 |
|---|---|
| analytic (chain rule) | −0.309196 |
| central difference (ε = 1e−6) | −0.309196 |
| difference | < 1e−9 |
Concept
Longer chains are the same move, repeated. Take a = 2, then u = a², then v = sin(u). The chain rule multiplies every local rate along the path from v back to a.
\[ \frac{dv}{da} = \frac{dv}{du}\cdot\frac{du}{da} = \cos(u)\cdot 2a \]
Notice the two-phase structure that IS backprop: first sweep forward to get the intermediate values (u = 4), then sweep backward multiplying rates (cos(4)·2·2).
Step zero
Discussion prompt
Forward then backward on v = sin(a²) — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Forward: u = a² = 4.0, then v = sin(4.0) = −0.756802
Answer:
Worked example
Push a = 2 forward to values, then multiply rates backward. Every number verified:
import numpy as np
a = 2.0
u = a**2 # forward: intermediate
v = np.sin(u) # forward: output
dvdu = np.cos(u) # backward: local rate of outer
duda = 2*a # backward: local rate of inner
dvda = dvdu * duda
print(u, round(v,6), round(dvdu,6), duda, round(dvda,6))Forward: u = a² = 4.0, then v = sin(4.0) = −0.756802
Why: Cache u — the backward pass needs it to evaluate cos(u). This caching is the seed of the whole backprop idea.
Backward: dv/du = cos(4.0) = −0.653644, du/da = 2·2 = 4.0
Why: Each machine reports its own local rate at the value flowing through it.
Multiply along the path: dv/da = −0.653644·4.0 = −2.614574
Why: Verified against a central difference: numeric dv/da = −2.614574 as well.
| stage | value / rate |
|---|---|
| u = a² | 4.0 |
| v = sin(u) | −0.756802 |
| dv/du = cos(u) | −0.653644 |
| du/da = 2a | 4.0 |
| dv/da (product) | −2.614574 |
Section
Part 2 of 6 — sum over paths
Concept
Now let one input t flow through two intermediates and then combine: u = u(t), w = w(t), and L = L(u, w). A nudge in t reaches L along two different paths.
Figure (svg): A node t branches to two nodes u and w, which both feed into a node L. Two paths t to u to L and t to w to L.
Concept
The multivariable chain rule says: multiply the rates along each path, then add the paths. Each route contributes its own product of local rates.
\[ \frac{dL}{dt} = \frac{\partial L}{\partial u}\frac{du}{dt} + \frac{\partial L}{\partial w}\frac{dw}{dt} \]
The single-variable rule was the one-path special case. The + is what shows up whenever a quantity is used in more than one place — and in a network, a shared weight or activation is used in many places.
Intuition
If t gets a little bigger, u shifts and pushes L one way, AND w shifts and pushes L another way. Both pushes happen at once, so the total change in L is the sum of the two contributions.
In backprop this becomes: a node that fans out to several children collects a gradient from each child and adds them. 'Sum the incoming gradients' is this rule wearing work clothes.
Step zero
Discussion prompt
Two-path example, checked numerically — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Path through u: (∂L/∂u)(du/dt) = w · 2t = (6)(4) = 24
Answer:
Worked example
Let u = t², w = 3t, and L = u·w. Then L(t) = t²·3t = 3t³, so we already know dL/dt = 9t². Let's get the same thing by summing paths at t = 2.
Path through u: (∂L/∂u)(du/dt) = w · 2t = (6)(4) = 24
Why: ∂L/∂u = w = 3t = 6 at t = 2; du/dt = 2t = 4. Their product is this path's contribution.
Path through w: (∂L/∂w)(dw/dt) = u · 3 = (4)(3) = 12
Why: ∂L/∂w = u = t² = 4; dw/dt = 3. Product is 12.
Add the paths: dL/dt = 24 + 12 = 36
Why: And the closed form 9t² = 9·4 = 36 agrees exactly. The sum-over-paths rule reproduces the direct answer.
| path | rate product | value at t = 2 |
|---|---|---|
| through u | w · 2t | 24 |
| through w | u · 3 | 12 |
| total dL/dt | sum | 36 |
| check: 9t² | direct | 36 |
Discrimination
Sort into buckets
Sort these by value at t = 2, from memory, without looking back at Two-path example, checked numerically. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Section
Part 3 of 6 — vectors & matrices
Concept
For a vector map f: ℝⁿ → ℝᵐ, the Jacobian J collects every partial derivative: row i, column j is ∂fᵢ/∂xⱼ. It is m × n — outputs down the rows, inputs across the columns.
\[ J_f = \begin{bmatrix} \dfrac{\partial f_1}{\partial x_1} & \cdots & \dfrac{\partial f_1}{\partial x_n} \\[4pt] \vdots & \ddots & \vdots \\[4pt] \dfrac{\partial f_m}{\partial x_1} & \cdots & \dfrac{\partial f_m}{\partial x_n} \end{bmatrix} \]
It is the best linear approximation of f at a point: f(x + Δ) ≈ f(x) + J_f Δ. The scalar f'(x) was the 1×1 Jacobian all along.
Concept
Stack the sum-over-paths rule for every output-input pair and it packs into one clean matrix product — the Jacobian of a composition is the product of the Jacobians:
\[ J_{f\circ g}(x) = J_f\big(g(x)\big)\, J_g(x) \]
Entry (i,k) of that product is Σⱼ (∂fᵢ/∂uⱼ)(∂uⱼ/∂xₖ) — exactly 'sum the rates over every intermediate path.' The matrix multiply IS the summation, done in bulk.
Concept
Every layer in our net is one of two maps. Their Jacobians are worth memorizing:
\[ g(x) = Wx + b \;\Longrightarrow\; J_g = W \]
\[ \phi(z) = \sigma(z)\ \text{(elementwise)} \;\Longrightarrow\; J_\phi = \operatorname{diag}\big(\sigma'(z)\big) \]
A linear layer's Jacobian is just its weight matrix W. An elementwise activation's Jacobian is diagonal — output i depends only on input i, so off-diagonal partials are zero.
Sorting
Sort into buckets
These are the pieces of Lesson 9: The Chain Rule & Backpropagation, out of order. Put each one back under the part of the lesson it belongs to.
Missing information
Discussion prompt
Compose the two: g(x) = W₁x + b₁ then φ(z) = σ(z), so h = σ(W₁x + b₁). With our real W₁ and the σ'(z₁) = h(1−h) from the forward pass, J = diag(σ') · W₁. Runnable:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
diag(σ') is 3×3, W₁ is 3×2; inner dimensions (3 and 3) match only in this order, giving the correct 3×2 Jacobian dh/dx.
Worked example
Compose the two: g(x) = W₁x + b₁ then φ(z) = σ(z), so h = σ(W₁x + b₁). With our real W₁ and the σ'(z₁) = h(1−h) from the forward pass, J = diag(σ') · W₁. Runnable:
import numpy as np
x = np.array([1., 2.])
W1 = np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1 = np.array([0.1,0.2,0.3])
sig = lambda z: 1/(1+np.exp(-z))
z1 = W1@x + b1; h = sig(z1)
Jf = np.diag(h*(1-h)) # 3x3 diagonal
Jcomp = Jf @ W1 # (3x3)(3x2) = 3x2
print(Jcomp[:,0]) # dh/dx0Shapes: (3×3)·(3×2) = (3×2) — outer Jacobian on the LEFT
Why: diag(σ') is 3×3, W₁ is 3×2; inner dimensions (3 and 3) match only in this order, giving the correct 3×2 Jacobian dh/dx.
Column 0 = dh/dx₀ = [0.022878, 0.05049, 0.052497]
Why: A central finite difference on x₀ gives the identical [0.022878, 0.05049, 0.052497]. The Jacobian product is exactly the numerical sensitivity.
| source | dh/dx₀ |
|---|---|
| Jf @ W1, column 0 | [0.022878, 0.05049, 0.052497] |
| central difference on x₀ | [0.022878, 0.05049, 0.052497] |
| match? | yes, to 1e−6 |
Error analysis
Annotate
Walk the callouts on Multiply two layer Jacobians, verified. Each one is a place this is easy to get subtly wrong.
Ranking
Put in order
These are the steps of The chain-rule pattern, all three forms, scrambled. Put them back in order before the next slide shows you.
dy/dx = f'(g(x))·g'(x)dL/dt = Σ (∂L/∂uⱼ)(duⱼ/dt)J_{f∘g} = J_f · J_g, outer Jacobian firstε ≈ 1e−5 to 1e−6Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
dy/dx = f'(g(x))·g'(x)dL/dt = Σ (∂L/∂uⱼ)(duⱼ/dt)J_{f∘g} = J_f · J_g, outer Jacobian firstε ≈ 1e−5 to 1e−6Anomaly
Predict first
A student writes this, and it looks reasonable:
We apply g first, then f, so write the Jacobians in that left-to-right reading order: J_g · J_f.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: With g: ℝ² → ℝ³ (J_g is 3×2) and f: ℝ³ → ℝ³ (J_f is 3×3), the product J_g·J_f is (3×2)(3×3) — inner dims 2 and 3 don't match.
The outer function's Jacobian comes first (on the left), mirroring f'(g(x))·g'(x).
Why: With g: ℝ² → ℝ³ (J_g is 3×2) and f: ℝ³ → ℝ³ (J_f is 3×3), the product J_g·J_f is (3×2)(3×3) — inner dims 2 and 3 don't match. It is not even a legal multiply.
Trap
We apply g first, then f, so write the Jacobians in that left-to-right reading order: J_g · J_f.
J_{f∘g} = J_g J_f ✗
Why: With g: ℝ² → ℝ³ (J_g is 3×2) and f: ℝ³ → ℝ³ (J_f is 3×3), the product J_g·J_f is (3×2)(3×3) — inner dims 2 and 3 don't match. It is not even a legal multiply.
The outer function's Jacobian comes first (on the left), mirroring f'(g(x))·g'(x).
J_{f∘g} = J_f(g(x)) · J_g(x) ✓
Why: J_f (3×3) times J_g (3×2) = (3×2). Inner dims match, shape is right, and it agrees with the finite difference. Reading order (g then f) is NOT multiplication order — backprop walks from the output inward, outer Jacobian first.
Section
Part 4 of 6 — the running network
Intuition
You could compute J_f·J_g·J_h·… left-to-right (forward mode) or right-to-left (reverse mode). Same product, different cost. The order that wins depends on the shapes.
In learning, the very last node is the scalar loss L. Starting from L means the running quantity is always a row vector (a gradient), and each multiply by a layer Jacobian keeps it small. One backward sweep hands back ∂L/∂(everything) at once.
Forward mode would need one sweep per input parameter — hopeless for a net with millions of weights. Reverse mode (backprop) is the whole reason deep learning is affordable.
Concept
One hidden layer with sigmoid, a linear output, squared-error loss — the homework's exact setup. We carry this one example through the entire lesson.
\[ \begin{aligned} z_1 &= W_1 x + b_1 \\ h &= \sigma(z_1) \\ \hat y &= W_2 h + b_2 \\ L &= (\hat y - y)^2 \end{aligned} \]
Concrete values, fixed for the whole deck: x = [1, 2], target y = 1, W₁ is 3×2, b₁ is length 3, W₂ is 1×3, b₂ is length 1.
Concept
Small, hand-checkable numbers so every product is verifiable on paper:
\[ W_1 = \begin{bmatrix} 0.1 & 0.2 \\ 0.3 & 0.4 \\ 0.5 & 0.6 \end{bmatrix}, \quad b_1 = \begin{bmatrix} 0.1 \\ 0.2 \\ 0.3 \end{bmatrix}, \quad W_2 = \begin{bmatrix} 0.2 & 0.3 & 0.4 \end{bmatrix}, \quad b_2 = \begin{bmatrix} 0.5 \end{bmatrix} \]
The computational graph is x → z₁ → h → ŷ → L. Forward fills these in; backward walks the same arrows in reverse. Keep this picture in view for the next dozen slides.
Hypothesis
Predict first
Forward pass — the linear hidden layer z₁ is about to be worked. State your hypothesis first: which rule or definition decides this one, and what is the first move it forces? Then watch whether the example agrees with you.
Correct: z₁[0] = 0.1·1 + 0.2·2 + 0.1 = 0.1 + 0.4 + 0.1 = 0.6
Why: Row 0 of W₁ is [0.1, 0.2]; dot with [1,2] then add b₁[0] = 0.1.
A hypothesis you wrote down is falsifiable; a vague sense of how it will go is not. If the example opens somewhere else, that gap is the thing worth chasing.
Worked example
z₁ = W₁x + b₁. Each z₁[i] is row i of W₁ dotted with x = [1, 2], plus b₁[i]. Do all three:
z₁[0] = 0.1·1 + 0.2·2 + 0.1 = 0.1 + 0.4 + 0.1 = 0.6
Why: Row 0 of W₁ is [0.1, 0.2]; dot with [1,2] then add b₁[0] = 0.1.
z₁[1] = 0.3·1 + 0.4·2 + 0.2 = 0.3 + 0.8 + 0.2 = 1.3
Why: Row 1 is [0.3, 0.4]; add b₁[1] = 0.2.
z₁[2] = 0.5·1 + 0.6·2 + 0.3 = 0.5 + 1.2 + 0.3 = 2.0
Why: Row 2 is [0.5, 0.6]; add b₁[2] = 0.3.
import numpy as np
x = np.array([1., 2.]); y = np.array([1.])
W1 = np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1 = np.array([0.1,0.2,0.3])
W2 = np.array([[0.2,0.3,0.4]]); b2 = np.array([0.5])
sig = lambda z: 1/(1+np.exp(-z))
z1 = W1@x + b1
print(z1) # [0.6 1.3 2. ]| unit i | row of W₁ · x | + b₁[i] | z₁[i] |
|---|---|---|---|
| 0 | 0.1 + 0.4 = 0.5 | + 0.1 | 0.6 |
| 1 | 0.3 + 0.8 = 1.1 | + 0.2 | 1.3 |
| 2 | 0.5 + 1.2 = 1.7 | + 0.3 | 2.0 |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
z₁[2] = 0.5·1 + 0.6·2 + 0.3 = 0.5 + 1.2 + 0.3 = 2.0
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
z₁ = W₁x + b₁. Each z₁[i] is row i of W₁ dotted with x = [1, 2], plus b₁[i]. Do all three:
Pattern
Predict first
The table runs: 0 | 0.6 | 0.645656 · 1 | 1.3 | 0.785835
In Forward pass — squash through the sigmoid, given the rows so far: what is the next one — the row where unit i is 2?
Correct: 2 | 2.0 | 0.880797
| unit i | z₁[i] | h[i] = σ(z₁[i]) |
|---|---|---|
| 0 | 0.6 | 0.645656 |
| 1 | 1.3 | 0.785835 |
| 2 | 2.0 | 0.880797 |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. e^{−0.6} ≈ 0.5488, so 1/(1.5488) ≈ 0.645656.
Worked example
Apply h = σ(z₁) elementwise, where σ(z) = 1/(1 + e^{−z}). Each hidden unit is squeezed into (0, 1):
h[0] = σ(0.6) = 1/(1 + e^{−0.6}) = 0.645656
Why: e^{−0.6} ≈ 0.5488, so 1/(1.5488) ≈ 0.645656. Above 0.5 because z₁[0] > 0.
h[1] = σ(1.3) = 0.785835, h[2] = σ(2.0) = 0.880797
Why: Larger pre-activations push closer to 1. All three verified by execution.
import numpy as np
x = np.array([1.,2.])
W1 = np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1 = np.array([0.1,0.2,0.3])
sig = lambda z: 1/(1+np.exp(-z))
z1 = W1@x + b1
h = sig(z1)
print(h) # [0.645656 0.785835 0.880797]| unit i | z₁[i] | h[i] = σ(z₁[i]) |
|---|---|---|
| 0 | 0.6 | 0.645656 |
| 1 | 1.3 | 0.785835 |
| 2 | 2.0 | 0.880797 |
Pattern
Step through it
Step through Forward pass — squash through the sigmoid one row at a time. What is driving the change, and what would the row after the last one be?
Ranking
Put in order
Put the moves of Forward pass — output ŷ and loss L into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Dot the 1×3 W₂ with the length-3 h.
Worked example
Linear output ŷ = W₂h + b₂ (a single number), then squared-error loss L = (ŷ − y)² against target y = 1.
W₂h = 0.2·0.645656 + 0.3·0.785835 + 0.4·0.880797 = 0.717201
Why: Dot the 1×3 W₂ with the length-3 h. This is the model's raw signal before the bias.
ŷ = 0.717201 + b₂ = 0.717201 + 0.5 = 1.217201
Why: Add the output bias. The prediction overshoots the target of 1.
L = (1.217201 − 1)² = (0.217201)² = 0.047176
Why: Squared error. This single scalar is the top of the graph — the quantity every gradient is 'with respect to'.
import numpy as np
x = np.array([1.,2.]); y = np.array([1.])
W1 = np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1 = np.array([0.1,0.2,0.3])
W2 = np.array([[0.2,0.3,0.4]]); b2 = np.array([0.5])
sig = lambda z: 1/(1+np.exp(-z))
h = sig(W1@x + b1)
yhat = W2@h + b2
L = float(((yhat - y)**2).sum())
print(yhat, round(L, 6)) # [1.217201] 0.047176| quantity | value |
|---|---|
| W₂h | 0.717201 |
| ŷ = W₂h + b₂ | 1.217201 |
| ŷ − y | 0.217201 |
| L = (ŷ − y)² | 0.047176 |
Pattern
Step through it
Step through Forward pass — output ŷ and loss L one row at a time. What is driving the change, and what would the row after the last one be?
Concept
Forward we went x → z₁ → h → ŷ → L. Backward we compute a gradient at each node, in the reverse order L → ŷ → h → z₁, multiplying the local Jacobian at each arrow.
\[ \frac{\partial L}{\partial \hat y} \to \frac{\partial L}{\partial W_2},\, \frac{\partial L}{\partial b_2},\, \frac{\partial L}{\partial h} \to \frac{\partial L}{\partial z_1} \to \frac{\partial L}{\partial W_1},\, \frac{\partial L}{\partial b_1} \]
At each layer we peel off the parameter gradients (what we want) and pass an upstream gradient to the next layer down. Let's take it one arrow at a time.
Fill the middle
Fill in the blanks
From Backward step 1 — seed ∂L/∂ŷ — one line has had its right-hand side removed. Put it back.
import numpy as np
x = np.array([1.,2.]); y = np.array([1.])
W1 = np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1 = np.array([0.1,0.2,0.3])
W2 = np.array([[0.2,0.3,0.4]]); b2 = np.array([0.5])
sig = lambda z: 1/(1+np.exp(-z))
h = sig(W1@x + b1); yhat = W2@h + b2
dyhat = 2*(yhat - y)
print(dyhat) # [0.434401]
Why: h is what everything below it consumes, so the wrong expression here fails later and somewhere else. Chain rule on (ŷ − y)²: outer 2(·) times inner ∂(ŷ−y)/∂ŷ = 1.
Worked example
L = (ŷ − y)². Differentiate with respect to ŷ — the outer square times the inner derivative (which is 1):
∂L/∂ŷ = 2(ŷ − y) = 2·0.217201 = 0.434401
Why: Chain rule on (ŷ − y)²: outer 2(·) times inner ∂(ŷ−y)/∂ŷ = 1. This scalar is the seed the whole backward pass flows from.
import numpy as np
x = np.array([1.,2.]); y = np.array([1.])
W1 = np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1 = np.array([0.1,0.2,0.3])
W2 = np.array([[0.2,0.3,0.4]]); b2 = np.array([0.5])
sig = lambda z: 1/(1+np.exp(-z))
h = sig(W1@x + b1); yhat = W2@h + b2
dyhat = 2*(yhat - y)
print(dyhat) # [0.434401]| node | gradient ∂L/∂(·) |
|---|---|
| L | 1 (by definition) |
| ŷ | 2(ŷ − y) = 0.434401 |
Step zero
Discussion prompt
Backward step 2 — gradients of the output layer — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: ∂L/∂W₂ = ∂L/∂ŷ · hᵀ = outer(0.434401, h)
Answer:
Worked example
The output arrow is ŷ = W₂h + b₂. Its local partials: ∂ŷ/∂W₂ = h (as a row), ∂ŷ/∂b₂ = 1, ∂ŷ/∂h = W₂. Multiply each by the upstream ∂L/∂ŷ.
∂L/∂W₂ = ∂L/∂ŷ · hᵀ = outer(0.434401, h)
Why: A weight's gradient is (upstream signal) × (the input it multiplied). W₂ multiplied h, so its gradient is the outer product of the upstream 0.434401 with h.
= [0.280474, 0.341368, 0.382619]
Why: 0.434401·[0.645656, 0.785835, 0.880797], entrywise. These are the three ∂L/∂W₂ components.
∂L/∂b₂ = ∂L/∂ŷ · 1 = 0.434401
Why: The bias multiplied a constant 1, so its gradient is just the upstream signal, unchanged.
import numpy as np
x = np.array([1.,2.]); y = np.array([1.])
W1 = np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1 = np.array([0.1,0.2,0.3])
W2 = np.array([[0.2,0.3,0.4]]); b2 = np.array([0.5])
sig = lambda z: 1/(1+np.exp(-z))
h = sig(W1@x + b1); yhat = W2@h + b2; dyhat = 2*(yhat - y)
dW2 = np.outer(dyhat, h); db2 = dyhat
print(dW2) # [[0.280474 0.341368 0.382619]]
print(db2) # [0.434401]| gradient | formula | value |
|---|---|---|
| ∂L/∂W₂ | outer(∂L/∂ŷ, h) | [0.280474, 0.341368, 0.382619] |
| ∂L/∂b₂ | ∂L/∂ŷ | 0.434401 |
Blank canvas
Draw it
Draw what Backward step 2 — gradients of the output layer just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Fill the middle
Fill in the blanks
From Backward step 3 — push through W₂ to reach h — one line has had its right-hand side removed. Put it back.
import numpy as np
x = np.array([1.,2.]); y = np.array([1.])
W1 = np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1 = np.array([0.1,0.2,0.3])
W2 = np.array([[0.2,0.3,0.4]]); b2 = np.array([0.5])
sig = lambda z: 1/(1+np.exp(-z))
h = sig(W1@x + b1); yhat = W2@h + b2; dyhat = 2*(yhat - y)
dh = W2.T @ dyhat
print(dh.ravel()) # [0.08688 0.13032 0.17376]
Why: dh is what everything below it consumes, so the wrong expression here fails later and somewhere else. Crossing a LINEAR layer backward multiplies by the transpose of its weight matrix — the Jacobian form of the chain rule for y = Wh.
Worked example
To keep going back we need ∂L/∂h. The arrow ŷ = W₂h + b₂ has Jacobian ∂ŷ/∂h = W₂, so multiply the upstream by W₂ᵀ:
∂L/∂h = W₂ᵀ · ∂L/∂ŷ = W₂ᵀ · 0.434401
Why: Crossing a LINEAR layer backward multiplies by the transpose of its weight matrix — the Jacobian form of the chain rule for y = Wh.
= 0.434401 · [0.2, 0.3, 0.4] = [0.08688, 0.13032, 0.17376]
Why: W₂ is the single row [0.2, 0.3, 0.4]; each hidden unit's share of the blame is its weight times the upstream signal.
import numpy as np
x = np.array([1.,2.]); y = np.array([1.])
W1 = np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1 = np.array([0.1,0.2,0.3])
W2 = np.array([[0.2,0.3,0.4]]); b2 = np.array([0.5])
sig = lambda z: 1/(1+np.exp(-z))
h = sig(W1@x + b1); yhat = W2@h + b2; dyhat = 2*(yhat - y)
dh = W2.T @ dyhat
print(dh.ravel()) # [0.08688 0.13032 0.17376]| hidden unit i | W₂[i] | ∂L/∂h[i] = W₂[i]·0.434401 |
|---|---|---|
| 0 | 0.2 | 0.08688 |
| 1 | 0.3 | 0.13032 |
| 2 | 0.4 | 0.17376 |
Blank canvas
Draw it
Draw what Backward step 3 — push through W₂ to reach h just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Concept
The next arrow back is the activation h = σ(z₁). Its derivative has a famously clean form in terms of the output h we already cached:
\[ \sigma'(z) = \sigma(z)\big(1 - \sigma(z)\big) = h(1-h) \]
Because the activation is elementwise, its Jacobian is diagonal, so crossing it backward is an elementwise multiply by h(1−h) — no matrix needed.
Step zero
Discussion prompt
Backward step 4 — cross the sigmoid to z₁ — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: h(1−h) = [0.228784, 0.168298, 0.104994]
Answer:
Worked example
Multiply the upstream ∂L/∂h elementwise by σ'(z₁) = h(1−h). Compute the activation slopes first:
h(1−h) = [0.228784, 0.168298, 0.104994]
Why: 0.645656·0.354344, 0.785835·0.214165, 0.880797·0.119203. Note the largest slope is on the unit nearest z = 0, all below the 0.25 ceiling.
∂L/∂z₁ = ∂L/∂h ⊙ h(1−h)
Why: Elementwise product ⊙, because the sigmoid Jacobian is diagonal. This is the single most-forgotten factor in a hand-rolled backward pass.
= [0.08688·0.228784, 0.13032·0.168298, 0.17376·0.104994] = [0.019877, 0.021933, 0.018244]
Why: Each hidden unit's gradient is shrunk by its own activation slope. Verified against autograd next.
import numpy as np
x = np.array([1.,2.]); y = np.array([1.])
W1 = np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1 = np.array([0.1,0.2,0.3])
W2 = np.array([[0.2,0.3,0.4]]); b2 = np.array([0.5])
sig = lambda z: 1/(1+np.exp(-z))
h = sig(W1@x + b1); yhat = W2@h + b2
dh = W2.T @ (2*(yhat - y))
dz1 = dh * (h*(1-h))
print(dz1) # [0.019877 0.021933 0.018244]| unit i | ∂L/∂h[i] | h(1−h) | ∂L/∂z₁[i] |
|---|---|---|---|
| 0 | 0.08688 | 0.228784 | 0.019877 |
| 1 | 0.13032 | 0.168298 | 0.021933 |
| 2 | 0.17376 | 0.104994 | 0.018244 |
Error analysis
Annotate
Walk the callouts on Backward step 4 — cross the sigmoid to z₁. Each one is a place this is easy to get subtly wrong.
Ranking
Put in order
Put the moves of Backward step 5 — gradients of the first layer into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. W₁ multiplied x, so its gradient is the outer product of the upstream ∂L/∂z₁ (length 3) with x = [1, 2] (length 2) — a 3×2 matrix, matching W₁'s shape.
Worked example
The first arrow is z₁ = W₁x + b₁, the same linear pattern as the output layer. Weight gradient = outer(upstream, input); bias gradient = upstream.
∂L/∂W₁ = outer(∂L/∂z₁, x)
Why: W₁ multiplied x, so its gradient is the outer product of the upstream ∂L/∂z₁ (length 3) with x = [1, 2] (length 2) — a 3×2 matrix, matching W₁'s shape.
Row i = ∂L/∂z₁[i]·[1, 2], e.g. row 0 = [0.019877, 0.039754]
Why: Column 0 uses x₀ = 1, column 1 uses x₁ = 2, so column 1 is exactly double column 0.
∂L/∂b₁ = ∂L/∂z₁ = [0.019877, 0.021933, 0.018244]
Why: Bias multiplied 1s, so its gradient equals the upstream, unchanged.
import numpy as np
x = np.array([1.,2.]); y = np.array([1.])
W1 = np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1 = np.array([0.1,0.2,0.3])
W2 = np.array([[0.2,0.3,0.4]]); b2 = np.array([0.5])
sig = lambda z: 1/(1+np.exp(-z))
h = sig(W1@x + b1); yhat = W2@h + b2
dz1 = (W2.T @ (2*(yhat - y))) * (h*(1-h))
dW1 = np.outer(dz1, x); db1 = dz1
print(dW1)
# [[0.019877 0.039754]
# [0.021933 0.043865]
# [0.018244 0.036487]]| row i | ∂L/∂z₁[i] | · x₀=1 | · x₁=2 |
|---|---|---|---|
| 0 | 0.019877 | 0.019877 | 0.039754 |
| 1 | 0.021933 | 0.021933 | 0.043865 |
| 2 | 0.018244 | 0.018244 | 0.036487 |
Comparison
Comparison matrix
From Backward step 5 — gradients of the first layer: refill the ∂L/∂z₁[i] column from what you know. The rest of the table is as it appeared.
| row i | ∂L/∂z₁[i] | · x₀=1 | · x₁=2 |
|---|---|---|---|
| 0 | 0.019877 | 0.019877 | 0.039754 |
| 1 | 0.021933 | 0.021933 | 0.043865 |
| 2 | 0.018244 | 0.018244 | 0.036487 |
Estimation
Predict first
Five arrows, five moves. Here is the complete manual backward pass — nothing hidden between the beats above and this code:
Commit before you compute: what does The whole backward pass in one block come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: The pattern repeats: linear layer → outer(upstream, input); activation → ⊙ its derivative
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Every layer is one of two moves.
Worked example
Five arrows, five moves. Here is the complete manual backward pass — nothing hidden between the beats above and this code:
import numpy as np
x = np.array([1.,2.]); y = np.array([1.])
W1 = np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1 = np.array([0.1,0.2,0.3])
W2 = np.array([[0.2,0.3,0.4]]); b2 = np.array([0.5])
sig = lambda z: 1/(1+np.exp(-z))
h = sig(W1@x + b1); yhat = W2@h + b2 # forward (cache h)
dyhat = 2*(yhat - y) # seed -> 0.434401
dW2 = np.outer(dyhat, h) # output weight grad
db2 = dyhat # output bias grad
dh = W2.T @ dyhat # push back through W2
dz1 = dh * (h*(1-h)) # cross sigmoid (elementwise)
dW1 = np.outer(dz1, x) # hidden weight grad
db1 = dz1 # hidden bias grad
print(dW1[0], dW2[0])The pattern repeats: linear layer → outer(upstream, input); activation → ⊙ its derivative
Why: Every layer is one of two moves. The dz1 and dW1 lines are the ONLY places the sigmoid factor and the outer product for W₁ appear — miss either and the gradient is wrong.
| gradient | value |
|---|---|
| dW1[0] | [0.019877, 0.039754] |
| dW2[0] | [0.280474, 0.341368, 0.382619] |
| db1 | [0.019877, 0.021933, 0.018244] |
| db2 | [0.434401] |
Fill the middle
Fill in the blanks
From Manual backward = torch.autograd — one line has had its right-hand side removed. Put it back.
import numpy as np, torch
x = np.array([1.,2.]); y = np.array([1.])
W1 = np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1 = np.array([0.1,0.2,0.3])
W2 = np.array([[0.2,0.3,0.4]]); b2 = np.array([0.5])
sig = lambda z: 1/(1+np.exp(-z))
h = sig(W1@x + b1); yhat = W2@h + b2 # forward
dyhat = 2*(yhat - y); dW2 = np.outer(dyhat, h) # manual backward
dz1 = (W2.T@dyhat)(h(1-h)); dW1 = np.outer(dz1, x)
W1t=torch.tensor(W1,requires_grad=True); b1t=torch.tensor(b1,requires_grad=True)
W2t=torch.tensor(W2,requires_grad=True); b2t=torch.tensor(b2,requires_grad=True)
h_t=torch.sigmoid(W1t@torch.tensor(x)+b1t)
L_t=((W2t@h_t+b2t-torch.tensor(y))**2).sum()
L_t.backward()
print(np.allclose(dW1, W1t.grad.numpy()),
np.allclose(dW2, W2t.grad.numpy()))
Why: h is what everything below it consumes, so the wrong expression here fails later and somewhere else. PyTorch runs the exact same reverse-mode chain rule we did by hand.
Worked example
The proof our hand derivation is correct: rebuild the same forward pass in PyTorch with requires_grad=True, call .backward(), and compare every gradient.
import numpy as np, torch
x = np.array([1.,2.]); y = np.array([1.])
W1 = np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1 = np.array([0.1,0.2,0.3])
W2 = np.array([[0.2,0.3,0.4]]); b2 = np.array([0.5])
sig = lambda z: 1/(1+np.exp(-z))
h = sig(W1@x + b1); yhat = W2@h + b2 # forward
dyhat = 2*(yhat - y); dW2 = np.outer(dyhat, h) # manual backward
dz1 = (W2.T@dyhat)*(h*(1-h)); dW1 = np.outer(dz1, x)
W1t=torch.tensor(W1,requires_grad=True); b1t=torch.tensor(b1,requires_grad=True)
W2t=torch.tensor(W2,requires_grad=True); b2t=torch.tensor(b2,requires_grad=True)
h_t=torch.sigmoid(W1t@torch.tensor(x)+b1t)
L_t=((W2t@h_t+b2t-torch.tensor(y))**2).sum()
L_t.backward()
print(np.allclose(dW1, W1t.grad.numpy()),
np.allclose(dW2, W2t.grad.numpy()))allclose(dW1, autograd) → True; allclose(dW2, autograd) → True
Why: PyTorch runs the exact same reverse-mode chain rule we did by hand. Agreement to machine precision confirms every factor and sign is right.
| gradient | manual | torch.autograd |
|---|---|---|
| dL/dŷ (=db2) | 0.434401 | 0.434401 |
| dW2 | [0.280474, 0.341368, 0.382619] | [0.280474, 0.341368, 0.382619] |
| dW1[0] | [0.019877, 0.039754] | [0.019877, 0.039754] |
| db1 | [0.019877, 0.021933, 0.018244] | [0.019877, 0.021933, 0.018244] |
Anomaly
Predict first
A student writes this, and it looks reasonable:
Gradient flows from h to z₁ unchanged — just pass ∂L/∂h straight through the sigmoid.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Skips the sigmoid's own Jacobian.
Crossing the sigmoid multiplies elementwise by its local slope σ'(z₁) = h(1−h).
Why: Skips the sigmoid's own Jacobian. Then dW1[0] comes out [0.08688, 0.17376] instead of [0.019877, 0.039754] — about 4.4× too large, because the σ'(z₁[0]) = 0.228784 factor was never applied.
Trap
Gradient flows from h to z₁ unchanged — just pass ∂L/∂h straight through the sigmoid.
dz1 = dh = [0.08688, 0.13032, 0.17376]
Why: Skips the sigmoid's own Jacobian. Then dW1[0] comes out [0.08688, 0.17376] instead of [0.019877, 0.039754] — about 4.4× too large, because the σ'(z₁[0]) = 0.228784 factor was never applied.
Crossing the sigmoid multiplies elementwise by its local slope σ'(z₁) = h(1−h).
dz1 = dh ⊙ h(1−h) = [0.019877, 0.021933, 0.018244]
Why: Now dW1[0] = [0.019877, 0.039754], matching autograd. This shrink factor (each ≤ 0.25 for sigmoid) is exactly what will make deep sigmoid nets stall — never omit it.
Break the constraint
Discussion prompt
The rule this trap just fixed:
Crossing the sigmoid multiplies elementwise by its local slope σ'(z₁) = h(1−h).
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
Skips the sigmoid's own Jacobian. Then dW1[0] comes out [0.08688, 0.17376] instead of [0.019877, 0.039754] — about 4.4× too large, because the σ'(z₁[0]) = 0.228784 factor was never applied.
Section
Part 5 of 6 — the safety net
Concept
The same central difference from Part 1 checks a network gradient: perturb one parameter by ±ε, recompute the loss, and compare the slope to your analytic value.
\[ \frac{\partial L}{\partial \theta} \approx \frac{L(\theta+\epsilon) - L(\theta-\epsilon)}{2\epsilon} \]
This is your correctness oracle: if analytic and numeric disagree, the backward pass has a sign or factor bug. It's slow (one forward pass per parameter) so you spot-check a few entries, not train with it.
Pattern
Predict first
The table runs: central difference (ε = 1e−5) | 0.019877 · analytic backprop | 0.019877
In Gradient check on W₁[0,0], given the rows so far: what is the next one — the row where method is also checked: ∂L/∂W₂[0,0]?
Correct: also checked: ∂L/∂W₂[0,0] | 0.280474 = 0.280474
| method | ∂L/∂W₁[0,0] |
|---|---|
| central difference (ε = 1e−5) | 0.019877 |
| analytic backprop | 0.019877 |
| also checked: ∂L/∂W₂[0,0] | 0.280474 = 0.280474 |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. Agreement to ~6 digits means no sign or factor bug on this path.
Worked example
Check the single entry ∂L/∂W₁[0,0], which our backprop says is 0.019877. Bump W₁[0,0] by ±ε = 1e−5 and read the loss slope:
import numpy as np
x = np.array([1.,2.]); y = np.array([1.])
W1 = np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1 = np.array([0.1,0.2,0.3])
W2 = np.array([[0.2,0.3,0.4]]); b2 = np.array([0.5])
sig = lambda z: 1/(1+np.exp(-z))
h = sig(W1@x + b1); yhat = W2@h + b2 # forward
dz1 = (W2.T@(2*(yhat-y)))*(h*(1-h)); dW1 = np.outer(dz1, x) # analytic
eps = 1e-5
def loss_with(W1mod):
hh = sig(W1mod @ x + b1)
return float(((W2 @ hh + b2 - y)**2).sum())
Wp = W1.copy(); Wp[0,0] += eps
Wm = W1.copy(); Wm[0,0] -= eps
num = (loss_with(Wp) - loss_with(Wm)) / (2*eps)
print(round(num,6), round(dW1[0,0],6))numeric = 0.019877, analytic = 0.019877 — match
Why: Agreement to ~6 digits means no sign or factor bug on this path. ε too small underflows in floating point; too large loses accuracy — 1e−5 is the sweet spot for a central difference.
| method | ∂L/∂W₁[0,0] |
|---|---|
| central difference (ε = 1e−5) | 0.019877 |
| analytic backprop | 0.019877 |
| also checked: ∂L/∂W₂[0,0] | 0.280474 = 0.280474 |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
numeric = 0.019877, analytic = 0.019877 — match
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Check the single entry ∂L/∂W₁[0,0], which our backprop says is 0.019877. Bump W₁[0,0] by ±ε = 1e−5 and read the loss slope:
Section
Part 6 of 6 — why depth hurts
Concept
The factor h(1−h) we multiplied by is σ'(z), and it is largest at the center and tiny in the tails:
\[ \sigma'(z) = \sigma(z)\big(1-\sigma(z)\big), \qquad 0 < \sigma'(z) \le \tfrac14 \]
The maximum is 0.25, hit at z = 0 where σ = 0.5. So every sigmoid layer multiplies the upstream gradient by at most a quarter — usually much less.
Missing information
Discussion prompt
Tabulate σ'(z) = σ(z)(1−σ(z)) across z to see the 0.25 ceiling and the tail collapse:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Saturated units (large |z|) have almost no slope — their gradient nearly vanishes even in a single layer. Our hidden units at z = 0.6, 1.3, 2.0 already gave 0.229, 0.168, 0.105.
Worked example
Tabulate σ'(z) = σ(z)(1−σ(z)) across z to see the 0.25 ceiling and the tail collapse:
import numpy as np
sig = lambda z: 1/(1+np.exp(-z))
for z in [-4,-2,-1,0,1,2,4]:
s = sig(z)
print(z, round(s,4), round(s*(1-s),4))σ'(0) = 0.5·0.5 = 0.25 is the peak; σ'(±4) ≈ 0.0177 in the tails
Why: Saturated units (large |z|) have almost no slope — their gradient nearly vanishes even in a single layer. Our hidden units at z = 0.6, 1.3, 2.0 already gave 0.229, 0.168, 0.105.
| z | σ(z) | σ'(z) = σ(1−σ) |
|---|---|---|
| −4 | 0.0180 | 0.0177 |
| −1 | 0.2689 | 0.1966 |
| 0 | 0.5000 | 0.2500 |
| 1 | 0.7311 | 0.1966 |
| 4 | 0.9820 | 0.0177 |
Pattern
Step through it
Step through Where the sigmoid slope peaks one row at a time. What is driving the change, and what would the row after the last one be?
Intuition
Backprop through d layers multiplies d local slopes. If each slope is at most 0.25, the gradient reaching the front is at most (0.25)^d of the signal at the top.
Multiplying numbers below 1 shrinks them fast: 0.25² = 0.0625, 0.25⁵ ≈ 0.001. By the time you reach an early layer of a deep sigmoid net, there is essentially no signal left — the vanishing gradient.
Explain it
Discussion prompt
Explain Sub-1 factors compound toward zero to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Backprop through d layers multiplies d local slopes. If each slope is at most 0.25, the gradient reaching the front is at most (0.25)^d of the signal at the top.
Estimation
Predict first
Compare the best-case sigmoid factor (0.25)^depth to ReLU's active-path factor 1^depth. Runnable:
Commit before you compute: what does The vanishing gradient, quantified come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: After 10 sigmoid layers the gradient is ≤ (0.25)¹⁰ ≈ 9.54e−7 — effectively zero
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. And this is the BEST case (every unit at z = 0).
Worked example
Compare the best-case sigmoid factor (0.25)^depth to ReLU's active-path factor 1^depth. Runnable:
for d in [1, 2, 5, 10, 20]:
print(d, format(0.25**d, '.3e'), 1.0**d) # sigmoid best-case vs ReLU activeAfter 10 sigmoid layers the gradient is ≤ (0.25)¹⁰ ≈ 9.54e−7 — effectively zero
Why: And this is the BEST case (every unit at z = 0). Real saturated units make it far worse. ReLU's active slope of 1 doesn't shrink at all: 1¹⁰ = 1.
| depth | sigmoid ≤ (0.25)ᵈ | ReLU (active) 1ᵈ |
|---|---|---|
| 1 | 2.500e−01 | 1 |
| 2 | 6.250e−02 | 1 |
| 5 | 9.766e−04 | 1 |
| 10 | 9.537e−07 | 1 |
| 20 | 9.095e−13 | 1 |
Invariant
Step through it
Step through The vanishing gradient, quantified one row at a time. One of these columns never changes — find it, and say why it cannot.
Concept
ReLU's derivative is 1 wherever the unit is active and 0 where it's off — no shrinkage on the active path, so gradients pass through many layers intact.
\[ \operatorname{ReLU}'(z) = \begin{cases} 1 & z > 0 \\ 0 & z < 0 \end{cases} \]
The cost is the dying-ReLU problem: a unit stuck at z < 0 gets zero gradient forever and never recovers. Most deep nets happily take that trade — and add residual connections and normalization for extra insurance.
Analogy
Discussion prompt
Explain ReLU: a gradient that survives depth by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
ReLU's derivative is 1 wherever the unit is active and 0 where it's off — no shrinkage on the active path, so gradients pass through many layers intact.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Stacking many sigmoid layers gives the network more capacity, so the early layers should train at least as fast as a shallow one.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Ignores that gradients MULTIPLY through layers.
Each layer multiplies the gradient by its local slope, so sub-1 factors compound toward zero.
Why: Ignores that gradients MULTIPLY through layers. With per-layer factors ≤ 0.25, depth makes the front-layer signal vanish, not grow — capacity is irrelevant if no gradient arrives.
Trap
Stacking many sigmoid layers gives the network more capacity, so the early layers should train at least as fast as a shallow one.
Assume the gradient magnitude is roughly constant with depth
Why: Ignores that gradients MULTIPLY through layers. With per-layer factors ≤ 0.25, depth makes the front-layer signal vanish, not grow — capacity is irrelevant if no gradient arrives.
Each layer multiplies the gradient by its local slope, so sub-1 factors compound toward zero.
Front-layer gradient ∝ (factor)^depth → (0.25)^depth for sigmoid
Why: That is why deep nets use ReLU (slope 1 when active), residual connections (add an identity path so the factor is ≥ 1), and normalization — all to keep the gradient product from collapsing.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
y = f(g(x)) is two machines wired in series: x goes into g to make an intermediate u = g(x), and u goes into f to make y = f(u).; A tiny bump Δx becomes Δu ≈ g'(x)·Δx after the first machine, and that becomes Δy ≈ f'(u)·Δu after the second. Substitute the first into the second.; Every analytic derivative can be sanity-checked numerically. A central finite difference nudges the input by ±ε and measures the output change:g first, then f, so write the Jacobians in that left-to-right reading order: J_g · J_f.; Gradient flows from h to z₁ unchanged — just pass ∂L/∂h straight through the sigmoid.Ranking
Put in order
These are the steps of The backprop recipe, scrambled. Put them back in order before the next slide shows you.
z₁, h, ŷ) — the backward pass needs them∂L/∂ŷ from the loss (here 2(ŷ − y))outer(upstream, input); bias grad = upstream; pass Wᵀ·upstream downh(1−h), ReLU → 1 if active else 0)Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
z₁, h, ŷ) — the backward pass needs them∂L/∂ŷ from the loss (here 2(ŷ − y))outer(upstream, input); bias grad = upstream; pass Wᵀ·upstream downh(1−h), ReLU → 1 if active else 0)Edge cases
Discussion prompt
The backprop recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
z₁, h, ŷ) — the backward pass needs them∂L/∂ŷ from the loss (here 2(ŷ − y))outer(upstream, input); bias grad = upstream; pass Wᵀ·upstream downh(1−h), ReLU → 1 if active else 0)Elimination
Eliminate the wrong options
For y = f(g(x)) with vector x, the Jacobian J of the composition is:
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: The chain rule is multiplicative with the OUTER function's Jacobian first: J_f evaluated at g(x), times J_g at x. Dimensions (m×k)(k×n) = (m×n) confirm the order.
Check
Match the matrix order to the derivative — and check the shapes.
Check your understanding
For y = f(g(x)) with vector x, the Jacobian J of the composition is:
Answer: A
Why: The chain rule is multiplicative with the OUTER function's Jacobian first: J_f evaluated at g(x), times J_g at x. Dimensions (m×k)(k×n) = (m×n) confirm the order.
Prediction
Predict first
Crossing a sigmoid layer in the backward pass, you convert ∂L/∂h into ∂L/∂z by multiplying (elementwise) by:
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: h(1 − h), the sigmoid derivative
Why: σ'(z) = σ(z)(1−σ(z)) = h(1−h), applied elementwise because the activation's Jacobian is diagonal. In our net this turned ∂L/∂h = [0.08688, 0.13032, 0.17376] into ∂L/∂z₁ = [0.019877, 0.021933, 0.018244].
Check
Trace one step of the backward pass across the sigmoid.
Check your understanding
Crossing a sigmoid layer in the backward pass, you convert ∂L/∂h into ∂L/∂z by multiplying (elementwise) by:
Answer: A
Why: σ'(z) = σ(z)(1−σ(z)) = h(1−h), applied elementwise because the activation's Jacobian is diagonal. In our net this turned ∂L/∂h = [0.08688, 0.13032, 0.17376] into ∂L/∂z₁ = [0.019877, 0.021933, 0.018244].
Elimination
Eliminate the wrong options
A 10-layer network of sigmoids trains very slowly in its early layers. The best explanation is:
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Every sigmoid contributes a factor σ' ≤ 0.25 to the gradient. Over 10 layers the product is at most (0.25)¹⁰ ≈ 9.5e−7, so the front layers get almost no signal — the vanishing-gradient problem.
Check
Think about what compounds as depth grows.
Check your understanding
A 10-layer network of sigmoids trains very slowly in its early layers. The best explanation is:
Answer: A
Why: Every sigmoid contributes a factor σ' ≤ 0.25 to the gradient. Over 10 layers the product is at most (0.25)¹⁰ ≈ 9.5e−7, so the front layers get almost no signal — the vanishing-gradient problem.
Section
The project — a 2-layer MLP by hand
Concept
Implement forward AND backward for our running network in pure NumPy, then prove every gradient matches torch.autograd and a finite difference. You've derived every piece — now assemble it yourself.
| # | requirement | tool |
|---|---|---|
| 1 | Forward: z₁, h, ŷ, L | @ and the sigmoid |
| 2 | Backward: dW1, db1, dW2, db2 | outer products + h(1−h) |
| 3 | Verify vs autograd and a central difference | torch + finite diff |
Build rules: cache the forward intermediates (you need h in the backward pass), type every line yourself, run after each line, and when something errors read the array shapes before you change anything.
Counterexample
Discussion prompt
Implement forward AND backward for our running network in pure NumPy, then prove every gradient matches torch.autograd and a finite difference. You've derived every piece — now assemble it yourself.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Worked example
Your turn: compute z₁, h, ŷ, L from the given weights. Predict out loud whether ŷ lands above or below the target of 1 before you print.
Hint: z1 = W1@x + b1, h = sig(z1), yhat = W2@h + b2, L = ((yhat-y)**2).sum(). Cache h — you'll need it going backward.
import numpy as np
x = np.array([1.,2.]); y = np.array([1.])
W1 = np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1 = np.array([0.1,0.2,0.3])
W2 = np.array([[0.2,0.3,0.4]]); b2 = np.array([0.5])
sig = lambda z: 1/(1+np.exp(-z))
z1 = W1@x + b1; h = sig(z1)
yhat = W2@h + b2; L = float(((yhat - y)**2).sum())
print(h, yhat, round(L,5))| value | result |
|---|---|
| z₁ | [0.6, 1.3, 2.0] |
| h | [0.645656, 0.785835, 0.880797] |
| ŷ | 1.217201 (above target 1) |
| L | 0.04718 |
Worked example
Your turn: seed ∂L/∂ŷ, then walk back to dW1. The one thing you must not drop is the h(1−h) factor when you cross the sigmoid.
Hint: dyhat = 2*(yhat-y); dW2 = np.outer(dyhat, h); dh = W2.T@dyhat; dz1 = dh*(h*(1-h)); dW1 = np.outer(dz1, x).
import numpy as np
x = np.array([1.,2.]); y = np.array([1.])
W1 = np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1 = np.array([0.1,0.2,0.3])
W2 = np.array([[0.2,0.3,0.4]]); b2 = np.array([0.5])
sig = lambda z: 1/(1+np.exp(-z))
h = sig(W1@x + b1); yhat = W2@h + b2 # forward pass from Milestone 1
# --- your backward pass ---
dyhat = 2*(yhat - y)
dW2 = np.outer(dyhat, h); db2 = dyhat
dh = W2.T @ dyhat
dz1 = dh * (h*(1-h))
dW1 = np.outer(dz1, x); db1 = dz1
print(dW1)| gradient | value |
|---|---|
| ∂L/∂ŷ | 0.434401 |
| dW2 | [0.280474, 0.341368, 0.382619] |
| dz1 | [0.019877, 0.021933, 0.018244] |
| dW1[0] | [0.019877, 0.039754] |
Worked example
Your turn: rebuild the forward pass in torch with requires_grad=True, call .backward(), and compare. Then spot-check one entry with a central difference. Predict: exact match or just close?
Hint: wrap W1,b1,W2,b2 as tensors with requires_grad=True, recompute L with torch.sigmoid, call L.backward(), and compare W1t.grad.numpy() to dW1 with np.allclose.
import numpy as np, torch
x = np.array([1.,2.]); y = np.array([1.])
W1 = np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1 = np.array([0.1,0.2,0.3])
W2 = np.array([[0.2,0.3,0.4]]); b2 = np.array([0.5])
sig = lambda z: 1/(1+np.exp(-z))
h = sig(W1@x + b1); yhat = W2@h + b2 # forward
dz1 = (W2.T@(2*(yhat-y)))*(h*(1-h)); dW1 = np.outer(dz1, x) # your dW1
W1t=torch.tensor(W1,requires_grad=True); b1t=torch.tensor(b1,requires_grad=True)
W2t=torch.tensor(W2,requires_grad=True); b2t=torch.tensor(b2,requires_grad=True)
h_t=torch.sigmoid(W1t@torch.tensor(x)+b1t)
L_t=((W2t@h_t+b2t-torch.tensor(y))**2).sum(); L_t.backward()
print(np.allclose(dW1, W1t.grad.numpy()))| compare | result |
|---|---|
| allclose(dW1, autograd) | True |
| allclose(dW2, autograd) | True |
| numeric vs analytic dW1[0,0] | 0.019877 = 0.019877 |
Trade off
Comparison matrix
From Milestone 3 — verify against autograd + finite diff: every row here is a choice with a cost. Fill the result column, then say which row you would actually pick and what you give up for it.
| compare | result |
|---|---|
| allclose(dW1, autograd) | True |
| allclose(dW2, autograd) | True |
| numeric vs analytic dW1[0,0] | 0.019877 = 0.019877 |
Concept
import numpy as np
x=np.array([1.,2.]); y=np.array([1.])
W1=np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1=np.array([0.1,0.2,0.3])
W2=np.array([[0.2,0.3,0.4]]); b2=np.array([0.5])
sig=lambda z:1/(1+np.exp(-z))
# forward (cache h for the backward pass)
h=sig(W1@x+b1); yhat=W2@h+b2; L=float(((yhat-y)**2).sum())
# backward
dyhat=2*(yhat-y); dW2=np.outer(dyhat,h); db2=dyhat
dz1=(W2.T@dyhat)*(h*(1-h)); dW1=np.outer(dz1,x); db1=dz1
print('L =',round(L,5),' dW1[0,0] =',round(dW1[0,0],6))| output | value |
|---|---|
| L | 0.04718 |
| dW1[0,0] | 0.019877 |
If your L reads 0.04718, dW1[0,0] reads 0.019877, and allclose against autograd returns True — you implemented backpropagation by hand.
Comparison
Comparison matrix
From The full program: refill the value column from what you know. The rest of the table is as it appeared.
| output | value |
|---|---|
| L | 0.04718 |
| dW1[0,0] | 0.019877 |
Concept
Out loud, slides closed: explain (1) why Jacobians multiply outer-first, (2) what the h(1−h) factor does to a gradient over 10 layers, and (3) why a central finite difference catches a sign bug your eyes miss.
Stretch (homework): prove the softmax Jacobian ∂sᵢ/∂xⱼ = sᵢ(δᵢⱼ − sⱼ), then swap MSE for softmax + cross-entropy and re-derive the famously clean ∂L/∂z = p − y. Confirm it against autograd the same way.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — The single-variable chain rule · The multivariable chain rule · The Jacobian form · Backprop as reverse-mode autodiff · Gradient checking · Activation gradients & vanishing. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
z₁ → h → ŷ → L) and all the way back, every number by handtorch.autograd to machine precision and gradient-check any entry with a central finite difference(0.25)^depth for sigmoid and explain why ReLU's slope-1 active path survives depth| move | the one thing to remember |
|---|---|
| chain rule | multiply local Jacobians, OUTER first |
| many paths | sum the contributions (fan-out adds gradients) |
| linear layer back | weight grad = outer(upstream, input); pass Wᵀ·upstream |
| activation back | multiply elementwise by its derivative — sigmoid → h(1−h) |
| depth | sub-1 factors compound → sigmoid vanishes; ReLU survives |
| verify | central difference, ε ≈ 1e−5 |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.