Lesson 9: The Chain Rule & Backpropagation

USAAIO Lesson 9, from Week 3 on calculus, fully worked. It builds the single-variable chain rule from local rates, then the multivariable sum-over-paths rule, then derives the Jacobian form and pins down the order in which the multiplications happen. From there it presents backpropagation as reverse-mode autodiff through a concrete two-layer sigmoid network: every forward and backward number is pushed through by hand, one beat at a time - z1, h, y-hat, L, then dy-hat, dW2, dh, sigma', dz1, dW1 - and checked against torch.autograd to machine precision. It quantifies vanishing gradients as (0.25)^depth, covers finite-difference gradient checking, and ends with a your-turn build of the whole backward pass. Every snippet runs as written, and every number came from real execution. The lesson runs to 64 slides.

Subject: Machine Learning · 116 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. The Chain Rule & Backpropagation

Title

USAAIO · Lesson 9 · Week 3 (Calculus)

How a network learns: multiply local derivatives backward through the graph. We derive the chain rule three ways — scalar, sum-over-paths, and Jacobian — then push one concrete 2-layer net all the way through by hand and match torch.autograd to the last digit.

2. By the end of this lesson you can

Objectives

  1. Apply the single-variable chain rule as a product of local rates, and check it against a finite difference
  2. Extend to the multivariable rule (sum over paths) and package it as a Jacobian matrix product — getting the multiply order right
  3. See backprop as reverse-mode autodiff: seed ∂L/∂L = 1 and sweep Jacobians from the loss back to every parameter
  4. Run a full manual backward pass through a 2-layer sigmoid net — every number — and match torch.autograd to machine precision
  5. Quantify the vanishing-gradient problem as (0.25)^depth, explain why ReLU survives, and gradient-check any analytic derivative

3. What survived from Maximum Likelihood Estimation?

Warm-up

Discussion prompt

Before we open Lesson 9: The Chain Rule & Backpropagation: without looking back, what was the main idea of Maximum Likelihood Estimation, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

MLE as an optimization principle, the log-likelihood trick, MLE for Gaussian/Bernoulli, and the big idea that MSE and cross-entropy ARE negative log-likelihoods — plus MAP as regularized MLE. Build maximum_likelihood_gaussian from scratch.

4. The single-variable chain rule

Section

Part 1 of 6 — local rates

5. A composition is a chain of machines

Concept

A composition y = f(g(x)) is two machines wired in series: x goes into g to make an intermediate u = g(x), and u goes into f to make y = f(u).

\[ x \;\xrightarrow{\;g\;}\; u = g(x) \;\xrightarrow{\;f\;}\; y = f(u) \]

To ask how y responds to a nudge in x, we ask two smaller questions: how does u respond to x, and how does y respond to u? Then we combine them.

6. Break it if you can: A composition is a chain of machines

Counterexample

Discussion prompt

A composition y = f(g(x)) is two machines wired in series: x goes into g to make an intermediate u = g(x), and u goes into f to make y = f(u).

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

To ask how y responds to a nudge in x, we ask two smaller questions: how does u respond to x, and how does y respond to u? Then we combine them.

7. Rates multiply, like gears

Intuition

Think of two meshed gears. If the first turns 3× as fast as its input and the second turns 2× as fast as the first, the last spins 6× the input — the rates multiply, they don't add.

A tiny bump Δx becomes Δu ≈ g'(x)·Δx after the first machine, and that becomes Δy ≈ f'(u)·Δu after the second. Substitute the first into the second.

So Δy ≈ f'(u)·g'(x)·Δx. Divide by Δx: the overall rate is the product of the two local rates. That single idea drives every gradient in this lesson.

8. By analogy: Rates multiply, like gears

Analogy

Discussion prompt

Explain Rates multiply, like gears by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Think of two meshed gears. If the first turns 3× as fast as its input and the second turns 2× as fast as the first, the last spins 6× the input — the rates multiply, they don't add.

9. The single-variable chain rule

Concept

Written cleanly, with u = g(x):

\[ \frac{dy}{dx} = \frac{dy}{du}\cdot\frac{du}{dx} = f'(g(x))\cdot g'(x) \]

The Leibniz form dy/dx = (dy/du)(du/dx) makes it look like the dus cancel. They don't literally cancel, but the mnemonic is exactly right: outer derivative (at the inner value) times inner derivative.

10. What has to happen first: Chain rule on y = sin(x²)

Ranking

Put in order

Put the moves of Chain rule on y = sin(x²) into the order they have to happen.

  1. Inner: u = x² = 1.3² = 1.69, and du/dx = 2x = 2.6
  2. Outer: dy/du = cos(u) = cos(1.69) = −0.118922
  3. Multiply: dy/dx = cos(1.69)·2.6 = −0.118922·2.6 = −0.309196

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. The inner machine squares its input; its local rate is 2x, which is 2.6 at x = 1.3.

11. Chain rule on y = sin(x²)

Worked example

Take y = sin(x²), evaluated at x = 1.3. Split it into inner u = g(x) = x² and outer f(u) = sin(u).

Inner: u = x² = 1.3² = 1.69, and du/dx = 2x = 2.6

Why: The inner machine squares its input; its local rate is 2x, which is 2.6 at x = 1.3.

Outer: dy/du = cos(u) = cos(1.69) = −0.118922

Why: The outer machine is sine; its local rate is cosine, evaluated at the INNER value u = 1.69, not at x.

Multiply: dy/dx = cos(1.69)·2.6 = −0.118922·2.6 = −0.309196

Why: Product of the two local rates. Watch the sign flip — the outer slope is negative here.

pieceexpressionvalue at x = 1.3
inner ux²1.69
du/dx2x2.6
dy/ducos(u)−0.118922
dy/dxcos(u)·2x−0.309196

12. Fill in: expression for Chain rule on y = sin(x²)

Comparison

Comparison matrix

From Chain rule on y = sin(x²): refill the expression column from what you know. The rest of the table is as it appeared.

pieceexpressionvalue at x = 1.3
inner ux²1.69
du/dx2x2.6
dy/ducos(u)−0.118922
dy/dxcos(u)·2x−0.309196

13. Trust, but verify: the finite difference

Concept

Every analytic derivative can be sanity-checked numerically. A central finite difference nudges the input by ±ε and measures the output change:

\[ \frac{dy}{dx} \approx \frac{y(x+\epsilon) - y(x-\epsilon)}{2\epsilon} \]

Central (both sides) is far more accurate than one-sided: its error shrinks like ε², not ε. We use this all lesson as the ground truth check.

14. Teach it back: Trust, but verify: the finite difference

Explain it

Discussion prompt

Explain Trust, but verify: the finite difference to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Every analytic derivative can be sanity-checked numerically. A central finite difference nudges the input by ±ε and measures the output change:

15. Guess the shape of the answer: Numeric check of the sin(x²) derivative

Estimation

Predict first

Confirm −0.309196 by nudging x = 1.3 by ε = 1e−6. Runnable as written:

Commit before you compute: what does Numeric check of the sin(x²) derivative come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: analytic = −0.309196, numeric = −0.309196 — they match

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Six-digit agreement means the split into inner/outer and the product were correct.

16. Numeric check of the sin(x²) derivative

Worked example

Confirm −0.309196 by nudging x = 1.3 by ε = 1e−6. Runnable as written:

import numpy as np
x0 = 1.3
analytic = np.cos(x0**2) * 2*x0        # cos(u)*2x
eps = 1e-6
num = (np.sin((x0+eps)**2) - np.sin((x0-eps)**2)) / (2*eps)
print(round(analytic, 6), round(num, 6))

analytic = −0.309196, numeric = −0.309196 — they match

Why: Six-digit agreement means the split into inner/outer and the product were correct. If these disagreed, one of the two local rates would be wrong.

methoddy/dx at x = 1.3
analytic (chain rule)−0.309196
central difference (ε = 1e−6)−0.309196
difference< 1e−9

17. What each one costs: Numeric check of the sin(x²) derivative

Trade off

Comparison matrix

From Numeric check of the sin(x²) derivative: every row here is a choice with a cost. Fill the dy/dx at x = 1.3 column, then say which row you would actually pick and what you give up for it.

methoddy/dx at x = 1.3
analytic (chain rule)−0.309196
central difference (ε = 1e−6)−0.309196
difference< 1e−9

18. A three-link chain: forward values, then backward rates

Concept

Longer chains are the same move, repeated. Take a = 2, then u = a², then v = sin(u). The chain rule multiplies every local rate along the path from v back to a.

\[ \frac{dv}{da} = \frac{dv}{du}\cdot\frac{du}{da} = \cos(u)\cdot 2a \]

Notice the two-phase structure that IS backprop: first sweep forward to get the intermediate values (u = 4), then sweep backward multiplying rates (cos(4)·2·2).

19. Plan first: Forward then backward on v = sin(a²)

Step zero

Discussion prompt

Forward then backward on v = sin(a²) — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Forward: u = a² = 4.0, then v = sin(4.0) = −0.756802

Answer:

  1. Forward: u = a² = 4.0, then v = sin(4.0) = −0.756802
  2. Backward: dv/du = cos(4.0) = −0.653644, du/da = 2·2 = 4.0
  3. Multiply along the path: dv/da = −0.653644·4.0 = −2.614574

20. Forward then backward on v = sin(a²)

Worked example

Push a = 2 forward to values, then multiply rates backward. Every number verified:

import numpy as np
a = 2.0
u = a**2          # forward: intermediate
v = np.sin(u)     # forward: output
dvdu = np.cos(u)  # backward: local rate of outer
duda = 2*a        # backward: local rate of inner
dvda = dvdu * duda
print(u, round(v,6), round(dvdu,6), duda, round(dvda,6))

Forward: u = a² = 4.0, then v = sin(4.0) = −0.756802

Why: Cache u — the backward pass needs it to evaluate cos(u). This caching is the seed of the whole backprop idea.

Backward: dv/du = cos(4.0) = −0.653644, du/da = 2·2 = 4.0

Why: Each machine reports its own local rate at the value flowing through it.

Multiply along the path: dv/da = −0.653644·4.0 = −2.614574

Why: Verified against a central difference: numeric dv/da = −2.614574 as well.

stagevalue / rate
u = a²4.0
v = sin(u)−0.756802
dv/du = cos(u)−0.653644
du/da = 2a4.0
dv/da (product)−2.614574

21. The multivariable chain rule

Section

Part 2 of 6 — sum over paths

22. When the input feeds several intermediates

Concept

Now let one input t flow through two intermediates and then combine: u = u(t), w = w(t), and L = L(u, w). A nudge in t reaches L along two different paths.

Figure (svg): A node t branches to two nodes u and w, which both feed into a node L. Two paths t to u to L and t to w to L.

t influences L through u AND through w — two paths to add up.

23. Sum over paths

Concept

The multivariable chain rule says: multiply the rates along each path, then add the paths. Each route contributes its own product of local rates.

\[ \frac{dL}{dt} = \frac{\partial L}{\partial u}\frac{du}{dt} + \frac{\partial L}{\partial w}\frac{dw}{dt} \]

The single-variable rule was the one-path special case. The + is what shows up whenever a quantity is used in more than one place — and in a network, a shared weight or activation is used in many places.

24. Why add — total influence accumulates

Intuition

If t gets a little bigger, u shifts and pushes L one way, AND w shifts and pushes L another way. Both pushes happen at once, so the total change in L is the sum of the two contributions.

In backprop this becomes: a node that fans out to several children collects a gradient from each child and adds them. 'Sum the incoming gradients' is this rule wearing work clothes.

25. Plan first: Two-path example, checked numerically

Step zero

Discussion prompt

Two-path example, checked numerically — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Path through u: (∂L/∂u)(du/dt) = w · 2t = (6)(4) = 24

Answer:

  1. Path through u: (∂L/∂u)(du/dt) = w · 2t = (6)(4) = 24
  2. Path through w: (∂L/∂w)(dw/dt) = u · 3 = (4)(3) = 12
  3. Add the paths: dL/dt = 24 + 12 = 36

26. Two-path example, checked numerically

Worked example

Let u = t², w = 3t, and L = u·w. Then L(t) = t²·3t = 3t³, so we already know dL/dt = 9t². Let's get the same thing by summing paths at t = 2.

Path through u: (∂L/∂u)(du/dt) = w · 2t = (6)(4) = 24

Why: ∂L/∂u = w = 3t = 6 at t = 2; du/dt = 2t = 4. Their product is this path's contribution.

Path through w: (∂L/∂w)(dw/dt) = u · 3 = (4)(3) = 12

Why: ∂L/∂w = u = t² = 4; dw/dt = 3. Product is 12.

Add the paths: dL/dt = 24 + 12 = 36

Why: And the closed form 9t² = 9·4 = 36 agrees exactly. The sum-over-paths rule reproduces the direct answer.

pathrate productvalue at t = 2
through uw · 2t24
through wu · 312
total dL/dtsum36
check: 9t²direct36

27. Which is which, by value at t = 2

Discrimination

Sort into buckets

Sort these by value at t = 2, from memory, without looking back at Two-path example, checked numerically. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

24
through u
12
through w
36
total dL/dt; check: 9t²
g1
value at t = 2 is "24" for through u — that is what the table on "Two-path example, checked numerically" records, and it is the single property separating this group from the rest.
g2
value at t = 2 is "12" for through w — that is what the table on "Two-path example, checked numerically" records, and it is the single property separating this group from the rest.
g3
value at t = 2 is "36" for total dL/dt, check: 9t² — that is what the table on "Two-path example, checked numerically" records, and it is the single property separating this group from the rest.

28. The Jacobian form

Section

Part 3 of 6 — vectors & matrices

29. The Jacobian: all partials in one matrix

Concept

For a vector map f: ℝⁿ → ℝᵐ, the Jacobian J collects every partial derivative: row i, column j is ∂fᵢ/∂xⱼ. It is m × n — outputs down the rows, inputs across the columns.

\[ J_f = \begin{bmatrix} \dfrac{\partial f_1}{\partial x_1} & \cdots & \dfrac{\partial f_1}{\partial x_n} \\[4pt] \vdots & \ddots & \vdots \\[4pt] \dfrac{\partial f_m}{\partial x_1} & \cdots & \dfrac{\partial f_m}{\partial x_n} \end{bmatrix} \]

It is the best linear approximation of f at a point: f(x + Δ) ≈ f(x) + J_f Δ. The scalar f'(x) was the 1×1 Jacobian all along.

30. The chain rule becomes matrix multiplication

Concept

Stack the sum-over-paths rule for every output-input pair and it packs into one clean matrix product — the Jacobian of a composition is the product of the Jacobians:

\[ J_{f\circ g}(x) = J_f\big(g(x)\big)\, J_g(x) \]

Entry (i,k) of that product is Σⱼ (∂fᵢ/∂uⱼ)(∂uⱼ/∂xₖ) — exactly 'sum the rates over every intermediate path.' The matrix multiply IS the summation, done in bulk.

31. Two building-block Jacobians we'll reuse

Concept

Every layer in our net is one of two maps. Their Jacobians are worth memorizing:

\[ g(x) = Wx + b \;\Longrightarrow\; J_g = W \]

\[ \phi(z) = \sigma(z)\ \text{(elementwise)} \;\Longrightarrow\; J_\phi = \operatorname{diag}\big(\sigma'(z)\big) \]

A linear layer's Jacobian is just its weight matrix W. An elementwise activation's Jacobian is diagonal — output i depends only on input i, so off-diagonal partials are zero.

32. Where does each piece belong: Lesson 9: The Chain Rule & Backpropagation

Sorting

Sort into buckets

These are the pieces of Lesson 9: The Chain Rule & Backpropagation, out of order. Put each one back under the part of the lesson it belongs to.

The single-variable chain rule
A composition is a chain of machines; Rates multiply, like gears; The single-variable chain rule
The multivariable chain rule
When the input feeds several intermediates; Sum over paths; Why add — total influence accumulates
The Jacobian form
The Jacobian: all partials in one matrix; The chain rule becomes matrix multiplication; Two building-block Jacobians we'll reuse
s1
The single-variable chain rule is where Lesson 9: The Chain Rule & Backpropagation puts A composition is a chain of machines, Rates multiply, like gears, The single-variable chain rule. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
The multivariable chain rule is where Lesson 9: The Chain Rule & Backpropagation puts When the input feeds several intermediates, Sum over paths, Why add — total influence accumulates. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
The Jacobian form is where Lesson 9: The Chain Rule & Backpropagation puts The Jacobian: all partials in one matrix, The chain rule becomes matrix multiplication, Two building-block Jacobians we'll reuse. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

33. What has to be given first: Multiply two layer Jacobians, verified

Missing information

Discussion prompt

Compose the two: g(x) = W₁x + b₁ then φ(z) = σ(z), so h = σ(W₁x + b₁). With our real W₁ and the σ'(z₁) = h(1−h) from the forward pass, J = diag(σ') · W₁. Runnable:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

diag(σ') is 3×3, W₁ is 3×2; inner dimensions (3 and 3) match only in this order, giving the correct 3×2 Jacobian dh/dx.

34. Multiply two layer Jacobians, verified

Worked example

Compose the two: g(x) = W₁x + b₁ then φ(z) = σ(z), so h = σ(W₁x + b₁). With our real W₁ and the σ'(z₁) = h(1−h) from the forward pass, J = diag(σ') · W₁. Runnable:

import numpy as np
x = np.array([1., 2.])
W1 = np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1 = np.array([0.1,0.2,0.3])
sig = lambda z: 1/(1+np.exp(-z))
z1 = W1@x + b1; h = sig(z1)
Jf = np.diag(h*(1-h))          # 3x3 diagonal
Jcomp = Jf @ W1                # (3x3)(3x2) = 3x2
print(Jcomp[:,0])              # dh/dx0

Shapes: (3×3)·(3×2) = (3×2) — outer Jacobian on the LEFT

Why: diag(σ') is 3×3, W₁ is 3×2; inner dimensions (3 and 3) match only in this order, giving the correct 3×2 Jacobian dh/dx.

Column 0 = dh/dx₀ = [0.022878, 0.05049, 0.052497]

Why: A central finite difference on x₀ gives the identical [0.022878, 0.05049, 0.052497]. The Jacobian product is exactly the numerical sensitivity.

sourcedh/dx₀
Jf @ W1, column 0[0.022878, 0.05049, 0.052497]
central difference on x₀[0.022878, 0.05049, 0.052497]
match?yes, to 1e−6

35. Inspect it line by line: Multiply two layer Jacobians, verified

Error analysis

Annotate

Walk the callouts on Multiply two layer Jacobians, verified. Each one is a place this is easy to get subtly wrong.

  • diag(σ') is 3×3, W₁ is 3×2; inner dimensions (3 and 3) match only in this order, giving the correct 3×2 Jacobian dh/dx.
  • A central finite difference on x₀ gives the identical [0.022878, 0.05049, 0.052497]. The Jacobian product is exactly the numerical sensitivity.

36. Rebuild the recipe: The chain-rule pattern, all three forms

Ranking

Put in order

These are the steps of The chain-rule pattern, all three forms, scrambled. Put them back in order before the next slide shows you.

  1. Scalar: one path → multiply local rates: dy/dx = f'(g(x))·g'(x)
  2. Multivariable: many paths → multiply along each, then add: dL/dt = Σ (∂L/∂uⱼ)(duⱼ/dt)
  3. Vector (Jacobian): the summation packed as a matrix product: J_{f∘g} = J_f · J_g, outer Jacobian first
  4. Always: sweep FORWARD to cache intermediate values, then sweep BACKWARD multiplying rates
  5. Verify any single entry with a central finite difference, ε ≈ 1e−5 to 1e−6

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

37. The chain-rule pattern, all three forms

Pattern

  1. Scalar: one path → multiply local rates: dy/dx = f'(g(x))·g'(x)
  2. Multivariable: many paths → multiply along each, then add: dL/dt = Σ (∂L/∂uⱼ)(duⱼ/dt)
  3. Vector (Jacobian): the summation packed as a matrix product: J_{f∘g} = J_f · J_g, outer Jacobian first
  4. Always: sweep FORWARD to cache intermediate values, then sweep BACKWARD multiplying rates
  5. Verify any single entry with a central finite difference, ε ≈ 1e−5 to 1e−6

38. Something is wrong here: which order do Jacobians multiply?

Anomaly

Predict first

A student writes this, and it looks reasonable:

We apply g first, then f, so write the Jacobians in that left-to-right reading order: J_g · J_f.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: With g: ℝ² → ℝ³ (J_g is 3×2) and f: ℝ³ → ℝ³ (J_f is 3×3), the product J_g·J_f is (3×2)(3×3) — inner dims 2 and 3 don't match.

The outer function's Jacobian comes first (on the left), mirroring f'(g(x))·g'(x).

Why: With g: ℝ² → ℝ³ (J_g is 3×2) and f: ℝ³ → ℝ³ (J_f is 3×3), the product J_g·J_f is (3×2)(3×3) — inner dims 2 and 3 don't match. It is not even a legal multiply.

39. Trap: which order do Jacobians multiply?

Trap

The trap

We apply g first, then f, so write the Jacobians in that left-to-right reading order: J_g · J_f.

J_{f∘g} = J_g J_f ✗

Why: With g: ℝ² → ℝ³ (J_g is 3×2) and f: ℝ³ → ℝ³ (J_f is 3×3), the product J_g·J_f is (3×2)(3×3) — inner dims 2 and 3 don't match. It is not even a legal multiply.

The fix

The outer function's Jacobian comes first (on the left), mirroring f'(g(x))·g'(x).

J_{f∘g} = J_f(g(x)) · J_g(x) ✓

Why: J_f (3×3) times J_g (3×2) = (3×2). Inner dims match, shape is right, and it agrees with the finite difference. Reading order (g then f) is NOT multiplication order — backprop walks from the output inward, outer Jacobian first.

40. Backprop as reverse-mode autodiff

Section

Part 4 of 6 — the running network

41. Why sweep backward, not forward?

Intuition

You could compute J_f·J_g·J_h·… left-to-right (forward mode) or right-to-left (reverse mode). Same product, different cost. The order that wins depends on the shapes.

In learning, the very last node is the scalar loss L. Starting from L means the running quantity is always a row vector (a gradient), and each multiply by a layer Jacobian keeps it small. One backward sweep hands back ∂L/∂(everything) at once.

Forward mode would need one sweep per input parameter — hopeless for a net with millions of weights. Reverse mode (backprop) is the whole reason deep learning is affordable.

42. Our running network

Concept

One hidden layer with sigmoid, a linear output, squared-error loss — the homework's exact setup. We carry this one example through the entire lesson.

\[ \begin{aligned} z_1 &= W_1 x + b_1 \\ h &= \sigma(z_1) \\ \hat y &= W_2 h + b_2 \\ L &= (\hat y - y)^2 \end{aligned} \]

Concrete values, fixed for the whole deck: x = [1, 2], target y = 1, W₁ is 3×2, b₁ is length 3, W₂ is 1×3, b₂ is length 1.

43. The exact weights we'll use

Concept

Small, hand-checkable numbers so every product is verifiable on paper:

\[ W_1 = \begin{bmatrix} 0.1 & 0.2 \\ 0.3 & 0.4 \\ 0.5 & 0.6 \end{bmatrix}, \quad b_1 = \begin{bmatrix} 0.1 \\ 0.2 \\ 0.3 \end{bmatrix}, \quad W_2 = \begin{bmatrix} 0.2 & 0.3 & 0.4 \end{bmatrix}, \quad b_2 = \begin{bmatrix} 0.5 \end{bmatrix} \]

The computational graph is x → z₁ → h → ŷ → L. Forward fills these in; backward walks the same arrows in reverse. Keep this picture in view for the next dozen slides.

44. State the rule before it runs: Forward pass — the linear hidden layer z₁

Hypothesis

Predict first

Forward pass — the linear hidden layer z₁ is about to be worked. State your hypothesis first: which rule or definition decides this one, and what is the first move it forces? Then watch whether the example agrees with you.

Correct: z₁[0] = 0.1·1 + 0.2·2 + 0.1 = 0.1 + 0.4 + 0.1 = 0.6

Why: Row 0 of W₁ is [0.1, 0.2]; dot with [1,2] then add b₁[0] = 0.1.

A hypothesis you wrote down is falsifiable; a vague sense of how it will go is not. If the example opens somewhere else, that gap is the thing worth chasing.

45. Forward pass — the linear hidden layer z₁

Worked example

z₁ = W₁x + b₁. Each z₁[i] is row i of W₁ dotted with x = [1, 2], plus b₁[i]. Do all three:

z₁[0] = 0.1·1 + 0.2·2 + 0.1 = 0.1 + 0.4 + 0.1 = 0.6

Why: Row 0 of W₁ is [0.1, 0.2]; dot with [1,2] then add b₁[0] = 0.1.

z₁[1] = 0.3·1 + 0.4·2 + 0.2 = 0.3 + 0.8 + 0.2 = 1.3

Why: Row 1 is [0.3, 0.4]; add b₁[1] = 0.2.

z₁[2] = 0.5·1 + 0.6·2 + 0.3 = 0.5 + 1.2 + 0.3 = 2.0

Why: Row 2 is [0.5, 0.6]; add b₁[2] = 0.3.

import numpy as np
x = np.array([1., 2.]); y = np.array([1.])
W1 = np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1 = np.array([0.1,0.2,0.3])
W2 = np.array([[0.2,0.3,0.4]]); b2 = np.array([0.5])
sig = lambda z: 1/(1+np.exp(-z))
z1 = W1@x + b1
print(z1)   # [0.6 1.3 2. ]
unit irow of W₁ · x+ b₁[i]z₁[i]
00.1 + 0.4 = 0.5+ 0.10.6
10.3 + 0.8 = 1.1+ 0.21.3
20.5 + 1.2 = 1.7+ 0.32.0

46. Work backwards from the answer: Forward pass — the linear hidden layer z₁

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

z₁[2] = 0.5·1 + 0.6·2 + 0.3 = 0.5 + 1.2 + 0.3 = 2.0

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

z₁ = W₁x + b₁. Each z₁[i] is row i of W₁ dotted with x = [1, 2], plus b₁[i]. Do all three:

47. Predict the next row: Forward pass — squash through the sigmoid

Pattern

Predict first

The table runs: 0 | 0.6 | 0.645656 · 1 | 1.3 | 0.785835

In Forward pass — squash through the sigmoid, given the rows so far: what is the next one — the row where unit i is 2?

Correct: 2 | 2.0 | 0.880797

unit iz₁[i]h[i] = σ(z₁[i])
00.60.645656
11.30.785835
22.00.880797

Why: The relationship between the columns, not the individual numbers, is what generates the next row. e^{−0.6} ≈ 0.5488, so 1/(1.5488) ≈ 0.645656.

48. Forward pass — squash through the sigmoid

Worked example

Apply h = σ(z₁) elementwise, where σ(z) = 1/(1 + e^{−z}). Each hidden unit is squeezed into (0, 1):

h[0] = σ(0.6) = 1/(1 + e^{−0.6}) = 0.645656

Why: e^{−0.6} ≈ 0.5488, so 1/(1.5488) ≈ 0.645656. Above 0.5 because z₁[0] > 0.

h[1] = σ(1.3) = 0.785835, h[2] = σ(2.0) = 0.880797

Why: Larger pre-activations push closer to 1. All three verified by execution.

import numpy as np
x = np.array([1.,2.])
W1 = np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1 = np.array([0.1,0.2,0.3])
sig = lambda z: 1/(1+np.exp(-z))
z1 = W1@x + b1
h = sig(z1)
print(h)   # [0.645656 0.785835 0.880797]
unit iz₁[i]h[i] = σ(z₁[i])
00.60.645656
11.30.785835
22.00.880797

49. Watch it run: Forward pass — squash through the sigmoid

Pattern

Step through it

Step through Forward pass — squash through the sigmoid one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: unit i is 0
  2. Step 2: unit i is 1
  3. Step 3: unit i is 2

50. What has to happen first: Forward pass — output ŷ and loss L

Ranking

Put in order

Put the moves of Forward pass — output ŷ and loss L into the order they have to happen.

  1. W₂h = 0.2·0.645656 + 0.3·0.785835 + 0.4·0.880797 = 0.717201
  2. ŷ = 0.717201 + b₂ = 0.717201 + 0.5 = 1.217201
  3. L = (1.217201 − 1)² = (0.217201)² = 0.047176

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Dot the 1×3 W₂ with the length-3 h.

51. Forward pass — output ŷ and loss L

Worked example

Linear output ŷ = W₂h + b₂ (a single number), then squared-error loss L = (ŷ − y)² against target y = 1.

W₂h = 0.2·0.645656 + 0.3·0.785835 + 0.4·0.880797 = 0.717201

Why: Dot the 1×3 W₂ with the length-3 h. This is the model's raw signal before the bias.

ŷ = 0.717201 + b₂ = 0.717201 + 0.5 = 1.217201

Why: Add the output bias. The prediction overshoots the target of 1.

L = (1.217201 − 1)² = (0.217201)² = 0.047176

Why: Squared error. This single scalar is the top of the graph — the quantity every gradient is 'with respect to'.

import numpy as np
x = np.array([1.,2.]); y = np.array([1.])
W1 = np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1 = np.array([0.1,0.2,0.3])
W2 = np.array([[0.2,0.3,0.4]]); b2 = np.array([0.5])
sig = lambda z: 1/(1+np.exp(-z))
h = sig(W1@x + b1)
yhat = W2@h + b2
L = float(((yhat - y)**2).sum())
print(yhat, round(L, 6))   # [1.217201] 0.047176
quantityvalue
W₂h0.717201
ŷ = W₂h + b₂1.217201
ŷ − y0.217201
L = (ŷ − y)²0.047176

52. Watch it run: Forward pass — output ŷ and loss L

Pattern

Step through it

Step through Forward pass — output ŷ and loss L one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: quantity is W₂h
  2. Step 2: quantity is ŷ = W₂h + b₂
  3. Step 3: quantity is ŷ − y
  4. Step 4: quantity is L = (ŷ − y)²

53. The backward plan: reverse every arrow

Concept

Forward we went x → z₁ → h → ŷ → L. Backward we compute a gradient at each node, in the reverse order L → ŷ → h → z₁, multiplying the local Jacobian at each arrow.

\[ \frac{\partial L}{\partial \hat y} \to \frac{\partial L}{\partial W_2},\, \frac{\partial L}{\partial b_2},\, \frac{\partial L}{\partial h} \to \frac{\partial L}{\partial z_1} \to \frac{\partial L}{\partial W_1},\, \frac{\partial L}{\partial b_1} \]

At each layer we peel off the parameter gradients (what we want) and pass an upstream gradient to the next layer down. Let's take it one arrow at a time.

54. Restore the missing line: Backward step 1 — seed ∂L/∂ŷ

Fill the middle

Fill in the blanks

From Backward step 1 — seed ∂L/∂ŷ — one line has had its right-hand side removed. Put it back.

import numpy as np
x = np.array([1.,2.]); y = np.array([1.])
W1 = np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1 = np.array([0.1,0.2,0.3])
W2 = np.array([[0.2,0.3,0.4]]); b2 = np.array([0.5])
sig = lambda z: 1/(1+np.exp(-z))
h = sig(W1@x + b1); yhat = W2@h + b2
dyhat = 2*(yhat - y)
print(dyhat) # [0.434401]

Why: h is what everything below it consumes, so the wrong expression here fails later and somewhere else. Chain rule on (ŷ − y)²: outer 2(·) times inner ∂(ŷ−y)/∂ŷ = 1.

55. Backward step 1 — seed ∂L/∂ŷ

Worked example

L = (ŷ − y)². Differentiate with respect to ŷ — the outer square times the inner derivative (which is 1):

∂L/∂ŷ = 2(ŷ − y) = 2·0.217201 = 0.434401

Why: Chain rule on (ŷ − y)²: outer 2(·) times inner ∂(ŷ−y)/∂ŷ = 1. This scalar is the seed the whole backward pass flows from.

import numpy as np
x = np.array([1.,2.]); y = np.array([1.])
W1 = np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1 = np.array([0.1,0.2,0.3])
W2 = np.array([[0.2,0.3,0.4]]); b2 = np.array([0.5])
sig = lambda z: 1/(1+np.exp(-z))
h = sig(W1@x + b1); yhat = W2@h + b2
dyhat = 2*(yhat - y)
print(dyhat)   # [0.434401]
nodegradient ∂L/∂(·)
L1 (by definition)
ŷ2(ŷ − y) = 0.434401

56. Plan first: Backward step 2 — gradients of the output layer

Step zero

Discussion prompt

Backward step 2 — gradients of the output layer — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: ∂L/∂W₂ = ∂L/∂ŷ · hᵀ = outer(0.434401, h)

Answer:

  1. ∂L/∂W₂ = ∂L/∂ŷ · hᵀ = outer(0.434401, h)
  2. = [0.280474, 0.341368, 0.382619]
  3. ∂L/∂b₂ = ∂L/∂ŷ · 1 = 0.434401

57. Backward step 2 — gradients of the output layer

Worked example

The output arrow is ŷ = W₂h + b₂. Its local partials: ∂ŷ/∂W₂ = h (as a row), ∂ŷ/∂b₂ = 1, ∂ŷ/∂h = W₂. Multiply each by the upstream ∂L/∂ŷ.

∂L/∂W₂ = ∂L/∂ŷ · hᵀ = outer(0.434401, h)

Why: A weight's gradient is (upstream signal) × (the input it multiplied). W₂ multiplied h, so its gradient is the outer product of the upstream 0.434401 with h.

= [0.280474, 0.341368, 0.382619]

Why: 0.434401·[0.645656, 0.785835, 0.880797], entrywise. These are the three ∂L/∂W₂ components.

∂L/∂b₂ = ∂L/∂ŷ · 1 = 0.434401

Why: The bias multiplied a constant 1, so its gradient is just the upstream signal, unchanged.

import numpy as np
x = np.array([1.,2.]); y = np.array([1.])
W1 = np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1 = np.array([0.1,0.2,0.3])
W2 = np.array([[0.2,0.3,0.4]]); b2 = np.array([0.5])
sig = lambda z: 1/(1+np.exp(-z))
h = sig(W1@x + b1); yhat = W2@h + b2; dyhat = 2*(yhat - y)
dW2 = np.outer(dyhat, h); db2 = dyhat
print(dW2)   # [[0.280474 0.341368 0.382619]]
print(db2)   # [0.434401]
gradientformulavalue
∂L/∂W₂outer(∂L/∂ŷ, h)[0.280474, 0.341368, 0.382619]
∂L/∂b₂∂L/∂ŷ0.434401

58. Draw the shape of it: Backward step 2 — gradients of the output…

Blank canvas

Draw it

Draw what Backward step 2 — gradients of the output layer just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.

59. Restore the missing line: Backward step 3 — push through W₂ to reach h

Fill the middle

Fill in the blanks

From Backward step 3 — push through W₂ to reach h — one line has had its right-hand side removed. Put it back.

import numpy as np
x = np.array([1.,2.]); y = np.array([1.])
W1 = np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1 = np.array([0.1,0.2,0.3])
W2 = np.array([[0.2,0.3,0.4]]); b2 = np.array([0.5])
sig = lambda z: 1/(1+np.exp(-z))
h = sig(W1@x + b1); yhat = W2@h + b2; dyhat = 2*(yhat - y)
dh = W2.T @ dyhat
print(dh.ravel()) # [0.08688 0.13032 0.17376]

Why: dh is what everything below it consumes, so the wrong expression here fails later and somewhere else. Crossing a LINEAR layer backward multiplies by the transpose of its weight matrix — the Jacobian form of the chain rule for y = Wh.

60. Backward step 3 — push through W₂ to reach h

Worked example

To keep going back we need ∂L/∂h. The arrow ŷ = W₂h + b₂ has Jacobian ∂ŷ/∂h = W₂, so multiply the upstream by W₂ᵀ:

∂L/∂h = W₂ᵀ · ∂L/∂ŷ = W₂ᵀ · 0.434401

Why: Crossing a LINEAR layer backward multiplies by the transpose of its weight matrix — the Jacobian form of the chain rule for y = Wh.

= 0.434401 · [0.2, 0.3, 0.4] = [0.08688, 0.13032, 0.17376]

Why: W₂ is the single row [0.2, 0.3, 0.4]; each hidden unit's share of the blame is its weight times the upstream signal.

import numpy as np
x = np.array([1.,2.]); y = np.array([1.])
W1 = np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1 = np.array([0.1,0.2,0.3])
W2 = np.array([[0.2,0.3,0.4]]); b2 = np.array([0.5])
sig = lambda z: 1/(1+np.exp(-z))
h = sig(W1@x + b1); yhat = W2@h + b2; dyhat = 2*(yhat - y)
dh = W2.T @ dyhat
print(dh.ravel())   # [0.08688 0.13032 0.17376]
hidden unit iW₂[i]∂L/∂h[i] = W₂[i]·0.434401
00.20.08688
10.30.13032
20.40.17376

61. Draw the shape of it: Backward step 3 — push through W₂ to reach h

Blank canvas

Draw it

Draw what Backward step 3 — push through W₂ to reach h just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.

62. The sigmoid's local Jacobian

Concept

The next arrow back is the activation h = σ(z₁). Its derivative has a famously clean form in terms of the output h we already cached:

\[ \sigma'(z) = \sigma(z)\big(1 - \sigma(z)\big) = h(1-h) \]

Because the activation is elementwise, its Jacobian is diagonal, so crossing it backward is an elementwise multiply by h(1−h) — no matrix needed.

63. Plan first: Backward step 4 — cross the sigmoid to z₁

Step zero

Discussion prompt

Backward step 4 — cross the sigmoid to z₁ — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: h(1−h) = [0.228784, 0.168298, 0.104994]

Answer:

  1. h(1−h) = [0.228784, 0.168298, 0.104994]
  2. ∂L/∂z₁ = ∂L/∂h ⊙ h(1−h)
  3. = [0.08688·0.228784, 0.13032·0.168298, 0.17376·0.104994] = [0.019877, 0.021933, 0.018244]

64. Backward step 4 — cross the sigmoid to z₁

Worked example

Multiply the upstream ∂L/∂h elementwise by σ'(z₁) = h(1−h). Compute the activation slopes first:

h(1−h) = [0.228784, 0.168298, 0.104994]

Why: 0.645656·0.354344, 0.785835·0.214165, 0.880797·0.119203. Note the largest slope is on the unit nearest z = 0, all below the 0.25 ceiling.

∂L/∂z₁ = ∂L/∂h ⊙ h(1−h)

Why: Elementwise product ⊙, because the sigmoid Jacobian is diagonal. This is the single most-forgotten factor in a hand-rolled backward pass.

= [0.08688·0.228784, 0.13032·0.168298, 0.17376·0.104994] = [0.019877, 0.021933, 0.018244]

Why: Each hidden unit's gradient is shrunk by its own activation slope. Verified against autograd next.

import numpy as np
x = np.array([1.,2.]); y = np.array([1.])
W1 = np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1 = np.array([0.1,0.2,0.3])
W2 = np.array([[0.2,0.3,0.4]]); b2 = np.array([0.5])
sig = lambda z: 1/(1+np.exp(-z))
h = sig(W1@x + b1); yhat = W2@h + b2
dh = W2.T @ (2*(yhat - y))
dz1 = dh * (h*(1-h))
print(dz1)   # [0.019877 0.021933 0.018244]
unit i∂L/∂h[i]h(1−h)∂L/∂z₁[i]
00.086880.2287840.019877
10.130320.1682980.021933
20.173760.1049940.018244

65. Inspect it line by line: Backward step 4 — cross the sigmoid to z₁

Error analysis

Annotate

Walk the callouts on Backward step 4 — cross the sigmoid to z₁. Each one is a place this is easy to get subtly wrong.

  • 0.645656·0.354344, 0.785835·0.214165, 0.880797·0.119203. Note the largest slope is on the unit nearest z = 0, all below the 0.25 ceiling.
  • Elementwise product ⊙, because the sigmoid Jacobian is diagonal. This is the single most-forgotten factor in a hand-rolled backward pass.
  • Each hidden unit's gradient is shrunk by its own activation slope. Verified against autograd next.

66. What has to happen first: Backward step 5 — gradients of the first layer

Ranking

Put in order

Put the moves of Backward step 5 — gradients of the first layer into the order they have to happen.

  1. ∂L/∂W₁ = outer(∂L/∂z₁, x)
  2. Row i = ∂L/∂z₁[i]·[1, 2], e.g. row 0 = [0.019877, 0.039754]
  3. ∂L/∂b₁ = ∂L/∂z₁ = [0.019877, 0.021933, 0.018244]

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. W₁ multiplied x, so its gradient is the outer product of the upstream ∂L/∂z₁ (length 3) with x = [1, 2] (length 2) — a 3×2 matrix, matching W₁'s shape.

67. Backward step 5 — gradients of the first layer

Worked example

The first arrow is z₁ = W₁x + b₁, the same linear pattern as the output layer. Weight gradient = outer(upstream, input); bias gradient = upstream.

∂L/∂W₁ = outer(∂L/∂z₁, x)

Why: W₁ multiplied x, so its gradient is the outer product of the upstream ∂L/∂z₁ (length 3) with x = [1, 2] (length 2) — a 3×2 matrix, matching W₁'s shape.

Row i = ∂L/∂z₁[i]·[1, 2], e.g. row 0 = [0.019877, 0.039754]

Why: Column 0 uses x₀ = 1, column 1 uses x₁ = 2, so column 1 is exactly double column 0.

∂L/∂b₁ = ∂L/∂z₁ = [0.019877, 0.021933, 0.018244]

Why: Bias multiplied 1s, so its gradient equals the upstream, unchanged.

import numpy as np
x = np.array([1.,2.]); y = np.array([1.])
W1 = np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1 = np.array([0.1,0.2,0.3])
W2 = np.array([[0.2,0.3,0.4]]); b2 = np.array([0.5])
sig = lambda z: 1/(1+np.exp(-z))
h = sig(W1@x + b1); yhat = W2@h + b2
dz1 = (W2.T @ (2*(yhat - y))) * (h*(1-h))
dW1 = np.outer(dz1, x); db1 = dz1
print(dW1)
# [[0.019877 0.039754]
#  [0.021933 0.043865]
#  [0.018244 0.036487]]
row i∂L/∂z₁[i]· x₀=1· x₁=2
00.0198770.0198770.039754
10.0219330.0219330.043865
20.0182440.0182440.036487

68. Fill in: ∂L/∂z₁[i] for Backward step 5 — gradients of the first…

Comparison

Comparison matrix

From Backward step 5 — gradients of the first layer: refill the ∂L/∂z₁[i] column from what you know. The rest of the table is as it appeared.

row i∂L/∂z₁[i]· x₀=1· x₁=2
00.0198770.0198770.039754
10.0219330.0219330.043865
20.0182440.0182440.036487

69. Guess the shape of the answer: The whole backward pass in one block

Estimation

Predict first

Five arrows, five moves. Here is the complete manual backward pass — nothing hidden between the beats above and this code:

Commit before you compute: what does The whole backward pass in one block come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: The pattern repeats: linear layer → outer(upstream, input); activation → ⊙ its derivative

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Every layer is one of two moves.

70. The whole backward pass in one block

Worked example

Five arrows, five moves. Here is the complete manual backward pass — nothing hidden between the beats above and this code:

import numpy as np
x = np.array([1.,2.]); y = np.array([1.])
W1 = np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1 = np.array([0.1,0.2,0.3])
W2 = np.array([[0.2,0.3,0.4]]); b2 = np.array([0.5])
sig = lambda z: 1/(1+np.exp(-z))
h = sig(W1@x + b1); yhat = W2@h + b2   # forward (cache h)
dyhat = 2*(yhat - y)          # seed  -> 0.434401
dW2 = np.outer(dyhat, h)      # output weight grad
db2 = dyhat                   # output bias grad
dh  = W2.T @ dyhat            # push back through W2
dz1 = dh * (h*(1-h))          # cross sigmoid (elementwise)
dW1 = np.outer(dz1, x)        # hidden weight grad
db1 = dz1                     # hidden bias grad
print(dW1[0], dW2[0])

The pattern repeats: linear layer → outer(upstream, input); activation → ⊙ its derivative

Why: Every layer is one of two moves. The dz1 and dW1 lines are the ONLY places the sigmoid factor and the outer product for W₁ appear — miss either and the gradient is wrong.

gradientvalue
dW1[0][0.019877, 0.039754]
dW2[0][0.280474, 0.341368, 0.382619]
db1[0.019877, 0.021933, 0.018244]
db2[0.434401]

71. Restore the missing line: Manual backward = torch.autograd

Fill the middle

Fill in the blanks

From Manual backward = torch.autograd — one line has had its right-hand side removed. Put it back.

import numpy as np, torch
x = np.array([1.,2.]); y = np.array([1.])
W1 = np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1 = np.array([0.1,0.2,0.3])
W2 = np.array([[0.2,0.3,0.4]]); b2 = np.array([0.5])
sig = lambda z: 1/(1+np.exp(-z))
h = sig(W1@x + b1); yhat = W2@h + b2 # forward
dyhat = 2*(yhat - y); dW2 = np.outer(dyhat, h) # manual backward
dz1 = (W2.T@dyhat)(h(1-h)); dW1 = np.outer(dz1, x)
W1t=torch.tensor(W1,requires_grad=True); b1t=torch.tensor(b1,requires_grad=True)
W2t=torch.tensor(W2,requires_grad=True); b2t=torch.tensor(b2,requires_grad=True)
h_t=torch.sigmoid(W1t@torch.tensor(x)+b1t)
L_t=((W2t@h_t+b2t-torch.tensor(y))**2).sum()
L_t.backward()
print(np.allclose(dW1, W1t.grad.numpy()),
np.allclose(dW2, W2t.grad.numpy()))

Why: h is what everything below it consumes, so the wrong expression here fails later and somewhere else. PyTorch runs the exact same reverse-mode chain rule we did by hand.

72. Manual backward = torch.autograd

Worked example

The proof our hand derivation is correct: rebuild the same forward pass in PyTorch with requires_grad=True, call .backward(), and compare every gradient.

import numpy as np, torch
x = np.array([1.,2.]); y = np.array([1.])
W1 = np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1 = np.array([0.1,0.2,0.3])
W2 = np.array([[0.2,0.3,0.4]]); b2 = np.array([0.5])
sig = lambda z: 1/(1+np.exp(-z))
h = sig(W1@x + b1); yhat = W2@h + b2          # forward
dyhat = 2*(yhat - y); dW2 = np.outer(dyhat, h)  # manual backward
dz1 = (W2.T@dyhat)*(h*(1-h)); dW1 = np.outer(dz1, x)
W1t=torch.tensor(W1,requires_grad=True); b1t=torch.tensor(b1,requires_grad=True)
W2t=torch.tensor(W2,requires_grad=True); b2t=torch.tensor(b2,requires_grad=True)
h_t=torch.sigmoid(W1t@torch.tensor(x)+b1t)
L_t=((W2t@h_t+b2t-torch.tensor(y))**2).sum()
L_t.backward()
print(np.allclose(dW1, W1t.grad.numpy()),
      np.allclose(dW2, W2t.grad.numpy()))

allclose(dW1, autograd) → True; allclose(dW2, autograd) → True

Why: PyTorch runs the exact same reverse-mode chain rule we did by hand. Agreement to machine precision confirms every factor and sign is right.

gradientmanualtorch.autograd
dL/dŷ (=db2)0.4344010.434401
dW2[0.280474, 0.341368, 0.382619][0.280474, 0.341368, 0.382619]
dW1[0][0.019877, 0.039754][0.019877, 0.039754]
db1[0.019877, 0.021933, 0.018244][0.019877, 0.021933, 0.018244]

73. Something is wrong here: forgetting the activation derivative

Anomaly

Predict first

A student writes this, and it looks reasonable:

Gradient flows from h to z₁ unchanged — just pass ∂L/∂h straight through the sigmoid.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Skips the sigmoid's own Jacobian.

Crossing the sigmoid multiplies elementwise by its local slope σ'(z₁) = h(1−h).

Why: Skips the sigmoid's own Jacobian. Then dW1[0] comes out [0.08688, 0.17376] instead of [0.019877, 0.039754] — about 4.4× too large, because the σ'(z₁[0]) = 0.228784 factor was never applied.

74. Trap: forgetting the activation derivative

Trap

The trap

Gradient flows from h to z₁ unchanged — just pass ∂L/∂h straight through the sigmoid.

dz1 = dh = [0.08688, 0.13032, 0.17376]

Why: Skips the sigmoid's own Jacobian. Then dW1[0] comes out [0.08688, 0.17376] instead of [0.019877, 0.039754] — about 4.4× too large, because the σ'(z₁[0]) = 0.228784 factor was never applied.

The fix

Crossing the sigmoid multiplies elementwise by its local slope σ'(z₁) = h(1−h).

dz1 = dh ⊙ h(1−h) = [0.019877, 0.021933, 0.018244]

Why: Now dW1[0] = [0.019877, 0.039754], matching autograd. This shrink factor (each ≤ 0.25 for sigmoid) is exactly what will make deep sigmoid nets stall — never omit it.

75. Break it on purpose: forgetting the activation derivative

Break the constraint

Discussion prompt

The rule this trap just fixed:

Crossing the sigmoid multiplies elementwise by its local slope σ'(z₁) = h(1−h).

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

Skips the sigmoid's own Jacobian. Then dW1[0] comes out [0.08688, 0.17376] instead of [0.019877, 0.039754] — about 4.4× too large, because the σ'(z₁[0]) = 0.228784 factor was never applied.

76. Gradient checking

Section

Part 5 of 6 — the safety net

77. The finite-difference check for a network

Concept

The same central difference from Part 1 checks a network gradient: perturb one parameter by ±ε, recompute the loss, and compare the slope to your analytic value.

\[ \frac{\partial L}{\partial \theta} \approx \frac{L(\theta+\epsilon) - L(\theta-\epsilon)}{2\epsilon} \]

This is your correctness oracle: if analytic and numeric disagree, the backward pass has a sign or factor bug. It's slow (one forward pass per parameter) so you spot-check a few entries, not train with it.

78. Predict the next row: Gradient check on W₁[0,0]

Pattern

Predict first

The table runs: central difference (ε = 1e−5) | 0.019877 · analytic backprop | 0.019877

In Gradient check on W₁[0,0], given the rows so far: what is the next one — the row where method is also checked: ∂L/∂W₂[0,0]?

Correct: also checked: ∂L/∂W₂[0,0] | 0.280474 = 0.280474

method∂L/∂W₁[0,0]
central difference (ε = 1e−5)0.019877
analytic backprop0.019877
also checked: ∂L/∂W₂[0,0]0.280474 = 0.280474

Why: The relationship between the columns, not the individual numbers, is what generates the next row. Agreement to ~6 digits means no sign or factor bug on this path.

79. Gradient check on W₁[0,0]

Worked example

Check the single entry ∂L/∂W₁[0,0], which our backprop says is 0.019877. Bump W₁[0,0] by ±ε = 1e−5 and read the loss slope:

import numpy as np
x = np.array([1.,2.]); y = np.array([1.])
W1 = np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1 = np.array([0.1,0.2,0.3])
W2 = np.array([[0.2,0.3,0.4]]); b2 = np.array([0.5])
sig = lambda z: 1/(1+np.exp(-z))
h = sig(W1@x + b1); yhat = W2@h + b2          # forward
dz1 = (W2.T@(2*(yhat-y)))*(h*(1-h)); dW1 = np.outer(dz1, x)  # analytic
eps = 1e-5
def loss_with(W1mod):
    hh = sig(W1mod @ x + b1)
    return float(((W2 @ hh + b2 - y)**2).sum())
Wp = W1.copy(); Wp[0,0] += eps
Wm = W1.copy(); Wm[0,0] -= eps
num = (loss_with(Wp) - loss_with(Wm)) / (2*eps)
print(round(num,6), round(dW1[0,0],6))

numeric = 0.019877, analytic = 0.019877 — match

Why: Agreement to ~6 digits means no sign or factor bug on this path. ε too small underflows in floating point; too large loses accuracy — 1e−5 is the sweet spot for a central difference.

method∂L/∂W₁[0,0]
central difference (ε = 1e−5)0.019877
analytic backprop0.019877
also checked: ∂L/∂W₂[0,0]0.280474 = 0.280474

80. Work backwards from the answer: Gradient check on W₁[0,0]

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

numeric = 0.019877, analytic = 0.019877 — match

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Check the single entry ∂L/∂W₁[0,0], which our backprop says is 0.019877. Bump W₁[0,0] by ±ε = 1e−5 and read the loss slope:

81. Activation gradients & vanishing

Section

Part 6 of 6 — why depth hurts

82. The sigmoid slope has a ceiling

Concept

The factor h(1−h) we multiplied by is σ'(z), and it is largest at the center and tiny in the tails:

\[ \sigma'(z) = \sigma(z)\big(1-\sigma(z)\big), \qquad 0 < \sigma'(z) \le \tfrac14 \]

The maximum is 0.25, hit at z = 0 where σ = 0.5. So every sigmoid layer multiplies the upstream gradient by at most a quarter — usually much less.

83. What has to be given first: Where the sigmoid slope peaks

Missing information

Discussion prompt

Tabulate σ'(z) = σ(z)(1−σ(z)) across z to see the 0.25 ceiling and the tail collapse:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

Saturated units (large |z|) have almost no slope — their gradient nearly vanishes even in a single layer. Our hidden units at z = 0.6, 1.3, 2.0 already gave 0.229, 0.168, 0.105.

84. Where the sigmoid slope peaks

Worked example

Tabulate σ'(z) = σ(z)(1−σ(z)) across z to see the 0.25 ceiling and the tail collapse:

import numpy as np
sig = lambda z: 1/(1+np.exp(-z))
for z in [-4,-2,-1,0,1,2,4]:
    s = sig(z)
    print(z, round(s,4), round(s*(1-s),4))

σ'(0) = 0.5·0.5 = 0.25 is the peak; σ'(±4) ≈ 0.0177 in the tails

Why: Saturated units (large |z|) have almost no slope — their gradient nearly vanishes even in a single layer. Our hidden units at z = 0.6, 1.3, 2.0 already gave 0.229, 0.168, 0.105.

zσ(z)σ'(z) = σ(1−σ)
−40.01800.0177
−10.26890.1966
00.50000.2500
10.73110.1966
40.98200.0177

85. Watch it run: Where the sigmoid slope peaks

Pattern

Step through it

Step through Where the sigmoid slope peaks one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: z is −4
  2. Step 2: z is −1
  3. Step 3: z is 0
  4. Step 4: z is 1
  5. Step 5: z is 4

86. Sub-1 factors compound toward zero

Intuition

Backprop through d layers multiplies d local slopes. If each slope is at most 0.25, the gradient reaching the front is at most (0.25)^d of the signal at the top.

Multiplying numbers below 1 shrinks them fast: 0.25² = 0.0625, 0.25⁵ ≈ 0.001. By the time you reach an early layer of a deep sigmoid net, there is essentially no signal left — the vanishing gradient.

87. Teach it back: Sub-1 factors compound toward zero

Explain it

Discussion prompt

Explain Sub-1 factors compound toward zero to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Backprop through d layers multiplies d local slopes. If each slope is at most 0.25, the gradient reaching the front is at most (0.25)^d of the signal at the top.

88. Guess the shape of the answer: The vanishing gradient, quantified

Estimation

Predict first

Compare the best-case sigmoid factor (0.25)^depth to ReLU's active-path factor 1^depth. Runnable:

Commit before you compute: what does The vanishing gradient, quantified come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: After 10 sigmoid layers the gradient is ≤ (0.25)¹⁰ ≈ 9.54e−7 — effectively zero

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. And this is the BEST case (every unit at z = 0).

89. The vanishing gradient, quantified

Worked example

Compare the best-case sigmoid factor (0.25)^depth to ReLU's active-path factor 1^depth. Runnable:

for d in [1, 2, 5, 10, 20]:
    print(d, format(0.25**d, '.3e'), 1.0**d)   # sigmoid best-case vs ReLU active

After 10 sigmoid layers the gradient is ≤ (0.25)¹⁰ ≈ 9.54e−7 — effectively zero

Why: And this is the BEST case (every unit at z = 0). Real saturated units make it far worse. ReLU's active slope of 1 doesn't shrink at all: 1¹⁰ = 1.

depthsigmoid ≤ (0.25)ᵈReLU (active) 1ᵈ
12.500e−011
26.250e−021
59.766e−041
109.537e−071
209.095e−131

90. What stays fixed: The vanishing gradient, quantified

Invariant

Step through it

Step through The vanishing gradient, quantified one row at a time. One of these columns never changes — find it, and say why it cannot.

  1. Step 1: depth is 1
  2. Step 2: depth is 2
  3. Step 3: depth is 5
  4. Step 4: depth is 10
  5. Step 5: depth is 20

91. ReLU: a gradient that survives depth

Concept

ReLU's derivative is 1 wherever the unit is active and 0 where it's off — no shrinkage on the active path, so gradients pass through many layers intact.

\[ \operatorname{ReLU}'(z) = \begin{cases} 1 & z > 0 \\ 0 & z < 0 \end{cases} \]

The cost is the dying-ReLU problem: a unit stuck at z < 0 gets zero gradient forever and never recovers. Most deep nets happily take that trade — and add residual connections and normalization for extra insurance.

92. By analogy: ReLU: a gradient that survives depth

Analogy

Discussion prompt

Explain ReLU: a gradient that survives depth by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

ReLU's derivative is 1 wherever the unit is active and 0 where it's off — no shrinkage on the active path, so gradients pass through many layers intact.

93. Something is wrong here: do more layers mean more gradient?

Anomaly

Predict first

A student writes this, and it looks reasonable:

Stacking many sigmoid layers gives the network more capacity, so the early layers should train at least as fast as a shallow one.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Ignores that gradients MULTIPLY through layers.

Each layer multiplies the gradient by its local slope, so sub-1 factors compound toward zero.

Why: Ignores that gradients MULTIPLY through layers. With per-layer factors ≤ 0.25, depth makes the front-layer signal vanish, not grow — capacity is irrelevant if no gradient arrives.

94. Trap: do more layers mean more gradient?

Trap

The trap

Stacking many sigmoid layers gives the network more capacity, so the early layers should train at least as fast as a shallow one.

Assume the gradient magnitude is roughly constant with depth

Why: Ignores that gradients MULTIPLY through layers. With per-layer factors ≤ 0.25, depth makes the front-layer signal vanish, not grow — capacity is irrelevant if no gradient arrives.

The fix

Each layer multiplies the gradient by its local slope, so sub-1 factors compound toward zero.

Front-layer gradient ∝ (factor)^depth → (0.25)^depth for sigmoid

Why: That is why deep nets use ReLU (slope 1 when active), residual connections (add an identity path so the factor is ≥ 1), and normalization — all to keep the gradient product from collapsing.

95. Which of these survive contact with Lesson 9: The Chain Rule & Backpropagation?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
A composition y = f(g(x)) is two machines wired in series: x goes into g to make an intermediate u = g(x), and u goes into f to make y = f(u).; A tiny bump Δx becomes Δu ≈ g'(x)·Δx after the first machine, and that becomes Δy ≈ f'(u)·Δu after the second. Substitute the first into the second.; Every analytic derivative can be sanity-checked numerically. A central finite difference nudges the input by ±ε and measures the output change:
Breaks
We apply g first, then f, so write the Jacobians in that left-to-right reading order: J_g · J_f.; Gradient flows from h to z₁ unchanged — just pass ∂L/∂h straight through the sigmoid.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 9: The Chain Rule & Backpropagation puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

96. Rebuild the recipe: The backprop recipe

Ranking

Put in order

These are the steps of The backprop recipe, scrambled. Put them back in order before the next slide shows you.

  1. Forward: compute and CACHE every intermediate (z₁, h, ŷ) — the backward pass needs them
  2. Seed: ∂L/∂ŷ from the loss (here 2(ŷ − y))
  3. Linear layer, backward: weight grad = outer(upstream, input); bias grad = upstream; pass Wᵀ·upstream down
  4. Activation, backward: multiply elementwise by its derivative (sigmoid → h(1−h), ReLU → 1 if active else 0)
  5. Verify any entry with a central finite difference, and expect vanishing gradients when many sub-1 factors stack

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

97. The backprop recipe

Pattern

  1. Forward: compute and CACHE every intermediate (z₁, h, ŷ) — the backward pass needs them
  2. Seed: ∂L/∂ŷ from the loss (here 2(ŷ − y))
  3. Linear layer, backward: weight grad = outer(upstream, input); bias grad = upstream; pass Wᵀ·upstream down
  4. Activation, backward: multiply elementwise by its derivative (sigmoid → h(1−h), ReLU → 1 if active else 0)
  5. Verify any entry with a central finite difference, and expect vanishing gradients when many sub-1 factors stack

98. Where does it stop working: The backprop recipe

Edge cases

Discussion prompt

The backprop recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Forward: compute and CACHE every intermediate (z₁, h, ŷ) — the backward pass needs them
  2. Seed: ∂L/∂ŷ from the loss (here 2(ŷ − y))
  3. Linear layer, backward: weight grad = outer(upstream, input); bias grad = upstream; pass Wᵀ·upstream down
  4. Activation, backward: multiply elementwise by its derivative (sigmoid → h(1−h), ReLU → 1 if active else 0)
  5. Verify any entry with a central finite difference, and expect vanishing gradients when many sub-1 factors stack

99. Rule out three: Check yourself — Jacobian order

Elimination

Eliminate the wrong options

For y = f(g(x)) with vector x, the Jacobian J of the composition is:

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. J_f(g(x)) · J_g(x) — outer Jacobian on the left
  • B. J_g(x) · J_f(g(x)) — inner Jacobian on the left
  • C. J_f(x) · J_g(x), both evaluated at x
  • D. J_f(g(x)) + J_g(x), summed

Survives elimination: A

Why: The chain rule is multiplicative with the OUTER function's Jacobian first: J_f evaluated at g(x), times J_g at x. Dimensions (m×k)(k×n) = (m×n) confirm the order.

100. Check yourself — Jacobian order

Check

Match the matrix order to the derivative — and check the shapes.

Check your understanding

For y = f(g(x)) with vector x, the Jacobian J of the composition is:

  • A. J_f(g(x)) · J_g(x) — outer Jacobian on the left (correct)
  • B. J_g(x) · J_f(g(x)) — inner Jacobian on the left
  • C. J_f(x) · J_g(x), both evaluated at x
  • D. J_f(g(x)) + J_g(x), summed

Answer: A

Why: The chain rule is multiplicative with the OUTER function's Jacobian first: J_f evaluated at g(x), times J_g at x. Dimensions (m×k)(k×n) = (m×n) confirm the order.

Why B tempts people
Reversed order — generally the wrong shape (inner dimensions won't match) and the wrong composition. Reading order (g then f) is not multiplication order.
Why C tempts people
J_f must be evaluated at the INNER value g(x), not at x. Evaluating the outer Jacobian at the raw input is a common slip that gives the wrong numbers even when the shape is right.
Why D tempts people
The chain rule multiplies local Jacobians, it never adds them. Addition appears in the multivariable sum-OVER-PATHS case (one input, several paths), not in a single composition.

101. Answer it before you see the options: Check yourself — the activation factor

Prediction

Predict first

Crossing a sigmoid layer in the backward pass, you convert ∂L/∂h into ∂L/∂z by multiplying (elementwise) by:

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: h(1 − h), the sigmoid derivative

Why: σ'(z) = σ(z)(1−σ(z)) = h(1−h), applied elementwise because the activation's Jacobian is diagonal. In our net this turned ∂L/∂h = [0.08688, 0.13032, 0.17376] into ∂L/∂z₁ = [0.019877, 0.021933, 0.018244].

102. Check yourself — the activation factor

Check

Trace one step of the backward pass across the sigmoid.

Check your understanding

Crossing a sigmoid layer in the backward pass, you convert ∂L/∂h into ∂L/∂z by multiplying (elementwise) by:

  • A. h(1 − h), the sigmoid derivative (correct)
  • B. W, the layer's weight matrix
  • C. 1 — nothing changes
  • D. h, the activation value itself

Answer: A

Why: σ'(z) = σ(z)(1−σ(z)) = h(1−h), applied elementwise because the activation's Jacobian is diagonal. In our net this turned ∂L/∂h = [0.08688, 0.13032, 0.17376] into ∂L/∂z₁ = [0.019877, 0.021933, 0.018244].

Why B tempts people
Multiplying by Wᵀ is how you cross the LINEAR part of a layer (∂L/∂ŷ → ∂L/∂h), not the nonlinearity. Different component of the same layer.
Why C tempts people
Passing the gradient through unchanged is exactly the Part-4 trap — it omits the activation's derivative and inflates the upstream weight gradients (here by ~4.4×).
Why D tempts people
It's h(1−h), not h. Using h alone is the right variable but the wrong function — you need the derivative σ', not the activation value.

103. Rule out three: Check yourself — vanishing gradients

Elimination

Eliminate the wrong options

A 10-layer network of sigmoids trains very slowly in its early layers. The best explanation is:

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Each layer multiplies the gradient by ≤ 0.25, so (0.25)¹⁰ ≈ 1e−6 reaches the front
  • B. The learning rate is too small
  • C. Sigmoid outputs are always exactly 0 or 1
  • D. There are too few parameters to fit the data

Survives elimination: A

Why: Every sigmoid contributes a factor σ' ≤ 0.25 to the gradient. Over 10 layers the product is at most (0.25)¹⁰ ≈ 9.5e−7, so the front layers get almost no signal — the vanishing-gradient problem.

104. Check yourself — vanishing gradients

Check

Think about what compounds as depth grows.

Check your understanding

A 10-layer network of sigmoids trains very slowly in its early layers. The best explanation is:

  • A. Each layer multiplies the gradient by ≤ 0.25, so (0.25)¹⁰ ≈ 1e−6 reaches the front (correct)
  • B. The learning rate is too small
  • C. Sigmoid outputs are always exactly 0 or 1
  • D. There are too few parameters to fit the data

Answer: A

Why: Every sigmoid contributes a factor σ' ≤ 0.25 to the gradient. Over 10 layers the product is at most (0.25)¹⁰ ≈ 9.5e−7, so the front layers get almost no signal — the vanishing-gradient problem.

Why B tempts people
A small learning rate slows ALL layers uniformly; here the FRONT layers specifically stall because their gradient is tiny — a depth effect from the multiplied σ' factors, not a global rate effect.
Why C tempts people
Sigmoid outputs lie strictly in (0,1), never pinned to 0/1. The culprit is the SLOPE σ' ≤ 0.25, not the output values.
Why D tempts people
Too few parameters causes underfitting everywhere at once, not a depth-dependent gradient collapse concentrated in the early layers.

105. Your turn: build backprop

Section

The project — a 2-layer MLP by hand

106. Project: a 2-layer MLP, forward and backward

Concept

Implement forward AND backward for our running network in pure NumPy, then prove every gradient matches torch.autograd and a finite difference. You've derived every piece — now assemble it yourself.

#requirementtool
1Forward: z₁, h, ŷ, L@ and the sigmoid
2Backward: dW1, db1, dW2, db2outer products + h(1−h)
3Verify vs autograd and a central differencetorch + finite diff

Build rules: cache the forward intermediates (you need h in the backward pass), type every line yourself, run after each line, and when something errors read the array shapes before you change anything.

107. Break it if you can: Project: a 2-layer MLP, forward and backward

Counterexample

Discussion prompt

Implement forward AND backward for our running network in pure NumPy, then prove every gradient matches torch.autograd and a finite difference. You've derived every piece — now assemble it yourself.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

108. Milestone 1 — forward pass

Worked example

Your turn: compute z₁, h, ŷ, L from the given weights. Predict out loud whether ŷ lands above or below the target of 1 before you print.

Hint: z1 = W1@x + b1, h = sig(z1), yhat = W2@h + b2, L = ((yhat-y)**2).sum(). Cache h — you'll need it going backward.

import numpy as np
x = np.array([1.,2.]); y = np.array([1.])
W1 = np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1 = np.array([0.1,0.2,0.3])
W2 = np.array([[0.2,0.3,0.4]]); b2 = np.array([0.5])
sig = lambda z: 1/(1+np.exp(-z))
z1 = W1@x + b1; h = sig(z1)
yhat = W2@h + b2; L = float(((yhat - y)**2).sum())
print(h, yhat, round(L,5))
valueresult
z₁[0.6, 1.3, 2.0]
h[0.645656, 0.785835, 0.880797]
ŷ1.217201 (above target 1)
L0.04718

109. Milestone 2 — backward pass

Worked example

Your turn: seed ∂L/∂ŷ, then walk back to dW1. The one thing you must not drop is the h(1−h) factor when you cross the sigmoid.

Hint: dyhat = 2*(yhat-y); dW2 = np.outer(dyhat, h); dh = W2.T@dyhat; dz1 = dh*(h*(1-h)); dW1 = np.outer(dz1, x).

import numpy as np
x = np.array([1.,2.]); y = np.array([1.])
W1 = np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1 = np.array([0.1,0.2,0.3])
W2 = np.array([[0.2,0.3,0.4]]); b2 = np.array([0.5])
sig = lambda z: 1/(1+np.exp(-z))
h = sig(W1@x + b1); yhat = W2@h + b2   # forward pass from Milestone 1
# --- your backward pass ---
dyhat = 2*(yhat - y)
dW2 = np.outer(dyhat, h); db2 = dyhat
dh  = W2.T @ dyhat
dz1 = dh * (h*(1-h))
dW1 = np.outer(dz1, x); db1 = dz1
print(dW1)
gradientvalue
∂L/∂ŷ0.434401
dW2[0.280474, 0.341368, 0.382619]
dz1[0.019877, 0.021933, 0.018244]
dW1[0][0.019877, 0.039754]

110. Milestone 3 — verify against autograd + finite diff

Worked example

Your turn: rebuild the forward pass in torch with requires_grad=True, call .backward(), and compare. Then spot-check one entry with a central difference. Predict: exact match or just close?

Hint: wrap W1,b1,W2,b2 as tensors with requires_grad=True, recompute L with torch.sigmoid, call L.backward(), and compare W1t.grad.numpy() to dW1 with np.allclose.

import numpy as np, torch
x = np.array([1.,2.]); y = np.array([1.])
W1 = np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1 = np.array([0.1,0.2,0.3])
W2 = np.array([[0.2,0.3,0.4]]); b2 = np.array([0.5])
sig = lambda z: 1/(1+np.exp(-z))
h = sig(W1@x + b1); yhat = W2@h + b2                    # forward
dz1 = (W2.T@(2*(yhat-y)))*(h*(1-h)); dW1 = np.outer(dz1, x)  # your dW1
W1t=torch.tensor(W1,requires_grad=True); b1t=torch.tensor(b1,requires_grad=True)
W2t=torch.tensor(W2,requires_grad=True); b2t=torch.tensor(b2,requires_grad=True)
h_t=torch.sigmoid(W1t@torch.tensor(x)+b1t)
L_t=((W2t@h_t+b2t-torch.tensor(y))**2).sum(); L_t.backward()
print(np.allclose(dW1, W1t.grad.numpy()))
compareresult
allclose(dW1, autograd)True
allclose(dW2, autograd)True
numeric vs analytic dW1[0,0]0.019877 = 0.019877

111. What each one costs: Milestone 3 — verify against autograd + finite…

Trade off

Comparison matrix

From Milestone 3 — verify against autograd + finite diff: every row here is a choice with a cost. Fill the result column, then say which row you would actually pick and what you give up for it.

compareresult
allclose(dW1, autograd)True
allclose(dW2, autograd)True
numeric vs analytic dW1[0,0]0.019877 = 0.019877

112. The full program

Concept

import numpy as np
x=np.array([1.,2.]); y=np.array([1.])
W1=np.array([[0.1,0.2],[0.3,0.4],[0.5,0.6]]); b1=np.array([0.1,0.2,0.3])
W2=np.array([[0.2,0.3,0.4]]); b2=np.array([0.5])
sig=lambda z:1/(1+np.exp(-z))
# forward (cache h for the backward pass)
h=sig(W1@x+b1); yhat=W2@h+b2; L=float(((yhat-y)**2).sum())
# backward
dyhat=2*(yhat-y); dW2=np.outer(dyhat,h); db2=dyhat
dz1=(W2.T@dyhat)*(h*(1-h)); dW1=np.outer(dz1,x); db1=dz1
print('L =',round(L,5),' dW1[0,0] =',round(dW1[0,0],6))
outputvalue
L0.04718
dW1[0,0]0.019877

If your L reads 0.04718, dW1[0,0] reads 0.019877, and allclose against autograd returns True — you implemented backpropagation by hand.

113. Fill in: value for The full program

Comparison

Comparison matrix

From The full program: refill the value column from what you know. The rest of the table is as it appeared.

outputvalue
L0.04718
dW1[0,0]0.019877

114. Show it off

Concept

Out loud, slides closed: explain (1) why Jacobians multiply outer-first, (2) what the h(1−h) factor does to a gradient over 10 layers, and (3) why a central finite difference catches a sign bug your eyes miss.

Stretch (homework): prove the softmax Jacobian ∂sᵢ/∂xⱼ = sᵢ(δᵢⱼ − sⱼ), then swap MSE for softmax + cross-entropy and re-derive the famously clean ∂L/∂z = p − y. Confirm it against autograd the same way.

115. Connect it up: Lesson 9: The Chain Rule & Backpropagation

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — The single-variable chain rule · The multivariable chain rule · The Jacobian form · Backprop as reverse-mode autodiff · Gradient checking · Activation gradients & vanishing. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

116. What you can do now

Recap

movethe one thing to remember
chain rulemultiply local Jacobians, OUTER first
many pathssum the contributions (fan-out adds gradients)
linear layer backweight grad = outer(upstream, input); pass Wᵀ·upstream
activation backmultiply elementwise by its derivative — sigmoid → h(1−h)
depthsub-1 factors compound → sigmoid vanishes; ReLU survives
verifycentral difference, ε ≈ 1e−5

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 9 (Week 3 — Chain Rule & Backprop) — Barron · USAAIO Round 2 Preparation, 2026
  2. PyTorch Autograd mechanics
  3. Michael Nielsen, Neural Networks and Deep Learning, Ch. 2 (How the backpropagation algorithm works) — neuralnetworksanddeeplearning.com
  4. Every forward value, gradient, Jacobian and finite-difference number produced by real execution — numpy 2.2.6 + torch 2.7.1, verification run July 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108