Lesson 6: Gradient Intuition & PyTorch Autograd

USAAIO Lesson 6, from Week 2 on calculus, fully worked. It builds partial derivatives one axis at a time, assembles the gradient and reads it as a steepest-ascent compass, and derives the directional derivative from the projection. It then derives the three ML gradient identities - 2w, x, and 2Ax - and checks each one against autograd, covers the computational graph and the chain rule behind .backward(), and shows the accumulation and no_grad traps in real execution. It ends with a from-scratch gradient-descent project verified line by line against PyTorch. Every snippet runs as written with the seeds set, and every number in a trace was produced by real execution. The lesson runs to 60 slides.

Subject: Machine Learning · 114 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Gradients & PyTorch Autograd

Title

USAAIO · Lesson 6 · Week 2 (Calculus)

The single object every optimizer, every backprop step, and every transformer gradient is built on. We derive the gradient from one partial at a time, prove the three identities that power all of ML, then let autograd reproduce every number — and end by building gradient descent from scratch.

2. By the end of this lesson you can

Objectives

  1. Compute a partial derivative by freezing every other variable, and stack partials into the gradient vector
  2. Explain — not just assert — why ∇f points in the direction of steepest ascent, and take a correct descent step
  3. Derive the directional derivative D_v f = ∇f · v̂ and read off its max, min, and zero directions
  4. Derive and apply the three ML identities ∇‖w‖² = 2w, ∇(wᵀx) = x, ∇(xᵀAx) = 2Ax, each checked against autograd
  5. Use requires_grad, .backward(), .grad, and torch.no_grad() — and explain why the update needs zero_grad()
  6. Build gradient descent from scratch on f(w) = (w−5)² and verify every gradient against PyTorch

3. The road ahead

Concept

Five parts, each building on the last. We start from a single slope and end holding a working optimizer you wrote yourself.

1. The gradient
Partials → a steepest-ascent vector.
2. Directional
Slope in ANY direction.
3. Identities
2w, x, 2Ax — the ML three.
4. Autograd
The chain rule, automated.
5. Build it
Gradient descent from scratch.

4. Which is which: The road ahead

Matching

Match the pairs

From The road ahead — match each one to what it actually does. The descriptions have been shuffled.

  • c1. 1. The gradient
  • c2. 3. Identities
  • c3. 4. Autograd
  • c4. 5. Build it
  • b1. Partials → a steepest-ascent vector.
  • b2. 2w, x, 2Ax — the ML three.
  • b3. The chain rule, automated.
  • b4. Gradient descent from scratch.

Why: 1. The gradient, 3. Identities, 4. Autograd, 5. Build it are easy to tell apart while they are sitting next to their descriptions and much harder afterwards, which is what this checks.

5. From one slope to many

Section

Part 1 of 5 — partial derivatives

6. Our running function

Concept

One function carries the whole first half of this lesson. It takes two inputs and returns one number — a surface floating over the (x, y) plane.

\[ f(x,y) = x^2 + 3xy + y^2 \]

In 1-D a derivative is a single slope. Here there is no single slope — the surface tilts differently depending on which way you walk. We need a slope per direction, starting with the two axes.

7. Break it if you can: Our running function

Counterexample

Discussion prompt

One function carries the whole first half of this lesson. It takes two inputs and returns one number — a surface floating over the (x, y) plane.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

In 1-D a derivative is a single slope. Here there is no single slope — the surface tilts differently depending on which way you walk. We need a slope per direction, starting with the two axes.

8. What a partial derivative means

Concept

The partial derivative ∂f/∂x is the slope you feel walking in the +x direction only — y held fixed, as if it were a constant.

Mechanically: to take ∂f/∂xᵢ, differentiate with respect to xᵢ and treat every other variable as a frozen number. Nothing else about differentiation changes.

\[ \frac{\partial f}{\partial x} = \lim_{h\to 0}\frac{f(x+h,\,y) - f(x,\,y)}{h} \]

9. By analogy: What a partial derivative means

Analogy

Discussion prompt

Explain What a partial derivative means by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

The partial derivative ∂f/∂x is the slope you feel walking in the +x direction only — y held fixed, as if it were a constant.

10. Freezing y is slicing the surface

Intuition

Holding y fixed is like cutting the surface with a vertical knife parallel to the x-axis. You expose a single curved slice — an ordinary 1-D curve.

∂f/∂x is just the everyday slope of that slice. The multivariable part is only bookkeeping: which knife you're using. The calculus underneath is the same one-variable calculus you already know.

So there is nothing new to learn to differentiate — only something to hold still.

11. Teach it back: Freezing y is slicing the surface

Explain it

Discussion prompt

Explain Freezing y is slicing the surface to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Holding y fixed is like cutting the surface with a vertical knife parallel to the x-axis. You expose a single curved slice — an ordinary 1-D curve.

12. What has to happen first: Partial in x, term by term

Ranking

Put in order

Put the moves of Partial in x, term by term into the order they have to happen.

  1. Term x² → 2x
  2. Term 3xy → 3y
  3. Add the three → ∂f/∂x = 2x + 3y

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Ordinary power rule in x: d(x²)/dx = 2x.

13. Partial in x, term by term

Worked example

Differentiate f = x² + 3xy + y² with respect to x, treating y as a constant. Take one term at a time — no shortcuts:

Term x² → 2x

Why: Ordinary power rule in x: d(x²)/dx = 2x. No y here to worry about.

Term 3xy → 3y

Why: 3y is a constant coefficient on x, so d(3y·x)/dx = 3y. The x differentiates away, the frozen y stays.

Term y² → 0

Why: y² has no x in it at all, so as a function of x it is a constant — and the derivative of a constant is 0.

Add the three → ∂f/∂x = 2x + 3y

Why: Sum the term-by-term results. This is the slope of the x-slice at any point (x, y).

\[ \frac{\partial f}{\partial x} = 2x + 3y \]

14. Decode the notation: Partial in x, term by term

Notation

Annotate

From Partial in x, term by term — read this one piece at a time. What is each part doing?

On: \( \frac{\partial f}{\partial x} = 2x + 3y \)

  • Ordinary power rule in x: d(x²)/dx = 2x. No y here to worry about.
  • 3y is a constant coefficient on x, so d(3y·x)/dx = 3y. The x differentiates away, the frozen y stays.
  • y² has no x in it at all, so as a function of x it is a constant — and the derivative of a constant is 0.

15. Plan first: Partial in y, term by term

Step zero

Discussion prompt

Partial in y, term by term — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Term x² → 0

Answer:

  1. Term x² → 0
  2. Term 3xy → 3x
  3. Term y² → 2y
  4. Add the three → ∂f/∂y = 3x + 2y

16. Partial in y, term by term

Worked example

Now freeze x and differentiate with respect to y. Same function, other knife:

Term x² → 0

Why: x² has no y, so it is constant in y and drops to 0.

Term 3xy → 3x

Why: 3x is the constant coefficient on y, so d(3x·y)/dy = 3x.

Term y² → 2y

Why: Ordinary power rule in y: d(y²)/dy = 2y.

Add the three → ∂f/∂y = 3x + 2y

Why: The other partial. Notice the two partials are DIFFERENT functions — the surface really does tilt differently along each axis.

\[ \frac{\partial f}{\partial y} = 3x + 2y \]

17. Work backwards from the answer: Partial in y, term by term

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Add the three → ∂f/∂y = 3x + 2y

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Now freeze x and differentiate with respect to y. Same function, other knife:

18. Stack the partials: the gradient

Concept

Collect the two partials into a single vector. That vector is the gradient, written ∇f. It is not a number — it is an arrow living in the input space.

\[ \nabla f = \begin{bmatrix} \partial f / \partial x \\[2pt] \partial f / \partial y \end{bmatrix} = \begin{bmatrix} 2x + 3y \\ 3x + 2y \end{bmatrix} \]

gradient ∇f — The vector of all first partial derivatives of a scalar function. For f: ℝⁿ → ℝ it is an n-vector. It is a FUNCTION of position — a different arrow at every point of the input space.

19. Guess the shape of the answer: Evaluate the gradient at (1, 2)

Estimation

Predict first

The gradient is a formula in x and y. To get an actual arrow, plug in a point. Use (1, 2):

Commit before you compute: what does Evaluate the gradient at (1, 2) come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: ∇f(1,2) = [8, 7]

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The arrow at (1,2). Both components positive, so f increases as you move in +x and +y from this point.

20. Evaluate the gradient at (1, 2)

Worked example

The gradient is a formula in x and y. To get an actual arrow, plug in a point. Use (1, 2):

∂f/∂x at (1,2) = 2(1) + 3(2) = 2 + 6 = 8

Why: Substitute x=1, y=2 into 2x+3y.

∂f/∂y at (1,2) = 3(1) + 2(2) = 3 + 4 = 7

Why: Substitute x=1, y=2 into 3x+2y.

∇f(1,2) = [8, 7]

Why: The arrow at (1,2). Both components positive, so f increases as you move in +x and +y from this point.

\[ \nabla f(1,2) = \begin{bmatrix} 8 \\ 7 \end{bmatrix} \]

componentformulavalue at (1,2)
∂f/∂x2x + 3y8
∂f/∂y3x + 2y7
f itselfx² + 3xy + y²1 + 6 + 4 = 11

21. Fill in: formula for Evaluate the gradient at (1, 2)

Comparison

Comparison matrix

From Evaluate the gradient at (1, 2): refill the formula column from what you know. The rest of the table is as it appeared.

componentformulavalue at (1,2)
∂f/∂x2x + 3y8
∂f/∂y3x + 2y7
f itselfx² + 3xy + y²1 + 6 + 4 = 11

22. Plan first: A second gradient, with a mixed term

Step zero

Discussion prompt

A second gradient, with a mixed term — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: ∂g/∂x: freeze y, differentiate x²y + y³ in x → 2xy + 0 = 2xy

Answer:

  1. ∂g/∂x: freeze y, differentiate x²y + y³ in x → 2xy + 0 = 2xy
  2. ∂g/∂y: freeze x, differentiate x²y + y³ in y → x² + 3y²
  3. At (2, 3): ∇g = [2·2·3, 2² + 3·3²] = [12, 31]

23. A second gradient, with a mixed term

Worked example

One more, because the exam loves functions where a variable hides inside a product. Take g(x, y) = x²y + y³ and find ∇g at (2, 3):

∂g/∂x: freeze y, differentiate x²y + y³ in x → 2xy + 0 = 2xy

Why: x²y has coefficient y on x², giving 2xy. y³ has no x and drops.

∂g/∂y: freeze x, differentiate x²y + y³ in y → x² + 3y²

Why: x²y has coefficient x² on y, giving x². y³ gives 3y².

At (2, 3): ∇g = [2·2·3, 2² + 3·3²] = [12, 31]

Why: Substitute. The two partials are wildly different sizes — a reminder that the gradient's DIRECTION matters, not just one component.

\[ \nabla g(2,3) = \begin{bmatrix} 2xy \\ x^2 + 3y^2 \end{bmatrix}_{(2,3)} = \begin{bmatrix} 12 \\ 31 \end{bmatrix} \]

import torch
x = torch.tensor(2.0, requires_grad=True)
y = torch.tensor(3.0, requires_grad=True)
g = x**2 * y + y**3
g.backward()
print(x.grad.item(), y.grad.item())
partialformulavalue (verified)
∂g/∂x2xy12.0
∂g/∂yx² + 3y²31.0

24. What each one costs: A second gradient, with a mixed term

Trade off

Comparison matrix

From A second gradient, with a mixed term: every row here is a choice with a cost. Fill the formula column, then say which row you would actually pick and what you give up for it.

partialformulavalue (verified)
∂g/∂x2xy12.0
∂g/∂yx² + 3y²31.0

25. One arrow per point: a vector field

Concept

Because ∇f is a formula in x and y, it gives a different arrow at every point. The gradient isn't one vector — it's a whole field of them draped over the input plane.

Evaluate it at a few points to feel the field. Notice (2, 1) swaps the components of (1, 2) — same length, different heading — and at the flat point (0, 0) the arrow vanishes:

point (x, y)∇f = [2x+3y, 3x+2y]‖∇f‖ (steepness)
(0, 0)[0, 0]0.000
(1, 0)[2, 3]3.606
(0, 1)[3, 2]3.606
(1, 2)[8, 7]10.630
(2, 1)[7, 8]10.630

A point where ∇f = 0 is a critical point — flat in every direction. That is exactly what an optimizer is hunting for: the place the gradient shuts off.

26. Which is which, by ‖∇f‖ (steepness)

Discrimination

Sort into buckets

Sort these by ‖∇f‖ (steepness), from memory, without looking back at One arrow per point: a vector field. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

0.000
(0, 0)
3.606
(1, 0); (0, 1)
10.630
(1, 2); (2, 1)
g1
‖∇f‖ (steepness) is "0.000" for (0, 0) — that is what the table on "One arrow per point: a vector field" records, and it is the single property separating this group from the rest.
g2
‖∇f‖ (steepness) is "3.606" for (1, 0), (0, 1) — that is what the table on "One arrow per point: a vector field" records, and it is the single property separating this group from the rest.
g3
‖∇f‖ (steepness) is "10.630" for (1, 2), (2, 1) — that is what the table on "One arrow per point: a vector field" records, and it is the single property separating this group from the rest.

27. The gradient is a compass

Intuition

Stand anywhere on the surface. The gradient arrow points in the direction where f climbs fastest, and its length is how steep that fastest climb is.

At (1, 2) the compass reads [8, 7]: the steepest uphill direction, pointing mostly along +x and a bit less along +y. Its length √(8² + 7²) = √113 ≈ 10.63 is the maximum rate of climb there.

Training walks the opposite way: take the gradient, flip its sign, step downhill. That one idea is every optimizer you will meet this year.

28. Why steepest ascent (the argument)

Concept

This isn't a slogan — it falls out of one inequality. The change in f for a tiny unit step v̂ is ∇f · v̂, and the dot product obeys:

\[ \nabla f \cdot \hat v = \lVert \nabla f \rVert\,\lVert \hat v \rVert \cos\theta = \lVert \nabla f \rVert \cos\theta \]

‖v̂‖ = 1, so the step's payoff is ‖∇f‖·cos θ, where θ is the angle between your step and the gradient. This is maximized when cos θ = 1 — that is, when you walk exactly along ∇f. Any other direction gives up a factor of cos θ < 1.

29. The gradient is the local linear map

Concept

Zoom in far enough and any smooth surface looks flat — a tangent plane. The gradient is the slope of that plane: it tells you how f changes for a small step Δ = [Δx, Δy].

\[ f(x + \Delta) \approx f(x) + \nabla f \cdot \Delta \]

This is the first-order Taylor approximation. Every gradient step trusts it: move a little in the direction the linear map says is downhill, then re-measure. It is exact only in the limit of tiny steps — which is why too-large a learning rate breaks the approximation and diverges.

30. Something is wrong here: which way does the gradient point?

Anomaly

Predict first

A student writes this, and it looks reasonable:

Goal: minimize the loss, so the gradient must point downhill toward the minimum — follow it.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Treats the gradient as a downhill arrow and walks ALONG it.

The gradient points toward steepest ascent — uphill, toward larger loss. To go down, negate it.

Why: Treats the gradient as a downhill arrow and walks ALONG it. But ∇f is steepest ASCENT, so this climbs toward LARGER loss — the loss goes up every step and training diverges.

31. Trap: which way does the gradient point?

Trap

The trap

Goal: minimize the loss, so the gradient must point downhill toward the minimum — follow it.

Update: w = w + lr · grad

Why: Treats the gradient as a downhill arrow and walks ALONG it. But ∇f is steepest ASCENT, so this climbs toward LARGER loss — the loss goes up every step and training diverges.

The fix

The gradient points toward steepest ascent — uphill, toward larger loss. To go down, negate it.

Update: w = w − lr · grad

Why: Step in the NEGATIVE gradient direction (cos θ = −1 minimizes ∇f · v̂). The single minus sign is the entire difference between training and diverging — memorize it.

32. The directional derivative

Section

Part 2 of 5 — any direction, not just the axes

33. Slope in an arbitrary direction

Concept

Partials give the slope along +x and +y. But what if you walk diagonally — some direction v that is not an axis? How fast does f change then?

The answer is the directional derivative D_v f: project the gradient onto the unit vector pointing along v.

\[ D_v f = \nabla f \cdot \frac{v}{\lVert v \rVert} = \nabla f \cdot \hat v \]

34. Why we must normalize v

Intuition

A direction is about which way, not how far. If you dotted with a long vector v instead of the unit v̂, you'd multiply the slope by v's length — measuring speed and distance tangled together.

Dividing by ‖v‖ strips out the length, leaving a pure rate: units of f per unit of distance moved. That is what 'slope in a direction' has to mean.

35. Predict the next row: Directional derivative along v = [1, 1]

Pattern

Predict first

The table runs: ∇f at (1,2) | [8, 7] · v̂ = v/‖v‖ | [0.7071, 0.7071] · D_v f = ∇f · v̂ | 10.6066

In Directional derivative along v = [1, 1], given the rows so far: what is the next one — the row where quantity is ‖∇f‖ (the maximum)?

Correct: ‖∇f‖ (the maximum) | 10.6301

quantityvalue (verified)
∇f at (1,2)[8, 7]
v̂ = v/‖v‖[0.7071, 0.7071]
D_v f = ∇f · v̂10.6066
‖∇f‖ (the maximum)10.6301

Why: The relationship between the columns, not the individual numbers, is what generates the next row. Divide each component by ‖v‖ = √2 to get a unit vector.

36. Directional derivative along v = [1, 1]

Worked example

Using ∇f(1,2) = [8, 7], find the rate of change if you move along v = [1, 1] (the diagonal):

Length: ‖v‖ = √(1² + 1²) = √2

Why: Ordinary Euclidean norm of [1, 1].

Normalize: v̂ = [1/√2, 1/√2] ≈ [0.707, 0.707]

Why: Divide each component by ‖v‖ = √2 to get a unit vector.

Dot with the gradient: D_v f = 8·(1/√2) + 7·(1/√2) = 15/√2

Why: ∇f · v̂. Both terms share 1/√2, so it factors to (8+7)/√2 = 15/√2.

\[ D_v f = \frac{8 + 7}{\sqrt2} = \frac{15}{\sqrt2} \approx 10.607 \]

Compare to the max: ‖∇f‖ = √113 ≈ 10.630

Why: 10.607 < 10.630 — climbing the diagonal is ALMOST as steep as the true steepest direction, but slightly less, exactly as the cos θ argument predicts.

quantityvalue (verified)
∇f at (1,2)[8, 7]
v̂ = v/‖v‖[0.7071, 0.7071]
D_v f = ∇f · v̂10.6066
‖∇f‖ (the maximum)10.6301

37. Say it in words: Directional derivative along v = [1, 1]

Translation

\( D_v f = \frac{8 + 7}{\sqrt2} = \frac{15}{\sqrt2} \approx 10.607 \)

Draw it

Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.

38. The three special directions

Concept

D_v f = ‖∇f‖ cos θ makes the whole picture fall out of one angle θ between your step and the gradient:

That perpendicular case is why gradients are always drawn at right angles to contour lines: along a contour, f doesn't change, so its directional derivative is zero.

39. What has to happen first: Descent direction has slope −‖∇f‖

Ranking

Put in order

Put the moves of Descent direction has slope −‖∇f‖ into the order they have to happen.

  1. Descent unit vector: v̂ = −∇f / ‖∇f‖ = [−8, −7]/√113
  2. D_v f = ∇f · v̂ = −(8² + 7²)/√113 = −113/√113 = −√113
  3. Compare to the level-curve direction

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Point exactly opposite the gradient, normalized to unit length so we read a pure slope.

40. Descent direction has slope −‖∇f‖

Worked example

The direction training actually uses is −∇f. Confirm its directional derivative is the most-negative possible, −‖∇f‖, using ∇f(1,2) = [8, 7]:

Descent unit vector: v̂ = −∇f / ‖∇f‖ = [−8, −7]/√113

Why: Point exactly opposite the gradient, normalized to unit length so we read a pure slope.

D_v f = ∇f · v̂ = −(8² + 7²)/√113 = −113/√113 = −√113

Why: The dot of ∇f with its own negative unit vector is −‖∇f‖ — the steepest possible DESCENT, mirror image of the steepest ascent.

\[ D_{-\nabla f}\,f = -\lVert \nabla f \rVert = -\sqrt{113} \approx -10.630 \]

Compare to the level-curve direction

Why: Perpendicular to ∇f is [7, −8]: moving along it, D_v f = (8·7 + 7·(−8))/√113 = 0. Same point, three directions, three signs — max, min, zero.

direction from (1,2)v̂D_v f (verified)
along +∇f (ascent)[0.753, 0.658]+10.630
along −∇f (descent)[−0.753, −0.658]−10.630
⟂ ∇f (level curve)[0.658, −0.753]0.000

41. Gradients sit at right angles to contours

Intuition

A level curve (contour) is the set of points where f has one fixed value. Walk along it and f never changes, so the directional derivative along it is zero — meaning your step is perpendicular to ∇f.

Figure (svg): Three nested oval contour lines with a gradient arrow at a point pointing outward perpendicular to the contour through that point.

∇f (the arrow) crosses every contour at a right angle — uphill, toward higher-value ovals.

This is why loss landscapes are drawn as contour maps with gradient arrows piercing them at 90°: the arrow always points the steepest way off the current level, and descent runs straight back down the arrow.

42. Three identities that power all of ML

Section

Part 3 of 5 — vector-calculus results to memorize

43. The three you must know cold

Concept

Almost every gradient in machine learning reduces to three vector-calculus results. They reappear in least squares, ridge regularization, and attention — so we derive each, then confirm it with autograd.

\[ \frac{d\,\lVert w \rVert^2}{dw} = 2w, \qquad \frac{d\,(w^\top x)}{dw} = x \]

\[ \frac{d\,(x^\top A x)}{dx} = (A + A^\top)\,x \;=\; 2Ax \quad\text{when } A = A^\top \]

44. What has to be given first: Identity 1: ∇‖w‖² = 2w, derived

Missing information

Discussion prompt

‖w‖² is the sum of squared components. Differentiate with respect to one component wₖ, then reassemble the vector:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

The squared norm is just the sum of each coordinate squared — no vectors needed to differentiate it.

45. Identity 1: ∇‖w‖² = 2w, derived

Worked example

‖w‖² is the sum of squared components. Differentiate with respect to one component wₖ, then reassemble the vector:

Write it as a sum: ‖w‖² = Σⱼ wⱼ²

Why: The squared norm is just the sum of each coordinate squared — no vectors needed to differentiate it.

\[ \lVert w \rVert^2 = \sum_j w_j^2 \]

Partial in wₖ: ∂/∂wₖ Σⱼ wⱼ² = 2wₖ

Why: Only the j = k term contains wₖ; every other term is constant in wₖ and drops. That one term gives 2wₖ.

Stack over all k: ∇‖w‖² = 2w

Why: Component k of the gradient is 2wₖ, so the whole gradient vector is 2w. This is the vector version of d(x²)/dx = 2x.

import torch
w = torch.tensor([3.0, -1.0, 2.0], requires_grad=True)
loss = (w**2).sum()   # ||w||^2
loss.backward()
print((2*w).detach())  # hand result 2w
print(w.grad)          # autograd result
sourcevalue at w = [3, −1, 2]
hand: 2w[6.0, −2.0, 4.0]
autograd w.grad[6.0, −2.0, 4.0]
match?yes ✓

46. Inspect it line by line: Identity 1: ∇‖w‖² = 2w, derived

Error analysis

Annotate

Walk the callouts on Identity 1: ∇‖w‖² = 2w, derived. Each one is a place this is easy to get subtly wrong.

  • The squared norm is just the sum of each coordinate squared — no vectors needed to differentiate it.
  • Only the j = k term contains wₖ; every other term is constant in wₖ and drops. That one term gives 2wₖ.
  • Component k of the gradient is 2wₖ, so the whole gradient vector is 2w. This is the vector version of d(x²)/dx = 2x.

47. Plan first: Identity 2: ∇(wᵀx) = x, derived

Step zero

Discussion prompt

Identity 2: ∇(wᵀx) = x, derived — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Write it out: wᵀx = Σⱼ wⱼ xⱼ

Answer:

  1. Write it out: wᵀx = Σⱼ wⱼ xⱼ
  2. Partial in wₖ: ∂/∂wₖ Σⱼ wⱼ xⱼ = xₖ
  3. Stack over k: ∇(wᵀx) = x

48. Identity 2: ∇(wᵀx) = x, derived

Worked example

wᵀx is the dot product Σⱼ wⱼ xⱼ — linear in w. Differentiate in one component again:

Write it out: wᵀx = Σⱼ wⱼ xⱼ

Why: A dot product is a sum of pairwise products. Here x is the constant, w is the variable.

\[ w^\top x = \sum_j w_j x_j \]

Partial in wₖ: ∂/∂wₖ Σⱼ wⱼ xⱼ = xₖ

Why: Only the j = k term has wₖ; its derivative is the constant coefficient xₖ. Everything else drops.

Stack over k: ∇(wᵀx) = x

Why: Component k is xₖ, so the gradient is the vector x itself — the vector version of d(bx)/dx = b.

import torch
w = torch.tensor([1.0, 2.0, 3.0], requires_grad=True)
x = torch.tensor([5.0, -2.0, 4.0])
loss = w @ x     # w^T x
loss.backward()
print(x)         # hand result: the gradient IS x
print(w.grad)    # autograd result
sourcevalue
hand: x[5.0, −2.0, 4.0]
autograd w.grad[5.0, −2.0, 4.0]
match?yes ✓

49. Inspect it line by line: Identity 2: ∇(wᵀx) = x, derived

Error analysis

Annotate

Walk the callouts on Identity 2: ∇(wᵀx) = x, derived. Each one is a place this is easy to get subtly wrong.

  • A dot product is a sum of pairwise products. Here x is the constant, w is the variable.
  • Only the j = k term has wₖ; its derivative is the constant coefficient xₖ. Everything else drops.
  • Component k is xₖ, so the gradient is the vector x itself — the vector version of d(bx)/dx = b.

50. State the rule before it runs: Identity 3: ∇(xᵀAx), the quadratic form

Hypothesis

Predict first

Identity 3: ∇(xᵀAx), the quadratic form is about to be worked. State your hypothesis first: which rule or definition decides this one, and what is the first move it forces? Then watch whether the example agrees with you.

Correct: Partial in xₖ hits the i = k copy and the j = k copy

Why: In the double sum, xₖ shows up wherever i = k (giving Σⱼ Aₖⱼ xⱼ = row k of Ax) and wherever j = k (giving Σᵢ xᵢ Aᵢₖ = row k of Aᵀx).

A hypothesis you wrote down is falsifiable; a vague sense of how it will go is not. If the example opens somewhere else, that gap is the thing worth chasing.

51. Identity 3: ∇(xᵀAx), the quadratic form

Worked example

The quadratic form xᵀAx = Σᵢ Σⱼ xᵢ Aᵢⱼ xⱼ. Here x appears twice, so differentiating hits both copies via the product rule:

Partial in xₖ hits the i = k copy and the j = k copy

Why: In the double sum, xₖ shows up wherever i = k (giving Σⱼ Aₖⱼ xⱼ = row k of Ax) and wherever j = k (giving Σᵢ xᵢ Aᵢₖ = row k of Aᵀx).

Add the two contributions: ∇ = Ax + Aᵀx = (A + Aᵀ)x

Why: One term from each occurrence of x. This is the fully general answer — no symmetry assumed yet.

\[ \frac{d\,(x^\top A x)}{dx} = (A + A^\top)\,x \]

If A is symmetric (A = Aᵀ), it collapses to 2Ax

Why: Then A + Aᵀ = 2A. Loss Hessians and XᵀX are symmetric, so 2Ax is the form you meet in practice.

\[ A = A^\top \;\Longrightarrow\; \frac{d\,(x^\top A x)}{dx} = 2Ax \]

import torch
A = torch.tensor([[2.0, 1.0], [1.0, 3.0]])   # symmetric
x = torch.tensor([1.0, 2.0], requires_grad=True)
loss = x @ (A @ x)   # x^T A x
loss.backward()
print(2 * (A @ x))   # hand result 2Ax
print(x.grad)        # autograd result
quantityvalue (verified)
xᵀAx18.0
hand: 2Ax[8.0, 14.0]
autograd x.grad[8.0, 14.0]
match?yes ✓

52. Decode the notation: Identity 3: ∇(xᵀAx), the quadratic form

Notation

Annotate

From Identity 3: ∇(xᵀAx), the quadratic form — read this one piece at a time. What is each part doing?

On: \( A = A^\top \;\Longrightarrow\; \frac{d\,(x^\top A x)}{dx} = 2Ax \)

  • In the double sum, xₖ shows up wherever i = k (giving Σⱼ Aₖⱼ xⱼ = row k of Ax) and wherever j = k (giving Σᵢ xᵢ Aᵢₖ = row k of Aᵀx).
  • One term from each occurrence of x. This is the fully general answer — no symmetry assumed yet.
  • Then A + Aᵀ = 2A. Loss Hessians and XᵀX are symmetric, so 2Ax is the form you meet in practice.

53. Plan first: All three identities in one real loss

Step zero

Discussion prompt

All three identities in one real loss — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Quadratic term wᵀ(XᵀX)w → 2XᵀXw

Answer:

  1. Quadratic term wᵀ(XᵀX)w → 2XᵀXw
  2. Linear term −2(Xᵀy)ᵀw → −2Xᵀy
  3. Constant yᵀy → 0, then combine: ∇L = 2Xᵀ(Xw − y)

54. All three identities in one real loss

Worked example

Watch the identities combine on the loss you'll meet next lesson: the least-squares loss L = ‖Xw − y‖². Expanded, it is wᵀ(XᵀX)w − 2(Xᵀy)ᵀw + yᵀy — a quadratic form, a linear term, and a constant.

Quadratic term wᵀ(XᵀX)w → 2XᵀXw

Why: Identity 3 with the symmetric A = XᵀX.

Linear term −2(Xᵀy)ᵀw → −2Xᵀy

Why: Identity 2 with x = Xᵀy, times the −2 constant.

Constant yᵀy → 0, then combine: ∇L = 2Xᵀ(Xw − y)

Why: Identity 1's cousin (constants vanish). All three rules fired in one gradient — the exact formula behind every regression fit.

\[ \nabla_w \lVert Xw - y \rVert^2 = 2X^\top(Xw - y) \]

import torch
X = torch.tensor([[1.,1.],[1.,2.],[1.,3.]])
y = torch.tensor([2.,4.,5.])
w = torch.tensor([0.5, 1.0], requires_grad=True)
L = ((X@w - y)**2).sum()
L.backward()
print((2*X.t()@(X@w-y)).detach())  # hand
print(w.grad)                       # autograd
quantityvalue at w = [0.5, 1.0]
residual Xw − y[−0.5, −1.5, −1.5]
L = ‖Xw − y‖²4.75
hand: 2Xᵀ(Xw − y)[−7.0, −16.0]
autograd w.grad[−7.0, −16.0]

55. Which is which, by value at w = [0.5, 1.0]

Discrimination

Sort into buckets

Sort these by value at w = [0.5, 1.0], from memory, without looking back at All three identities in one real loss. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

[−0.5, −1.5, −1.5]
residual Xw − y
4.75
L = ‖Xw − y‖²
[−7.0, −16.0]
hand: 2Xᵀ(Xw − y); autograd w.grad
g1
value at w = [0.5, 1.0] is "[−0.5, −1.5, −1.5]" for residual Xw − y — that is what the table on "All three identities in one real loss" records, and it is the single property separating this group from the rest.
g2
value at w = [0.5, 1.0] is "4.75" for L = ‖Xw − y‖² — that is what the table on "All three identities in one real loss" records, and it is the single property separating this group from the rest.
g3
value at w = [0.5, 1.0] is "[−7.0, −16.0]" for hand: 2Xᵀ(Xw − y), autograd w.grad — that is what the table on "All three identities in one real loss" records, and it is the single property separating this group from the rest.

56. Something is wrong here: the quadratic-form factor of 2

Anomaly

Predict first

A student writes this, and it looks reasonable:

Differentiate xᵀAx. By analogy with the linear term wᵀx → x, guess the gradient is just Ax.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Misses that x appears TWICE in the quadratic form.

x sits on both sides of A, so the product rule yields two terms: Ax + Aᵀx.

Why: Misses that x appears TWICE in the quadratic form. The product rule gives a contribution from EACH copy, so you are short by exactly the Aᵀx term — off by a factor near 2.

57. Trap: the quadratic-form factor of 2

Trap

The trap

Differentiate xᵀAx. By analogy with the linear term wᵀx → x, guess the gradient is just Ax.

d(xᵀAx)/dx = Ax ✗

Why: Misses that x appears TWICE in the quadratic form. The product rule gives a contribution from EACH copy, so you are short by exactly the Aᵀx term — off by a factor near 2.

The fix

x sits on both sides of A, so the product rule yields two terms: Ax + Aᵀx.

d(xᵀAx)/dx = (A + Aᵀ)x = 2Ax for symmetric A ✓

Why: Verified: A = [[2,1],[1,3]], x = [1,2] → 2Ax = [8, 14], matching autograd exactly. Guessing Ax would give [4, 7] — half the true gradient, and training would crawl at half speed.

58. Break it on purpose: the quadratic-form factor of 2

Break the constraint

Discussion prompt

The rule this trap just fixed:

Verified: A = [[2,1],[1,3]], x = [1,2] → 2Ax = [8, 14], matching autograd exactly. Guessing Ax would give [4, 7] — half the true gradient, and training would crawl at half speed.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

Misses that x appears TWICE in the quadratic form. The product rule gives a contribution from EACH copy, so you are short by exactly the Aᵀx term — off by a factor near 2.

59. Autograd: gradients for free

Section

Part 4 of 5 — the engine

60. Picture it first: The computational graph

Picture it

Figure (svg): Three nodes x, x squared, and f connected left to right, labeled forward in one direction and backward for the gradient pass.

Discussion prompt

Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.

Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.

Answer:

Every operation on a tensor with requires_grad=True records a node in a graph, storing the operation's local derivative as the forward pass runs.

61. The computational graph

Concept

Every operation on a tensor with requires_grad=True records a node in a graph, storing the operation's local derivative as the forward pass runs.

.backward() walks that graph backward, multiplying the local derivatives via the chain rule, and deposits ∂(output)/∂(leaf) into each leaf's .grad.

Figure (svg): Three nodes x, x squared, and f connected left to right, labeled forward in one direction and backward for the gradient pass.

62. Leaves, requires_grad, and .grad

Concept

A leaf is a tensor you created directly (not the result of an op). Setting requires_grad=True on a leaf tells autograd to track everything downstream of it and to fill its .grad on backward.

63. Restore the missing line: Autograd reproduces our hand gradient

Fill the middle

Fill in the blanks

From Autograd reproduces our hand gradient — one line has had its right-hand side removed. Put it back.

import torch
x = torch.tensor(1.0, requires_grad=True)
y = torch.tensor(2.0, requires_grad=True)
f = x2 + 3xy + y2
f.backward()
print(f.item(), x.grad.item(), y.grad.item())

Why: f is what everything below it consumes, so the wrong expression here fails later and somewhere else. The scalar f triggers a single reverse traversal; each leaf receives its partial derivative.

64. Autograd reproduces our hand gradient

Worked example

Recompute ∇f at (1, 2) for f = x² + 3xy + y² — but let PyTorch do the calculus, and compare to our by-hand [8, 7]:

import torch
x = torch.tensor(1.0, requires_grad=True)
y = torch.tensor(2.0, requires_grad=True)
f = x**2 + 3*x*y + y**2
f.backward()
print(f.item(), x.grad.item(), y.grad.item())

f.backward() fills x.grad and y.grad in one pass

Why: The scalar f triggers a single reverse traversal; each leaf receives its partial derivative. No formula for the gradient was ever written.

tensorwhat it holdsvalue (verified)
fx² + 3xy + y²11.0
x.grad∂f/∂x = 2x + 3y8.0
y.grad∂f/∂y = 3x + 2y7.0

65. Autograd is just the chain rule, automated

Intuition

There is no magic and no symbolic algebra. As the forward pass runs, PyTorch stores the local derivative of each operation it performs.

On .backward() it multiplies those stored local derivatives together, from output back to input — exactly the chain rule you'll do by hand in Week 14's backprop. autograd just never forgets a factor.

66. What has to happen first: The chain rule autograd runs, by hand

Ranking

Put in order

Put the moves of The chain rule autograd runs, by hand into the order they have to happen.

  1. Inner: u = x² + 1, so du/dx = 2x
  2. Outer: f = u³, so df/du = 3u²
  3. Chain them: df/dx = (df/du)(du/dx) = 3u² · 2x = 75 · 4 = 300

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. The first recorded node. At x=2, u = 4 + 1 = 5 and du/dx = 4.

67. The chain rule autograd runs, by hand

Worked example

See the chain rule autograd applies. Take f = (x² + 1)³ at x = 2. Split it into an inner and outer step, exactly as PyTorch records nodes:

Inner: u = x² + 1, so du/dx = 2x

Why: The first recorded node. At x=2, u = 4 + 1 = 5 and du/dx = 4.

Outer: f = u³, so df/du = 3u²

Why: The second node. At u=5, df/du = 3·25 = 75.

Chain them: df/dx = (df/du)(du/dx) = 3u² · 2x = 75 · 4 = 300

Why: The reverse pass multiplies the stored local derivatives from output back to input — that product IS the chain rule.

\[ \frac{df}{dx} = 3u^2 \cdot 2x = 3(5)^2 \cdot 2(2) = 300 \]

import torch
x = torch.tensor(2.0, requires_grad=True)
u = x**2 + 1     # inner node
f = u**3         # outer node
f.backward()
print(f.item(), x.grad.item())
quantityhand valueautograd
u = x² + 15.05.0
f = u³125.0125.0
df/dx = 3u²·2x300.0300.0

68. Fill in: hand value for The chain rule autograd runs, by hand

Comparison

Comparison matrix

From The chain rule autograd runs, by hand: refill the hand value column from what you know. The rest of the table is as it appeared.

quantityhand valueautograd
u = x² + 15.05.0
f = u³125.0125.0
df/dx = 3u²·2x300.0300.0

69. Second order: the Hessian (preview)

Concept

Differentiate the gradient again and you get the Hessian — the matrix of second partials. For our f = x² + 3xy + y² it is constant:

\[ \nabla^2 f = \begin{bmatrix} \partial^2 f/\partial x^2 & \partial^2 f/\partial x\,\partial y \\ \partial^2 f/\partial y\,\partial x & \partial^2 f/\partial y^2 \end{bmatrix} = \begin{bmatrix} 2 & 3 \\ 3 & 2 \end{bmatrix} \]

The Hessian is symmetric (mixed partials are equal, ∂²f/∂x∂y = ∂²f/∂y∂x = 3) — which is exactly why the ∇(xᵀAx) = 2Ax identity shows up so often: loss curvature is a symmetric quadratic form. You'll use the Hessian directly in Lesson 7's convexity argument.

70. The gradient descent update

Concept

Given any differentiable loss, the recipe is fixed: compute the gradient, then nudge each parameter against it.

\[ \theta \leftarrow \theta - \eta\,\nabla_\theta L \]

η (the learning rate) sets the step size. Too small crawls; too large overshoots and diverges. Below we'll use η = 0.1 and watch the parameters march toward the minimum.

71. Restore the missing line: Descending a loss bowl with autograd

Fill the middle

Fill in the blanks

From Descending a loss bowl with autograd — one line has had its right-hand side removed. Put it back.

import torch
w = torch.tensor(0.0, requires_grad=True)
b = torch.tensor(0.0, requires_grad=True)
lr = 0.1
for step in range(5):
loss = (w-3)2 + (b+1)2
loss.backward()
with torch.no_grad():
w -= lr * w.grad
b -= lr * b.grad
w.grad.zero_(); b.grad.zero_()

Why: loss is what everything below it consumes, so the wrong expression here fails later and somewhere else. This four-line rhythm is identical for a 2-parameter bowl and a 70-billion-parameter model.

72. Descending a loss bowl with autograd

Worked example

Minimize f(w, b) = (w−3)² + (b+1)² from (0, 0). The minimum is obviously (3, −1) — watch the steps march toward it, five passes, seed set:

import torch
w = torch.tensor(0.0, requires_grad=True)
b = torch.tensor(0.0, requires_grad=True)
lr = 0.1
for step in range(5):
    loss = (w-3)**2 + (b+1)**2
    loss.backward()
    with torch.no_grad():
        w -= lr * w.grad
        b -= lr * b.grad
        w.grad.zero_(); b.grad.zero_()

Each pass: forward loss → backward → step under no_grad → zero the grads

Why: This four-line rhythm is identical for a 2-parameter bowl and a 70-billion-parameter model. Learn it here at a size you can trace by hand.

stepw→b→loss∂w∂b
10.600−0.20010.000−6.002.00
21.080−0.3606.400−4.801.60
31.464−0.4884.096−3.841.28
41.771−0.5902.621−3.071.02
52.017−0.6721.678−2.460.82

73. Watch it run: Descending a loss bowl with autograd

Pattern

Step through it

Step through Descending a loss bowl with autograd one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: step is 1
  2. Step 2: step is 2
  3. Step 3: step is 3
  4. Step 4: step is 4
  5. Step 5: step is 5

74. The graph is rebuilt every iteration

Intuition

Notice loss = (w-3)**2 + (b+1)**2 sits inside the loop. Each pass builds a fresh computational graph from the current w, b. PyTorch's graph is dynamic — it exists only for one forward/backward cycle, then is discarded.

That's why calling .backward() twice on the same graph errors (the buffers are freed after the first), but calling it once per freshly-built loss is fine. Rebuild, backward, step, zero — every iteration, from scratch.

75. no_grad also speeds up inference

Concept

torch.no_grad() isn't only for the update. At prediction time you never call .backward(), so tracking the graph is wasted memory and time. Wrapping inference in no_grad() turns tracking off entirely.

Under the guard, results carry requires_grad = False — no graph is built at all. Same numbers, less overhead. It's the standard wrapper around any evaluation or deployment code path.

76. Why the update lives under no_grad()

Concept

The update w -= lr * w.grad is itself a tensor operation on a leaf that requires grad. Left unguarded, PyTorch either records it into next step's graph or refuses the in-place edit outright.

with torch.no_grad(): says 'don't track this'. The weights change, but the change is not part of any future derivative — which is exactly what a parameter update should be.

In fact, editing a grad-requiring leaf in place without no_grad() raises a RuntimeError — the next slide shows the exact message.

77. Teach it back: Why the update lives under no_grad()

Explain it

Discussion prompt

Explain Why the update lives under no_grad() to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

The update w -= lr * w.grad is itself a tensor operation on a leaf that requires grad. Left unguarded, PyTorch either records it into next step's graph or refuses the in-place edit outright.

78. Guess the shape of the answer: The no_grad error, live

Estimation

Predict first

Try the update in place without the guard, then with it. The first raises; the second works:

Commit before you compute: what does The no_grad error, live come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: The unguarded in-place update raises RuntimeError

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. PyTorch forbids in-place edits of a leaf that requires grad, because it would corrupt the recorded graph.

79. The no_grad error, live

Worked example

Try the update in place without the guard, then with it. The first raises; the second works:

import torch
w = torch.tensor(1.0, requires_grad=True)
try:
    w -= 0.1 * torch.tensor(2.0)   # in-place on a grad leaf
except RuntimeError as e:
    print('ERROR:', str(e)[:55])
with torch.no_grad():
    w -= 0.1 * torch.tensor(2.0)   # now allowed
print('ok, w =', round(w.item(), 3))

The unguarded in-place update raises RuntimeError

Why: PyTorch forbids in-place edits of a leaf that requires grad, because it would corrupt the recorded graph. The guard removes the tensor from tracking, so the edit is legal.

code pathresult (verified)
w -= ... (no guard)RuntimeError: a leaf Variable that requires grad…
with torch.no_grad(): w -= ...ok, w = 0.8

80. Something is wrong here: forgetting zero_grad()

Anomaly

Predict first

A student writes this, and it looks reasonable:

.grad holds this step's gradient, so just call .backward() each loop and read it — no cleanup needed.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: PyTorch ACCUMULATES (adds) into .grad.

.grad is a running sum — you must reset it to zero every iteration.

Why: PyTorch ACCUMULATES (adds) into .grad. The true gradient is −6.0, but the stale −6.0 from the previous call was summed on top, giving −12.0. Steps get bigger every iteration and training silently corrupts.

81. Trap: forgetting zero_grad()

Trap

The trap

.grad holds this step's gradient, so just call .backward() each loop and read it — no cleanup needed.

Loop without zeroing: on (w−5)² at w=2, backward twice → w.grad reads −12.0 ✗

Why: PyTorch ACCUMULATES (adds) into .grad. The true gradient is −6.0, but the stale −6.0 from the previous call was summed on top, giving −12.0. Steps get bigger every iteration and training silently corrupts.

The fix

.grad is a running sum — you must reset it to zero every iteration.

Call w.grad.zero_() after each step → w.grad correctly reads −6.0 ✓

Why: Verified: with zeroing the gradient stays −6.0 each pass; without it, it grows −6, −12, −18… The accumulation is a feature (it lets you sum grads across mini-batches) but only if you zero it when you don't want it.

82. Which of these survive contact with Lesson 6: Gradient Intuition & PyTorch…?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Five parts, each building on the last. We start from a single slope and end holding a working optimizer you wrote yourself.; One function carries the whole first half of this lesson. It takes two inputs and returns one number — a surface floating over the (x, y) plane.; The partial derivative ∂f/∂x is the slope you feel walking in the +x direction only — y held fixed, as if it were a constant.
Breaks
Goal: minimize the loss, so the gradient must point downhill toward the minimum — follow it.; Differentiate xᵀAx. By analogy with the linear term wᵀx → x, guess the gradient is just Ax.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 6: Gradient Intuition & PyTorch Autograd puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

83. Rebuild the recipe: The autograd training loop — the recipe

Ranking

Put in order

These are the steps of The autograd training loop — the recipe, scrambled. Put them back in order before the next slide shows you.

  1. Create parameters with requires_grad=True
  2. Forward: compute the scalar loss from the parameters
  3. Backward: loss.backward() fills every .grad
  4. Update under torch.no_grad(): θ -= lr * θ.grad
  5. Zero the gradients: θ.grad.zero_() — or the next step's gradient is corrupted
  6. Repeat until the loss stops improving

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

84. The autograd training loop — the recipe

Pattern

  1. Create parameters with requires_grad=True
  2. Forward: compute the scalar loss from the parameters
  3. Backward: loss.backward() fills every .grad
  4. Update under torch.no_grad(): θ -= lr * θ.grad
  5. Zero the gradients: θ.grad.zero_() — or the next step's gradient is corrupted
  6. Repeat until the loss stops improving

85. Where does it stop working: The autograd training loop — the recipe

Edge cases

Discussion prompt

The autograd training loop — the recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Create parameters with requires_grad=True
  2. Forward: compute the scalar loss from the parameters
  3. Backward: loss.backward() fills every .grad
  4. Update under torch.no_grad(): θ -= lr * θ.grad
  5. Zero the gradients: θ.grad.zero_() — or the next step's gradient is corrupted
  6. Repeat until the loss stops improving

86. The same loop, at every scale

Intuition

Everything so far — two-variable gradients, the three identities, the four-line loop — is exactly what a modern network runs, just with more parameters. A transformer's loss.backward() fills millions of .grad entries by the same chain rule you traced by hand.

Nothing conceptually new is added when you scale up. Master the 2-parameter bowl and you understand the 70-billion-parameter one. The checks below make sure the fundamentals are solid before you trust them at scale.

87. Rule out three: Check yourself — partial derivatives

Elimination

Eliminate the wrong options

For f(x, y) = x²y + y³, what is ∂f/∂x at (2, 3)?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. 12
  • B. 31
  • C. 4
  • D. 27

Survives elimination: A

Why: Treat y as a constant: ∂f/∂x = 2xy. At (2, 3): 2·2·3 = 12. The y³ term has no x, so it drops to 0.

88. Check yourself — partial derivatives

Check

Solve it on paper before clicking. Freeze the other variable.

Check your understanding

For f(x, y) = x²y + y³, what is ∂f/∂x at (2, 3)?

  • A. 12 (correct)
  • B. 31
  • C. 4
  • D. 27

Answer: A

Why: Treat y as a constant: ∂f/∂x = 2xy. At (2, 3): 2·2·3 = 12. The y³ term has no x, so it drops to 0.

Why B tempts people
This is ∂f/∂y = x² + 3y² = 4 + 27 = 31 — differentiated with respect to the WRONG variable (y instead of x).
Why C tempts people
Treated y as if it were the constant 1 instead of its value 3, giving 2x = 4. The frozen variable keeps its actual value.
Why D tempts people
This is just the y³ term's y-derivative (3y² = 27) — differentiated the wrong term and the wrong variable.

89. Answer it before you see the options: Check yourself — the descent direction

Prediction

Predict first

You are minimizing a loss L(θ). At the current θ, ∇L = [4, −2]. Which step decreases L fastest for a small fixed step size?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: Move along [−4, 2]

Why: Steepest descent is the negative gradient: −∇L = [−4, 2]. The gradient points to steepest ASCENT, so flip its sign to go downhill fastest.

90. Check yourself — the descent direction

Check

Think about the sign, and about perpendicularity, before you pick.

Check your understanding

You are minimizing a loss L(θ). At the current θ, ∇L = [4, −2]. Which step decreases L fastest for a small fixed step size?

  • A. Move along [−4, 2] (correct)
  • B. Move along [4, −2]
  • C. Move along [2, 4]
  • D. Move along [−2, 4]

Answer: A

Why: Steepest descent is the negative gradient: −∇L = [−4, 2]. The gradient points to steepest ASCENT, so flip its sign to go downhill fastest.

Why B tempts people
This is +∇L itself — the direction of steepest ASCENT. It increases L fastest, the exact opposite of minimization.
Why C tempts people
This is perpendicular to ∇L (dot product 4·2 + (−2)·4 = 0), so it's along a level curve — to first order it doesn't change L at all, let alone decrease it fastest.
Why D tempts people
Swapped the components of −∇L = [−4, 2] into [−2, 4]. It's neither the descent direction nor perpendicular — a plausible-looking rearrangement that points the wrong way.

91. How sure are you: Check yourself — the directional derivative

Commit first

Predict first

At a point, ∇f = [3, 4]. What is the directional derivative of f in the direction v = [4, 3]?

Commit to an answer, then rate it — certain, fairly sure, or guessing — and write the rating down before you turn the page.

Correct: 24/5 = 4.8

Why: Normalize v: ‖v‖ = √(16+9) = 5, so v̂ = [0.8, 0.6]. Then D_v f = ∇f · v̂ = 3·0.8 + 4·0.6 = 2.4 + 2.4 = 4.8.

The rating matters as much as the answer: confident-and-wrong is the combination that survives revision, because nothing about it feels like it needs revisiting.

92. Check yourself — the directional derivative

Check

Remember to normalize the direction first.

Check your understanding

At a point, ∇f = [3, 4]. What is the directional derivative of f in the direction v = [4, 3]?

  • A. 24/5 = 4.8 (correct)
  • B. 24
  • C. 5
  • D. 7

Answer: A

Why: Normalize v: ‖v‖ = √(16+9) = 5, so v̂ = [0.8, 0.6]. Then D_v f = ∇f · v̂ = 3·0.8 + 4·0.6 = 2.4 + 2.4 = 4.8.

Why B tempts people
Dotted with the raw v = [4, 3] (3·4 + 4·3 = 24) WITHOUT normalizing. That mixes the slope with the length of v — you must divide by ‖v‖ = 5.
Why C tempts people
This is ‖∇f‖ = √(9+16) = 5, the MAXIMUM directional derivative (along ∇f itself). But v = [4, 3] is not aligned with ∇f = [3, 4], so the actual rate is less.
Why D tempts people
Added the components of ∇f (3 + 4 = 7) — not a dot product at all, and no normalization.

93. Answer it before you see the options: Check yourself — the quadratic-form…

Prediction

Predict first

For symmetric A = [[2, 1], [1, 3]] and x = [1, 2], what is d(xᵀAx)/dx?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: [8, 14]

Why: For symmetric A the gradient is 2Ax. Ax = [2·1 + 1·2, 1·1 + 3·2] = [4, 7], so 2Ax = [8, 14]. Confirmed by autograd.

94. Check yourself — the quadratic-form gradient

Check

Count how many times x appears. That's the whole question.

Check your understanding

For symmetric A = [[2, 1], [1, 3]] and x = [1, 2], what is d(xᵀAx)/dx?

  • A. [8, 14] (correct)
  • B. [4, 7]
  • C. [3, 4]
  • D. [2, 6]

Answer: A

Why: For symmetric A the gradient is 2Ax. Ax = [2·1 + 1·2, 1·1 + 3·2] = [4, 7], so 2Ax = [8, 14]. Confirmed by autograd.

Why B tempts people
This is Ax without the factor of 2 — the classic 'forgot x appears twice' mistake. The product rule doubles it to 2Ax = [8, 14].
Why C tempts people
This is Ax + x componentwise or a similar slip — it uses neither the doubling nor the full rows of A. The correct product Ax = [4, 7] uses whole rows, then doubles to [8, 14].
Why D tempts people
This is the diagonal of A applied to x ([2·1, 3·2] = [2, 6]) — ignoring the off-diagonal 1s that couple the two components. The full matrix product Ax uses entire rows, not just diagonals.

95. Rule out three: Check yourself — the accumulation gotcha

Elimination

Eliminate the wrong options

w = torch.tensor(2.0, requires_grad=True). You run loss=(w-5)**2; loss.backward() TWICE in a row with no zero_grad() between. What does w.grad read?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. −12.0
  • B. −6.0
  • C. 0.0
  • D. An error — backward can only be called once

Survives elimination: A

Why: Each backward computes ∂loss/∂w = 2(w−5) = 2(2−5) = −6, and PyTorch ADDS it into the existing .grad. Two calls → −6 + −6 = −12.0.

96. Check yourself — the accumulation gotcha

Check

This one bites everyone once. Trace it carefully.

Check your understanding

w = torch.tensor(2.0, requires_grad=True). You run loss=(w-5)**2; loss.backward() TWICE in a row with no zero_grad() between. What does w.grad read?

  • A. −12.0 (correct)
  • B. −6.0
  • C. 0.0
  • D. An error — backward can only be called once

Answer: A

Why: Each backward computes ∂loss/∂w = 2(w−5) = 2(2−5) = −6, and PyTorch ADDS it into the existing .grad. Two calls → −6 + −6 = −12.0.

Why B tempts people
This is the gradient from a SINGLE backward pass. It's correct only if you zero_grad() between calls — which the question explicitly skips.
Why C tempts people
.grad starts at None, not 0, and it accumulates; it is never reset to zero unless you call zero_grad() yourself.
Why D tempts people
Re-running the forward line (loss=(w-5)**2) builds a fresh graph each time, so a second backward is legal here. The 'call backward once' error only appears if you reuse the SAME graph without retain_graph=True.

97. Your turn: build gradient descent

Section

Part 5 of 5 — the project

98. Project: gradient descent from scratch

Concept

Minimize f(w) = (w − 5)² two ways — hand-coded calculus, then verified against autograd. The minimum is obviously w = 5; you'll watch your loop crawl toward it.

#requirementtool
1Define f(w) and its analytic gradient 2(w−5)a Python function
2Loop the update w ← w − 0.1·grad from w = 0a for-loop
3Confirm one gradient against torch autogradrequires_grad / .backward()

Build rules: type every line yourself, run after each line, and when something errors, read the error — don't delete it. The struggle is the encoding.

99. By analogy: Project: gradient descent from scratch

Analogy

Discussion prompt

Explain Project: gradient descent from scratch by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Minimize f(w) = (w − 5)² two ways — hand-coded calculus, then verified against autograd. The minimum is obviously w = 5; you'll watch your loop crawl toward it.

100. Milestone 1 — the function and its gradient

Worked example

Your turn: write f(w) and a separate grad(w). Say out loud what grad(0) should be before you run it.

Hint: f(w) = (w-5)**2, and by the chain rule its derivative is 2*(w-5) — no PyTorch yet, just arithmetic.

def f(w):
    return (w - 5)**2

def grad(w):
    return 2 * (w - 5)

print(f(0), grad(0))

Predict then check: grad(0) = 2(0−5) = −10

Why: At w=0 the slope is steep and negative — f falls as w increases, so descent will push w UP toward 5. The sign tells you the direction before you loop.

callvalue (verified)
f(0) = (0−5)²25
grad(0) = 2(0−5)−10

101. Work backwards from the answer: Milestone 1 — the function and its gradient

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Predict then check: grad(0) = 2(0−5) = −10

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Your turn: write f(w) and a separate grad(w). Say out loud what grad(0) should be before you run it.

102. Predict the next row: Milestone 2 — the descent loop

Pattern

Predict first

The table runs: 1 | 0.000 | −10.000 | 1.0000 | 16.00000 · 2 | 1.000 | −8.000 | 1.8000 | 10.24000 · 3 | 1.800 | −6.400 | 2.4400 | 6.55360 · 4 | 2.440 | −5.120 | 2.9520 | 4.19430 · 5 | 2.952 | −4.096 | 3.3616 | 2.68435 · 6 | 3.362 | −3.277 | 3.6893 | 1.71799

In Milestone 2 — the descent loop, given the rows so far: what is the next one — the row where step is 7?

Correct: 7 | 3.689 | −2.621 | 3.9514 | 1.09951

stepw_startgradw_newf(w_new)
10.000−10.0001.000016.00000
21.000−8.0001.800010.24000
31.800−6.4002.44006.55360
42.440−5.1202.95204.19430
52.952−4.0963.36162.68435
63.362−3.2773.68931.71799
73.689−2.6213.95141.09951

Why: The relationship between the columns, not the individual numbers, is what generates the next row. grad(w) = 2(w−5) → 0 as w → 5, so each step is smaller than the last.

103. Milestone 2 — the descent loop

Worked example

Your turn: loop 7 times from w = 0 with lr = 0.1, updating w each pass. Predict: does w reach exactly 5, or only approach it?

Hint: inside the loop compute g = grad(w), then w = w - lr * g. Print w and f(w) each pass so you can watch the loss shrink.

def f(w):    return (w - 5)**2
def grad(w): return 2 * (w - 5)

w, lr = 0.0, 0.1
for step in range(1, 8):
    g = grad(w)
    w = w - lr * g
    print(step, round(w, 4), round(f(w), 5))

The steps SHRINK as w nears 5

Why: grad(w) = 2(w−5) → 0 as w → 5, so each step is smaller than the last. Gradient descent approaches the minimum but (with fixed lr) never lands exactly on it — it converges geometrically.

stepw_startgradw_newf(w_new)
10.000−10.0001.000016.00000
21.000−8.0001.800010.24000
31.800−6.4002.44006.55360
42.440−5.1202.95204.19430
52.952−4.0963.36162.68435
63.362−3.2773.68931.71799
73.689−2.6213.95141.09951

104. Watch it run: Milestone 2 — the descent loop

Pattern

Step through it

Step through Milestone 2 — the descent loop one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: step is 1
  2. Step 2: step is 2
  3. Step 3: step is 3
  4. Step 4: step is 4
  5. Step 5: step is 5
  6. Step 6: step is 6
  7. Step 7: step is 7

105. What has to be given first: Milestone 3 — verify against autograd

Missing information

Discussion prompt

Your turn: at w = 2, confirm your hand gradient grad(2) equals what PyTorch computes. Predict the number first (it's 2(2−5)).

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

grad(2) = 2(2−5) = −6, and PyTorch's reverse pass computes the identical −6.0. That agreement is the whole point: you understand what autograd is doing.

106. Milestone 3 — verify against autograd

Worked example

Your turn: at w = 2, confirm your hand gradient grad(2) equals what PyTorch computes. Predict the number first (it's 2(2−5)).

Hint: make wt = torch.tensor(2.0, requires_grad=True), build loss = (wt-5)**2, call loss.backward(), then read wt.grad.

import torch
def grad(w): return 2 * (w - 5)

wt = torch.tensor(2.0, requires_grad=True)
loss = (wt - 5)**2
loss.backward()
print(grad(2), wt.grad.item())

Both read −6.0 — your calculus and autograd agree

Why: grad(2) = 2(2−5) = −6, and PyTorch's reverse pass computes the identical −6.0. That agreement is the whole point: you understand what autograd is doing.

sourcegradient at w = 2
your grad(2) = 2(2−5)−6.0
wt.grad (autograd)−6.0
match?yes ✓

107. What each one costs: Milestone 3 — verify against autograd

Trade off

Comparison matrix

From Milestone 3 — verify against autograd: every row here is a choice with a cost. Fill the gradient at w = 2 column, then say which row you would actually pick and what you give up for it.

sourcegradient at w = 2
your grad(2) = 2(2−5)−6.0
wt.grad (autograd)−6.0
match?yes ✓

108. Guess the shape of the answer: The full program

Estimation

Predict first

All three milestones assembled into one runnable script:

Commit before you compute: what does The full program come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Success criterion: w marches 0 → ~3.95, and the check line prints two matching −6.0

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. If your run matches the table below, you built gradient descent from the math up and proved it against PyTorch's engine.

109. The full program

Worked example

All three milestones assembled into one runnable script:

import torch

def f(w):    return (w - 5)**2
def grad(w): return 2 * (w - 5)

w, lr = 0.0, 0.1
for step in range(1, 8):
    g = grad(w)
    w = w - lr * g
    print(step, round(w, 4), round(f(w), 5))

# autograd cross-check at w = 2
wt = torch.tensor(2.0, requires_grad=True)
((wt - 5)**2).backward()
print('check:', grad(2), wt.grad.item())

Success criterion: w marches 0 → ~3.95, and the check line prints two matching −6.0

Why: If your run matches the table below, you built gradient descent from the math up and proved it against PyTorch's engine.

printed linevalue (verified)
step 11.0 16.0
step 42.952 4.1943
step 73.9514 1.09951
check−6.0 −6.0

110. Fill in: value (verified) for The full program

Comparison

Comparison matrix

From The full program: refill the value (verified) column from what you know. The rest of the table is as it appeared.

printed linevalue (verified)
step 11.0 16.0
step 42.952 4.1943
step 73.9514 1.09951
check−6.0 −6.0

111. Show it off

Concept

Slides closed, out loud: explain (1) why grad(0) is negative, (2) why the steps shrink as w nears 5, and (3) what zero_grad() would prevent if you looped the autograd version.

Stretch (the homework CHALLENGE): the softmax Jacobian ∂sᵢ/∂xⱼ = sᵢ(δᵢⱼ − sⱼ). For x = [1, 2, 3], softmax = [0.090, 0.245, 0.665] — this exact derivative drives backprop through every classifier head. Build the 3×3 Jacobian and check it against torch.autograd.

112. Break it if you can: Show it off

Counterexample

Discussion prompt

Slides closed, out loud: explain (1) why grad(0) is negative, (2) why the steps shrink as w nears 5, and (3) what zero_grad() would prevent if you looped the autograd version.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

113. Connect it up: Lesson 6: Gradient Intuition & PyTorch Autograd

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — From one slope to many · The directional derivative · Three identities that power all of ML · Autograd: gradients for free · Your turn: build gradient descent. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

114. What you can do now

Recap

movethe one thing to remember
partialfreeze the other variables
gradient signpoints uphill — SUBTRACT it to train
directional derivdot with the UNIT vector; zero along a level curve
xᵀAxfactor of 2 (symmetric A): 2Ax, not Ax
.gradaccumulates — zero it every step; update under no_grad()

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 6 (Week 2 — Gradient Intuition) — Barron · USAAIO Round 2 Preparation, 2026
  2. PyTorch Autograd mechanics
  3. torch.autograd — automatic differentiation
  4. Gilbert Strang, Calculus, Ch. 13 (Partial Derivatives, Gradient, Directional Derivative) — Wellesley-Cambridge Press
  5. Every trace table, gradient, and coefficient produced by real execution — torch 2.7.1 + numpy 2.2.6, seeds set, verification run July 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108