USAAIO Lesson 6, from Week 2 on calculus, fully worked. It builds partial derivatives one axis at a time, assembles the gradient and reads it as a steepest-ascent compass, and derives the directional derivative from the projection. It then derives the three ML gradient identities - 2w, x, and 2Ax - and checks each one against autograd, covers the computational graph and the chain rule behind .backward(), and shows the accumulation and no_grad traps in real execution. It ends with a from-scratch gradient-descent project verified line by line against PyTorch. Every snippet runs as written with the seeds set, and every number in a trace was produced by real execution. The lesson runs to 60 slides.
Subject: Machine Learning · 114 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 6 · Week 2 (Calculus)
The single object every optimizer, every backprop step, and every transformer gradient is built on. We derive the gradient from one partial at a time, prove the three identities that power all of ML, then let autograd reproduce every number — and end by building gradient descent from scratch.
Objectives
∇f points in the direction of steepest ascent, and take a correct descent stepD_v f = ∇f · v̂ and read off its max, min, and zero directions∇‖w‖² = 2w, ∇(wᵀx) = x, ∇(xᵀAx) = 2Ax, each checked against autogradrequires_grad, .backward(), .grad, and torch.no_grad() — and explain why the update needs zero_grad()f(w) = (w−5)² and verify every gradient against PyTorchConcept
Five parts, each building on the last. We start from a single slope and end holding a working optimizer you wrote yourself.
Matching
Match the pairs
From The road ahead — match each one to what it actually does. The descriptions have been shuffled.
Why: 1. The gradient, 3. Identities, 4. Autograd, 5. Build it are easy to tell apart while they are sitting next to their descriptions and much harder afterwards, which is what this checks.
Section
Part 1 of 5 — partial derivatives
Concept
One function carries the whole first half of this lesson. It takes two inputs and returns one number — a surface floating over the (x, y) plane.
\[ f(x,y) = x^2 + 3xy + y^2 \]
In 1-D a derivative is a single slope. Here there is no single slope — the surface tilts differently depending on which way you walk. We need a slope per direction, starting with the two axes.
Counterexample
Discussion prompt
One function carries the whole first half of this lesson. It takes two inputs and returns one number — a surface floating over the (x, y) plane.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
In 1-D a derivative is a single slope. Here there is no single slope — the surface tilts differently depending on which way you walk. We need a slope per direction, starting with the two axes.
Concept
The partial derivative ∂f/∂x is the slope you feel walking in the +x direction only — y held fixed, as if it were a constant.
Mechanically: to take ∂f/∂xᵢ, differentiate with respect to xᵢ and treat every other variable as a frozen number. Nothing else about differentiation changes.
\[ \frac{\partial f}{\partial x} = \lim_{h\to 0}\frac{f(x+h,\,y) - f(x,\,y)}{h} \]
Analogy
Discussion prompt
Explain What a partial derivative means by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
The partial derivative ∂f/∂x is the slope you feel walking in the +x direction only — y held fixed, as if it were a constant.
Intuition
Holding y fixed is like cutting the surface with a vertical knife parallel to the x-axis. You expose a single curved slice — an ordinary 1-D curve.
∂f/∂x is just the everyday slope of that slice. The multivariable part is only bookkeeping: which knife you're using. The calculus underneath is the same one-variable calculus you already know.
So there is nothing new to learn to differentiate — only something to hold still.
Explain it
Discussion prompt
Explain Freezing y is slicing the surface to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Holding y fixed is like cutting the surface with a vertical knife parallel to the x-axis. You expose a single curved slice — an ordinary 1-D curve.
Ranking
Put in order
Put the moves of Partial in x, term by term into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Ordinary power rule in x: d(x²)/dx = 2x.
Worked example
Differentiate f = x² + 3xy + y² with respect to x, treating y as a constant. Take one term at a time — no shortcuts:
Term x² → 2x
Why: Ordinary power rule in x: d(x²)/dx = 2x. No y here to worry about.
Term 3xy → 3y
Why: 3y is a constant coefficient on x, so d(3y·x)/dx = 3y. The x differentiates away, the frozen y stays.
Term y² → 0
Why: y² has no x in it at all, so as a function of x it is a constant — and the derivative of a constant is 0.
Add the three → ∂f/∂x = 2x + 3y
Why: Sum the term-by-term results. This is the slope of the x-slice at any point (x, y).
\[ \frac{\partial f}{\partial x} = 2x + 3y \]
Notation
Annotate
From Partial in x, term by term — read this one piece at a time. What is each part doing?
On: \( \frac{\partial f}{\partial x} = 2x + 3y \)
Step zero
Discussion prompt
Partial in y, term by term — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Term x² → 0
Answer:
Worked example
Now freeze x and differentiate with respect to y. Same function, other knife:
Term x² → 0
Why: x² has no y, so it is constant in y and drops to 0.
Term 3xy → 3x
Why: 3x is the constant coefficient on y, so d(3x·y)/dy = 3x.
Term y² → 2y
Why: Ordinary power rule in y: d(y²)/dy = 2y.
Add the three → ∂f/∂y = 3x + 2y
Why: The other partial. Notice the two partials are DIFFERENT functions — the surface really does tilt differently along each axis.
\[ \frac{\partial f}{\partial y} = 3x + 2y \]
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Add the three → ∂f/∂y = 3x + 2y
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Now freeze x and differentiate with respect to y. Same function, other knife:
Concept
Collect the two partials into a single vector. That vector is the gradient, written ∇f. It is not a number — it is an arrow living in the input space.
\[ \nabla f = \begin{bmatrix} \partial f / \partial x \\[2pt] \partial f / \partial y \end{bmatrix} = \begin{bmatrix} 2x + 3y \\ 3x + 2y \end{bmatrix} \]
gradient ∇f — The vector of all first partial derivatives of a scalar function. For f: ℝⁿ → ℝ it is an n-vector. It is a FUNCTION of position — a different arrow at every point of the input space.
Estimation
Predict first
The gradient is a formula in x and y. To get an actual arrow, plug in a point. Use (1, 2):
Commit before you compute: what does Evaluate the gradient at (1, 2) come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: ∇f(1,2) = [8, 7]
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The arrow at (1,2). Both components positive, so f increases as you move in +x and +y from this point.
Worked example
The gradient is a formula in x and y. To get an actual arrow, plug in a point. Use (1, 2):
∂f/∂x at (1,2) = 2(1) + 3(2) = 2 + 6 = 8
Why: Substitute x=1, y=2 into 2x+3y.
∂f/∂y at (1,2) = 3(1) + 2(2) = 3 + 4 = 7
Why: Substitute x=1, y=2 into 3x+2y.
∇f(1,2) = [8, 7]
Why: The arrow at (1,2). Both components positive, so f increases as you move in +x and +y from this point.
\[ \nabla f(1,2) = \begin{bmatrix} 8 \\ 7 \end{bmatrix} \]
| component | formula | value at (1,2) |
|---|---|---|
| ∂f/∂x | 2x + 3y | 8 |
| ∂f/∂y | 3x + 2y | 7 |
| f itself | x² + 3xy + y² | 1 + 6 + 4 = 11 |
Comparison
Comparison matrix
From Evaluate the gradient at (1, 2): refill the formula column from what you know. The rest of the table is as it appeared.
| component | formula | value at (1,2) |
|---|---|---|
| ∂f/∂x | 2x + 3y | 8 |
| ∂f/∂y | 3x + 2y | 7 |
| f itself | x² + 3xy + y² | 1 + 6 + 4 = 11 |
Step zero
Discussion prompt
A second gradient, with a mixed term — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: ∂g/∂x: freeze y, differentiate x²y + y³ in x → 2xy + 0 = 2xy
Answer:
Worked example
One more, because the exam loves functions where a variable hides inside a product. Take g(x, y) = x²y + y³ and find ∇g at (2, 3):
∂g/∂x: freeze y, differentiate x²y + y³ in x → 2xy + 0 = 2xy
Why: x²y has coefficient y on x², giving 2xy. y³ has no x and drops.
∂g/∂y: freeze x, differentiate x²y + y³ in y → x² + 3y²
Why: x²y has coefficient x² on y, giving x². y³ gives 3y².
At (2, 3): ∇g = [2·2·3, 2² + 3·3²] = [12, 31]
Why: Substitute. The two partials are wildly different sizes — a reminder that the gradient's DIRECTION matters, not just one component.
\[ \nabla g(2,3) = \begin{bmatrix} 2xy \\ x^2 + 3y^2 \end{bmatrix}_{(2,3)} = \begin{bmatrix} 12 \\ 31 \end{bmatrix} \]
import torch
x = torch.tensor(2.0, requires_grad=True)
y = torch.tensor(3.0, requires_grad=True)
g = x**2 * y + y**3
g.backward()
print(x.grad.item(), y.grad.item())| partial | formula | value (verified) |
|---|---|---|
| ∂g/∂x | 2xy | 12.0 |
| ∂g/∂y | x² + 3y² | 31.0 |
Trade off
Comparison matrix
From A second gradient, with a mixed term: every row here is a choice with a cost. Fill the formula column, then say which row you would actually pick and what you give up for it.
| partial | formula | value (verified) |
|---|---|---|
| ∂g/∂x | 2xy | 12.0 |
| ∂g/∂y | x² + 3y² | 31.0 |
Concept
Because ∇f is a formula in x and y, it gives a different arrow at every point. The gradient isn't one vector — it's a whole field of them draped over the input plane.
Evaluate it at a few points to feel the field. Notice (2, 1) swaps the components of (1, 2) — same length, different heading — and at the flat point (0, 0) the arrow vanishes:
| point (x, y) | ∇f = [2x+3y, 3x+2y] | ‖∇f‖ (steepness) |
|---|---|---|
| (0, 0) | [0, 0] | 0.000 |
| (1, 0) | [2, 3] | 3.606 |
| (0, 1) | [3, 2] | 3.606 |
| (1, 2) | [8, 7] | 10.630 |
| (2, 1) | [7, 8] | 10.630 |
A point where ∇f = 0 is a critical point — flat in every direction. That is exactly what an optimizer is hunting for: the place the gradient shuts off.
Discrimination
Sort into buckets
Sort these by ‖∇f‖ (steepness), from memory, without looking back at One arrow per point: a vector field. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Intuition
Stand anywhere on the surface. The gradient arrow points in the direction where f climbs fastest, and its length is how steep that fastest climb is.
At (1, 2) the compass reads [8, 7]: the steepest uphill direction, pointing mostly along +x and a bit less along +y. Its length √(8² + 7²) = √113 ≈ 10.63 is the maximum rate of climb there.
Training walks the opposite way: take the gradient, flip its sign, step downhill. That one idea is every optimizer you will meet this year.
Concept
This isn't a slogan — it falls out of one inequality. The change in f for a tiny unit step v̂ is ∇f · v̂, and the dot product obeys:
\[ \nabla f \cdot \hat v = \lVert \nabla f \rVert\,\lVert \hat v \rVert \cos\theta = \lVert \nabla f \rVert \cos\theta \]
‖v̂‖ = 1, so the step's payoff is ‖∇f‖·cos θ, where θ is the angle between your step and the gradient. This is maximized when cos θ = 1 — that is, when you walk exactly along ∇f. Any other direction gives up a factor of cos θ < 1.
Concept
Zoom in far enough and any smooth surface looks flat — a tangent plane. The gradient is the slope of that plane: it tells you how f changes for a small step Δ = [Δx, Δy].
\[ f(x + \Delta) \approx f(x) + \nabla f \cdot \Delta \]
This is the first-order Taylor approximation. Every gradient step trusts it: move a little in the direction the linear map says is downhill, then re-measure. It is exact only in the limit of tiny steps — which is why too-large a learning rate breaks the approximation and diverges.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Goal: minimize the loss, so the gradient must point downhill toward the minimum — follow it.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Treats the gradient as a downhill arrow and walks ALONG it.
The gradient points toward steepest ascent — uphill, toward larger loss. To go down, negate it.
Why: Treats the gradient as a downhill arrow and walks ALONG it. But ∇f is steepest ASCENT, so this climbs toward LARGER loss — the loss goes up every step and training diverges.
Trap
Goal: minimize the loss, so the gradient must point downhill toward the minimum — follow it.
Update: w = w + lr · grad
Why: Treats the gradient as a downhill arrow and walks ALONG it. But ∇f is steepest ASCENT, so this climbs toward LARGER loss — the loss goes up every step and training diverges.
The gradient points toward steepest ascent — uphill, toward larger loss. To go down, negate it.
Update: w = w − lr · grad
Why: Step in the NEGATIVE gradient direction (cos θ = −1 minimizes ∇f · v̂). The single minus sign is the entire difference between training and diverging — memorize it.
Section
Part 2 of 5 — any direction, not just the axes
Concept
Partials give the slope along +x and +y. But what if you walk diagonally — some direction v that is not an axis? How fast does f change then?
The answer is the directional derivative D_v f: project the gradient onto the unit vector pointing along v.
\[ D_v f = \nabla f \cdot \frac{v}{\lVert v \rVert} = \nabla f \cdot \hat v \]
Intuition
A direction is about which way, not how far. If you dotted with a long vector v instead of the unit v̂, you'd multiply the slope by v's length — measuring speed and distance tangled together.
Dividing by ‖v‖ strips out the length, leaving a pure rate: units of f per unit of distance moved. That is what 'slope in a direction' has to mean.
Pattern
Predict first
The table runs: ∇f at (1,2) | [8, 7] · v̂ = v/‖v‖ | [0.7071, 0.7071] · D_v f = ∇f · v̂ | 10.6066
In Directional derivative along v = [1, 1], given the rows so far: what is the next one — the row where quantity is ‖∇f‖ (the maximum)?
Correct: ‖∇f‖ (the maximum) | 10.6301
| quantity | value (verified) |
|---|---|
| ∇f at (1,2) | [8, 7] |
| v̂ = v/‖v‖ | [0.7071, 0.7071] |
| D_v f = ∇f · v̂ | 10.6066 |
| ‖∇f‖ (the maximum) | 10.6301 |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. Divide each component by ‖v‖ = √2 to get a unit vector.
Worked example
Using ∇f(1,2) = [8, 7], find the rate of change if you move along v = [1, 1] (the diagonal):
Length: ‖v‖ = √(1² + 1²) = √2
Why: Ordinary Euclidean norm of [1, 1].
Normalize: v̂ = [1/√2, 1/√2] ≈ [0.707, 0.707]
Why: Divide each component by ‖v‖ = √2 to get a unit vector.
Dot with the gradient: D_v f = 8·(1/√2) + 7·(1/√2) = 15/√2
Why: ∇f · v̂. Both terms share 1/√2, so it factors to (8+7)/√2 = 15/√2.
\[ D_v f = \frac{8 + 7}{\sqrt2} = \frac{15}{\sqrt2} \approx 10.607 \]
Compare to the max: ‖∇f‖ = √113 ≈ 10.630
Why: 10.607 < 10.630 — climbing the diagonal is ALMOST as steep as the true steepest direction, but slightly less, exactly as the cos θ argument predicts.
| quantity | value (verified) |
|---|---|
| ∇f at (1,2) | [8, 7] |
| v̂ = v/‖v‖ | [0.7071, 0.7071] |
| D_v f = ∇f · v̂ | 10.6066 |
| ‖∇f‖ (the maximum) | 10.6301 |
Translation
\( D_v f = \frac{8 + 7}{\sqrt2} = \frac{15}{\sqrt2} \approx 10.607 \)
Draw it
Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.
Concept
D_v f = ‖∇f‖ cos θ makes the whole picture fall out of one angle θ between your step and the gradient:
cos θ = 1, D_v f = ‖∇f‖ — the steepest ascent, the maximum possible slopecos θ = −1, D_v f = −‖∇f‖ — steepest descent, what training usescos θ = 0, D_v f = 0 — no change; you are moving along a level curveThat perpendicular case is why gradients are always drawn at right angles to contour lines: along a contour, f doesn't change, so its directional derivative is zero.
Ranking
Put in order
Put the moves of Descent direction has slope −‖∇f‖ into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Point exactly opposite the gradient, normalized to unit length so we read a pure slope.
Worked example
The direction training actually uses is −∇f. Confirm its directional derivative is the most-negative possible, −‖∇f‖, using ∇f(1,2) = [8, 7]:
Descent unit vector: v̂ = −∇f / ‖∇f‖ = [−8, −7]/√113
Why: Point exactly opposite the gradient, normalized to unit length so we read a pure slope.
D_v f = ∇f · v̂ = −(8² + 7²)/√113 = −113/√113 = −√113
Why: The dot of ∇f with its own negative unit vector is −‖∇f‖ — the steepest possible DESCENT, mirror image of the steepest ascent.
\[ D_{-\nabla f}\,f = -\lVert \nabla f \rVert = -\sqrt{113} \approx -10.630 \]
Compare to the level-curve direction
Why: Perpendicular to ∇f is [7, −8]: moving along it, D_v f = (8·7 + 7·(−8))/√113 = 0. Same point, three directions, three signs — max, min, zero.
| direction from (1,2) | v̂ | D_v f (verified) |
|---|---|---|
| along +∇f (ascent) | [0.753, 0.658] | +10.630 |
| along −∇f (descent) | [−0.753, −0.658] | −10.630 |
| ⟂ ∇f (level curve) | [0.658, −0.753] | 0.000 |
Intuition
A level curve (contour) is the set of points where f has one fixed value. Walk along it and f never changes, so the directional derivative along it is zero — meaning your step is perpendicular to ∇f.
Figure (svg): Three nested oval contour lines with a gradient arrow at a point pointing outward perpendicular to the contour through that point.
This is why loss landscapes are drawn as contour maps with gradient arrows piercing them at 90°: the arrow always points the steepest way off the current level, and descent runs straight back down the arrow.
Section
Part 3 of 5 — vector-calculus results to memorize
Concept
Almost every gradient in machine learning reduces to three vector-calculus results. They reappear in least squares, ridge regularization, and attention — so we derive each, then confirm it with autograd.
\[ \frac{d\,\lVert w \rVert^2}{dw} = 2w, \qquad \frac{d\,(w^\top x)}{dw} = x \]
\[ \frac{d\,(x^\top A x)}{dx} = (A + A^\top)\,x \;=\; 2Ax \quad\text{when } A = A^\top \]
Missing information
Discussion prompt
‖w‖² is the sum of squared components. Differentiate with respect to one component wₖ, then reassemble the vector:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
The squared norm is just the sum of each coordinate squared — no vectors needed to differentiate it.
Worked example
‖w‖² is the sum of squared components. Differentiate with respect to one component wₖ, then reassemble the vector:
Write it as a sum: ‖w‖² = Σⱼ wⱼ²
Why: The squared norm is just the sum of each coordinate squared — no vectors needed to differentiate it.
\[ \lVert w \rVert^2 = \sum_j w_j^2 \]
Partial in wₖ: ∂/∂wₖ Σⱼ wⱼ² = 2wₖ
Why: Only the j = k term contains wₖ; every other term is constant in wₖ and drops. That one term gives 2wₖ.
Stack over all k: ∇‖w‖² = 2w
Why: Component k of the gradient is 2wₖ, so the whole gradient vector is 2w. This is the vector version of d(x²)/dx = 2x.
import torch
w = torch.tensor([3.0, -1.0, 2.0], requires_grad=True)
loss = (w**2).sum() # ||w||^2
loss.backward()
print((2*w).detach()) # hand result 2w
print(w.grad) # autograd result| source | value at w = [3, −1, 2] |
|---|---|
| hand: 2w | [6.0, −2.0, 4.0] |
| autograd w.grad | [6.0, −2.0, 4.0] |
| match? | yes ✓ |
Error analysis
Annotate
Walk the callouts on Identity 1: ∇‖w‖² = 2w, derived. Each one is a place this is easy to get subtly wrong.
Step zero
Discussion prompt
Identity 2: ∇(wᵀx) = x, derived — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Write it out: wᵀx = Σⱼ wⱼ xⱼ
Answer:
Worked example
wᵀx is the dot product Σⱼ wⱼ xⱼ — linear in w. Differentiate in one component again:
Write it out: wᵀx = Σⱼ wⱼ xⱼ
Why: A dot product is a sum of pairwise products. Here x is the constant, w is the variable.
\[ w^\top x = \sum_j w_j x_j \]
Partial in wₖ: ∂/∂wₖ Σⱼ wⱼ xⱼ = xₖ
Why: Only the j = k term has wₖ; its derivative is the constant coefficient xₖ. Everything else drops.
Stack over k: ∇(wᵀx) = x
Why: Component k is xₖ, so the gradient is the vector x itself — the vector version of d(bx)/dx = b.
import torch
w = torch.tensor([1.0, 2.0, 3.0], requires_grad=True)
x = torch.tensor([5.0, -2.0, 4.0])
loss = w @ x # w^T x
loss.backward()
print(x) # hand result: the gradient IS x
print(w.grad) # autograd result| source | value |
|---|---|
| hand: x | [5.0, −2.0, 4.0] |
| autograd w.grad | [5.0, −2.0, 4.0] |
| match? | yes ✓ |
Error analysis
Annotate
Walk the callouts on Identity 2: ∇(wᵀx) = x, derived. Each one is a place this is easy to get subtly wrong.
Hypothesis
Predict first
Identity 3: ∇(xᵀAx), the quadratic form is about to be worked. State your hypothesis first: which rule or definition decides this one, and what is the first move it forces? Then watch whether the example agrees with you.
Correct: Partial in xₖ hits the i = k copy and the j = k copy
Why: In the double sum, xₖ shows up wherever i = k (giving Σⱼ Aₖⱼ xⱼ = row k of Ax) and wherever j = k (giving Σᵢ xᵢ Aᵢₖ = row k of Aᵀx).
A hypothesis you wrote down is falsifiable; a vague sense of how it will go is not. If the example opens somewhere else, that gap is the thing worth chasing.
Worked example
The quadratic form xᵀAx = Σᵢ Σⱼ xᵢ Aᵢⱼ xⱼ. Here x appears twice, so differentiating hits both copies via the product rule:
Partial in xₖ hits the i = k copy and the j = k copy
Why: In the double sum, xₖ shows up wherever i = k (giving Σⱼ Aₖⱼ xⱼ = row k of Ax) and wherever j = k (giving Σᵢ xᵢ Aᵢₖ = row k of Aᵀx).
Add the two contributions: ∇ = Ax + Aᵀx = (A + Aᵀ)x
Why: One term from each occurrence of x. This is the fully general answer — no symmetry assumed yet.
\[ \frac{d\,(x^\top A x)}{dx} = (A + A^\top)\,x \]
If A is symmetric (A = Aᵀ), it collapses to 2Ax
Why: Then A + Aᵀ = 2A. Loss Hessians and XᵀX are symmetric, so 2Ax is the form you meet in practice.
\[ A = A^\top \;\Longrightarrow\; \frac{d\,(x^\top A x)}{dx} = 2Ax \]
import torch
A = torch.tensor([[2.0, 1.0], [1.0, 3.0]]) # symmetric
x = torch.tensor([1.0, 2.0], requires_grad=True)
loss = x @ (A @ x) # x^T A x
loss.backward()
print(2 * (A @ x)) # hand result 2Ax
print(x.grad) # autograd result| quantity | value (verified) |
|---|---|
| xᵀAx | 18.0 |
| hand: 2Ax | [8.0, 14.0] |
| autograd x.grad | [8.0, 14.0] |
| match? | yes ✓ |
Notation
Annotate
From Identity 3: ∇(xᵀAx), the quadratic form — read this one piece at a time. What is each part doing?
On: \( A = A^\top \;\Longrightarrow\; \frac{d\,(x^\top A x)}{dx} = 2Ax \)
Step zero
Discussion prompt
All three identities in one real loss — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Quadratic term wᵀ(XᵀX)w → 2XᵀXw
Answer:
Worked example
Watch the identities combine on the loss you'll meet next lesson: the least-squares loss L = ‖Xw − y‖². Expanded, it is wᵀ(XᵀX)w − 2(Xᵀy)ᵀw + yᵀy — a quadratic form, a linear term, and a constant.
Quadratic term wᵀ(XᵀX)w → 2XᵀXw
Why: Identity 3 with the symmetric A = XᵀX.
Linear term −2(Xᵀy)ᵀw → −2Xᵀy
Why: Identity 2 with x = Xᵀy, times the −2 constant.
Constant yᵀy → 0, then combine: ∇L = 2Xᵀ(Xw − y)
Why: Identity 1's cousin (constants vanish). All three rules fired in one gradient — the exact formula behind every regression fit.
\[ \nabla_w \lVert Xw - y \rVert^2 = 2X^\top(Xw - y) \]
import torch
X = torch.tensor([[1.,1.],[1.,2.],[1.,3.]])
y = torch.tensor([2.,4.,5.])
w = torch.tensor([0.5, 1.0], requires_grad=True)
L = ((X@w - y)**2).sum()
L.backward()
print((2*X.t()@(X@w-y)).detach()) # hand
print(w.grad) # autograd| quantity | value at w = [0.5, 1.0] |
|---|---|
| residual Xw − y | [−0.5, −1.5, −1.5] |
| L = ‖Xw − y‖² | 4.75 |
| hand: 2Xᵀ(Xw − y) | [−7.0, −16.0] |
| autograd w.grad | [−7.0, −16.0] |
Discrimination
Sort into buckets
Sort these by value at w = [0.5, 1.0], from memory, without looking back at All three identities in one real loss. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Differentiate xᵀAx. By analogy with the linear term wᵀx → x, guess the gradient is just Ax.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Misses that x appears TWICE in the quadratic form.
x sits on both sides of A, so the product rule yields two terms: Ax + Aᵀx.
Why: Misses that x appears TWICE in the quadratic form. The product rule gives a contribution from EACH copy, so you are short by exactly the Aᵀx term — off by a factor near 2.
Trap
Differentiate xᵀAx. By analogy with the linear term wᵀx → x, guess the gradient is just Ax.
d(xᵀAx)/dx = Ax ✗
Why: Misses that x appears TWICE in the quadratic form. The product rule gives a contribution from EACH copy, so you are short by exactly the Aᵀx term — off by a factor near 2.
x sits on both sides of A, so the product rule yields two terms: Ax + Aᵀx.
d(xᵀAx)/dx = (A + Aᵀ)x = 2Ax for symmetric A ✓
Why: Verified: A = [[2,1],[1,3]], x = [1,2] → 2Ax = [8, 14], matching autograd exactly. Guessing Ax would give [4, 7] — half the true gradient, and training would crawl at half speed.
Break the constraint
Discussion prompt
The rule this trap just fixed:
Verified: A = [[2,1],[1,3]], x = [1,2] → 2Ax = [8, 14], matching autograd exactly. Guessing Ax would give [4, 7] — half the true gradient, and training would crawl at half speed.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
Misses that x appears TWICE in the quadratic form. The product rule gives a contribution from EACH copy, so you are short by exactly the Aᵀx term — off by a factor near 2.
Section
Part 4 of 5 — the engine
Picture it
Figure (svg): Three nodes x, x squared, and f connected left to right, labeled forward in one direction and backward for the gradient pass.
Discussion prompt
Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.
Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.
Answer:
Every operation on a tensor with requires_grad=True records a node in a graph, storing the operation's local derivative as the forward pass runs.
Concept
Every operation on a tensor with requires_grad=True records a node in a graph, storing the operation's local derivative as the forward pass runs.
.backward() walks that graph backward, multiplying the local derivatives via the chain rule, and deposits ∂(output)/∂(leaf) into each leaf's .grad.
Figure (svg): Three nodes x, x squared, and f connected left to right, labeled forward in one direction and backward for the gradient pass.
Concept
A leaf is a tensor you created directly (not the result of an op). Setting requires_grad=True on a leaf tells autograd to track everything downstream of it and to fill its .grad on backward.
requires_grad=True — mark a leaf as a parameter to differentiate with respect to.backward() — run the reverse pass from a scalar output.grad — where the partial ∂output/∂leaf lands; starts as None.item() — pull a Python float out of a 0-d tensorFill the middle
Fill in the blanks
From Autograd reproduces our hand gradient — one line has had its right-hand side removed. Put it back.
import torch
x = torch.tensor(1.0, requires_grad=True)
y = torch.tensor(2.0, requires_grad=True)
f = x2 + 3xy + y2
f.backward()
print(f.item(), x.grad.item(), y.grad.item())
Why: f is what everything below it consumes, so the wrong expression here fails later and somewhere else. The scalar f triggers a single reverse traversal; each leaf receives its partial derivative.
Worked example
Recompute ∇f at (1, 2) for f = x² + 3xy + y² — but let PyTorch do the calculus, and compare to our by-hand [8, 7]:
import torch
x = torch.tensor(1.0, requires_grad=True)
y = torch.tensor(2.0, requires_grad=True)
f = x**2 + 3*x*y + y**2
f.backward()
print(f.item(), x.grad.item(), y.grad.item())f.backward() fills x.grad and y.grad in one pass
Why: The scalar f triggers a single reverse traversal; each leaf receives its partial derivative. No formula for the gradient was ever written.
| tensor | what it holds | value (verified) |
|---|---|---|
| f | x² + 3xy + y² | 11.0 |
| x.grad | ∂f/∂x = 2x + 3y | 8.0 |
| y.grad | ∂f/∂y = 3x + 2y | 7.0 |
Intuition
There is no magic and no symbolic algebra. As the forward pass runs, PyTorch stores the local derivative of each operation it performs.
On .backward() it multiplies those stored local derivatives together, from output back to input — exactly the chain rule you'll do by hand in Week 14's backprop. autograd just never forgets a factor.
Ranking
Put in order
Put the moves of The chain rule autograd runs, by hand into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. The first recorded node. At x=2, u = 4 + 1 = 5 and du/dx = 4.
Worked example
See the chain rule autograd applies. Take f = (x² + 1)³ at x = 2. Split it into an inner and outer step, exactly as PyTorch records nodes:
Inner: u = x² + 1, so du/dx = 2x
Why: The first recorded node. At x=2, u = 4 + 1 = 5 and du/dx = 4.
Outer: f = u³, so df/du = 3u²
Why: The second node. At u=5, df/du = 3·25 = 75.
Chain them: df/dx = (df/du)(du/dx) = 3u² · 2x = 75 · 4 = 300
Why: The reverse pass multiplies the stored local derivatives from output back to input — that product IS the chain rule.
\[ \frac{df}{dx} = 3u^2 \cdot 2x = 3(5)^2 \cdot 2(2) = 300 \]
import torch
x = torch.tensor(2.0, requires_grad=True)
u = x**2 + 1 # inner node
f = u**3 # outer node
f.backward()
print(f.item(), x.grad.item())| quantity | hand value | autograd |
|---|---|---|
| u = x² + 1 | 5.0 | 5.0 |
| f = u³ | 125.0 | 125.0 |
| df/dx = 3u²·2x | 300.0 | 300.0 |
Comparison
Comparison matrix
From The chain rule autograd runs, by hand: refill the hand value column from what you know. The rest of the table is as it appeared.
| quantity | hand value | autograd |
|---|---|---|
| u = x² + 1 | 5.0 | 5.0 |
| f = u³ | 125.0 | 125.0 |
| df/dx = 3u²·2x | 300.0 | 300.0 |
Concept
Differentiate the gradient again and you get the Hessian — the matrix of second partials. For our f = x² + 3xy + y² it is constant:
\[ \nabla^2 f = \begin{bmatrix} \partial^2 f/\partial x^2 & \partial^2 f/\partial x\,\partial y \\ \partial^2 f/\partial y\,\partial x & \partial^2 f/\partial y^2 \end{bmatrix} = \begin{bmatrix} 2 & 3 \\ 3 & 2 \end{bmatrix} \]
The Hessian is symmetric (mixed partials are equal, ∂²f/∂x∂y = ∂²f/∂y∂x = 3) — which is exactly why the ∇(xᵀAx) = 2Ax identity shows up so often: loss curvature is a symmetric quadratic form. You'll use the Hessian directly in Lesson 7's convexity argument.
Concept
Given any differentiable loss, the recipe is fixed: compute the gradient, then nudge each parameter against it.
\[ \theta \leftarrow \theta - \eta\,\nabla_\theta L \]
η (the learning rate) sets the step size. Too small crawls; too large overshoots and diverges. Below we'll use η = 0.1 and watch the parameters march toward the minimum.
Fill the middle
Fill in the blanks
From Descending a loss bowl with autograd — one line has had its right-hand side removed. Put it back.
import torch
w = torch.tensor(0.0, requires_grad=True)
b = torch.tensor(0.0, requires_grad=True)
lr = 0.1
for step in range(5):
loss = (w-3)2 + (b+1)2
loss.backward()
with torch.no_grad():
w -= lr * w.grad
b -= lr * b.grad
w.grad.zero_(); b.grad.zero_()
Why: loss is what everything below it consumes, so the wrong expression here fails later and somewhere else. This four-line rhythm is identical for a 2-parameter bowl and a 70-billion-parameter model.
Worked example
Minimize f(w, b) = (w−3)² + (b+1)² from (0, 0). The minimum is obviously (3, −1) — watch the steps march toward it, five passes, seed set:
import torch
w = torch.tensor(0.0, requires_grad=True)
b = torch.tensor(0.0, requires_grad=True)
lr = 0.1
for step in range(5):
loss = (w-3)**2 + (b+1)**2
loss.backward()
with torch.no_grad():
w -= lr * w.grad
b -= lr * b.grad
w.grad.zero_(); b.grad.zero_()Each pass: forward loss → backward → step under no_grad → zero the grads
Why: This four-line rhythm is identical for a 2-parameter bowl and a 70-billion-parameter model. Learn it here at a size you can trace by hand.
| step | w→ | b→ | loss | ∂w | ∂b |
|---|---|---|---|---|---|
| 1 | 0.600 | −0.200 | 10.000 | −6.00 | 2.00 |
| 2 | 1.080 | −0.360 | 6.400 | −4.80 | 1.60 |
| 3 | 1.464 | −0.488 | 4.096 | −3.84 | 1.28 |
| 4 | 1.771 | −0.590 | 2.621 | −3.07 | 1.02 |
| 5 | 2.017 | −0.672 | 1.678 | −2.46 | 0.82 |
Pattern
Step through it
Step through Descending a loss bowl with autograd one row at a time. What is driving the change, and what would the row after the last one be?
Intuition
Notice loss = (w-3)**2 + (b+1)**2 sits inside the loop. Each pass builds a fresh computational graph from the current w, b. PyTorch's graph is dynamic — it exists only for one forward/backward cycle, then is discarded.
That's why calling .backward() twice on the same graph errors (the buffers are freed after the first), but calling it once per freshly-built loss is fine. Rebuild, backward, step, zero — every iteration, from scratch.
Concept
torch.no_grad() isn't only for the update. At prediction time you never call .backward(), so tracking the graph is wasted memory and time. Wrapping inference in no_grad() turns tracking off entirely.
Under the guard, results carry requires_grad = False — no graph is built at all. Same numbers, less overhead. It's the standard wrapper around any evaluation or deployment code path.
Concept
The update w -= lr * w.grad is itself a tensor operation on a leaf that requires grad. Left unguarded, PyTorch either records it into next step's graph or refuses the in-place edit outright.
with torch.no_grad(): says 'don't track this'. The weights change, but the change is not part of any future derivative — which is exactly what a parameter update should be.
In fact, editing a grad-requiring leaf in place without no_grad() raises a RuntimeError — the next slide shows the exact message.
Explain it
Discussion prompt
Explain Why the update lives under no_grad() to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
The update w -= lr * w.grad is itself a tensor operation on a leaf that requires grad. Left unguarded, PyTorch either records it into next step's graph or refuses the in-place edit outright.
Estimation
Predict first
Try the update in place without the guard, then with it. The first raises; the second works:
Commit before you compute: what does The no_grad error, live come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: The unguarded in-place update raises RuntimeError
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. PyTorch forbids in-place edits of a leaf that requires grad, because it would corrupt the recorded graph.
Worked example
Try the update in place without the guard, then with it. The first raises; the second works:
import torch
w = torch.tensor(1.0, requires_grad=True)
try:
w -= 0.1 * torch.tensor(2.0) # in-place on a grad leaf
except RuntimeError as e:
print('ERROR:', str(e)[:55])
with torch.no_grad():
w -= 0.1 * torch.tensor(2.0) # now allowed
print('ok, w =', round(w.item(), 3))The unguarded in-place update raises RuntimeError
Why: PyTorch forbids in-place edits of a leaf that requires grad, because it would corrupt the recorded graph. The guard removes the tensor from tracking, so the edit is legal.
| code path | result (verified) |
|---|---|
| w -= ... (no guard) | RuntimeError: a leaf Variable that requires grad… |
| with torch.no_grad(): w -= ... | ok, w = 0.8 |
Anomaly
Predict first
A student writes this, and it looks reasonable:
.grad holds this step's gradient, so just call .backward() each loop and read it — no cleanup needed.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: PyTorch ACCUMULATES (adds) into .grad.
.grad is a running sum — you must reset it to zero every iteration.
Why: PyTorch ACCUMULATES (adds) into .grad. The true gradient is −6.0, but the stale −6.0 from the previous call was summed on top, giving −12.0. Steps get bigger every iteration and training silently corrupts.
Trap
.grad holds this step's gradient, so just call .backward() each loop and read it — no cleanup needed.
Loop without zeroing: on (w−5)² at w=2, backward twice → w.grad reads −12.0 ✗
Why: PyTorch ACCUMULATES (adds) into .grad. The true gradient is −6.0, but the stale −6.0 from the previous call was summed on top, giving −12.0. Steps get bigger every iteration and training silently corrupts.
.grad is a running sum — you must reset it to zero every iteration.
Call w.grad.zero_() after each step → w.grad correctly reads −6.0 ✓
Why: Verified: with zeroing the gradient stays −6.0 each pass; without it, it grows −6, −12, −18… The accumulation is a feature (it lets you sum grads across mini-batches) but only if you zero it when you don't want it.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
(x, y) plane.; The partial derivative ∂f/∂x is the slope you feel walking in the +x direction only — y held fixed, as if it were a constant.xᵀAx. By analogy with the linear term wᵀx → x, guess the gradient is just Ax.Ranking
Put in order
These are the steps of The autograd training loop — the recipe, scrambled. Put them back in order before the next slide shows you.
requires_grad=Trueloss from the parametersloss.backward() fills every .gradtorch.no_grad(): θ -= lr * θ.gradθ.grad.zero_() — or the next step's gradient is corruptedWhy: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
requires_grad=Trueloss from the parametersloss.backward() fills every .gradtorch.no_grad(): θ -= lr * θ.gradθ.grad.zero_() — or the next step's gradient is corruptedEdge cases
Discussion prompt
The autograd training loop — the recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
requires_grad=Trueloss from the parametersloss.backward() fills every .gradtorch.no_grad(): θ -= lr * θ.gradθ.grad.zero_() — or the next step's gradient is corruptedIntuition
Everything so far — two-variable gradients, the three identities, the four-line loop — is exactly what a modern network runs, just with more parameters. A transformer's loss.backward() fills millions of .grad entries by the same chain rule you traced by hand.
Nothing conceptually new is added when you scale up. Master the 2-parameter bowl and you understand the 70-billion-parameter one. The checks below make sure the fundamentals are solid before you trust them at scale.
Elimination
Eliminate the wrong options
For f(x, y) = x²y + y³, what is ∂f/∂x at (2, 3)?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Treat y as a constant: ∂f/∂x = 2xy. At (2, 3): 2·2·3 = 12. The y³ term has no x, so it drops to 0.
Check
Solve it on paper before clicking. Freeze the other variable.
Check your understanding
For f(x, y) = x²y + y³, what is ∂f/∂x at (2, 3)?
Answer: A
Why: Treat y as a constant: ∂f/∂x = 2xy. At (2, 3): 2·2·3 = 12. The y³ term has no x, so it drops to 0.
Prediction
Predict first
You are minimizing a loss L(θ). At the current θ, ∇L = [4, −2]. Which step decreases L fastest for a small fixed step size?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Move along [−4, 2]
Why: Steepest descent is the negative gradient: −∇L = [−4, 2]. The gradient points to steepest ASCENT, so flip its sign to go downhill fastest.
Check
Think about the sign, and about perpendicularity, before you pick.
Check your understanding
You are minimizing a loss L(θ). At the current θ, ∇L = [4, −2]. Which step decreases L fastest for a small fixed step size?
Answer: A
Why: Steepest descent is the negative gradient: −∇L = [−4, 2]. The gradient points to steepest ASCENT, so flip its sign to go downhill fastest.
Commit first
Predict first
At a point, ∇f = [3, 4]. What is the directional derivative of f in the direction v = [4, 3]?
Commit to an answer, then rate it — certain, fairly sure, or guessing — and write the rating down before you turn the page.
Correct: 24/5 = 4.8
Why: Normalize v: ‖v‖ = √(16+9) = 5, so v̂ = [0.8, 0.6]. Then D_v f = ∇f · v̂ = 3·0.8 + 4·0.6 = 2.4 + 2.4 = 4.8.
The rating matters as much as the answer: confident-and-wrong is the combination that survives revision, because nothing about it feels like it needs revisiting.
Check
Remember to normalize the direction first.
Check your understanding
At a point, ∇f = [3, 4]. What is the directional derivative of f in the direction v = [4, 3]?
Answer: A
Why: Normalize v: ‖v‖ = √(16+9) = 5, so v̂ = [0.8, 0.6]. Then D_v f = ∇f · v̂ = 3·0.8 + 4·0.6 = 2.4 + 2.4 = 4.8.
Prediction
Predict first
For symmetric A = [[2, 1], [1, 3]] and x = [1, 2], what is d(xᵀAx)/dx?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: [8, 14]
Why: For symmetric A the gradient is 2Ax. Ax = [2·1 + 1·2, 1·1 + 3·2] = [4, 7], so 2Ax = [8, 14]. Confirmed by autograd.
Check
Count how many times x appears. That's the whole question.
Check your understanding
For symmetric A = [[2, 1], [1, 3]] and x = [1, 2], what is d(xᵀAx)/dx?
Answer: A
Why: For symmetric A the gradient is 2Ax. Ax = [2·1 + 1·2, 1·1 + 3·2] = [4, 7], so 2Ax = [8, 14]. Confirmed by autograd.
Elimination
Eliminate the wrong options
w = torch.tensor(2.0, requires_grad=True). You run loss=(w-5)**2; loss.backward() TWICE in a row with no zero_grad() between. What does w.grad read?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Each backward computes ∂loss/∂w = 2(w−5) = 2(2−5) = −6, and PyTorch ADDS it into the existing .grad. Two calls → −6 + −6 = −12.0.
Check
This one bites everyone once. Trace it carefully.
Check your understanding
w = torch.tensor(2.0, requires_grad=True). You run loss=(w-5)**2; loss.backward() TWICE in a row with no zero_grad() between. What does w.grad read?
Answer: A
Why: Each backward computes ∂loss/∂w = 2(w−5) = 2(2−5) = −6, and PyTorch ADDS it into the existing .grad. Two calls → −6 + −6 = −12.0.
Section
Part 5 of 5 — the project
Concept
Minimize f(w) = (w − 5)² two ways — hand-coded calculus, then verified against autograd. The minimum is obviously w = 5; you'll watch your loop crawl toward it.
| # | requirement | tool |
|---|---|---|
| 1 | Define f(w) and its analytic gradient 2(w−5) | a Python function |
| 2 | Loop the update w ← w − 0.1·grad from w = 0 | a for-loop |
| 3 | Confirm one gradient against torch autograd | requires_grad / .backward() |
Build rules: type every line yourself, run after each line, and when something errors, read the error — don't delete it. The struggle is the encoding.
Analogy
Discussion prompt
Explain Project: gradient descent from scratch by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Minimize f(w) = (w − 5)² two ways — hand-coded calculus, then verified against autograd. The minimum is obviously w = 5; you'll watch your loop crawl toward it.
Worked example
Your turn: write f(w) and a separate grad(w). Say out loud what grad(0) should be before you run it.
Hint: f(w) = (w-5)**2, and by the chain rule its derivative is 2*(w-5) — no PyTorch yet, just arithmetic.
def f(w):
return (w - 5)**2
def grad(w):
return 2 * (w - 5)
print(f(0), grad(0))Predict then check: grad(0) = 2(0−5) = −10
Why: At w=0 the slope is steep and negative — f falls as w increases, so descent will push w UP toward 5. The sign tells you the direction before you loop.
| call | value (verified) |
|---|---|
| f(0) = (0−5)² | 25 |
| grad(0) = 2(0−5) | −10 |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Predict then check: grad(0) = 2(0−5) = −10
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Your turn: write f(w) and a separate grad(w). Say out loud what grad(0) should be before you run it.
Pattern
Predict first
The table runs: 1 | 0.000 | −10.000 | 1.0000 | 16.00000 · 2 | 1.000 | −8.000 | 1.8000 | 10.24000 · 3 | 1.800 | −6.400 | 2.4400 | 6.55360 · 4 | 2.440 | −5.120 | 2.9520 | 4.19430 · 5 | 2.952 | −4.096 | 3.3616 | 2.68435 · 6 | 3.362 | −3.277 | 3.6893 | 1.71799
In Milestone 2 — the descent loop, given the rows so far: what is the next one — the row where step is 7?
Correct: 7 | 3.689 | −2.621 | 3.9514 | 1.09951
| step | w_start | grad | w_new | f(w_new) |
|---|---|---|---|---|
| 1 | 0.000 | −10.000 | 1.0000 | 16.00000 |
| 2 | 1.000 | −8.000 | 1.8000 | 10.24000 |
| 3 | 1.800 | −6.400 | 2.4400 | 6.55360 |
| 4 | 2.440 | −5.120 | 2.9520 | 4.19430 |
| 5 | 2.952 | −4.096 | 3.3616 | 2.68435 |
| 6 | 3.362 | −3.277 | 3.6893 | 1.71799 |
| 7 | 3.689 | −2.621 | 3.9514 | 1.09951 |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. grad(w) = 2(w−5) → 0 as w → 5, so each step is smaller than the last.
Worked example
Your turn: loop 7 times from w = 0 with lr = 0.1, updating w each pass. Predict: does w reach exactly 5, or only approach it?
Hint: inside the loop compute g = grad(w), then w = w - lr * g. Print w and f(w) each pass so you can watch the loss shrink.
def f(w): return (w - 5)**2
def grad(w): return 2 * (w - 5)
w, lr = 0.0, 0.1
for step in range(1, 8):
g = grad(w)
w = w - lr * g
print(step, round(w, 4), round(f(w), 5))The steps SHRINK as w nears 5
Why: grad(w) = 2(w−5) → 0 as w → 5, so each step is smaller than the last. Gradient descent approaches the minimum but (with fixed lr) never lands exactly on it — it converges geometrically.
| step | w_start | grad | w_new | f(w_new) |
|---|---|---|---|---|
| 1 | 0.000 | −10.000 | 1.0000 | 16.00000 |
| 2 | 1.000 | −8.000 | 1.8000 | 10.24000 |
| 3 | 1.800 | −6.400 | 2.4400 | 6.55360 |
| 4 | 2.440 | −5.120 | 2.9520 | 4.19430 |
| 5 | 2.952 | −4.096 | 3.3616 | 2.68435 |
| 6 | 3.362 | −3.277 | 3.6893 | 1.71799 |
| 7 | 3.689 | −2.621 | 3.9514 | 1.09951 |
Pattern
Step through it
Step through Milestone 2 — the descent loop one row at a time. What is driving the change, and what would the row after the last one be?
Missing information
Discussion prompt
Your turn: at w = 2, confirm your hand gradient grad(2) equals what PyTorch computes. Predict the number first (it's 2(2−5)).
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
grad(2) = 2(2−5) = −6, and PyTorch's reverse pass computes the identical −6.0. That agreement is the whole point: you understand what autograd is doing.
Worked example
Your turn: at w = 2, confirm your hand gradient grad(2) equals what PyTorch computes. Predict the number first (it's 2(2−5)).
Hint: make wt = torch.tensor(2.0, requires_grad=True), build loss = (wt-5)**2, call loss.backward(), then read wt.grad.
import torch
def grad(w): return 2 * (w - 5)
wt = torch.tensor(2.0, requires_grad=True)
loss = (wt - 5)**2
loss.backward()
print(grad(2), wt.grad.item())Both read −6.0 — your calculus and autograd agree
Why: grad(2) = 2(2−5) = −6, and PyTorch's reverse pass computes the identical −6.0. That agreement is the whole point: you understand what autograd is doing.
| source | gradient at w = 2 |
|---|---|
| your grad(2) = 2(2−5) | −6.0 |
| wt.grad (autograd) | −6.0 |
| match? | yes ✓ |
Trade off
Comparison matrix
From Milestone 3 — verify against autograd: every row here is a choice with a cost. Fill the gradient at w = 2 column, then say which row you would actually pick and what you give up for it.
| source | gradient at w = 2 |
|---|---|
| your grad(2) = 2(2−5) | −6.0 |
| wt.grad (autograd) | −6.0 |
| match? | yes ✓ |
Estimation
Predict first
All three milestones assembled into one runnable script:
Commit before you compute: what does The full program come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Success criterion: w marches 0 → ~3.95, and the check line prints two matching −6.0
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. If your run matches the table below, you built gradient descent from the math up and proved it against PyTorch's engine.
Worked example
All three milestones assembled into one runnable script:
import torch
def f(w): return (w - 5)**2
def grad(w): return 2 * (w - 5)
w, lr = 0.0, 0.1
for step in range(1, 8):
g = grad(w)
w = w - lr * g
print(step, round(w, 4), round(f(w), 5))
# autograd cross-check at w = 2
wt = torch.tensor(2.0, requires_grad=True)
((wt - 5)**2).backward()
print('check:', grad(2), wt.grad.item())Success criterion: w marches 0 → ~3.95, and the check line prints two matching −6.0
Why: If your run matches the table below, you built gradient descent from the math up and proved it against PyTorch's engine.
| printed line | value (verified) |
|---|---|
| step 1 | 1.0 16.0 |
| step 4 | 2.952 4.1943 |
| step 7 | 3.9514 1.09951 |
| check | −6.0 −6.0 |
Comparison
Comparison matrix
From The full program: refill the value (verified) column from what you know. The rest of the table is as it appeared.
| printed line | value (verified) |
|---|---|
| step 1 | 1.0 16.0 |
| step 4 | 2.952 4.1943 |
| step 7 | 3.9514 1.09951 |
| check | −6.0 −6.0 |
Concept
Slides closed, out loud: explain (1) why grad(0) is negative, (2) why the steps shrink as w nears 5, and (3) what zero_grad() would prevent if you looped the autograd version.
Stretch (the homework CHALLENGE): the softmax Jacobian ∂sᵢ/∂xⱼ = sᵢ(δᵢⱼ − sⱼ). For x = [1, 2, 3], softmax = [0.090, 0.245, 0.665] — this exact derivative drives backprop through every classifier head. Build the 3×3 Jacobian and check it against torch.autograd.
Counterexample
Discussion prompt
Slides closed, out loud: explain (1) why grad(0) is negative, (2) why the steps shrink as w nears 5, and (3) what zero_grad() would prevent if you looped the autograd version.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — From one slope to many · The directional derivative · Three identities that power all of ML · Autograd: gradients for free · Your turn: build gradient descent. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
∇f∇f as a steepest-ascent compass, and prove it via ∇f · v̂ = ‖∇f‖ cos θD_v f = ∇f · v̂, and know its max, min, and zero (level-curve) directions∇‖w‖² = 2w, ∇(wᵀx) = x, ∇(xᵀAx) = 2Ax — the identities behind least squares and attentionbackward() → update under no_grad() → zero_grad() — and build gradient descent from scratch| move | the one thing to remember |
|---|---|
| partial | freeze the other variables |
| gradient sign | points uphill — SUBTRACT it to train |
| directional deriv | dot with the UNIT vector; zero along a level curve |
| xᵀAx | factor of 2 (symmetric A): 2Ax, not Ax |
| .grad | accumulates — zero it every step; update under no_grad() |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.