Lesson 15: Modern Optimizers

USAAIO Lesson 15, from Week 5 on optimization, fully worked. It starts by showing why plain gradient descent crawls on an ill-conditioned bowl with a condition number of 100, then derives momentum, RMSprop, and Adam and traces each of them step by step on ONE running problem, f(x) = ½(100x₀² + x₁²). Every update rule is walked out by hand, and bias correction is dissected at t=1, where the raw step is 3.16 times too big while the corrected one is exactly 1.0. All four optimizers are then raced with real loss numbers, and AdamW's decoupled weight decay and learning-rate warmup are derived. In the your-turn project you build momentum, RMSprop, and Adam from scratch. Every snippet runs standalone, and every number came from real numpy execution. The lesson runs to 61 slides.

Subject: Machine Learning · 107 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Modern Optimizers

Title

USAAIO · Lesson 15 · Week 5 (Optimization)

Why nobody trains with plain SGD anymore. We take one stretched bowl and race momentum, RMSprop, and Adam across it — each rule derived, hand-traced step by step, and built from scratch. Every loss number here came from real execution.

2. By the end of this lesson you can

Objectives

  1. Say exactly why plain gradient descent crawls on an ill-conditioned bowl — its step is capped by the steep axis while the flat axis inches
  2. Derive and hand-trace momentum v ← βv + g as a velocity that cancels oscillation and accumulates the consistent direction
  3. Derive and hand-trace RMSprop s ← ρs + (1−ρ)g² as a per-parameter adaptive rate that rescales each axis
  4. Assemble Adam = momentum + RMSprop + bias correction, and show at t=1 the raw step is 3.16× too big while the corrected step is exactly 1
  5. Distinguish Adam vs AdamW (decoupled weight decay), explain warmup, and implement all three optimizers from scratch

3. What survived from Information Theory?

Warm-up

Discussion prompt

Before we open Lesson 15: Modern Optimizers: without looking back, what was the main idea of Information Theory, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

entropy, KL divergence (asymmetric, ≥0), cross-entropy as H(P)+KL, mutual information, and the ELBO that underpins the VAE. Implement entropy, KL, and cross-entropy from scratch and verify the identities.

4. The problem: ill-conditioning

Section

Part 1 of 8

5. One bowl for the whole lesson

Concept

Every optimizer today races across the same quadratic bowl — a stretched paraboloid. We minimize:

\[ f(x_0, x_1) = \tfrac12\left(100\,x_0^2 + x_1^2\right) \]

Its gradient is one line, and we start every run from x = [1, 1], where the loss is ½(100 + 1) = 50.5.

\[ \nabla f(x) = \begin{bmatrix} 100\,x_0 \\ x_1 \end{bmatrix}, \qquad x^{(0)} = \begin{bmatrix}1\\1\end{bmatrix},\quad f(x^{(0)}) = 50.5 \]

6. Break it if you can: One bowl for the whole lesson

Counterexample

Discussion prompt

Every optimizer today races across the same quadratic bowl — a stretched paraboloid. We minimize:

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Its gradient is one line, and we start every run from x = [1, 1], where the loss is ½(100 + 1) = 50.5.

7. The Hessian and the condition number

Concept

The curvature of this bowl is its Hessian — the matrix of second derivatives. For our f it is diagonal and constant everywhere:

\[ H = \nabla^2 f = \begin{bmatrix} 100 & 0 \\ 0 & 1 \end{bmatrix} \]

condition number κ — The ratio of largest to smallest Hessian eigenvalue, κ = λ_max / λ_min. Here the eigenvalues are 100 and 1, so κ = 100. Big κ = a long, narrow valley — steep across one axis, nearly flat along the other.

κ = 100 is the entire villain of this lesson. Everything modern optimizers do is a response to large κ.

8. By analogy: The Hessian and the condition number

Analogy

Discussion prompt

Explain The Hessian and the condition number by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

The curvature of this bowl is its Hessian — the matrix of second derivatives. For our f it is diagonal and constant everywhere:

9. A taco shell, not a soup bowl

Intuition

Picture the surface. Along x₀ the walls shoot up steeply (curvature 100); along x₁ the floor is nearly level (curvature 1). It's a long narrow taco shell, not a round soup bowl.

A ball dropped in rolls hard across the narrow direction and barely drifts down the long one. That mismatch — fast one way, slow the other — is what makes a single global step size fail.

10. Teach it back: A taco shell, not a soup bowl

Explain it

Discussion prompt

Explain A taco shell, not a soup bowl to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Picture the surface. Along x₀ the walls shoot up steeply (curvature 100); along x₁ the floor is nearly level (curvature 1). It's a long narrow taco shell, not a round soup bowl.

11. Plain gradient descent, and its speed limit

Concept

Plain gradient descent (GD) takes one global step size η down the negative gradient:

\[ x \leftarrow x - \eta\,\nabla f(x) \]

On a quadratic, GD is stable only while η < 2/L, where L is the largest curvature (the steep axis). Here L = 100, so:

\[ \eta < \frac{2}{L} = \frac{2}{100} = 0.02 \]

The steep axis sets the ceiling for the whole vector

Why: One η must serve both axes. Push η above 0.02 and the steep x₀ axis diverges — so η is hostage to the steepest direction, even though the flat axis would happily take a step 100× larger.

12. Say it in words: Plain gradient descent, and its speed limit

Translation

\( \eta < \frac{2}{L} = \frac{2}{100} = 0.02 \)

Draw it

Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.

13. Guess the shape of the answer: Where the 2/L limit comes from

Estimation

Predict first

On our separable bowl each axis is its own 1-D problem f = ½Lx² with ∇ = Lx. One GD step multiplies x by a fixed factor — derive it:

Commit before you compute: what does Where the 2/L limit comes from come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Iterating gives xₜ = (1 − ηL)ᵗ x₀

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The factor compounds. The sequence shrinks to 0 only if the factor has magnitude below 1: |1 − ηL| < 1.

14. Where the 2/L limit comes from

Worked example

On our separable bowl each axis is its own 1-D problem f = ½Lx² with ∇ = Lx. One GD step multiplies x by a fixed factor — derive it:

Write one step on a single axis

Why: Substitute the gradient Lx into the update x ← x − ηLx and factor out x.

\[ x_{t+1} = x_t - \eta\,(L x_t) = (1 - \eta L)\,x_t \]

Iterating gives xₜ = (1 − ηL)ᵗ x₀

Why: The factor compounds. The sequence shrinks to 0 only if the factor has magnitude below 1: |1 − ηL| < 1.

\[ |1 - \eta L| < 1 \;\Longleftrightarrow\; 0 < \eta < \frac{2}{L} \]

axisLfactor 1−ηL at η=0.018behavior
x₀ (steep)100−0.800oscillates, shrinks 20%/step
x₁ (flat)1+0.982crawls, shrinks <2%/step
x₀ at η=0.021100−1.100|factor|>1 → DIVERGES

15. Fill in: L for Where the 2/L limit comes from

Comparison

Comparison matrix

From Where the 2/L limit comes from: refill the L column from what you know. The rest of the table is as it appeared.

axisLfactor 1−ηL at η=0.018behavior
x₀ (steep)100−0.800oscillates, shrinks 20%/step
x₁ (flat)1+0.982crawls, shrinks <2%/step
x₀ at η=0.021100−1.100|factor|>1 → DIVERGES

16. What has to be given first: Watch GD crawl — full trace

Missing information

Discussion prompt

Run GD at η = 0.018 (just under the 0.02 ceiling). Because f is separable, each axis updates independently by a fixed factor: x₀ ← (1 − 100η)x₀ and x₁ ← (1 − η)x₁.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

The steep axis multiplies by −0.8 each step: it OSCILLATES in sign (that minus) and shrinks only 20% per step. The flat axis multiplies by 0.982 — it barely moves, losing under 2% per step.

17. Watch GD crawl — full trace

Worked example

Run GD at η = 0.018 (just under the 0.02 ceiling). Because f is separable, each axis updates independently by a fixed factor: x₀ ← (1 − 100η)x₀ and x₁ ← (1 − η)x₁.

import numpy as np
def grad(x): return np.array([100*x[0], x[1]])
x, eta = np.array([1., 1.]), 0.018
for t in range(1, 11):
    x = x - eta*grad(x)
    loss = 0.5*(100*x[0]**2 + x[1]**2)
    print(t, np.round(x, 4), round(loss, 4))

x₀ factor is 1 − 100·0.018 = −0.8; x₁ factor is 1 − 0.018 = 0.982

Why: The steep axis multiplies by −0.8 each step: it OSCILLATES in sign (that minus) and shrinks only 20% per step. The flat axis multiplies by 0.982 — it barely moves, losing under 2% per step.

step tx₀x₁loss
1−0.80000.982032.4822
2+0.64000.964320.9450
3−0.51200.947013.5556
5−0.32770.91325.7857
10+0.10740.83390.9242

18. What each one costs: Watch GD crawl — full trace

Trade off

Comparison matrix

From Watch GD crawl — full trace: every row here is a choice with a cost. Fill the loss column, then say which row you would actually pick and what you give up for it.

step tx₀x₁loss
1−0.80000.982032.4822
2+0.64000.964320.9450
3−0.51200.947013.5556
5−0.32770.91325.7857
10+0.10740.83390.9242

19. The diagnosis: two failures at once

Concept

The trace shows both symptoms of ill-conditioning in one run:

  1. x₀ zig-zags — the sign flips every step (−0.8, +0.64, −0.51…) because its update factor is negative. Energy wasted bouncing across the valley.
  2. x₁ crawls — from 1.0 to 0.83 in ten steps. It needs a step 100× bigger, but η is pinned by x₀ and can't grow.

Momentum attacks the zig-zag; adaptive methods (RMSprop, Adam) attack the crawl. The next six parts build each fix and race them here.

20. Momentum

Section

Part 2 of 8 — a velocity

21. A heavy ball with mass

Intuition

Plain GD is a massless particle: it stops the instant the gradient does, and it turns on a dime every time the gradient flips. That's why it zig-zags.

Momentum gives the ball mass. Across the narrow valley the pushes alternate left–right and largely cancel; down the long axis the pushes all point the same way and accumulate into speed. The ball smooths the bounce and rolls straight down the valley.

22. The momentum update

Concept

Keep a velocity v — an exponential moving average of past gradients — and step along it instead of along the raw gradient:

\[ \begin{aligned} v &\leftarrow \beta\, v + \nabla f(x) \\ x &\leftarrow x - \eta\, v \end{aligned} \]

The friction/decay β (typically 0.9) says how much old velocity survives each step. β = 0 recovers plain GD; larger β gives the ball more inertia. We start v at the zero vector.

23. Why the velocity cancels and accumulates

Concept

Unroll the EMA: v is a decaying sum of every past gradient, v_t = Σ βᵏ g_{t−k}. On the steep axis the g's alternate sign, so the sum partly cancels — velocity stays small.

On the flat axis every g has the same sign, so the terms reinforce. A geometric series of same-sign pushes sums toward 1/(1−β) = 10× the single-step pull. The flat axis effectively gets a 10× boost — exactly where GD was crawling.

24. Predict the next row: Momentum, hand-traced (β=0.9, η=0.003)

Pattern

Predict first

The table runs: 1 | 100.000 | 1.000 | 24.9970 · 2 | 160.000 | 1.897 | 2.9113 · 4 | 121.600 | 3.412 | 21.1329 · 6 | −37.184 | 4.600 | 22.6748

In Momentum, hand-traced (β=0.9, η=0.003), given the rows so far: what is the next one — the row where step t is 8?

Correct: 8 | −126.756 | 5.510 | 0.4286

step tv[0]v[1]loss
1100.0001.00024.9970
2160.0001.8972.9113
4121.6003.41221.1329
6−37.1844.60022.6748
8−126.7565.5100.4286

Why: The relationship between the columns, not the individual numbers, is what generates the next row. The flat-axis velocity accumulates monotonically (same-sign pushes), so x₁ finally makes progress.

25. Momentum, hand-traced (β=0.9, η=0.003)

Worked example

Trace the velocity accumulating. Watch v[1] (flat axis) grow steadily from 1.0 while v[0] (steep axis) swings and averages toward zero.

import numpy as np
def grad(x): return np.array([100*x[0], x[1]])
x, v = np.array([1., 1.]), np.zeros(2)
for t in range(1, 9):
    g = grad(x)
    v = 0.9*v + g
    x = x - 0.003*v
    print(t, np.round(v, 3), np.round(x, 4))

v[1] climbs 1.00 → 1.90 → 2.70 → …; v[0] swings 100 → 160 → 166 → 122 → 45 → −37

Why: The flat-axis velocity accumulates monotonically (same-sign pushes), so x₁ finally makes progress. The steep-axis velocity peaks then reverses as the sign flips — the oscillation is being averaged out inside v.

step tv[0]v[1]loss
1100.0001.00024.9970
2160.0001.8972.9113
4121.6003.41221.1329
6−37.1844.60022.6748
8−126.7565.5100.4286

26. Watch it run: Momentum, hand-traced (β=0.9, η=0.003)

Pattern

Step through it

Step through Momentum, hand-traced (β=0.9, η=0.003) one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: step t is 1
  2. Step 2: step t is 2
  3. Step 3: step t is 4
  4. Step 4: step t is 6
  5. Step 5: step t is 8

27. Momentum overshoots before it settles

Intuition

Notice the loss in that trace is not monotone: 25.0 → 2.9 → … → 21.1 → 22.7 → 0.43. A heavy ball builds speed on the steep axis and coasts past the bottom before inertia turns it around.

That's the price of mass: momentum can overshoot early, then converges fast once the transient settles. By step 80 it reaches loss 0.0008 — far past where GD stalled.

28. The two momentum conventions

Concept

You will see momentum written two ways, and they are not the same. Ours accumulates the raw gradient into v; the other (PyTorch's SGD(momentum=...)) blends with a (1−β) weight:

\[ \underbrace{v \leftarrow \beta v + g}_{\text{ours (raw sum)}} \qquad\text{vs}\qquad \underbrace{v \leftarrow \beta v + (1-\beta)\,g}_{\text{normalized average}} \]

The normalized form keeps v on the scale of a single gradient, so it needs an η about 1/(1−β) = 10× larger to match. Same trajectory, different η bookkeeping — know which one your library uses before you tune.

29. RMSprop

Section

Part 3 of 8 — per-axis rates

30. Why squared gradients, not absolute

Concept

RMSprop averages g², then divides by √s. Why the square-then-root instead of just averaging |g|? Because s is a running mean of squares, so √s is a root-mean-square — that's the RMS in the name.

Squaring is also what connects it to curvature: on a quadratic, the typical squared gradient on an axis scales with that axis's curvature, so √s estimates the local steepness and g/√s is a Newton-like, curvature-normalized step — without ever forming the Hessian.

31. RMSprop is a cheap second-order stand-in

Intuition

A true second-order method (Lesson 21) divides the step by the curvature itself — but that needs the Hessian, which is huge and expensive for real networks.

RMSprop fakes it: √s is a running, per-coordinate proxy for curvature built only from gradients you already computed. You get most of the axis-rescaling benefit of a second-order method at first-order cost. That trade is why adaptive methods dominate deep learning.

32. Give each axis its own step size

Intuition

Momentum still uses one global η. But the real disease is that the two axes want wildly different step sizes — the steep one small, the flat one large.

RMSprop's idea: measure how big each axis's gradients typically are, and divide each coordinate's step by that size. Loud axes get quieted; quiet axes get amplified. Every coordinate ends up moving at a comparable pace.

33. Where does each piece belong: Lesson 15: Modern Optimizers

Sorting

Sort into buckets

These are the pieces of Lesson 15: Modern Optimizers, out of order. Put each one back under the part of the lesson it belongs to.

The problem: ill-conditioning
One bowl for the whole lesson; The Hessian and the condition number; A taco shell, not a soup bowl
Momentum
A heavy ball with mass; The momentum update; Why the velocity cancels and accumulates
RMSprop
Why squared gradients, not absolute; RMSprop is a cheap second-order stand-in; Give each axis its own step size
s1
The problem: ill-conditioning is where Lesson 15: Modern Optimizers puts One bowl for the whole lesson, The Hessian and the condition number, A taco shell, not a soup bowl. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
Momentum is where Lesson 15: Modern Optimizers puts A heavy ball with mass, The momentum update, Why the velocity cancels and accumulates. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
RMSprop is where Lesson 15: Modern Optimizers puts Why squared gradients, not absolute, RMSprop is a cheap second-order stand-in, Give each axis its own step size. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

34. The RMSprop update

Concept

Track a running average of squared gradients, coordinate by coordinate, then divide the step by its square root:

\[ \begin{aligned} s &\leftarrow \rho\, s + (1-\rho)\, g^{\odot 2} \\ x &\leftarrow x - \eta\,\frac{g}{\sqrt{s} + \epsilon} \end{aligned} \]

Here g² and the division are elementwise (that's the ⊙). Defaults ρ = 0.9, ε = 1e-8 (a floor that stops division by zero). s starts at the zero vector.

√s is a per-coordinate estimate of gradient magnitude, so g/√s is roughly a unit-size step on every axis — the rescaling an ill-conditioned bowl needs.

35. Restore the missing line: RMSprop, hand-traced (ρ=0.9, η=0.05)

Fill the middle

Fill in the blanks

From RMSprop, hand-traced (ρ=0.9, η=0.05) — one line has had its right-hand side removed. Put it back.

import numpy as np
def grad(x): return np.array([100*x[0], x[1]])
x, s = np.array([1., 1.]), np.zeros(2)
for t in range(1, 6):
g = grad(x)
s = 0.9s + 0.1g*g
step = g/(np.sqrt(s) + 1e-8)
x = x - 0.05*step
print(t, np.round(s, 3), np.round(step, 4), np.round(x, 5))

Why: step is what everything below it consumes, so the wrong expression here fails later and somewhere else. s[0]=0.1·100²=1000 and s[1]=0.1·1²=0.1.

36. RMSprop, hand-traced (ρ=0.9, η=0.05)

Worked example

At step 1 the raw gradient is [100, 1] — a 100:1 mismatch. Watch RMSprop erase it: after dividing by √s, both axes move by the identical amount.

import numpy as np
def grad(x): return np.array([100*x[0], x[1]])
x, s = np.array([1., 1.]), np.zeros(2)
for t in range(1, 6):
    g = grad(x)
    s = 0.9*s + 0.1*g*g
    step = g/(np.sqrt(s) + 1e-8)
    x = x - 0.05*step
    print(t, np.round(s, 3), np.round(step, 4), np.round(x, 5))

step 1: s = [1000, 0.1], and step = g/√s = [3.162, 3.162] — identical on both axes

Why: s[0]=0.1·100²=1000 and s[1]=0.1·1²=0.1. Then 100/√1000 = 1/√0.1 = 3.162. The 100:1 gradient ratio is completely cancelled: each axis moves by 0.05·3.162 = 0.1581, so both x₀ and x₁ go 1 → 0.84189.

step ts[0]step[0]=step[1]x₀ = x₁loss
11000.0003.16230.8418935.7930
21608.7722.09900.7369427.4254
31990.9721.65160.6543621.6234
52340.1861.20910.5244613.8906

37. What happens as it grows: RMSprop, hand-traced (ρ=0.9, η=0.05)

Scale up

Step through it

Step through RMSprop, hand-traced (ρ=0.9, η=0.05) and watch the numbers move. Now imagine the input ten times bigger: which column is the one that stops this being practical?

  1. Step 1: step t is 1
  2. Step 2: step t is 2
  3. Step 3: step t is 3
  4. Step 4: step t is 5

38. RMSprop's payoff

Concept

Because both axes descend in lockstep, the crawl is gone. By step 20 the loss is 0.25, by step 40 it is ≈ 1.5e-7, and by step 80 it has effectively hit machine zero — a completely different regime from GD's 0.03.

The catch: RMSprop rescales but has no memory of direction — it doesn't smooth the path. Combine its per-axis rescaling with momentum's smoothing and you get Adam.

39. Adam

Section

Part 4 of 8 — assemble the pieces

40. The two fixes are orthogonal

Intuition

Momentum fixes the direction of the step (smooth the zig-zag by averaging gradients over time). RMSprop fixes the scale of the step (rescale each axis by its own gradient size). These are independent knobs.

So there's nothing stopping you from turning both at once: average the gradient for direction, average the squared gradient for scale. That combination is Adam — and the only new subtlety is that both averages start biased at zero.

41. Adam = momentum + RMSprop

Concept

Adam keeps both running averages: a first moment m (momentum's velocity) and a second moment v (RMSprop's squared-gradient average).

\[ \begin{aligned} m &\leftarrow \beta_1\, m + (1-\beta_1)\, g \\ v &\leftarrow \beta_2\, v + (1-\beta_2)\, g^{\odot 2} \end{aligned} \]

Both start at the zero vector. Defaults β₁ = 0.9, β₂ = 0.999, ε = 1e-8. Note m uses the (1−β₁) weighting (a true average), unlike our bare momentum — a cosmetic difference that just rescales η.

42. The zero-init bias problem

Concept

Because m and v start at 0, early estimates are pulled toward zero. At t=1, with β₂ = 0.999:

\[ v_1 = 0.999\cdot 0 + 0.001\, g^2 = 0.001\,g^2 \]

That is 1000× too small as an estimate of g²! Its square root, √v ≈ 0.0316\,|g|, is 31.6× too small — so dividing by it would make the step 31.6× too large. The estimate hasn't warmed up yet.

43. Finish it with less help: Bias correction, dissected at t=1

Faded example

Fill in the blanks

Bias correction, dissected at t=1, with the scaffolding fading: two lines are gone now — fill both.

import numpy as np
g = 100.0 # steep-axis gradient at t=1
m, v = 0.1g, 0.001g*g # m=(1-b1)g, v=(1-b2)g^2
raw = m/(np.sqrt(v) + 1e-8) # NO correction
mh, vh = m/(1-0.91), v/(1-0.9991) # bias-corrected
cor = mh/(np.sqrt(vh) + 1e-8)
print('raw step dir:', round(raw, 4))
print('corrected :', round(cor, 4))

Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. Correction rescales m by 1/(1−0.9)=10× and v by 1/(1−0.999)=1000×.

44. Bias correction, dissected at t=1

Worked example

The fix: divide each moment by (1 − βᵗ) to undo the shrink. Watch what it does to the actual step magnitude on the steep axis at t=1, where g₀ = 100:

import numpy as np
g = 100.0                      # steep-axis gradient at t=1
m, v = 0.1*g, 0.001*g*g        # m=(1-b1)g, v=(1-b2)g^2
raw = m/(np.sqrt(v) + 1e-8)               # NO correction
mh, vh = m/(1-0.9**1), v/(1-0.999**1)     # bias-corrected
cor = mh/(np.sqrt(vh) + 1e-8)
print('raw step dir:', round(raw, 4))
print('corrected   :', round(cor, 4))

raw = m/√v = 3.1623, but corrected = m̂/√v̂ = 1.0000

Why: Correction rescales m by 1/(1−0.9)=10× and v by 1/(1−0.999)=1000×. After correction m̂=g and v̂=g², so m̂/√v̂ = g/|g| = sign(g), magnitude exactly 1. Uncorrected it is 3.16× too big — the un-warmed v underestimates g².

quantity at t=1raw (no BC)corrected
m (first moment)10.0m̂ = 100.0
v (second moment)10.0v̂ = 10000.0
step dir = m/√v3.16231.0000
effective step η·(…), η=0.10.31620.1000

45. Which is which, by raw (no BC)

Discrimination

Sort into buckets

Sort these by raw (no BC), from memory, without looking back at Bias correction, dissected at t=1. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

10.0
m (first moment); v (second moment)
3.1623
step dir = m/√v
0.3162
effective step η·(…), η=0.1
g1
raw (no BC) is "10.0" for m (first moment), v (second moment) — that is what the table on "Bias correction, dissected at t=1" records, and it is the single property separating this group from the rest.
g2
raw (no BC) is "3.1623" for step dir = m/√v — that is what the table on "Bias correction, dissected at t=1" records, and it is the single property separating this group from the rest.
g3
raw (no BC) is "0.3162" for effective step η·(…), η=0.1 — that is what the table on "Bias correction, dissected at t=1" records, and it is the single property separating this group from the rest.

46. The full Adam update

Concept

Correct both moments with their own factor, then take the RMSprop-style step:

\[ \hat m = \frac{m}{1-\beta_1^{\,t}}, \qquad \hat v = \frac{v}{1-\beta_2^{\,t}}, \qquad x \leftarrow x - \eta\,\frac{\hat m}{\sqrt{\hat v} + \epsilon} \]

As t grows, βᵗ → 0, so (1 − βᵗ) → 1 and the correction quietly fades away — it only matters in the first dozen-or-so steps. Adam is the default optimizer for nearly all deep learning.

47. Adam's step size is self-bounding

Intuition

Here's why Adam is so forgiving to tune. The update per coordinate is η·m̂/√v̂, and both m̂ and √v̂ are averages of the same gradients — so their ratio is roughly sign(g), of magnitude near 1.

That means each parameter moves by about ±η regardless of its gradient's scale. A coordinate with a gradient of 100 and one with a gradient of 0.001 both step by ~η. The raw gradient magnitude is divided out — which is exactly the cure for a 100:1 bowl, and why one η now works for all axes.

48. Adam, hand-traced (η=0.1)

Worked example

Full Adam on the bowl. The bias-corrected m̂[0] starts at exactly 100 (the true gradient), and the step marches both axes down together — no zig-zag, no crawl.

import numpy as np
def grad(x): return np.array([100*x[0], x[1]])
x, m, v = np.array([1., 1.]), np.zeros(2), np.zeros(2)
for t in range(1, 11):
    g = grad(x)
    m = 0.9*m + 0.1*g
    v = 0.999*v + 0.001*g*g
    mh, vh = m/(1-0.9**t), v/(1-0.999**t)
    x = x - 0.1*mh/(np.sqrt(vh) + 1e-8)
    print(t, np.round(x, 5), round(0.5*(100*x[0]**2+x[1]**2), 4))

x moves in lockstep on both axes: [0.9,0.9] → [0.8,0.8] → … → [0.076,0.076]

Why: Because m̂/√v̂ ≈ sign(g) early, each axis takes an ≈η=0.1 step regardless of its 100:1 gradient gap. By step 10 the loss is 0.294 — already below where GD sat at step 10 (0.924).

step tm̂[0]x₀ = x₁loss
1100.0000.9000040.9050
389.3140.7015924.8573
672.2270.414248.6654
1048.3720.076250.2936

49. Watch it run: Adam, hand-traced (η=0.1)

Pattern

Step through it

Step through Adam, hand-traced (η=0.1) one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: step t is 1
  2. Step 2: step t is 3
  3. Step 3: step t is 6
  4. Step 4: step t is 10

50. Something is wrong here: skipping Adam's bias correction

Anomaly

Predict first

A student writes this, and it looks reasonable:

m and v are moving averages of the gradient, so just plug them straight into the update — the /(1−βᵗ) looks like decoration.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: At t=1, v = 0.001·g² underestimates g² by 1000×, so √v is 31.6× too small and the step direction blows up to 3.16 instead of 1.

Divide each moment by (1 − βᵗ) to undo the zero-init shrink before you step.

Why: At t=1, v = 0.001·g² underestimates g² by 1000×, so √v is 31.6× too small and the step direction blows up to 3.16 instead of 1. The un-warmed second moment makes the earliest steps unreliably large.

51. Trap: skipping Adam's bias correction

Trap

The trap

m and v are moving averages of the gradient, so just plug them straight into the update — the /(1−βᵗ) looks like decoration.

x -= η·m/(√v + ε), no correction

Why: At t=1, v = 0.001·g² underestimates g² by 1000×, so √v is 31.6× too small and the step direction blows up to 3.16 instead of 1. The un-warmed second moment makes the earliest steps unreliably large.

On our bowl it converges worse: loss 1.12 at step 30

Why: Verified: uncorrected Adam (η=0.1) sits at loss 1.12 by step 30, while corrected Adam is at 0.058 — the bad early steps cost you the whole run.

The fix

Divide each moment by (1 − βᵗ) to undo the zero-init shrink before you step.

m̂ = m/(1−β₁ᵗ), v̂ = v/(1−β₂ᵗ), then step

Why: At t=1 this rescales v by 1000× and m by 10×, so m̂/√v̂ = sign(g), a clean unit step. The correction fades as t grows (βᵗ→0), so it only touches the fragile early steps.

Corrected Adam reaches loss 0.058 by step 30

Why: The right early-step magnitude keeps the trajectory on track. Verified against real execution.

52. Break it on purpose: skipping Adam's bias correction

Break the constraint

Discussion prompt

The rule this trap just fixed:

At t=1 this rescales v by 1000× and m by 10×, so m̂/√v̂ = sign(g), a clean unit step. The correction fades as t grows (βᵗ→0), so it only touches the fragile early steps.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

At t=1, v = 0.001·g² underestimates g² by 1000×, so √v is 31.6× too small and the step direction blows up to 3.16 instead of 1. The un-warmed second moment makes the earliest steps unreliably large.

53. Restore the missing line: Bias correction on vs off — head to head

Fill the middle

Fill in the blanks

From Bias correction on vs off — head to head — one line has had its right-hand side removed. Put it back.

import numpy as np
def grad(x): return np.array([100*x[0], x[1]])
def loss(x): return 0.5(100x[0]2 + x[1]2)
def adam(bc, steps=30, lr=0.1, b1=0.9, b2=0.999, eps=1e-8):
x, m, v = np.array([1.,1.]), np.zeros(2), np.zeros(2)
hist = *b2v + (1-b2)gg**
for t in range(1, steps+1):
g = grad(x)
m = b1m + (1-b1)g
v = ___
if bc: mh, vh = m/(1-b1t), v/(1-b2t)
else: mh, vh = m, v
x = x - lr*mh/(np.sqrt(vh)+eps)
hist[t] = loss(x)
return hist
on, off = adam(True), adam(False)
print('with BC :', [round(on[k], 4) for k in (1, 5, 30)])
print('no BC :', [round(off[k], 4) for k in (1, 5, 30)])

Why: v is what everything below it consumes, so the wrong expression here fails later and somewhere else. The uncorrected step direction is 3.16× too big at t=1, which happens to cut the first loss more — a false win.

54. Bias correction on vs off — head to head

Worked example

One flag toggles the /(1−βᵗ) terms. Run both variants on the bowl at η = 0.1 and compare the loss at three checkpoints.

import numpy as np
def grad(x): return np.array([100*x[0], x[1]])
def loss(x): return 0.5*(100*x[0]**2 + x[1]**2)
def adam(bc, steps=30, lr=0.1, b1=0.9, b2=0.999, eps=1e-8):
    x, m, v = np.array([1.,1.]), np.zeros(2), np.zeros(2)
    hist = {}
    for t in range(1, steps+1):
        g = grad(x)
        m = b1*m + (1-b1)*g
        v = b2*v + (1-b2)*g*g
        if bc: mh, vh = m/(1-b1**t), v/(1-b2**t)
        else:  mh, vh = m, v
        x = x - lr*mh/(np.sqrt(vh)+eps)
        hist[t] = loss(x)
    return hist
on, off = adam(True), adam(False)
print('with BC :', [round(on[k], 4)  for k in (1, 5, 30)])
print('no BC   :', [round(off[k], 4) for k in (1, 5, 30)])

no-BC starts smaller (23.6 vs 40.9 at t=1) but ends WORSE (1.12 vs 0.058 at t=30)

Why: The uncorrected step direction is 3.16× too big at t=1, which happens to cut the first loss more — a false win. Those oversized early steps mistrack the valley and the run never recovers, ending 19× worse by step 30.

variantloss @1@5@30
with bias correction40.90513.0300.058
no bias correction23.61123.0821.120

55. Work backwards from the answer: Bias correction on vs off — head to head

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

no-BC starts smaller (23.6 vs 40.9 at t=1) but ends WORSE (1.12 vs 0.058 at t=30)

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

One flag toggles the /(1−βᵗ) terms. Run both variants on the bowl at η = 0.1 and compare the loss at three checkpoints.

56. Race the four

Section

Part 5 of 8 — same bowl, same start

57. Guess the shape of the answer: All four optimizers, one bowl

Estimation

Predict first

Same problem, same start [1,1], start loss 50.5. Each optimizer uses its own tuned η. This block runs GD, momentum, RMSprop, and Adam and prints the loss at four checkpoints.

Commit before you compute: what does All four optimizers, one bowl come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: GD stalls near 0.03; momentum & Adam reach ~0.001; RMSprop hits machine zero

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. GD is capped by the steep axis and never escapes the slow flat-axis regime.

58. All four optimizers, one bowl

Worked example

Same problem, same start [1,1], start loss 50.5. Each optimizer uses its own tuned η. This block runs GD, momentum, RMSprop, and Adam and prints the loss at four checkpoints.

import numpy as np
def grad(x): return np.array([100*x[0], x[1]])
def loss(x): return 0.5*(100*x[0]**2 + x[1]**2)

def adam(steps, lr=0.1, b1=0.9, b2=0.999, eps=1e-8):
    x, m, v = np.array([1.,1.]), np.zeros(2), np.zeros(2)
    hist = {}
    for t in range(1, steps+1):
        g = grad(x)
        m = b1*m + (1-b1)*g
        v = b2*v + (1-b2)*g*g
        mh, vh = m/(1-b1**t), v/(1-b2**t)
        x = x - lr*mh/(np.sqrt(vh)+eps)
        hist[t] = loss(x)
    return hist

h = adam(80)
print([round(h[k], 4) for k in (10, 20, 40, 80)])

GD stalls near 0.03; momentum & Adam reach ~0.001; RMSprop hits machine zero

Why: GD is capped by the steep axis and never escapes the slow flat-axis regime. The adaptive/momentum methods give the flat axis the large effective step it always needed. Adam prints [0.2936, 3.713, 0.513, 0.0015].

optimizer (η)loss @10@20@40@80
GD (0.018)0.92420.24840.11690.0273
momentum (0.003)15.551.9320.34410.0008
RMSprop (0.05)4.5870.2512≈1.5e-7≈0
Adam (0.1)0.29363.7130.51300.0015

59. Reading the race honestly

Concept

The table isn't a clean 'Adam wins' story, and that's the point. Momentum and Adam both overshoot mid-run (momentum 15.5 at step 10; Adam 3.7 at step 20) as their velocity coasts past the bottom, then recover.

RMSprop, with no momentum to overshoot, descends smoothly straight to machine zero on this pure quadratic. Every method beats GD's 0.027 floor by orders of magnitude — because every method frees the flat axis.

60. The shape of each trajectory

Concept

The loss curve reveals each method's personality on the same bowl:

The lesson: momentum buys speed at the cost of an early overshoot; adaptivity buys smoothness. Adam takes both, so it inherits a little of both traits.

61. Something is wrong here: is Adam always the best choice?

Anomaly

Predict first

A student writes this, and it looks reasonable:

Adam has the most machinery — momentum AND adaptive rates AND bias correction — so it must always converge fastest and best.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: On our own bowl Adam overshoots (loss 0.29 → 3.71 between steps 10 and 20) and RMSprop actually reaches a lower loss.

Treat Adam as a robust default, not a guaranteed winner — match the optimizer and its schedule to the problem.

Why: On our own bowl Adam overshoots (loss 0.29 → 3.71 between steps 10 and 20) and RMSprop actually reaches a lower loss. 'Most components' does not mean 'always best'.

62. Trap: is Adam always the best choice?

Trap

The trap

Adam has the most machinery — momentum AND adaptive rates AND bias correction — so it must always converge fastest and best.

Reach for Adam everywhere and never tune

Why: On our own bowl Adam overshoots (loss 0.29 → 3.71 between steps 10 and 20) and RMSprop actually reaches a lower loss. 'Most components' does not mean 'always best'.

Ignore that η and the schedule still matter

Why: Adam with a bad η diverges like anything else; the warmup/decay schedule often decides success more than the optimizer choice does.

The fix

Treat Adam as a robust default, not a guaranteed winner — match the optimizer and its schedule to the problem.

Adam/AdamW for transformers & sparse gradients; well-tuned SGD+momentum for many CNNs

Why: Each has regimes where it wins: on some vision tasks a tuned SGD+momentum generalizes better than Adam. Pick per problem, and tune η.

Always pair the optimizer with a schedule

Why: Warmup then cosine decay stabilizes and sharpens convergence — the schedule is part of the choice, not an afterthought.

63. Which of these survive contact with Lesson 15: Modern Optimizers?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Every optimizer today races across the same quadratic bowl — a stretched paraboloid. We minimize:; The curvature of this bowl is its Hessian — the matrix of second derivatives. For our f it is diagonal and constant everywhere:; Plain gradient descent (GD) takes one global step size η down the negative gradient:
Breaks
m and v are moving averages of the gradient, so just plug them straight into the update — the /(1−βᵗ) looks like decoration.; Adam has the most machinery — momentum AND adaptive rates AND bias correction — so it must always converge fastest and best.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 15: Modern Optimizers puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

64. AdamW & warmup

Section

Part 6 of 8 — the modern defaults

65. Weight decay, and where Adam puts it

Concept

Weight decay shrinks each weight toward 0 every step — the L2 regularizer of Lesson 12. The classic way folds it into the gradient as g + λθ, so it flows through Adam's adaptive √v̂ divisor.

\[ \text{Adam+L2:}\quad g \leftarrow \nabla f + \lambda\theta, \qquad x \leftarrow x - \eta\,\frac{\hat m}{\sqrt{\hat v}+\epsilon} \]

The problem: dividing the decay by √v̂ means weights on loud (large-gradient) axes decay less than weights on quiet ones. The regularization strength gets distorted per-coordinate — not what you asked for.

66. AdamW: decouple the decay

Concept

AdamW applies the decay as its own term, outside the adaptive step, so every weight decays by the same fraction η·λ:

\[ x \leftarrow x - \eta\,\frac{\hat m}{\sqrt{\hat v}+\epsilon} - \eta\,\lambda\,\theta \]

The decay η·λ·θ no longer passes through √v̂, so it is decoupled from the gradient geometry. This is the standard recipe for training transformers.

67. Finish it with less help: AdamW vs Adam+L2 differ on step one

Faded example

Fill in the blanks

AdamW vs Adam+L2 differ on step one, with the scaffolding fading: two lines are gone now — fill both.

import numpy as np
def grad(x): return np.array([100*x[0], x[1]])
b1,b2,eta,eps,lam = 0.9,0.999,0.1,1e-8,0.1
# Adam + L2: decay folded into the gradient
x = np.array([1.,1.]); g = grad(x) + lam*x
m, v = (1-b1)g, (1-b2)g*g
mh, vh = m/(1-b1), v/(1-b2)
x_l2 = *x - etamh/(np.sqrt(vh)+eps)
# AdamW: decay applied separately
x =
np.array([1.,1.]); g = grad(x)**
m, v = (1-b1)g, (1-b2)g*g
mh, vh = m/(1-b1), v/(1-b2)
x_w = x - etamh/(np.sqrt(vh)+eps) - etalam*x
print('Adam+L2:', np.round(x_l2,4), ' AdamW:', np.round(x_w,4))

Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. In Adam+L2 the λθ term is swallowed by the √v̂ normalization (m̂/√v̂ ≈ sign, so the decay's size is lost).

68. AdamW vs Adam+L2 differ on step one

Worked example

One step from θ = [1, 1] with λ = 0.1, η = 0.1. The two recipes land on different points — proof the placement matters.

import numpy as np
def grad(x): return np.array([100*x[0], x[1]])
b1,b2,eta,eps,lam = 0.9,0.999,0.1,1e-8,0.1
# Adam + L2: decay folded into the gradient
x = np.array([1.,1.]); g = grad(x) + lam*x
m, v = (1-b1)*g, (1-b2)*g*g
mh, vh = m/(1-b1), v/(1-b2)
x_l2 = x - eta*mh/(np.sqrt(vh)+eps)
# AdamW: decay applied separately
x = np.array([1.,1.]); g = grad(x)
m, v = (1-b1)*g, (1-b2)*g*g
mh, vh = m/(1-b1), v/(1-b2)
x_w = x - eta*mh/(np.sqrt(vh)+eps) - eta*lam*x
print('Adam+L2:', np.round(x_l2,4), ' AdamW:', np.round(x_w,4))

Adam+L2 → [0.90, 0.90]; AdamW → [0.89, 0.89]; they differ by 0.01

Why: In Adam+L2 the λθ term is swallowed by the √v̂ normalization (m̂/√v̂ ≈ sign, so the decay's size is lost). In AdamW the −η·λ·θ = −0.01 is applied cleanly on top. Same λ, different weight decay.

recipex₀ after 1 stepx₁ after 1 step
Adam + L2 (folded)0.90000.9000
AdamW (decoupled)0.89000.8900
difference0.01000.0100

69. Fill in: x₀ after 1 step for AdamW vs Adam+L2 differ on step one

Comparison

Comparison matrix

From AdamW vs Adam+L2 differ on step one: refill the x₀ after 1 step column from what you know. The rest of the table is as it appeared.

recipex₀ after 1 stepx₁ after 1 step
Adam + L2 (folded)0.90000.9000
AdamW (decoupled)0.89000.8900
difference0.01000.0100

70. Learning-rate warmup

Concept

We saw the flip side of bias correction: very early, the second moment v̂ is built from just a few gradients and is noisy. Even corrected, a big η on a shaky estimate can destabilize training.

Warmup ramps η linearly from ≈0 over the first few hundred steps, then cosine-anneals it back down toward 0. The ramp buys time for v̂ to settle before full-size steps hit; transformers essentially require it.

71. Teach it back: Learning-rate warmup

Explain it

Discussion prompt

Explain Learning-rate warmup to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

We saw the flip side of bias correction: very early, the second moment v̂ is built from just a few gradients and is noisy. Even corrected, a big η on a shaky estimate can destabilize training.

72. A warmup + cosine schedule

Worked example

Linear ramp over the first W steps to base, then a cosine decay to 0 across the remaining T − W. This block prints the η the optimizer would use at several steps.

import math
def lr_sched(t, base=0.1, W=10, T=80):
    if t < W:
        return base * (t+1)/W          # linear warmup
    prog = (t - W)/(T - W)
    return base * 0.5*(1 + math.cos(math.pi*prog))   # cosine decay
for t in [0, 1, 4, 9, 20, 40, 79]:
    print(t, round(lr_sched(t), 5))

η climbs 0.01 → 0.10 across the first 10 steps, then decays 0.10 → ~0 by step 79

Why: During warmup (t<10) the ramp keeps steps small while v̂ is noisy; at t=9 it reaches the base 0.10; then cosine annealing shrinks it to 0.00005 near the end for a gentle landing.

step tphaseη
0warmup0.01000
9warmup end0.10000
20cosine decay0.09505
40cosine decay0.06113
79near end0.00005

73. Work backwards from the answer: A warmup + cosine schedule

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

η climbs 0.01 → 0.10 across the first 10 steps, then decays 0.10 → ~0 by step 79

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Linear ramp over the first W steps to base, then a cosine decay to 0 across the remaining T − W. This block prints the η the optimizer would use at several steps.

74. Warmup and bias correction solve the same era, differently

Intuition

Both target the fragile first few steps, where v̂ is estimated from almost no data. Bias correction fixes the systematic shrink (the 1000× underestimate at t=1); warmup handles the random noise that correction can't remove.

Correction makes the early step the right size on average; warmup keeps it small so a single unlucky, noisy v̂ can't send the weights flying before the averages settle. On a huge transformer, one bad early step can poison the whole run — so you use both.

75. The toolkit

Section

Part 7 of 8 — pattern & checks

76. Rebuild the recipe: The optimizer toolkit

Ranking

Put in order

These are the steps of The optimizer toolkit, scrambled. Put them back in order before the next slide shows you.

  1. Momentum — v = βv + g; x -= ηv. A velocity EMA: cancels oscillation, accumulates the consistent direction. β ≈ 0.9.
  2. RMSprop — s = ρs + (1−ρ)g²; x -= ηg/(√s+ε). Per-parameter rate: divide each axis by its own gradient size.
  3. Adam — both moments, then bias-correct m̂ = m/(1−β₁ᵗ), v̂ = v/(1−β₂ᵗ), then x -= η m̂/(√v̂+ε).
  4. AdamW — decouple weight decay: x -= η m̂/(√v̂+ε) − ηλθ (decay outside the adaptive divisor).
  5. Schedule — warmup η up, then cosine-decay it down; near-mandatory for transformers.

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

77. The optimizer toolkit

Pattern

  1. Momentum — v = βv + g; x -= ηv. A velocity EMA: cancels oscillation, accumulates the consistent direction. β ≈ 0.9.
  2. RMSprop — s = ρs + (1−ρ)g²; x -= ηg/(√s+ε). Per-parameter rate: divide each axis by its own gradient size.
  3. Adam — both moments, then bias-correct m̂ = m/(1−β₁ᵗ), v̂ = v/(1−β₂ᵗ), then x -= η m̂/(√v̂+ε).
  4. AdamW — decouple weight decay: x -= η m̂/(√v̂+ε) − ηλθ (decay outside the adaptive divisor).
  5. Schedule — warmup η up, then cosine-decay it down; near-mandatory for transformers.

78. Where does it stop working: The optimizer toolkit

Edge cases

Discussion prompt

The optimizer toolkit works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Momentum — v = βv + g; x -= ηv. A velocity EMA: cancels oscillation, accumulates the consistent direction. β ≈ 0.9.
  2. RMSprop — s = ρs + (1−ρ)g²; x -= ηg/(√s+ε). Per-parameter rate: divide each axis by its own gradient size.
  3. Adam — both moments, then bias-correct m̂ = m/(1−β₁ᵗ), v̂ = v/(1−β₂ᵗ), then x -= η m̂/(√v̂+ε).
  4. AdamW — decouple weight decay: x -= η m̂/(√v̂+ε) − ηλθ (decay outside the adaptive divisor).
  5. Schedule — warmup η up, then cosine-decay it down; near-mandatory for transformers.

79. Rule out three: Check yourself — why GD crawls

Elimination

Eliminate the wrong options

On f(x) = ½(100x₀² + x₁²), why does plain gradient descent barely move the x₁ axis?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. One global η must stay below 2/100 for the steep x₀ axis to be stable, leaving the flat x₁ axis a tiny effective step
  • B. The gradient of x₁ is exactly zero, so it never updates
  • C. GD ignores x₁ because it has the smaller index
  • D. The x₁ axis has already converged, so no further motion is needed

Survives elimination: A

Why: Stability requires η < 2/L with L = 100 (the steep curvature). That same small η, applied to the flat axis whose curvature is only 1, shrinks x₁ by a factor of just (1 − η) ≈ 0.982 per step — a crawl. One step size cannot serve a 100:1 curvature gap.

80. Check yourself — why GD crawls

Check

Think about what caps the step size on our bowl.

Check your understanding

On f(x) = ½(100x₀² + x₁²), why does plain gradient descent barely move the x₁ axis?

  • A. One global η must stay below 2/100 for the steep x₀ axis to be stable, leaving the flat x₁ axis a tiny effective step (correct)
  • B. The gradient of x₁ is exactly zero, so it never updates
  • C. GD ignores x₁ because it has the smaller index
  • D. The x₁ axis has already converged, so no further motion is needed

Answer: A

Why: Stability requires η < 2/L with L = 100 (the steep curvature). That same small η, applied to the flat axis whose curvature is only 1, shrinks x₁ by a factor of just (1 − η) ≈ 0.982 per step — a crawl. One step size cannot serve a 100:1 curvature gap.

Why B tempts people
The x₁ gradient is x₁ itself, which is 1 at the start — nonzero. It updates every step; it just updates by a tiny amount because η is pinned small.
Why C tempts people
Gradient descent has no notion of index priority; it updates all coordinates simultaneously by the same η. The asymmetry comes from curvature, not ordering.
Why D tempts people
x₁ starts at 1 with loss contribution ½·1 = 0.5, far from converged. It is still crawling toward 0, not finished.

81. Answer it before you see the options: Check yourself — Adam's parts

Prediction

Predict first

Adam combines which two mechanisms, plus a correction?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: Momentum (first moment m) + RMSprop (second moment v), with bias correction /(1−βᵗ)

Why: Adam tracks a first moment m (the momentum velocity) and a second moment v (RMSprop's squared-gradient average), then bias-corrects both with /(1−βᵗ) because they start at zero. That's the entire recipe.

82. Check yourself — Adam's parts

Check

What two ideas is Adam built from?

Check your understanding

Adam combines which two mechanisms, plus a correction?

  • A. Momentum (first moment m) + RMSprop (second moment v), with bias correction /(1−βᵗ) (correct)
  • B. Newton's method (the Hessian) + a line search
  • C. Weight decay + dropout
  • D. Gradient clipping + warmup

Answer: A

Why: Adam tracks a first moment m (the momentum velocity) and a second moment v (RMSprop's squared-gradient average), then bias-corrects both with /(1−βᵗ) because they start at zero. That's the entire recipe.

Why B tempts people
Adam is strictly first-order: it uses only gradients and their running averages. It never forms or inverts the Hessian, and it does no line search.
Why C tempts people
Weight decay and dropout are regularizers, not the moment estimates that define Adam's update direction. (AdamW does add decoupled decay, but that's an add-on, not what Adam is made of.)
Why D tempts people
Clipping and warmup are training tricks wrapped around an optimizer; they are not the two moments Adam is assembled from.

83. Answer it before you see the options: Check yourself — bias correction

Prediction

Predict first

Without bias correction, what goes wrong at Adam's very first step (β₂ = 0.999)?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: v = 0.001g² underestimates g², so √v is too small and the step is too large / unreliable

Why: At t=1, v = (1−0.999)g² = 0.001g², about 1000× below g². Its root √v ≈ 0.0316|g| is 31.6× too small, so g/√v ≈ 3.16·sign(g) instead of 1 — the first step is ~3× too big. Correction divides v by (1−0.999) = 0.001, restoring v̂ = g² and a unit-size step.

84. Check yourself — bias correction

Check

Recall what happens at t=1 with β₂ = 0.999.

Check your understanding

Without bias correction, what goes wrong at Adam's very first step (β₂ = 0.999)?

  • A. v = 0.001g² underestimates g², so √v is too small and the step is too large / unreliable (correct)
  • B. v = 0.001g² overestimates g², so the step is far too small and training stalls
  • C. The gradient is computed incorrectly at t=1
  • D. Nothing — bias correction never affects the first step

Answer: A

Why: At t=1, v = (1−0.999)g² = 0.001g², about 1000× below g². Its root √v ≈ 0.0316|g| is 31.6× too small, so g/√v ≈ 3.16·sign(g) instead of 1 — the first step is ~3× too big. Correction divides v by (1−0.999) = 0.001, restoring v̂ = g² and a unit-size step.

Why B tempts people
It's the reverse: v starts too SMALL (biased toward its zero init), not too large. Too-small v makes the divisor small and the step big, not small.
Why C tempts people
The gradient g = ∇f(x⁽⁰⁾) is exact at t=1; the bias is in the moving-average estimate v, not in g itself.
Why D tempts people
Bias correction has its LARGEST effect at t=1, where (1−βᵗ) is smallest (0.001 for v). It fades as t grows, not the other way around.

85. Rule out three: Check yourself — AdamW

Elimination

Eliminate the wrong options

How does AdamW differ from folding L2 into the gradient (Adam+L2)?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. It applies the decay −ηλθ separately, outside the adaptive √v̂ divisor, so every weight decays by the same fraction
  • B. It uses a larger λ than Adam+L2 but is otherwise identical
  • C. It removes weight decay entirely to speed up training
  • D. It divides the decay by √v̂ twice for stronger regularization

Survives elimination: A

Why: AdamW's update is x ← x − η·m̂/(√v̂+ε) − η·λ·θ. The decay term is decoupled from the adaptive step, so it isn't distorted by √v̂ and each weight shrinks by the same fraction ηλ. On our one-step demo AdamW landed at 0.89 vs Adam+L2's 0.90.

86. Check yourself — AdamW

Check

Where does AdamW put the weight decay?

Check your understanding

How does AdamW differ from folding L2 into the gradient (Adam+L2)?

  • A. It applies the decay −ηλθ separately, outside the adaptive √v̂ divisor, so every weight decays by the same fraction (correct)
  • B. It uses a larger λ than Adam+L2 but is otherwise identical
  • C. It removes weight decay entirely to speed up training
  • D. It divides the decay by √v̂ twice for stronger regularization

Answer: A

Why: AdamW's update is x ← x − η·m̂/(√v̂+ε) − η·λ·θ. The decay term is decoupled from the adaptive step, so it isn't distorted by √v̂ and each weight shrinks by the same fraction ηλ. On our one-step demo AdamW landed at 0.89 vs Adam+L2's 0.90.

Why B tempts people
The difference is structural (where the decay is applied), not a matter of a bigger λ. With the SAME λ the two give different updates.
Why C tempts people
AdamW keeps weight decay — it's designed to make decay behave correctly, not to drop it. Removing regularization would hurt generalization.
Why D tempts people
The whole point is the opposite: AdamW takes the decay OUT of the √v̂ path so it isn't divided by the adaptive term at all.

87. Your turn: build the optimizers

Section

Part 8 of 8 — the project

88. Project: momentum, RMSprop, Adam from scratch

Concept

Implement all three update rules yourself and run each on the same ill-conditioned bowl f(x) = ½(100x₀² + x₁²) from [1, 1]. You've traced every rule by hand — now assemble them in code.

#requirementkey line
1momentum: velocity EMAv = 0.9v + g; x -= lrv
2RMSprop: per-axis rates = 0.9s + 0.1gg; x -= lrg/(√s+ε)
3Adam: two moments + bias correctionmh=m/(1-b1t); vh=v/(1-b2t)

Build rules: type every line yourself, initialize all state to zeros, and for Adam loop t from 1 and never drop the /(1−βᵗ) correction. Run after each milestone and read the loss.

89. By analogy: Project: momentum, RMSprop, Adam from scratch

Analogy

Discussion prompt

Explain Project: momentum, RMSprop, Adam from scratch by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Implement all three update rules yourself and run each on the same ill-conditioned bowl f(x) = ½(100x₀² + x₁²) from [1, 1]. You've traced every rule by hand — now assemble them in code.

90. Restore the missing line: Milestone 1 — momentum

Fill the middle

Fill in the blanks

From Milestone 1 — momentum — one line has had its right-hand side removed. Put it back.

import numpy as np
def grad(x): return np.array([100*x[0], x[1]])
x, v, lr = np.array([1., 1.]), np.zeros(2), 0.003
for t in range(1, 81):
g = grad(x)
v = *0.9v + g**
x = x - lr*v
if t in (10, 40, 80):
print(t, round(0.5(100x[0]2 + x[1]2), 4))

Why: v is what everything below it consumes, so the wrong expression here fails later and somewhere else. The heavy ball builds speed on the steep axis and coasts past the minimum around step 10, then inertia turns it and it settles well below GD's 0.027 by step 80.

91. Milestone 1 — momentum

Worked example

Your turn: implement momentum and run 80 steps at η = 0.003. Predict aloud whether it beats GD's step-80 loss of 0.027 before you print.

Hint: keep v = np.zeros(2); each step g = grad(x); v = 0.9*v + g; x = x - lr*v. Expect an early overshoot in the loss, then a fast finish.

import numpy as np
def grad(x): return np.array([100*x[0], x[1]])
x, v, lr = np.array([1., 1.]), np.zeros(2), 0.003
for t in range(1, 81):
    g = grad(x)
    v = 0.9*v + g
    x = x - lr*v
    if t in (10, 40, 80):
        print(t, round(0.5*(100*x[0]**2 + x[1]**2), 4))

Loss is NON-monotone: 15.55 at step 10 (overshoot), 0.344 at 40, 0.0008 at 80

Why: The heavy ball builds speed on the steep axis and coasts past the minimum around step 10, then inertia turns it and it settles well below GD's 0.027 by step 80.

steploss
0 (start)50.5
1015.55
400.344
800.0008

92. Watch it run: Milestone 1 — momentum

Pattern

Step through it

Step through Milestone 1 — momentum one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: step is 0 (start)
  2. Step 2: step is 10
  3. Step 3: step is 40
  4. Step 4: step is 80

93. Predict the next row: Milestone 2 — RMSprop

Pattern

Predict first

The table runs: 0 (start) | 50.5 · 20 | 0.2512 · 40 | ≈1.5e-7

In Milestone 2 — RMSprop, given the rows so far: what is the next one — the row where step is 80?

Correct: 80 | ≈0

steploss
0 (start)50.5
200.2512
40≈1.5e-7
80≈0

Why: The relationship between the columns, not the individual numbers, is what generates the next row. No momentum means no overshoot — RMSprop rescales both axes to move in lockstep and drives this pure quadratic to machine zero.

94. Milestone 2 — RMSprop

Worked example

Your turn: implement RMSprop at η = 0.05. Predict which axis its √s division speeds up most before you run.

Hint: s = np.zeros(2); each step s = 0.9*s + 0.1*g*g; x = x - lr*g/(np.sqrt(s) + 1e-8). On step 1 both axes should move by the identical amount.

import numpy as np
def grad(x): return np.array([100*x[0], x[1]])
x, s, lr = np.array([1., 1.]), np.zeros(2), 0.05
for t in range(1, 81):
    g = grad(x)
    s = 0.9*s + 0.1*g*g
    x = x - lr*g/(np.sqrt(s) + 1e-8)
    if t in (20, 40, 80):
        print(t, round(0.5*(100*x[0]**2 + x[1]**2), 8))

Smooth, monotone descent: 0.251 at 20, ~1.5e-7 at 40, ~0 at 80

Why: No momentum means no overshoot — RMSprop rescales both axes to move in lockstep and drives this pure quadratic to machine zero. It speeds up the flat x₁ axis the most (its tiny gradient gets divided by a tiny √s).

steploss
0 (start)50.5
200.2512
40≈1.5e-7
80≈0

95. What happens as it grows: Milestone 2 — RMSprop

Scale up

Step through it

Step through Milestone 2 — RMSprop and watch the numbers move. Now imagine the input ten times bigger: which column is the one that stops this being practical?

  1. Step 1: step is 0 (start)
  2. Step 2: step is 20
  3. Step 3: step is 40
  4. Step 4: step is 80

96. What has to be given first: Milestone 3 — Adam

Missing information

Discussion prompt

Your turn: implement Adam with bias correction at η = 0.1. Predict what breaks if you delete the mh, vh lines and step with raw m, v instead.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

Bias correction makes the first step a clean unit-size move (m̂/√v̂ = sign(g)), so Adam is already at 0.294 by step 10 — ahead of GD. Drop the mh/vh lines and the un-warmed v makes early steps 3.16× too big, converging worse.

97. Milestone 3 — Adam

Worked example

Your turn: implement Adam with bias correction at η = 0.1. Predict what breaks if you delete the mh, vh lines and step with raw m, v instead.

Hint: track m, v = np.zeros(2), np.zeros(2); each step compute mh = m/(1-0.9**t), vh = v/(1-0.999**t); then x -= lr*mh/(np.sqrt(vh)+1e-8). Loop t from 1.

import numpy as np
def grad(x): return np.array([100*x[0], x[1]])
x, m, v, lr = np.array([1., 1.]), np.zeros(2), np.zeros(2), 0.1
for t in range(1, 81):
    g = grad(x)
    m = 0.9*m + 0.1*g
    v = 0.999*v + 0.001*g*g
    mh, vh = m/(1-0.9**t), v/(1-0.999**t)
    x = x - lr*mh/(np.sqrt(vh) + 1e-8)
    if t in (10, 20, 80):
        print(t, round(0.5*(100*x[0]**2 + x[1]**2), 5))

Loss 0.294 at 10, a mild overshoot 3.71 at 20, then 0.0015 at 80

Why: Bias correction makes the first step a clean unit-size move (m̂/√v̂ = sign(g)), so Adam is already at 0.294 by step 10 — ahead of GD. Drop the mh/vh lines and the un-warmed v makes early steps 3.16× too big, converging worse.

steploss
0 (start)50.5
100.294
203.713
800.0015

98. Watch it run: Milestone 3 — Adam

Pattern

Step through it

Step through Milestone 3 — Adam one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: step is 0 (start)
  2. Step 2: step is 10
  3. Step 3: step is 20
  4. Step 4: step is 80

99. Guess the shape of the answer: Milestone 4 — race your three vs GD

Estimation

Predict first

Your turn: wrap each rule in a function returning the final loss, and print all four side by side. Predict the ranking at step 80 before you run.

Commit before you compute: what does Milestone 4 — race your three vs GD come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: GD 0.02734 (worst); momentum 0.00075; Adam 0.00147; RMSprop ≈0 (best)

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Every adaptive/momentum method beats GD's floor by 1–30×.

100. Milestone 4 — race your three vs GD

Worked example

Your turn: wrap each rule in a function returning the final loss, and print all four side by side. Predict the ranking at step 80 before you run.

Hint: each function starts x=np.array([1.,1.]) and its own zero state, loops 80 steps, returns loss(x). Reuse grad and loss helpers.

import numpy as np
def grad(x): return np.array([100*x[0], x[1]])
def loss(x): return 0.5*(100*x[0]**2 + x[1]**2)
def run(kind, lr, steps=80):
    x, v, s, m, w = np.array([1.,1.]), np.zeros(2), np.zeros(2), np.zeros(2), np.zeros(2)
    for t in range(1, steps+1):
        g = grad(x)
        if kind == 'gd':   x = x - lr*g
        elif kind == 'mom': v = 0.9*v + g;            x = x - lr*v
        elif kind == 'rms': s = 0.9*s + 0.1*g*g;      x = x - lr*g/(np.sqrt(s)+1e-8)
        elif kind == 'adam':
            m = 0.9*m + 0.1*g; w = 0.999*w + 0.001*g*g
            mh, vh = m/(1-0.9**t), w/(1-0.999**t)
            x = x - lr*mh/(np.sqrt(vh)+1e-8)
    return round(loss(x), 5)
for k, lr in [('gd',0.018),('mom',0.003),('rms',0.05),('adam',0.1)]:
    print(k, run(k, lr))

GD 0.02734 (worst); momentum 0.00075; Adam 0.00147; RMSprop ≈0 (best)

Why: Every adaptive/momentum method beats GD's floor by 1–30×. RMSprop wins on this pure quadratic; GD is stuck because its η is hostage to the steep axis.

optimizerfinal loss @80
gd (0.018)0.02734
mom (0.003)0.00075
adam (0.1)0.00147
rms (0.05)≈0

101. What each one costs: Milestone 4 — race your three vs GD

Trade off

Comparison matrix

From Milestone 4 — race your three vs GD: every row here is a choice with a cost. Fill the final loss @80 column, then say which row you would actually pick and what you give up for it.

optimizerfinal loss @80
gd (0.018)0.02734
mom (0.003)0.00075
adam (0.1)0.00147
rms (0.05)≈0

102. The full program

Concept

import numpy as np
def grad(x): return np.array([100*x[0], x[1]])
def loss(x): return 0.5*(100*x[0]**2 + x[1]**2)

def adam(steps=80, lr=0.1, b1=0.9, b2=0.999, eps=1e-8):
    x, m, v = np.array([1., 1.]), np.zeros(2), np.zeros(2)
    for t in range(1, steps+1):
        g = grad(x)
        m = b1*m + (1-b1)*g              # momentum piece
        v = b2*v + (1-b2)*g*g            # RMSprop piece
        mh, vh = m/(1-b1**t), v/(1-b2**t)  # bias correction
        x = x - lr*mh/(np.sqrt(vh)+eps)
    return loss(x)

print(round(adam(), 5))     # 0.00147
outputvalue
start loss50.5
adam() final loss0.00147
plain GD @80 (for contrast)0.02734

If your Adam drives the loss from 50.5 to about 0.0015 while plain GD is still stuck near 0.027 — you've built the optimizer that trains modern networks.

103. Fill in: value for The full program

Comparison

Comparison matrix

From The full program: refill the value column from what you know. The rest of the table is as it appeared.

outputvalue
start loss50.5
adam() final loss0.00147
plain GD @80 (for contrast)0.02734

104. Show it off

Concept

Slides closed, out loud: explain (1) why one global η crawls on a κ = 100 bowl, (2) what RMSprop's √s division does to the steep axis, and (3) why bias correction matters most at step 1.

Stretch (homework): add a warmup + cosine schedule to your Adam and re-race; then swap in AdamW's decoupled decay. These optimizers reappear in the full training loop (Week 14) and transformer training (Week 29).

105. Break it if you can: Show it off

Counterexample

Discussion prompt

Slides closed, out loud: explain (1) why one global η crawls on a κ = 100 bowl, (2) what RMSprop's √s division does to the steep axis, and (3) why bias correction matters most at step 1.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

106. Connect it up: Lesson 15: Modern Optimizers

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — The problem: ill-conditioning · Momentum · RMSprop · Adam · Race the four · AdamW & warmup. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

107. What you can do now

Recap

optimizerthe one thing to remember
momentumv = βv + g — velocity smooths the zig-zag
RMSpropdivide by √(avg g²) — per-axis step size
Adamboth moments + /(1−βᵗ) bias correction
AdamWdecay outside the √v̂ divisor: −ηλθ
warmupramp η up while v̂ is still noisy, then decay

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 15 (Week 5 — Modern Optimizers) — Barron · USAAIO Round 2 Preparation, 2026
  2. Kingma & Ba, Adam: A Method for Stochastic Optimization
  3. Loshchilov & Hutter, Decoupled Weight Decay Regularization (AdamW)
  4. Tieleman & Hinton, RMSprop (Lecture 6e, Coursera Neural Networks) — University of Toronto, 2012
  5. Every trajectory, loss checkpoint, and step magnitude produced by real execution — numpy 2.2.6, verification run July 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108