USAAIO Lesson 15, from Week 5 on optimization, fully worked. It starts by showing why plain gradient descent crawls on an ill-conditioned bowl with a condition number of 100, then derives momentum, RMSprop, and Adam and traces each of them step by step on ONE running problem, f(x) = ½(100x₀² + x₁²). Every update rule is walked out by hand, and bias correction is dissected at t=1, where the raw step is 3.16 times too big while the corrected one is exactly 1.0. All four optimizers are then raced with real loss numbers, and AdamW's decoupled weight decay and learning-rate warmup are derived. In the your-turn project you build momentum, RMSprop, and Adam from scratch. Every snippet runs standalone, and every number came from real numpy execution. The lesson runs to 61 slides.
Subject: Machine Learning · 107 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 15 · Week 5 (Optimization)
Why nobody trains with plain SGD anymore. We take one stretched bowl and race momentum, RMSprop, and Adam across it — each rule derived, hand-traced step by step, and built from scratch. Every loss number here came from real execution.
Objectives
v ← βv + g as a velocity that cancels oscillation and accumulates the consistent directions ← ρs + (1−ρ)g² as a per-parameter adaptive rate that rescales each axist=1 the raw step is 3.16× too big while the corrected step is exactly 1Warm-up
Discussion prompt
Before we open Lesson 15: Modern Optimizers: without looking back, what was the main idea of Information Theory, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
entropy, KL divergence (asymmetric, ≥0), cross-entropy as H(P)+KL, mutual information, and the ELBO that underpins the VAE. Implement entropy, KL, and cross-entropy from scratch and verify the identities.
Section
Part 1 of 8
Concept
Every optimizer today races across the same quadratic bowl — a stretched paraboloid. We minimize:
\[ f(x_0, x_1) = \tfrac12\left(100\,x_0^2 + x_1^2\right) \]
Its gradient is one line, and we start every run from x = [1, 1], where the loss is ½(100 + 1) = 50.5.
\[ \nabla f(x) = \begin{bmatrix} 100\,x_0 \\ x_1 \end{bmatrix}, \qquad x^{(0)} = \begin{bmatrix}1\\1\end{bmatrix},\quad f(x^{(0)}) = 50.5 \]
Counterexample
Discussion prompt
Every optimizer today races across the same quadratic bowl — a stretched paraboloid. We minimize:
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Its gradient is one line, and we start every run from x = [1, 1], where the loss is ½(100 + 1) = 50.5.
Concept
The curvature of this bowl is its Hessian — the matrix of second derivatives. For our f it is diagonal and constant everywhere:
\[ H = \nabla^2 f = \begin{bmatrix} 100 & 0 \\ 0 & 1 \end{bmatrix} \]
condition number κ — The ratio of largest to smallest Hessian eigenvalue, κ = λ_max / λ_min. Here the eigenvalues are 100 and 1, so κ = 100. Big κ = a long, narrow valley — steep across one axis, nearly flat along the other.
κ = 100 is the entire villain of this lesson. Everything modern optimizers do is a response to large κ.
Analogy
Discussion prompt
Explain The Hessian and the condition number by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
The curvature of this bowl is its Hessian — the matrix of second derivatives. For our f it is diagonal and constant everywhere:
Intuition
Picture the surface. Along x₀ the walls shoot up steeply (curvature 100); along x₁ the floor is nearly level (curvature 1). It's a long narrow taco shell, not a round soup bowl.
A ball dropped in rolls hard across the narrow direction and barely drifts down the long one. That mismatch — fast one way, slow the other — is what makes a single global step size fail.
Explain it
Discussion prompt
Explain A taco shell, not a soup bowl to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Picture the surface. Along x₀ the walls shoot up steeply (curvature 100); along x₁ the floor is nearly level (curvature 1). It's a long narrow taco shell, not a round soup bowl.
Concept
Plain gradient descent (GD) takes one global step size η down the negative gradient:
\[ x \leftarrow x - \eta\,\nabla f(x) \]
On a quadratic, GD is stable only while η < 2/L, where L is the largest curvature (the steep axis). Here L = 100, so:
\[ \eta < \frac{2}{L} = \frac{2}{100} = 0.02 \]
The steep axis sets the ceiling for the whole vector
Why: One η must serve both axes. Push η above 0.02 and the steep x₀ axis diverges — so η is hostage to the steepest direction, even though the flat axis would happily take a step 100× larger.
Translation
\( \eta < \frac{2}{L} = \frac{2}{100} = 0.02 \)
Draw it
Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.
Estimation
Predict first
On our separable bowl each axis is its own 1-D problem f = ½Lx² with ∇ = Lx. One GD step multiplies x by a fixed factor — derive it:
Commit before you compute: what does Where the 2/L limit comes from come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Iterating gives xₜ = (1 − ηL)ᵗ x₀
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The factor compounds. The sequence shrinks to 0 only if the factor has magnitude below 1: |1 − ηL| < 1.
Worked example
On our separable bowl each axis is its own 1-D problem f = ½Lx² with ∇ = Lx. One GD step multiplies x by a fixed factor — derive it:
Write one step on a single axis
Why: Substitute the gradient Lx into the update x ← x − ηLx and factor out x.
\[ x_{t+1} = x_t - \eta\,(L x_t) = (1 - \eta L)\,x_t \]
Iterating gives xₜ = (1 − ηL)ᵗ x₀
Why: The factor compounds. The sequence shrinks to 0 only if the factor has magnitude below 1: |1 − ηL| < 1.
\[ |1 - \eta L| < 1 \;\Longleftrightarrow\; 0 < \eta < \frac{2}{L} \]
| axis | L | factor 1−ηL at η=0.018 | behavior |
|---|---|---|---|
| x₀ (steep) | 100 | −0.800 | oscillates, shrinks 20%/step |
| x₁ (flat) | 1 | +0.982 | crawls, shrinks <2%/step |
| x₀ at η=0.021 | 100 | −1.100 | |factor|>1 → DIVERGES |
Comparison
Comparison matrix
From Where the 2/L limit comes from: refill the L column from what you know. The rest of the table is as it appeared.
| axis | L | factor 1−ηL at η=0.018 | behavior |
|---|---|---|---|
| x₀ (steep) | 100 | −0.800 | oscillates, shrinks 20%/step |
| x₁ (flat) | 1 | +0.982 | crawls, shrinks <2%/step |
| x₀ at η=0.021 | 100 | −1.100 | |factor|>1 → DIVERGES |
Missing information
Discussion prompt
Run GD at η = 0.018 (just under the 0.02 ceiling). Because f is separable, each axis updates independently by a fixed factor: x₀ ← (1 − 100η)x₀ and x₁ ← (1 − η)x₁.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
The steep axis multiplies by −0.8 each step: it OSCILLATES in sign (that minus) and shrinks only 20% per step. The flat axis multiplies by 0.982 — it barely moves, losing under 2% per step.
Worked example
Run GD at η = 0.018 (just under the 0.02 ceiling). Because f is separable, each axis updates independently by a fixed factor: x₀ ← (1 − 100η)x₀ and x₁ ← (1 − η)x₁.
import numpy as np
def grad(x): return np.array([100*x[0], x[1]])
x, eta = np.array([1., 1.]), 0.018
for t in range(1, 11):
x = x - eta*grad(x)
loss = 0.5*(100*x[0]**2 + x[1]**2)
print(t, np.round(x, 4), round(loss, 4))x₀ factor is 1 − 100·0.018 = −0.8; x₁ factor is 1 − 0.018 = 0.982
Why: The steep axis multiplies by −0.8 each step: it OSCILLATES in sign (that minus) and shrinks only 20% per step. The flat axis multiplies by 0.982 — it barely moves, losing under 2% per step.
| step t | x₀ | x₁ | loss |
|---|---|---|---|
| 1 | −0.8000 | 0.9820 | 32.4822 |
| 2 | +0.6400 | 0.9643 | 20.9450 |
| 3 | −0.5120 | 0.9470 | 13.5556 |
| 5 | −0.3277 | 0.9132 | 5.7857 |
| 10 | +0.1074 | 0.8339 | 0.9242 |
Trade off
Comparison matrix
From Watch GD crawl — full trace: every row here is a choice with a cost. Fill the loss column, then say which row you would actually pick and what you give up for it.
| step t | x₀ | x₁ | loss |
|---|---|---|---|
| 1 | −0.8000 | 0.9820 | 32.4822 |
| 2 | +0.6400 | 0.9643 | 20.9450 |
| 3 | −0.5120 | 0.9470 | 13.5556 |
| 5 | −0.3277 | 0.9132 | 5.7857 |
| 10 | +0.1074 | 0.8339 | 0.9242 |
Concept
The trace shows both symptoms of ill-conditioning in one run:
Momentum attacks the zig-zag; adaptive methods (RMSprop, Adam) attack the crawl. The next six parts build each fix and race them here.
Section
Part 2 of 8 — a velocity
Intuition
Plain GD is a massless particle: it stops the instant the gradient does, and it turns on a dime every time the gradient flips. That's why it zig-zags.
Momentum gives the ball mass. Across the narrow valley the pushes alternate left–right and largely cancel; down the long axis the pushes all point the same way and accumulate into speed. The ball smooths the bounce and rolls straight down the valley.
Concept
Keep a velocity v — an exponential moving average of past gradients — and step along it instead of along the raw gradient:
\[ \begin{aligned} v &\leftarrow \beta\, v + \nabla f(x) \\ x &\leftarrow x - \eta\, v \end{aligned} \]
The friction/decay β (typically 0.9) says how much old velocity survives each step. β = 0 recovers plain GD; larger β gives the ball more inertia. We start v at the zero vector.
Concept
Unroll the EMA: v is a decaying sum of every past gradient, v_t = Σ βᵏ g_{t−k}. On the steep axis the g's alternate sign, so the sum partly cancels — velocity stays small.
On the flat axis every g has the same sign, so the terms reinforce. A geometric series of same-sign pushes sums toward 1/(1−β) = 10× the single-step pull. The flat axis effectively gets a 10× boost — exactly where GD was crawling.
Pattern
Predict first
The table runs: 1 | 100.000 | 1.000 | 24.9970 · 2 | 160.000 | 1.897 | 2.9113 · 4 | 121.600 | 3.412 | 21.1329 · 6 | −37.184 | 4.600 | 22.6748
In Momentum, hand-traced (β=0.9, η=0.003), given the rows so far: what is the next one — the row where step t is 8?
Correct: 8 | −126.756 | 5.510 | 0.4286
| step t | v[0] | v[1] | loss |
|---|---|---|---|
| 1 | 100.000 | 1.000 | 24.9970 |
| 2 | 160.000 | 1.897 | 2.9113 |
| 4 | 121.600 | 3.412 | 21.1329 |
| 6 | −37.184 | 4.600 | 22.6748 |
| 8 | −126.756 | 5.510 | 0.4286 |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. The flat-axis velocity accumulates monotonically (same-sign pushes), so x₁ finally makes progress.
Worked example
Trace the velocity accumulating. Watch v[1] (flat axis) grow steadily from 1.0 while v[0] (steep axis) swings and averages toward zero.
import numpy as np
def grad(x): return np.array([100*x[0], x[1]])
x, v = np.array([1., 1.]), np.zeros(2)
for t in range(1, 9):
g = grad(x)
v = 0.9*v + g
x = x - 0.003*v
print(t, np.round(v, 3), np.round(x, 4))v[1] climbs 1.00 → 1.90 → 2.70 → …; v[0] swings 100 → 160 → 166 → 122 → 45 → −37
Why: The flat-axis velocity accumulates monotonically (same-sign pushes), so x₁ finally makes progress. The steep-axis velocity peaks then reverses as the sign flips — the oscillation is being averaged out inside v.
| step t | v[0] | v[1] | loss |
|---|---|---|---|
| 1 | 100.000 | 1.000 | 24.9970 |
| 2 | 160.000 | 1.897 | 2.9113 |
| 4 | 121.600 | 3.412 | 21.1329 |
| 6 | −37.184 | 4.600 | 22.6748 |
| 8 | −126.756 | 5.510 | 0.4286 |
Pattern
Step through it
Step through Momentum, hand-traced (β=0.9, η=0.003) one row at a time. What is driving the change, and what would the row after the last one be?
Intuition
Notice the loss in that trace is not monotone: 25.0 → 2.9 → … → 21.1 → 22.7 → 0.43. A heavy ball builds speed on the steep axis and coasts past the bottom before inertia turns it around.
That's the price of mass: momentum can overshoot early, then converges fast once the transient settles. By step 80 it reaches loss 0.0008 — far past where GD stalled.
Concept
You will see momentum written two ways, and they are not the same. Ours accumulates the raw gradient into v; the other (PyTorch's SGD(momentum=...)) blends with a (1−β) weight:
\[ \underbrace{v \leftarrow \beta v + g}_{\text{ours (raw sum)}} \qquad\text{vs}\qquad \underbrace{v \leftarrow \beta v + (1-\beta)\,g}_{\text{normalized average}} \]
The normalized form keeps v on the scale of a single gradient, so it needs an η about 1/(1−β) = 10× larger to match. Same trajectory, different η bookkeeping — know which one your library uses before you tune.
Section
Part 3 of 8 — per-axis rates
Concept
RMSprop averages g², then divides by √s. Why the square-then-root instead of just averaging |g|? Because s is a running mean of squares, so √s is a root-mean-square — that's the RMS in the name.
Squaring is also what connects it to curvature: on a quadratic, the typical squared gradient on an axis scales with that axis's curvature, so √s estimates the local steepness and g/√s is a Newton-like, curvature-normalized step — without ever forming the Hessian.
Intuition
A true second-order method (Lesson 21) divides the step by the curvature itself — but that needs the Hessian, which is huge and expensive for real networks.
RMSprop fakes it: √s is a running, per-coordinate proxy for curvature built only from gradients you already computed. You get most of the axis-rescaling benefit of a second-order method at first-order cost. That trade is why adaptive methods dominate deep learning.
Intuition
Momentum still uses one global η. But the real disease is that the two axes want wildly different step sizes — the steep one small, the flat one large.
RMSprop's idea: measure how big each axis's gradients typically are, and divide each coordinate's step by that size. Loud axes get quieted; quiet axes get amplified. Every coordinate ends up moving at a comparable pace.
Sorting
Sort into buckets
These are the pieces of Lesson 15: Modern Optimizers, out of order. Put each one back under the part of the lesson it belongs to.
Concept
Track a running average of squared gradients, coordinate by coordinate, then divide the step by its square root:
\[ \begin{aligned} s &\leftarrow \rho\, s + (1-\rho)\, g^{\odot 2} \\ x &\leftarrow x - \eta\,\frac{g}{\sqrt{s} + \epsilon} \end{aligned} \]
Here g² and the division are elementwise (that's the ⊙). Defaults ρ = 0.9, ε = 1e-8 (a floor that stops division by zero). s starts at the zero vector.
√s is a per-coordinate estimate of gradient magnitude, so g/√s is roughly a unit-size step on every axis — the rescaling an ill-conditioned bowl needs.
Fill the middle
Fill in the blanks
From RMSprop, hand-traced (ρ=0.9, η=0.05) — one line has had its right-hand side removed. Put it back.
import numpy as np
def grad(x): return np.array([100*x[0], x[1]])
x, s = np.array([1., 1.]), np.zeros(2)
for t in range(1, 6):
g = grad(x)
s = 0.9s + 0.1g*g
step = g/(np.sqrt(s) + 1e-8)
x = x - 0.05*step
print(t, np.round(s, 3), np.round(step, 4), np.round(x, 5))
Why: step is what everything below it consumes, so the wrong expression here fails later and somewhere else. s[0]=0.1·100²=1000 and s[1]=0.1·1²=0.1.
Worked example
At step 1 the raw gradient is [100, 1] — a 100:1 mismatch. Watch RMSprop erase it: after dividing by √s, both axes move by the identical amount.
import numpy as np
def grad(x): return np.array([100*x[0], x[1]])
x, s = np.array([1., 1.]), np.zeros(2)
for t in range(1, 6):
g = grad(x)
s = 0.9*s + 0.1*g*g
step = g/(np.sqrt(s) + 1e-8)
x = x - 0.05*step
print(t, np.round(s, 3), np.round(step, 4), np.round(x, 5))step 1: s = [1000, 0.1], and step = g/√s = [3.162, 3.162] — identical on both axes
Why: s[0]=0.1·100²=1000 and s[1]=0.1·1²=0.1. Then 100/√1000 = 1/√0.1 = 3.162. The 100:1 gradient ratio is completely cancelled: each axis moves by 0.05·3.162 = 0.1581, so both x₀ and x₁ go 1 → 0.84189.
| step t | s[0] | step[0]=step[1] | x₀ = x₁ | loss |
|---|---|---|---|---|
| 1 | 1000.000 | 3.1623 | 0.84189 | 35.7930 |
| 2 | 1608.772 | 2.0990 | 0.73694 | 27.4254 |
| 3 | 1990.972 | 1.6516 | 0.65436 | 21.6234 |
| 5 | 2340.186 | 1.2091 | 0.52446 | 13.8906 |
Scale up
Step through it
Step through RMSprop, hand-traced (ρ=0.9, η=0.05) and watch the numbers move. Now imagine the input ten times bigger: which column is the one that stops this being practical?
Concept
Because both axes descend in lockstep, the crawl is gone. By step 20 the loss is 0.25, by step 40 it is ≈ 1.5e-7, and by step 80 it has effectively hit machine zero — a completely different regime from GD's 0.03.
The catch: RMSprop rescales but has no memory of direction — it doesn't smooth the path. Combine its per-axis rescaling with momentum's smoothing and you get Adam.
Section
Part 4 of 8 — assemble the pieces
Intuition
Momentum fixes the direction of the step (smooth the zig-zag by averaging gradients over time). RMSprop fixes the scale of the step (rescale each axis by its own gradient size). These are independent knobs.
So there's nothing stopping you from turning both at once: average the gradient for direction, average the squared gradient for scale. That combination is Adam — and the only new subtlety is that both averages start biased at zero.
Concept
Adam keeps both running averages: a first moment m (momentum's velocity) and a second moment v (RMSprop's squared-gradient average).
\[ \begin{aligned} m &\leftarrow \beta_1\, m + (1-\beta_1)\, g \\ v &\leftarrow \beta_2\, v + (1-\beta_2)\, g^{\odot 2} \end{aligned} \]
Both start at the zero vector. Defaults β₁ = 0.9, β₂ = 0.999, ε = 1e-8. Note m uses the (1−β₁) weighting (a true average), unlike our bare momentum — a cosmetic difference that just rescales η.
Concept
Because m and v start at 0, early estimates are pulled toward zero. At t=1, with β₂ = 0.999:
\[ v_1 = 0.999\cdot 0 + 0.001\, g^2 = 0.001\,g^2 \]
That is 1000× too small as an estimate of g²! Its square root, √v ≈ 0.0316\,|g|, is 31.6× too small — so dividing by it would make the step 31.6× too large. The estimate hasn't warmed up yet.
Faded example
Fill in the blanks
Bias correction, dissected at t=1, with the scaffolding fading: two lines are gone now — fill both.
import numpy as np
g = 100.0 # steep-axis gradient at t=1
m, v = 0.1g, 0.001g*g # m=(1-b1)g, v=(1-b2)g^2
raw = m/(np.sqrt(v) + 1e-8) # NO correction
mh, vh = m/(1-0.91), v/(1-0.9991) # bias-corrected
cor = mh/(np.sqrt(vh) + 1e-8)
print('raw step dir:', round(raw, 4))
print('corrected :', round(cor, 4))
Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. Correction rescales m by 1/(1−0.9)=10× and v by 1/(1−0.999)=1000×.
Worked example
The fix: divide each moment by (1 − βᵗ) to undo the shrink. Watch what it does to the actual step magnitude on the steep axis at t=1, where g₀ = 100:
import numpy as np
g = 100.0 # steep-axis gradient at t=1
m, v = 0.1*g, 0.001*g*g # m=(1-b1)g, v=(1-b2)g^2
raw = m/(np.sqrt(v) + 1e-8) # NO correction
mh, vh = m/(1-0.9**1), v/(1-0.999**1) # bias-corrected
cor = mh/(np.sqrt(vh) + 1e-8)
print('raw step dir:', round(raw, 4))
print('corrected :', round(cor, 4))raw = m/√v = 3.1623, but corrected = m̂/√v̂ = 1.0000
Why: Correction rescales m by 1/(1−0.9)=10× and v by 1/(1−0.999)=1000×. After correction m̂=g and v̂=g², so m̂/√v̂ = g/|g| = sign(g), magnitude exactly 1. Uncorrected it is 3.16× too big — the un-warmed v underestimates g².
| quantity at t=1 | raw (no BC) | corrected |
|---|---|---|
| m (first moment) | 10.0 | m̂ = 100.0 |
| v (second moment) | 10.0 | v̂ = 10000.0 |
| step dir = m/√v | 3.1623 | 1.0000 |
| effective step η·(…), η=0.1 | 0.3162 | 0.1000 |
Discrimination
Sort into buckets
Sort these by raw (no BC), from memory, without looking back at Bias correction, dissected at t=1. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Concept
Correct both moments with their own factor, then take the RMSprop-style step:
\[ \hat m = \frac{m}{1-\beta_1^{\,t}}, \qquad \hat v = \frac{v}{1-\beta_2^{\,t}}, \qquad x \leftarrow x - \eta\,\frac{\hat m}{\sqrt{\hat v} + \epsilon} \]
As t grows, βᵗ → 0, so (1 − βᵗ) → 1 and the correction quietly fades away — it only matters in the first dozen-or-so steps. Adam is the default optimizer for nearly all deep learning.
Intuition
Here's why Adam is so forgiving to tune. The update per coordinate is η·m̂/√v̂, and both m̂ and √v̂ are averages of the same gradients — so their ratio is roughly sign(g), of magnitude near 1.
That means each parameter moves by about ±η regardless of its gradient's scale. A coordinate with a gradient of 100 and one with a gradient of 0.001 both step by ~η. The raw gradient magnitude is divided out — which is exactly the cure for a 100:1 bowl, and why one η now works for all axes.
Worked example
Full Adam on the bowl. The bias-corrected m̂[0] starts at exactly 100 (the true gradient), and the step marches both axes down together — no zig-zag, no crawl.
import numpy as np
def grad(x): return np.array([100*x[0], x[1]])
x, m, v = np.array([1., 1.]), np.zeros(2), np.zeros(2)
for t in range(1, 11):
g = grad(x)
m = 0.9*m + 0.1*g
v = 0.999*v + 0.001*g*g
mh, vh = m/(1-0.9**t), v/(1-0.999**t)
x = x - 0.1*mh/(np.sqrt(vh) + 1e-8)
print(t, np.round(x, 5), round(0.5*(100*x[0]**2+x[1]**2), 4))x moves in lockstep on both axes: [0.9,0.9] → [0.8,0.8] → … → [0.076,0.076]
Why: Because m̂/√v̂ ≈ sign(g) early, each axis takes an ≈η=0.1 step regardless of its 100:1 gradient gap. By step 10 the loss is 0.294 — already below where GD sat at step 10 (0.924).
| step t | m̂[0] | x₀ = x₁ | loss |
|---|---|---|---|
| 1 | 100.000 | 0.90000 | 40.9050 |
| 3 | 89.314 | 0.70159 | 24.8573 |
| 6 | 72.227 | 0.41424 | 8.6654 |
| 10 | 48.372 | 0.07625 | 0.2936 |
Pattern
Step through it
Step through Adam, hand-traced (η=0.1) one row at a time. What is driving the change, and what would the row after the last one be?
Anomaly
Predict first
A student writes this, and it looks reasonable:
m and v are moving averages of the gradient, so just plug them straight into the update — the /(1−βᵗ) looks like decoration.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: At t=1, v = 0.001·g² underestimates g² by 1000×, so √v is 31.6× too small and the step direction blows up to 3.16 instead of 1.
Divide each moment by (1 − βᵗ) to undo the zero-init shrink before you step.
Why: At t=1, v = 0.001·g² underestimates g² by 1000×, so √v is 31.6× too small and the step direction blows up to 3.16 instead of 1. The un-warmed second moment makes the earliest steps unreliably large.
Trap
m and v are moving averages of the gradient, so just plug them straight into the update — the /(1−βᵗ) looks like decoration.
x -= η·m/(√v + ε), no correction
Why: At t=1, v = 0.001·g² underestimates g² by 1000×, so √v is 31.6× too small and the step direction blows up to 3.16 instead of 1. The un-warmed second moment makes the earliest steps unreliably large.
On our bowl it converges worse: loss 1.12 at step 30
Why: Verified: uncorrected Adam (η=0.1) sits at loss 1.12 by step 30, while corrected Adam is at 0.058 — the bad early steps cost you the whole run.
Divide each moment by (1 − βᵗ) to undo the zero-init shrink before you step.
m̂ = m/(1−β₁ᵗ), v̂ = v/(1−β₂ᵗ), then step
Why: At t=1 this rescales v by 1000× and m by 10×, so m̂/√v̂ = sign(g), a clean unit step. The correction fades as t grows (βᵗ→0), so it only touches the fragile early steps.
Corrected Adam reaches loss 0.058 by step 30
Why: The right early-step magnitude keeps the trajectory on track. Verified against real execution.
Break the constraint
Discussion prompt
The rule this trap just fixed:
At t=1 this rescales v by 1000× and m by 10×, so m̂/√v̂ = sign(g), a clean unit step. The correction fades as t grows (βᵗ→0), so it only touches the fragile early steps.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
At t=1, v = 0.001·g² underestimates g² by 1000×, so √v is 31.6× too small and the step direction blows up to 3.16 instead of 1. The un-warmed second moment makes the earliest steps unreliably large.
Fill the middle
Fill in the blanks
From Bias correction on vs off — head to head — one line has had its right-hand side removed. Put it back.
import numpy as np
def grad(x): return np.array([100*x[0], x[1]])
def loss(x): return 0.5(100x[0]2 + x[1]2)
def adam(bc, steps=30, lr=0.1, b1=0.9, b2=0.999, eps=1e-8):
x, m, v = np.array([1.,1.]), np.zeros(2), np.zeros(2)
hist = *b2v + (1-b2)gg**
for t in range(1, steps+1):
g = grad(x)
m = b1m + (1-b1)g
v = ___
if bc: mh, vh = m/(1-b1t), v/(1-b2t)
else: mh, vh = m, v
x = x - lr*mh/(np.sqrt(vh)+eps)
hist[t] = loss(x)
return hist
on, off = adam(True), adam(False)
print('with BC :', [round(on[k], 4) for k in (1, 5, 30)])
print('no BC :', [round(off[k], 4) for k in (1, 5, 30)])
Why: v is what everything below it consumes, so the wrong expression here fails later and somewhere else. The uncorrected step direction is 3.16× too big at t=1, which happens to cut the first loss more — a false win.
Worked example
One flag toggles the /(1−βᵗ) terms. Run both variants on the bowl at η = 0.1 and compare the loss at three checkpoints.
import numpy as np
def grad(x): return np.array([100*x[0], x[1]])
def loss(x): return 0.5*(100*x[0]**2 + x[1]**2)
def adam(bc, steps=30, lr=0.1, b1=0.9, b2=0.999, eps=1e-8):
x, m, v = np.array([1.,1.]), np.zeros(2), np.zeros(2)
hist = {}
for t in range(1, steps+1):
g = grad(x)
m = b1*m + (1-b1)*g
v = b2*v + (1-b2)*g*g
if bc: mh, vh = m/(1-b1**t), v/(1-b2**t)
else: mh, vh = m, v
x = x - lr*mh/(np.sqrt(vh)+eps)
hist[t] = loss(x)
return hist
on, off = adam(True), adam(False)
print('with BC :', [round(on[k], 4) for k in (1, 5, 30)])
print('no BC :', [round(off[k], 4) for k in (1, 5, 30)])no-BC starts smaller (23.6 vs 40.9 at t=1) but ends WORSE (1.12 vs 0.058 at t=30)
Why: The uncorrected step direction is 3.16× too big at t=1, which happens to cut the first loss more — a false win. Those oversized early steps mistrack the valley and the run never recovers, ending 19× worse by step 30.
| variant | loss @1 | @5 | @30 |
|---|---|---|---|
| with bias correction | 40.905 | 13.030 | 0.058 |
| no bias correction | 23.611 | 23.082 | 1.120 |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
no-BC starts smaller (23.6 vs 40.9 at t=1) but ends WORSE (1.12 vs 0.058 at t=30)
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
One flag toggles the /(1−βᵗ) terms. Run both variants on the bowl at η = 0.1 and compare the loss at three checkpoints.
Section
Part 5 of 8 — same bowl, same start
Estimation
Predict first
Same problem, same start [1,1], start loss 50.5. Each optimizer uses its own tuned η. This block runs GD, momentum, RMSprop, and Adam and prints the loss at four checkpoints.
Commit before you compute: what does All four optimizers, one bowl come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: GD stalls near 0.03; momentum & Adam reach ~0.001; RMSprop hits machine zero
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. GD is capped by the steep axis and never escapes the slow flat-axis regime.
Worked example
Same problem, same start [1,1], start loss 50.5. Each optimizer uses its own tuned η. This block runs GD, momentum, RMSprop, and Adam and prints the loss at four checkpoints.
import numpy as np
def grad(x): return np.array([100*x[0], x[1]])
def loss(x): return 0.5*(100*x[0]**2 + x[1]**2)
def adam(steps, lr=0.1, b1=0.9, b2=0.999, eps=1e-8):
x, m, v = np.array([1.,1.]), np.zeros(2), np.zeros(2)
hist = {}
for t in range(1, steps+1):
g = grad(x)
m = b1*m + (1-b1)*g
v = b2*v + (1-b2)*g*g
mh, vh = m/(1-b1**t), v/(1-b2**t)
x = x - lr*mh/(np.sqrt(vh)+eps)
hist[t] = loss(x)
return hist
h = adam(80)
print([round(h[k], 4) for k in (10, 20, 40, 80)])GD stalls near 0.03; momentum & Adam reach ~0.001; RMSprop hits machine zero
Why: GD is capped by the steep axis and never escapes the slow flat-axis regime. The adaptive/momentum methods give the flat axis the large effective step it always needed. Adam prints [0.2936, 3.713, 0.513, 0.0015].
| optimizer (η) | loss @10 | @20 | @40 | @80 |
|---|---|---|---|---|
| GD (0.018) | 0.9242 | 0.2484 | 0.1169 | 0.0273 |
| momentum (0.003) | 15.55 | 1.932 | 0.3441 | 0.0008 |
| RMSprop (0.05) | 4.587 | 0.2512 | ≈1.5e-7 | ≈0 |
| Adam (0.1) | 0.2936 | 3.713 | 0.5130 | 0.0015 |
Concept
The table isn't a clean 'Adam wins' story, and that's the point. Momentum and Adam both overshoot mid-run (momentum 15.5 at step 10; Adam 3.7 at step 20) as their velocity coasts past the bottom, then recover.
RMSprop, with no momentum to overshoot, descends smoothly straight to machine zero on this pure quadratic. Every method beats GD's 0.027 floor by orders of magnitude — because every method frees the flat axis.
Concept
The loss curve reveals each method's personality on the same bowl:
0.03.15.5) as the ball overshoots, then a fast plunge to 0.0008.0.29 → 3.7) from its momentum half, then it settles like RMSprop.The lesson: momentum buys speed at the cost of an early overshoot; adaptivity buys smoothness. Adam takes both, so it inherits a little of both traits.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Adam has the most machinery — momentum AND adaptive rates AND bias correction — so it must always converge fastest and best.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: On our own bowl Adam overshoots (loss 0.29 → 3.71 between steps 10 and 20) and RMSprop actually reaches a lower loss.
Treat Adam as a robust default, not a guaranteed winner — match the optimizer and its schedule to the problem.
Why: On our own bowl Adam overshoots (loss 0.29 → 3.71 between steps 10 and 20) and RMSprop actually reaches a lower loss. 'Most components' does not mean 'always best'.
Trap
Adam has the most machinery — momentum AND adaptive rates AND bias correction — so it must always converge fastest and best.
Reach for Adam everywhere and never tune
Why: On our own bowl Adam overshoots (loss 0.29 → 3.71 between steps 10 and 20) and RMSprop actually reaches a lower loss. 'Most components' does not mean 'always best'.
Ignore that η and the schedule still matter
Why: Adam with a bad η diverges like anything else; the warmup/decay schedule often decides success more than the optimizer choice does.
Treat Adam as a robust default, not a guaranteed winner — match the optimizer and its schedule to the problem.
Adam/AdamW for transformers & sparse gradients; well-tuned SGD+momentum for many CNNs
Why: Each has regimes where it wins: on some vision tasks a tuned SGD+momentum generalizes better than Adam. Pick per problem, and tune η.
Always pair the optimizer with a schedule
Why: Warmup then cosine decay stabilizes and sharpens convergence — the schedule is part of the choice, not an afterthought.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
f it is diagonal and constant everywhere:; Plain gradient descent (GD) takes one global step size η down the negative gradient:m and v are moving averages of the gradient, so just plug them straight into the update — the /(1−βᵗ) looks like decoration.; Adam has the most machinery — momentum AND adaptive rates AND bias correction — so it must always converge fastest and best.Section
Part 6 of 8 — the modern defaults
Concept
Weight decay shrinks each weight toward 0 every step — the L2 regularizer of Lesson 12. The classic way folds it into the gradient as g + λθ, so it flows through Adam's adaptive √v̂ divisor.
\[ \text{Adam+L2:}\quad g \leftarrow \nabla f + \lambda\theta, \qquad x \leftarrow x - \eta\,\frac{\hat m}{\sqrt{\hat v}+\epsilon} \]
The problem: dividing the decay by √v̂ means weights on loud (large-gradient) axes decay less than weights on quiet ones. The regularization strength gets distorted per-coordinate — not what you asked for.
Concept
AdamW applies the decay as its own term, outside the adaptive step, so every weight decays by the same fraction η·λ:
\[ x \leftarrow x - \eta\,\frac{\hat m}{\sqrt{\hat v}+\epsilon} - \eta\,\lambda\,\theta \]
The decay η·λ·θ no longer passes through √v̂, so it is decoupled from the gradient geometry. This is the standard recipe for training transformers.
Faded example
Fill in the blanks
AdamW vs Adam+L2 differ on step one, with the scaffolding fading: two lines are gone now — fill both.
import numpy as np
def grad(x): return np.array([100*x[0], x[1]])
b1,b2,eta,eps,lam = 0.9,0.999,0.1,1e-8,0.1
# Adam + L2: decay folded into the gradient
x = np.array([1.,1.]); g = grad(x) + lam*x
m, v = (1-b1)g, (1-b2)g*g
mh, vh = m/(1-b1), v/(1-b2)
x_l2 = *x - etamh/(np.sqrt(vh)+eps)
# AdamW: decay applied separately
x = np.array([1.,1.]); g = grad(x)**
m, v = (1-b1)g, (1-b2)g*g
mh, vh = m/(1-b1), v/(1-b2)
x_w = x - etamh/(np.sqrt(vh)+eps) - etalam*x
print('Adam+L2:', np.round(x_l2,4), ' AdamW:', np.round(x_w,4))
Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. In Adam+L2 the λθ term is swallowed by the √v̂ normalization (m̂/√v̂ ≈ sign, so the decay's size is lost).
Worked example
One step from θ = [1, 1] with λ = 0.1, η = 0.1. The two recipes land on different points — proof the placement matters.
import numpy as np
def grad(x): return np.array([100*x[0], x[1]])
b1,b2,eta,eps,lam = 0.9,0.999,0.1,1e-8,0.1
# Adam + L2: decay folded into the gradient
x = np.array([1.,1.]); g = grad(x) + lam*x
m, v = (1-b1)*g, (1-b2)*g*g
mh, vh = m/(1-b1), v/(1-b2)
x_l2 = x - eta*mh/(np.sqrt(vh)+eps)
# AdamW: decay applied separately
x = np.array([1.,1.]); g = grad(x)
m, v = (1-b1)*g, (1-b2)*g*g
mh, vh = m/(1-b1), v/(1-b2)
x_w = x - eta*mh/(np.sqrt(vh)+eps) - eta*lam*x
print('Adam+L2:', np.round(x_l2,4), ' AdamW:', np.round(x_w,4))Adam+L2 → [0.90, 0.90]; AdamW → [0.89, 0.89]; they differ by 0.01
Why: In Adam+L2 the λθ term is swallowed by the √v̂ normalization (m̂/√v̂ ≈ sign, so the decay's size is lost). In AdamW the −η·λ·θ = −0.01 is applied cleanly on top. Same λ, different weight decay.
| recipe | x₀ after 1 step | x₁ after 1 step |
|---|---|---|
| Adam + L2 (folded) | 0.9000 | 0.9000 |
| AdamW (decoupled) | 0.8900 | 0.8900 |
| difference | 0.0100 | 0.0100 |
Comparison
Comparison matrix
From AdamW vs Adam+L2 differ on step one: refill the x₀ after 1 step column from what you know. The rest of the table is as it appeared.
| recipe | x₀ after 1 step | x₁ after 1 step |
|---|---|---|
| Adam + L2 (folded) | 0.9000 | 0.9000 |
| AdamW (decoupled) | 0.8900 | 0.8900 |
| difference | 0.0100 | 0.0100 |
Concept
We saw the flip side of bias correction: very early, the second moment v̂ is built from just a few gradients and is noisy. Even corrected, a big η on a shaky estimate can destabilize training.
Warmup ramps η linearly from ≈0 over the first few hundred steps, then cosine-anneals it back down toward 0. The ramp buys time for v̂ to settle before full-size steps hit; transformers essentially require it.
Explain it
Discussion prompt
Explain Learning-rate warmup to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
We saw the flip side of bias correction: very early, the second moment v̂ is built from just a few gradients and is noisy. Even corrected, a big η on a shaky estimate can destabilize training.
Worked example
Linear ramp over the first W steps to base, then a cosine decay to 0 across the remaining T − W. This block prints the η the optimizer would use at several steps.
import math
def lr_sched(t, base=0.1, W=10, T=80):
if t < W:
return base * (t+1)/W # linear warmup
prog = (t - W)/(T - W)
return base * 0.5*(1 + math.cos(math.pi*prog)) # cosine decay
for t in [0, 1, 4, 9, 20, 40, 79]:
print(t, round(lr_sched(t), 5))η climbs 0.01 → 0.10 across the first 10 steps, then decays 0.10 → ~0 by step 79
Why: During warmup (t<10) the ramp keeps steps small while v̂ is noisy; at t=9 it reaches the base 0.10; then cosine annealing shrinks it to 0.00005 near the end for a gentle landing.
| step t | phase | η |
|---|---|---|
| 0 | warmup | 0.01000 |
| 9 | warmup end | 0.10000 |
| 20 | cosine decay | 0.09505 |
| 40 | cosine decay | 0.06113 |
| 79 | near end | 0.00005 |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
η climbs 0.01 → 0.10 across the first 10 steps, then decays 0.10 → ~0 by step 79
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Linear ramp over the first W steps to base, then a cosine decay to 0 across the remaining T − W. This block prints the η the optimizer would use at several steps.
Intuition
Both target the fragile first few steps, where v̂ is estimated from almost no data. Bias correction fixes the systematic shrink (the 1000× underestimate at t=1); warmup handles the random noise that correction can't remove.
Correction makes the early step the right size on average; warmup keeps it small so a single unlucky, noisy v̂ can't send the weights flying before the averages settle. On a huge transformer, one bad early step can poison the whole run — so you use both.
Section
Part 7 of 8 — pattern & checks
Ranking
Put in order
These are the steps of The optimizer toolkit, scrambled. Put them back in order before the next slide shows you.
v = βv + g; x -= ηv. A velocity EMA: cancels oscillation, accumulates the consistent direction. β ≈ 0.9.s = ρs + (1−ρ)g²; x -= ηg/(√s+ε). Per-parameter rate: divide each axis by its own gradient size.m̂ = m/(1−β₁ᵗ), v̂ = v/(1−β₂ᵗ), then x -= η m̂/(√v̂+ε).x -= η m̂/(√v̂+ε) − ηλθ (decay outside the adaptive divisor).η up, then cosine-decay it down; near-mandatory for transformers.Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
v = βv + g; x -= ηv. A velocity EMA: cancels oscillation, accumulates the consistent direction. β ≈ 0.9.s = ρs + (1−ρ)g²; x -= ηg/(√s+ε). Per-parameter rate: divide each axis by its own gradient size.m̂ = m/(1−β₁ᵗ), v̂ = v/(1−β₂ᵗ), then x -= η m̂/(√v̂+ε).x -= η m̂/(√v̂+ε) − ηλθ (decay outside the adaptive divisor).η up, then cosine-decay it down; near-mandatory for transformers.Edge cases
Discussion prompt
The optimizer toolkit works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
v = βv + g; x -= ηv. A velocity EMA: cancels oscillation, accumulates the consistent direction. β ≈ 0.9.s = ρs + (1−ρ)g²; x -= ηg/(√s+ε). Per-parameter rate: divide each axis by its own gradient size.m̂ = m/(1−β₁ᵗ), v̂ = v/(1−β₂ᵗ), then x -= η m̂/(√v̂+ε).x -= η m̂/(√v̂+ε) − ηλθ (decay outside the adaptive divisor).η up, then cosine-decay it down; near-mandatory for transformers.Elimination
Eliminate the wrong options
On f(x) = ½(100x₀² + x₁²), why does plain gradient descent barely move the x₁ axis?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Stability requires η < 2/L with L = 100 (the steep curvature). That same small η, applied to the flat axis whose curvature is only 1, shrinks x₁ by a factor of just (1 − η) ≈ 0.982 per step — a crawl. One step size cannot serve a 100:1 curvature gap.
Check
Think about what caps the step size on our bowl.
Check your understanding
On f(x) = ½(100x₀² + x₁²), why does plain gradient descent barely move the x₁ axis?
Answer: A
Why: Stability requires η < 2/L with L = 100 (the steep curvature). That same small η, applied to the flat axis whose curvature is only 1, shrinks x₁ by a factor of just (1 − η) ≈ 0.982 per step — a crawl. One step size cannot serve a 100:1 curvature gap.
Prediction
Predict first
Adam combines which two mechanisms, plus a correction?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Momentum (first moment m) + RMSprop (second moment v), with bias correction /(1−βᵗ)
Why: Adam tracks a first moment m (the momentum velocity) and a second moment v (RMSprop's squared-gradient average), then bias-corrects both with /(1−βᵗ) because they start at zero. That's the entire recipe.
Check
What two ideas is Adam built from?
Check your understanding
Adam combines which two mechanisms, plus a correction?
Answer: A
Why: Adam tracks a first moment m (the momentum velocity) and a second moment v (RMSprop's squared-gradient average), then bias-corrects both with /(1−βᵗ) because they start at zero. That's the entire recipe.
Prediction
Predict first
Without bias correction, what goes wrong at Adam's very first step (β₂ = 0.999)?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: v = 0.001g² underestimates g², so √v is too small and the step is too large / unreliable
Why: At t=1, v = (1−0.999)g² = 0.001g², about 1000× below g². Its root √v ≈ 0.0316|g| is 31.6× too small, so g/√v ≈ 3.16·sign(g) instead of 1 — the first step is ~3× too big. Correction divides v by (1−0.999) = 0.001, restoring v̂ = g² and a unit-size step.
Check
Recall what happens at t=1 with β₂ = 0.999.
Check your understanding
Without bias correction, what goes wrong at Adam's very first step (β₂ = 0.999)?
Answer: A
Why: At t=1, v = (1−0.999)g² = 0.001g², about 1000× below g². Its root √v ≈ 0.0316|g| is 31.6× too small, so g/√v ≈ 3.16·sign(g) instead of 1 — the first step is ~3× too big. Correction divides v by (1−0.999) = 0.001, restoring v̂ = g² and a unit-size step.
Elimination
Eliminate the wrong options
How does AdamW differ from folding L2 into the gradient (Adam+L2)?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: AdamW's update is x ← x − η·m̂/(√v̂+ε) − η·λ·θ. The decay term is decoupled from the adaptive step, so it isn't distorted by √v̂ and each weight shrinks by the same fraction ηλ. On our one-step demo AdamW landed at 0.89 vs Adam+L2's 0.90.
Check
Where does AdamW put the weight decay?
Check your understanding
How does AdamW differ from folding L2 into the gradient (Adam+L2)?
Answer: A
Why: AdamW's update is x ← x − η·m̂/(√v̂+ε) − η·λ·θ. The decay term is decoupled from the adaptive step, so it isn't distorted by √v̂ and each weight shrinks by the same fraction ηλ. On our one-step demo AdamW landed at 0.89 vs Adam+L2's 0.90.
Section
Part 8 of 8 — the project
Concept
Implement all three update rules yourself and run each on the same ill-conditioned bowl f(x) = ½(100x₀² + x₁²) from [1, 1]. You've traced every rule by hand — now assemble them in code.
| # | requirement | key line |
|---|---|---|
| 1 | momentum: velocity EMA | v = 0.9v + g; x -= lrv |
| 2 | RMSprop: per-axis rate | s = 0.9s + 0.1gg; x -= lrg/(√s+ε) |
| 3 | Adam: two moments + bias correction | mh=m/(1-b1t); vh=v/(1-b2t) |
Build rules: type every line yourself, initialize all state to zeros, and for Adam loop t from 1 and never drop the /(1−βᵗ) correction. Run after each milestone and read the loss.
Analogy
Discussion prompt
Explain Project: momentum, RMSprop, Adam from scratch by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Implement all three update rules yourself and run each on the same ill-conditioned bowl f(x) = ½(100x₀² + x₁²) from [1, 1]. You've traced every rule by hand — now assemble them in code.
Fill the middle
Fill in the blanks
From Milestone 1 — momentum — one line has had its right-hand side removed. Put it back.
import numpy as np
def grad(x): return np.array([100*x[0], x[1]])
x, v, lr = np.array([1., 1.]), np.zeros(2), 0.003
for t in range(1, 81):
g = grad(x)
v = *0.9v + g**
x = x - lr*v
if t in (10, 40, 80):
print(t, round(0.5(100x[0]2 + x[1]2), 4))
Why: v is what everything below it consumes, so the wrong expression here fails later and somewhere else. The heavy ball builds speed on the steep axis and coasts past the minimum around step 10, then inertia turns it and it settles well below GD's 0.027 by step 80.
Worked example
Your turn: implement momentum and run 80 steps at η = 0.003. Predict aloud whether it beats GD's step-80 loss of 0.027 before you print.
Hint: keep v = np.zeros(2); each step g = grad(x); v = 0.9*v + g; x = x - lr*v. Expect an early overshoot in the loss, then a fast finish.
import numpy as np
def grad(x): return np.array([100*x[0], x[1]])
x, v, lr = np.array([1., 1.]), np.zeros(2), 0.003
for t in range(1, 81):
g = grad(x)
v = 0.9*v + g
x = x - lr*v
if t in (10, 40, 80):
print(t, round(0.5*(100*x[0]**2 + x[1]**2), 4))Loss is NON-monotone: 15.55 at step 10 (overshoot), 0.344 at 40, 0.0008 at 80
Why: The heavy ball builds speed on the steep axis and coasts past the minimum around step 10, then inertia turns it and it settles well below GD's 0.027 by step 80.
| step | loss |
|---|---|
| 0 (start) | 50.5 |
| 10 | 15.55 |
| 40 | 0.344 |
| 80 | 0.0008 |
Pattern
Step through it
Step through Milestone 1 — momentum one row at a time. What is driving the change, and what would the row after the last one be?
Pattern
Predict first
The table runs: 0 (start) | 50.5 · 20 | 0.2512 · 40 | ≈1.5e-7
In Milestone 2 — RMSprop, given the rows so far: what is the next one — the row where step is 80?
Correct: 80 | ≈0
| step | loss |
|---|---|
| 0 (start) | 50.5 |
| 20 | 0.2512 |
| 40 | ≈1.5e-7 |
| 80 | ≈0 |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. No momentum means no overshoot — RMSprop rescales both axes to move in lockstep and drives this pure quadratic to machine zero.
Worked example
Your turn: implement RMSprop at η = 0.05. Predict which axis its √s division speeds up most before you run.
Hint: s = np.zeros(2); each step s = 0.9*s + 0.1*g*g; x = x - lr*g/(np.sqrt(s) + 1e-8). On step 1 both axes should move by the identical amount.
import numpy as np
def grad(x): return np.array([100*x[0], x[1]])
x, s, lr = np.array([1., 1.]), np.zeros(2), 0.05
for t in range(1, 81):
g = grad(x)
s = 0.9*s + 0.1*g*g
x = x - lr*g/(np.sqrt(s) + 1e-8)
if t in (20, 40, 80):
print(t, round(0.5*(100*x[0]**2 + x[1]**2), 8))Smooth, monotone descent: 0.251 at 20, ~1.5e-7 at 40, ~0 at 80
Why: No momentum means no overshoot — RMSprop rescales both axes to move in lockstep and drives this pure quadratic to machine zero. It speeds up the flat x₁ axis the most (its tiny gradient gets divided by a tiny √s).
| step | loss |
|---|---|
| 0 (start) | 50.5 |
| 20 | 0.2512 |
| 40 | ≈1.5e-7 |
| 80 | ≈0 |
Scale up
Step through it
Step through Milestone 2 — RMSprop and watch the numbers move. Now imagine the input ten times bigger: which column is the one that stops this being practical?
Missing information
Discussion prompt
Your turn: implement Adam with bias correction at η = 0.1. Predict what breaks if you delete the mh, vh lines and step with raw m, v instead.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Bias correction makes the first step a clean unit-size move (m̂/√v̂ = sign(g)), so Adam is already at 0.294 by step 10 — ahead of GD. Drop the mh/vh lines and the un-warmed v makes early steps 3.16× too big, converging worse.
Worked example
Your turn: implement Adam with bias correction at η = 0.1. Predict what breaks if you delete the mh, vh lines and step with raw m, v instead.
Hint: track m, v = np.zeros(2), np.zeros(2); each step compute mh = m/(1-0.9**t), vh = v/(1-0.999**t); then x -= lr*mh/(np.sqrt(vh)+1e-8). Loop t from 1.
import numpy as np
def grad(x): return np.array([100*x[0], x[1]])
x, m, v, lr = np.array([1., 1.]), np.zeros(2), np.zeros(2), 0.1
for t in range(1, 81):
g = grad(x)
m = 0.9*m + 0.1*g
v = 0.999*v + 0.001*g*g
mh, vh = m/(1-0.9**t), v/(1-0.999**t)
x = x - lr*mh/(np.sqrt(vh) + 1e-8)
if t in (10, 20, 80):
print(t, round(0.5*(100*x[0]**2 + x[1]**2), 5))Loss 0.294 at 10, a mild overshoot 3.71 at 20, then 0.0015 at 80
Why: Bias correction makes the first step a clean unit-size move (m̂/√v̂ = sign(g)), so Adam is already at 0.294 by step 10 — ahead of GD. Drop the mh/vh lines and the un-warmed v makes early steps 3.16× too big, converging worse.
| step | loss |
|---|---|
| 0 (start) | 50.5 |
| 10 | 0.294 |
| 20 | 3.713 |
| 80 | 0.0015 |
Pattern
Step through it
Step through Milestone 3 — Adam one row at a time. What is driving the change, and what would the row after the last one be?
Estimation
Predict first
Your turn: wrap each rule in a function returning the final loss, and print all four side by side. Predict the ranking at step 80 before you run.
Commit before you compute: what does Milestone 4 — race your three vs GD come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: GD 0.02734 (worst); momentum 0.00075; Adam 0.00147; RMSprop ≈0 (best)
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Every adaptive/momentum method beats GD's floor by 1–30×.
Worked example
Your turn: wrap each rule in a function returning the final loss, and print all four side by side. Predict the ranking at step 80 before you run.
Hint: each function starts x=np.array([1.,1.]) and its own zero state, loops 80 steps, returns loss(x). Reuse grad and loss helpers.
import numpy as np
def grad(x): return np.array([100*x[0], x[1]])
def loss(x): return 0.5*(100*x[0]**2 + x[1]**2)
def run(kind, lr, steps=80):
x, v, s, m, w = np.array([1.,1.]), np.zeros(2), np.zeros(2), np.zeros(2), np.zeros(2)
for t in range(1, steps+1):
g = grad(x)
if kind == 'gd': x = x - lr*g
elif kind == 'mom': v = 0.9*v + g; x = x - lr*v
elif kind == 'rms': s = 0.9*s + 0.1*g*g; x = x - lr*g/(np.sqrt(s)+1e-8)
elif kind == 'adam':
m = 0.9*m + 0.1*g; w = 0.999*w + 0.001*g*g
mh, vh = m/(1-0.9**t), w/(1-0.999**t)
x = x - lr*mh/(np.sqrt(vh)+1e-8)
return round(loss(x), 5)
for k, lr in [('gd',0.018),('mom',0.003),('rms',0.05),('adam',0.1)]:
print(k, run(k, lr))GD 0.02734 (worst); momentum 0.00075; Adam 0.00147; RMSprop ≈0 (best)
Why: Every adaptive/momentum method beats GD's floor by 1–30×. RMSprop wins on this pure quadratic; GD is stuck because its η is hostage to the steep axis.
| optimizer | final loss @80 |
|---|---|
| gd (0.018) | 0.02734 |
| mom (0.003) | 0.00075 |
| adam (0.1) | 0.00147 |
| rms (0.05) | ≈0 |
Trade off
Comparison matrix
From Milestone 4 — race your three vs GD: every row here is a choice with a cost. Fill the final loss @80 column, then say which row you would actually pick and what you give up for it.
| optimizer | final loss @80 |
|---|---|
| gd (0.018) | 0.02734 |
| mom (0.003) | 0.00075 |
| adam (0.1) | 0.00147 |
| rms (0.05) | ≈0 |
Concept
import numpy as np
def grad(x): return np.array([100*x[0], x[1]])
def loss(x): return 0.5*(100*x[0]**2 + x[1]**2)
def adam(steps=80, lr=0.1, b1=0.9, b2=0.999, eps=1e-8):
x, m, v = np.array([1., 1.]), np.zeros(2), np.zeros(2)
for t in range(1, steps+1):
g = grad(x)
m = b1*m + (1-b1)*g # momentum piece
v = b2*v + (1-b2)*g*g # RMSprop piece
mh, vh = m/(1-b1**t), v/(1-b2**t) # bias correction
x = x - lr*mh/(np.sqrt(vh)+eps)
return loss(x)
print(round(adam(), 5)) # 0.00147| output | value |
|---|---|
| start loss | 50.5 |
| adam() final loss | 0.00147 |
| plain GD @80 (for contrast) | 0.02734 |
If your Adam drives the loss from 50.5 to about 0.0015 while plain GD is still stuck near 0.027 — you've built the optimizer that trains modern networks.
Comparison
Comparison matrix
From The full program: refill the value column from what you know. The rest of the table is as it appeared.
| output | value |
|---|---|
| start loss | 50.5 |
| adam() final loss | 0.00147 |
| plain GD @80 (for contrast) | 0.02734 |
Concept
Slides closed, out loud: explain (1) why one global η crawls on a κ = 100 bowl, (2) what RMSprop's √s division does to the steep axis, and (3) why bias correction matters most at step 1.
Stretch (homework): add a warmup + cosine schedule to your Adam and re-race; then swap in AdamW's decoupled decay. These optimizers reappear in the full training loop (Week 14) and transformer training (Week 29).
Counterexample
Discussion prompt
Slides closed, out loud: explain (1) why one global η crawls on a κ = 100 bowl, (2) what RMSprop's √s division does to the steep axis, and (3) why bias correction matters most at step 1.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — The problem: ill-conditioning · Momentum · RMSprop · Adam · Race the four · AdamW & warmup. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
η is capped by the steep axis (η < 2/L), so the flat axis crawls — the whole reason for modern optimizersg/√s rescaling)3.16× too big at t=1 while the corrected step is exactly 1−ηλθ decay) from Adam+L2, and explain why warmup stabilizes the noisy early v̂0.027 floor on our bowl| optimizer | the one thing to remember |
|---|---|
| momentum | v = βv + g — velocity smooths the zig-zag |
| RMSprop | divide by √(avg g²) — per-axis step size |
| Adam | both moments + /(1−βᵗ) bias correction |
| AdamW | decay outside the √v̂ divisor: −ηλθ |
| warmup | ramp η up while v̂ is still noisy, then decay |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.