USAAIO Lesson 62. It covers linear warmup, cosine annealing, SGDR with cosine restarts, the one-cycle policy for both learning rate and momentum, and the inverse-square-root Transformer schedule. Each is derived from first principles and verified in PyTorch with LambdaLR and the built-in schedulers, and there is a project implementing LambdaLR from scratch. The lesson runs to 30 slides.
Subject: Machine Learning · 59 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 62
Why a fixed LR leaves performance on the table — warmup, cosine annealing, SGDR, one-cycle, and the Transformer schedule, all from scratch.
Objectives
LambdaLRWarm-up
Discussion prompt
Before we open Lesson 62: Learning Rate Schedules: without looking back, what was the main idea of Multiclass Classification — Softmax, OvR, OvO, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
softmax regression (P(y=k|x), cross-entropy gradient = P − Y_one_hot), One-vs-Rest and One-vs-One multiclass strategies, multinomial vs multiple binary logistic, and weighted cross-entropy for class imbalance. Implement multinomial logistic from scratch and benchmark OvR/OvO/multinomial on an imbalanced 10-class dataset.
Section
Section 1
Concept
A large LR covers ground fast early, but overshoots sharp valleys late. A small LR lands precisely but wastes most of the budget crawling.
| LR choice | early training | late training |
|---|---|---|
| too large | fast descent | oscillates, diverges near minimum |
| too small | stable | painfully slow, under-trains |
| scheduled | large early | small late — best of both |
A schedule decouples the two regimes: explore aggressively, then exploit precisely. Every competitive USAAIO submission uses one.
Comparison
Comparison matrix
From Fixed LR leaves a trade-off unresolved: refill the early training column from what you know. The rest of the table is as it appeared.
| LR choice | early training | late training |
|---|---|---|
| too large | fast descent | oscillates, diverges near minimum |
| too small | stable | painfully slow, under-trains |
| scheduled | large early | small late — best of both |
Section
Section 2
Concept
Adam tracks a per-parameter variance estimate v. At step 0 that estimate is 0 — the effective step size is garbage. Warmup lets v stabilize before the full LR kicks in.
\[ \text{lr}(t) = \text{lr}_{\text{target}} \cdot \frac{\min(t,\, N_{\text{warm}})}{N_{\text{warm}}} \]
| step t | lr (target = 1e-3, N=100) |
|---|---|
| 0 | 0.0000e+00 |
| 25 | 2.5000e-04 |
| 50 | 5.0000e-04 |
| 75 | 7.5000e-04 |
| 100 | 1.0000e-03 |
| > 100 | 1.0000e-03 (capped) |
Trade off
Comparison matrix
From Warmup: ramp from 0 to target_lr: every row here is a choice with a cost. Fill the lr (target = 1e-3, N=100) column, then say which row you would actually pick and what you give up for it.
| step t | lr (target = 1e-3, N=100) |
|---|---|
| 0 | 0.0000e+00 |
| 25 | 2.5000e-04 |
| 50 | 5.0000e-04 |
| 75 | 7.5000e-04 |
| 100 | 1.0000e-03 |
| > 100 | 1.0000e-03 (capped) |
Estimation
Predict first
Implement warmup with LambdaLR — the lambda returns a multiplier on base_lr.
Commit before you compute: what does Warmup via LambdaLR come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: lr at step 100 = 1.0000e-03 — LR reached target
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The lambda returns 1.0 at step=N_warm, so the multiplier is exactly 1.0 and the full base_lr is active.
Worked example
Implement warmup with LambdaLR — the lambda returns a multiplier on base_lr.
import torch, torch.nn as nn, torch.optim as optim
N_warm = 100
def warmup_lambda(step):
return min(step / N_warm, 1.0)
model = nn.Linear(4, 1)
opt = optim.Adam(model.parameters(), lr=1e-3)
sched = optim.lr_scheduler.LambdaLR(opt, lr_lambda=warmup_lambda)
for step in range(1, 101):
opt.zero_grad(); opt.step(); sched.step()
print(f'lr at step 100: {opt.param_groups[0]["lr"]:.4e}')lr at step 100 = 1.0000e-03 — LR reached target
Why: The lambda returns 1.0 at step=N_warm, so the multiplier is exactly 1.0 and the full base_lr is active.
| step | multiplier | effective lr |
|---|---|---|
| 0 | 0.00 | 0.0000e+00 |
| 25 | 0.25 | 2.5000e-04 |
| 50 | 0.50 | 5.0000e-04 |
| 100 | 1.00 | 1.0000e-03 |
Pattern
Step through it
Step through Warmup via LambdaLR one row at a time. What is driving the change, and what would the row after the last one be?
Section
Section 3
Concept
Decay LR along a cosine curve: starts at LR_max, ends at LR_min. Smooth means the gradient never sees a sudden LR cliff.
\[ \text{lr}(t) = \text{lr}_{\min} + \tfrac{1}{2}(\text{lr}_{\max} - \text{lr}_{\min})\Bigl(1 + \cos\!\Bigl(\frac{\pi\, t}{T}\Bigr)\Bigr) \]
| t (of T=100) | lr (max=1e-3, min=0) |
|---|---|
| 0 | 1.0000e-03 |
| 25 | 8.5355e-04 |
| 50 | 5.0000e-04 |
| 75 | 1.4645e-04 |
| 100 | 0.0000e+00 |
Pattern
Step through it
Step through Cosine annealing — smooth decay one row at a time. What is driving the change, and what would the row after the last one be?
Missing information
Discussion prompt
PyTorch's built-in CosineAnnealingLR matches the formula exactly with T_max=T, eta_min=LR_min.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Matches the formula exactly — cosine half-period from LR_max to LR_min.
Worked example
PyTorch's built-in CosineAnnealingLR matches the formula exactly with T_max=T, eta_min=LR_min.
import torch, torch.nn as nn, torch.optim as optim, math
model = nn.Linear(4, 1)
opt = optim.SGD(model.parameters(), lr=1e-3)
sched = optim.lr_scheduler.CosineAnnealingLR(opt, T_max=100, eta_min=0)
for step in range(101):
opt.zero_grad(); opt.step(); sched.step()
if step in [0, 25, 50, 75, 100]:
print(f't={step:3d} lr={opt.param_groups[0]["lr"]:.4e}')t=0: 1.00e-3, t=25: 8.54e-4, t=50: 5.00e-4, t=75: 1.46e-4, t=100: 0.00e+0
Why: Matches the formula exactly — cosine half-period from LR_max to LR_min.
| t | lr |
|---|---|
| 0 | 1.0000e-03 |
| 25 | 8.5355e-04 |
| 50 | 5.0000e-04 |
| 75 | 1.4645e-04 |
| 100 | 0.0000e+00 |
Scale up
Step through it
Step through CosineAnnealingLR in PyTorch and watch the numbers move. Now imagine the input ten times bigger: which column is the one that stops this being practical?
Concept
Add a restart every T_0 steps: after reaching LR_min, snap back to LR_max. Each restart helps the optimizer escape a sharp local minimum (Lesson 50's loss-surface intuition).
| step (T_0=50) | lr after restart |
|---|---|
| 0 | 1.0000e-03 (start) |
| 25 | 5.0000e-04 (midpoint) |
| 50 | 1.0000e-03 (restart 1) |
| 75 | 5.0000e-04 (midpoint) |
| 100 | 1.0000e-03 (restart 2) |
T_mult > 1 doubles the period each restart — rare at first, then rare-but-long cooldowns. Verified with CosineAnnealingWarmRestarts(T_0=50, T_mult=1).
Counterexample
Discussion prompt
Add a restart every T_0 steps: after reaching LR_min, snap back to LR_max. Each restart helps the optimizer escape a sharp local minimum (Lesson 50's loss-surface intuition).
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
T_mult > 1 doubles the period each restart — rare at first, then rare-but-long cooldowns. Verified with CosineAnnealingWarmRestarts(T_0=50, T_mult=1).
Anomaly
Predict first
A student writes this, and it looks reasonable:
The scheduler is called at the top of the loop, before optimizer.step().
It is wrong. Say what breaks — and say it before you turn the page.
Correct: PyTorch raises a warning: 'Detected call of lr_scheduler.step() before optimizer.step().' The scheduler skips the first LR value, so step 0 uses step 1's LR and the whole trace shifts by one.
Call scheduler.step() AFTER optimizer.step(), at the end of each iteration.
Why: PyTorch raises a warning: 'Detected call of lr_scheduler.step() before optimizer.step().' The scheduler skips the first LR value, so step 0 uses step 1's LR and the whole trace shifts by one.
Trap
The scheduler is called at the top of the loop, before optimizer.step().
sched.step() → opt.zero_grad() → loss.backward() → opt.step()
Why: PyTorch raises a warning: 'Detected call of lr_scheduler.step() before optimizer.step().' The scheduler skips the first LR value, so step 0 uses step 1's LR and the whole trace shifts by one.
Call scheduler.step() AFTER optimizer.step(), at the end of each iteration.
opt.zero_grad() → forward → loss → loss.backward() → opt.step() → sched.step()
Why: The optimizer applies the current LR first, then the scheduler advances to the next LR. Order matters — PyTorch enforces it with a warning in 1.1+ and will eventually error.
Break the constraint
Discussion prompt
The rule this trap just fixed:
The optimizer applies the current LR first, then the scheduler advances to the next LR. Order matters — PyTorch enforces it with a warning in 1.1+ and will eventually error.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
PyTorch raises a warning: 'Detected call of lr_scheduler.step() before optimizer.step().' The scheduler skips the first LR value, so step 0 uses step 1's LR and the whole trace shifts by one.
Section
Section 4
Concept
Leslie Smith (2019): ramp LR from low to max_lr over the first pct_start fraction of training, then decay to near 0. Momentum moves inversely — high when LR is low, low when LR is high.
| phase | LR | momentum |
|---|---|---|
| warmup (0–30%) | 0.004 → 0.100 | 0.950 → 0.850 |
| decay (30–100%) | 0.100 → 0.000 | 0.850 → 0.950 |
The anti-correlation is deliberate: high momentum averages many gradients (like a large effective batch), compensating for the noisy high-LR phase. Low momentum at peak LR prevents overshooting.
Analogy
Discussion prompt
Explain One-cycle: LR up then down, momentum inverse by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Leslie Smith (2019): ramp LR from low to max_lr over the first pct_start fraction of training, then decay to near 0. Momentum moves inversely — high when LR is low, low when LR is high.
Estimation
Predict first
Verify the LR/momentum anti-correlation with PyTorch OneCycleLR on a 100-step schedule, max_lr=0.1, pct_start=0.3.
Commit before you compute: what does OneCycleLR — LR and momentum trace come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: step 0: lr=4.28e-3, mom=0.9497 → step 30: lr=9.98e-2, mom=0.8502 → step 99: lr=5.07e-5, mom=0.9499
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. LR peaks at step 30 (= 100 × 0.3) while momentum hits its minimum — confirming the inverse relationship Smith described.
Worked example
Verify the LR/momentum anti-correlation with PyTorch OneCycleLR on a 100-step schedule, max_lr=0.1, pct_start=0.3.
import torch, torch.nn as nn, torch.optim as optim
model = nn.Linear(4, 1)
opt = optim.SGD(model.parameters(), lr=0.01,
momentum=0.9)
oc = optim.lr_scheduler.OneCycleLR(
opt, max_lr=0.1, total_steps=100,
pct_start=0.3, cycle_momentum=True,
base_momentum=0.85, max_momentum=0.95)
for step in range(100):
opt.step(); oc.step()
if step in [0, 15, 30, 50, 70, 99]:
lr = opt.param_groups[0]['lr']
mom = opt.param_groups[0]['momentum']
print(f'step {step:2d} lr={lr:.3e} mom={mom:.4f}')step 0: lr=4.28e-3, mom=0.9497 → step 30: lr=9.98e-2, mom=0.8502 → step 99: lr=5.07e-5, mom=0.9499
Why: LR peaks at step 30 (= 100 × 0.3) while momentum hits its minimum — confirming the inverse relationship Smith described.
| step | lr | momentum |
|---|---|---|
| 0 | 4.28e-03 | 0.9497 |
| 15 | 5.98e-02 | 0.8919 |
| 30 | 9.98e-02 | 0.8502 |
| 50 | 7.75e-02 | 0.8725 |
| 70 | 3.45e-02 | 0.9155 |
| 99 | 5.07e-05 | 0.9499 |
Comparison
Comparison matrix
From OneCycleLR — LR and momentum trace: refill the momentum column from what you know. The rest of the table is as it appeared.
| step | lr | momentum |
|---|---|---|
| 0 | 4.28e-03 | 0.9497 |
| 15 | 5.98e-02 | 0.8919 |
| 30 | 9.98e-02 | 0.8502 |
| 50 | 7.75e-02 | 0.8725 |
| 70 | 3.45e-02 | 0.9155 |
| 99 | 5.07e-05 | 0.9499 |
Section
Section 5
Concept
Vaswani et al. (2017) introduced a schedule that combines warmup with a slow 1/√step decay — no manual T or cycle needed.
\[ \text{lr}(t) = d_{\text{model}}^{-0.5} \cdot \min\!\Bigl(t^{-0.5},\; t \cdot N_{\text{warm}}^{-1.5}\Bigr) \]
| step t | lr (d_model=512, N_warm=4000) |
|---|---|
| 1 | 1.70e-07 |
| 100 | 1.75e-05 |
| 1000 | 1.75e-04 |
| 4000 | 6.99e-04 (peak) |
| 8000 | 4.94e-04 |
| 16000 | 3.49e-04 |
Analogy
Discussion prompt
Explain Inverse square root — the Transformer schedule by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Vaswani et al. (2017) introduced a schedule that combines warmup with a slow 1/√step decay — no manual T or cycle needed.
Estimation
Predict first
Implement the exact Vaswani formula in LambdaLR — one function, both warmup and decay.
Commit before you compute: what does Transformer schedule from scratch with LambdaLR come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Peak at step=4000: lr=6.99e-04 = d_model^-0.5 × N_warm^-0.5 = 512^-0.5 × 4000^-0.5
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The two terms in the min() cross exactly at t = N_warm.
Worked example
Implement the exact Vaswani formula in LambdaLR — one function, both warmup and decay.
import math, torch, torch.nn as nn, torch.optim as optim
d_model, N_warm = 512, 4000
def transformer_lambda(step):
step = max(step, 1) # avoid div/0 at step 0
return d_model**(-0.5) * min(step**(-0.5),
step * N_warm**(-1.5))
model = nn.Linear(512, 512)
opt = optim.Adam(model.parameters(), lr=1.0)
sched = optim.lr_scheduler.LambdaLR(opt, lr_lambda=transformer_lambda)
for step in range(1, 16001):
opt.step(); sched.step()
if step in [1, 100, 1000, 4000, 8000, 16000]:
print(f'step {step:6d} lr={opt.param_groups[0]["lr"]:.4e}')Peak at step=4000: lr=6.99e-04 = d_model^-0.5 × N_warm^-0.5 = 512^-0.5 × 4000^-0.5
Why: The two terms in the min() cross exactly at t = N_warm. Before the crossover the line grows; after it the sqrt decay dominates. No manual T_max needed — the formula self-terminates.
| step | lr |
|---|---|
| 1 | 1.70e-07 |
| 1000 | 1.75e-04 |
| 4000 | 6.99e-04 |
| 8000 | 4.94e-04 |
| 16000 | 3.49e-04 |
Pattern
Step through it
Step through Transformer schedule from scratch with LambdaLR one row at a time. What is driving the change, and what would the row after the last one be?
Anomaly
Predict first
A student writes this, and it looks reasonable:
Set base_lr=1e-3 in Adam, then define the lambda to return the full formula value.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: LambdaLR multiplies base_lr by the lambda.
Set base_lr=1.0 in Adam so the lambda IS the effective LR.
Why: LambdaLR multiplies base_lr by the lambda. If base_lr is not 1.0, the effective LR is base_lr × lambda — it will be 1000× too small (or too large) compared to the intended schedule.
Trap
Set base_lr=1e-3 in Adam, then define the lambda to return the full formula value.
lambda(step) = d_model**-0.5 * min(step**-0.5, step * N_warm**-1.5)
Why: LambdaLR multiplies base_lr by the lambda. If base_lr is not 1.0, the effective LR is base_lr × lambda — it will be 1000× too small (or too large) compared to the intended schedule.
Set base_lr=1.0 in Adam so the lambda IS the effective LR.
opt = Adam(..., lr=1.0); LambdaLR(opt, lr_lambda=transformer_lambda)
Why: When base_lr=1.0, the multiplier equals the final LR directly. The Vaswani formula is an absolute LR, not a scale factor — base_lr must be 1.0 to use it as-is.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
v. At step 0 that estimate is 0 — the effective step size is garbage. Warmup lets v stabilize before the full LR kicks in.; Decay LR along a cosine curve: starts at LR_max, ends at LR_min. Smooth means the gradient never sees a sudden LR cliff.base_lr=1e-3 in Adam, then define the lambda to return the full formula value.Ranking
Put in order
These are the steps of LR schedule selection recipe, scrambled. Put them back in order before the next slide shows you.
N_warm ≈ 5–10 % of total steps; use LambdaLRCosineAnnealingWarmRestarts when plateau is a risk (non-convex losses)max_lr via LR-finder, pct_start=0.3base_lr=1.0, tune d_model & N_warmLambdaLR — the lambda is just a function of stepWhy: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
N_warm ≈ 5–10 % of total steps; use LambdaLRCosineAnnealingWarmRestarts when plateau is a risk (non-convex losses)max_lr via LR-finder, pct_start=0.3base_lr=1.0, tune d_model & N_warmLambdaLR — the lambda is just a function of stepEdge cases
Discussion prompt
LR schedule selection recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
N_warm ≈ 5–10 % of total steps; use LambdaLRCosineAnnealingWarmRestarts when plateau is a risk (non-convex losses)max_lr via LR-finder, pct_start=0.3base_lr=1.0, tune d_model & N_warmLambdaLR — the lambda is just a function of stepElimination
Eliminate the wrong options
Why does linear warmup stabilize Adam in early training?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Adam's per-parameter second moment estimate v(0) = 0. The effective step size is gradient / sqrt(v + ε) — when v ≈ 0, ε dominates and the step is large and unstable. Warmup keeps LR tiny until v is meaningful.
Check
Think before clicking.
Check your understanding
Why does linear warmup stabilize Adam in early training?
Answer: A
Why: Adam's per-parameter second moment estimate v(0) = 0. The effective step size is gradient / sqrt(v + ε) — when v ≈ 0, ε dominates and the step is large and unstable. Warmup keeps LR tiny until v is meaningful.
Prediction
Predict first
What does SGDR add over plain cosine annealing?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Periodic restarts back to LR_max, helping escape sharp local minima
Why: SGDR (Loshchilov & Hutter 2017) appends a snap-back to LR_max at the end of each T_0-step cycle. This lets the optimizer escape narrow minima it would otherwise converge into — the restart provides the perturbation.
Check
Distinguish the two cosine variants.
Check your understanding
What does SGDR add over plain cosine annealing?
Answer: A
Why: SGDR (Loshchilov & Hutter 2017) appends a snap-back to LR_max at the end of each T_0-step cycle. This lets the optimizer escape narrow minima it would otherwise converge into — the restart provides the perturbation.
Elimination
Eliminate the wrong options
For d_model=512 and N_warm=4000, at which step does the Transformer LR reach its peak, and what is that peak?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: The two terms cross when t^-0.5 = t × N_warm^-1.5, which gives t = N_warm = 4000. Peak = d_model^-0.5 × N_warm^-0.5 = 512^-0.5 × 4000^-0.5 ≈ 6.99e-04. Verified in Python.
Check
Read the formula, not memory.
Check your understanding
For d_model=512 and N_warm=4000, at which step does the Transformer LR reach its peak, and what is that peak?
Answer: A
Why: The two terms cross when t^-0.5 = t × N_warm^-1.5, which gives t = N_warm = 4000. Peak = d_model^-0.5 × N_warm^-0.5 = 512^-0.5 × 4000^-0.5 ≈ 6.99e-04. Verified in Python.
Section
Project
Concept
Implement all five schedules from scratch using LambdaLR. Each lambda is a pure Python function — no scheduler classes except LambdaLR itself.
| # | schedule | lambda signature |
|---|---|---|
| 1 | Linear warmup | warmup_lambda(step) |
| 2 | Cosine annealing | cosine_lambda(step) |
| 3 | SGDR | sgdr_lambda(step) |
| 4 | One-cycle LR | onecycle_lambda(step) |
| 5 | Transformer (inv-sqrt) | transformer_lambda(step) |
Build rule: set base_lr=1.0 in your optimizer so the lambda directly controls the effective LR — no accidental scaling.
Counterexample
Discussion prompt
Implement all five schedules from scratch using LambdaLR. Each lambda is a pure Python function — no scheduler classes except LambdaLR itself.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Build rule: set base_lr=1.0 in your optimizer so the lambda directly controls the effective LR — no accidental scaling.
Worked example
Your turn: write warmup_lambda and cosine_lambda. Predict what cosine_lambda(50) returns for LR_max=1e-3, LR_min=0, T=100.
Hint: warmup multiplier = min(step/N_warm, 1.0); cosine = 0.5*(1 + cos(pi*step/T)).
import math
N_warm = 100; T = 100
def warmup_lambda(step):
return min(step / N_warm, 1.0)
def cosine_lambda(step):
return 0.5 * (1.0 + math.cos(math.pi * step / T))
for s in [0, 25, 50, 75, 100]:
print(f's={s:3d} warmup={warmup_lambda(s):.3f} cosine={cosine_lambda(s):.4f}')| step | warmup mult | cosine mult |
|---|---|---|
| 0 | 0.000 | 1.0000 |
| 25 | 0.250 | 0.8536 |
| 50 | 0.500 | 0.5000 |
| 75 | 0.750 | 0.1464 |
| 100 | 1.000 | 0.0000 |
Pattern
Step through it
Step through Milestone 1 — warmup + cosine lambdas one row at a time. What is driving the change, and what would the row after the last one be?
Worked example
Your turn: implement sgdr_lambda with T_0=50 restarts using only math.cos and %. Predict the LR at steps 0, 25, 50, 75.
Hint: t_in_cycle = step % T_0; return cosine over [0, T_0]. At t_in_cycle=0 the LR snaps back to 1.0.
import math
T_0 = 50
def sgdr_lambda(step):
t = step % T_0
return 0.5 * (1.0 + math.cos(math.pi * t / T_0))
for s in [0, 25, 50, 75, 100]:
print(f'step {s:3d}: mult={sgdr_lambda(s):.4f}')| step | t_in_cycle | multiplier (= lr if base=1) |
|---|---|---|
| 0 | 0 | 1.0000 |
| 25 | 25 | 0.5000 |
| 50 | 0 | 1.0000 (restart) |
| 75 | 25 | 0.5000 |
| 100 | 0 | 1.0000 (restart) |
Trade off
Comparison matrix
From Milestone 2 — SGDR lambda: every row here is a choice with a cost. Fill the multiplier (= lr if base=1) column, then say which row you would actually pick and what you give up for it.
| step | t_in_cycle | multiplier (= lr if base=1) |
|---|---|---|
| 0 | 0 | 1.0000 |
| 25 | 25 | 0.5000 |
| 50 | 0 | 1.0000 (restart) |
| 75 | 25 | 0.5000 |
| 100 | 0 | 1.0000 (restart) |
Concept
import math, torch, torch.optim as optim
# shared hyperparams
N_warm, T, T_0 = 100, 200, 50
N_total, pct = 200, 0.3
d_model, N_w = 512, 4000
def warmup_fn(s): return min(s/N_warm, 1.0)
def cosine_fn(s): return 0.5*(1+math.cos(math.pi*s/T))
def sgdr_fn(s): return 0.5*(1+math.cos(math.pi*(s%T_0)/T_0))
def onecycle_fn(s):
k = int(N_total*pct)
if s < k: return s/k
return 0.5*(1+math.cos(math.pi*(s-k)/(N_total-k)))
def transformer_fn(s):
s = max(s,1)
return d_model**-0.5*min(s**-0.5, s*N_w**-1.5)
for name, fn in [('warmup', warmup_fn), ('cosine', cosine_fn),
('sgdr', sgdr_fn), ('onecycle', onecycle_fn),
('transf', transformer_fn)]:
print(f'{name:10s} step50={fn(50):.4f} step100={fn(100):.4f}')| schedule | step 50 | step 100 |
|---|---|---|
| warmup | 0.5000 | 1.0000 |
| cosine | 0.5000 | 0.0000 |
| sgdr | 1.0000 | 1.0000 |
| onecycle | 0.8333 | 0.4268 |
| transf | 0.0001 | 0.0002 |
All five lambdas implemented, no scheduler classes needed except LambdaLR — you own every schedule from first principles.
Comparison
Comparison matrix
From The full program: all five schedules: refill the step 100 column from what you know. The rest of the table is as it appeared.
| schedule | step 50 | step 100 |
|---|---|---|
| warmup | 0.5000 | 1.0000 |
| cosine | 0.5000 | 0.0000 |
| sgdr | 1.0000 | 1.0000 |
| onecycle | 0.8333 | 0.4268 |
| transf | 0.0001 | 0.0002 |
Concept
Out loud, slides closed: (1) derive the cosine formula from scratch and state what happens at t=0, T/2, T; (2) explain why the Transformer peak is at step=N_warm; (3) describe the LR/momentum anti-correlation in one-cycle and why it helps.
Stretch (homework from lesson plan): implement the LR finder (Leslie Smith) — sweep LR from 1e-7 to 10 over 100 steps, log loss vs LR, find the inflection point. That point is your max_lr for one-cycle. Compare constant LR vs cosine annealing vs one-cycle on a real training run and plot curves.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Why schedules? The fixed-LR problem · Schedule 1: Linear Warmup · Schedule 2: Cosine Annealing · Schedule 3: One-Cycle Policy · Schedule 4: Transformer / Inverse Sqrt · Your turn: implement all five schedules. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
vLambdaLRpct_startbase_lr=1.0, function of step| schedule | key formula | use when |
|---|---|---|
| warmup | lr × min(t/N, 1) | always (Adam) |
| cosine | 0.5·(1+cos(πt/T)) | default decay |
| SGDR | cosine + restart at T_0 | non-convex, plateau risk |
| one-cycle | LR↑ then↓; momentum inverse | single fast run |
| transformer | d^-0.5 · min(t^-0.5, t·N^-1.5) | self-attention models |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.