Lesson 62: Learning Rate Schedules

USAAIO Lesson 62. It covers linear warmup, cosine annealing, SGDR with cosine restarts, the one-cycle policy for both learning rate and momentum, and the inverse-square-root Transformer schedule. Each is derived from first principles and verified in PyTorch with LambdaLR and the built-in schedulers, and there is a project implementing LambdaLR from scratch. The lesson runs to 30 slides.

Subject: Machine Learning · 59 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Learning Rate Schedules

Title

USAAIO · Lesson 62

Why a fixed LR leaves performance on the table — warmup, cosine annealing, SGDR, one-cycle, and the Transformer schedule, all from scratch.

2. By the end of this lesson you can

Objectives

  1. Explain why linear warmup stabilizes Adam in early training
  2. Derive and implement cosine annealing (with and without restarts)
  3. Describe the one-cycle policy — LR up then down, momentum down then up
  4. Reproduce the Transformer inverse-square-root schedule from the formula
  5. Implement any schedule from scratch using PyTorch LambdaLR

3. What survived from Multiclass Classification — Softmax, OvR, OvO?

Warm-up

Discussion prompt

Before we open Lesson 62: Learning Rate Schedules: without looking back, what was the main idea of Multiclass Classification — Softmax, OvR, OvO, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

softmax regression (P(y=k|x), cross-entropy gradient = P − Y_one_hot), One-vs-Rest and One-vs-One multiclass strategies, multinomial vs multiple binary logistic, and weighted cross-entropy for class imbalance. Implement multinomial logistic from scratch and benchmark OvR/OvO/multinomial on an imbalanced 10-class dataset.

4. Why schedules? The fixed-LR problem

Section

Section 1

5. Fixed LR leaves a trade-off unresolved

Concept

A large LR covers ground fast early, but overshoots sharp valleys late. A small LR lands precisely but wastes most of the budget crawling.

LR choiceearly traininglate training
too largefast descentoscillates, diverges near minimum
too smallstablepainfully slow, under-trains
scheduledlarge earlysmall late — best of both

A schedule decouples the two regimes: explore aggressively, then exploit precisely. Every competitive USAAIO submission uses one.

6. Fill in: early training for Fixed LR leaves a trade-off unresolved

Comparison

Comparison matrix

From Fixed LR leaves a trade-off unresolved: refill the early training column from what you know. The rest of the table is as it appeared.

LR choiceearly traininglate training
too largefast descentoscillates, diverges near minimum
too smallstablepainfully slow, under-trains
scheduledlarge earlysmall late — best of both

7. Schedule 1: Linear Warmup

Section

Section 2

8. Warmup: ramp from 0 to target_lr

Concept

Adam tracks a per-parameter variance estimate v. At step 0 that estimate is 0 — the effective step size is garbage. Warmup lets v stabilize before the full LR kicks in.

\[ \text{lr}(t) = \text{lr}_{\text{target}} \cdot \frac{\min(t,\, N_{\text{warm}})}{N_{\text{warm}}} \]

step tlr (target = 1e-3, N=100)
00.0000e+00
252.5000e-04
505.0000e-04
757.5000e-04
1001.0000e-03
> 1001.0000e-03 (capped)

9. What each one costs: Warmup: ramp from 0 to target_lr

Trade off

Comparison matrix

From Warmup: ramp from 0 to target_lr: every row here is a choice with a cost. Fill the lr (target = 1e-3, N=100) column, then say which row you would actually pick and what you give up for it.

step tlr (target = 1e-3, N=100)
00.0000e+00
252.5000e-04
505.0000e-04
757.5000e-04
1001.0000e-03
> 1001.0000e-03 (capped)

10. Guess the shape of the answer: Warmup via LambdaLR

Estimation

Predict first

Implement warmup with LambdaLR — the lambda returns a multiplier on base_lr.

Commit before you compute: what does Warmup via LambdaLR come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: lr at step 100 = 1.0000e-03 — LR reached target

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The lambda returns 1.0 at step=N_warm, so the multiplier is exactly 1.0 and the full base_lr is active.

11. Warmup via LambdaLR

Worked example

Implement warmup with LambdaLR — the lambda returns a multiplier on base_lr.

import torch, torch.nn as nn, torch.optim as optim

N_warm = 100
def warmup_lambda(step):
    return min(step / N_warm, 1.0)

model = nn.Linear(4, 1)
opt   = optim.Adam(model.parameters(), lr=1e-3)
sched = optim.lr_scheduler.LambdaLR(opt, lr_lambda=warmup_lambda)

for step in range(1, 101):
    opt.zero_grad(); opt.step(); sched.step()

print(f'lr at step 100: {opt.param_groups[0]["lr"]:.4e}')

lr at step 100 = 1.0000e-03 — LR reached target

Why: The lambda returns 1.0 at step=N_warm, so the multiplier is exactly 1.0 and the full base_lr is active.

stepmultipliereffective lr
00.000.0000e+00
250.252.5000e-04
500.505.0000e-04
1001.001.0000e-03

12. Watch it run: Warmup via LambdaLR

Pattern

Step through it

Step through Warmup via LambdaLR one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: step is 0
  2. Step 2: step is 25
  3. Step 3: step is 50
  4. Step 4: step is 100

13. Schedule 2: Cosine Annealing

Section

Section 3

14. Cosine annealing — smooth decay

Concept

Decay LR along a cosine curve: starts at LR_max, ends at LR_min. Smooth means the gradient never sees a sudden LR cliff.

\[ \text{lr}(t) = \text{lr}_{\min} + \tfrac{1}{2}(\text{lr}_{\max} - \text{lr}_{\min})\Bigl(1 + \cos\!\Bigl(\frac{\pi\, t}{T}\Bigr)\Bigr) \]

t (of T=100)lr (max=1e-3, min=0)
01.0000e-03
258.5355e-04
505.0000e-04
751.4645e-04
1000.0000e+00

15. Watch it run: Cosine annealing — smooth decay

Pattern

Step through it

Step through Cosine annealing — smooth decay one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: t (of T=100) is 0
  2. Step 2: t (of T=100) is 25
  3. Step 3: t (of T=100) is 50
  4. Step 4: t (of T=100) is 75
  5. Step 5: t (of T=100) is 100

16. What has to be given first: CosineAnnealingLR in PyTorch

Missing information

Discussion prompt

PyTorch's built-in CosineAnnealingLR matches the formula exactly with T_max=T, eta_min=LR_min.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

Matches the formula exactly — cosine half-period from LR_max to LR_min.

17. CosineAnnealingLR in PyTorch

Worked example

PyTorch's built-in CosineAnnealingLR matches the formula exactly with T_max=T, eta_min=LR_min.

import torch, torch.nn as nn, torch.optim as optim, math

model = nn.Linear(4, 1)
opt   = optim.SGD(model.parameters(), lr=1e-3)
sched = optim.lr_scheduler.CosineAnnealingLR(opt, T_max=100, eta_min=0)

for step in range(101):
    opt.zero_grad(); opt.step(); sched.step()
    if step in [0, 25, 50, 75, 100]:
        print(f't={step:3d}  lr={opt.param_groups[0]["lr"]:.4e}')

t=0: 1.00e-3, t=25: 8.54e-4, t=50: 5.00e-4, t=75: 1.46e-4, t=100: 0.00e+0

Why: Matches the formula exactly — cosine half-period from LR_max to LR_min.

tlr
01.0000e-03
258.5355e-04
505.0000e-04
751.4645e-04
1000.0000e+00

18. What happens as it grows: CosineAnnealingLR in PyTorch

Scale up

Step through it

Step through CosineAnnealingLR in PyTorch and watch the numbers move. Now imagine the input ten times bigger: which column is the one that stops this being practical?

  1. Step 1: t is 0
  2. Step 2: t is 25
  3. Step 3: t is 50
  4. Step 4: t is 75
  5. Step 5: t is 100

19. SGDR: cosine with periodic restarts

Concept

Add a restart every T_0 steps: after reaching LR_min, snap back to LR_max. Each restart helps the optimizer escape a sharp local minimum (Lesson 50's loss-surface intuition).

step (T_0=50)lr after restart
01.0000e-03 (start)
255.0000e-04 (midpoint)
501.0000e-03 (restart 1)
755.0000e-04 (midpoint)
1001.0000e-03 (restart 2)

T_mult > 1 doubles the period each restart — rare at first, then rare-but-long cooldowns. Verified with CosineAnnealingWarmRestarts(T_0=50, T_mult=1).

20. Break it if you can: SGDR: cosine with periodic restarts

Counterexample

Discussion prompt

Add a restart every T_0 steps: after reaching LR_min, snap back to LR_max. Each restart helps the optimizer escape a sharp local minimum (Lesson 50's loss-surface intuition).

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

T_mult > 1 doubles the period each restart — rare at first, then rare-but-long cooldowns. Verified with CosineAnnealingWarmRestarts(T_0=50, T_mult=1).

21. Something is wrong here: calling sched.step() before opt.step()

Anomaly

Predict first

A student writes this, and it looks reasonable:

The scheduler is called at the top of the loop, before optimizer.step().

It is wrong. Say what breaks — and say it before you turn the page.

Correct: PyTorch raises a warning: 'Detected call of lr_scheduler.step() before optimizer.step().' The scheduler skips the first LR value, so step 0 uses step 1's LR and the whole trace shifts by one.

Call scheduler.step() AFTER optimizer.step(), at the end of each iteration.

Why: PyTorch raises a warning: 'Detected call of lr_scheduler.step() before optimizer.step().' The scheduler skips the first LR value, so step 0 uses step 1's LR and the whole trace shifts by one.

22. Trap: calling sched.step() before opt.step()

Trap

The trap

The scheduler is called at the top of the loop, before optimizer.step().

sched.step() → opt.zero_grad() → loss.backward() → opt.step()

Why: PyTorch raises a warning: 'Detected call of lr_scheduler.step() before optimizer.step().' The scheduler skips the first LR value, so step 0 uses step 1's LR and the whole trace shifts by one.

The fix

Call scheduler.step() AFTER optimizer.step(), at the end of each iteration.

opt.zero_grad() → forward → loss → loss.backward() → opt.step() → sched.step()

Why: The optimizer applies the current LR first, then the scheduler advances to the next LR. Order matters — PyTorch enforces it with a warning in 1.1+ and will eventually error.

23. Break it on purpose: calling sched.step() before opt.step()

Break the constraint

Discussion prompt

The rule this trap just fixed:

The optimizer applies the current LR first, then the scheduler advances to the next LR. Order matters — PyTorch enforces it with a warning in 1.1+ and will eventually error.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

PyTorch raises a warning: 'Detected call of lr_scheduler.step() before optimizer.step().' The scheduler skips the first LR value, so step 0 uses step 1's LR and the whole trace shifts by one.

24. Schedule 3: One-Cycle Policy

Section

Section 4

25. One-cycle: LR up then down, momentum inverse

Concept

Leslie Smith (2019): ramp LR from low to max_lr over the first pct_start fraction of training, then decay to near 0. Momentum moves inversely — high when LR is low, low when LR is high.

phaseLRmomentum
warmup (0–30%)0.004 → 0.1000.950 → 0.850
decay (30–100%)0.100 → 0.0000.850 → 0.950

The anti-correlation is deliberate: high momentum averages many gradients (like a large effective batch), compensating for the noisy high-LR phase. Low momentum at peak LR prevents overshooting.

26. By analogy: One-cycle: LR up then down, momentum inverse

Analogy

Discussion prompt

Explain One-cycle: LR up then down, momentum inverse by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Leslie Smith (2019): ramp LR from low to max_lr over the first pct_start fraction of training, then decay to near 0. Momentum moves inversely — high when LR is low, low when LR is high.

27. Guess the shape of the answer: OneCycleLR — LR and momentum trace

Estimation

Predict first

Verify the LR/momentum anti-correlation with PyTorch OneCycleLR on a 100-step schedule, max_lr=0.1, pct_start=0.3.

Commit before you compute: what does OneCycleLR — LR and momentum trace come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: step 0: lr=4.28e-3, mom=0.9497 → step 30: lr=9.98e-2, mom=0.8502 → step 99: lr=5.07e-5, mom=0.9499

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. LR peaks at step 30 (= 100 × 0.3) while momentum hits its minimum — confirming the inverse relationship Smith described.

28. OneCycleLR — LR and momentum trace

Worked example

Verify the LR/momentum anti-correlation with PyTorch OneCycleLR on a 100-step schedule, max_lr=0.1, pct_start=0.3.

import torch, torch.nn as nn, torch.optim as optim

model = nn.Linear(4, 1)
opt   = optim.SGD(model.parameters(), lr=0.01,
                  momentum=0.9)
oc    = optim.lr_scheduler.OneCycleLR(
    opt, max_lr=0.1, total_steps=100,
    pct_start=0.3, cycle_momentum=True,
    base_momentum=0.85, max_momentum=0.95)

for step in range(100):
    opt.step(); oc.step()
    if step in [0, 15, 30, 50, 70, 99]:
        lr  = opt.param_groups[0]['lr']
        mom = opt.param_groups[0]['momentum']
        print(f'step {step:2d}  lr={lr:.3e}  mom={mom:.4f}')

step 0: lr=4.28e-3, mom=0.9497 → step 30: lr=9.98e-2, mom=0.8502 → step 99: lr=5.07e-5, mom=0.9499

Why: LR peaks at step 30 (= 100 × 0.3) while momentum hits its minimum — confirming the inverse relationship Smith described.

steplrmomentum
04.28e-030.9497
155.98e-020.8919
309.98e-020.8502
507.75e-020.8725
703.45e-020.9155
995.07e-050.9499

29. Fill in: momentum for OneCycleLR — LR and momentum trace

Comparison

Comparison matrix

From OneCycleLR — LR and momentum trace: refill the momentum column from what you know. The rest of the table is as it appeared.

steplrmomentum
04.28e-030.9497
155.98e-020.8919
309.98e-020.8502
507.75e-020.8725
703.45e-020.9155
995.07e-050.9499

30. Schedule 4: Transformer / Inverse Sqrt

Section

Section 5

31. Inverse square root — the Transformer schedule

Concept

Vaswani et al. (2017) introduced a schedule that combines warmup with a slow 1/√step decay — no manual T or cycle needed.

\[ \text{lr}(t) = d_{\text{model}}^{-0.5} \cdot \min\!\Bigl(t^{-0.5},\; t \cdot N_{\text{warm}}^{-1.5}\Bigr) \]

step tlr (d_model=512, N_warm=4000)
11.70e-07
1001.75e-05
10001.75e-04
40006.99e-04 (peak)
80004.94e-04
160003.49e-04

32. By analogy: Inverse square root — the Transformer schedule

Analogy

Discussion prompt

Explain Inverse square root — the Transformer schedule by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Vaswani et al. (2017) introduced a schedule that combines warmup with a slow 1/√step decay — no manual T or cycle needed.

33. Guess the shape of the answer: Transformer schedule from scratch with…

Estimation

Predict first

Implement the exact Vaswani formula in LambdaLR — one function, both warmup and decay.

Commit before you compute: what does Transformer schedule from scratch with LambdaLR come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Peak at step=4000: lr=6.99e-04 = d_model^-0.5 × N_warm^-0.5 = 512^-0.5 × 4000^-0.5

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The two terms in the min() cross exactly at t = N_warm.

34. Transformer schedule from scratch with LambdaLR

Worked example

Implement the exact Vaswani formula in LambdaLR — one function, both warmup and decay.

import math, torch, torch.nn as nn, torch.optim as optim

d_model, N_warm = 512, 4000

def transformer_lambda(step):
    step = max(step, 1)   # avoid div/0 at step 0
    return d_model**(-0.5) * min(step**(-0.5),
                                  step * N_warm**(-1.5))

model = nn.Linear(512, 512)
opt   = optim.Adam(model.parameters(), lr=1.0)
sched = optim.lr_scheduler.LambdaLR(opt, lr_lambda=transformer_lambda)

for step in range(1, 16001):
    opt.step(); sched.step()
    if step in [1, 100, 1000, 4000, 8000, 16000]:
        print(f'step {step:6d}  lr={opt.param_groups[0]["lr"]:.4e}')

Peak at step=4000: lr=6.99e-04 = d_model^-0.5 × N_warm^-0.5 = 512^-0.5 × 4000^-0.5

Why: The two terms in the min() cross exactly at t = N_warm. Before the crossover the line grows; after it the sqrt decay dominates. No manual T_max needed — the formula self-terminates.

steplr
11.70e-07
10001.75e-04
40006.99e-04
80004.94e-04
160003.49e-04

35. Watch it run: Transformer schedule from scratch with LambdaLR

Pattern

Step through it

Step through Transformer schedule from scratch with LambdaLR one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: step is 1
  2. Step 2: step is 1000
  3. Step 3: step is 4000
  4. Step 4: step is 8000
  5. Step 5: step is 16000

36. Something is wrong here: using LR_scale = 1 with LambdaLR for the Transformer

Anomaly

Predict first

A student writes this, and it looks reasonable:

Set base_lr=1e-3 in Adam, then define the lambda to return the full formula value.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: LambdaLR multiplies base_lr by the lambda.

Set base_lr=1.0 in Adam so the lambda IS the effective LR.

Why: LambdaLR multiplies base_lr by the lambda. If base_lr is not 1.0, the effective LR is base_lr × lambda — it will be 1000× too small (or too large) compared to the intended schedule.

37. Trap: using LR_scale = 1 with LambdaLR for the Transformer

Trap

The trap

Set base_lr=1e-3 in Adam, then define the lambda to return the full formula value.

lambda(step) = d_model**-0.5 * min(step**-0.5, step * N_warm**-1.5)

Why: LambdaLR multiplies base_lr by the lambda. If base_lr is not 1.0, the effective LR is base_lr × lambda — it will be 1000× too small (or too large) compared to the intended schedule.

The fix

Set base_lr=1.0 in Adam so the lambda IS the effective LR.

opt = Adam(..., lr=1.0); LambdaLR(opt, lr_lambda=transformer_lambda)

Why: When base_lr=1.0, the multiplier equals the final LR directly. The Vaswani formula is an absolute LR, not a scale factor — base_lr must be 1.0 to use it as-is.

38. Which of these survive contact with Lesson 62: Learning Rate Schedules?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
A large LR covers ground fast early, but overshoots sharp valleys late. A small LR lands precisely but wastes most of the budget crawling.; Adam tracks a per-parameter variance estimate v. At step 0 that estimate is 0 — the effective step size is garbage. Warmup lets v stabilize before the full LR kicks in.; Decay LR along a cosine curve: starts at LR_max, ends at LR_min. Smooth means the gradient never sees a sudden LR cliff.
Breaks
The scheduler is called at the top of the loop, before optimizer.step().; Set base_lr=1e-3 in Adam, then define the lambda to return the full formula value.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 62: Learning Rate Schedules puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

39. Rebuild the recipe: LR schedule selection recipe

Ranking

Put in order

These are the steps of LR schedule selection recipe, scrambled. Put them back in order before the next slide shows you.

  1. Always warmup Adam — N_warm ≈ 5–10 % of total steps; use LambdaLR
  2. Cosine annealing — default decay choice; smooth, no hyperparameter tuning
  3. SGDR — add CosineAnnealingWarmRestarts when plateau is a risk (non-convex losses)
  4. One-cycle — for fast single-run training; set max_lr via LR-finder, pct_start=0.3
  5. Transformer schedule — for self-attention models; set base_lr=1.0, tune d_model & N_warm
  6. Implement any schedule with LambdaLR — the lambda is just a function of step

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

40. LR schedule selection recipe

Pattern

  1. Always warmup Adam — N_warm ≈ 5–10 % of total steps; use LambdaLR
  2. Cosine annealing — default decay choice; smooth, no hyperparameter tuning
  3. SGDR — add CosineAnnealingWarmRestarts when plateau is a risk (non-convex losses)
  4. One-cycle — for fast single-run training; set max_lr via LR-finder, pct_start=0.3
  5. Transformer schedule — for self-attention models; set base_lr=1.0, tune d_model & N_warm
  6. Implement any schedule with LambdaLR — the lambda is just a function of step

41. Where does it stop working: LR schedule selection recipe

Edge cases

Discussion prompt

LR schedule selection recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Always warmup Adam — N_warm ≈ 5–10 % of total steps; use LambdaLR
  2. Cosine annealing — default decay choice; smooth, no hyperparameter tuning
  3. SGDR — add CosineAnnealingWarmRestarts when plateau is a risk (non-convex losses)
  4. One-cycle — for fast single-run training; set max_lr via LR-finder, pct_start=0.3
  5. Transformer schedule — for self-attention models; set base_lr=1.0, tune d_model & N_warm
  6. Implement any schedule with LambdaLR — the lambda is just a function of step

42. Rule out three: Check yourself — warmup purpose

Elimination

Eliminate the wrong options

Why does linear warmup stabilize Adam in early training?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Adam's variance estimate v starts at 0; a small LR lets v accumulate before the full step size is trusted
  • B. A small LR reduces gradient noise by averaging more mini-batches
  • C. Warmup prevents the loss from overshooting in cosine annealing
  • D. Adam's bias correction makes early steps too small without warmup

Survives elimination: A

Why: Adam's per-parameter second moment estimate v(0) = 0. The effective step size is gradient / sqrt(v + ε) — when v ≈ 0, ε dominates and the step is large and unstable. Warmup keeps LR tiny until v is meaningful.

43. Check yourself — warmup purpose

Check

Think before clicking.

Check your understanding

Why does linear warmup stabilize Adam in early training?

  • A. Adam's variance estimate v starts at 0; a small LR lets v accumulate before the full step size is trusted (correct)
  • B. A small LR reduces gradient noise by averaging more mini-batches
  • C. Warmup prevents the loss from overshooting in cosine annealing
  • D. Adam's bias correction makes early steps too small without warmup

Answer: A

Why: Adam's per-parameter second moment estimate v(0) = 0. The effective step size is gradient / sqrt(v + ε) — when v ≈ 0, ε dominates and the step is large and unstable. Warmup keeps LR tiny until v is meaningful.

Why B tempts people
Gradient noise is reduced by a larger batch, not a smaller LR. Mini-batch size is unchanged during warmup.
Why C tempts people
Warmup is independent of cosine annealing — it's a concern about Adam's internal state, not about LR decay.
Why D tempts people
Adam's bias correction (×1/(1-β^t)) actually inflates early steps, not deflates them — warmup counters this inflation.

44. Answer it before you see the options: Check yourself — cosine vs SGDR

Prediction

Predict first

What does SGDR add over plain cosine annealing?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: Periodic restarts back to LR_max, helping escape sharp local minima

Why: SGDR (Loshchilov & Hutter 2017) appends a snap-back to LR_max at the end of each T_0-step cycle. This lets the optimizer escape narrow minima it would otherwise converge into — the restart provides the perturbation.

45. Check yourself — cosine vs SGDR

Check

Distinguish the two cosine variants.

Check your understanding

What does SGDR add over plain cosine annealing?

  • A. Periodic restarts back to LR_max, helping escape sharp local minima (correct)
  • B. An inverse-cosine momentum schedule coupled to the LR
  • C. A warmup phase at the start of each cosine period
  • D. A linear final decay after the cosine reaches LR_min

Answer: A

Why: SGDR (Loshchilov & Hutter 2017) appends a snap-back to LR_max at the end of each T_0-step cycle. This lets the optimizer escape narrow minima it would otherwise converge into — the restart provides the perturbation.

Why B tempts people
Momentum coupling is the one-cycle policy (Smith 2019), not SGDR.
Why C tempts people
SGDR does not add a warmup within each cosine period — the restart jumps directly to LR_max.
Why D tempts people
SGDR does not append a linear decay; it snaps back to LR_max and starts a new cosine cycle.

46. Rule out three: Check yourself — Transformer schedule peak

Elimination

Eliminate the wrong options

For d_model=512 and N_warm=4000, at which step does the Transformer LR reach its peak, and what is that peak?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. step 4000, lr ≈ 6.99e-04
  • B. step 512, lr ≈ 1.0e-03
  • C. step 1, lr ≈ 1.75e-04
  • D. step 8000, lr ≈ 4.94e-04

Survives elimination: A

Why: The two terms cross when t^-0.5 = t × N_warm^-1.5, which gives t = N_warm = 4000. Peak = d_model^-0.5 × N_warm^-0.5 = 512^-0.5 × 4000^-0.5 ≈ 6.99e-04. Verified in Python.

47. Check yourself — Transformer schedule peak

Check

Read the formula, not memory.

Check your understanding

For d_model=512 and N_warm=4000, at which step does the Transformer LR reach its peak, and what is that peak?

  • A. step 4000, lr ≈ 6.99e-04 (correct)
  • B. step 512, lr ≈ 1.0e-03
  • C. step 1, lr ≈ 1.75e-04
  • D. step 8000, lr ≈ 4.94e-04

Answer: A

Why: The two terms cross when t^-0.5 = t × N_warm^-1.5, which gives t = N_warm = 4000. Peak = d_model^-0.5 × N_warm^-0.5 = 512^-0.5 × 4000^-0.5 ≈ 6.99e-04. Verified in Python.

Why B tempts people
d_model appears as a scale factor (d_model^-0.5), not as a step index — the peak is always at step=N_warm.
Why C tempts people
Step 1 is the first warmup step, where lr ≈ 1.70e-07 (nearly zero), far from the peak.
Why D tempts people
Step 8000 is past the peak — by then the sqrt-decay dominates and lr ≈ 4.94e-04, lower than the peak.

48. Your turn: implement all five schedules

Section

Project

49. Project: LR schedule toolkit with LambdaLR

Concept

Implement all five schedules from scratch using LambdaLR. Each lambda is a pure Python function — no scheduler classes except LambdaLR itself.

#schedulelambda signature
1Linear warmupwarmup_lambda(step)
2Cosine annealingcosine_lambda(step)
3SGDRsgdr_lambda(step)
4One-cycle LRonecycle_lambda(step)
5Transformer (inv-sqrt)transformer_lambda(step)

Build rule: set base_lr=1.0 in your optimizer so the lambda directly controls the effective LR — no accidental scaling.

50. Break it if you can: Project: LR schedule toolkit with LambdaLR

Counterexample

Discussion prompt

Implement all five schedules from scratch using LambdaLR. Each lambda is a pure Python function — no scheduler classes except LambdaLR itself.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Build rule: set base_lr=1.0 in your optimizer so the lambda directly controls the effective LR — no accidental scaling.

51. Milestone 1 — warmup + cosine lambdas

Worked example

Your turn: write warmup_lambda and cosine_lambda. Predict what cosine_lambda(50) returns for LR_max=1e-3, LR_min=0, T=100.

Hint: warmup multiplier = min(step/N_warm, 1.0); cosine = 0.5*(1 + cos(pi*step/T)).

import math
N_warm = 100; T = 100

def warmup_lambda(step):
    return min(step / N_warm, 1.0)

def cosine_lambda(step):
    return 0.5 * (1.0 + math.cos(math.pi * step / T))

for s in [0, 25, 50, 75, 100]:
    print(f's={s:3d}  warmup={warmup_lambda(s):.3f}  cosine={cosine_lambda(s):.4f}')
stepwarmup multcosine mult
00.0001.0000
250.2500.8536
500.5000.5000
750.7500.1464
1001.0000.0000

52. Watch it run: Milestone 1 — warmup + cosine lambdas

Pattern

Step through it

Step through Milestone 1 — warmup + cosine lambdas one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: step is 0
  2. Step 2: step is 25
  3. Step 3: step is 50
  4. Step 4: step is 75
  5. Step 5: step is 100

53. Milestone 2 — SGDR lambda

Worked example

Your turn: implement sgdr_lambda with T_0=50 restarts using only math.cos and %. Predict the LR at steps 0, 25, 50, 75.

Hint: t_in_cycle = step % T_0; return cosine over [0, T_0]. At t_in_cycle=0 the LR snaps back to 1.0.

import math
T_0 = 50

def sgdr_lambda(step):
    t = step % T_0
    return 0.5 * (1.0 + math.cos(math.pi * t / T_0))

for s in [0, 25, 50, 75, 100]:
    print(f'step {s:3d}: mult={sgdr_lambda(s):.4f}')
stept_in_cyclemultiplier (= lr if base=1)
001.0000
25250.5000
5001.0000 (restart)
75250.5000
10001.0000 (restart)

54. What each one costs: Milestone 2 — SGDR lambda

Trade off

Comparison matrix

From Milestone 2 — SGDR lambda: every row here is a choice with a cost. Fill the multiplier (= lr if base=1) column, then say which row you would actually pick and what you give up for it.

stept_in_cyclemultiplier (= lr if base=1)
001.0000
25250.5000
5001.0000 (restart)
75250.5000
10001.0000 (restart)

55. The full program: all five schedules

Concept

import math, torch, torch.optim as optim

# shared hyperparams
N_warm, T, T_0 = 100, 200, 50
N_total, pct = 200, 0.3
d_model, N_w = 512, 4000

def warmup_fn(s):      return min(s/N_warm, 1.0)
def cosine_fn(s):      return 0.5*(1+math.cos(math.pi*s/T))
def sgdr_fn(s):        return 0.5*(1+math.cos(math.pi*(s%T_0)/T_0))
def onecycle_fn(s):
    k = int(N_total*pct)
    if s < k:  return s/k
    return 0.5*(1+math.cos(math.pi*(s-k)/(N_total-k)))
def transformer_fn(s):
    s = max(s,1)
    return d_model**-0.5*min(s**-0.5, s*N_w**-1.5)

for name, fn in [('warmup', warmup_fn), ('cosine', cosine_fn),
                  ('sgdr',   sgdr_fn),   ('onecycle', onecycle_fn),
                  ('transf', transformer_fn)]:
    print(f'{name:10s}  step50={fn(50):.4f}  step100={fn(100):.4f}')
schedulestep 50step 100
warmup0.50001.0000
cosine0.50000.0000
sgdr1.00001.0000
onecycle0.83330.4268
transf0.00010.0002

All five lambdas implemented, no scheduler classes needed except LambdaLR — you own every schedule from first principles.

56. Fill in: step 100 for The full program: all five schedules

Comparison

Comparison matrix

From The full program: all five schedules: refill the step 100 column from what you know. The rest of the table is as it appeared.

schedulestep 50step 100
warmup0.50001.0000
cosine0.50000.0000
sgdr1.00001.0000
onecycle0.83330.4268
transf0.00010.0002

57. Show it off

Concept

Out loud, slides closed: (1) derive the cosine formula from scratch and state what happens at t=0, T/2, T; (2) explain why the Transformer peak is at step=N_warm; (3) describe the LR/momentum anti-correlation in one-cycle and why it helps.

Stretch (homework from lesson plan): implement the LR finder (Leslie Smith) — sweep LR from 1e-7 to 10 over 100 steps, log loss vs LR, find the inflection point. That point is your max_lr for one-cycle. Compare constant LR vs cosine annealing vs one-cycle on a real training run and plot curves.

58. Connect it up: Lesson 62: Learning Rate Schedules

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Why schedules? The fixed-LR problem · Schedule 1: Linear Warmup · Schedule 2: Cosine Annealing · Schedule 3: One-Cycle Policy · Schedule 4: Transformer / Inverse Sqrt · Your turn: implement all five schedules. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

59. What you can do now

Recap

schedulekey formulause when
warmuplr × min(t/N, 1)always (Adam)
cosine0.5·(1+cos(πt/T))default decay
SGDRcosine + restart at T_0non-convex, plateau risk
one-cycleLR↑ then↓; momentum inversesingle fast run
transformerd^-0.5 · min(t^-0.5, t·N^-1.5)self-attention models

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 62 (LR Schedules) — Barron · USAAIO Round 2 Preparation, 2026
  2. PyTorch LR scheduler docs — CosineAnnealingLR, CosineAnnealingWarmRestarts, OneCycleLR, LambdaLR
  3. Loshchilov & Hutter, SGDR: Stochastic Gradient Descent with Warm Restarts, ICLR 2017 — arXiv 1608.03983
  4. Smith & Topin, Super-Convergence: Very Fast Training of Neural Networks Using Large Learning Rates, 2019 — arXiv 1708.07120
  5. Vaswani et al., Attention Is All You Need (LR schedule), NeurIPS 2017 — arXiv 1706.03762
  6. All schedule values verified with torch 2.7.1 + numpy 2.2.6, real execution, June 2026 — local run

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108