Lesson 57: Transfer Learning & Fine-Tuning

USAAIO Lesson 57, from Week 20. It compares feature extraction with fine-tuning, covers freezing the pretrained layers, and gives the four-quadrant decision rule that crosses data size with domain similarity. It then covers progressive layer unfreezing, differential learning rates - typically ten times lower for the pretrained layers - and the data-augmentation strategies mixup and cutmix. You build a frozen-backbone classifier and then progressively unfreeze it. The lesson runs to 30 slides.

Subject: Machine Learning · 57 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Transfer Learning & Fine-Tuning

Title

USAAIO · Lesson 57 · Week 20

Freeze, adapt, unfreeze: reuse pretrained representations to hit high accuracy with far less data. The strategy behind every modern vision and NLP pipeline.

2. By the end of this lesson you can

Objectives

  1. Explain feature extraction (frozen backbone) vs full fine-tuning (all layers updated)
  2. Apply the 4-quadrant decision rule (data size × domain similarity) to choose a strategy
  3. Implement progressive unfreezing: unfreeze layers top-to-bottom with lower LR each level
  4. Set differential learning rates so pretrained layers update 10x slower than the new head
  5. Choose among data augmentation strategies (flip, crop, color jitter, mixup, cutmix) for fine-tuning

3. What survived from ResNet, Depthwise Separable Convolutions & EfficientNet?

Warm-up

Discussion prompt

Before we open Lesson 57: Transfer Learning & Fine-Tuning: without looking back, what was the main idea of ResNet, Depthwise Separable Convolutions & EfficientNet, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

ResNet identity shortcuts and vanishing-gradient fix, BasicBlock vs Bottleneck designs, pre-activation vs post-activation BatchNorm, depthwise separable convolutions (MobileNet ~8x parameter reduction), and EfficientNet compound scaling of depth+width+resolution.

4. Feature extraction vs fine-tuning

Section

Part 1 of 4

5. What a pretrained model already knows

Concept

A CNN pretrained on ImageNet (1.2M images, 1000 classes) has learned reusable low-level features: edges, textures, colors in early layers; object parts in later layers. These cost millions of GPU-hours.

layer depthwhat it detectstransferability
early (conv1–2)edges, gradients, blobshigh — universal
middle (conv3–4)textures, patternshigh — broadly useful
late (fc / head)class-specific scoreslow — task-specific

Transfer learning reuses the costly early/middle representations and replaces only the task-specific head.

6. Fill in: what it detects for What a pretrained model already knows

Comparison

Comparison matrix

From What a pretrained model already knows: refill the what it detects column from what you know. The rest of the table is as it appeared.

layer depthwhat it detectstransferability
early (conv1–2)edges, gradients, blobshigh — universal
middle (conv3–4)textures, patternshigh — broadly useful
late (fc / head)class-specific scoreslow — task-specific

7. Feature extraction: freeze the backbone

Concept

Feature extraction freezes all pretrained layers (requires_grad = False) and adds a new classification head. Gradients never flow into the backbone — it is a fixed feature extractor.

layer grouprequires_gradupdated by optimizer
backbone (pretrained)Falseno
new head (random init)Trueyes

With backbone params = 3424 and head params = 165, you are training only 165 of 3589 total parameters — far less risk of overfitting on small data.

8. What each one costs: Feature extraction: freeze the backbone

Trade off

Comparison matrix

From Feature extraction: freeze the backbone: every row here is a choice with a cost. Fill the requires_grad column, then say which row you would actually pick and what you give up for it.

layer grouprequires_gradupdated by optimizer
backbone (pretrained)Falseno
new head (random init)Trueyes

9. Guess the shape of the answer: Freeze backbone, train new head

Estimation

Predict first

Pretrained backbone (64-d -> 32-d). Freeze it, attach a 5-class head, and train on a small target dataset (100 samples).

Commit before you compute: what does Freeze backbone, train new head come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Only head.parameters() are passed to the optimizer

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. backbone params have requires_grad=False so autograd skips them entirely — freezing is enforced at the graph level, not just by the optimizer.

10. Freeze backbone, train new head

Worked example

Pretrained backbone (64-d -> 32-d). Freeze it, attach a 5-class head, and train on a small target dataset (100 samples).

import torch, torch.nn as nn, torch.optim as optim

# backbone loaded from checkpoint (weights already learned)
for p in backbone.parameters():
    p.requires_grad = False     # freeze

head = nn.Linear(32, 5)         # new random head
model = nn.Sequential(backbone, head)
opt  = optim.Adam(head.parameters(), lr=1e-3)

for _ in range(400):
    opt.zero_grad()
    nn.CrossEntropyLoss()(model(X_tr), y_tr).backward()
    opt.step()

Only head.parameters() are passed to the optimizer

Why: backbone params have requires_grad=False so autograd skips them entirely — freezing is enforced at the graph level, not just by the optimizer.

modeparams trainedtest acc (small, n=100)
feature extraction (frozen)1650.452
full fine-tuning (all layers)35890.595
from scratch (no pretraining)35890.790

11. Work backwards from the answer: Freeze backbone, train new head

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Only head.parameters() are passed to the optimizer

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Pretrained backbone (64-d -> 32-d). Freeze it, attach a 5-class head, and train on a small target dataset (100 samples).

12. Full fine-tuning: unfreeze everything

Concept

Fine-tuning sets requires_grad = True on the entire network and trains all layers end-to-end with a small learning rate (e.g. 1e-4) to avoid overwriting the pretrained representations.

With large target data, fine-tuning consistently beats feature extraction because every layer adapts to the new distribution. With small data it can overfit — pretrained weights are too aggressively overwritten.

13. Something is wrong here: using the default LR when fine-tuning

Anomaly

Predict first

A student writes this, and it looks reasonable:

Fine-tune with the same LR used during pretraining (lr=1e-2). The backbone has already converged — just kick off the optimizer.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: A large LR catastrophically overwrites the pretrained weights.

Fine-tune with a much smaller LR (1e-4 to 1e-5) to nudge, not overwrite, the pretrained weights.

Why: A large LR catastrophically overwrites the pretrained weights. The backbone forgets everything and the model trains from noise — often worse than scratch.

14. Trap: using the default LR when fine-tuning

Trap

The trap

Fine-tune with the same LR used during pretraining (lr=1e-2). The backbone has already converged — just kick off the optimizer.

opt = Adam(model.parameters(), lr=1e-2)

Why: A large LR catastrophically overwrites the pretrained weights. The backbone forgets everything and the model trains from noise — often worse than scratch.

The fix

Fine-tune with a much smaller LR (1e-4 to 1e-5) to nudge, not overwrite, the pretrained weights.

opt = Adam(model.parameters(), lr=1e-4)

Why: Pretrained weights are near a good solution already. A small LR makes tiny adjustments toward the target distribution while preserving the learned features.

15. Break it on purpose: using the default LR when fine-tuning

Break the constraint

Discussion prompt

The rule this trap just fixed:

Pretrained weights are near a good solution already. A small LR makes tiny adjustments toward the target distribution while preserving the learned features.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

A large LR catastrophically overwrites the pretrained weights. The backbone forgets everything and the model trains from noise — often worse than scratch.

16. The 4-quadrant decision rule

Section

Part 2 of 4

17. When to use which strategy

Concept

Two axes decide the strategy: dataset size (few vs many labeled target examples) and domain similarity (ImageNet-like photos vs X-rays vs audio spectrograms).

data sizesimilar domainrecommended strategy
smallyesfeature extraction — frozen backbone, train head
smallnorisky; try feature extraction + heavy augmentation
largeyesfull fine-tuning — all layers, lr=1e-4
largenofine-tune later layers only, or train from scratch

The intuition: the more data you have, the more you can afford to adjust the backbone. The more similar the domains, the more the pretrained features generalize.

18. Break it if you can: When to use which strategy

Counterexample

Discussion prompt

Two axes decide the strategy: dataset size (few vs many labeled target examples) and domain similarity (ImageNet-like photos vs X-rays vs audio spectrograms).

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

The intuition: the more data you have, the more you can afford to adjust the backbone. The more similar the domains, the more the pretrained features generalize.

19. Measured decision: data size matters most

Concept

Verified on a similar-domain synthetic task (source: 4-class 20-feature, target: 5-class, overlapping informative features; pretrain acc 0.962):

strategysmall (n=100)large (n=1600)
feature extraction0.4520.580
full fine-tuning0.5950.695
from scratch0.7900.925

From-scratch wins here because the two tasks share the feature space and enough data is available. In real CV with millions of ImageNet parameters and 1000 target images, feature extraction is often the only viable option.

20. By analogy: Measured decision: data size matters most

Analogy

Discussion prompt

Explain Measured decision: data size matters most by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Verified on a similar-domain synthetic task (source: 4-class 20-feature, target: 5-class, overlapping informative features; pretrain acc 0.962):

21. Progressive unfreezing + differential LR

Section

Part 3 of 4

22. Progressive unfreezing: top-down

Concept

Progressive unfreezing (Howard & Ruder, ULMFiT) thaws layers from top (near head) to bottom (early layers), training each phase before proceeding.

  1. Phase 1: freeze backbone entirely, train head only
  2. Phase 2: unfreeze top backbone layer, use lower LR for it
  3. Phase 3: unfreeze next layer with even lower LR; head keeps its LR
  4. Repeat until all layers are unfrozen or validation loss stops improving

The key insight: later layers contain the most task-specific knowledge; early layers contain universal features. Unfreeze the task-specific layers first so they adapt before the universal ones move.

23. Guess the shape of the answer: Three-phase progressive unfreezing

Estimation

Predict first

Backbone: fc1(20->64) + fc2(64->32). Phase by phase — each with a lower LR for the newly unfrozen layer.

Commit before you compute: what does Three-phase progressive unfreezing come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Accuracy climbs 0.470 -> 0.630 -> 0.717 as layers unfreeze

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Each phase lets the newly unfrozen layer adapt with a carefully chosen small LR, preventing the early layers from overwriting their pretrained low-level features.

24. Three-phase progressive unfreezing

Worked example

Backbone: fc1(20->64) + fc2(64->32). Phase by phase — each with a lower LR for the newly unfrozen layer.

# Phase 1: head only
for p in model.backbone.parameters(): p.requires_grad = False
opt1 = optim.Adam(model.head.parameters(), lr=1e-3)
train(model, opt1, epochs=150)   # acc: 0.470

# Phase 2: unfreeze fc2 (top backbone layer)
for p in model.backbone.fc2.parameters(): p.requires_grad = True
opt2 = optim.Adam([
    {'params': model.backbone.fc2.parameters(), 'lr': 1e-4},
    {'params': model.head.parameters(),         'lr': 1e-3}
])
train(model, opt2, epochs=100)   # acc: 0.630

# Phase 3: unfreeze fc1 (bottom backbone layer)
for p in model.backbone.fc1.parameters(): p.requires_grad = True
opt3 = optim.Adam([
    {'params': model.backbone.fc1.parameters(), 'lr': 1e-5},
    {'params': model.backbone.fc2.parameters(), 'lr': 1e-4},
    {'params': model.head.parameters(),         'lr': 1e-3}
])
train(model, opt3, epochs=100)   # acc: 0.717

Accuracy climbs 0.470 -> 0.630 -> 0.717 as layers unfreeze

Why: Each phase lets the newly unfrozen layer adapt with a carefully chosen small LR, preventing the early layers from overwriting their pretrained low-level features.

phasefrozenunfrozen (LR)test acc
1fc1, fc2head (1e-3)0.470
2fc1fc2 (1e-4), head (1e-3)0.630
3nonefc1 (1e-5), fc2 (1e-4), head (1e-3)0.717

25. Work backwards from the answer: Three-phase progressive unfreezing

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Accuracy climbs 0.470 -> 0.630 -> 0.717 as layers unfreeze

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Backbone: fc1(20->64) + fc2(64->32). Phase by phase — each with a lower LR for the newly unfrozen layer.

26. Differential learning rates

Concept

Differential LRs skip the per-phase orchestration by assigning a lower LR to pretrained layer groups from the start.

The rule of thumb: 10x lower LR for each group of pretrained layers vs the new head. For a 3-group network: head 1e-3, top backbone layer 1e-4, bottom backbone layer 1e-5.

layer groupLRrationale
head (random init)1e-3needs to learn from scratch
backbone fc2 (top)1e-4 (10x lower)task-specific features need some adaptation
backbone fc1 (bottom)1e-5 (100x lower)universal features should barely move

27. Fill in: LR for Differential learning rates

Comparison

Comparison matrix

From Differential learning rates: refill the LR column from what you know. The rest of the table is as it appeared.

layer groupLRrationale
head (random init)1e-3needs to learn from scratch
backbone fc2 (top)1e-4 (10x lower)task-specific features need some adaptation
backbone fc1 (bottom)1e-5 (100x lower)universal features should barely move

28. Guess the shape of the answer: Differential LR with param groups

Estimation

Predict first

Set up a single optimizer with per-group LRs. PyTorch Adam accepts a list of {'params': ..., 'lr': ...} dicts.

Commit before you compute: what does Differential LR with param groups come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: zero_grad clears ALL param groups in one call

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The optimizer internally iterates over every param group.

29. Differential LR with param groups

Worked example

Set up a single optimizer with per-group LRs. PyTorch Adam accepts a list of {'params': ..., 'lr': ...} dicts.

opt = optim.Adam([
    {'params': model.backbone.fc1.parameters(), 'lr': 1e-5},
    {'params': model.backbone.fc2.parameters(), 'lr': 1e-4},
    {'params': model.head.parameters(),         'lr': 1e-3}
])

# All 3589 params train together — just at different speeds
for _ in range(300):
    opt.zero_grad()
    nn.CrossEntropyLoss()(model(X_tr), y_tr).backward()
    opt.step()
# test acc: 0.717 (matches progressive unfreezing final phase)

zero_grad clears ALL param groups in one call

Why: The optimizer internally iterates over every param group. There is no need to call zero_grad separately per group.

param groupparamsLR
backbone.fc113441e-5
backbone.fc220801e-4
head1651e-3

30. Watch it run: Differential LR with param groups

Pattern

Step through it

Step through Differential LR with param groups one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: param group is backbone.fc1
  2. Step 2: param group is backbone.fc2
  3. Step 3: param group is head

31. Something is wrong here: applying weight_decay to pretrained layers

Anomaly

Predict first

A student writes this, and it looks reasonable:

Add weight_decay=1e-2 uniformly across all param groups to regularize the fine-tuning.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: weight_decay shrinks ALL weights toward 0 including the pretrained backbone.

Apply weight_decay only to the new head (or not at all if the frozen phases already provided implicit regularization).

Why: weight_decay shrinks ALL weights toward 0 including the pretrained backbone. This aggressively erodes the useful representations that took millions of examples to learn.

32. Trap: applying weight_decay to pretrained layers

Trap

The trap

Add weight_decay=1e-2 uniformly across all param groups to regularize the fine-tuning.

opt = Adam(model.parameters(), lr=1e-4, weight_decay=1e-2)

Why: weight_decay shrinks ALL weights toward 0 including the pretrained backbone. This aggressively erodes the useful representations that took millions of examples to learn.

The fix

Apply weight_decay only to the new head (or not at all if the frozen phases already provided implicit regularization).

opt = Adam([{'params': backbone_params, 'lr': 1e-5, 'weight_decay': 0}, {'params': head_params, 'lr': 1e-3, 'weight_decay': 1e-2}])

Why: Regularize the random-init head to prevent overfitting while leaving the carefully pretrained backbone weights near their learned values.

33. Which of these survive contact with Lesson 57: Transfer Learning & Fine-Tuning?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Transfer learning reuses the costly early/middle representations and replaces only the task-specific head.; With backbone params = 3424 and head params = 165, you are training only 165 of 3589 total parameters — far less risk of overfitting on small data.; Two axes decide the strategy: dataset size (few vs many labeled target examples) and domain similarity (ImageNet-like photos vs X-rays vs audio spectrograms).
Breaks
Fine-tune with the same LR used during pretraining (lr=1e-2). The backbone has already converged — just kick off the optimizer.; Add weight_decay=1e-2 uniformly across all param groups to regularize the fine-tuning.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 57: Transfer Learning & Fine-Tuning puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

34. Data augmentation for fine-tuning

Section

Part 4 of 4

35. Why augmentation matters more in fine-tuning

Concept

Fine-tuning on small datasets risks catastrophic forgetting and overfitting simultaneously. Augmentation creates effective diversity from limited labeled samples.

techniquewhat it doesgood for
random crop + flipspatial perturbationsstandard baseline; ResNet training
color jitterbrightness/contrast/saturation shiftslighting variability
mixuplinear blend: x=lamx_i+(1-lam)x_j, y similarlycalibration; inter-class boundaries
cutmixpaste a patch from one image into another, blend labels by areastronger spatial regularization than mixup

Mixup and CutMix are label-level augmentations: the target is also a convex combination of one-hot vectors, so CrossEntropyLoss must support soft targets (pass the mixed label vector directly).

36. Teach it back: Why augmentation matters more in fine-tuning

Explain it

Discussion prompt

Explain Why augmentation matters more in fine-tuning to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Fine-tuning on small datasets risks catastrophic forgetting and overfitting simultaneously. Augmentation creates effective diversity from limited labeled samples.

37. Mixup: the math

Concept

Sample mixing coefficient lam ~ Beta(alpha, alpha) (typically alpha=0.2). Blend a pair of training examples:

\[ \tilde{x} = \lambda x_i + (1-\lambda) x_j \qquad \tilde{y} = \lambda y_i + (1-\lambda) y_j \]

At alpha=0.2, lam concentrates near 0 and 1 (Beta distribution), so most blends are 80/20 mixtures. At alpha=1.0, lam~Uniform(0,1) — fully uniform blends.

alphaBeta shapeeffect
0.2U-shaped (mass at 0,1)mild mixing — mostly one image
1.0flatuniform mixing — 50/50 average common
2.0bell (mass at 0.5)aggressive mixing — always near 50/50

38. By analogy: Mixup: the math

Analogy

Discussion prompt

Explain Mixup: the math by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Sample mixing coefficient lam ~ Beta(alpha, alpha) (typically alpha=0.2). Blend a pair of training examples:

39. Rebuild the recipe: Transfer learning strategy recipe

Ranking

Put in order

These are the steps of Transfer learning strategy recipe, scrambled. Put them back in order before the next slide shows you.

  1. Load pretrained backbone; replace head with nn.Linear(feat_dim, n_classes)
  2. Freeze backbone (p.requires_grad = False); train head for N epochs
  3. Decide: small data OR dissimilar domain? Stop here. Large/similar data? Continue
  4. Unfreeze layers top-to-bottom; use 10x lower LR per level deeper (differential LR param groups)
  5. Augment with flip/crop/color jitter; add mixup or cutmix for small datasets
  6. Monitor val loss per phase — stop unfreezing when val loss stops improving

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

40. Transfer learning strategy recipe

Pattern

  1. Load pretrained backbone; replace head with nn.Linear(feat_dim, n_classes)
  2. Freeze backbone (p.requires_grad = False); train head for N epochs
  3. Decide: small data OR dissimilar domain? Stop here. Large/similar data? Continue
  4. Unfreeze layers top-to-bottom; use 10x lower LR per level deeper (differential LR param groups)
  5. Augment with flip/crop/color jitter; add mixup or cutmix for small datasets
  6. Monitor val loss per phase — stop unfreezing when val loss stops improving

41. Where does it stop working: Transfer learning strategy recipe

Edge cases

Discussion prompt

Transfer learning strategy recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Load pretrained backbone; replace head with nn.Linear(feat_dim, n_classes)
  2. Freeze backbone (p.requires_grad = False); train head for N epochs
  3. Decide: small data OR dissimilar domain? Stop here. Large/similar data? Continue
  4. Unfreeze layers top-to-bottom; use 10x lower LR per level deeper (differential LR param groups)
  5. Augment with flip/crop/color jitter; add mixup or cutmix for small datasets
  6. Monitor val loss per phase — stop unfreezing when val loss stops improving

42. Rule out three: Check yourself — freezing

Elimination

Eliminate the wrong options

You call for p in model.backbone.parameters(): p.requires_grad = False and then opt = Adam(model.parameters(), lr=1e-3). What happens during backprop?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Autograd skips computing gradients for the frozen layers; the optimizer step leaves those parameters unchanged
  • B. Gradients are computed for all layers; the optimizer just ignores the frozen ones
  • C. An error is raised because the optimizer received frozen parameters
  • D. The backbone parameters are updated but with zero gradient, so they stay the same numerically

Survives elimination: A

Why: requires_grad=False is a graph-level flag. Autograd stops building the computation graph at those tensors, saving both compute and memory. No gradient is ever computed for the backbone — the optimizer has nothing to apply, so the weights are truly unchanged.

43. Check yourself — freezing

Check

Predict the behavior before running.

Check your understanding

You call for p in model.backbone.parameters(): p.requires_grad = False and then opt = Adam(model.parameters(), lr=1e-3). What happens during backprop?

  • A. Autograd skips computing gradients for the frozen layers; the optimizer step leaves those parameters unchanged (correct)
  • B. Gradients are computed for all layers; the optimizer just ignores the frozen ones
  • C. An error is raised because the optimizer received frozen parameters
  • D. The backbone parameters are updated but with zero gradient, so they stay the same numerically

Answer: A

Why: requires_grad=False is a graph-level flag. Autograd stops building the computation graph at those tensors, saving both compute and memory. No gradient is ever computed for the backbone — the optimizer has nothing to apply, so the weights are truly unchanged.

Why B tempts people
Incorrect — autograd does not compute the gradient at all for requires_grad=False tensors. Skipping is at the graph-build stage, not the optimizer stage.
Why C tempts people
PyTorch silently ignores parameters with requires_grad=False in the optimizer; no error is raised. Only trainable parameters participate in step().
Why D tempts people
If the gradient were computed (even as zero), it would still count as a useless pass. But the whole point is that the gradient is never computed — it is a graph-level stop.

44. Answer it before you see the options: Check yourself — progressive unfreezing

Prediction

Predict first

In a 3-phase progressive unfreezing schedule, your val loss stops improving at Phase 2 (top backbone layer unfrozen). What is the correct next step?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: Stop training — unfreezing further would overwrite useful low-level features without benefit

Why: Progressive unfreezing stops when val loss no longer improves. Continuing to unfreeze risks catastrophic forgetting of the low-level features that make the backbone valuable. The plateau is the signal to stop, not to add capacity.

45. Check yourself — progressive unfreezing

Check

Think about which phase to change.

Check your understanding

In a 3-phase progressive unfreezing schedule, your val loss stops improving at Phase 2 (top backbone layer unfrozen). What is the correct next step?

  • A. Stop training — unfreezing further would overwrite useful low-level features without benefit (correct)
  • B. Immediately raise the LR on the already-unfrozen layers to break out of the plateau
  • C. Unfreeze Phase 3 (deeper layers) with an even lower LR — the plateau may resolve with more capacity
  • D. Restart from Phase 1 with a fresh head initialization

Answer: A

Why: Progressive unfreezing stops when val loss no longer improves. Continuing to unfreeze risks catastrophic forgetting of the low-level features that make the backbone valuable. The plateau is the signal to stop, not to add capacity.

Why B tempts people
Raising LR on pretrained layers is the exact opposite of the safe-update principle — it risks catastrophically overwriting the features that are still working.
Why C tempts people
Unfreezing Phase 3 when Phase 2 has already plateaued adds parameters that are unlikely to help and increases risk of forgetting. Val loss improvement is the gating criterion.
Why D tempts people
Resetting the head wastes the Phase 1 and 2 training. The plateau does not mean the features are wrong — it means the current unfrozen layers have saturated their contribution.

46. Rule out three: Check yourself — mixup labels

Elimination

Eliminate the wrong options

You apply mixup with lam=0.7 to two training examples (class 2 and class 4 in a 5-class problem). What label vector do you pass to the loss?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. [0, 0, 0.7, 0, 0.3] (soft label: 0.7 for class 2, 0.3 for class 4)
  • B. [0, 0, 1, 0, 0] (round to the dominant class 2)
  • C. [0, 0, 0.5, 0, 0.5] (always equal blend regardless of lam)
  • D. [0.14, 0.14, 0.28, 0.14, 0.3] (uniform uncertainty)

Survives elimination: A

Why: Mixup blends labels with the same lambda as the input: y_tilde = 0.7one_hot(2) + 0.3one_hot(4) = [0,0,0.7,0,0.3]. This soft label is passed directly to CrossEntropyLoss (which handles non-integer targets).

47. Check yourself — mixup labels

Check

A subtle but exam-worthy detail.

Check your understanding

You apply mixup with lam=0.7 to two training examples (class 2 and class 4 in a 5-class problem). What label vector do you pass to the loss?

  • A. [0, 0, 0.7, 0, 0.3] (soft label: 0.7 for class 2, 0.3 for class 4) (correct)
  • B. [0, 0, 1, 0, 0] (round to the dominant class 2)
  • C. [0, 0, 0.5, 0, 0.5] (always equal blend regardless of lam)
  • D. [0.14, 0.14, 0.28, 0.14, 0.3] (uniform uncertainty)

Answer: A

Why: Mixup blends labels with the same lambda as the input: y_tilde = 0.7one_hot(2) + 0.3one_hot(4) = [0,0,0.7,0,0.3]. This soft label is passed directly to CrossEntropyLoss (which handles non-integer targets).

Why B tempts people
Rounding to the argmax discards the mixing information and defeats the purpose of mixup — you would be training on the original hard label with a corrupted input.
Why C tempts people
lam=0.5 is a special case of uniform mixing. For lam=0.7, the blend is asymmetric by design — that is the information in the sampled lam.
Why D tempts people
Distributing probability uniformly ignores both lam and the two specific classes. This is not what mixup computes — it is a convex combination of the two one-hot vectors.

48. Your turn: frozen head to progressive unfreezing

Section

Project

49. Project: Transfer Learning Pipeline

Concept

Build a transfer learning pipeline from scratch: pretrain a backbone on a source task, then transfer it to a 5-class target task using all three strategies.

#requirementkey API
1Freeze backbone, train head only (feature extraction)p.requires_grad = False
2Full fine-tuning from the same pretrained backboneall params, lr=1e-4
3Progressive unfreezing (3 phases, differential LR)param groups, 10x LR decay per level

Use make_classification(n_samples=2000, n_features=20, n_classes=5) as target data. Backbone: Linear(20,64) -> Linear(64,32). Head: Linear(32,5).

50. Break it if you can: Project: Transfer Learning Pipeline

Counterexample

Discussion prompt

Build a transfer learning pipeline from scratch: pretrain a backbone on a source task, then transfer it to a 5-class target task using all three strategies.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Use make_classification(n_samples=2000, n_features=20, n_classes=5) as target data. Backbone: Linear(20,64) -> Linear(64,32). Head: Linear(32,5).

51. Milestone 1 — feature extraction

Worked example

Your turn: freeze the backbone and train the head for 400 epochs. Predict: will test accuracy be higher or lower than training from scratch?

Hint: for p in model.backbone.parameters(): p.requires_grad = False. Pass only model.head.parameters() to the optimizer.

for p in model.backbone.parameters():
    p.requires_grad = False

opt = optim.Adam(model.head.parameters(), lr=1e-3)
for _ in range(400):
    opt.zero_grad()
    nn.CrossEntropyLoss()(model(X_tr), y_tr).backward()
    opt.step()

# evaluate
model.eval()
with torch.no_grad():
    acc = (model(X_te).argmax(1) == y_te).float().mean()
print(f'FE acc: {acc:.3f}')  # 0.452 on small (100) data
strategytrained paramstest acc
feature extraction165 / 35890.452
from scratch (reference)3589 / 35890.790

52. What each one costs: Milestone 1 — feature extraction

Trade off

Comparison matrix

From Milestone 1 — feature extraction: every row here is a choice with a cost. Fill the trained params column, then say which row you would actually pick and what you give up for it.

strategytrained paramstest acc
feature extraction165 / 35890.452
from scratch (reference)3589 / 35890.790

53. Milestone 2 — progressive unfreezing

Worked example

Your turn: implement the 3-phase schedule. Predict how accuracy changes at each phase.

Hint: between phases, set requires_grad = True for the newly unfrozen layer's parameters and rebuild the optimizer with a param-group dict.

# Phase 1 already done (acc 0.470)
# Phase 2: unfreeze fc2
for p in model.backbone.fc2.parameters():
    p.requires_grad = True
opt2 = optim.Adam([
    {'params': model.backbone.fc2.parameters(), 'lr': 1e-4},
    {'params': model.head.parameters(),         'lr': 1e-3}
])
for _ in range(100):
    opt2.zero_grad()
    nn.CrossEntropyLoss()(model(X_tr), y_tr).backward()
    opt2.step()
# acc: 0.630

# Phase 3: unfreeze fc1
for p in model.backbone.fc1.parameters():
    p.requires_grad = True
opt3 = optim.Adam([
    {'params': model.backbone.fc1.parameters(), 'lr': 1e-5},
    {'params': model.backbone.fc2.parameters(), 'lr': 1e-4},
    {'params': model.head.parameters(),         'lr': 1e-3}
])
for _ in range(100):
    opt3.zero_grad()
    nn.CrossEntropyLoss()(model(X_tr), y_tr).backward()
    opt3.step()
# acc: 0.717
phaseunlocked (LR)test acc
1head (1e-3)0.470
2+fc2 (1e-4)0.630
3+fc1 (1e-5)0.717

54. Fill in: unlocked (LR) for Milestone 2 — progressive unfreezing

Comparison

Comparison matrix

From Milestone 2 — progressive unfreezing: refill the unlocked (LR) column from what you know. The rest of the table is as it appeared.

phaseunlocked (LR)test acc
1head (1e-3)0.470
2+fc2 (1e-4)0.630
3+fc1 (1e-5)0.717

55. Show it off

Concept

Slides closed, out loud: (1) explain why requires_grad=False saves compute beyond just skipping the optimizer step, (2) justify the 10x LR rule between layer groups, (3) describe when you would stop at Phase 1 vs run all three phases.

Stretch (homework): implement mixup in the training loop for Phase 1 (blend two samples with lam~Beta(0.2,0.2)); verify the soft label is used. Code a ResNet-50 feature extractor on torchvision.datasets.FakeData and confirm frozen vs unfrozen weight norms differ after training.

56. Connect it up: Lesson 57: Transfer Learning & Fine-Tuning

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Feature extraction vs fine-tuning · The 4-quadrant decision rule · Progressive unfreezing + differential LR · Data augmentation for fine-tuning · Your turn: frozen head to progressive unfreezing. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

57. What you can do now

Recap

ideathe one thing to remember
feature extractionfreeze backbone; train only head (165/3589 params)
fine-tuning LR10x-100x smaller than pretraining LR — never full LR
progressive unfreezingtop layer first; stop when val loss plateaus
differential LRparam groups in Adam: 1e-5 / 1e-4 / 1e-3 from bottom to head
mixuplamx_i + (1-lam)x_j; same blend for labels

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 57 (Week 20 — PCA + Transfer Learning + Model Selection) — Barron · USAAIO Round 2 Preparation, 2026
  2. Feature extraction vs fine-tuning accuracy comparison verified with torch 2.7.1 + sklearn, make_classification, June 2026 — Real execution: FE small=0.452, FT small=0.595, FE large=0.580, FT large=0.695
  3. Progressive unfreezing phase accuracy verified: p1=0.470, p2=0.630, p3=0.717 — torch 2.7.1+cpu, June 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108