USAAIO Lesson 57, from Week 20. It compares feature extraction with fine-tuning, covers freezing the pretrained layers, and gives the four-quadrant decision rule that crosses data size with domain similarity. It then covers progressive layer unfreezing, differential learning rates - typically ten times lower for the pretrained layers - and the data-augmentation strategies mixup and cutmix. You build a frozen-backbone classifier and then progressively unfreeze it. The lesson runs to 30 slides.
Subject: Machine Learning · 57 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 57 · Week 20
Freeze, adapt, unfreeze: reuse pretrained representations to hit high accuracy with far less data. The strategy behind every modern vision and NLP pipeline.
Objectives
Warm-up
Discussion prompt
Before we open Lesson 57: Transfer Learning & Fine-Tuning: without looking back, what was the main idea of ResNet, Depthwise Separable Convolutions & EfficientNet, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
ResNet identity shortcuts and vanishing-gradient fix, BasicBlock vs Bottleneck designs, pre-activation vs post-activation BatchNorm, depthwise separable convolutions (MobileNet ~8x parameter reduction), and EfficientNet compound scaling of depth+width+resolution.
Section
Part 1 of 4
Concept
A CNN pretrained on ImageNet (1.2M images, 1000 classes) has learned reusable low-level features: edges, textures, colors in early layers; object parts in later layers. These cost millions of GPU-hours.
| layer depth | what it detects | transferability |
|---|---|---|
| early (conv1–2) | edges, gradients, blobs | high — universal |
| middle (conv3–4) | textures, patterns | high — broadly useful |
| late (fc / head) | class-specific scores | low — task-specific |
Transfer learning reuses the costly early/middle representations and replaces only the task-specific head.
Comparison
Comparison matrix
From What a pretrained model already knows: refill the what it detects column from what you know. The rest of the table is as it appeared.
| layer depth | what it detects | transferability |
|---|---|---|
| early (conv1–2) | edges, gradients, blobs | high — universal |
| middle (conv3–4) | textures, patterns | high — broadly useful |
| late (fc / head) | class-specific scores | low — task-specific |
Concept
Feature extraction freezes all pretrained layers (requires_grad = False) and adds a new classification head. Gradients never flow into the backbone — it is a fixed feature extractor.
| layer group | requires_grad | updated by optimizer |
|---|---|---|
| backbone (pretrained) | False | no |
| new head (random init) | True | yes |
With backbone params = 3424 and head params = 165, you are training only 165 of 3589 total parameters — far less risk of overfitting on small data.
Trade off
Comparison matrix
From Feature extraction: freeze the backbone: every row here is a choice with a cost. Fill the requires_grad column, then say which row you would actually pick and what you give up for it.
| layer group | requires_grad | updated by optimizer |
|---|---|---|
| backbone (pretrained) | False | no |
| new head (random init) | True | yes |
Estimation
Predict first
Pretrained backbone (64-d -> 32-d). Freeze it, attach a 5-class head, and train on a small target dataset (100 samples).
Commit before you compute: what does Freeze backbone, train new head come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Only head.parameters() are passed to the optimizer
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. backbone params have requires_grad=False so autograd skips them entirely — freezing is enforced at the graph level, not just by the optimizer.
Worked example
Pretrained backbone (64-d -> 32-d). Freeze it, attach a 5-class head, and train on a small target dataset (100 samples).
import torch, torch.nn as nn, torch.optim as optim
# backbone loaded from checkpoint (weights already learned)
for p in backbone.parameters():
p.requires_grad = False # freeze
head = nn.Linear(32, 5) # new random head
model = nn.Sequential(backbone, head)
opt = optim.Adam(head.parameters(), lr=1e-3)
for _ in range(400):
opt.zero_grad()
nn.CrossEntropyLoss()(model(X_tr), y_tr).backward()
opt.step()Only head.parameters() are passed to the optimizer
Why: backbone params have requires_grad=False so autograd skips them entirely — freezing is enforced at the graph level, not just by the optimizer.
| mode | params trained | test acc (small, n=100) |
|---|---|---|
| feature extraction (frozen) | 165 | 0.452 |
| full fine-tuning (all layers) | 3589 | 0.595 |
| from scratch (no pretraining) | 3589 | 0.790 |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Only head.parameters() are passed to the optimizer
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Pretrained backbone (64-d -> 32-d). Freeze it, attach a 5-class head, and train on a small target dataset (100 samples).
Concept
Fine-tuning sets requires_grad = True on the entire network and trains all layers end-to-end with a small learning rate (e.g. 1e-4) to avoid overwriting the pretrained representations.
With large target data, fine-tuning consistently beats feature extraction because every layer adapts to the new distribution. With small data it can overfit — pretrained weights are too aggressively overwritten.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Fine-tune with the same LR used during pretraining (lr=1e-2). The backbone has already converged — just kick off the optimizer.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: A large LR catastrophically overwrites the pretrained weights.
Fine-tune with a much smaller LR (1e-4 to 1e-5) to nudge, not overwrite, the pretrained weights.
Why: A large LR catastrophically overwrites the pretrained weights. The backbone forgets everything and the model trains from noise — often worse than scratch.
Trap
Fine-tune with the same LR used during pretraining (lr=1e-2). The backbone has already converged — just kick off the optimizer.
opt = Adam(model.parameters(), lr=1e-2)
Why: A large LR catastrophically overwrites the pretrained weights. The backbone forgets everything and the model trains from noise — often worse than scratch.
Fine-tune with a much smaller LR (1e-4 to 1e-5) to nudge, not overwrite, the pretrained weights.
opt = Adam(model.parameters(), lr=1e-4)
Why: Pretrained weights are near a good solution already. A small LR makes tiny adjustments toward the target distribution while preserving the learned features.
Break the constraint
Discussion prompt
The rule this trap just fixed:
Pretrained weights are near a good solution already. A small LR makes tiny adjustments toward the target distribution while preserving the learned features.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
A large LR catastrophically overwrites the pretrained weights. The backbone forgets everything and the model trains from noise — often worse than scratch.
Section
Part 2 of 4
Concept
Two axes decide the strategy: dataset size (few vs many labeled target examples) and domain similarity (ImageNet-like photos vs X-rays vs audio spectrograms).
| data size | similar domain | recommended strategy |
|---|---|---|
| small | yes | feature extraction — frozen backbone, train head |
| small | no | risky; try feature extraction + heavy augmentation |
| large | yes | full fine-tuning — all layers, lr=1e-4 |
| large | no | fine-tune later layers only, or train from scratch |
The intuition: the more data you have, the more you can afford to adjust the backbone. The more similar the domains, the more the pretrained features generalize.
Counterexample
Discussion prompt
Two axes decide the strategy: dataset size (few vs many labeled target examples) and domain similarity (ImageNet-like photos vs X-rays vs audio spectrograms).
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
The intuition: the more data you have, the more you can afford to adjust the backbone. The more similar the domains, the more the pretrained features generalize.
Concept
Verified on a similar-domain synthetic task (source: 4-class 20-feature, target: 5-class, overlapping informative features; pretrain acc 0.962):
| strategy | small (n=100) | large (n=1600) |
|---|---|---|
| feature extraction | 0.452 | 0.580 |
| full fine-tuning | 0.595 | 0.695 |
| from scratch | 0.790 | 0.925 |
From-scratch wins here because the two tasks share the feature space and enough data is available. In real CV with millions of ImageNet parameters and 1000 target images, feature extraction is often the only viable option.
Analogy
Discussion prompt
Explain Measured decision: data size matters most by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Verified on a similar-domain synthetic task (source: 4-class 20-feature, target: 5-class, overlapping informative features; pretrain acc 0.962):
Section
Part 3 of 4
Concept
Progressive unfreezing (Howard & Ruder, ULMFiT) thaws layers from top (near head) to bottom (early layers), training each phase before proceeding.
The key insight: later layers contain the most task-specific knowledge; early layers contain universal features. Unfreeze the task-specific layers first so they adapt before the universal ones move.
Estimation
Predict first
Backbone: fc1(20->64) + fc2(64->32). Phase by phase — each with a lower LR for the newly unfrozen layer.
Commit before you compute: what does Three-phase progressive unfreezing come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Accuracy climbs 0.470 -> 0.630 -> 0.717 as layers unfreeze
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Each phase lets the newly unfrozen layer adapt with a carefully chosen small LR, preventing the early layers from overwriting their pretrained low-level features.
Worked example
Backbone: fc1(20->64) + fc2(64->32). Phase by phase — each with a lower LR for the newly unfrozen layer.
# Phase 1: head only
for p in model.backbone.parameters(): p.requires_grad = False
opt1 = optim.Adam(model.head.parameters(), lr=1e-3)
train(model, opt1, epochs=150) # acc: 0.470
# Phase 2: unfreeze fc2 (top backbone layer)
for p in model.backbone.fc2.parameters(): p.requires_grad = True
opt2 = optim.Adam([
{'params': model.backbone.fc2.parameters(), 'lr': 1e-4},
{'params': model.head.parameters(), 'lr': 1e-3}
])
train(model, opt2, epochs=100) # acc: 0.630
# Phase 3: unfreeze fc1 (bottom backbone layer)
for p in model.backbone.fc1.parameters(): p.requires_grad = True
opt3 = optim.Adam([
{'params': model.backbone.fc1.parameters(), 'lr': 1e-5},
{'params': model.backbone.fc2.parameters(), 'lr': 1e-4},
{'params': model.head.parameters(), 'lr': 1e-3}
])
train(model, opt3, epochs=100) # acc: 0.717Accuracy climbs 0.470 -> 0.630 -> 0.717 as layers unfreeze
Why: Each phase lets the newly unfrozen layer adapt with a carefully chosen small LR, preventing the early layers from overwriting their pretrained low-level features.
| phase | frozen | unfrozen (LR) | test acc |
|---|---|---|---|
| 1 | fc1, fc2 | head (1e-3) | 0.470 |
| 2 | fc1 | fc2 (1e-4), head (1e-3) | 0.630 |
| 3 | none | fc1 (1e-5), fc2 (1e-4), head (1e-3) | 0.717 |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Accuracy climbs 0.470 -> 0.630 -> 0.717 as layers unfreeze
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Backbone: fc1(20->64) + fc2(64->32). Phase by phase — each with a lower LR for the newly unfrozen layer.
Concept
Differential LRs skip the per-phase orchestration by assigning a lower LR to pretrained layer groups from the start.
The rule of thumb: 10x lower LR for each group of pretrained layers vs the new head. For a 3-group network: head 1e-3, top backbone layer 1e-4, bottom backbone layer 1e-5.
| layer group | LR | rationale |
|---|---|---|
| head (random init) | 1e-3 | needs to learn from scratch |
| backbone fc2 (top) | 1e-4 (10x lower) | task-specific features need some adaptation |
| backbone fc1 (bottom) | 1e-5 (100x lower) | universal features should barely move |
Comparison
Comparison matrix
From Differential learning rates: refill the LR column from what you know. The rest of the table is as it appeared.
| layer group | LR | rationale |
|---|---|---|
| head (random init) | 1e-3 | needs to learn from scratch |
| backbone fc2 (top) | 1e-4 (10x lower) | task-specific features need some adaptation |
| backbone fc1 (bottom) | 1e-5 (100x lower) | universal features should barely move |
Estimation
Predict first
Set up a single optimizer with per-group LRs. PyTorch Adam accepts a list of {'params': ..., 'lr': ...} dicts.
Commit before you compute: what does Differential LR with param groups come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: zero_grad clears ALL param groups in one call
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The optimizer internally iterates over every param group.
Worked example
Set up a single optimizer with per-group LRs. PyTorch Adam accepts a list of {'params': ..., 'lr': ...} dicts.
opt = optim.Adam([
{'params': model.backbone.fc1.parameters(), 'lr': 1e-5},
{'params': model.backbone.fc2.parameters(), 'lr': 1e-4},
{'params': model.head.parameters(), 'lr': 1e-3}
])
# All 3589 params train together — just at different speeds
for _ in range(300):
opt.zero_grad()
nn.CrossEntropyLoss()(model(X_tr), y_tr).backward()
opt.step()
# test acc: 0.717 (matches progressive unfreezing final phase)zero_grad clears ALL param groups in one call
Why: The optimizer internally iterates over every param group. There is no need to call zero_grad separately per group.
| param group | params | LR |
|---|---|---|
| backbone.fc1 | 1344 | 1e-5 |
| backbone.fc2 | 2080 | 1e-4 |
| head | 165 | 1e-3 |
Pattern
Step through it
Step through Differential LR with param groups one row at a time. What is driving the change, and what would the row after the last one be?
Anomaly
Predict first
A student writes this, and it looks reasonable:
Add weight_decay=1e-2 uniformly across all param groups to regularize the fine-tuning.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: weight_decay shrinks ALL weights toward 0 including the pretrained backbone.
Apply weight_decay only to the new head (or not at all if the frozen phases already provided implicit regularization).
Why: weight_decay shrinks ALL weights toward 0 including the pretrained backbone. This aggressively erodes the useful representations that took millions of examples to learn.
Trap
Add weight_decay=1e-2 uniformly across all param groups to regularize the fine-tuning.
opt = Adam(model.parameters(), lr=1e-4, weight_decay=1e-2)
Why: weight_decay shrinks ALL weights toward 0 including the pretrained backbone. This aggressively erodes the useful representations that took millions of examples to learn.
Apply weight_decay only to the new head (or not at all if the frozen phases already provided implicit regularization).
opt = Adam([{'params': backbone_params, 'lr': 1e-5, 'weight_decay': 0}, {'params': head_params, 'lr': 1e-3, 'weight_decay': 1e-2}])
Why: Regularize the random-init head to prevent overfitting while leaving the carefully pretrained backbone weights near their learned values.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
backbone params = 3424 and head params = 165, you are training only 165 of 3589 total parameters — far less risk of overfitting on small data.; Two axes decide the strategy: dataset size (few vs many labeled target examples) and domain similarity (ImageNet-like photos vs X-rays vs audio spectrograms).lr=1e-2). The backbone has already converged — just kick off the optimizer.; Add weight_decay=1e-2 uniformly across all param groups to regularize the fine-tuning.Section
Part 4 of 4
Concept
Fine-tuning on small datasets risks catastrophic forgetting and overfitting simultaneously. Augmentation creates effective diversity from limited labeled samples.
| technique | what it does | good for |
|---|---|---|
| random crop + flip | spatial perturbations | standard baseline; ResNet training |
| color jitter | brightness/contrast/saturation shifts | lighting variability |
| mixup | linear blend: x=lamx_i+(1-lam)x_j, y similarly | calibration; inter-class boundaries |
| cutmix | paste a patch from one image into another, blend labels by area | stronger spatial regularization than mixup |
Mixup and CutMix are label-level augmentations: the target is also a convex combination of one-hot vectors, so CrossEntropyLoss must support soft targets (pass the mixed label vector directly).
Explain it
Discussion prompt
Explain Why augmentation matters more in fine-tuning to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Fine-tuning on small datasets risks catastrophic forgetting and overfitting simultaneously. Augmentation creates effective diversity from limited labeled samples.
Concept
Sample mixing coefficient lam ~ Beta(alpha, alpha) (typically alpha=0.2). Blend a pair of training examples:
\[ \tilde{x} = \lambda x_i + (1-\lambda) x_j \qquad \tilde{y} = \lambda y_i + (1-\lambda) y_j \]
At alpha=0.2, lam concentrates near 0 and 1 (Beta distribution), so most blends are 80/20 mixtures. At alpha=1.0, lam~Uniform(0,1) — fully uniform blends.
| alpha | Beta shape | effect |
|---|---|---|
| 0.2 | U-shaped (mass at 0,1) | mild mixing — mostly one image |
| 1.0 | flat | uniform mixing — 50/50 average common |
| 2.0 | bell (mass at 0.5) | aggressive mixing — always near 50/50 |
Analogy
Discussion prompt
Explain Mixup: the math by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Sample mixing coefficient lam ~ Beta(alpha, alpha) (typically alpha=0.2). Blend a pair of training examples:
Ranking
Put in order
These are the steps of Transfer learning strategy recipe, scrambled. Put them back in order before the next slide shows you.
nn.Linear(feat_dim, n_classes)p.requires_grad = False); train head for N epochsWhy: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
nn.Linear(feat_dim, n_classes)p.requires_grad = False); train head for N epochsEdge cases
Discussion prompt
Transfer learning strategy recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
nn.Linear(feat_dim, n_classes)p.requires_grad = False); train head for N epochsElimination
Eliminate the wrong options
You call for p in model.backbone.parameters(): p.requires_grad = False and then opt = Adam(model.parameters(), lr=1e-3). What happens during backprop?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: requires_grad=False is a graph-level flag. Autograd stops building the computation graph at those tensors, saving both compute and memory. No gradient is ever computed for the backbone — the optimizer has nothing to apply, so the weights are truly unchanged.
Check
Predict the behavior before running.
Check your understanding
You call for p in model.backbone.parameters(): p.requires_grad = False and then opt = Adam(model.parameters(), lr=1e-3). What happens during backprop?
Answer: A
Why: requires_grad=False is a graph-level flag. Autograd stops building the computation graph at those tensors, saving both compute and memory. No gradient is ever computed for the backbone — the optimizer has nothing to apply, so the weights are truly unchanged.
Prediction
Predict first
In a 3-phase progressive unfreezing schedule, your val loss stops improving at Phase 2 (top backbone layer unfrozen). What is the correct next step?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Stop training — unfreezing further would overwrite useful low-level features without benefit
Why: Progressive unfreezing stops when val loss no longer improves. Continuing to unfreeze risks catastrophic forgetting of the low-level features that make the backbone valuable. The plateau is the signal to stop, not to add capacity.
Check
Think about which phase to change.
Check your understanding
In a 3-phase progressive unfreezing schedule, your val loss stops improving at Phase 2 (top backbone layer unfrozen). What is the correct next step?
Answer: A
Why: Progressive unfreezing stops when val loss no longer improves. Continuing to unfreeze risks catastrophic forgetting of the low-level features that make the backbone valuable. The plateau is the signal to stop, not to add capacity.
Elimination
Eliminate the wrong options
You apply mixup with lam=0.7 to two training examples (class 2 and class 4 in a 5-class problem). What label vector do you pass to the loss?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Mixup blends labels with the same lambda as the input: y_tilde = 0.7one_hot(2) + 0.3one_hot(4) = [0,0,0.7,0,0.3]. This soft label is passed directly to CrossEntropyLoss (which handles non-integer targets).
Check
A subtle but exam-worthy detail.
Check your understanding
You apply mixup with lam=0.7 to two training examples (class 2 and class 4 in a 5-class problem). What label vector do you pass to the loss?
Answer: A
Why: Mixup blends labels with the same lambda as the input: y_tilde = 0.7one_hot(2) + 0.3one_hot(4) = [0,0,0.7,0,0.3]. This soft label is passed directly to CrossEntropyLoss (which handles non-integer targets).
Section
Project
Concept
Build a transfer learning pipeline from scratch: pretrain a backbone on a source task, then transfer it to a 5-class target task using all three strategies.
| # | requirement | key API |
|---|---|---|
| 1 | Freeze backbone, train head only (feature extraction) | p.requires_grad = False |
| 2 | Full fine-tuning from the same pretrained backbone | all params, lr=1e-4 |
| 3 | Progressive unfreezing (3 phases, differential LR) | param groups, 10x LR decay per level |
Use make_classification(n_samples=2000, n_features=20, n_classes=5) as target data. Backbone: Linear(20,64) -> Linear(64,32). Head: Linear(32,5).
Counterexample
Discussion prompt
Build a transfer learning pipeline from scratch: pretrain a backbone on a source task, then transfer it to a 5-class target task using all three strategies.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Use make_classification(n_samples=2000, n_features=20, n_classes=5) as target data. Backbone: Linear(20,64) -> Linear(64,32). Head: Linear(32,5).
Worked example
Your turn: freeze the backbone and train the head for 400 epochs. Predict: will test accuracy be higher or lower than training from scratch?
Hint: for p in model.backbone.parameters(): p.requires_grad = False. Pass only model.head.parameters() to the optimizer.
for p in model.backbone.parameters():
p.requires_grad = False
opt = optim.Adam(model.head.parameters(), lr=1e-3)
for _ in range(400):
opt.zero_grad()
nn.CrossEntropyLoss()(model(X_tr), y_tr).backward()
opt.step()
# evaluate
model.eval()
with torch.no_grad():
acc = (model(X_te).argmax(1) == y_te).float().mean()
print(f'FE acc: {acc:.3f}') # 0.452 on small (100) data| strategy | trained params | test acc |
|---|---|---|
| feature extraction | 165 / 3589 | 0.452 |
| from scratch (reference) | 3589 / 3589 | 0.790 |
Trade off
Comparison matrix
From Milestone 1 — feature extraction: every row here is a choice with a cost. Fill the trained params column, then say which row you would actually pick and what you give up for it.
| strategy | trained params | test acc |
|---|---|---|
| feature extraction | 165 / 3589 | 0.452 |
| from scratch (reference) | 3589 / 3589 | 0.790 |
Worked example
Your turn: implement the 3-phase schedule. Predict how accuracy changes at each phase.
Hint: between phases, set requires_grad = True for the newly unfrozen layer's parameters and rebuild the optimizer with a param-group dict.
# Phase 1 already done (acc 0.470)
# Phase 2: unfreeze fc2
for p in model.backbone.fc2.parameters():
p.requires_grad = True
opt2 = optim.Adam([
{'params': model.backbone.fc2.parameters(), 'lr': 1e-4},
{'params': model.head.parameters(), 'lr': 1e-3}
])
for _ in range(100):
opt2.zero_grad()
nn.CrossEntropyLoss()(model(X_tr), y_tr).backward()
opt2.step()
# acc: 0.630
# Phase 3: unfreeze fc1
for p in model.backbone.fc1.parameters():
p.requires_grad = True
opt3 = optim.Adam([
{'params': model.backbone.fc1.parameters(), 'lr': 1e-5},
{'params': model.backbone.fc2.parameters(), 'lr': 1e-4},
{'params': model.head.parameters(), 'lr': 1e-3}
])
for _ in range(100):
opt3.zero_grad()
nn.CrossEntropyLoss()(model(X_tr), y_tr).backward()
opt3.step()
# acc: 0.717| phase | unlocked (LR) | test acc |
|---|---|---|
| 1 | head (1e-3) | 0.470 |
| 2 | +fc2 (1e-4) | 0.630 |
| 3 | +fc1 (1e-5) | 0.717 |
Comparison
Comparison matrix
From Milestone 2 — progressive unfreezing: refill the unlocked (LR) column from what you know. The rest of the table is as it appeared.
| phase | unlocked (LR) | test acc |
|---|---|---|
| 1 | head (1e-3) | 0.470 |
| 2 | +fc2 (1e-4) | 0.630 |
| 3 | +fc1 (1e-5) | 0.717 |
Concept
Slides closed, out loud: (1) explain why requires_grad=False saves compute beyond just skipping the optimizer step, (2) justify the 10x LR rule between layer groups, (3) describe when you would stop at Phase 1 vs run all three phases.
Stretch (homework): implement mixup in the training loop for Phase 1 (blend two samples with lam~Beta(0.2,0.2)); verify the soft label is used. Code a ResNet-50 feature extractor on torchvision.datasets.FakeData and confirm frozen vs unfrozen weight norms differ after training.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Feature extraction vs fine-tuning · The 4-quadrant decision rule · Progressive unfreezing + differential LR · Data augmentation for fine-tuning · Your turn: frozen head to progressive unfreezing. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
requires_grad = False)| idea | the one thing to remember |
|---|---|
| feature extraction | freeze backbone; train only head (165/3589 params) |
| fine-tuning LR | 10x-100x smaller than pretraining LR — never full LR |
| progressive unfreezing | top layer first; stop when val loss plateaus |
| differential LR | param groups in Adam: 1e-5 / 1e-4 / 1e-3 from bottom to head |
| mixup | lamx_i + (1-lam)x_j; same blend for labels |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.