USAAIO Lesson 77, from Phase 4 on deep learning. It covers the MLP, CNN, and ResNet architectures; the ReLU, sigmoid, and tanh activations and the mechanics of the vanishing gradient; BatchNorm, dropout, and Xavier and Kaiming initialization; the PyTorch training loop line by line; backpropagation through the chain rule, with a concrete two-layer trace; the SGD, Momentum, Adam, and AdamW optimizers and learning-rate scheduling; and Dataset and DataLoader for mini-batch pipelines. You build a complete digits classifier with every one of these techniques applied. The lesson runs to 32 slides.
Subject: Machine Learning · 61 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 77 · Phase 4
MLP → CNN → ResNet. Activations, normalization, dropout, initialization, backprop, optimizers, and the full PyTorch pipeline — every concept Barron needs to reason about any deep model on the exam.
Objectives
Warm-up
Discussion prompt
Before we open Lesson 77: Deep Learning Foundations: without looking back, what was the main idea of Algorithm Review & Selection, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
flash-review of 7 core classifiers (logistic regression, SVM, decision trees, random forest, GBM, kNN, Naive Bayes) — loss functions, gradients, hyperparameters, and bias-variance profiles; whiteboard-style derivations; USAAIO exam question patterns; and an algorithm-selection framework. Implement and benchmark all 7 on a synthetic dataset.
Section
Part 1 of 5
Concept
A multi-layer perceptron (Lesson 40) stacks Linear → activation blocks. Every unit sees every input — no spatial structure is exploited.
| layer | operation | output shape |
|---|---|---|
| input | — | (N, 64) |
| Linear(64→128) | xW^T + b | (N, 128) |
| ReLU | max(0, z) | (N, 128) |
| Linear(128→10) | xW^T + b | (N, 10) |
For 64-input → 128 → 10: 64×128 + 128 + 128×10 + 10 = 9610 parameters. All learnable via autograd.
Discrimination
Sort into buckets
Sort these by operation, from memory, without looking back at MLP: the baseline deep architecture. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Concept
A Conv2d(1, 8, kernel_size=3, padding=1) on a 1×8×8 input slides an 8 × (1×3×3) filter bank across the image, outputting 8×8×8 feature maps.
| block | output shape | params |
|---|---|---|
| input | 1×8×8 | — |
| Conv2d(1,8,k=3,p=1) | 8×8×8 | 80 |
| ReLU | 8×8×8 | — |
| MaxPool2d(2) | 8×4×4 | — |
| Linear(128→32) | (N,32) | 4128 |
| Linear(32→10) | (N,10) | 330 |
Total CNN params: 80 + 4128 + 330 = 4538 — 2× fewer than MLP, yet the conv layer is translation-equivariant (a digit at any position fires the same filter).
Comparison
Comparison matrix
From CNN: exploit spatial structure: refill the params column from what you know. The rest of the table is as it appeared.
| block | output shape | params |
|---|---|---|
| input | 1×8×8 | — |
| Conv2d(1,8,k=3,p=1) | 8×8×8 | 80 |
| ReLU | 8×8×8 | — |
| MaxPool2d(2) | 8×4×4 | — |
| Linear(128→32) | (N,32) | 4128 |
| Linear(32→10) | (N,10) | 330 |
Concept
A residual block outputs relu(F(x) + x) where x is the skip. The gradient of the sum is dF/dx + 1 — the +1 ensures a gradient highway past the block.
\[ \frac{\partial\, \text{loss}}{\partial x} = \frac{\partial\, \text{loss}}{\partial y}\left(\frac{\partial F(x)}{\partial x} + 1\right) \]
Even if dF/dx → 0 (saturated layers), the gradient remains at least dL/dy — a direct path back to the input. This is why ResNets can be 50–152+ layers deep.
Counterexample
Discussion prompt
A residual block outputs relu(F(x) + x) where x is the skip. The gradient of the sum is dF/dx + 1 — the +1 ensures a gradient highway past the block.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Even if dF/dx → 0 (saturated layers), the gradient remains at least dL/dy — a direct path back to the input. This is why ResNets can be 50–152+ layers deep.
Section
Part 2 of 5
Concept
| activation | range | peak |d/dx| | dead neurons? |
|---|---|---|---|
| sigmoid | (0, 1) | 0.25 at x=0 | yes (large |x|) |
| tanh | (-1, 1) | 1.0 at x=0 | yes (large |x|) |
| ReLU | [0, ∞) | 1 for x>0 | yes (x<0 permanently) |
| Leaky ReLU | (-∞, ∞) | 1 or α for x<0 | no |
Verified: sigmoid([-2,-1,0,1,2]) = [0.1192, 0.2689, 0.5, 0.7311, 0.8808]. The peak derivative at x=0 is 0.25 — each sigmoid layer multiplies the gradient by at most 0.25.
Trade off
Comparison matrix
From Activation functions: output & gradient ranges: every row here is a choice with a cost. Fill the range column, then say which row you would actually pick and what you give up for it.
| activation | range | peak |d/dx| | dead neurons? |
|---|---|---|---|
| sigmoid | (0, 1) | 0.25 at x=0 | yes (large |x|) |
| tanh | (-1, 1) | 1.0 at x=0 | yes (large |x|) |
| ReLU | [0, ∞) | 1 for x>0 | yes (x<0 permanently) |
| Leaky ReLU | (-∞, ∞) | 1 or α for x<0 | no |
Concept
In a chain of sigmoid layers, each backward pass multiplies the gradient by σ'(z) ≤ 0.25. After L layers the gradient shrinks by 0.25^L.
| depth L | gradient scale (≤ 0.25^L) | practical effect |
|---|---|---|
| 1 | 2.50 × 10⁻¹ | fine |
| 4 | 3.91 × 10⁻³ | weak signal |
| 8 | 1.53 × 10⁻⁵ | near zero |
| 16 | 2.33 × 10⁻¹⁰ | effectively zero |
ReLU has derivative 1 for positive activations — the gradient doesn't shrink. This is why deep nets use ReLU (or its variants), not sigmoid, in hidden layers.
Analogy
Discussion prompt
Explain Vanishing gradient: the math by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
In a chain of sigmoid layers, each backward pass multiplies the gradient by σ'(z) ≤ 0.25. After L layers the gradient shrinks by 0.25^L.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Sigmoid is a 'proper' probability output, so it should work well everywhere — use it for all hidden layers in a 10-layer network.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Each sigmoid multiplies the backward gradient by ≤0.25.
Use ReLU (or LeakyReLU / GELU) in hidden layers; reserve sigmoid for the output of a binary classifier.
Why: Each sigmoid multiplies the backward gradient by ≤0.25. After 10 layers: 0.25^10 ≈ 9.5×10⁻⁷. Early layers receive essentially zero gradient and their weights stop learning.
Trap
Sigmoid is a 'proper' probability output, so it should work well everywhere — use it for all hidden layers in a 10-layer network.
Place sigmoid activations after every hidden Linear layer
Why: Each sigmoid multiplies the backward gradient by ≤0.25. After 10 layers: 0.25^10 ≈ 9.5×10⁻⁷. Early layers receive essentially zero gradient and their weights stop learning.
Use ReLU (or LeakyReLU / GELU) in hidden layers; reserve sigmoid for the output of a binary classifier.
Linear → ReLU → Linear → ReLU → ... → Linear → sigmoid (output only)
Why: ReLU's derivative is 1 for positive activations, so the gradient can flow through arbitrarily many layers without shrinking. Sigmoid/tanh belong at the output, not inside the stack.
Break the constraint
Discussion prompt
The rule this trap just fixed:
ReLU's derivative is 1 for positive activations, so the gradient can flow through arbitrarily many layers without shrinking. Sigmoid/tanh belong at the output, not inside the stack.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
Each sigmoid multiplies the backward gradient by ≤0.25. After 10 layers: 0.25^10 ≈ 9.5×10⁻⁷. Early layers receive essentially zero gradient and their weights stop learning.
Section
Part 3 of 5
Estimation
Predict first
BatchNorm standardizes a mini-batch of pre-activations to mean=0, std=1, then rescales with learnable γ, β. Input batch: [1, 3, 5, 7].
Commit before you compute: what does Batch normalization — step-by-step come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: μ = 4.0, σ = 2.2361; x̂ values verified from execution
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. BatchNorm eliminates internal covariate shift: each layer always sees inputs near (0,1), stabilizing training and allowing higher learning rates.
Worked example
BatchNorm standardizes a mini-batch of pre-activations to mean=0, std=1, then rescales with learnable γ, β. Input batch: [1, 3, 5, 7].
import torch
x = torch.tensor([[1.0],[3.0],[5.0],[7.0]]) # batch of 4
mu = x.mean() # mean over batch
var = ((x - mu)**2).mean() # population variance
xhat = (x - mu) / (var**0.5 + 1e-5)
print(f'mu={mu:.4f} var={var:.4f} std={var**0.5:.4f}')
print(xhat.squeeze().tolist())| x | x - μ | x̂ = (x-μ)/σ |
|---|---|---|
| 1.0 | -3.0 | -1.3416 |
| 3.0 | -1.0 | -0.4472 |
| 5.0 | +1.0 | +0.4472 |
| 7.0 | +3.0 | +1.3416 |
μ = 4.0, σ = 2.2361; x̂ values verified from execution
Why: BatchNorm eliminates internal covariate shift: each layer always sees inputs near (0,1), stabilizing training and allowing higher learning rates.
Pattern
Step through it
Step through Batch normalization — step-by-step one row at a time. What is driving the change, and what would the row after the last one be?
Concept
Dropout randomly zeroes each activation with probability p during training, then scales surviving activations up by 1/(1−p) so the expected value is unchanged at test time.
| p | fraction kept | scale factor | expected output |
|---|---|---|---|
| 0.3 | 70% | 1/0.7 ≈ 1.43 | 1.0 |
| 0.5 | 50% | 2.0 | 1.0 |
Verified: Dropout(p=0.5) on ones(10000) gives mean ≈ 0.9908 (≈1.0). At model.eval(), dropout is a no-op — no scaling, no masking.
Explain it
Discussion prompt
Explain Dropout: inverted scaling to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Dropout randomly zeroes each activation with probability p during training, then scales surviving activations up by 1/(1−p) so the expected value is unchanged at test time.
Concept
Poor initialization causes vanishing (weights too small) or exploding (weights too large) activations from the first forward pass — before any gradient update.
| method | formula for std | use with | fan_in=fan_out=256 → std |
|---|---|---|---|
| Xavier (Glorot) | √(2 / (fan_in + fan_out)) | sigmoid / tanh | 0.0625 |
| Kaiming (He) | √(2 / fan_in) | ReLU family | 0.0884 |
Kaiming is wider because ReLU kills half the neurons — we need larger initial weights to compensate and keep the signal magnitude stable layer-to-layer.
Sorting
Sort into buckets
These are the pieces of Lesson 77: Deep Learning Foundations, out of order. Put each one back under the part of the lesson it belongs to.
Section
Part 4 of 5
Estimation
Predict first
2-layer network. x=[1,2], w1=[[0.5,-0.3],[0.2,0.8]], w2=[[0.4,-0.1]]. Loss = z2.sum(). Trace forward then backward.
Commit before you compute: what does Backpropagation: chain-rule trace come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Row 0 of dL/dw1 is [0,0] because z1[0]=-0.1 < 0: ReLU killed it, its gradient is 0
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Chain rule: dL/dw1[i,j] = dL/dz2 · w2[i] · relu'(z1[i]) · x[j].
Worked example
2-layer network. x=[1,2], w1=[[0.5,-0.3],[0.2,0.8]], w2=[[0.4,-0.1]]. Loss = z2.sum(). Trace forward then backward.
import torch
x = torch.tensor([1.0, 2.0])
w1 = torch.tensor([[0.5,-0.3],[0.2,0.8]], requires_grad=True)
w2 = torch.tensor([[0.4,-0.1]], requires_grad=True)
z1 = w1 @ x # [-0.1, 1.8]
a1 = torch.relu(z1) # [ 0.0, 1.8]
z2 = w2 @ a1 # [-0.18]
loss = z2.sum()
loss.backward()
print('dL/dw2:', w2.grad.tolist())
print('dL/dw1:', w1.grad.tolist())| quantity | value | note |
|---|---|---|
| z1 = w1@x | [-0.1, 1.8] | pre-activation |
| a1 = relu(z1) | [0.0, 1.8] | z1[0]<0 → killed |
| z2 = w2@a1 | [-0.18] | output |
| dL/dw2 | [0.0, 1.8] | = a1 (z1[0] dead) |
| dL/dw1[1,:] | [-0.1, -0.2] | w2[1]relu_gradx |
Row 0 of dL/dw1 is [0,0] because z1[0]=-0.1 < 0: ReLU killed it, its gradient is 0
Why: Chain rule: dL/dw1[i,j] = dL/dz2 · w2[i] · relu'(z1[i]) · x[j]. When relu'=0, the whole chain is zero — dead neuron blocks all gradient to that row.
Comparison
Comparison matrix
From Backpropagation: chain-rule trace: refill the note column from what you know. The rest of the table is as it appeared.
| quantity | value | note |
|---|---|---|
| z1 = w1@x | [-0.1, 1.8] | pre-activation |
| a1 = relu(z1) | [0.0, 1.8] | z1[0]<0 → killed |
| z2 = w2@a1 | [-0.18] | output |
| dL/dw2 | [0.0, 1.8] | = a1 (z1[0] dead) |
| dL/dw1[1,:] | [-0.1, -0.2] | w2[1]relu_gradx |
Concept
| optimizer | update rule (sketch) | extra cost | typical use |
|---|---|---|---|
| SGD | w -= lr·g | none | simple; needs tuned lr |
| Momentum | v = β·v + g; w -= lr·v | +v buffer | smoother; overshoots |
| Adam | uses m̂ₜ/√v̂ₜ | +m,v buffers | adaptive lr; default |
| AdamW | Adam + weight decay separate | +m,v buffers | prevents L2 coupling |
On f(w)=w² from w=5: SGD (lr=0.1) reaches 0.54 after 10 steps; Momentum (β=0.9) reaches 0.022 (overshoots then converges); Adam (lr=0.5) reaches 0.38 — verified by execution.
Discrimination
Sort into buckets
Sort these by extra cost, from memory, without looking back at Optimizer review: SGD → Adam → AdamW. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Concept
A fixed lr is rarely optimal. LR schedulers decay the rate according to a rule, allowing aggressive early learning and fine-grained convergence later.
| scheduler | rule | lr at epoch 1,6,11 (start=0.1) |
|---|---|---|
| StepLR(step=5, γ=0.5) | multiply by γ every step_size epochs | 0.10000, 0.05000, 0.02500 |
| CosineAnnealingLR(T=10) | cosine half-cycle | 0.10000, 0.07955, 0.02061 |
| ReduceLROnPlateau | divide by γ if val loss stalls | dynamic, needs val loop |
Cosine annealing (T_max=10) values verified: [0.1, 0.0976, 0.0905, 0.0794, 0.0655, 0.05, 0.0345, 0.0206, 0.0095, 0.0024].
Analogy
Discussion prompt
Explain Learning rate scheduling by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Cosine annealing (T_max=10) values verified: [0.1, 0.0976, 0.0905, 0.0794, 0.0655, 0.05, 0.0345, 0.0206, 0.0095, 0.0024].
Anomaly
Predict first
A student writes this, and it looks reasonable:
Each iteration: sched.step() first to set the rate, then opt.step() to update weights with that rate.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: PyTorch's LR schedulers expect to be called AFTER the optimizer step.
Call sched.step() at the end of each epoch, after opt.step().
Why: PyTorch's LR schedulers expect to be called AFTER the optimizer step. Calling sched first shifts the schedule by one epoch, meaning epoch 1 uses the epoch-0 rate, epoch 2 uses epoch-1, etc. — subtle and silent.
Trap
Each iteration: sched.step() first to set the rate, then opt.step() to update weights with that rate.
sched.step() → opt.zero_grad() → loss.backward() → opt.step()
Why: PyTorch's LR schedulers expect to be called AFTER the optimizer step. Calling sched first shifts the schedule by one epoch, meaning epoch 1 uses the epoch-0 rate, epoch 2 uses epoch-1, etc. — subtle and silent.
Call sched.step() at the end of each epoch, after opt.step().
opt.zero_grad() → forward → loss → backward → opt.step() → sched.step()
Why: opt.step() updates parameters with the current lr; sched.step() then updates lr for the NEXT epoch. The standard PyTorch training loop always ends the epoch with sched.step().
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
Linear → activation blocks. Every unit sees every input — no spatial structure is exploited.; A Conv2d(1, 8, kernel_size=3, padding=1) on a 1×8×8 input slides an 8 × (1×3×3) filter bank across the image, outputting 8×8×8 feature maps.; A residual block outputs relu(F(x) + x) where x is the skip. The gradient of the sum is dF/dx + 1 — the +1 ensures a gradient highway past the block.sched.step() first to set the rate, then opt.step() to update weights with that rate.Section
Part 5 of 5
Estimation
Predict first
Wrap raw data in a Dataset subclass, then a DataLoader handles batching, shuffling, and (optionally) multi-process prefetch.
Commit before you compute: what does Dataset + DataLoader come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: 1797 samples ÷ batch 64 = 28 full batches + 1 batch of 5
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. DataLoader automatically handles the partial last batch; set drop_last=True if your model requires a fixed batch size (e.g.
Worked example
Wrap raw data in a Dataset subclass, then a DataLoader handles batching, shuffling, and (optionally) multi-process prefetch.
from torch.utils.data import Dataset, DataLoader
from sklearn.datasets import load_digits
import torch
class DigitsDS(Dataset):
def __init__(self):
d = load_digits()
self.X = torch.tensor(d.data/16.0, dtype=torch.float32)
self.y = torch.tensor(d.target, dtype=torch.long)
def __len__(self): return len(self.X)
def __getitem__(self, i): return self.X[i], self.y[i]
ds = DigitsDS()
dl = DataLoader(ds, batch_size=64, shuffle=True)
print(len(ds), len(dl)) # 1797 29| batch # | X.shape | y.shape | note |
|---|---|---|---|
| 1–28 | (64, 64) | (64,) | full batch |
| 29 (last) | (5, 64) | (5,) | 1797 mod 64 = 5 |
1797 samples ÷ batch 64 = 28 full batches + 1 batch of 5
Why: DataLoader automatically handles the partial last batch; set drop_last=True if your model requires a fixed batch size (e.g. BatchNorm1d crashes on batch_size=1).
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
1797 samples ÷ batch 64 = 28 full batches + 1 batch of 5
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Wrap raw data in a Dataset subclass, then a DataLoader handles batching, shuffling, and (optionally) multi-process prefetch.
Constraint
Discussion prompt
Run The complete deep-learning pipeline with this step confiscated:
Regularize: BatchNorm (stable activations) + Dropout (generalization)
Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.
Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.
Answer:
Dataset.__getitem__ + DataLoader(shuffle=True, batch_size=...)zero_grad → forward → loss → backward → clip_grad? → opt.step() → sched.step()Pattern
Dataset.__getitem__ + DataLoader(shuffle=True, batch_size=...)zero_grad → forward → loss → backward → clip_grad? → opt.step() → sched.step()Edge cases
Discussion prompt
The complete deep-learning pipeline works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
Dataset.__getitem__ + DataLoader(shuffle=True, batch_size=...)zero_grad → forward → loss → backward → clip_grad? → opt.step() → sched.step()Elimination
Eliminate the wrong options
A 16-layer network trained with sigmoid hidden activations shows near-zero weight updates in the first few layers. Which change MOST directly fixes this?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Sigmoid's max derivative is 0.25, so 16 layers multiply the gradient by ≤0.25^16 ≈ 2.3×10⁻¹⁰. ReLU has derivative 1 for positive activations, preventing this multiplicative decay. A higher lr or more units doesn't fix the root cause.
Check
Think before clicking.
Check your understanding
A 16-layer network trained with sigmoid hidden activations shows near-zero weight updates in the first few layers. Which change MOST directly fixes this?
Answer: A
Why: Sigmoid's max derivative is 0.25, so 16 layers multiply the gradient by ≤0.25^16 ≈ 2.3×10⁻¹⁰. ReLU has derivative 1 for positive activations, preventing this multiplicative decay. A higher lr or more units doesn't fix the root cause.
Prediction
Predict first
During training, BatchNorm normalizes each mini-batch using the batch's mean and variance. At inference (model.eval()), what does BatchNorm use instead?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Running (exponential moving) mean and variance accumulated during training
Why: PyTorch's BatchNorm tracks running_mean and running_var via momentum during training. At eval() it switches to these fixed statistics so the output is deterministic for a single sample — you don't need a batch.
Check
What changes at test time?
Check your understanding
During training, BatchNorm normalizes each mini-batch using the batch's mean and variance. At inference (model.eval()), what does BatchNorm use instead?
Answer: A
Why: PyTorch's BatchNorm tracks running_mean and running_var via momentum during training. At eval() it switches to these fixed statistics so the output is deterministic for a single sample — you don't need a batch.
Elimination
Eliminate the wrong options
After calling loss.backward(), the next two calls in the PyTorch training loop should be, in order:
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: backward() fills the .grad tensors; opt.step() reads those grads and updates parameters; then zero_grad() clears them so the NEXT iteration starts clean. sched.step() goes at the end of the epoch, not each batch.
Check
Recall the canonical loop (Lesson 40).
Check your understanding
After calling loss.backward(), the next two calls in the PyTorch training loop should be, in order:
Answer: A
Why: backward() fills the .grad tensors; opt.step() reads those grads and updates parameters; then zero_grad() clears them so the NEXT iteration starts clean. sched.step() goes at the end of the epoch, not each batch.
Section
Project
Concept
Build a complete digits classifier applying every technique from today: BatchNorm, Dropout, Kaiming init, Adam optimizer, StepLR schedule, and a DataLoader mini-batch loop.
| # | requirement | tool |
|---|---|---|
| 1 | Dataset + DataLoader (batch_size=64) | DigitsDS, DataLoader |
| 2 | MLP with BatchNorm + Dropout(0.3) | nn.Sequential |
| 3 | Training loop: zero_grad→forward→loss→backward→step→sched | Adam + StepLR |
| 4 | Measure accuracy at epoch 1, 10, 20, 50 | argmax + mean |
Expected: loss ≈ 2.31 at epoch 1, below 1.0 by epoch 50, accuracy above 87% on training set.
Counterexample
Discussion prompt
Build a complete digits classifier applying every technique from today: BatchNorm, Dropout, Kaiming init, Adam optimizer, StepLR schedule, and a DataLoader mini-batch loop.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Worked example
Your turn: wrap load_digits in a Dataset and confirm 29 batches of size 64 (last = 5).
Hint: __len__ returns len(self.X); __getitem__ returns (self.X[i], self.y[i]); DataLoader(ds, batch_size=64, shuffle=True).
from torch.utils.data import Dataset, DataLoader
from sklearn.datasets import load_digits
import torch
class DigitsDS(Dataset):
def __init__(self):
d = load_digits()
self.X = torch.tensor(d.data/16.0, dtype=torch.float32)
self.y = torch.tensor(d.target, dtype=torch.long)
def __len__(self): return len(self.X)
def __getitem__(self, i): return self.X[i], self.y[i]
ds = DigitsDS()
dl = DataLoader(ds, batch_size=64, shuffle=True)
print(len(ds), len(dl)) # 1797 29| stat | value |
|---|---|
| total samples | 1797 |
| batches | 29 |
| last batch size | 5 (1797 mod 64 = 5) |
Worked example
Your turn: define an MLP with BatchNorm and Dropout, train for 50 epochs with Adam + StepLR, and print loss at epochs 1 and 50.
Hint: nn.Sequential(Linear(64,128), BatchNorm1d(128), ReLU(), Dropout(0.3), Linear(128,64), ReLU(), Linear(64,10)).
import torch.nn as nn, torch.optim as optim
torch.manual_seed(0)
model = nn.Sequential(
nn.Linear(64,128), nn.BatchNorm1d(128), nn.ReLU(), nn.Dropout(0.3),
nn.Linear(128,64), nn.ReLU(), nn.Linear(64,10))
opt = optim.Adam(model.parameters(), lr=1e-3)
sched = optim.lr_scheduler.StepLR(opt, step_size=10, gamma=0.5)
lossfn = nn.CrossEntropyLoss()
for ep in range(1, 51):
model.train()
for xb, yb in dl:
opt.zero_grad()
lossfn(model(xb), yb).backward()
opt.step()
sched.step()
model.eval()
with torch.no_grad():
Xall = torch.tensor(load_digits().data/16.0, dtype=torch.float32)
yall = torch.tensor(load_digits().target, dtype=torch.long)
acc = (model(Xall).argmax(1)==yall).float().mean().item()
print(f'acc={acc:.4f}')| checkpoint | expected loss | expected acc |
|---|---|---|
| epoch 1 | ~2.31 | ~0.09 |
| epoch 10 | ~2.21 | ~0.46 |
| epoch 20 | ~2.04 | ~0.68 |
| epoch 50 | ~0.89 | ~0.88 |
Trade off
Comparison matrix
From Milestone 2 — MLP + training loop: every row here is a choice with a cost. Fill the expected loss column, then say which row you would actually pick and what you give up for it.
| checkpoint | expected loss | expected acc |
|---|---|---|
| epoch 1 | ~2.31 | ~0.09 |
| epoch 10 | ~2.21 | ~0.46 |
| epoch 20 | ~2.04 | ~0.68 |
| epoch 50 | ~0.89 | ~0.88 |
Concept
from sklearn.datasets import load_digits
import torch, torch.nn as nn, torch.optim as optim
from torch.utils.data import Dataset, DataLoader
class DigitsDS(Dataset):
def __init__(self):
d = load_digits()
self.X = torch.tensor(d.data/16.0, dtype=torch.float32)
self.y = torch.tensor(d.target, dtype=torch.long)
def __len__(self): return len(self.X)
def __getitem__(self, i): return self.X[i], self.y[i]
torch.manual_seed(0)
ds = DigitsDS()
dl = DataLoader(ds, batch_size=64, shuffle=True)
model = nn.Sequential(
nn.Linear(64,128), nn.BatchNorm1d(128), nn.ReLU(), nn.Dropout(0.3),
nn.Linear(128,64), nn.ReLU(), nn.Linear(64,10))
opt = optim.Adam(model.parameters(), lr=1e-3)
sched = optim.lr_scheduler.StepLR(opt, step_size=10, gamma=0.5)
lossfn = nn.CrossEntropyLoss()
for ep in range(1, 51):
model.train()
for xb, yb in dl:
opt.zero_grad()
lossfn(model(xb), yb).backward()
opt.step()
sched.step()
model.eval()
Xa = torch.tensor(load_digits().data/16.0, dtype=torch.float32)
ya = torch.tensor(load_digits().target, dtype=torch.long)
with torch.no_grad():
acc = (model(Xa).argmax(1)==ya).float().mean().item()
print(f'acc={acc:.4f}') # ~0.88| component | role |
|---|---|
| DigitsDS | Dataset subclass — wraps numpy arrays |
| DataLoader | mini-batch iterator, shuffle each epoch |
| BatchNorm1d | stabilize activations per batch |
| Dropout(0.3) | regularize; disabled at eval() |
| Adam + StepLR | adaptive lr, halved every 10 epochs |
| sched.step() | called after opt.step(), end of epoch |
If your accuracy lands near 0.88 on the training set after 50 epochs — every technique from today is working correctly.
Comparison
Comparison matrix
From The full program: refill the role column from what you know. The rest of the table is as it appeared.
| component | role |
|---|---|
| DigitsDS | Dataset subclass — wraps numpy arrays |
| DataLoader | mini-batch iterator, shuffle each epoch |
| BatchNorm1d | stabilize activations per batch |
| Dropout(0.3) | regularize; disabled at eval() |
| Adam + StepLR | adaptive lr, halved every 10 epochs |
| sched.step() | called after opt.step(), end of epoch |
Concept
Slides closed: explain (1) why a 16-layer sigmoid net fails and two architectural fixes, (2) the exact sequence of the training loop including where sched.step() goes, and (3) what BatchNorm computes at training vs inference.
Stretch (homework): implement gradient clipping (clip_grad_norm_(model.parameters(), 1.0)) before opt.step(); swap in a CNN (Conv2d on 1×8×8 inputs); derive on paper that 0.25^16 ≈ 2.3×10⁻¹⁰. Next up: attention mechanisms (Vaswani et al. 2017 — read the abstract and architecture diagram before Lesson 78).
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Architecture: MLP → CNN → ResNet · Activations & Vanishing Gradients · BatchNorm, Dropout & Initialization · Backprop & Optimizers · Data Pipeline & Pattern · Your turn: full digits classifier. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
≤ 0.25^L — and why ReLU fixes itzero_grad → forward → loss → backward → opt.step() → sched.step()| concept | the one thing to remember |
|---|---|
| vanishing gradient | sigmoid kills it; ReLU or skip connections fix it |
| BatchNorm | batch stats at train; running stats at eval |
| Dropout | zero with prob p, scale by 1/(1-p); off at eval |
| Adam vs SGD | Adam adapts lr per-param; SGD needs tuned global lr |
| sched.step() | end of epoch, AFTER opt.step() |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.