Lesson 77: Deep Learning Foundations

USAAIO Lesson 77, from Phase 4 on deep learning. It covers the MLP, CNN, and ResNet architectures; the ReLU, sigmoid, and tanh activations and the mechanics of the vanishing gradient; BatchNorm, dropout, and Xavier and Kaiming initialization; the PyTorch training loop line by line; backpropagation through the chain rule, with a concrete two-layer trace; the SGD, Momentum, Adam, and AdamW optimizers and learning-rate scheduling; and Dataset and DataLoader for mini-batch pipelines. You build a complete digits classifier with every one of these techniques applied. The lesson runs to 32 slides.

Subject: Machine Learning · 61 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Deep Learning Foundations

Title

USAAIO · Lesson 77 · Phase 4

MLP → CNN → ResNet. Activations, normalization, dropout, initialization, backprop, optimizers, and the full PyTorch pipeline — every concept Barron needs to reason about any deep model on the exam.

2. By the end of this lesson you can

Objectives

  1. Explain how MLP → CNN → ResNet extend each other and why skip connections fix vanishing gradients
  2. Trace backpropagation through a 2-layer network using the chain rule (with real numbers)
  3. Predict activation function behavior and its effect on gradient magnitude across layers
  4. Apply BatchNorm, Dropout, Xavier/Kaiming init and explain what breaks without each
  5. Implement the complete PyTorch pipeline: Dataset → DataLoader → training loop → optimizer → LR schedule

3. What survived from Algorithm Review & Selection?

Warm-up

Discussion prompt

Before we open Lesson 77: Deep Learning Foundations: without looking back, what was the main idea of Algorithm Review & Selection, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

flash-review of 7 core classifiers (logistic regression, SVM, decision trees, random forest, GBM, kNN, Naive Bayes) — loss functions, gradients, hyperparameters, and bias-variance profiles; whiteboard-style derivations; USAAIO exam question patterns; and an algorithm-selection framework. Implement and benchmark all 7 on a synthetic dataset.

4. Architecture: MLP → CNN → ResNet

Section

Part 1 of 5

5. MLP: the baseline deep architecture

Concept

A multi-layer perceptron (Lesson 40) stacks Linear → activation blocks. Every unit sees every input — no spatial structure is exploited.

layeroperationoutput shape
input—(N, 64)
Linear(64→128)xW^T + b(N, 128)
ReLUmax(0, z)(N, 128)
Linear(128→10)xW^T + b(N, 10)

For 64-input → 128 → 10: 64×128 + 128 + 128×10 + 10 = 9610 parameters. All learnable via autograd.

6. Which is which, by operation

Discrimination

Sort into buckets

Sort these by operation, from memory, without looking back at MLP: the baseline deep architecture. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

—
input
xW^T + b
Linear(64→128); Linear(128→10)
max(0, z)
ReLU
g1
operation is "—" for input — that is what the table on "MLP: the baseline deep architecture" records, and it is the single property separating this group from the rest.
g2
operation is "xW^T + b" for Linear(64→128), Linear(128→10) — that is what the table on "MLP: the baseline deep architecture" records, and it is the single property separating this group from the rest.
g3
operation is "max(0, z)" for ReLU — that is what the table on "MLP: the baseline deep architecture" records, and it is the single property separating this group from the rest.

7. CNN: exploit spatial structure

Concept

A Conv2d(1, 8, kernel_size=3, padding=1) on a 1×8×8 input slides an 8 × (1×3×3) filter bank across the image, outputting 8×8×8 feature maps.

blockoutput shapeparams
input1×8×8—
Conv2d(1,8,k=3,p=1)8×8×880
ReLU8×8×8—
MaxPool2d(2)8×4×4—
Linear(128→32)(N,32)4128
Linear(32→10)(N,10)330

Total CNN params: 80 + 4128 + 330 = 4538 — 2× fewer than MLP, yet the conv layer is translation-equivariant (a digit at any position fires the same filter).

8. Fill in: params for CNN: exploit spatial structure

Comparison

Comparison matrix

From CNN: exploit spatial structure: refill the params column from what you know. The rest of the table is as it appeared.

blockoutput shapeparams
input1×8×8—
Conv2d(1,8,k=3,p=1)8×8×880
ReLU8×8×8—
MaxPool2d(2)8×4×4—
Linear(128→32)(N,32)4128
Linear(32→10)(N,10)330

9. ResNet: skip connections defeat vanishing gradients

Concept

A residual block outputs relu(F(x) + x) where x is the skip. The gradient of the sum is dF/dx + 1 — the +1 ensures a gradient highway past the block.

\[ \frac{\partial\, \text{loss}}{\partial x} = \frac{\partial\, \text{loss}}{\partial y}\left(\frac{\partial F(x)}{\partial x} + 1\right) \]

Even if dF/dx → 0 (saturated layers), the gradient remains at least dL/dy — a direct path back to the input. This is why ResNets can be 50–152+ layers deep.

10. Break it if you can: ResNet: skip connections defeat vanishing…

Counterexample

Discussion prompt

A residual block outputs relu(F(x) + x) where x is the skip. The gradient of the sum is dF/dx + 1 — the +1 ensures a gradient highway past the block.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Even if dF/dx → 0 (saturated layers), the gradient remains at least dL/dy — a direct path back to the input. This is why ResNets can be 50–152+ layers deep.

11. Activations & Vanishing Gradients

Section

Part 2 of 5

12. Activation functions: output & gradient ranges

Concept

activationrangepeak |d/dx|dead neurons?
sigmoid(0, 1)0.25 at x=0yes (large |x|)
tanh(-1, 1)1.0 at x=0yes (large |x|)
ReLU[0, ∞)1 for x>0yes (x<0 permanently)
Leaky ReLU(-∞, ∞)1 or α for x<0no

Verified: sigmoid([-2,-1,0,1,2]) = [0.1192, 0.2689, 0.5, 0.7311, 0.8808]. The peak derivative at x=0 is 0.25 — each sigmoid layer multiplies the gradient by at most 0.25.

13. What each one costs: Activation functions: output & gradient ranges

Trade off

Comparison matrix

From Activation functions: output & gradient ranges: every row here is a choice with a cost. Fill the range column, then say which row you would actually pick and what you give up for it.

activationrangepeak |d/dx|dead neurons?
sigmoid(0, 1)0.25 at x=0yes (large |x|)
tanh(-1, 1)1.0 at x=0yes (large |x|)
ReLU[0, ∞)1 for x>0yes (x<0 permanently)
Leaky ReLU(-∞, ∞)1 or α for x<0no

14. Vanishing gradient: the math

Concept

In a chain of sigmoid layers, each backward pass multiplies the gradient by σ'(z) ≤ 0.25. After L layers the gradient shrinks by 0.25^L.

depth Lgradient scale (≤ 0.25^L)practical effect
12.50 × 10⁻¹fine
43.91 × 10⁻³weak signal
81.53 × 10⁻⁵near zero
162.33 × 10⁻¹⁰effectively zero

ReLU has derivative 1 for positive activations — the gradient doesn't shrink. This is why deep nets use ReLU (or its variants), not sigmoid, in hidden layers.

15. By analogy: Vanishing gradient: the math

Analogy

Discussion prompt

Explain Vanishing gradient: the math by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

In a chain of sigmoid layers, each backward pass multiplies the gradient by σ'(z) ≤ 0.25. After L layers the gradient shrinks by 0.25^L.

16. Something is wrong here: using sigmoid in deep hidden layers

Anomaly

Predict first

A student writes this, and it looks reasonable:

Sigmoid is a 'proper' probability output, so it should work well everywhere — use it for all hidden layers in a 10-layer network.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Each sigmoid multiplies the backward gradient by ≤0.25.

Use ReLU (or LeakyReLU / GELU) in hidden layers; reserve sigmoid for the output of a binary classifier.

Why: Each sigmoid multiplies the backward gradient by ≤0.25. After 10 layers: 0.25^10 ≈ 9.5×10⁻⁷. Early layers receive essentially zero gradient and their weights stop learning.

17. Trap: using sigmoid in deep hidden layers

Trap

The trap

Sigmoid is a 'proper' probability output, so it should work well everywhere — use it for all hidden layers in a 10-layer network.

Place sigmoid activations after every hidden Linear layer

Why: Each sigmoid multiplies the backward gradient by ≤0.25. After 10 layers: 0.25^10 ≈ 9.5×10⁻⁷. Early layers receive essentially zero gradient and their weights stop learning.

The fix

Use ReLU (or LeakyReLU / GELU) in hidden layers; reserve sigmoid for the output of a binary classifier.

Linear → ReLU → Linear → ReLU → ... → Linear → sigmoid (output only)

Why: ReLU's derivative is 1 for positive activations, so the gradient can flow through arbitrarily many layers without shrinking. Sigmoid/tanh belong at the output, not inside the stack.

18. Break it on purpose: using sigmoid in deep hidden layers

Break the constraint

Discussion prompt

The rule this trap just fixed:

ReLU's derivative is 1 for positive activations, so the gradient can flow through arbitrarily many layers without shrinking. Sigmoid/tanh belong at the output, not inside the stack.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

Each sigmoid multiplies the backward gradient by ≤0.25. After 10 layers: 0.25^10 ≈ 9.5×10⁻⁷. Early layers receive essentially zero gradient and their weights stop learning.

19. BatchNorm, Dropout & Initialization

Section

Part 3 of 5

20. Guess the shape of the answer: Batch normalization — step-by-step

Estimation

Predict first

BatchNorm standardizes a mini-batch of pre-activations to mean=0, std=1, then rescales with learnable γ, β. Input batch: [1, 3, 5, 7].

Commit before you compute: what does Batch normalization — step-by-step come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: μ = 4.0, σ = 2.2361; x̂ values verified from execution

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. BatchNorm eliminates internal covariate shift: each layer always sees inputs near (0,1), stabilizing training and allowing higher learning rates.

21. Batch normalization — step-by-step

Worked example

BatchNorm standardizes a mini-batch of pre-activations to mean=0, std=1, then rescales with learnable γ, β. Input batch: [1, 3, 5, 7].

import torch
x = torch.tensor([[1.0],[3.0],[5.0],[7.0]])  # batch of 4
mu  = x.mean()            # mean over batch
var = ((x - mu)**2).mean() # population variance
xhat = (x - mu) / (var**0.5 + 1e-5)
print(f'mu={mu:.4f}  var={var:.4f}  std={var**0.5:.4f}')
print(xhat.squeeze().tolist())
xx - μx̂ = (x-μ)/σ
1.0-3.0-1.3416
3.0-1.0-0.4472
5.0+1.0+0.4472
7.0+3.0+1.3416

μ = 4.0, σ = 2.2361; x̂ values verified from execution

Why: BatchNorm eliminates internal covariate shift: each layer always sees inputs near (0,1), stabilizing training and allowing higher learning rates.

22. Watch it run: Batch normalization — step-by-step

Pattern

Step through it

Step through Batch normalization — step-by-step one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: x is 1.0
  2. Step 2: x is 3.0
  3. Step 3: x is 5.0
  4. Step 4: x is 7.0

23. Dropout: inverted scaling

Concept

Dropout randomly zeroes each activation with probability p during training, then scales surviving activations up by 1/(1−p) so the expected value is unchanged at test time.

pfraction keptscale factorexpected output
0.370%1/0.7 ≈ 1.431.0
0.550%2.01.0

Verified: Dropout(p=0.5) on ones(10000) gives mean ≈ 0.9908 (≈1.0). At model.eval(), dropout is a no-op — no scaling, no masking.

24. Teach it back: Dropout: inverted scaling

Explain it

Discussion prompt

Explain Dropout: inverted scaling to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Dropout randomly zeroes each activation with probability p during training, then scales surviving activations up by 1/(1−p) so the expected value is unchanged at test time.

25. Weight initialization: Xavier vs Kaiming

Concept

Poor initialization causes vanishing (weights too small) or exploding (weights too large) activations from the first forward pass — before any gradient update.

methodformula for stduse withfan_in=fan_out=256 → std
Xavier (Glorot)√(2 / (fan_in + fan_out))sigmoid / tanh0.0625
Kaiming (He)√(2 / fan_in)ReLU family0.0884

Kaiming is wider because ReLU kills half the neurons — we need larger initial weights to compensate and keep the signal magnitude stable layer-to-layer.

26. Where does each piece belong: Lesson 77: Deep Learning Foundations

Sorting

Sort into buckets

These are the pieces of Lesson 77: Deep Learning Foundations, out of order. Put each one back under the part of the lesson it belongs to.

Architecture: MLP → CNN → ResNet
MLP: the baseline deep architecture; CNN: exploit spatial structure; ResNet: skip connections defeat vanishing gradients
Activations & Vanishing Gradients
Activation functions: output & gradient ranges; Vanishing gradient: the math
BatchNorm, Dropout & Initialization
Batch normalization — step-by-step; Dropout: inverted scaling; Weight initialization: Xavier vs Kaiming
s1
Architecture: MLP → CNN → ResNet is where Lesson 77: Deep Learning Foundations puts MLP: the baseline deep architecture, CNN: exploit spatial structure, ResNet: skip connections defeat vanishing gradients. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
Activations & Vanishing Gradients is where Lesson 77: Deep Learning Foundations puts Activation functions: output & gradient ranges, Vanishing gradient: the math. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
BatchNorm, Dropout & Initialization is where Lesson 77: Deep Learning Foundations puts Batch normalization — step-by-step, Dropout: inverted scaling, Weight initialization: Xavier vs Kaiming. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

27. Backprop & Optimizers

Section

Part 4 of 5

28. Guess the shape of the answer: Backpropagation: chain-rule trace

Estimation

Predict first

2-layer network. x=[1,2], w1=[[0.5,-0.3],[0.2,0.8]], w2=[[0.4,-0.1]]. Loss = z2.sum(). Trace forward then backward.

Commit before you compute: what does Backpropagation: chain-rule trace come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Row 0 of dL/dw1 is [0,0] because z1[0]=-0.1 < 0: ReLU killed it, its gradient is 0

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Chain rule: dL/dw1[i,j] = dL/dz2 · w2[i] · relu'(z1[i]) · x[j].

29. Backpropagation: chain-rule trace

Worked example

2-layer network. x=[1,2], w1=[[0.5,-0.3],[0.2,0.8]], w2=[[0.4,-0.1]]. Loss = z2.sum(). Trace forward then backward.

import torch
x  = torch.tensor([1.0, 2.0])
w1 = torch.tensor([[0.5,-0.3],[0.2,0.8]], requires_grad=True)
w2 = torch.tensor([[0.4,-0.1]], requires_grad=True)
z1 = w1 @ x            # [-0.1, 1.8]
a1 = torch.relu(z1)    # [ 0.0, 1.8]
z2 = w2 @ a1           # [-0.18]
loss = z2.sum()
loss.backward()
print('dL/dw2:', w2.grad.tolist())
print('dL/dw1:', w1.grad.tolist())
quantityvaluenote
z1 = w1@x[-0.1, 1.8]pre-activation
a1 = relu(z1)[0.0, 1.8]z1[0]<0 → killed
z2 = w2@a1[-0.18]output
dL/dw2[0.0, 1.8]= a1 (z1[0] dead)
dL/dw1[1,:][-0.1, -0.2]w2[1]relu_gradx

Row 0 of dL/dw1 is [0,0] because z1[0]=-0.1 < 0: ReLU killed it, its gradient is 0

Why: Chain rule: dL/dw1[i,j] = dL/dz2 · w2[i] · relu'(z1[i]) · x[j]. When relu'=0, the whole chain is zero — dead neuron blocks all gradient to that row.

30. Fill in: note for Backpropagation: chain-rule trace

Comparison

Comparison matrix

From Backpropagation: chain-rule trace: refill the note column from what you know. The rest of the table is as it appeared.

quantityvaluenote
z1 = w1@x[-0.1, 1.8]pre-activation
a1 = relu(z1)[0.0, 1.8]z1[0]<0 → killed
z2 = w2@a1[-0.18]output
dL/dw2[0.0, 1.8]= a1 (z1[0] dead)
dL/dw1[1,:][-0.1, -0.2]w2[1]relu_gradx

31. Optimizer review: SGD → Adam → AdamW

Concept

optimizerupdate rule (sketch)extra costtypical use
SGDw -= lr·gnonesimple; needs tuned lr
Momentumv = β·v + g; w -= lr·v+v buffersmoother; overshoots
Adamuses m̂ₜ/√v̂ₜ+m,v buffersadaptive lr; default
AdamWAdam + weight decay separate+m,v buffersprevents L2 coupling

On f(w)=w² from w=5: SGD (lr=0.1) reaches 0.54 after 10 steps; Momentum (β=0.9) reaches 0.022 (overshoots then converges); Adam (lr=0.5) reaches 0.38 — verified by execution.

32. Which is which, by extra cost

Discrimination

Sort into buckets

Sort these by extra cost, from memory, without looking back at Optimizer review: SGD → Adam → AdamW. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

none
SGD
+v buffer
Momentum
+m,v buffers
Adam; AdamW
g1
extra cost is "none" for SGD — that is what the table on "Optimizer review: SGD → Adam → AdamW" records, and it is the single property separating this group from the rest.
g2
extra cost is "+v buffer" for Momentum — that is what the table on "Optimizer review: SGD → Adam → AdamW" records, and it is the single property separating this group from the rest.
g3
extra cost is "+m,v buffers" for Adam, AdamW — that is what the table on "Optimizer review: SGD → Adam → AdamW" records, and it is the single property separating this group from the rest.

33. Learning rate scheduling

Concept

A fixed lr is rarely optimal. LR schedulers decay the rate according to a rule, allowing aggressive early learning and fine-grained convergence later.

schedulerrulelr at epoch 1,6,11 (start=0.1)
StepLR(step=5, γ=0.5)multiply by γ every step_size epochs0.10000, 0.05000, 0.02500
CosineAnnealingLR(T=10)cosine half-cycle0.10000, 0.07955, 0.02061
ReduceLROnPlateaudivide by γ if val loss stallsdynamic, needs val loop

Cosine annealing (T_max=10) values verified: [0.1, 0.0976, 0.0905, 0.0794, 0.0655, 0.05, 0.0345, 0.0206, 0.0095, 0.0024].

34. By analogy: Learning rate scheduling

Analogy

Discussion prompt

Explain Learning rate scheduling by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Cosine annealing (T_max=10) values verified: [0.1, 0.0976, 0.0905, 0.0794, 0.0655, 0.05, 0.0345, 0.0206, 0.0095, 0.0024].

35. Something is wrong here: calling sched.step() before opt.step()

Anomaly

Predict first

A student writes this, and it looks reasonable:

Each iteration: sched.step() first to set the rate, then opt.step() to update weights with that rate.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: PyTorch's LR schedulers expect to be called AFTER the optimizer step.

Call sched.step() at the end of each epoch, after opt.step().

Why: PyTorch's LR schedulers expect to be called AFTER the optimizer step. Calling sched first shifts the schedule by one epoch, meaning epoch 1 uses the epoch-0 rate, epoch 2 uses epoch-1, etc. — subtle and silent.

36. Trap: calling sched.step() before opt.step()

Trap

The trap

Each iteration: sched.step() first to set the rate, then opt.step() to update weights with that rate.

sched.step() → opt.zero_grad() → loss.backward() → opt.step()

Why: PyTorch's LR schedulers expect to be called AFTER the optimizer step. Calling sched first shifts the schedule by one epoch, meaning epoch 1 uses the epoch-0 rate, epoch 2 uses epoch-1, etc. — subtle and silent.

The fix

Call sched.step() at the end of each epoch, after opt.step().

opt.zero_grad() → forward → loss → backward → opt.step() → sched.step()

Why: opt.step() updates parameters with the current lr; sched.step() then updates lr for the NEXT epoch. The standard PyTorch training loop always ends the epoch with sched.step().

37. Which of these survive contact with Lesson 77: Deep Learning Foundations?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
A multi-layer perceptron (Lesson 40) stacks Linear → activation blocks. Every unit sees every input — no spatial structure is exploited.; A Conv2d(1, 8, kernel_size=3, padding=1) on a 1×8×8 input slides an 8 × (1×3×3) filter bank across the image, outputting 8×8×8 feature maps.; A residual block outputs relu(F(x) + x) where x is the skip. The gradient of the sum is dF/dx + 1 — the +1 ensures a gradient highway past the block.
Breaks
Sigmoid is a 'proper' probability output, so it should work well everywhere — use it for all hidden layers in a 10-layer network.; Each iteration: sched.step() first to set the rate, then opt.step() to update weights with that rate.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 77: Deep Learning Foundations puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

38. Data Pipeline & Pattern

Section

Part 5 of 5

39. Guess the shape of the answer: Dataset + DataLoader

Estimation

Predict first

Wrap raw data in a Dataset subclass, then a DataLoader handles batching, shuffling, and (optionally) multi-process prefetch.

Commit before you compute: what does Dataset + DataLoader come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: 1797 samples ÷ batch 64 = 28 full batches + 1 batch of 5

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. DataLoader automatically handles the partial last batch; set drop_last=True if your model requires a fixed batch size (e.g.

40. Dataset + DataLoader

Worked example

Wrap raw data in a Dataset subclass, then a DataLoader handles batching, shuffling, and (optionally) multi-process prefetch.

from torch.utils.data import Dataset, DataLoader
from sklearn.datasets import load_digits
import torch

class DigitsDS(Dataset):
    def __init__(self):
        d = load_digits()
        self.X = torch.tensor(d.data/16.0, dtype=torch.float32)
        self.y = torch.tensor(d.target, dtype=torch.long)
    def __len__(self): return len(self.X)
    def __getitem__(self, i): return self.X[i], self.y[i]

ds = DigitsDS()
dl = DataLoader(ds, batch_size=64, shuffle=True)
print(len(ds), len(dl))   # 1797  29
batch #X.shapey.shapenote
1–28(64, 64)(64,)full batch
29 (last)(5, 64)(5,)1797 mod 64 = 5

1797 samples ÷ batch 64 = 28 full batches + 1 batch of 5

Why: DataLoader automatically handles the partial last batch; set drop_last=True if your model requires a fixed batch size (e.g. BatchNorm1d crashes on batch_size=1).

41. Work backwards from the answer: Dataset + DataLoader

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

1797 samples ÷ batch 64 = 28 full batches + 1 batch of 5

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Wrap raw data in a Dataset subclass, then a DataLoader handles batching, shuffling, and (optionally) multi-process prefetch.

42. Without one step: The complete deep-learning pipeline

Constraint

Discussion prompt

Run The complete deep-learning pipeline with this step confiscated:

Regularize: BatchNorm (stable activations) + Dropout (generalization)

Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.

Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.

Answer:

  1. Architecture: choose MLP / CNN / ResNet based on data structure (tabular / image / deep image)
  2. Activations: ReLU (or GELU) in hidden layers; sigmoid/softmax at output only
  3. Init: Kaiming for ReLU stacks; Xavier for tanh/sigmoid
  4. Regularize: BatchNorm (stable activations) + Dropout (generalization)
  5. Data: Dataset.__getitem__ + DataLoader(shuffle=True, batch_size=...)
  6. Loop: zero_grad → forward → loss → backward → clip_grad? → opt.step() → sched.step()
  7. Optimizer: Adam / AdamW default; tune lr with a scheduler

43. The complete deep-learning pipeline

Pattern

  1. Architecture: choose MLP / CNN / ResNet based on data structure (tabular / image / deep image)
  2. Activations: ReLU (or GELU) in hidden layers; sigmoid/softmax at output only
  3. Init: Kaiming for ReLU stacks; Xavier for tanh/sigmoid
  4. Regularize: BatchNorm (stable activations) + Dropout (generalization)
  5. Data: Dataset.__getitem__ + DataLoader(shuffle=True, batch_size=...)
  6. Loop: zero_grad → forward → loss → backward → clip_grad? → opt.step() → sched.step()
  7. Optimizer: Adam / AdamW default; tune lr with a scheduler

44. Where does it stop working: The complete deep-learning pipeline

Edge cases

Discussion prompt

The complete deep-learning pipeline works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Architecture: choose MLP / CNN / ResNet based on data structure (tabular / image / deep image)
  2. Activations: ReLU (or GELU) in hidden layers; sigmoid/softmax at output only
  3. Init: Kaiming for ReLU stacks; Xavier for tanh/sigmoid
  4. Regularize: BatchNorm (stable activations) + Dropout (generalization)
  5. Data: Dataset.__getitem__ + DataLoader(shuffle=True, batch_size=...)
  6. Loop: zero_grad → forward → loss → backward → clip_grad? → opt.step() → sched.step()
  7. Optimizer: Adam / AdamW default; tune lr with a scheduler

45. Rule out three: Check — vanishing gradient fix

Elimination

Eliminate the wrong options

A 16-layer network trained with sigmoid hidden activations shows near-zero weight updates in the first few layers. Which change MOST directly fixes this?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Replace sigmoid hidden activations with ReLU
  • B. Increase the learning rate by 10×
  • C. Add more hidden units to each layer
  • D. Switch from Adam to SGD

Survives elimination: A

Why: Sigmoid's max derivative is 0.25, so 16 layers multiply the gradient by ≤0.25^16 ≈ 2.3×10⁻¹⁰. ReLU has derivative 1 for positive activations, preventing this multiplicative decay. A higher lr or more units doesn't fix the root cause.

46. Check — vanishing gradient fix

Check

Think before clicking.

Check your understanding

A 16-layer network trained with sigmoid hidden activations shows near-zero weight updates in the first few layers. Which change MOST directly fixes this?

  • A. Replace sigmoid hidden activations with ReLU (correct)
  • B. Increase the learning rate by 10×
  • C. Add more hidden units to each layer
  • D. Switch from Adam to SGD

Answer: A

Why: Sigmoid's max derivative is 0.25, so 16 layers multiply the gradient by ≤0.25^16 ≈ 2.3×10⁻¹⁰. ReLU has derivative 1 for positive activations, preventing this multiplicative decay. A higher lr or more units doesn't fix the root cause.

Why B tempts people
A higher lr scales the already-near-zero gradient — it multiplies zero, not un-vanishes it. It can also cause divergence in the later layers where gradients are healthy.
Why C tempts people
More units increase capacity but don't change the per-unit gradient magnitude. The vanishing problem is about the chain-rule multiplier, not the number of neurons.
Why D tempts people
The choice of optimizer (Adam vs SGD) does not fix vanishing gradients — both multiply the same near-zero gradient by a step-size. The problem is in the activation function, not the optimizer.

47. Answer it before you see the options: Check — BatchNorm inference behavior

Prediction

Predict first

During training, BatchNorm normalizes each mini-batch using the batch's mean and variance. At inference (model.eval()), what does BatchNorm use instead?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: Running (exponential moving) mean and variance accumulated during training

Why: PyTorch's BatchNorm tracks running_mean and running_var via momentum during training. At eval() it switches to these fixed statistics so the output is deterministic for a single sample — you don't need a batch.

48. Check — BatchNorm inference behavior

Check

What changes at test time?

Check your understanding

During training, BatchNorm normalizes each mini-batch using the batch's mean and variance. At inference (model.eval()), what does BatchNorm use instead?

  • A. Running (exponential moving) mean and variance accumulated during training (correct)
  • B. Mean and variance of the current test batch
  • C. All zeros for mean and all ones for variance (identity)
  • D. The learnable parameters γ and β directly, with no normalization

Answer: A

Why: PyTorch's BatchNorm tracks running_mean and running_var via momentum during training. At eval() it switches to these fixed statistics so the output is deterministic for a single sample — you don't need a batch.

Why B tempts people
Using test-batch statistics would make inference non-deterministic for single samples (batch_size=1 gives var=0) and would require knowing the test distribution — exactly what we don't want.
Why C tempts people
Passing through unnormalized (identity) would undo all the benefits of BatchNorm and produce a distribution mismatch between training and inference.
Why D tempts people
γ and β are the post-normalization scale and shift; they're applied after the standardization step, not instead of it. Skipping normalization entirely breaks the layer.

49. Rule out three: Check — training loop order

Elimination

Eliminate the wrong options

After calling loss.backward(), the next two calls in the PyTorch training loop should be, in order:

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. opt.step(), then opt.zero_grad() for the NEXT iteration
  • B. opt.zero_grad(), then opt.step()
  • C. opt.step(), then loss.backward() again for momentum
  • D. sched.step(), then opt.step()

Survives elimination: A

Why: backward() fills the .grad tensors; opt.step() reads those grads and updates parameters; then zero_grad() clears them so the NEXT iteration starts clean. sched.step() goes at the end of the epoch, not each batch.

50. Check — training loop order

Check

Recall the canonical loop (Lesson 40).

Check your understanding

After calling loss.backward(), the next two calls in the PyTorch training loop should be, in order:

  • A. opt.step(), then opt.zero_grad() for the NEXT iteration (correct)
  • B. opt.zero_grad(), then opt.step()
  • C. opt.step(), then loss.backward() again for momentum
  • D. sched.step(), then opt.step()

Answer: A

Why: backward() fills the .grad tensors; opt.step() reads those grads and updates parameters; then zero_grad() clears them so the NEXT iteration starts clean. sched.step() goes at the end of the epoch, not each batch.

Why B tempts people
zero_grad() before opt.step() would wipe the gradients backward() just computed — the optimizer would step with all-zero gradients and parameters would not update.
Why C tempts people
Calling backward() twice without zero_grad() accumulates gradients, doubling the effective gradient. There is no 'momentum backward' — momentum lives entirely inside the optimizer.
Why D tempts people
sched.step() updates the learning rate for the next epoch, not the next batch. Calling it before opt.step() shifts the schedule by one epoch and applies the wrong rate to the current step.

51. Your turn: full digits classifier

Section

Project

52. Project: MLP with all the techniques

Concept

Build a complete digits classifier applying every technique from today: BatchNorm, Dropout, Kaiming init, Adam optimizer, StepLR schedule, and a DataLoader mini-batch loop.

#requirementtool
1Dataset + DataLoader (batch_size=64)DigitsDS, DataLoader
2MLP with BatchNorm + Dropout(0.3)nn.Sequential
3Training loop: zero_grad→forward→loss→backward→step→schedAdam + StepLR
4Measure accuracy at epoch 1, 10, 20, 50argmax + mean

Expected: loss ≈ 2.31 at epoch 1, below 1.0 by epoch 50, accuracy above 87% on training set.

53. Break it if you can: Project: MLP with all the techniques

Counterexample

Discussion prompt

Build a complete digits classifier applying every technique from today: BatchNorm, Dropout, Kaiming init, Adam optimizer, StepLR schedule, and a DataLoader mini-batch loop.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

54. Milestone 1 — Dataset and DataLoader

Worked example

Your turn: wrap load_digits in a Dataset and confirm 29 batches of size 64 (last = 5).

Hint: __len__ returns len(self.X); __getitem__ returns (self.X[i], self.y[i]); DataLoader(ds, batch_size=64, shuffle=True).

from torch.utils.data import Dataset, DataLoader
from sklearn.datasets import load_digits
import torch

class DigitsDS(Dataset):
    def __init__(self):
        d = load_digits()
        self.X = torch.tensor(d.data/16.0, dtype=torch.float32)
        self.y = torch.tensor(d.target, dtype=torch.long)
    def __len__(self): return len(self.X)
    def __getitem__(self, i): return self.X[i], self.y[i]

ds = DigitsDS()
dl = DataLoader(ds, batch_size=64, shuffle=True)
print(len(ds), len(dl))   # 1797  29
statvalue
total samples1797
batches29
last batch size5 (1797 mod 64 = 5)

55. Milestone 2 — MLP + training loop

Worked example

Your turn: define an MLP with BatchNorm and Dropout, train for 50 epochs with Adam + StepLR, and print loss at epochs 1 and 50.

Hint: nn.Sequential(Linear(64,128), BatchNorm1d(128), ReLU(), Dropout(0.3), Linear(128,64), ReLU(), Linear(64,10)).

import torch.nn as nn, torch.optim as optim
torch.manual_seed(0)
model = nn.Sequential(
    nn.Linear(64,128), nn.BatchNorm1d(128), nn.ReLU(), nn.Dropout(0.3),
    nn.Linear(128,64), nn.ReLU(), nn.Linear(64,10))
opt  = optim.Adam(model.parameters(), lr=1e-3)
sched = optim.lr_scheduler.StepLR(opt, step_size=10, gamma=0.5)
lossfn = nn.CrossEntropyLoss()
for ep in range(1, 51):
    model.train()
    for xb, yb in dl:
        opt.zero_grad()
        lossfn(model(xb), yb).backward()
        opt.step()
    sched.step()
model.eval()
with torch.no_grad():
    Xall = torch.tensor(load_digits().data/16.0, dtype=torch.float32)
    yall = torch.tensor(load_digits().target, dtype=torch.long)
    acc = (model(Xall).argmax(1)==yall).float().mean().item()
print(f'acc={acc:.4f}')
checkpointexpected lossexpected acc
epoch 1~2.31~0.09
epoch 10~2.21~0.46
epoch 20~2.04~0.68
epoch 50~0.89~0.88

56. What each one costs: Milestone 2 — MLP + training loop

Trade off

Comparison matrix

From Milestone 2 — MLP + training loop: every row here is a choice with a cost. Fill the expected loss column, then say which row you would actually pick and what you give up for it.

checkpointexpected lossexpected acc
epoch 1~2.31~0.09
epoch 10~2.21~0.46
epoch 20~2.04~0.68
epoch 50~0.89~0.88

57. The full program

Concept

from sklearn.datasets import load_digits
import torch, torch.nn as nn, torch.optim as optim
from torch.utils.data import Dataset, DataLoader

class DigitsDS(Dataset):
    def __init__(self):
        d = load_digits()
        self.X = torch.tensor(d.data/16.0, dtype=torch.float32)
        self.y = torch.tensor(d.target, dtype=torch.long)
    def __len__(self): return len(self.X)
    def __getitem__(self, i): return self.X[i], self.y[i]

torch.manual_seed(0)
ds = DigitsDS()
dl = DataLoader(ds, batch_size=64, shuffle=True)

model = nn.Sequential(
    nn.Linear(64,128), nn.BatchNorm1d(128), nn.ReLU(), nn.Dropout(0.3),
    nn.Linear(128,64), nn.ReLU(), nn.Linear(64,10))
opt   = optim.Adam(model.parameters(), lr=1e-3)
sched = optim.lr_scheduler.StepLR(opt, step_size=10, gamma=0.5)
lossfn = nn.CrossEntropyLoss()

for ep in range(1, 51):
    model.train()
    for xb, yb in dl:
        opt.zero_grad()
        lossfn(model(xb), yb).backward()
        opt.step()
    sched.step()

model.eval()
Xa = torch.tensor(load_digits().data/16.0, dtype=torch.float32)
ya = torch.tensor(load_digits().target, dtype=torch.long)
with torch.no_grad():
    acc = (model(Xa).argmax(1)==ya).float().mean().item()
print(f'acc={acc:.4f}')   # ~0.88
componentrole
DigitsDSDataset subclass — wraps numpy arrays
DataLoadermini-batch iterator, shuffle each epoch
BatchNorm1dstabilize activations per batch
Dropout(0.3)regularize; disabled at eval()
Adam + StepLRadaptive lr, halved every 10 epochs
sched.step()called after opt.step(), end of epoch

If your accuracy lands near 0.88 on the training set after 50 epochs — every technique from today is working correctly.

58. Fill in: role for The full program

Comparison

Comparison matrix

From The full program: refill the role column from what you know. The rest of the table is as it appeared.

componentrole
DigitsDSDataset subclass — wraps numpy arrays
DataLoadermini-batch iterator, shuffle each epoch
BatchNorm1dstabilize activations per batch
Dropout(0.3)regularize; disabled at eval()
Adam + StepLRadaptive lr, halved every 10 epochs
sched.step()called after opt.step(), end of epoch

59. Show it off

Concept

Slides closed: explain (1) why a 16-layer sigmoid net fails and two architectural fixes, (2) the exact sequence of the training loop including where sched.step() goes, and (3) what BatchNorm computes at training vs inference.

Stretch (homework): implement gradient clipping (clip_grad_norm_(model.parameters(), 1.0)) before opt.step(); swap in a CNN (Conv2d on 1×8×8 inputs); derive on paper that 0.25^16 ≈ 2.3×10⁻¹⁰. Next up: attention mechanisms (Vaswani et al. 2017 — read the abstract and architecture diagram before Lesson 78).

60. Connect it up: Lesson 77: Deep Learning Foundations

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Architecture: MLP → CNN → ResNet · Activations & Vanishing Gradients · BatchNorm, Dropout & Initialization · Backprop & Optimizers · Data Pipeline & Pattern · Your turn: full digits classifier. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

61. What you can do now

Recap

conceptthe one thing to remember
vanishing gradientsigmoid kills it; ReLU or skip connections fix it
BatchNormbatch stats at train; running stats at eval
Dropoutzero with prob p, scale by 1/(1-p); off at eval
Adam vs SGDAdam adapts lr per-param; SGD needs tuned global lr
sched.step()end of epoch, AFTER opt.step()

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 77 (Phase 4 — Deep Learning Foundations) — Barron · USAAIO Round 2 Preparation, 2026
  2. Activation outputs, vanishing-gradient magnitudes, BatchNorm trace, optimizer trajectories, Adam internals, LR scheduling, DataLoader batch shapes, backprop gradients, and MLP/CNN training curves verified — torch 2.7.1 + numpy 2.2.6 + scikit-learn, real execution, June 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108