USAAIO Lesson 55, from Phase 2. It builds 2D convolution from scratch, covering the output-shape formula, parameter sharing, and translation equivariance, then max pooling and the growth of the receptive field, and finishes with a complete CNN forward pass in PyTorch using nn.Conv2d, nn.MaxPool2d, and nn.ReLU. It was verified on the sklearn digits dataset, reaching 97% test accuracy after 100 epochs. The lesson runs to 30 slides.
Subject: Machine Learning · 59 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 55 · Phase 2
Output shape formula, parameter sharing, translation equivariance, max pooling, receptive field, and the full PyTorch CNN forward pass — everything you need to reason about any convolutional architecture.
Objectives
N, K, P, SWarm-up
Discussion prompt
Before we open Lesson 55: 2D Convolutions & CNNs: without looking back, what was the main idea of DBSCAN & Hierarchical Clustering, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
DBSCAN (core/border/noise, epsilon/min_samples), agglomerative hierarchical clustering with dendrograms, linkage methods (single/complete/average/Ward), cluster validation without labels (silhouette) and with labels (ARI, NMI), and a practical guide matching dataset geometry to the right algorithm. Verified with sklearn + scipy, June 2026.
Section
Part 1 of 4
Concept
A convolutional layer slides a small kernel over the input. At every position (i, j) it computes a dot product between the kernel and the overlapping input patch.
\[ \text{output}[i,j] = \sum_{m=0}^{K-1}\sum_{n=0}^{K-1} \text{input}[i+m,\; j+n]\cdot\text{kernel}[m,n] \]
Stacking multiple kernels gives multiple output channels (feature maps). Bias is added per output channel after the sum.
Counterexample
Discussion prompt
A convolutional layer slides a small kernel over the input. At every position (i, j) it computes a dot product between the kernel and the overlapping input patch.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Stacking multiple kernels gives multiple output channels (feature maps). Bias is added per output channel after the sum.
Concept
Given input size N, kernel size K, padding P, and stride S:
\[ \text{out} = \left\lfloor \frac{N + 2P - K}{S} \right\rfloor + 1 \]
| N | K | P | S | out |
|---|---|---|---|---|
| 28 | 3 | 0 | 1 | 26 |
| 28 | 3 | 1 | 1 | 28 (same padding) |
| 32 | 3 | 0 | 2 | 15 |
| 32 | 3 | 1 | 2 | 16 |
Comparison
Comparison matrix
From Output shape formula: refill the K column from what you know. The rest of the table is as it appeared.
| N | K | P | S | out |
|---|---|---|---|---|
| 28 | 3 | 0 | 1 | 26 |
| 28 | 3 | 1 | 1 | 28 (same padding) |
| 32 | 3 | 0 | 2 | 15 |
| 32 | 3 | 1 | 2 | 16 |
Estimation
Predict first
Apply a vertical-edge kernel to a 5×5 input with no padding, stride 1. Kernel detects left-minus-right gradient (bright on left, dark on right).
Commit before you compute: what does Manual 2D convolution (nested loop) come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: output shape = (3,3), matches floor((5+0-3)/1)+1 = 3
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Each kernel placement covers a 3×3 patch; with 5-input and 3-kernel and no padding there are exactly 3 valid positions per axis.
Worked example
Apply a vertical-edge kernel to a 5×5 input with no padding, stride 1. Kernel detects left-minus-right gradient (bright on left, dark on right).
import numpy as np
input_mat = np.array([
[1,2,3,0,0],[4,5,6,1,0],[7,8,9,2,0],
[0,1,2,3,1],[0,0,1,2,3]], dtype=float)
kernel = np.array([[1,0,-1],[1,0,-1],[1,0,-1]], dtype=float)
out_n = (5 + 2*0 - 3) // 1 + 1 # = 3
output = np.zeros((out_n, out_n))
for i in range(out_n):
for j in range(out_n):
output[i, j] = np.sum(input_mat[i:i+3, j:j+3] * kernel)
print(output)output shape = (3,3), matches floor((5+0-3)/1)+1 = 3
Why: Each kernel placement covers a 3×3 patch; with 5-input and 3-kernel and no padding there are exactly 3 valid positions per axis.
| (i,j) | patch left-col sum | patch right-col sum | output[i,j] |
|---|---|---|---|
| (0,0) | 1+4+7 = 12 | 3+6+9 = 18 | -6.0 |
| (0,1) | 2+5+8 = 15 | 0+1+2 = 3 | 12.0 |
| (0,2) | 3+6+9 = 18 | 0+0+0 = 0 | 18.0 |
Trade off
Comparison matrix
From Manual 2D convolution (nested loop): every row here is a choice with a cost. Fill the output[i,j] column, then say which row you would actually pick and what you give up for it.
| (i,j) | patch left-col sum | patch right-col sum | output[i,j] |
|---|---|---|---|
| (0,0) | 1+4+7 = 12 | 3+6+9 = 18 | -6.0 |
| (0,1) | 2+5+8 = 15 | 0+1+2 = 3 | 12.0 |
| (0,2) | 3+6+9 = 18 | 0+0+0 = 0 | 18.0 |
Concept
The same kernel weights are reused at every spatial position. A 3×3 conv with C_in input channels and C_out output channels has exactly K²·C_in·C_out + C_out parameters — regardless of the input size.
\[ \text{params}_{\text{conv}} = K^2 \cdot C_{\text{in}} \cdot C_{\text{out}} + C_{\text{out}} \]
| layer type | setting | parameter count |
|---|---|---|
| Conv2d(3→64, 3×3) | K=3, C_in=3, C_out=64 | 1,792 |
| FC equivalent | 32×32×3 → 30×30×64 | 176,947,200 |
| ratio | FC / Conv | ≈ 98,743× |
Concept
Equivariance: if the input shifts by (Δi, Δj), the output feature map shifts by the same amount (scaled by stride). The detector is location-agnostic.
This is why CNNs work for images: a filter that detects an edge at position (5,5) detects the same edge at (50,50) with the same weights — no extra parameters needed.
Note: equivariance ≠ invariance. The output moves with the input shift. It is max pooling that discards position information and provides (approximate) invariance.
Analogy
Discussion prompt
Explain Translation equivariance by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Note: equivariance ≠ invariance. The output moves with the input shift. It is max pooling that discards position information and provides (approximate) invariance.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Convolution gives translation invariance: the output is the same regardless of where the feature appears.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: a shifted input produces a shifted output map.
Convolution gives equivariance; max-pooling adds invariance.
Why: a shifted input produces a shifted output map. Convolution is equivariant (output shifts too), not invariant (output stays identical).
Trap
Convolution gives translation invariance: the output is the same regardless of where the feature appears.
Shift the input right by 1 pixel — the output feature map is unchanged
Why: This is wrong: a shifted input produces a shifted output map. Convolution is equivariant (output shifts too), not invariant (output stays identical).
Convolution gives equivariance; max-pooling adds invariance.
Shift the input right by 1 pixel — the output feature map shifts right by 1 too
Why: Convolution tracks where the feature is (equivariant). Max pooling then discards small position differences by taking the largest activation in each window — that's the invariance.
Break the constraint
Discussion prompt
The rule this trap just fixed:
Convolution tracks where the feature is (equivariant). Max pooling then discards small position differences by taking the largest activation in each window — that's the invariance.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
a shifted input produces a shifted output map. Convolution is equivariant (output shifts too), not invariant (output stays identical).
Section
Part 2 of 4
Concept
Max pooling slides a window (typically 2×2, stride 2) over a feature map and outputs the maximum value in each window. It provides two things:
Average pooling is an alternative (smoother gradient) but max pooling dominates in classification CNNs.
Explain it
Discussion prompt
Explain Max pooling to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Max pooling slides a window (typically 2×2, stride 2) over a feature map and outputs the maximum value in each window. It provides two things:
Estimation
Predict first
Pool a 4×4 feature map with a 2×2 window, stride 2.
Commit before you compute: what does Max pooling trace (2×2, stride 2) come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: top-left window max([1,3,5,6])=6; top-right max([2,4,1,2])=4; bottom-left max([4,2,0,3])=4; bottom-right max([7,1,2,8])=8
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Each 2×2 block contributes one scalar — the activations that survived, i.e.
Worked example
Pool a 4×4 feature map with a 2×2 window, stride 2.
import numpy as np, torch, torch.nn.functional as F
fm = np.array([[1,3,2,4],[5,6,1,2],[4,2,7,1],[0,3,2,8]], dtype=np.float32)
fm_t = torch.tensor(fm).unsqueeze(0).unsqueeze(0) # (1,1,4,4)
pooled = F.max_pool2d(fm_t, 2, stride=2).squeeze()
print(pooled.numpy())top-left window max([1,3,5,6])=6; top-right max([2,4,1,2])=4; bottom-left max([4,2,0,3])=4; bottom-right max([7,1,2,8])=8
Why: Each 2×2 block contributes one scalar — the activations that survived, i.e. the strongest detector responses. Spatial size 4×4 → 2×2.
| window | values | max |
|---|---|---|
| top-left | 1, 3, 5, 6 | 6 |
| top-right | 2, 4, 1, 2 | 4 |
| bottom-left | 4, 2, 0, 3 | 4 |
| bottom-right | 7, 1, 2, 8 | 8 |
Discrimination
Sort into buckets
Sort these by max, from memory, without looking back at Max pooling trace (2×2, stride 2). Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Concept
The receptive field of a neuron is the set of input pixels that can influence its value. It grows with depth: each additional conv layer adds K-1 on each side.
\[ \text{RF after } L \text{ layers of } K{\times}K \text{ (stride 1)} = 1 + L \cdot (K-1) \]
| layers L | kernel K | receptive field |
|---|---|---|
| 1 | 3×3 | 3×3 |
| 2 | 3×3 | 5×5 |
| 3 | 3×3 | 7×7 |
| 5 | 3×3 | 11×11 |
Invariant
Step through it
Step through Receptive field one row at a time. One of these columns never changes — find it, and say why it cannot.
Concept
Two stacked 3×3 conv layers have a 5×5 receptive field — the same as one 5×5 layer — but with fewer parameters and two nonlinearities instead of one.
| design | params (C_in=C_out=C) | nonlinearities |
|---|---|---|
| one 5×5 conv | 25·C² | 1 |
| two 3×3 convs | 18·C² | 2 |
| savings vs 5×5 | 28% fewer | +1 ReLU layer |
This is why VGG16 (2014) and almost every modern CNN uses 3×3 kernels exclusively (Lesson 60 context).
Pattern
Step through it
Step through Why small kernels stack well one row at a time. What is driving the change, and what would the row after the last one be?
Anomaly
Predict first
A student writes this, and it looks reasonable:
With a 32×32 input, kernel 3×3, padding 1, stride 2: output = (32+2·1-3)/2 = 31/2 ≈ 15.5 → 16... and since it's exact, output is 16.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: This is actually correct for this case — but students often mis-apply the formula as (N-K)/S + 1 (dropping the +2P term) and get (32-3)/2+1 = 15 instead of 16.
Always use the full formula: floor((N + 2P - K) / S) + 1.
Why: This is actually correct for this case — but students often mis-apply the formula as (N-K)/S + 1 (dropping the +2P term) and get (32-3)/2+1 = 15 instead of 16.
Trap
With a 32×32 input, kernel 3×3, padding 1, stride 2: output = (32+2·1-3)/2 = 31/2 ≈ 15.5 → 16... and since it's exact, output is 16.
Apply the formula and get 16 for stride=2, pad=1
Why: This is actually correct for this case — but students often mis-apply the formula as (N-K)/S + 1 (dropping the +2P term) and get (32-3)/2+1 = 15 instead of 16.
Always use the full formula: floor((N + 2P - K) / S) + 1.
N=32, K=3, P=1, S=2: floor((32 + 2 - 3)/2) + 1 = floor(31/2) + 1 = 15 + 1 = 16
Why: Including the +2P term is critical: without padding, (32-3)/2+1 = 15; with pad=1, it's 16. Drop the padding term and every shape calculation is wrong.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
(i, j) it computes a dot product between the kernel and the overlapping input patch.; Given input size N, kernel size K, padding P, and stride S:; Note: equivariance ≠ invariance. The output moves with the input shift. It is max pooling that discards position information and provides (approximate) invariance.Section
Part 3 of 4
Concept
nn.Conv2d(C_in, C_out, kernel_size, padding, stride) — learnable kernelsnn.ReLU() — element-wise activation (applied after each conv)nn.MaxPool2d(kernel_size, stride) — spatial downsamplingnn.Flatten() or .flatten(1) — reshape (N, C, H, W) → (N, C*H*W) before the FC headnn.Linear(in_features, out_features) — fully connected classifier headThe standard stack: conv → ReLU → pool, repeated. Then flatten → FC → output logits. Training is the same loop as Lesson 40: zero_grad → forward → loss → backward → step.
Analogy
Discussion prompt
Explain CNN building blocks in PyTorch by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
The standard stack: conv → ReLU → pool, repeated. Then flatten → FC → output logits. Training is the same loop as Lesson 40: zero_grad → forward → loss → backward → step.
Estimation
Predict first
Build a two-conv CNN and train it on sklearn load_digits (1797 samples, 8×8 grayscale, 10 classes). Observe the activation shape at each layer.
Commit before you compute: what does SmallCNN on load_digits (8×8 images) come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: test accuracy = 97.2% after 100 epochs, total params = 38,282
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Two conv layers (1→16→32) + one max pool + two FC layers.
Worked example
Build a two-conv CNN and train it on sklearn load_digits (1797 samples, 8×8 grayscale, 10 classes). Observe the activation shape at each layer.
import torch, torch.nn as nn, torch.nn.functional as F
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
import numpy as np
digits = load_digits()
X = digits.data.reshape(-1,1,8,8).astype(np.float32) / 16.0
y = digits.target.astype(np.int64)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, random_state=42)
Xt = torch.tensor(X_tr); yt = torch.tensor(y_tr)
class SmallCNN(nn.Module):
def __init__(self):
super().__init__()
self.conv1 = nn.Conv2d(1, 16, 3, padding=1) # 8→8
self.conv2 = nn.Conv2d(16, 32, 3, padding=1) # 8→8
self.pool = nn.MaxPool2d(2, 2) # 8→4
self.fc1 = nn.Linear(32*4*4, 64)
self.fc2 = nn.Linear(64, 10)
def forward(self, x):
x = F.relu(self.conv1(x)) # (N,16,8,8)
x = F.relu(self.conv2(x)) # (N,32,8,8)
x = self.pool(x) # (N,32,4,4)
x = F.relu(self.fc1(x.flatten(1))) # (N,64)
return self.fc2(x) # (N,10)
torch.manual_seed(42)
model = SmallCNN()
opt = torch.optim.Adam(model.parameters(), lr=1e-3)
lossfn = nn.CrossEntropyLoss()
for epoch in range(100):
opt.zero_grad()
loss = lossfn(model(Xt), yt)
loss.backward(); opt.step()
with torch.no_grad():
acc = (model(torch.tensor(X_te)).argmax(1) == torch.tensor(y_te)).float().mean()
print(f'test acc: {acc.item()*100:.1f}%')test accuracy = 97.2% after 100 epochs, total params = 38,282
Why: Two conv layers (1→16→32) + one max pool + two FC layers. Same five-step training loop as Lesson 40 — zero_grad → forward → loss → backward → step.
| layer | output shape | params |
|---|---|---|
| Input | (N, 1, 8, 8) | — |
| Conv2d(1,16,3,p=1) | (N, 16, 8, 8) | 160 |
| Conv2d(16,32,3,p=1) | (N, 32, 8, 8) | 4,640 |
| MaxPool2d(2,2) | (N, 32, 4, 4) | — |
| Linear(512,64) | (N, 64) | 32,832 |
| Linear(64,10) | (N, 10) | 650 |
Comparison
Comparison matrix
From SmallCNN on load_digits (8×8 images): refill the output shape column from what you know. The rest of the table is as it appeared.
| layer | output shape | params |
|---|---|---|
| Input | (N, 1, 8, 8) | — |
| Conv2d(1,16,3,p=1) | (N, 16, 8, 8) | 160 |
| Conv2d(16,32,3,p=1) | (N, 32, 8, 8) | 4,640 |
| MaxPool2d(2,2) | (N, 32, 4, 4) | — |
| Linear(512,64) | (N, 64) | 32,832 |
| Linear(64,10) | (N, 10) | 650 |
Ranking
Put in order
These are the steps of The CNN design recipe, scrambled. Put them back in order before the next slide shows you.
floor((N + 2P − K) / S) + 1 — check every layer before codingK²·C_in·C_out + C_outConv2d → ReLU → MaxPool2d (repeat n times); depth builds receptive field1 + L·(K−1) after L layers of K×K stride-1 convflatten → Linear → ReLU → Linear → logits; train with standard loop (L40)Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
floor((N + 2P − K) / S) + 1 — check every layer before codingK²·C_in·C_out + C_outConv2d → ReLU → MaxPool2d (repeat n times); depth builds receptive field1 + L·(K−1) after L layers of K×K stride-1 convflatten → Linear → ReLU → Linear → logits; train with standard loop (L40)Sorting
Sort into buckets
These are the pieces of Lesson 55: 2D Convolutions & CNNs, out of order. Put each one back under the part of the lesson it belongs to.
Elimination
Eliminate the wrong options
A Conv2d layer has input size N=28, kernel K=5, padding P=0, stride S=1. What is the output size?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: floor((28 + 2·0 − 5)/1) + 1 = floor(23/1) + 1 = 24. Verified in the shape table above (N=28, K=5, P=0, S=1 → 24).
Check
Work it out before clicking.
Check your understanding
A Conv2d layer has input size N=28, kernel K=5, padding P=0, stride S=1. What is the output size?
Answer: A
Why: floor((28 + 2·0 − 5)/1) + 1 = floor(23/1) + 1 = 24. Verified in the shape table above (N=28, K=5, P=0, S=1 → 24).
Prediction
Predict first
How many learnable parameters does Conv2d(C_in=3, C_out=32, kernel_size=3) have (including bias)?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: 896
Why: K²·C_in·C_out + C_out = 9·3·32 + 32 = 864 + 32 = 896. Verified directly: sum(p.numel() for p in nn.Conv2d(3,32,3).parameters()) = 896.
Check
Recall the count formula.
Check your understanding
How many learnable parameters does Conv2d(C_in=3, C_out=32, kernel_size=3) have (including bias)?
Answer: A
Why: K²·C_in·C_out + C_out = 9·3·32 + 32 = 864 + 32 = 896. Verified directly: sum(p.numel() for p in nn.Conv2d(3,32,3).parameters()) = 896.
Elimination
Eliminate the wrong options
A CNN has 5 conv layers, each with a 3×3 kernel and stride 1 (no pooling). What is the receptive field of a neuron in the final layer?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: RF = 1 + L·(K−1) = 1 + 5·2 = 11. Verified in the receptive field table: 5 layers of 3×3 → 11×11.
Check
Apply the formula.
Check your understanding
A CNN has 5 conv layers, each with a 3×3 kernel and stride 1 (no pooling). What is the receptive field of a neuron in the final layer?
Answer: A
Why: RF = 1 + L·(K−1) = 1 + 5·2 = 11. Verified in the receptive field table: 5 layers of 3×3 → 11×11.
Section
Project
Concept
Build a CNN from scratch on load_digits, predicting digit class from an 8×8 image. Three milestones: manual conv → PyTorch forward pass → full training loop.
| # | milestone | key tool |
|---|---|---|
| 1 | Manual nested-loop 2D conv on a small array | numpy |
| 2 | Build the SmallCNN forward pass, verify shapes | nn.Conv2d, nn.MaxPool2d |
| 3 | Train for 100 epochs, reach >90% test accuracy | Adam, CrossEntropyLoss |
Build rules: type every line, compute the output shape formula for your architecture before writing code, and confirm shapes with print(x.shape) at each layer.
Counterexample
Discussion prompt
Build a CNN from scratch on load_digits, predicting digit class from an 8×8 image. Three milestones: manual conv → PyTorch forward pass → full training loop.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Build rules: type every line, compute the output shape formula for your architecture before writing code, and confirm shapes with print(x.shape) at each layer.
Worked example
Your turn: implement 2D conv as a nested loop on the 5×5 array with the vertical-edge kernel. Predict the output at position (0,1).
Hint: for i in range(out_n): for j in range(out_n): output[i,j] = np.sum(input[i:i+K, j:j+K] * kernel). Output size = floor((5+0-3)/1)+1 = 3.
import numpy as np
input_mat = np.array([
[1,2,3,0,0],[4,5,6,1,0],[7,8,9,2,0],
[0,1,2,3,1],[0,0,1,2,3]], dtype=float)
kernel = np.array([[1,0,-1],[1,0,-1],[1,0,-1]], dtype=float)
out_n = 3; output = np.zeros((out_n, out_n))
for i in range(out_n):
for j in range(out_n):
output[i, j] = np.sum(input_mat[i:i+3, j:j+3] * kernel)
print(output)| (i,j) | output[i,j] |
|---|---|
| (0,0) | -6.0 |
| (0,1) | 12.0 |
| (0,2) | 18.0 |
| (1,0) | -6.0 |
| (1,1) | 8.0 |
| (1,2) | 16.0 |
| (2,0) | -5.0 |
| (2,1) | 2.0 |
| (2,2) | 8.0 |
Pattern
Step through it
Step through Milestone 1 — manual convolution one row at a time. What is driving the change, and what would the row after the last one be?
Worked example
Your turn: build SmallCNN, pass a dummy batch of shape (4, 1, 8, 8) through each layer, and print the shape. Predict each shape before running.
Hint: use x = F.relu(self.conv1(x)) etc. Check after every layer. MaxPool2d(2,2) halves H and W.
import torch, torch.nn as nn, torch.nn.functional as F
model = nn.Sequential()
# just print shapes step by step
x = torch.zeros(4, 1, 8, 8)
print('input: ', x.shape)
conv1 = nn.Conv2d(1, 16, 3, padding=1)
conv2 = nn.Conv2d(16, 32, 3, padding=1)
pool = nn.MaxPool2d(2, 2)
x = F.relu(conv1(x)); print('conv1: ', x.shape)
x = F.relu(conv2(x)); print('conv2: ', x.shape)
x = pool(x); print('pool: ', x.shape)
x = x.flatten(1); print('flatten:', x.shape)| layer | output shape |
|---|---|
| input | (4, 1, 8, 8) |
| conv1 | (4, 16, 8, 8) |
| conv2 | (4, 32, 8, 8) |
| pool | (4, 32, 4, 4) |
| flatten | (4, 512) |
Pattern
Step through it
Step through Milestone 2 — CNN forward pass, shape check one row at a time. What is driving the change, and what would the row after the last one be?
Worked example
Your turn: complete the training loop (Lesson 40 pattern: zero_grad → forward → loss → backward → step) for 100 epochs. Predict whether Adam at lr=1e-3 converges.
Hint: opt = torch.optim.Adam(model.parameters(), lr=1e-3). The same five-step loop from Lesson 40 applies unchanged — only the model and loss function differ.
torch.manual_seed(42)
model = SmallCNN()
opt = torch.optim.Adam(model.parameters(), lr=1e-3)
lossfn = nn.CrossEntropyLoss()
for epoch in range(100):
opt.zero_grad()
loss = lossfn(model(Xt), yt)
loss.backward(); opt.step()
with torch.no_grad():
acc = (model(torch.tensor(X_te)).argmax(1)
== torch.tensor(y_te)).float().mean()
print(f'test acc: {acc.item()*100:.1f}%')| epoch | train loss | test acc |
|---|---|---|
| 0 | 2.3015 | 10.3% |
| 9 | 2.2118 | 33.9% |
| 24 | 1.7577 | 70.6% |
| 49 | 0.4197 | 90.6% |
| 99 | 0.0764 | 97.2% |
Trade off
Comparison matrix
From Milestone 3 — train to >90% accuracy: every row here is a choice with a cost. Fill the test acc column, then say which row you would actually pick and what you give up for it.
| epoch | train loss | test acc |
|---|---|---|
| 0 | 2.3015 | 10.3% |
| 9 | 2.2118 | 33.9% |
| 24 | 1.7577 | 70.6% |
| 49 | 0.4197 | 90.6% |
| 99 | 0.0764 | 97.2% |
Concept
import torch, torch.nn as nn, torch.nn.functional as F
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
import numpy as np
digits = load_digits()
X = digits.data.reshape(-1,1,8,8).astype(np.float32)/16.0
y = digits.target.astype(np.int64)
X_tr,X_te,y_tr,y_te = train_test_split(X,y,test_size=0.2,random_state=42)
Xt=torch.tensor(X_tr); yt=torch.tensor(y_tr)
class SmallCNN(nn.Module):
def __init__(self):
super().__init__()
self.conv1=nn.Conv2d(1,16,3,padding=1)
self.conv2=nn.Conv2d(16,32,3,padding=1)
self.pool=nn.MaxPool2d(2,2)
self.fc1=nn.Linear(32*4*4,64)
self.fc2=nn.Linear(64,10)
def forward(self,x):
x=F.relu(self.conv1(x))
x=F.relu(self.conv2(x))
x=self.pool(x)
return self.fc2(F.relu(self.fc1(x.flatten(1))))
torch.manual_seed(42)
m=SmallCNN(); opt=torch.optim.Adam(m.parameters(),lr=1e-3)
for _ in range(100):
opt.zero_grad(); nn.CrossEntropyLoss()(m(Xt),yt).backward(); opt.step()
with torch.no_grad():
acc=(m(torch.tensor(X_te)).argmax(1)==torch.tensor(y_te)).float().mean()
print(f'{acc.item()*100:.1f}%') # 97.2%| component | choice | why |
|---|---|---|
| architecture | conv→ReLU→conv→ReLU→pool→FC | standard stack; pool halves spatial size once |
| optimizer | Adam, lr=1e-3 | faster than SGD on small data (L40 note) |
| loss | CrossEntropyLoss | 10-class logits → NLL (not MSE — see L44) |
| accuracy | 97.2% @ epoch 100 | verified on 20% held-out test split |
38,282 parameters total — vs. ~176 million for a fully-connected equivalent. Parameter sharing is the core efficiency of CNNs.
Comparison
Comparison matrix
From The full program: refill the why column from what you know. The rest of the table is as it appeared.
| component | choice | why |
|---|---|---|
| architecture | conv→ReLU→conv→ReLU→pool→FC | standard stack; pool halves spatial size once |
| optimizer | Adam, lr=1e-3 | faster than SGD on small data (L40 note) |
| loss | CrossEntropyLoss | 10-class logits → NLL (not MSE — see L44) |
| accuracy | 97.2% @ epoch 100 | verified on 20% held-out test split |
Concept
Out loud, slides closed: explain (1) the output shape formula and what each term does, (2) why parameter sharing matters, and (3) the difference between translation equivariance (convolution) and translation invariance (pooling).
Stretch (homework): implement 2D convolution with np.lib.stride_tricks; compute the receptive field for your 5-layer design; build a CNN for CIFAR-10 (32×32×3 images) targeting >70% test accuracy. Next up: Lesson 56 — deeper CNN architectures, batch normalization, and residual connections.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — The convolution operation · Max pooling & receptive field · PyTorch CNN forward pass · Your turn: build a CNN. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
floor((N + 2P − K) / S) + 1 to compute any conv output shapeK²·C_in·C_out + C_out (independent of input size)1 + L·(K−1) and build a CNN in PyTorch| concept | the one thing to remember |
|---|---|
| output shape | floor((N+2P−K)/S)+1; verify every layer |
| parameter sharing | same kernel everywhere; params don't scale with image size |
| equivariance | conv output shifts with input — not the same as invariance |
| max pooling | picks strongest activation per window; halves spatial dims |
| receptive field | grows by K−1 per stride-1 layer; depth = reach |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.