Lesson 55: 2D Convolutions & CNNs

USAAIO Lesson 55, from Phase 2. It builds 2D convolution from scratch, covering the output-shape formula, parameter sharing, and translation equivariance, then max pooling and the growth of the receptive field, and finishes with a complete CNN forward pass in PyTorch using nn.Conv2d, nn.MaxPool2d, and nn.ReLU. It was verified on the sklearn digits dataset, reaching 97% test accuracy after 100 epochs. The lesson runs to 30 slides.

Subject: Machine Learning · 59 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. 2D Convolutions & CNNs from Scratch

Title

USAAIO · Lesson 55 · Phase 2

Output shape formula, parameter sharing, translation equivariance, max pooling, receptive field, and the full PyTorch CNN forward pass — everything you need to reason about any convolutional architecture.

2. By the end of this lesson you can

Objectives

  1. Compute the output shape of a conv layer given N, K, P, S
  2. Implement 2D convolution as a nested loop and explain parameter sharing
  3. State what translation equivariance means and why CNNs exploit it
  4. Trace a max pool output and state what spatial invariance it provides
  5. Calculate the receptive field of a multi-layer CNN and build one in PyTorch

3. What survived from DBSCAN & Hierarchical Clustering?

Warm-up

Discussion prompt

Before we open Lesson 55: 2D Convolutions & CNNs: without looking back, what was the main idea of DBSCAN & Hierarchical Clustering, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

DBSCAN (core/border/noise, epsilon/min_samples), agglomerative hierarchical clustering with dendrograms, linkage methods (single/complete/average/Ward), cluster validation without labels (silhouette) and with labels (ARI, NMI), and a practical guide matching dataset geometry to the right algorithm. Verified with sklearn + scipy, June 2026.

4. The convolution operation

Section

Part 1 of 4

5. 2D convolution: the formula

Concept

A convolutional layer slides a small kernel over the input. At every position (i, j) it computes a dot product between the kernel and the overlapping input patch.

\[ \text{output}[i,j] = \sum_{m=0}^{K-1}\sum_{n=0}^{K-1} \text{input}[i+m,\; j+n]\cdot\text{kernel}[m,n] \]

Stacking multiple kernels gives multiple output channels (feature maps). Bias is added per output channel after the sum.

6. Break it if you can: 2D convolution: the formula

Counterexample

Discussion prompt

A convolutional layer slides a small kernel over the input. At every position (i, j) it computes a dot product between the kernel and the overlapping input patch.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Stacking multiple kernels gives multiple output channels (feature maps). Bias is added per output channel after the sum.

7. Output shape formula

Concept

Given input size N, kernel size K, padding P, and stride S:

\[ \text{out} = \left\lfloor \frac{N + 2P - K}{S} \right\rfloor + 1 \]

NKPSout
2830126
2831128 (same padding)
3230215
3231216

8. Fill in: K for Output shape formula

Comparison

Comparison matrix

From Output shape formula: refill the K column from what you know. The rest of the table is as it appeared.

NKPSout
2830126
2831128 (same padding)
3230215
3231216

9. Guess the shape of the answer: Manual 2D convolution (nested loop)

Estimation

Predict first

Apply a vertical-edge kernel to a 5×5 input with no padding, stride 1. Kernel detects left-minus-right gradient (bright on left, dark on right).

Commit before you compute: what does Manual 2D convolution (nested loop) come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: output shape = (3,3), matches floor((5+0-3)/1)+1 = 3

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Each kernel placement covers a 3×3 patch; with 5-input and 3-kernel and no padding there are exactly 3 valid positions per axis.

10. Manual 2D convolution (nested loop)

Worked example

Apply a vertical-edge kernel to a 5×5 input with no padding, stride 1. Kernel detects left-minus-right gradient (bright on left, dark on right).

import numpy as np
input_mat = np.array([
    [1,2,3,0,0],[4,5,6,1,0],[7,8,9,2,0],
    [0,1,2,3,1],[0,0,1,2,3]], dtype=float)
kernel = np.array([[1,0,-1],[1,0,-1],[1,0,-1]], dtype=float)
out_n = (5 + 2*0 - 3) // 1 + 1          # = 3
output = np.zeros((out_n, out_n))
for i in range(out_n):
    for j in range(out_n):
        output[i, j] = np.sum(input_mat[i:i+3, j:j+3] * kernel)
print(output)

output shape = (3,3), matches floor((5+0-3)/1)+1 = 3

Why: Each kernel placement covers a 3×3 patch; with 5-input and 3-kernel and no padding there are exactly 3 valid positions per axis.

(i,j)patch left-col sumpatch right-col sumoutput[i,j]
(0,0)1+4+7 = 123+6+9 = 18-6.0
(0,1)2+5+8 = 150+1+2 = 312.0
(0,2)3+6+9 = 180+0+0 = 018.0

11. What each one costs: Manual 2D convolution (nested loop)

Trade off

Comparison matrix

From Manual 2D convolution (nested loop): every row here is a choice with a cost. Fill the output[i,j] column, then say which row you would actually pick and what you give up for it.

(i,j)patch left-col sumpatch right-col sumoutput[i,j]
(0,0)1+4+7 = 123+6+9 = 18-6.0
(0,1)2+5+8 = 150+1+2 = 312.0
(0,2)3+6+9 = 180+0+0 = 018.0

12. Parameter sharing

Concept

The same kernel weights are reused at every spatial position. A 3×3 conv with C_in input channels and C_out output channels has exactly K²·C_in·C_out + C_out parameters — regardless of the input size.

\[ \text{params}_{\text{conv}} = K^2 \cdot C_{\text{in}} \cdot C_{\text{out}} + C_{\text{out}} \]

layer typesettingparameter count
Conv2d(3→64, 3×3)K=3, C_in=3, C_out=641,792
FC equivalent32×32×3 → 30×30×64176,947,200
ratioFC / Conv≈ 98,743×

13. Translation equivariance

Concept

Equivariance: if the input shifts by (Δi, Δj), the output feature map shifts by the same amount (scaled by stride). The detector is location-agnostic.

This is why CNNs work for images: a filter that detects an edge at position (5,5) detects the same edge at (50,50) with the same weights — no extra parameters needed.

Note: equivariance ≠ invariance. The output moves with the input shift. It is max pooling that discards position information and provides (approximate) invariance.

14. By analogy: Translation equivariance

Analogy

Discussion prompt

Explain Translation equivariance by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Note: equivariance ≠ invariance. The output moves with the input shift. It is max pooling that discards position information and provides (approximate) invariance.

15. Something is wrong here: confusing equivariance with invariance

Anomaly

Predict first

A student writes this, and it looks reasonable:

Convolution gives translation invariance: the output is the same regardless of where the feature appears.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: a shifted input produces a shifted output map.

Convolution gives equivariance; max-pooling adds invariance.

Why: a shifted input produces a shifted output map. Convolution is equivariant (output shifts too), not invariant (output stays identical).

16. Trap: confusing equivariance with invariance

Trap

The trap

Convolution gives translation invariance: the output is the same regardless of where the feature appears.

Shift the input right by 1 pixel — the output feature map is unchanged

Why: This is wrong: a shifted input produces a shifted output map. Convolution is equivariant (output shifts too), not invariant (output stays identical).

The fix

Convolution gives equivariance; max-pooling adds invariance.

Shift the input right by 1 pixel — the output feature map shifts right by 1 too

Why: Convolution tracks where the feature is (equivariant). Max pooling then discards small position differences by taking the largest activation in each window — that's the invariance.

17. Break it on purpose: confusing equivariance with invariance

Break the constraint

Discussion prompt

The rule this trap just fixed:

Convolution tracks where the feature is (equivariant). Max pooling then discards small position differences by taking the largest activation in each window — that's the invariance.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

a shifted input produces a shifted output map. Convolution is equivariant (output shifts too), not invariant (output stays identical).

18. Max pooling & receptive field

Section

Part 2 of 4

19. Max pooling

Concept

Max pooling slides a window (typically 2×2, stride 2) over a feature map and outputs the maximum value in each window. It provides two things:

Average pooling is an alternative (smoother gradient) but max pooling dominates in classification CNNs.

20. Teach it back: Max pooling

Explain it

Discussion prompt

Explain Max pooling to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Max pooling slides a window (typically 2×2, stride 2) over a feature map and outputs the maximum value in each window. It provides two things:

21. Guess the shape of the answer: Max pooling trace (2×2, stride 2)

Estimation

Predict first

Pool a 4×4 feature map with a 2×2 window, stride 2.

Commit before you compute: what does Max pooling trace (2×2, stride 2) come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: top-left window max([1,3,5,6])=6; top-right max([2,4,1,2])=4; bottom-left max([4,2,0,3])=4; bottom-right max([7,1,2,8])=8

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Each 2×2 block contributes one scalar — the activations that survived, i.e.

22. Max pooling trace (2×2, stride 2)

Worked example

Pool a 4×4 feature map with a 2×2 window, stride 2.

import numpy as np, torch, torch.nn.functional as F
fm = np.array([[1,3,2,4],[5,6,1,2],[4,2,7,1],[0,3,2,8]], dtype=np.float32)
fm_t = torch.tensor(fm).unsqueeze(0).unsqueeze(0)  # (1,1,4,4)
pooled = F.max_pool2d(fm_t, 2, stride=2).squeeze()
print(pooled.numpy())

top-left window max([1,3,5,6])=6; top-right max([2,4,1,2])=4; bottom-left max([4,2,0,3])=4; bottom-right max([7,1,2,8])=8

Why: Each 2×2 block contributes one scalar — the activations that survived, i.e. the strongest detector responses. Spatial size 4×4 → 2×2.

windowvaluesmax
top-left1, 3, 5, 66
top-right2, 4, 1, 24
bottom-left4, 2, 0, 34
bottom-right7, 1, 2, 88

23. Which is which, by max

Discrimination

Sort into buckets

Sort these by max, from memory, without looking back at Max pooling trace (2×2, stride 2). Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

6
top-left
4
top-right; bottom-left
8
bottom-right
g1
max is "6" for top-left — that is what the table on "Max pooling trace (2×2, stride 2)" records, and it is the single property separating this group from the rest.
g2
max is "4" for top-right, bottom-left — that is what the table on "Max pooling trace (2×2, stride 2)" records, and it is the single property separating this group from the rest.
g3
max is "8" for bottom-right — that is what the table on "Max pooling trace (2×2, stride 2)" records, and it is the single property separating this group from the rest.

24. Receptive field

Concept

The receptive field of a neuron is the set of input pixels that can influence its value. It grows with depth: each additional conv layer adds K-1 on each side.

\[ \text{RF after } L \text{ layers of } K{\times}K \text{ (stride 1)} = 1 + L \cdot (K-1) \]

layers Lkernel Kreceptive field
13×33×3
23×35×5
33×37×7
53×311×11

25. What stays fixed: Receptive field

Invariant

Step through it

Step through Receptive field one row at a time. One of these columns never changes — find it, and say why it cannot.

  1. Step 1: layers L is 1
  2. Step 2: layers L is 2
  3. Step 3: layers L is 3
  4. Step 4: layers L is 5

26. Why small kernels stack well

Concept

Two stacked 3×3 conv layers have a 5×5 receptive field — the same as one 5×5 layer — but with fewer parameters and two nonlinearities instead of one.

designparams (C_in=C_out=C)nonlinearities
one 5×5 conv25·C²1
two 3×3 convs18·C²2
savings vs 5×528% fewer+1 ReLU layer

This is why VGG16 (2014) and almost every modern CNN uses 3×3 kernels exclusively (Lesson 60 context).

27. Watch it run: Why small kernels stack well

Pattern

Step through it

Step through Why small kernels stack well one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: design is one 5×5 conv
  2. Step 2: design is two 3×3 convs
  3. Step 3: design is savings vs 5×5

28. Something is wrong here: output shape off-by-one with stride

Anomaly

Predict first

A student writes this, and it looks reasonable:

With a 32×32 input, kernel 3×3, padding 1, stride 2: output = (32+2·1-3)/2 = 31/2 ≈ 15.5 → 16... and since it's exact, output is 16.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: This is actually correct for this case — but students often mis-apply the formula as (N-K)/S + 1 (dropping the +2P term) and get (32-3)/2+1 = 15 instead of 16.

Always use the full formula: floor((N + 2P - K) / S) + 1.

Why: This is actually correct for this case — but students often mis-apply the formula as (N-K)/S + 1 (dropping the +2P term) and get (32-3)/2+1 = 15 instead of 16.

29. Trap: output shape off-by-one with stride

Trap

The trap

With a 32×32 input, kernel 3×3, padding 1, stride 2: output = (32+2·1-3)/2 = 31/2 ≈ 15.5 → 16... and since it's exact, output is 16.

Apply the formula and get 16 for stride=2, pad=1

Why: This is actually correct for this case — but students often mis-apply the formula as (N-K)/S + 1 (dropping the +2P term) and get (32-3)/2+1 = 15 instead of 16.

The fix

Always use the full formula: floor((N + 2P - K) / S) + 1.

N=32, K=3, P=1, S=2: floor((32 + 2 - 3)/2) + 1 = floor(31/2) + 1 = 15 + 1 = 16

Why: Including the +2P term is critical: without padding, (32-3)/2+1 = 15; with pad=1, it's 16. Drop the padding term and every shape calculation is wrong.

30. Which of these survive contact with Lesson 55: 2D Convolutions & CNNs?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
A convolutional layer slides a small kernel over the input. At every position (i, j) it computes a dot product between the kernel and the overlapping input patch.; Given input size N, kernel size K, padding P, and stride S:; Note: equivariance ≠ invariance. The output moves with the input shift. It is max pooling that discards position information and provides (approximate) invariance.
Breaks
Convolution gives translation invariance: the output is the same regardless of where the feature appears.; With a 32×32 input, kernel 3×3, padding 1, stride 2: output = (32+2·1-3)/2 = 31/2 ≈ 15.5 → 16... and since it's exact, output is 16.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 55: 2D Convolutions & CNNs puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

31. PyTorch CNN forward pass

Section

Part 3 of 4

32. CNN building blocks in PyTorch

Concept

The standard stack: conv → ReLU → pool, repeated. Then flatten → FC → output logits. Training is the same loop as Lesson 40: zero_grad → forward → loss → backward → step.

33. By analogy: CNN building blocks in PyTorch

Analogy

Discussion prompt

Explain CNN building blocks in PyTorch by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

The standard stack: conv → ReLU → pool, repeated. Then flatten → FC → output logits. Training is the same loop as Lesson 40: zero_grad → forward → loss → backward → step.

34. Guess the shape of the answer: SmallCNN on load_digits (8×8 images)

Estimation

Predict first

Build a two-conv CNN and train it on sklearn load_digits (1797 samples, 8×8 grayscale, 10 classes). Observe the activation shape at each layer.

Commit before you compute: what does SmallCNN on load_digits (8×8 images) come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: test accuracy = 97.2% after 100 epochs, total params = 38,282

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Two conv layers (1→16→32) + one max pool + two FC layers.

35. SmallCNN on load_digits (8×8 images)

Worked example

Build a two-conv CNN and train it on sklearn load_digits (1797 samples, 8×8 grayscale, 10 classes). Observe the activation shape at each layer.

import torch, torch.nn as nn, torch.nn.functional as F
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
import numpy as np

digits = load_digits()
X = digits.data.reshape(-1,1,8,8).astype(np.float32) / 16.0
y = digits.target.astype(np.int64)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, random_state=42)
Xt = torch.tensor(X_tr); yt = torch.tensor(y_tr)

class SmallCNN(nn.Module):
    def __init__(self):
        super().__init__()
        self.conv1 = nn.Conv2d(1, 16, 3, padding=1)   # 8→8
        self.conv2 = nn.Conv2d(16, 32, 3, padding=1)  # 8→8
        self.pool  = nn.MaxPool2d(2, 2)               # 8→4
        self.fc1   = nn.Linear(32*4*4, 64)
        self.fc2   = nn.Linear(64, 10)
    def forward(self, x):
        x = F.relu(self.conv1(x))     # (N,16,8,8)
        x = F.relu(self.conv2(x))     # (N,32,8,8)
        x = self.pool(x)              # (N,32,4,4)
        x = F.relu(self.fc1(x.flatten(1)))  # (N,64)
        return self.fc2(x)            # (N,10)

torch.manual_seed(42)
model = SmallCNN()
opt   = torch.optim.Adam(model.parameters(), lr=1e-3)
lossfn = nn.CrossEntropyLoss()
for epoch in range(100):
    opt.zero_grad()
    loss = lossfn(model(Xt), yt)
    loss.backward(); opt.step()
with torch.no_grad():
    acc = (model(torch.tensor(X_te)).argmax(1) == torch.tensor(y_te)).float().mean()
print(f'test acc: {acc.item()*100:.1f}%')

test accuracy = 97.2% after 100 epochs, total params = 38,282

Why: Two conv layers (1→16→32) + one max pool + two FC layers. Same five-step training loop as Lesson 40 — zero_grad → forward → loss → backward → step.

layeroutput shapeparams
Input(N, 1, 8, 8)—
Conv2d(1,16,3,p=1)(N, 16, 8, 8)160
Conv2d(16,32,3,p=1)(N, 32, 8, 8)4,640
MaxPool2d(2,2)(N, 32, 4, 4)—
Linear(512,64)(N, 64)32,832
Linear(64,10)(N, 10)650

36. Fill in: output shape for SmallCNN on load_digits (8×8 images)

Comparison

Comparison matrix

From SmallCNN on load_digits (8×8 images): refill the output shape column from what you know. The rest of the table is as it appeared.

layeroutput shapeparams
Input(N, 1, 8, 8)—
Conv2d(1,16,3,p=1)(N, 16, 8, 8)160
Conv2d(16,32,3,p=1)(N, 32, 8, 8)4,640
MaxPool2d(2,2)(N, 32, 4, 4)—
Linear(512,64)(N, 64)32,832
Linear(64,10)(N, 10)650

37. Rebuild the recipe: The CNN design recipe

Ranking

Put in order

These are the steps of The CNN design recipe, scrambled. Put them back in order before the next slide shows you.

  1. Output shape: floor((N + 2P − K) / S) + 1 — check every layer before coding
  2. Parameter sharing: kernels are reused spatially; params = K²·C_in·C_out + C_out
  3. Stack: Conv2d → ReLU → MaxPool2d (repeat n times); depth builds receptive field
  4. Receptive field: 1 + L·(K−1) after L layers of K×K stride-1 conv
  5. Head: flatten → Linear → ReLU → Linear → logits; train with standard loop (L40)

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

38. The CNN design recipe

Pattern

  1. Output shape: floor((N + 2P − K) / S) + 1 — check every layer before coding
  2. Parameter sharing: kernels are reused spatially; params = K²·C_in·C_out + C_out
  3. Stack: Conv2d → ReLU → MaxPool2d (repeat n times); depth builds receptive field
  4. Receptive field: 1 + L·(K−1) after L layers of K×K stride-1 conv
  5. Head: flatten → Linear → ReLU → Linear → logits; train with standard loop (L40)

39. Where does each piece belong: Lesson 55: 2D Convolutions & CNNs

Sorting

Sort into buckets

These are the pieces of Lesson 55: 2D Convolutions & CNNs, out of order. Put each one back under the part of the lesson it belongs to.

The convolution operation
2D convolution: the formula; Output shape formula; Manual 2D convolution (nested loop)
Max pooling & receptive field
Max pooling; Max pooling trace (2×2, stride 2); Receptive field
PyTorch CNN forward pass
CNN building blocks in PyTorch; SmallCNN on load_digits (8×8 images); The CNN design recipe
s1
The convolution operation is where Lesson 55: 2D Convolutions & CNNs puts 2D convolution: the formula, Output shape formula, Manual 2D convolution (nested loop). Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
Max pooling & receptive field is where Lesson 55: 2D Convolutions & CNNs puts Max pooling, Max pooling trace (2×2, stride 2), Receptive field. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
PyTorch CNN forward pass is where Lesson 55: 2D Convolutions & CNNs puts CNN building blocks in PyTorch, SmallCNN on load_digits (8×8 images), The CNN design recipe. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

40. Rule out three: Check yourself — output shape

Elimination

Eliminate the wrong options

A Conv2d layer has input size N=28, kernel K=5, padding P=0, stride S=1. What is the output size?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. 24
  • B. 28
  • C. 25
  • D. 23

Survives elimination: A

Why: floor((28 + 2·0 − 5)/1) + 1 = floor(23/1) + 1 = 24. Verified in the shape table above (N=28, K=5, P=0, S=1 → 24).

41. Check yourself — output shape

Check

Work it out before clicking.

Check your understanding

A Conv2d layer has input size N=28, kernel K=5, padding P=0, stride S=1. What is the output size?

  • A. 24 (correct)
  • B. 28
  • C. 25
  • D. 23

Answer: A

Why: floor((28 + 2·0 − 5)/1) + 1 = floor(23/1) + 1 = 24. Verified in the shape table above (N=28, K=5, P=0, S=1 → 24).

Why B tempts people
28 would require 'same' padding (P=2 for K=5, S=1). Without padding, the valid-pad conv shrinks the spatial size.
Why C tempts people
Off by one — using (N−K)/S instead of the full formula, giving (28−5)/1 = 23 before the +1, but forgetting +1 entirely gives 23; including +1 gives 24, not 25.
Why D tempts people
Forgot the +1 term: floor((28−5)/1) = 23 rather than 23+1 = 24.

42. Answer it before you see the options: Check yourself — parameter sharing

Prediction

Predict first

How many learnable parameters does Conv2d(C_in=3, C_out=32, kernel_size=3) have (including bias)?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: 896

Why: K²·C_in·C_out + C_out = 9·3·32 + 32 = 864 + 32 = 896. Verified directly: sum(p.numel() for p in nn.Conv2d(3,32,3).parameters()) = 896.

43. Check yourself — parameter sharing

Check

Recall the count formula.

Check your understanding

How many learnable parameters does Conv2d(C_in=3, C_out=32, kernel_size=3) have (including bias)?

  • A. 896 (correct)
  • B. 9 × 32 = 288
  • C. 3 × 32 = 96
  • D. 9 × 3 = 27

Answer: A

Why: K²·C_in·C_out + C_out = 9·3·32 + 32 = 864 + 32 = 896. Verified directly: sum(p.numel() for p in nn.Conv2d(3,32,3).parameters()) = 896.

Why B tempts people
9×32 = 288 forgets to multiply by C_in=3 — each of the 32 kernels has 9 weights per input channel, so 9·3 = 27 weights per kernel.
Why C tempts people
3×32 = 96 treats each kernel as a single scalar — one weight per (C_in, C_out) pair rather than K²=9 weights.
Why D tempts people
9×3 = 27 counts only the weights of one output kernel and ignores C_out=32 and the biases.

44. Rule out three: Check yourself — receptive field

Elimination

Eliminate the wrong options

A CNN has 5 conv layers, each with a 3×3 kernel and stride 1 (no pooling). What is the receptive field of a neuron in the final layer?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. 11×11
  • B. 15×15
  • C. 9×9
  • D. 25×25

Survives elimination: A

Why: RF = 1 + L·(K−1) = 1 + 5·2 = 11. Verified in the receptive field table: 5 layers of 3×3 → 11×11.

45. Check yourself — receptive field

Check

Apply the formula.

Check your understanding

A CNN has 5 conv layers, each with a 3×3 kernel and stride 1 (no pooling). What is the receptive field of a neuron in the final layer?

  • A. 11×11 (correct)
  • B. 15×15
  • C. 9×9
  • D. 25×25

Answer: A

Why: RF = 1 + L·(K−1) = 1 + 5·2 = 11. Verified in the receptive field table: 5 layers of 3×3 → 11×11.

Why B tempts people
15 = 3×5 applies K directly (5 layers × 3) instead of adding K−1=2 per layer to a starting RF of 1.
Why C tempts people
9 = 1 + 4·2 counts only 4 layers — an off-by-one on L, perhaps starting the layer count at 0.
Why D tempts people
25 = 5² treats the RF as K^L (exponential), which holds for dilated convolutions but not standard stride-1 convs.

46. Your turn: build a CNN

Section

Project

47. Project: 2D CNN on load_digits

Concept

Build a CNN from scratch on load_digits, predicting digit class from an 8×8 image. Three milestones: manual conv → PyTorch forward pass → full training loop.

#milestonekey tool
1Manual nested-loop 2D conv on a small arraynumpy
2Build the SmallCNN forward pass, verify shapesnn.Conv2d, nn.MaxPool2d
3Train for 100 epochs, reach >90% test accuracyAdam, CrossEntropyLoss

Build rules: type every line, compute the output shape formula for your architecture before writing code, and confirm shapes with print(x.shape) at each layer.

48. Break it if you can: Project: 2D CNN on load_digits

Counterexample

Discussion prompt

Build a CNN from scratch on load_digits, predicting digit class from an 8×8 image. Three milestones: manual conv → PyTorch forward pass → full training loop.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Build rules: type every line, compute the output shape formula for your architecture before writing code, and confirm shapes with print(x.shape) at each layer.

49. Milestone 1 — manual convolution

Worked example

Your turn: implement 2D conv as a nested loop on the 5×5 array with the vertical-edge kernel. Predict the output at position (0,1).

Hint: for i in range(out_n): for j in range(out_n): output[i,j] = np.sum(input[i:i+K, j:j+K] * kernel). Output size = floor((5+0-3)/1)+1 = 3.

import numpy as np
input_mat = np.array([
    [1,2,3,0,0],[4,5,6,1,0],[7,8,9,2,0],
    [0,1,2,3,1],[0,0,1,2,3]], dtype=float)
kernel = np.array([[1,0,-1],[1,0,-1],[1,0,-1]], dtype=float)
out_n = 3; output = np.zeros((out_n, out_n))
for i in range(out_n):
    for j in range(out_n):
        output[i, j] = np.sum(input_mat[i:i+3, j:j+3] * kernel)
print(output)
(i,j)output[i,j]
(0,0)-6.0
(0,1)12.0
(0,2)18.0
(1,0)-6.0
(1,1)8.0
(1,2)16.0
(2,0)-5.0
(2,1)2.0
(2,2)8.0

50. Watch it run: Milestone 1 — manual convolution

Pattern

Step through it

Step through Milestone 1 — manual convolution one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: (i,j) is (0,0)
  2. Step 2: (i,j) is (0,1)
  3. Step 3: (i,j) is (0,2)
  4. Step 4: (i,j) is (1,0)
  5. Step 5: (i,j) is (1,1)
  6. Step 6: (i,j) is (1,2)
  7. Step 7: (i,j) is (2,0)
  8. Step 8: (i,j) is (2,1)
  9. Step 9: (i,j) is (2,2)

51. Milestone 2 — CNN forward pass, shape check

Worked example

Your turn: build SmallCNN, pass a dummy batch of shape (4, 1, 8, 8) through each layer, and print the shape. Predict each shape before running.

Hint: use x = F.relu(self.conv1(x)) etc. Check after every layer. MaxPool2d(2,2) halves H and W.

import torch, torch.nn as nn, torch.nn.functional as F
model = nn.Sequential()
# just print shapes step by step
x = torch.zeros(4, 1, 8, 8)
print('input:  ', x.shape)
conv1 = nn.Conv2d(1, 16, 3, padding=1)
conv2 = nn.Conv2d(16, 32, 3, padding=1)
pool  = nn.MaxPool2d(2, 2)
x = F.relu(conv1(x));  print('conv1:  ', x.shape)
x = F.relu(conv2(x));  print('conv2:  ', x.shape)
x = pool(x);           print('pool:   ', x.shape)
x = x.flatten(1);      print('flatten:', x.shape)
layeroutput shape
input(4, 1, 8, 8)
conv1(4, 16, 8, 8)
conv2(4, 32, 8, 8)
pool(4, 32, 4, 4)
flatten(4, 512)

52. Watch it run: Milestone 2 — CNN forward pass, shape check

Pattern

Step through it

Step through Milestone 2 — CNN forward pass, shape check one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: layer is input
  2. Step 2: layer is conv1
  3. Step 3: layer is conv2
  4. Step 4: layer is pool
  5. Step 5: layer is flatten

53. Milestone 3 — train to >90% accuracy

Worked example

Your turn: complete the training loop (Lesson 40 pattern: zero_grad → forward → loss → backward → step) for 100 epochs. Predict whether Adam at lr=1e-3 converges.

Hint: opt = torch.optim.Adam(model.parameters(), lr=1e-3). The same five-step loop from Lesson 40 applies unchanged — only the model and loss function differ.

torch.manual_seed(42)
model = SmallCNN()
opt = torch.optim.Adam(model.parameters(), lr=1e-3)
lossfn = nn.CrossEntropyLoss()
for epoch in range(100):
    opt.zero_grad()
    loss = lossfn(model(Xt), yt)
    loss.backward(); opt.step()
with torch.no_grad():
    acc = (model(torch.tensor(X_te)).argmax(1)
           == torch.tensor(y_te)).float().mean()
print(f'test acc: {acc.item()*100:.1f}%')
epochtrain losstest acc
02.301510.3%
92.211833.9%
241.757770.6%
490.419790.6%
990.076497.2%

54. What each one costs: Milestone 3 — train to >90% accuracy

Trade off

Comparison matrix

From Milestone 3 — train to >90% accuracy: every row here is a choice with a cost. Fill the test acc column, then say which row you would actually pick and what you give up for it.

epochtrain losstest acc
02.301510.3%
92.211833.9%
241.757770.6%
490.419790.6%
990.076497.2%

55. The full program

Concept

import torch, torch.nn as nn, torch.nn.functional as F
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
import numpy as np

digits = load_digits()
X = digits.data.reshape(-1,1,8,8).astype(np.float32)/16.0
y = digits.target.astype(np.int64)
X_tr,X_te,y_tr,y_te = train_test_split(X,y,test_size=0.2,random_state=42)
Xt=torch.tensor(X_tr); yt=torch.tensor(y_tr)

class SmallCNN(nn.Module):
    def __init__(self):
        super().__init__()
        self.conv1=nn.Conv2d(1,16,3,padding=1)
        self.conv2=nn.Conv2d(16,32,3,padding=1)
        self.pool=nn.MaxPool2d(2,2)
        self.fc1=nn.Linear(32*4*4,64)
        self.fc2=nn.Linear(64,10)
    def forward(self,x):
        x=F.relu(self.conv1(x))
        x=F.relu(self.conv2(x))
        x=self.pool(x)
        return self.fc2(F.relu(self.fc1(x.flatten(1))))

torch.manual_seed(42)
m=SmallCNN(); opt=torch.optim.Adam(m.parameters(),lr=1e-3)
for _ in range(100):
    opt.zero_grad(); nn.CrossEntropyLoss()(m(Xt),yt).backward(); opt.step()
with torch.no_grad():
    acc=(m(torch.tensor(X_te)).argmax(1)==torch.tensor(y_te)).float().mean()
print(f'{acc.item()*100:.1f}%')   # 97.2%
componentchoicewhy
architectureconv→ReLU→conv→ReLU→pool→FCstandard stack; pool halves spatial size once
optimizerAdam, lr=1e-3faster than SGD on small data (L40 note)
lossCrossEntropyLoss10-class logits → NLL (not MSE — see L44)
accuracy97.2% @ epoch 100verified on 20% held-out test split

38,282 parameters total — vs. ~176 million for a fully-connected equivalent. Parameter sharing is the core efficiency of CNNs.

56. Fill in: why for The full program

Comparison

Comparison matrix

From The full program: refill the why column from what you know. The rest of the table is as it appeared.

componentchoicewhy
architectureconv→ReLU→conv→ReLU→pool→FCstandard stack; pool halves spatial size once
optimizerAdam, lr=1e-3faster than SGD on small data (L40 note)
lossCrossEntropyLoss10-class logits → NLL (not MSE — see L44)
accuracy97.2% @ epoch 100verified on 20% held-out test split

57. Show it off

Concept

Out loud, slides closed: explain (1) the output shape formula and what each term does, (2) why parameter sharing matters, and (3) the difference between translation equivariance (convolution) and translation invariance (pooling).

Stretch (homework): implement 2D convolution with np.lib.stride_tricks; compute the receptive field for your 5-layer design; build a CNN for CIFAR-10 (32×32×3 images) targeting >70% test accuracy. Next up: Lesson 56 — deeper CNN architectures, batch normalization, and residual connections.

58. Connect it up: Lesson 55: 2D Convolutions & CNNs

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — The convolution operation · Max pooling & receptive field · PyTorch CNN forward pass · Your turn: build a CNN. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

59. What you can do now

Recap

conceptthe one thing to remember
output shapefloor((N+2P−K)/S)+1; verify every layer
parameter sharingsame kernel everywhere; params don't scale with image size
equivarianceconv output shifts with input — not the same as invariance
max poolingpicks strongest activation per window; halves spatial dims
receptive fieldgrows by K−1 per stride-1 layer; depth = reach

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 55 — 2D Convolutions & CNNs — Barron · USAAIO Round 2 Preparation, 2026
  2. PyTorch CNN on sklearn load_digits; receptive field and parameter-sharing counts verified with numpy 2.2.6 and torch 2.7.1+cpu, June 2026 — Real execution, verified

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108