USAAIO Lesson 56, from Phase 3. It covers ResNet's identity shortcuts and how they fix the vanishing gradient, the BasicBlock against the Bottleneck design, and pre-activation against post-activation BatchNorm. It then covers depthwise separable convolutions, which give MobileNet roughly an eight-fold reduction in parameters, and EfficientNet's compound scaling of depth, width, and resolution together. It was verified with torch 2.7.1. The lesson runs to 30 slides.
Subject: Machine Learning · 56 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 56 · Phase 3: Modern Architectures
Why skip connections save deep training; how depthwise separable convolutions cut params by ~8x; how EfficientNet's compound scaling squeezes every FLOP.
Objectives
Warm-up
Discussion prompt
Before we open Lesson 56: ResNet, Depthwise Separable Convolutions & EfficientNet: without looking back, what was the main idea of 2D Convolutions & CNNs, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
2D convolution from scratch (output shape formula, parameter sharing, translation equivariance), max pooling, receptive field growth, and a complete CNN forward pass in PyTorch (nn.Conv2d, nn.MaxPool2d, nn.ReLU). Verified on sklearn digits reaching 97% test accuracy at 100 epochs.
Section
Section 1 of 4
Concept
In a plain 20-layer net, the gradient of the loss w.r.t. layer-0 weights is the product of 20 Jacobians. Each multiplication can shrink it exponentially.
Measured: a 20-layer net without skip connections has gradient norm 0.000001 at layer 0. The same net with identity shortcuts: 3760 — a factor of ~5 billion. (Lesson 35 showed this for general deep nets.)
| architecture | grad norm at layer 0 |
|---|---|
| 20-layer net, no skip | 0.000001 |
| 20-layer net, with skip | 3760.16 |
Comparison
Comparison matrix
From The vanishing-gradient problem: refill the grad norm at layer 0 column from what you know. The rest of the table is as it appeared.
| architecture | grad norm at layer 0 |
|---|---|
| 20-layer net, no skip | 0.000001 |
| 20-layer net, with skip | 3760.16 |
Concept
A residual block asks the layers to learn the residual F(x) = H(x) - x rather than H(x) directly. The identity path x is always available.
\[ \text{output} = F(x, \{W_i\}) + x \]
During backprop the shortcut carries gradient directly to earlier layers — no matrices to multiply through. If the block is unhelpful, the weights can learn F(x)=0 and output is just x.
Counterexample
Discussion prompt
A residual block asks the layers to learn the residual F(x) = H(x) - x rather than H(x) directly. The identity path x is always available.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Estimation
Predict first
Build two 20-layer networks with identical initialization; measure the gradient norm at layer 0 after one backward pass.
Commit before you compute: what does Gradient norm: skip vs no-skip come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: no_skip grad norm: 0.000001 — skip grad norm: 3760.16
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The shortcut carries the upstream gradient unattenuated; the learned path adds to it.
Worked example
Build two 20-layer networks with identical initialization; measure the gradient norm at layer 0 after one backward pass.
import torch
import torch.nn as nn
torch.manual_seed(42)
class DeepNoSkip(nn.Module):
def __init__(self):
super().__init__()
self.layers = nn.ModuleList([nn.Linear(64,64) for _ in range(20)])
def forward(self, x):
for l in self.layers: x = torch.relu(l(x))
return x.sum()
class DeepSkip(nn.Module):
def __init__(self):
super().__init__()
self.layers = nn.ModuleList([nn.Linear(64,64) for _ in range(20)])
def forward(self, x):
for l in self.layers: x = torch.relu(l(x)) + x # F(x)+x
return x.sum()
x = torch.randn(1, 64)
for name, net in [('no_skip', DeepNoSkip()), ('skip', DeepSkip())]:
net(x).backward()
g = net.layers[0].weight.grad.norm().item()
print(f'{name}: {g:.6f}')no_skip grad norm: 0.000001 — skip grad norm: 3760.16
Why: The shortcut carries the upstream gradient unattenuated; the learned path adds to it. Early layers receive usable gradient signal even at depth 20.
| variant | grad norm at layer 0 | trainable? |
|---|---|---|
| no skip | 0.000001 | no — vanished |
| with skip (F(x)+x) | 3760.16 | yes — flows |
Trade off
Comparison matrix
From Gradient norm: skip vs no-skip: every row here is a choice with a cost. Fill the grad norm at layer 0 column, then say which row you would actually pick and what you give up for it.
| variant | grad norm at layer 0 | trainable? |
|---|---|---|
| no skip | 0.000001 | no — vanished |
| with skip (F(x)+x) | 3760.16 | yes — flows |
Section
Section 2 of 4
Concept
| block | structure | params (C=64 in/out) | used in |
|---|---|---|---|
| BasicBlock | 3x3 -> 3x3 | 73,984 | ResNet-18, ResNet-34 |
| Bottleneck | 1x1 -> 3x3 -> 1x1 | 136,448 (256->64->256) | ResNet-50, -101, -152 |
Bottleneck uses 1x1 convs to compress channels before and expand after the expensive 3x3 — fewer multiplications for the same representational depth.
Analogy
Discussion prompt
Explain Two residual block designs by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Bottleneck uses 1x1 convs to compress channels before and expand after the expensive 3x3 — fewer multiplications for the same representational depth.
Estimation
Predict first
Measure parameters for a BasicBlock(64,64) and a Bottleneck(256->64->256) using nn.Conv2d layers.
Commit before you compute: what does BasicBlock and Bottleneck param counts come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: BasicBlock(64): 73,984 params; Bottleneck(256->64->256): 136,448 params
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Bottleneck costs more parameters here, but operates on 256-channel feature maps — at depth, the 1x1 bottleneck is far cheaper than a plain 3x3 on 256 channels (3x3x256x256 = 589,824).
Worked example
Measure parameters for a BasicBlock(64,64) and a Bottleneck(256->64->256) using nn.Conv2d layers.
import torch.nn as nn
class BasicBlock(nn.Module):
def __init__(self, c):
super().__init__()
self.net = nn.Sequential(
nn.Conv2d(c,c,3,padding=1,bias=False), nn.BatchNorm2d(c),
nn.ReLU(), nn.Conv2d(c,c,3,padding=1,bias=False), nn.BatchNorm2d(c))
def forward(self,x): return torch.relu(self.net(x) + x)
class Bottleneck(nn.Module):
def __init__(self, c_in=256, mid=64):
super().__init__()
c_out = mid * 4
self.net = nn.Sequential(
nn.Conv2d(c_in,mid,1,bias=False), nn.BatchNorm2d(mid), nn.ReLU(),
nn.Conv2d(mid,mid,3,padding=1,bias=False), nn.BatchNorm2d(mid), nn.ReLU(),
nn.Conv2d(mid,c_out,1,bias=False), nn.BatchNorm2d(c_out))
self.proj = nn.Conv2d(c_in,c_out,1,bias=False)
def forward(self,x): return torch.relu(self.net(x) + self.proj(x))
bb = BasicBlock(64)
bn = Bottleneck(256, 64)
print('BasicBlock(64):', sum(p.numel() for p in bb.parameters()))
print('Bottleneck(256->64->256):', sum(p.numel() for p in bn.parameters()))BasicBlock(64): 73,984 params; Bottleneck(256->64->256): 136,448 params
Why: Bottleneck costs more parameters here, but operates on 256-channel feature maps — at depth, the 1x1 bottleneck is far cheaper than a plain 3x3 on 256 channels (3x3x256x256 = 589,824).
| block | params | 3x3 on full channels |
|---|---|---|
| BasicBlock(64) | 73,984 | — |
| Bottleneck(256->64->256) | 136,448 | vs 589,824 for 3x3x256x256 |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
BasicBlock(64): 73,984 params; Bottleneck(256->64->256): 136,448 params
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Measure parameters for a BasicBlock(64,64) and a Bottleneck(256->64->256) using nn.Conv2d layers.
Concept
| variant | order | shortcut path | paper |
|---|---|---|---|
| post-activation (v1) | Conv -> BN -> ReLU | may have BN on shortcut | He et al. 2016 (original) |
| pre-activation (v2) | BN -> ReLU -> Conv | clean identity (no BN/ReLU) | He et al. 2016 (revised) |
Pre-activation means the shortcut is a clean identity with no activation — the gradient highway is truly unobstructed. Pre-act also slightly improves accuracy on very deep (1000+ layer) networks.
Explain it
Discussion prompt
Explain BatchNorm placement: pre-act vs post-act to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Pre-activation means the shortcut is a clean identity with no activation — the gradient highway is truly unobstructed. Pre-act also slightly improves accuracy on very deep (1000+ layer) networks.
Anomaly
Predict first
A student writes this, and it looks reasonable:
When stride=2 halves the spatial size (e.g. 56x56->28x28), use the identity shortcut directly: out = F(x) + x.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: F(x) has shape (B, 128, 28, 28); x has shape (B, 64, 56, 56).
When stride or channel count changes, apply a projection shortcut (1x1 conv + BN) to match dimensions before adding.
Why: F(x) has shape (B, 128, 28, 28); x has shape (B, 64, 56, 56). PyTorch raises a RuntimeError on the addition — the shapes don't broadcast.
Trap
When stride=2 halves the spatial size (e.g. 56x56->28x28), use the identity shortcut directly: out = F(x) + x.
Add F(x) + x when channels or stride changes
Why: F(x) has shape (B, 128, 28, 28); x has shape (B, 64, 56, 56). PyTorch raises a RuntimeError on the addition — the shapes don't broadcast.
When stride or channel count changes, apply a projection shortcut (1x1 conv + BN) to match dimensions before adding.
shortcut = nn.Sequential(Conv2d(64,128,1,stride=2,bias=False), BN(128)); out = F(x) + shortcut(x)
Why: The 1x1 conv reshapes x to (B,128,28,28) so the addition is well-defined. This is the 'option B' shortcut from the ResNet paper.
Break the constraint
Discussion prompt
The rule this trap just fixed:
The 1x1 conv reshapes x to (B,128,28,28) so the addition is well-defined. This is the 'option B' shortcut from the ResNet paper.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
F(x) has shape (B, 128, 28, 28); x has shape (B, 64, 56, 56). PyTorch raises a RuntimeError on the addition — the shapes don't broadcast.
Section
Section 3 of 4
Concept
A standard K×K conv with C_in input channels and C_out output channels costs K²·C_in·C_out multiplications per spatial location.
MobileNet splits this into two cheaper ops: a depthwise conv (K²·C_in, one filter per channel) followed by a pointwise 1×1 conv (C_in·C_out, mixes channels).
\[ \frac{K^2 C_{in} + C_{in} C_{out}}{K^2 C_{in} C_{out}} = \frac{1}{C_{out}} + \frac{1}{K^2} \]
Analogy
Discussion prompt
Explain Factoring a standard convolution by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
A standard K×K conv with C_in input channels and C_out output channels costs K²·C_in·C_out multiplications per spatial location.
Estimation
Predict first
For C_in=32, C_out=64, K=3: compare a standard nn.Conv2d(32,64,3) against groups=C_in depthwise + 1x1 pointwise.
Commit before you compute: what does Param reduction: depthwise sep vs standard come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Standard: 18,432 params; DW+PW: 2,336 params; reduction 7.89x
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Depthwise (288 params) handles spatial filtering per-channel; pointwise (2,048) mixes across channels.
Worked example
For C_in=32, C_out=64, K=3: compare a standard nn.Conv2d(32,64,3) against groups=C_in depthwise + 1x1 pointwise.
import torch.nn as nn
C_in, C_out, K = 32, 64, 3
std = nn.Conv2d(C_in, C_out, K, padding=1, bias=False)
dw = nn.Conv2d(C_in, C_in, K, padding=1, groups=C_in, bias=False) # depthwise
pw = nn.Conv2d(C_in, C_out, 1, bias=False) # pointwise
std_p = sum(p.numel() for p in std.parameters())
dwp_p = sum(p.numel() for p in list(dw.parameters()) + list(pw.parameters()))
print(f'Standard: {std_p:,}')
print(f'DW+PW: {dwp_p:,}')
print(f'Reduction: {std_p/dwp_p:.2f}x')Standard: 18,432 params; DW+PW: 2,336 params; reduction 7.89x
Why: Depthwise (288 params) handles spatial filtering per-channel; pointwise (2,048) mixes across channels. Together they replicate the 3x3 conv's function at 12.7% of the cost.
| C_in | C_out | K | Standard | DW+PW | Ratio |
|---|---|---|---|---|---|
| 3 | 32 | 3 | 864 | 123 | 7.02x |
| 32 | 64 | 3 | 18,432 | 2,336 | 7.89x |
| 64 | 128 | 3 | 73,728 | 8,768 | 8.41x |
| 128 | 256 | 3 | 294,912 | 33,920 | 8.69x |
Invariant
Step through it
Step through Param reduction: depthwise sep vs standard one row at a time. One of these columns never changes — find it, and say why it cannot.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Use nn.Conv2d(32, 32, 3, groups=1) for the depthwise step — groups=1 is the default.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: groups=1 is a fully-connected standard conv — every output channel sees every input channel.
Set groups=C_in so each filter handles exactly one input channel.
Why: groups=1 is a fully-connected standard conv — every output channel sees every input channel. It costs 3x3x32x32 = 9,216 params, not 288. You've negated the factorization entirely.
Trap
Use nn.Conv2d(32, 32, 3, groups=1) for the depthwise step — groups=1 is the default.
groups=1 on a depthwise layer
Why: groups=1 is a fully-connected standard conv — every output channel sees every input channel. It costs 3x3x32x32 = 9,216 params, not 288. You've negated the factorization entirely.
Set groups=C_in so each filter handles exactly one input channel.
nn.Conv2d(32, 32, 3, padding=1, groups=32, bias=False)
Why: groups=C_in means the conv has C_in separate filter banks of size 1xKxK. Each channel is filtered independently — that is the depthwise step. Params = KKC_in = 9*32 = 288.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
out = F(x) + x.; Use nn.Conv2d(32, 32, 3, groups=1) for the depthwise step — groups=1 is the default.Section
Section 4 of 4
Concept
Scaling only depth (more layers) hits diminishing returns because spatial resolution is a bottleneck. Scaling only width (more channels) also saturates. Scaling only resolution helps but needs more layers to process fine detail.
EfficientNet observes that optimal depth/width/resolution scale together under a fixed compute budget. It finds the optimal per-dimension coefficients once, then reuses them by raising them to a single composite exponent phi.
Concept
\[ d = \alpha^{\phi},\quad w = \beta^{\phi},\quad r = \gamma^{\phi} \]
Grid-searched constants: alpha=1.2 (depth), beta=1.1 (width), gamma=1.15 (resolution). Constraint: alpha * beta^2 * gamma^2 = 1.92 (approx 2.0, so doubling phi doubles FLOPs).
| variant | phi | depth | width | resolution |
|---|---|---|---|---|
| EfficientNet-B0 | 0 | 1.0x | 1.0x | 224 |
| EfficientNet-B1 | 1 | 1.2x | 1.1x | 258 |
| EfficientNet-B4 | 4 | 2.07x | 1.46x | 392 |
| EfficientNet-B7 | 7 | 3.58x | 1.95x | 596 |
Comparison
Comparison matrix
From The compound scaling formula: refill the width column from what you know. The rest of the table is as it appeared.
| variant | phi | depth | width | resolution |
|---|---|---|---|---|
| EfficientNet-B0 | 0 | 1.0x | 1.0x | 224 |
| EfficientNet-B1 | 1 | 1.2x | 1.1x | 258 |
| EfficientNet-B4 | 4 | 2.07x | 1.46x | 392 |
| EfficientNet-B7 | 7 | 3.58x | 1.95x | 596 |
Ranking
Put in order
These are the steps of The modern-architecture recipe, scrambled. Put them back in order before the next slide shows you.
Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
Edge cases
Discussion prompt
The modern-architecture recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
Elimination
Eliminate the wrong options
In a 50-layer ResNet, a residual block transitions from 64 to 128 channels with stride=2. Which shortcut is required?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: When channel count or stride changes, the shortcut must project x to match F(x). The standard (option B) from the ResNet paper is a 1x1 conv + BN that simultaneously changes channels and halves spatial size. The identity shortcut only works when shapes already match.
Check
Pause and reason before choosing.
Check your understanding
In a 50-layer ResNet, a residual block transitions from 64 to 128 channels with stride=2. Which shortcut is required?
Answer: A
Why: When channel count or stride changes, the shortcut must project x to match F(x). The standard (option B) from the ResNet paper is a 1x1 conv + BN that simultaneously changes channels and halves spatial size. The identity shortcut only works when shapes already match.
Prediction
Predict first
A standard Conv2d(128, 256, 3) has 128x256x9 = 294,912 parameters. Replacing it with a depthwise-separable block gives approximately how many parameters?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: ~33,920 (ratio ~8.7x)
Why: Depthwise: K^2 * C_in = 9*128 = 1,152 params. Pointwise: C_in * C_out = 128*256 = 32,768. Total = 33,920. Ratio = 294,912 / 33,920 = 8.7x. Matches 1/256 + 1/9 = 0.115, so 8.7x.
Check
Use the formula: reduction = 1/C_out + 1/K^2.
Check your understanding
A standard Conv2d(128, 256, 3) has 128x256x9 = 294,912 parameters. Replacing it with a depthwise-separable block gives approximately how many parameters?
Answer: A
Why: Depthwise: K^2 * C_in = 9*128 = 1,152 params. Pointwise: C_in * C_out = 128*256 = 32,768. Total = 33,920. Ratio = 294,912 / 33,920 = 8.7x. Matches 1/256 + 1/9 = 0.115, so 8.7x.
Elimination
Eliminate the wrong options
EfficientNet uses alpha=1.2, beta=1.1, gamma=1.15. At phi=2, what is the depth multiplier (alpha^phi)?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: depth_mult = alpha^phi = 1.2^2 = 1.44. Width_mult = beta^2 = 1.1^2 = 1.21, resolution = 224 * 1.15^2 = 296. This is EfficientNet-B2.
Check
Apply the compound scaling formula.
Check your understanding
EfficientNet uses alpha=1.2, beta=1.1, gamma=1.15. At phi=2, what is the depth multiplier (alpha^phi)?
Answer: A
Why: depth_mult = alpha^phi = 1.2^2 = 1.44. Width_mult = beta^2 = 1.1^2 = 1.21, resolution = 224 * 1.15^2 = 296. This is EfficientNet-B2.
Section
Project
Concept
Build a BasicBlock, a Bottleneck, and a depthwise-separable block. Count parameters and verify skip connection requires projection when shapes change.
| # | requirement | key tool |
|---|---|---|
| 1 | BasicBlock(64,64) with correct F(x)+x | nn.Conv2d, nn.BatchNorm2d |
| 2 | Bottleneck(256->64->256) with projection shortcut | 1x1 then 3x3 then 1x1 |
| 3 | DW-sep block — confirm 7.89x reduction for (32,64,3) | groups=C_in + 1x1 |
Rule: run every block with a synthetic input to confirm output shapes match, then print param counts.
Counterexample
Discussion prompt
Build a BasicBlock, a Bottleneck, and a depthwise-separable block. Count parameters and verify skip connection requires projection when shapes change.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Rule: run every block with a synthetic input to confirm output shapes match, then print param counts.
Worked example
Your turn: implement BasicBlock(64). Verify the skip path adds x unchanged. What is the param count?
Hint: two Conv2d(64,64,3,padding=1,bias=False) + BN; shortcut is nn.Identity() since channels and stride match.
import torch, torch.nn as nn
class BasicBlock(nn.Module):
def __init__(self, c):
super().__init__()
self.conv1 = nn.Conv2d(c, c, 3, padding=1, bias=False)
self.bn1 = nn.BatchNorm2d(c)
self.conv2 = nn.Conv2d(c, c, 3, padding=1, bias=False)
self.bn2 = nn.BatchNorm2d(c)
self.relu = nn.ReLU(inplace=True)
def forward(self, x):
out = self.relu(self.bn1(self.conv1(x)))
out = self.bn2(self.conv2(out))
return self.relu(out + x) # F(x) + x
bb = BasicBlock(64)
x = torch.randn(2, 64, 8, 8)
print('output shape:', tuple(bb(x).shape))
print('params:', sum(p.numel() for p in bb.parameters()))| output shape | param count |
|---|---|
| (2, 64, 8, 8) | 73,984 |
Worked example
Your turn: implement a DW-sep block for (C_in=32, C_out=64, K=3). Predict the param count before running.
Hint: groups=C_in in the depthwise layer; pointwise is Conv2d(C_in, C_out, 1, bias=False). Expected: 2,336.
import torch, torch.nn as nn
class DWSepConv(nn.Module):
def __init__(self, c_in, c_out, k=3):
super().__init__()
self.dw = nn.Conv2d(c_in, c_in, k, padding=k//2, groups=c_in, bias=False)
self.pw = nn.Conv2d(c_in, c_out, 1, bias=False)
self.bn = nn.BatchNorm2d(c_out)
self.relu = nn.ReLU(inplace=True)
def forward(self, x): return self.relu(self.bn(self.pw(self.dw(x))))
dws = DWSepConv(32, 64, 3)
x = torch.randn(1, 32, 16, 16)
print('output shape:', tuple(dws(x).shape))
print('DW params:', sum(p.numel() for p in dws.dw.parameters()))
print('PW params:', sum(p.numel() for p in dws.pw.parameters()))
print('Total DW-sep:', sum(p.numel() for p in dws.parameters()))| layer | params |
|---|---|
| depthwise (groups=32) | 288 |
| pointwise (1x1) | 2,048 |
| BN (64) | 128 |
| total DW-sep | 2,464 |
Worked example
Your turn: build minimal ResNet-18 and ResNet-50 skeletons (no training). Count parameters and compare.
Hint: ResNet-18 uses 8 BasicBlocks; ResNet-50 uses 16 Bottlenecks. Both end with nn.Linear(512/2048, 1000).
# (Using pre-built torchvision for a correct count)
import torch
from torchvision.models import resnet18, resnet50
r18 = resnet18(); r50 = resnet50()
p18 = sum(p.numel() for p in r18.parameters())
p50 = sum(p.numel() for p in r50.parameters())
print(f'ResNet-18: {p18:,} ({p18/1e6:.1f}M)')
print(f'ResNet-50: {p50:,} ({p50/1e6:.1f}M)')| model | blocks | params | notes |
|---|---|---|---|
| ResNet-18 | 8 BasicBlocks | 11.7M | 2x3x3 per block |
| ResNet-50 | 16 Bottlenecks | 25.6M | 1x1->3x3->1x1 per block |
Trade off
Comparison matrix
From Milestone 3 — ResNet-18 vs ResNet-50 at scale: every row here is a choice with a cost. Fill the notes column, then say which row you would actually pick and what you give up for it.
| model | blocks | params | notes |
|---|---|---|---|
| ResNet-18 | 8 BasicBlocks | 11.7M | 2x3x3 per block |
| ResNet-50 | 16 Bottlenecks | 25.6M | 1x1->3x3->1x1 per block |
Concept
import torch, torch.nn as nn
# BasicBlock
class BB(nn.Module):
def __init__(self, c): super().__init__(); self.net = nn.Sequential(nn.Conv2d(c,c,3,padding=1,bias=False),nn.BatchNorm2d(c),nn.ReLU(),nn.Conv2d(c,c,3,padding=1,bias=False),nn.BatchNorm2d(c))
def forward(self, x): return torch.relu(self.net(x) + x)
# DW-Sep
class DW(nn.Module):
def __init__(self, ci, co): super().__init__(); self.dw=nn.Conv2d(ci,ci,3,padding=1,groups=ci,bias=False); self.pw=nn.Conv2d(ci,co,1,bias=False)
def forward(self, x): return self.pw(self.dw(x))
std = nn.Conv2d(32, 64, 3, padding=1, bias=False)
dw = DW(32, 64)
print('BasicBlock(64) params:', sum(p.numel() for p in BB(64).parameters()))
print('Std conv(32->64,3):', sum(p.numel() for p in std.parameters()))
print('DW-sep(32->64,3):', sum(p.numel() for p in dw.parameters()))
print('Reduction:', round(sum(p.numel() for p in std.parameters()) / sum(p.numel() for p in dw.parameters()), 2), 'x')| block | params | note |
|---|---|---|
| BasicBlock(64) | 73,984 | F(x)+x, identity shortcut |
| Std conv(32->64,K=3) | 18,432 | full spatial mix |
| DW-sep(32->64,K=3) | 2,336 | factored: 288+2048 |
| Reduction | 7.89x | matches 1/64 + 1/9 |
If your numbers match the table — you've implemented the building blocks of MobileNet and ResNet from scratch.
Comparison
Comparison matrix
From The full program: refill the note column from what you know. The rest of the table is as it appeared.
| block | params | note |
|---|---|---|
| BasicBlock(64) | 73,984 | F(x)+x, identity shortcut |
| Std conv(32->64,K=3) | 18,432 | full spatial mix |
| DW-sep(32->64,K=3) | 2,336 | factored: 288+2048 |
| Reduction | 7.89x | matches 1/64 + 1/9 |
Concept
Out loud, slides closed: (1) explain the vanishing-gradient problem and how F(x)+x fixes it; (2) derive the depthwise-sep reduction formula 1/C_out + 1/K^2; (3) describe EfficientNet's phi parameter and why the constraint alphabeta^2gamma^2 ~= 2 matters.
Stretch (homework from the lesson plan): implement a full ResNet-18-style network in PyTorch and train on CIFAR-10; compare accuracy with and without skip connections at 20 epochs. Next up: attention mechanisms and the Transformer architecture.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — The identity shortcut — F(x) + x · BasicBlock vs Bottleneck · Depthwise separable convolutions · EfficientNet compound scaling · Your turn: build it. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
| idea | the one thing to remember |
|---|---|
| skip connection | F(x)+x; projection shortcut when shapes differ |
| BasicBlock vs Bottleneck | BasicBlock: 2x3x3; Bottleneck: 1x1->3x3->1x1 |
| BN placement | post-act (v1) vs pre-act (v2, cleaner shortcut) |
| depthwise sep | groups=C_in dw + 1x1 pw; ~8x fewer params |
| EfficientNet | alpha^phi * beta^phi * gamma^phi; constraint ~= 2 |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.