Lesson 56: ResNet, Depthwise Separable Convolutions & EfficientNet

USAAIO Lesson 56, from Phase 3. It covers ResNet's identity shortcuts and how they fix the vanishing gradient, the BasicBlock against the Bottleneck design, and pre-activation against post-activation BatchNorm. It then covers depthwise separable convolutions, which give MobileNet roughly an eight-fold reduction in parameters, and EfficientNet's compound scaling of depth, width, and resolution together. It was verified with torch 2.7.1. The lesson runs to 30 slides.

Subject: Machine Learning · 56 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. ResNet, Depthwise Separable Convolutions & EfficientNet

Title

USAAIO · Lesson 56 · Phase 3: Modern Architectures

Why skip connections save deep training; how depthwise separable convolutions cut params by ~8x; how EfficientNet's compound scaling squeezes every FLOP.

2. By the end of this lesson you can

Objectives

  1. Explain why F(x)+x lets gradients bypass saturated layers
  2. Distinguish BasicBlock (2 convs) from Bottleneck (1x1→3x3→1x1)
  3. Choose pre-activation vs post-activation BatchNorm placement
  4. Compute the ~8x parameter reduction of depthwise separable convolutions
  5. Apply EfficientNet compound scaling (depth, width, resolution together)

3. What survived from 2D Convolutions & CNNs?

Warm-up

Discussion prompt

Before we open Lesson 56: ResNet, Depthwise Separable Convolutions & EfficientNet: without looking back, what was the main idea of 2D Convolutions & CNNs, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

2D convolution from scratch (output shape formula, parameter sharing, translation equivariance), max pooling, receptive field growth, and a complete CNN forward pass in PyTorch (nn.Conv2d, nn.MaxPool2d, nn.ReLU). Verified on sklearn digits reaching 97% test accuracy at 100 epochs.

4. The identity shortcut — F(x) + x

Section

Section 1 of 4

5. The vanishing-gradient problem

Concept

In a plain 20-layer net, the gradient of the loss w.r.t. layer-0 weights is the product of 20 Jacobians. Each multiplication can shrink it exponentially.

Measured: a 20-layer net without skip connections has gradient norm 0.000001 at layer 0. The same net with identity shortcuts: 3760 — a factor of ~5 billion. (Lesson 35 showed this for general deep nets.)

architecturegrad norm at layer 0
20-layer net, no skip0.000001
20-layer net, with skip3760.16

6. Fill in: grad norm at layer 0 for The vanishing-gradient problem

Comparison

Comparison matrix

From The vanishing-gradient problem: refill the grad norm at layer 0 column from what you know. The rest of the table is as it appeared.

architecturegrad norm at layer 0
20-layer net, no skip0.000001
20-layer net, with skip3760.16

7. The residual shortcut: F(x) + x

Concept

A residual block asks the layers to learn the residual F(x) = H(x) - x rather than H(x) directly. The identity path x is always available.

\[ \text{output} = F(x, \{W_i\}) + x \]

During backprop the shortcut carries gradient directly to earlier layers — no matrices to multiply through. If the block is unhelpful, the weights can learn F(x)=0 and output is just x.

8. Break it if you can: The residual shortcut: F(x) + x

Counterexample

Discussion prompt

A residual block asks the layers to learn the residual F(x) = H(x) - x rather than H(x) directly. The identity path x is always available.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

9. Guess the shape of the answer: Gradient norm: skip vs no-skip

Estimation

Predict first

Build two 20-layer networks with identical initialization; measure the gradient norm at layer 0 after one backward pass.

Commit before you compute: what does Gradient norm: skip vs no-skip come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: no_skip grad norm: 0.000001 — skip grad norm: 3760.16

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The shortcut carries the upstream gradient unattenuated; the learned path adds to it.

10. Gradient norm: skip vs no-skip

Worked example

Build two 20-layer networks with identical initialization; measure the gradient norm at layer 0 after one backward pass.

import torch
import torch.nn as nn
torch.manual_seed(42)

class DeepNoSkip(nn.Module):
    def __init__(self):
        super().__init__()
        self.layers = nn.ModuleList([nn.Linear(64,64) for _ in range(20)])
    def forward(self, x):
        for l in self.layers: x = torch.relu(l(x))
        return x.sum()

class DeepSkip(nn.Module):
    def __init__(self):
        super().__init__()
        self.layers = nn.ModuleList([nn.Linear(64,64) for _ in range(20)])
    def forward(self, x):
        for l in self.layers: x = torch.relu(l(x)) + x  # F(x)+x
        return x.sum()

x = torch.randn(1, 64)
for name, net in [('no_skip', DeepNoSkip()), ('skip', DeepSkip())]:
    net(x).backward()
    g = net.layers[0].weight.grad.norm().item()
    print(f'{name}: {g:.6f}')

no_skip grad norm: 0.000001 — skip grad norm: 3760.16

Why: The shortcut carries the upstream gradient unattenuated; the learned path adds to it. Early layers receive usable gradient signal even at depth 20.

variantgrad norm at layer 0trainable?
no skip0.000001no — vanished
with skip (F(x)+x)3760.16yes — flows

11. What each one costs: Gradient norm: skip vs no-skip

Trade off

Comparison matrix

From Gradient norm: skip vs no-skip: every row here is a choice with a cost. Fill the grad norm at layer 0 column, then say which row you would actually pick and what you give up for it.

variantgrad norm at layer 0trainable?
no skip0.000001no — vanished
with skip (F(x)+x)3760.16yes — flows

12. BasicBlock vs Bottleneck

Section

Section 2 of 4

13. Two residual block designs

Concept

blockstructureparams (C=64 in/out)used in
BasicBlock3x3 -> 3x373,984ResNet-18, ResNet-34
Bottleneck1x1 -> 3x3 -> 1x1136,448 (256->64->256)ResNet-50, -101, -152

Bottleneck uses 1x1 convs to compress channels before and expand after the expensive 3x3 — fewer multiplications for the same representational depth.

14. By analogy: Two residual block designs

Analogy

Discussion prompt

Explain Two residual block designs by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Bottleneck uses 1x1 convs to compress channels before and expand after the expensive 3x3 — fewer multiplications for the same representational depth.

15. Guess the shape of the answer: BasicBlock and Bottleneck param counts

Estimation

Predict first

Measure parameters for a BasicBlock(64,64) and a Bottleneck(256->64->256) using nn.Conv2d layers.

Commit before you compute: what does BasicBlock and Bottleneck param counts come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: BasicBlock(64): 73,984 params; Bottleneck(256->64->256): 136,448 params

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Bottleneck costs more parameters here, but operates on 256-channel feature maps — at depth, the 1x1 bottleneck is far cheaper than a plain 3x3 on 256 channels (3x3x256x256 = 589,824).

16. BasicBlock and Bottleneck param counts

Worked example

Measure parameters for a BasicBlock(64,64) and a Bottleneck(256->64->256) using nn.Conv2d layers.

import torch.nn as nn

class BasicBlock(nn.Module):
    def __init__(self, c):
        super().__init__()
        self.net = nn.Sequential(
            nn.Conv2d(c,c,3,padding=1,bias=False), nn.BatchNorm2d(c),
            nn.ReLU(), nn.Conv2d(c,c,3,padding=1,bias=False), nn.BatchNorm2d(c))
    def forward(self,x): return torch.relu(self.net(x) + x)

class Bottleneck(nn.Module):
    def __init__(self, c_in=256, mid=64):
        super().__init__()
        c_out = mid * 4
        self.net = nn.Sequential(
            nn.Conv2d(c_in,mid,1,bias=False), nn.BatchNorm2d(mid), nn.ReLU(),
            nn.Conv2d(mid,mid,3,padding=1,bias=False), nn.BatchNorm2d(mid), nn.ReLU(),
            nn.Conv2d(mid,c_out,1,bias=False), nn.BatchNorm2d(c_out))
        self.proj = nn.Conv2d(c_in,c_out,1,bias=False)
    def forward(self,x): return torch.relu(self.net(x) + self.proj(x))

bb = BasicBlock(64)
bn = Bottleneck(256, 64)
print('BasicBlock(64):', sum(p.numel() for p in bb.parameters()))
print('Bottleneck(256->64->256):', sum(p.numel() for p in bn.parameters()))

BasicBlock(64): 73,984 params; Bottleneck(256->64->256): 136,448 params

Why: Bottleneck costs more parameters here, but operates on 256-channel feature maps — at depth, the 1x1 bottleneck is far cheaper than a plain 3x3 on 256 channels (3x3x256x256 = 589,824).

blockparams3x3 on full channels
BasicBlock(64)73,984—
Bottleneck(256->64->256)136,448vs 589,824 for 3x3x256x256

17. Work backwards from the answer: BasicBlock and Bottleneck param counts

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

BasicBlock(64): 73,984 params; Bottleneck(256->64->256): 136,448 params

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Measure parameters for a BasicBlock(64,64) and a Bottleneck(256->64->256) using nn.Conv2d layers.

18. BatchNorm placement: pre-act vs post-act

Concept

variantordershortcut pathpaper
post-activation (v1)Conv -> BN -> ReLUmay have BN on shortcutHe et al. 2016 (original)
pre-activation (v2)BN -> ReLU -> Convclean identity (no BN/ReLU)He et al. 2016 (revised)

Pre-activation means the shortcut is a clean identity with no activation — the gradient highway is truly unobstructed. Pre-act also slightly improves accuracy on very deep (1000+ layer) networks.

19. Teach it back: BatchNorm placement: pre-act vs post-act

Explain it

Discussion prompt

Explain BatchNorm placement: pre-act vs post-act to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Pre-activation means the shortcut is a clean identity with no activation — the gradient highway is truly unobstructed. Pre-act also slightly improves accuracy on very deep (1000+ layer) networks.

20. Something is wrong here: projection shortcut shape mismatch

Anomaly

Predict first

A student writes this, and it looks reasonable:

When stride=2 halves the spatial size (e.g. 56x56->28x28), use the identity shortcut directly: out = F(x) + x.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: F(x) has shape (B, 128, 28, 28); x has shape (B, 64, 56, 56).

When stride or channel count changes, apply a projection shortcut (1x1 conv + BN) to match dimensions before adding.

Why: F(x) has shape (B, 128, 28, 28); x has shape (B, 64, 56, 56). PyTorch raises a RuntimeError on the addition — the shapes don't broadcast.

21. Trap: projection shortcut shape mismatch

Trap

The trap

When stride=2 halves the spatial size (e.g. 56x56->28x28), use the identity shortcut directly: out = F(x) + x.

Add F(x) + x when channels or stride changes

Why: F(x) has shape (B, 128, 28, 28); x has shape (B, 64, 56, 56). PyTorch raises a RuntimeError on the addition — the shapes don't broadcast.

The fix

When stride or channel count changes, apply a projection shortcut (1x1 conv + BN) to match dimensions before adding.

shortcut = nn.Sequential(Conv2d(64,128,1,stride=2,bias=False), BN(128)); out = F(x) + shortcut(x)

Why: The 1x1 conv reshapes x to (B,128,28,28) so the addition is well-defined. This is the 'option B' shortcut from the ResNet paper.

22. Break it on purpose: projection shortcut shape mismatch

Break the constraint

Discussion prompt

The rule this trap just fixed:

The 1x1 conv reshapes x to (B,128,28,28) so the addition is well-defined. This is the 'option B' shortcut from the ResNet paper.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

F(x) has shape (B, 128, 28, 28); x has shape (B, 64, 56, 56). PyTorch raises a RuntimeError on the addition — the shapes don't broadcast.

23. Depthwise separable convolutions

Section

Section 3 of 4

24. Factoring a standard convolution

Concept

A standard K×K conv with C_in input channels and C_out output channels costs K²·C_in·C_out multiplications per spatial location.

MobileNet splits this into two cheaper ops: a depthwise conv (K²·C_in, one filter per channel) followed by a pointwise 1×1 conv (C_in·C_out, mixes channels).

\[ \frac{K^2 C_{in} + C_{in} C_{out}}{K^2 C_{in} C_{out}} = \frac{1}{C_{out}} + \frac{1}{K^2} \]

25. By analogy: Factoring a standard convolution

Analogy

Discussion prompt

Explain Factoring a standard convolution by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

A standard K×K conv with C_in input channels and C_out output channels costs K²·C_in·C_out multiplications per spatial location.

26. Guess the shape of the answer: Param reduction: depthwise sep vs standard

Estimation

Predict first

For C_in=32, C_out=64, K=3: compare a standard nn.Conv2d(32,64,3) against groups=C_in depthwise + 1x1 pointwise.

Commit before you compute: what does Param reduction: depthwise sep vs standard come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Standard: 18,432 params; DW+PW: 2,336 params; reduction 7.89x

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Depthwise (288 params) handles spatial filtering per-channel; pointwise (2,048) mixes across channels.

27. Param reduction: depthwise sep vs standard

Worked example

For C_in=32, C_out=64, K=3: compare a standard nn.Conv2d(32,64,3) against groups=C_in depthwise + 1x1 pointwise.

import torch.nn as nn

C_in, C_out, K = 32, 64, 3
std  = nn.Conv2d(C_in, C_out, K, padding=1, bias=False)
dw   = nn.Conv2d(C_in, C_in, K, padding=1, groups=C_in, bias=False)  # depthwise
pw   = nn.Conv2d(C_in, C_out, 1, bias=False)                          # pointwise

std_p  = sum(p.numel() for p in std.parameters())
dwp_p  = sum(p.numel() for p in list(dw.parameters()) + list(pw.parameters()))
print(f'Standard:    {std_p:,}')
print(f'DW+PW:       {dwp_p:,}')
print(f'Reduction:   {std_p/dwp_p:.2f}x')

Standard: 18,432 params; DW+PW: 2,336 params; reduction 7.89x

Why: Depthwise (288 params) handles spatial filtering per-channel; pointwise (2,048) mixes across channels. Together they replicate the 3x3 conv's function at 12.7% of the cost.

C_inC_outKStandardDW+PWRatio
33238641237.02x
3264318,4322,3367.89x
64128373,7288,7688.41x
1282563294,91233,9208.69x

28. What stays fixed: Param reduction: depthwise sep vs standard

Invariant

Step through it

Step through Param reduction: depthwise sep vs standard one row at a time. One of these columns never changes — find it, and say why it cannot.

  1. Step 1: C_in is 3
  2. Step 2: C_in is 32
  3. Step 3: C_in is 64
  4. Step 4: C_in is 128

29. Something is wrong here: depthwise groups != C_in

Anomaly

Predict first

A student writes this, and it looks reasonable:

Use nn.Conv2d(32, 32, 3, groups=1) for the depthwise step — groups=1 is the default.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: groups=1 is a fully-connected standard conv — every output channel sees every input channel.

Set groups=C_in so each filter handles exactly one input channel.

Why: groups=1 is a fully-connected standard conv — every output channel sees every input channel. It costs 3x3x32x32 = 9,216 params, not 288. You've negated the factorization entirely.

30. Trap: depthwise groups != C_in

Trap

The trap

Use nn.Conv2d(32, 32, 3, groups=1) for the depthwise step — groups=1 is the default.

groups=1 on a depthwise layer

Why: groups=1 is a fully-connected standard conv — every output channel sees every input channel. It costs 3x3x32x32 = 9,216 params, not 288. You've negated the factorization entirely.

The fix

Set groups=C_in so each filter handles exactly one input channel.

nn.Conv2d(32, 32, 3, padding=1, groups=32, bias=False)

Why: groups=C_in means the conv has C_in separate filter banks of size 1xKxK. Each channel is filtered independently — that is the depthwise step. Params = KKC_in = 9*32 = 288.

31. Which of these survive contact with Lesson 56: ResNet, Depthwise Separable…?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
In a plain 20-layer net, the gradient of the loss w.r.t. layer-0 weights is the product of 20 Jacobians. Each multiplication can shrink it exponentially.; A residual block asks the layers to learn the residual F(x) = H(x) - x rather than H(x) directly. The identity path x is always available.; Bottleneck uses 1x1 convs to compress channels before and expand after the expensive 3x3 — fewer multiplications for the same representational depth.
Breaks
When stride=2 halves the spatial size (e.g. 56x56->28x28), use the identity shortcut directly: out = F(x) + x.; Use nn.Conv2d(32, 32, 3, groups=1) for the depthwise step — groups=1 is the default.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 56: ResNet, Depthwise Separable Convolutions & EfficientNet puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

32. EfficientNet compound scaling

Section

Section 4 of 4

33. Why scale all three dimensions together

Concept

Scaling only depth (more layers) hits diminishing returns because spatial resolution is a bottleneck. Scaling only width (more channels) also saturates. Scaling only resolution helps but needs more layers to process fine detail.

EfficientNet observes that optimal depth/width/resolution scale together under a fixed compute budget. It finds the optimal per-dimension coefficients once, then reuses them by raising them to a single composite exponent phi.

34. The compound scaling formula

Concept

\[ d = \alpha^{\phi},\quad w = \beta^{\phi},\quad r = \gamma^{\phi} \]

Grid-searched constants: alpha=1.2 (depth), beta=1.1 (width), gamma=1.15 (resolution). Constraint: alpha * beta^2 * gamma^2 = 1.92 (approx 2.0, so doubling phi doubles FLOPs).

variantphidepthwidthresolution
EfficientNet-B001.0x1.0x224
EfficientNet-B111.2x1.1x258
EfficientNet-B442.07x1.46x392
EfficientNet-B773.58x1.95x596

35. Fill in: width for The compound scaling formula

Comparison

Comparison matrix

From The compound scaling formula: refill the width column from what you know. The rest of the table is as it appeared.

variantphidepthwidthresolution
EfficientNet-B001.0x1.0x224
EfficientNet-B111.2x1.1x258
EfficientNet-B442.07x1.46x392
EfficientNet-B773.58x1.95x596

36. Rebuild the recipe: The modern-architecture recipe

Ranking

Put in order

These are the steps of The modern-architecture recipe, scrambled. Put them back in order before the next slide shows you.

  1. Skip connections: F(x)+x; use projection shortcut (1x1+BN) when stride or channels change
  2. Block choice: BasicBlock for 18/34; Bottleneck (1x1->3x3->1x1) for 50/101/152+
  3. BatchNorm placement: post-act (Conv->BN->ReLU) by default; pre-act for very deep nets
  4. Depthwise separable: replace 3x3 conv with (DW, groups=C_in) + (PW, 1x1); expect ~8x fewer params
  5. EfficientNet: compute alpha/beta/gamma once at phi=1; scale with phi for B0..B7

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

37. The modern-architecture recipe

Pattern

  1. Skip connections: F(x)+x; use projection shortcut (1x1+BN) when stride or channels change
  2. Block choice: BasicBlock for 18/34; Bottleneck (1x1->3x3->1x1) for 50/101/152+
  3. BatchNorm placement: post-act (Conv->BN->ReLU) by default; pre-act for very deep nets
  4. Depthwise separable: replace 3x3 conv with (DW, groups=C_in) + (PW, 1x1); expect ~8x fewer params
  5. EfficientNet: compute alpha/beta/gamma once at phi=1; scale with phi for B0..B7

38. Where does it stop working: The modern-architecture recipe

Edge cases

Discussion prompt

The modern-architecture recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Skip connections: F(x)+x; use projection shortcut (1x1+BN) when stride or channels change
  2. Block choice: BasicBlock for 18/34; Bottleneck (1x1->3x3->1x1) for 50/101/152+
  3. BatchNorm placement: post-act (Conv->BN->ReLU) by default; pre-act for very deep nets
  4. Depthwise separable: replace 3x3 conv with (DW, groups=C_in) + (PW, 1x1); expect ~8x fewer params
  5. EfficientNet: compute alpha/beta/gamma once at phi=1; scale with phi for B0..B7

39. Rule out three: Check yourself — skip connection

Elimination

Eliminate the wrong options

In a 50-layer ResNet, a residual block transitions from 64 to 128 channels with stride=2. Which shortcut is required?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. A 1x1 Conv2d(64, 128, stride=2) followed by BatchNorm — a projection shortcut
  • B. The identity shortcut (no extra layer) — shapes broadcast automatically
  • C. A 3x3 Conv2d(64, 128, stride=2) to match shapes
  • D. Zero-pad the 64-channel tensor to 128 channels, then downsample

Survives elimination: A

Why: When channel count or stride changes, the shortcut must project x to match F(x). The standard (option B) from the ResNet paper is a 1x1 conv + BN that simultaneously changes channels and halves spatial size. The identity shortcut only works when shapes already match.

40. Check yourself — skip connection

Check

Pause and reason before choosing.

Check your understanding

In a 50-layer ResNet, a residual block transitions from 64 to 128 channels with stride=2. Which shortcut is required?

  • A. A 1x1 Conv2d(64, 128, stride=2) followed by BatchNorm — a projection shortcut (correct)
  • B. The identity shortcut (no extra layer) — shapes broadcast automatically
  • C. A 3x3 Conv2d(64, 128, stride=2) to match shapes
  • D. Zero-pad the 64-channel tensor to 128 channels, then downsample

Answer: A

Why: When channel count or stride changes, the shortcut must project x to match F(x). The standard (option B) from the ResNet paper is a 1x1 conv + BN that simultaneously changes channels and halves spatial size. The identity shortcut only works when shapes already match.

Why B tempts people
The identity shortcut has no parameters and cannot change the tensor shape. PyTorch would raise a RuntimeError when adding mismatched shapes.
Why C tempts people
A 3x3 projection shortcut adds extra spatial filtering to the skip path, introducing non-linearity into the 'gradient highway'. The 1x1 projection is preferred precisely because it keeps the skip path cheap and linear.
Why D tempts people
Zero-padding preserves old channel values and can work (ResNet option A), but it mixes zeros with real features — the 1x1 learned projection (option B) is empirically stronger.

41. Answer it before you see the options: Check yourself — depthwise sep reduction

Prediction

Predict first

A standard Conv2d(128, 256, 3) has 128x256x9 = 294,912 parameters. Replacing it with a depthwise-separable block gives approximately how many parameters?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: ~33,920 (ratio ~8.7x)

Why: Depthwise: K^2 * C_in = 9*128 = 1,152 params. Pointwise: C_in * C_out = 128*256 = 32,768. Total = 33,920. Ratio = 294,912 / 33,920 = 8.7x. Matches 1/256 + 1/9 = 0.115, so 8.7x.

42. Check yourself — depthwise sep reduction

Check

Use the formula: reduction = 1/C_out + 1/K^2.

Check your understanding

A standard Conv2d(128, 256, 3) has 128x256x9 = 294,912 parameters. Replacing it with a depthwise-separable block gives approximately how many parameters?

  • A. ~33,920 (ratio ~8.7x) (correct)
  • B. ~147,456 (ratio ~2x)
  • C. ~1,152 (ratio ~256x)
  • D. ~73,728 (ratio ~4x)

Answer: A

Why: Depthwise: K^2 * C_in = 9*128 = 1,152 params. Pointwise: C_in * C_out = 128*256 = 32,768. Total = 33,920. Ratio = 294,912 / 33,920 = 8.7x. Matches 1/256 + 1/9 = 0.115, so 8.7x.

Why B tempts people
A 2x reduction would come from halving either C_in or C_out — not from the depthwise-separable factorization. The formula gives ~8.7x not 2x for C_out=256, K=3.
Why C tempts people
1,152 is only the depthwise step; the pointwise 1x1 (32,768 params) must be added. Total is 33,920, not 1,152.
Why D tempts people
A 4x reduction corresponds roughly to K=2 (1/4 spatial) — for K=3, C_out=256 the reduction is 8.7x, not 4x.

43. Rule out three: Check yourself — EfficientNet scaling

Elimination

Eliminate the wrong options

EfficientNet uses alpha=1.2, beta=1.1, gamma=1.15. At phi=2, what is the depth multiplier (alpha^phi)?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. 1.44
  • B. 2.4
  • C. 1.2
  • D. 1.21

Survives elimination: A

Why: depth_mult = alpha^phi = 1.2^2 = 1.44. Width_mult = beta^2 = 1.1^2 = 1.21, resolution = 224 * 1.15^2 = 296. This is EfficientNet-B2.

44. Check yourself — EfficientNet scaling

Check

Apply the compound scaling formula.

Check your understanding

EfficientNet uses alpha=1.2, beta=1.1, gamma=1.15. At phi=2, what is the depth multiplier (alpha^phi)?

  • A. 1.44 (correct)
  • B. 2.4
  • C. 1.2
  • D. 1.21

Answer: A

Why: depth_mult = alpha^phi = 1.2^2 = 1.44. Width_mult = beta^2 = 1.1^2 = 1.21, resolution = 224 * 1.15^2 = 296. This is EfficientNet-B2.

Why B tempts people
2.4 = 1.2 * 2 — linear scaling, not exponential. EfficientNet uses alpha^phi, not alpha * phi.
Why C tempts people
1.2 = alpha^1, which would be phi=1 (B1). At phi=2 the multiplier is 1.2^2 = 1.44.
Why D tempts people
1.21 = 1.1^2 — that is the WIDTH multiplier (beta^2), not depth (alpha^2).

45. Your turn: build it

Section

Project

46. Project: ResNet-style block + depthwise-sep block

Concept

Build a BasicBlock, a Bottleneck, and a depthwise-separable block. Count parameters and verify skip connection requires projection when shapes change.

#requirementkey tool
1BasicBlock(64,64) with correct F(x)+xnn.Conv2d, nn.BatchNorm2d
2Bottleneck(256->64->256) with projection shortcut1x1 then 3x3 then 1x1
3DW-sep block — confirm 7.89x reduction for (32,64,3)groups=C_in + 1x1

Rule: run every block with a synthetic input to confirm output shapes match, then print param counts.

47. Break it if you can: Project: ResNet-style block + depthwise-sep block

Counterexample

Discussion prompt

Build a BasicBlock, a Bottleneck, and a depthwise-separable block. Count parameters and verify skip connection requires projection when shapes change.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Rule: run every block with a synthetic input to confirm output shapes match, then print param counts.

48. Milestone 1 — BasicBlock with skip

Worked example

Your turn: implement BasicBlock(64). Verify the skip path adds x unchanged. What is the param count?

Hint: two Conv2d(64,64,3,padding=1,bias=False) + BN; shortcut is nn.Identity() since channels and stride match.

import torch, torch.nn as nn

class BasicBlock(nn.Module):
    def __init__(self, c):
        super().__init__()
        self.conv1 = nn.Conv2d(c, c, 3, padding=1, bias=False)
        self.bn1   = nn.BatchNorm2d(c)
        self.conv2 = nn.Conv2d(c, c, 3, padding=1, bias=False)
        self.bn2   = nn.BatchNorm2d(c)
        self.relu  = nn.ReLU(inplace=True)
    def forward(self, x):
        out = self.relu(self.bn1(self.conv1(x)))
        out = self.bn2(self.conv2(out))
        return self.relu(out + x)        # F(x) + x

bb = BasicBlock(64)
x  = torch.randn(2, 64, 8, 8)
print('output shape:', tuple(bb(x).shape))
print('params:', sum(p.numel() for p in bb.parameters()))
output shapeparam count
(2, 64, 8, 8)73,984

49. Milestone 2 — depthwise-separable block

Worked example

Your turn: implement a DW-sep block for (C_in=32, C_out=64, K=3). Predict the param count before running.

Hint: groups=C_in in the depthwise layer; pointwise is Conv2d(C_in, C_out, 1, bias=False). Expected: 2,336.

import torch, torch.nn as nn

class DWSepConv(nn.Module):
    def __init__(self, c_in, c_out, k=3):
        super().__init__()
        self.dw = nn.Conv2d(c_in, c_in, k, padding=k//2, groups=c_in, bias=False)
        self.pw = nn.Conv2d(c_in, c_out, 1, bias=False)
        self.bn = nn.BatchNorm2d(c_out)
        self.relu = nn.ReLU(inplace=True)
    def forward(self, x): return self.relu(self.bn(self.pw(self.dw(x))))

dws = DWSepConv(32, 64, 3)
x   = torch.randn(1, 32, 16, 16)
print('output shape:', tuple(dws(x).shape))
print('DW params:', sum(p.numel() for p in dws.dw.parameters()))
print('PW params:', sum(p.numel() for p in dws.pw.parameters()))
print('Total DW-sep:', sum(p.numel() for p in dws.parameters()))
layerparams
depthwise (groups=32)288
pointwise (1x1)2,048
BN (64)128
total DW-sep2,464

50. Milestone 3 — ResNet-18 vs ResNet-50 at scale

Worked example

Your turn: build minimal ResNet-18 and ResNet-50 skeletons (no training). Count parameters and compare.

Hint: ResNet-18 uses 8 BasicBlocks; ResNet-50 uses 16 Bottlenecks. Both end with nn.Linear(512/2048, 1000).

# (Using pre-built torchvision for a correct count)
import torch
from torchvision.models import resnet18, resnet50
r18 = resnet18(); r50 = resnet50()
p18 = sum(p.numel() for p in r18.parameters())
p50 = sum(p.numel() for p in r50.parameters())
print(f'ResNet-18: {p18:,}  ({p18/1e6:.1f}M)')
print(f'ResNet-50: {p50:,}  ({p50/1e6:.1f}M)')
modelblocksparamsnotes
ResNet-188 BasicBlocks11.7M2x3x3 per block
ResNet-5016 Bottlenecks25.6M1x1->3x3->1x1 per block

51. What each one costs: Milestone 3 — ResNet-18 vs ResNet-50 at scale

Trade off

Comparison matrix

From Milestone 3 — ResNet-18 vs ResNet-50 at scale: every row here is a choice with a cost. Fill the notes column, then say which row you would actually pick and what you give up for it.

modelblocksparamsnotes
ResNet-188 BasicBlocks11.7M2x3x3 per block
ResNet-5016 Bottlenecks25.6M1x1->3x3->1x1 per block

52. The full program

Concept

import torch, torch.nn as nn

# BasicBlock
class BB(nn.Module):
    def __init__(self, c): super().__init__(); self.net = nn.Sequential(nn.Conv2d(c,c,3,padding=1,bias=False),nn.BatchNorm2d(c),nn.ReLU(),nn.Conv2d(c,c,3,padding=1,bias=False),nn.BatchNorm2d(c))
    def forward(self, x): return torch.relu(self.net(x) + x)

# DW-Sep
class DW(nn.Module):
    def __init__(self, ci, co): super().__init__(); self.dw=nn.Conv2d(ci,ci,3,padding=1,groups=ci,bias=False); self.pw=nn.Conv2d(ci,co,1,bias=False)
    def forward(self, x): return self.pw(self.dw(x))

std = nn.Conv2d(32, 64, 3, padding=1, bias=False)
dw  = DW(32, 64)

print('BasicBlock(64) params:', sum(p.numel() for p in BB(64).parameters()))
print('Std conv(32->64,3):', sum(p.numel() for p in std.parameters()))
print('DW-sep(32->64,3):', sum(p.numel() for p in dw.parameters()))
print('Reduction:', round(sum(p.numel() for p in std.parameters()) / sum(p.numel() for p in dw.parameters()), 2), 'x')
blockparamsnote
BasicBlock(64)73,984F(x)+x, identity shortcut
Std conv(32->64,K=3)18,432full spatial mix
DW-sep(32->64,K=3)2,336factored: 288+2048
Reduction7.89xmatches 1/64 + 1/9

If your numbers match the table — you've implemented the building blocks of MobileNet and ResNet from scratch.

53. Fill in: note for The full program

Comparison

Comparison matrix

From The full program: refill the note column from what you know. The rest of the table is as it appeared.

blockparamsnote
BasicBlock(64)73,984F(x)+x, identity shortcut
Std conv(32->64,K=3)18,432full spatial mix
DW-sep(32->64,K=3)2,336factored: 288+2048
Reduction7.89xmatches 1/64 + 1/9

54. Show it off

Concept

Out loud, slides closed: (1) explain the vanishing-gradient problem and how F(x)+x fixes it; (2) derive the depthwise-sep reduction formula 1/C_out + 1/K^2; (3) describe EfficientNet's phi parameter and why the constraint alphabeta^2gamma^2 ~= 2 matters.

Stretch (homework from the lesson plan): implement a full ResNet-18-style network in PyTorch and train on CIFAR-10; compare accuracy with and without skip connections at 20 epochs. Next up: attention mechanisms and the Transformer architecture.

55. Connect it up: Lesson 56: ResNet, Depthwise Separable Convolutions & EfficientNet

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — The identity shortcut — F(x) + x · BasicBlock vs Bottleneck · Depthwise separable convolutions · EfficientNet compound scaling · Your turn: build it. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

56. What you can do now

Recap

ideathe one thing to remember
skip connectionF(x)+x; projection shortcut when shapes differ
BasicBlock vs BottleneckBasicBlock: 2x3x3; Bottleneck: 1x1->3x3->1x1
BN placementpost-act (v1) vs pre-act (v2, cleaner shortcut)
depthwise sepgroups=C_in dw + 1x1 pw; ~8x fewer params
EfficientNetalpha^phi * beta^phi * gamma^phi; constraint ~= 2

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 56 (ResNet identity shortcut, depthwise separable, EfficientNet compound scaling) — Barron · USAAIO Round 2 Preparation, 2026
  2. ResNet-18 (11.7M params), ResNet-50 (25.7M), depthwise-sep ratio (7.89x for C_in=32->64 K=3), gradient norms with/without skip, and EfficientNet scaling table verified — torch 2.7.1+cpu, numpy 2.2.6, real execution, June 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108