Lesson 44: MLP Architectures — Depth, Width & Activations

USAAIO Lesson 44, from Week 15 of Phase 2, fully worked. One running 2→3→1 MLP is hand-traced neuron by neuron and confirmed against torch. It states the universal approximation theorem and demonstrates it on XOR, showing a single Linear layer fail, then covers depth against width as expressivity by counting linear regions. It derives the exact nn.Linear parameter formula and sums it layer by layer, computes all four activations - sigmoid, tanh, ReLU, and GELU - at five inputs with their gradient consequences, and rebuilds GELU from erf. It quantifies the vanishing-gradient and dying-ReLU failure modes, runs a fair wide-against-deep experiment at about 38k parameters on load_digits, and ends with a from-scratch build-it project. Every snippet runs standalone in a fresh interpreter, and every number came from real execution with torch 2.7.1 and numpy 2.2.6. The lesson runs to 63 slides.

Subject: Machine Learning · 97 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. MLP Architectures: Depth, Width & Activations

Title

USAAIO · Lesson 44 · Week 15 (Phase 2)

We trace one tiny network by hand — every neuron, every weight — then scale the same ideas up: universal approximation, why depth beats width, exact parameter counting, and the activation math that decides whether a deep net learns at all.

2. By the end of this lesson you can

Objectives

  1. Forward-propagate an MLP by hand — pre-activation, activation, next layer — and confirm it against torch
  2. State the universal approximation theorem, its existence-only catch, and demonstrate it on XOR
  3. Explain why depth buys exponentially many linear regions while width only interpolates
  4. Compute exact parameter counts with d_in × d_out + d_out and sum any nn.Linear stack
  5. Compare sigmoid / tanh / ReLU / GELU on outputs and gradients, and name the vanishing-gradient and dying-ReLU failures
  6. Run a fair wide-vs-deep experiment and apply USAAIO architecture-design rules

3. What survived from Support Vector Machines?

Warm-up

Discussion prompt

Before we open Lesson 44: MLP Architectures — Depth, Width & Activations: without looking back, what was the main idea of Support Vector Machines, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

hard-margin SVM geometric derivation, soft-margin slack variables and the role of C, kernel trick (RBF, polynomial), hyperparameter grid search with cross-validation, multi-class OvO vs OvR strategies, and SVM vs logistic regression trade-offs. Build and evaluate SVC on Iris and moons; verified with sklearn and real execution numbers.

4. What an MLP actually computes

Section

Part 1 of 8 — one running example

5. An MLP is stacked affine-then-nonlinear

Concept

A multilayer perceptron alternates two moves: an affine map z = Wh + b (a matrix multiply plus a bias), then a nonlinear activation h' = σ(z) applied element-wise. Repeat for each layer.

\[ h^{(0)} = x, \qquad z^{(\ell)} = W^{(\ell)} h^{(\ell-1)} + b^{(\ell)}, \qquad h^{(\ell)} = \sigma\!\big(z^{(\ell)}\big) \]

The activation σ is the only nonlinear part. Strip it out and the whole stack collapses to a single affine map — which is exactly the trap we'll prove on XOR.

6. Our running network: 2 → 3 → 1

Concept

We fix one concrete network for the whole lesson: input dimension 2, a hidden layer of 3 ReLU neurons, and a single linear output. Every part of this deck comes back to it.

\[ W^{(1)} = \begin{bmatrix} 1 & -1 \\ 0.5 & 0.5 \\ -2 & 1 \end{bmatrix},\; b^{(1)} = \begin{bmatrix} 0 \\ -1 \\ 0.5 \end{bmatrix},\; W^{(2)} = \begin{bmatrix} 1 & -1 & 2 \end{bmatrix},\; b^{(2)} = \begin{bmatrix} 0.5 \end{bmatrix} \]

We feed it the input x = [1, 2]. Row j of W⁽¹⁾ holds the two weights of hidden neuron j. Keep this table of numbers in view — the next four slides trace it end to end.

7. Break it if you can: Our running network: 2 → 3 → 1

Counterexample

Discussion prompt

We fix one concrete network for the whole lesson: input dimension 2, a hidden layer of 3 ReLU neurons, and a single linear output. Every part of this deck comes back to it.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

We feed it the input x = [1, 2]. Row j of W⁽¹⁾ holds the two weights of hidden neuron j. Keep this table of numbers in view — the next four slides trace it end to end.

8. Each hidden neuron is a little detector

Intuition

Think of every hidden neuron as a feature detector. It takes a weighted vote of the inputs (Wⱼ·x + bⱼ), and the activation decides how loudly it fires — ReLU stays silent (0) until its vote goes positive, then passes the vote straight through.

The output layer is a weighted committee of those detectors. Learning means tuning who votes, how strongly each detector responds, and how the committee combines them.

That is the whole mental model. Now let's turn the crank on real numbers — no hand-waving, one arithmetic step per neuron.

9. By analogy: Each hidden neuron is a little detector

Analogy

Discussion prompt

Explain Each hidden neuron is a little detector by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

The output layer is a weighted committee of those detectors. Learning means tuning who votes, how strongly each detector responds, and how the committee combines them.

10. What has to happen first: Forward pass, step 1 — the pre-activations z⁽¹⁾

Ranking

Put in order

Put the moves of Forward pass, step 1 — the pre-activations z⁽¹⁾ into the order they have to happen.

  1. Neuron 0: (1)(1) + (−1)(2) + 0 = 1 − 2 = −1
  2. Neuron 1: (0.5)(1) + (0.5)(2) − 1 = 0.5 + 1 − 1 = 0.5
  3. Neuron 2: (−2)(1) + (1)(2) + 0.5 = −2 + 2 + 0.5 = 0.5

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Row 0 of W⁽¹⁾ is [1, −1], bias 0.

11. Forward pass, step 1 — the pre-activations z⁽¹⁾

Worked example

First layer: z⁽¹⁾ = W⁽¹⁾x + b⁽¹⁾. Each hidden neuron dots its weight row with x = [1, 2] and adds its bias. Do all three, one at a time:

Neuron 0: (1)(1) + (−1)(2) + 0 = 1 − 2 = −1

Why: Row 0 of W⁽¹⁾ is [1, −1], bias 0. This vote came out negative — ReLU will silence it.

import numpy as np
x  = np.array([1.0, 2.0])
W1 = np.array([[ 1.0, -1.0],
               [ 0.5,  0.5],
               [-2.0,  1.0]])
b1 = np.array([0.0, -1.0, 0.5])
z1 = W1 @ x + b1
print(z1)   # [-1.  0.5  0.5]

Neuron 1: (0.5)(1) + (0.5)(2) − 1 = 0.5 + 1 − 1 = 0.5

Why: Row 1 is [0.5, 0.5], bias −1. Positive vote — survives ReLU.

Neuron 2: (−2)(1) + (1)(2) + 0.5 = −2 + 2 + 0.5 = 0.5

Why: Row 2 is [−2, 1], bias 0.5. Also positive. So z⁽¹⁾ = [−1, 0.5, 0.5], matching the printed array exactly.

neuron jrow WⱼWⱼ·x+ bⱼz⁽¹⁾ⱼ
0[1, −1]1 − 2 = −1+0−1.0
1[0.5, 0.5]0.5 + 1 = 1.5−10.5
2[−2, 1]−2 + 2 = 0+0.50.5

12. Fill in: row Wⱼ for Forward pass, step 1 — the pre-activations…

Comparison

Comparison matrix

From Forward pass, step 1 — the pre-activations z⁽¹⁾: refill the row Wⱼ column from what you know. The rest of the table is as it appeared.

neuron jrow WⱼWⱼ·x+ bⱼz⁽¹⁾ⱼ
0[1, −1]1 − 2 = −1+0−1.0
1[0.5, 0.5]0.5 + 1 = 1.5−10.5
2[−2, 1]−2 + 2 = 0+0.50.5

13. Guess the shape of the answer: Forward pass, step 2 — ReLU gives h⁽¹⁾

Estimation

Predict first

Apply ReLU(z) = max(0, z) element-wise to z⁽¹⁾ = [−1, 0.5, 0.5]. ReLU clamps negatives to zero and leaves positives untouched:

Commit before you compute: what does Forward pass, step 2 — ReLU gives h⁽¹⁾ come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Neurons 1, 2: max(0, 0.5) = 0.5 each → pass through

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Positive pre-activations are unchanged, so h⁽¹⁾ = [0, 0.5, 0.5].

14. Forward pass, step 2 — ReLU gives h⁽¹⁾

Worked example

Apply ReLU(z) = max(0, z) element-wise to z⁽¹⁾ = [−1, 0.5, 0.5]. ReLU clamps negatives to zero and leaves positives untouched:

import numpy as np
z1 = np.array([-1.0, 0.5, 0.5])
h = np.maximum(0.0, z1)   # ReLU
print(h)   # [0.  0.5 0.5]

Neuron 0: max(0, −1) = 0 → dead on this input

Why: Its negative vote is silenced. This neuron contributes nothing to the output for x = [1, 2] — a live illustration of ReLU gating.

Neurons 1, 2: max(0, 0.5) = 0.5 each → pass through

Why: Positive pre-activations are unchanged, so h⁽¹⁾ = [0, 0.5, 0.5].

neuron jz⁽¹⁾ⱼReLU(z⁽¹⁾ⱼ)state
0−1.00.0silenced
10.50.5firing
20.50.5firing

15. What each one costs: Forward pass, step 2 — ReLU gives h⁽¹⁾

Trade off

Comparison matrix

From Forward pass, step 2 — ReLU gives h⁽¹⁾: every row here is a choice with a cost. Fill the state column, then say which row you would actually pick and what you give up for it.

neuron jz⁽¹⁾ⱼReLU(z⁽¹⁾ⱼ)state
0−1.00.0silenced
10.50.5firing
20.50.5firing

16. What has to be given first: Forward pass, step 3 — the output z⁽²⁾

Missing information

Discussion prompt

Output layer is linear: z⁽²⁾ = W⁽²⁾h⁽¹⁾ + b⁽²⁾ with W⁽²⁾ = [1, −1, 2], b⁽²⁾ = 0.5, and h⁽¹⁾ = [0, 0.5, 0.5]. One dot product:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

Multiply each hidden output by its committee weight, sum, add the output bias.

17. Forward pass, step 3 — the output z⁽²⁾

Worked example

Output layer is linear: z⁽²⁾ = W⁽²⁾h⁽¹⁾ + b⁽²⁾ with W⁽²⁾ = [1, −1, 2], b⁽²⁾ = 0.5, and h⁽¹⁾ = [0, 0.5, 0.5]. One dot product:

(1)(0) + (−1)(0.5) + (2)(0.5) + 0.5

Why: Multiply each hidden output by its committee weight, sum, add the output bias.

import numpy as np
h  = np.array([0.0, 0.5, 0.5])
W2 = np.array([1.0, -1.0, 2.0])
b2 = 0.5
z2 = W2 @ h + b2
print(z2)   # 1.0

= 0 − 0.5 + 1.0 + 0.5 = 1.0

Why: The dead neuron drops out, neuron 1 subtracts 0.5, neuron 2 adds 1.0, bias adds 0.5. Network output for x = [1, 2] is exactly 1.0.

termwⱼ · hⱼvalue
neuron 01 × 00.0
neuron 1−1 × 0.5−0.5
neuron 22 × 0.5+1.0
bias—+0.5
output z⁽²⁾1.0

18. Inspect it line by line: Forward pass, step 3 — the output z⁽²⁾

Error analysis

Annotate

Walk the callouts on Forward pass, step 3 — the output z⁽²⁾. Each one is a place this is easy to get subtly wrong.

  • Multiply each hidden output by its committee weight, sum, add the output bias.
  • The dead neuron drops out, neuron 1 subtracts 0.5, neuron 2 adds 1.0, bias adds 0.5. Network output for x = [1, 2] is exactly 1.0.

19. Predict the next row: Confirm the hand trace against PyTorch

Pattern

Predict first

The table runs: hand trace (numpy) | 1.0 · nn.Sequential(torch) | 1.0

In Confirm the hand trace against PyTorch, given the rows so far: what is the next one — the row where source is match??

Correct: match? | yes

sourceoutput for x = [1, 2]
hand trace (numpy)1.0
nn.Sequential(torch)1.0
match?yes

Why: The relationship between the columns, not the individual numbers, is what generates the next row. nn.Linear stores weight as (out, in) and computes xWᵀ + b, which is exactly W⁽¹⁾x + b⁽¹⁾ for our row layout.

20. Confirm the hand trace against PyTorch

Worked example

Load the exact same weights into nn.Linear layers and run the network. It must print our hand-computed 1.0:

import torch, torch.nn as nn
x  = torch.tensor([1.0, 2.0])
lin1 = nn.Linear(2, 3)
lin1.weight.data = torch.tensor([[1.,-1.],[0.5,0.5],[-2.,1.]])
lin1.bias.data   = torch.tensor([0., -1., 0.5])
lin2 = nn.Linear(3, 1)
lin2.weight.data = torch.tensor([[1., -1., 2.]])
lin2.bias.data   = torch.tensor([0.5])
net = nn.Sequential(lin1, nn.ReLU(), lin2)
with torch.no_grad():
    print(net(x).numpy())   # [1.]

torch prints [1.] — identical to the hand trace

Why: nn.Linear stores weight as (out, in) and computes xWᵀ + b, which is exactly W⁽¹⁾x + b⁽¹⁾ for our row layout. The math and the library agree to the last digit.

sourceoutput for x = [1, 2]
hand trace (numpy)1.0
nn.Sequential(torch)1.0
match?yes

21. Work backwards from the answer: Confirm the hand trace against PyTorch

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

torch prints [1.] — identical to the hand trace

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Load the exact same weights into nn.Linear layers and run the network. It must print our hand-computed 1.0:

22. Universal approximation — and its catch

Section

Part 2 of 8 — what one hidden layer can do

23. The universal approximation theorem

Concept

A feedforward net with one hidden layer and a non-polynomial activation can approximate any continuous function on a compact domain, to any accuracy ε — given enough hidden neurons N.

\[ \forall\, f \in C([0,1]^d),\; \varepsilon > 0 \;\exists\; N,\, w,\, a,\, b :\; \sup_{x} \Big|\, f(x) - \sum_{i=1}^{N} w_i\, \sigma\!\big(a_i^{\top} x + b_i\big) \Big| < \varepsilon \]

It is a statement of existence (Cybenko 1989, Hornik 1991): the weights exist. It says nothing about how many neurons you need or whether gradient descent can find them.

24. Teach it back: The universal approximation theorem

Explain it

Discussion prompt

Explain The universal approximation theorem to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

A feedforward net with one hidden layer and a non-polynomial activation can approximate any continuous function on a compact domain, to any accuracy ε — given enough hidden neurons N.

25. The catch: 'enough' can be astronomically many

Intuition

'Enough neurons' can mean exponentially many. A single hidden layer approximates a wiggly function by stacking up many bump-shaped pieces — one small region at a time — so a function with lots of structure needs a huge, flat layer.

Depth is the escape hatch: by composing layers, a deep net reuses earlier features and represents the same function with far fewer parameters. Existence (one wide layer) and efficiency (depth) are different questions.

So the theorem is reassuring, not prescriptive. In practice you reach for depth. But first let's watch a single hidden layer actually solve a problem a linear model cannot.

26. XOR: one hidden layer solves it

Worked example

XOR is the classic not linearly separable problem: no straight line separates the 1s from the 0s. One hidden layer (width 4, tanh) fits it. Train and read off the predictions:

import torch, torch.nn as nn
torch.manual_seed(42)
X = torch.tensor([[0.,0.],[0.,1.],[1.,0.],[1.,1.]])
y = torch.tensor([[0.],[1.],[1.],[0.]])
model = nn.Sequential(nn.Linear(2,4), nn.Tanh(),
                      nn.Linear(4,1), nn.Sigmoid())
opt = torch.optim.Adam(model.parameters(), lr=0.1)
for _ in range(3000):
    opt.zero_grad()
    nn.MSELoss()(model(X), y).backward()
    opt.step()
with torch.no_grad():
    print(model(X).numpy().round(3).flatten())

Predictions [0.0, 0.998, 0.998, 0.003] round to [0, 1, 1, 0]

Why: The four hidden tanh units bend the input space so the output layer can separate the classes. Universal approximation made concrete: a hidden layer represents a function a single Linear cannot.

inputtargetpredictionrounds to
(0,0)00.0000 ✓
(0,1)10.9981 ✓
(1,0)10.9981 ✓
(1,1)00.0030 ✓

27. Which is which, by target

Discrimination

Sort into buckets

Sort these by target, from memory, without looking back at XOR: one hidden layer solves it. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

0
(0,0); (1,1)
1
(0,1); (1,0)
g1
target is "0" for (0,0), (1,1) — that is what the table on "XOR: one hidden layer solves it" records, and it is the single property separating this group from the rest.
g2
target is "1" for (0,1), (1,0) — that is what the table on "XOR: one hidden layer solves it" records, and it is the single property separating this group from the rest.

28. Something is wrong here: 'a bigger linear model will learn XOR'

Anomaly

Predict first

A student writes this, and it looks reasonable:

XOR just needs more capacity — drop the hidden layer and train a single Linear(2,1) harder.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: A Linear + Sigmoid is one linear boundary.

Add a nonlinear hidden layer. Nonlinearity — not size — is what defeats non-separability.

Why: A Linear + Sigmoid is one linear boundary. No line separates XOR, so the best it can do is sit on the fence at 0.5 for all four points. More epochs never fix a representational limit.

29. Trap: 'a bigger linear model will learn XOR'

Trap

The trap

XOR just needs more capacity — drop the hidden layer and train a single Linear(2,1) harder.

import torch, torch.nn as nn
torch.manual_seed(0)
X = torch.tensor([[0.,0.],[0.,1.],[1.,0.],[1.,1.]])
y = torch.tensor([[0.],[1.],[1.],[0.]])
model = nn.Sequential(nn.Linear(2,1), nn.Sigmoid())  # no hidden layer
opt = torch.optim.Adam(model.parameters(), lr=0.1)
for _ in range(3000):
    opt.zero_grad(); nn.BCELoss()(model(X), y).backward(); opt.step()
with torch.no_grad():
    print(model(X).numpy().round(3).flatten())  # [0.5 0.5 0.5 0.5]

Output [0.5, 0.5, 0.5, 0.5] — pure guessing

Why: A Linear + Sigmoid is one linear boundary. No line separates XOR, so the best it can do is sit on the fence at 0.5 for all four points. More epochs never fix a representational limit.

The fix

Add a nonlinear hidden layer. Nonlinearity — not size — is what defeats non-separability.

model = nn.Sequential(nn.Linear(2,4), nn.Tanh(),
                      nn.Linear(4,1), nn.Sigmoid())
# ... same training loop ...
# predictions -> [0.0, 0.998, 0.998, 0.003]

Hidden layer → [0, 1, 1, 0]

Why: Four tanh units warp the space until the classes are linearly separable in hidden coordinates. The fix is a nonlinearity between two Linear layers, exactly what the universal approximation theorem promises.

30. Break it on purpose: 'a bigger linear model will learn XOR'

Break the constraint

Discussion prompt

The rule this trap just fixed:

Add a nonlinear hidden layer. Nonlinearity — not size — is what defeats non-separability.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

A Linear + Sigmoid is one linear boundary. No line separates XOR, so the best it can do is sit on the fence at 0.5 for all four points. More epochs never fix a representational limit.

31. Depth vs width — why depth wins

Section

Part 3 of 8 — expressivity

32. Depth = composition; width = interpolation

Concept

A depth-L network composes L nonlinear maps: f = σ∘W_L ∘ … ∘ σ∘W_1. Each layer refines a feature hierarchy — edges from pixels, textures from edges, objects from textures.

strategymechanismpractical effect
one wide layerinterpolation over many bumpsmay need exponentially many neurons
many narrow layersfunction compositionexpressivity grows fast with depth
deep + moderate widthbothmost competition architectures

Rule of thumb: hierarchical targets (images, sequences) reward depth; a smooth scalar target with no hierarchy is fine with a moderately wide shallow net.

33. Counting linear regions: depth is exponential

Worked example

A ReLU network is piecewise linear — it carves input space into flat regions. A rough upper bound: one width-w layer makes ~w regions; composing L such layers can multiply them, giving ~wᴸ. Count it:

# ReLU regions grow ~w per layer when stacked -> ~w**L
w = 4
for L in [1, 2, 3, 4]:
    print(f'L={L} width {w}: up to ~{w**L} linear regions')

Width 4: 1 layer ≈ 4 regions, 4 layers ≈ 256 regions

Why: Adding a layer MULTIPLIES the region count; adding width only ADDS to it. That multiplicative-vs-additive gap is the formal reason depth is more parameter-efficient than width.

layers Lwᴸ (width 4)growth vs +width
14baseline
216×4
364×4
4256×4 (exponential)

34. Composition reuses work; width repeats it

Intuition

A wide shallow layer is like describing a face by listing every pixel pattern separately — no shared vocabulary. A deep net first learns 'edge', then reuses 'edge' everywhere to build 'eye', then reuses 'eye' to build 'face'.

Reuse is the whole advantage. The same features feed many higher-level features, so a deep net says far more with far fewer parameters — as long as the target really has that layered structure.

35. Where does each piece belong: Lesson 44: MLP Architectures — Depth, Width…

Sorting

Sort into buckets

These are the pieces of Lesson 44: MLP Architectures — Depth, Width & Activations, out of order. Put each one back under the part of the lesson it belongs to.

What an MLP actually computes
An MLP is stacked affine-then-nonlinear; Our running network: 2 → 3 → 1; Each hidden neuron is a little detector
Universal approximation — and its catch
The universal approximation theorem; The catch: 'enough' can be astronomically many; XOR: one hidden layer solves it
Depth vs width — why depth wins
Depth = composition; width = interpolation; Counting linear regions: depth is exponential; Composition reuses work; width repeats it
s1
What an MLP actually computes is where Lesson 44: MLP Architectures — Depth, Width & Activations puts An MLP is stacked affine-then-nonlinear, Our running network: 2 → 3 → 1, Each hidden neuron is a little detector. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
Universal approximation — and its catch is where Lesson 44: MLP Architectures — Depth, Width & Activations puts The universal approximation theorem, The catch: 'enough' can be astronomically many, XOR: one hidden layer solves it. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
Depth vs width — why depth wins is where Lesson 44: MLP Architectures — Depth, Width & Activations puts Depth = composition; width = interpolation, Counting linear regions: depth is exponential, Composition reuses work; width repeats it. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

36. Parameter counting — exact formula

Section

Part 4 of 8 — audit any architecture

37. Parameters of one nn.Linear

Concept

nn.Linear(d_in, d_out) holds a weight matrix of shape (d_out, d_in) and a bias vector of length d_out. Count both:

\[ \operatorname{params}\big(\text{Linear}(d_{\text{in}}, d_{\text{out}})\big) = \underbrace{d_{\text{in}}\, d_{\text{out}}}_{\text{weights}} + \underbrace{d_{\text{out}}}_{\text{bias}} \]

The + d_out is the bias — one scalar per output neuron. Activations (ReLU, Tanh, GELU, …) have zero parameters. A stack's total is the sum over its Linear layers.

38. What has to happen first: Params of our 2 → 3 → 1 network

Ranking

Put in order

Put the moves of Params of our 2 → 3 → 1 network into the order they have to happen.

  1. Layer 1 Linear(2, 3): 2×3 + 3 = 6 + 3 = 9
  2. Layer 2 Linear(3, 1): 3×1 + 1 = 3 + 1 = 4
  3. torch confirms total = 13

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. 6 weights (the 3×2 matrix) plus 3 biases (one per hidden neuron).

39. Params of our 2 → 3 → 1 network

Worked example

Apply the formula to the running network before trusting any library. Two Linear layers:

Layer 1 Linear(2, 3): 2×3 + 3 = 6 + 3 = 9

Why: 6 weights (the 3×2 matrix) plus 3 biases (one per hidden neuron).

Layer 2 Linear(3, 1): 3×1 + 1 = 3 + 1 = 4

Why: 3 weights (the committee) plus 1 output bias. Total = 9 + 4 = 13.

import torch.nn as nn
net = nn.Sequential(nn.Linear(2, 3), nn.ReLU(), nn.Linear(3, 1))
for name, p in net.named_parameters():
    print(name, p.numel())
print('total:', sum(p.numel() for p in net.parameters()))

torch confirms total = 13

Why: 0.weight=6, 0.bias=3, 2.weight=3, 2.bias=1 → 13. The ReLU at index 1 has no parameters, so it never appears in named_parameters().

tensorshapenumel
0.weight(3, 2)6
0.bias(3,)3
2.weight(1, 3)3
2.bias(1,)1
total13

40. Draw the shape of it: Params of our 2 → 3 → 1 network

Blank canvas

Draw it

Draw what Params of our 2 → 3 → 1 network just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.

41. Params of a real MLP: [64 → 128 → 64 → 10]

Worked example

Scale up to a digits-sized classifier (64 input features, 10 classes). Sum the three Linear layers:

import torch.nn as nn
model = nn.Sequential(nn.Linear(64,128), nn.ReLU(),
                      nn.Linear(128,64), nn.ReLU(),
                      nn.Linear(64,10))
for m in model:
    if isinstance(m, nn.Linear):
        print(m, sum(p.numel() for p in m.parameters()))
print('total:', sum(p.numel() for p in model.parameters()))

L1: 64×128+128 = 8320; L2: 128×64+64 = 8256; L3: 64×10+10 = 650

Why: Each layer by the formula. The activations add nothing; only the three Linear layers carry parameters.

Total = 8320 + 8256 + 650 = 17226

Why: torch prints 17226 — matching the hand sum exactly. You can now audit any architecture's size on paper.

layerformulaparams
Linear(64,128)64×128 + 1288 320
Linear(128,64)128×64 + 648 256
Linear(64,10)64×10 + 10650
TOTAL17 226

42. Draw the shape of it: Params of a real MLP: [64 → 128 → 64 → 10]

Blank canvas

Draw it

Draw what Params of a real MLP: [64 → 128 → 64 → 10] just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.

43. Something is wrong here: forgetting the bias in the count

Anomaly

Predict first

A student writes this, and it looks reasonable:

Linear(64, 128) is a 64×128 weight matrix, so 64 × 128 = 8192 parameters.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Misses the 128-element bias vector.

Add the bias: d_in × d_out + d_out = 64×128 + 128 = 8320.

Why: Misses the 128-element bias vector. That is an off-by-d_out error on EVERY layer; across a deep net it compounds into thousands of uncounted parameters and a wrong size budget.

44. Trap: forgetting the bias in the count

Trap

The trap

Linear(64, 128) is a 64×128 weight matrix, so 64 × 128 = 8192 parameters.

import torch.nn as nn
lin = nn.Linear(64, 128)
print(lin.weight.numel())   # 8192  (weights ONLY)

Count only the weight matrix → 8192

Why: Misses the 128-element bias vector. That is an off-by-d_out error on EVERY layer; across a deep net it compounds into thousands of uncounted parameters and a wrong size budget.

The fix

Add the bias: d_in × d_out + d_out = 64×128 + 128 = 8320.

import torch.nn as nn
lin = nn.Linear(64, 128)
w = lin.weight.numel(); b = lin.bias.numel()
print(w, b, w + b)   # 8192 128 8320

Weights + bias → 8320

Why: The bias adds one scalar per output neuron (128 of them). Verified: sum(p.numel() for p in Linear(64,128).parameters()) = 8320. Always add d_out.

45. Inspect it line by line: Trap: forgetting the bias in the count

Error analysis

Annotate

Walk the callouts on Trap: forgetting the bias in the count. Each one is a place this is easy to get subtly wrong.

  • Misses the 128-element bias vector. That is an off-by-d_out error on EVERY layer; across a deep net it compounds into thousands of uncounted parameters and a wrong size budget.
  • The bias adds one scalar per output neuron (128 of them). Verified: sum(p.numel() for p in Linear(64,128).parameters()) = 8320. Always add d_out.

46. Activation functions — outputs & gradients

Section

Part 5 of 8 — the nonlinear core

47. The four activations at a glance

Concept

All four map a scalar pre-activation to a scalar output. Their differences in range, zero-centering, and gradient behaviour decide which layer they belong in.

activationrangezero-centered?key issue
sigmoid σ(x)(0, 1)no (σ(0)=0.5)saturates → vanishing gradient
tanh(x)(−1, 1)yes (tanh(0)=0)saturates, but less than sigmoid
ReLU max(0,x)[0, ∞)nodead neurons if preact stays < 0
GELU x·Φ(x)≈(−0.17, ∞)≈ yes (GELU(0)=0)smooth; default in transformers

48. Guess the shape of the answer: Exact activation values — computed

Estimation

Predict first

Evaluate all four at x ∈ {−2, −1, 0, 1, 2}, and also the sigmoid derivative σ'(x) = σ(x)(1−σ(x)) — the quantity that governs gradient flow:

Commit before you compute: what does Exact activation values — computed come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: σ'(±2) = 0.1050 already; the max of σ' is only 0.25 (at x=0)

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The sigmoid derivative never exceeds 0.25 anywhere, and shrinks fast toward the tails.

49. Exact activation values — computed

Worked example

Evaluate all four at x ∈ {−2, −1, 0, 1, 2}, and also the sigmoid derivative σ'(x) = σ(x)(1−σ(x)) — the quantity that governs gradient flow:

import torch, torch.nn as nn
x = torch.tensor([-2., -1., 0., 1., 2.])
sig  = torch.sigmoid(x)
tanh = torch.tanh(x)
relu = torch.relu(x)
gelu = nn.GELU()(x)
sdv  = sig * (1 - sig)      # sigmoid derivative
for xi, s, ta, r, g, sd in zip(x.tolist(), sig, tanh, relu, gelu, sdv):
    print(f"x={xi:+.0f} sig={s:.4f} tanh={ta:.4f} relu={r:.4f} gelu={g:.4f} sig'={sd:.4f}")

σ'(±2) = 0.1050 already; the max of σ' is only 0.25 (at x=0)

Why: The sigmoid derivative never exceeds 0.25 anywhere, and shrinks fast toward the tails. In a deep stack these sub-0.25 factors multiply — the seed of the vanishing gradient.

xsigmoidtanhrelugeluσ'(x)
−2.00.1192−0.96400.0000−0.04550.1050
−1.00.2689−0.76160.0000−0.15870.1966
0.00.50000.00000.00000.00000.2500
1.00.73110.76161.00000.84130.1966
2.00.88080.96402.00001.95450.1050

50. Which is which, by relu

Discrimination

Sort into buckets

Sort these by relu, from memory, without looking back at Exact activation values — computed. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

0.0000
−2.0; −1.0; 0.0
1.0000
1.0
2.0000
2.0
g1
relu is "0.0000" for −2.0, −1.0, 0.0 — that is what the table on "Exact activation values — computed" records, and it is the single property separating this group from the rest.
g2
relu is "1.0000" for 1.0 — that is what the table on "Exact activation values — computed" records, and it is the single property separating this group from the rest.
g3
relu is "2.0000" for 2.0 — that is what the table on "Exact activation values — computed" records, and it is the single property separating this group from the rest.

51. GELU: a smooth, self-gating activation

Concept

GELU (used in GPT, BERT, ViT) multiplies the input by the probability that a standard normal is below it — softly gating small values while passing large positives almost unchanged.

\[ \operatorname{GELU}(x) = x\cdot \Phi(x) = x\cdot \tfrac{1}{2}\Big[\,1 + \operatorname{erf}\!\big(\tfrac{x}{\sqrt{2}}\big)\Big] \]

Unlike ReLU it is smooth everywhere and produces small negative outputs for slightly negative inputs (never a hard zero), so a GELU unit never fully 'dies'.

52. Restore the missing line: GELU from scratch — verified against nn.GELU

Fill the middle

Fill in the blanks

From GELU from scratch — verified against nn.GELU — one line has had its right-hand side removed. Put it back.

import numpy as np, torch, torch.nn as nn
from scipy.special import erf
def gelu_scratch(x):
return x * 0.5 * (1.0 + erf(x / np.sqrt(2)))
pts = np.array([-1., 0., 1., 2.])
manual = gelu_scratch(pts)
torch_v = nn.GELU()(torch.tensor(pts, dtype=torch.float32)).numpy()
for x, m, t in zip(pts, manual, torch_v):
print(f'x=___ manual=___ torch=___ match=___')

Why: manual is what everything below it consumes, so the wrong expression here fails later and somewhere else. GELU(−1)=−0.158655, GELU(0)=0, GELU(1)=0.841345, GELU(2)=1.954500.

53. GELU from scratch — verified against nn.GELU

Worked example

Implement GELU straight from the erf formula and confirm it matches PyTorch's exact (non-tanh) implementation:

import numpy as np, torch, torch.nn as nn
from scipy.special import erf
def gelu_scratch(x):
    return x * 0.5 * (1.0 + erf(x / np.sqrt(2)))
pts = np.array([-1., 0., 1., 2.])
manual = gelu_scratch(pts)
torch_v = nn.GELU()(torch.tensor(pts, dtype=torch.float32)).numpy()
for x, m, t in zip(pts, manual, torch_v):
    print(f'x={x:+.1f} manual={m:.6f} torch={t:.6f} match={abs(m-t)<1e-5}')

All four agree to < 1e-5

Why: GELU(−1)=−0.158655, GELU(0)=0, GELU(1)=0.841345, GELU(2)=1.954500. nn.GELU defaults to the exact erf form (not the tanh approximation), so a from-scratch scipy.special.erf build matches it.

xGELU (formula)nn.GELU()match
−1.0−0.158655−0.158655yes
0.00.0000000.000000yes
1.00.8413450.841345yes
2.01.9545001.954500yes

54. What stays fixed: GELU from scratch — verified against nn.GELU

Invariant

Step through it

Step through GELU from scratch — verified against nn.GELU one row at a time. One of these columns never changes — find it, and say why it cannot.

  1. Step 1: x is −1.0
  2. Step 2: x is 0.0
  3. Step 3: x is 1.0
  4. Step 4: x is 2.0

55. Failure mode 1: vanishing gradients

Concept

Backprop multiplies one activation derivative per layer. With sigmoid, each factor is ≤ 0.25, so across depth the gradient reaching the early layers shrinks geometrically — they stop learning.

This is the whole reason ReLU replaced sigmoid in hidden layers: ReLU'(z) = 1 for active neurons, so there is no shrinking factor to compound.

56. Teach it back: Failure mode 1: vanishing gradients

Explain it

Discussion prompt

Explain Failure mode 1: vanishing gradients to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Backprop multiplies one activation derivative per layer. With sigmoid, each factor is ≤ 0.25, so across depth the gradient reaching the early layers shrinks geometrically — they stop learning.

57. Quantify the vanishing gradient

Worked example

Take the best case σ' = 0.25 and a realistic case σ' ≈ 0.1966 (at activation ≈ 1), and raise each to the depth. Watch it collapse:

for L in [1, 4, 8, 12]:
    print(f'depth {L:2d}:  0.25**{L} = {0.25**L:.3e}')
print('---')
for L in [1, 4, 8, 12]:
    print(f'depth {L:2d}:  0.1966**{L} = {0.1966**L:.3e}')

At depth 12 the gradient factor is ~5.96e−08 (best case)

Why: Even the most generous σ' = 0.25 gives 0.25¹² ≈ 6×10⁻⁸ — the first layers receive essentially zero signal and training stalls. The realistic case is worse (~3.3×10⁻⁹).

depth L0.25ᴸ (best)0.1966ᴸ (typical)
12.500e−011.966e−01
43.906e−031.494e−03
81.526e−052.232e−06
125.960e−083.334e−09

58. Fill in: 0.25ᴸ (best) for Quantify the vanishing gradient

Comparison

Comparison matrix

From Quantify the vanishing gradient: refill the 0.25ᴸ (best) column from what you know. The rest of the table is as it appeared.

depth L0.25ᴸ (best)0.1966ᴸ (typical)
12.500e−011.966e−01
43.906e−031.494e−03
81.526e−052.232e−06
125.960e−083.334e−09

59. Predict the next row: Failure mode 2: the dying ReLU

Pattern

Predict first

The table runs: 0 | −5.164 | 0.0 | 0 · 1 | −6.220 | 0.0 | 0 · 2 | −3.890 | 0.0 | 0

In Failure mode 2: the dying ReLU, given the rows so far: what is the next one — the row where sample is … (all 6)?

Correct: … (all 6) | all < 0 | 0.0 | 0 (dead)

samplepre-activationReLU outgrad gate
0−5.1640.00
1−6.2200.00
2−3.8900.00
… (all 6)all < 00.00 (dead)

Why: The relationship between the columns, not the individual numbers, is what generates the next row. preacts ≈ [−5.16, −6.22, −3.89, −3.80, −3.31, −4.60], all < 0, so the gradient gate is 0 on every sample.

60. Failure mode 2: the dying ReLU

Worked example

A ReLU neuron dies when its pre-activation is negative for every input: output 0, and ReLU'(z) = 0 too, so no gradient ever flows to revive it. Construct one and check its gate on all data:

import numpy as np
np.random.seed(0)
data = np.random.randn(6, 2)          # 6 samples, 2 features
w = np.array([-1.0, -1.0]); b = -3.0  # forces preact very negative
preact = data @ w + b
gate = (preact > 0).astype(float)     # ReLU'(z): 1 if alive else 0
print('preacts:', preact.round(3))
print('gate   :', gate)
print('alive on any sample?', bool(gate.any()))

Every pre-activation is negative → gate is all zeros

Why: preacts ≈ [−5.16, −6.22, −3.89, −3.80, −3.31, −4.60], all < 0, so the gradient gate is 0 on every sample. The neuron outputs 0 forever and receives no update — permanently dead.

samplepre-activationReLU outgrad gate
0−5.1640.00
1−6.2200.00
2−3.8900.00
… (all 6)all < 00.00 (dead)

61. What stays fixed: Failure mode 2: the dying ReLU

Invariant

Step through it

Step through Failure mode 2: the dying ReLU one row at a time. One of these columns never changes — find it, and say why it cannot.

  1. Step 1: sample is 0
  2. Step 2: sample is 1
  3. Step 3: sample is 2
  4. Step 4: sample is … (all 6)

62. Something is wrong here: sigmoid in every hidden layer

Anomaly

Predict first

A student writes this, and it looks reasonable:

Sigmoid is a classic activation — use it in every hidden layer so activations stay in (0, 1).

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Each backprop step multiplies by σ'(x) ≤ 0.25.

Use ReLU or GELU in hidden layers; reserve sigmoid for the output of binary classification, where saturation is the desired probability clamp.

Why: Each backprop step multiplies by σ'(x) ≤ 0.25. After 12 layers the gradient is ≤ 0.25¹² ≈ 6×10⁻⁸ — the early layers get no signal and training freezes.

63. Trap: sigmoid in every hidden layer

Trap

The trap

Sigmoid is a classic activation — use it in every hidden layer so activations stay in (0, 1).

Stack 12 sigmoid hidden layers, train with SGD

Why: Each backprop step multiplies by σ'(x) ≤ 0.25. After 12 layers the gradient is ≤ 0.25¹² ≈ 6×10⁻⁸ — the early layers get no signal and training freezes.

# effective gradient scale into the first layer, 12 sigmoid layers
print(0.25**12)   # 5.96e-08  -> vanished

The fix

Use ReLU or GELU in hidden layers; reserve sigmoid for the output of binary classification, where saturation is the desired probability clamp.

Hidden: ReLU / GELU / tanh · Output: sigmoid (binary) or logits

Why: ReLU's gradient is 1 for active neurons — no compounding decay. Sigmoid is safe only in a single output layer where you WANT (0,1) squashing.

# ReLU gradient does not compound across depth
print(1.0**12)   # 1.0  -> signal survives

64. Which of these survive contact with Lesson 44: MLP Architectures — Depth, Width…?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
We feed it the input x = [1, 2]. Row j of W⁽¹⁾ holds the two weights of hidden neuron j. Keep this table of numbers in view — the next four slides trace it end to end.; The output layer is a weighted committee of those detectors. Learning means tuning who votes, how strongly each detector responds, and how the committee combines them.; Rule of thumb: hierarchical targets (images, sequences) reward depth; a smooth scalar target with no hierarchy is fine with a moderately wide shallow net.
Breaks
XOR just needs more capacity — drop the hidden layer and train a single Linear(2,1) harder.; Linear(64, 128) is a 64×128 weight matrix, so 64 × 128 = 8192 parameters.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 44: MLP Architectures — Depth, Width & Activations puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

65. Deep vs wide — a fair experiment

Section

Part 6 of 8 — equal parameter budgets

66. What has to be given first: Match the parameter budgets first

Missing information

Discussion prompt

A fair comparison holds the parameter count fixed and varies only the shape. Compute the budgets by hand: a wide net 64→512→10 vs a deep net (9 hidden layers of width 64, then 64→10).

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

Within 1% of each other — so any accuracy gap is due to shape, not size. The wide net spends its budget on one fat layer; the deep net spreads it across ten thin ones.

67. Match the parameter budgets first

Worked example

A fair comparison holds the parameter count fixed and varies only the shape. Compute the budgets by hand: a wide net 64→512→10 vs a deep net (9 hidden layers of width 64, then 64→10).

# wide: Linear(64,512) + Linear(512,10)
w1 = 64*512 + 512
w2 = 512*10 + 10
print('wide:', w1, '+', w2, '=', w1 + w2)
# deep: 9 x Linear(64,64) + Linear(64,10)
d_hidden = 9 * (64*64 + 64)
d_out = 64*10 + 10
print('deep:', d_hidden, '+', d_out, '=', d_hidden + d_out)

Wide = 33280 + 5130 = 38410; Deep = 37440 + 650 = 38090

Why: Within 1% of each other — so any accuracy gap is due to shape, not size. The wide net spends its budget on one fat layer; the deep net spreads it across ten thin ones.

modelshapeparams
wide64 → 512 → 1038 410
deep64 → (64)×9 → 1038 090
gap< 1%

68. Guess the shape of the answer: Wide vs deep on load_digits

Estimation

Predict first

Train both on load_digits (1 797 samples, 64 features, 10 classes), same optimizer, same seed, same 400 epochs. This block is self-contained:

Commit before you compute: what does Wide vs deep on load_digits come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: wide → 97.50% test acc; deep → 94.72% test acc

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. On flat tabular digits the wide net WINS.

69. Wide vs deep on load_digits

Worked example

Train both on load_digits (1 797 samples, 64 features, 10 classes), same optimizer, same seed, same 400 epochs. This block is self-contained:

import torch, torch.nn as nn, numpy as np
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
X, y = load_digits(return_X_y=True)
Xtr, Xte, ytr, yte = train_test_split(X.astype(np.float32), y, test_size=0.2, random_state=42)
sc = StandardScaler(); Xtr = sc.fit_transform(Xtr); Xte = sc.transform(Xte)
Xtr_t, ytr_t = torch.tensor(Xtr), torch.tensor(ytr, dtype=torch.long)
Xte_t, yte_t = torch.tensor(Xte), torch.tensor(yte, dtype=torch.long)
torch.manual_seed(0)
wide = nn.Sequential(nn.Linear(64,512), nn.ReLU(), nn.Linear(512,10))
torch.manual_seed(0)
layers = []
for i in range(10):
    layers += [nn.Linear(64,64), nn.ReLU()] if i < 9 else [nn.Linear(64,10)]
deep = nn.Sequential(*layers)
for name, m in [('wide', wide), ('deep', deep)]:
    opt = torch.optim.Adam(m.parameters(), lr=0.005)
    for _ in range(400):
        opt.zero_grad(); nn.CrossEntropyLoss()(m(Xtr_t), ytr_t).backward(); opt.step()
    acc = (m(Xte_t).argmax(1) == yte_t).float().mean().item()
    print(f'{name}: params={sum(p.numel() for p in m.parameters())} acc={acc*100:.2f}%')

wide → 97.50% test acc; deep → 94.72% test acc

Why: On flat tabular digits the wide net WINS. There is no spatial hierarchy for depth to exploit, and the deep ReLU stack is slightly harder to optimize. The result is task-dependent, not a universal law.

modelparamsinit lossfinal losstest acc
wide (1 hidden, w=512)38 4102.34810.000297.50%
deep (10 hidden, w=64)38 0902.30470.000194.72%

70. Work backwards from the answer: Wide vs deep on load_digits

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

wide → 97.50% test acc; deep → 94.72% test acc

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Train both on load_digits (1 797 samples, 64 features, 10 classes), same optimizer, same seed, same 400 epochs. This block is self-contained:

71. USAAIO architecture-design guidelines

Concept

scenariorecommended choicereason
flat tabular, < 1 000 features2–3 layers, moderate widthno hierarchy; a wide-ish shallow net interpolates well
image / sequence (hierarchical)deep, narrow (ResNet-style)each layer learns finer structure
tight parameter budgetprefer depthdepth is more parameter-efficient on structured data
deep sigmoid stallsswap to ReLU / add BatchNorm / residualsavoid compounding sub-0.25 gradients
very deep net divergesreduce depth first, then retune lrdepth × lr interact — start shallow, grow

The digits result is the guideline in action: match the architecture to the data's structure, not to a slogan like 'deeper is always better'.

72. By analogy: USAAIO architecture-design guidelines

Analogy

Discussion prompt

Explain USAAIO architecture-design guidelines by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

The digits result is the guideline in action: match the architecture to the data's structure, not to a slogan like 'deeper is always better'.

73. The recipe & self-check

Section

Part 7 of 8

74. Without one step: The MLP architecture recipe

Constraint

Discussion prompt

Run The MLP architecture recipe with this step confiscated:

Output activation: none/logits for CrossEntropyLoss · sigmoid for binary · softmax only if you need explicit probabilities

Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.

Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.

Answer:

  1. Trace forward once: z = Wh + b, then h = σ(z), layer by layer — confirm shapes and one numeric output against torch
  2. Count params per layer: d_in × d_out + d_out; sum across all nn.Linear (activations add zero)
  3. Hidden activation: ReLU (default) · GELU (transformers) · tanh (RNNs) — never sigmoid in hidden layers
  4. Output activation: none/logits for CrossEntropyLoss · sigmoid for binary · softmax only if you need explicit probabilities
  5. Depth vs width: hierarchical data → go deep; flat tabular → moderate width; hold the parameter budget fixed to compare fairly
  6. Watch the failure modes: sigmoid → vanishing gradient across depth; large lr / bad init → dying ReLU (permanently zero neurons)

75. The MLP architecture recipe

Pattern

  1. Trace forward once: z = Wh + b, then h = σ(z), layer by layer — confirm shapes and one numeric output against torch
  2. Count params per layer: d_in × d_out + d_out; sum across all nn.Linear (activations add zero)
  3. Hidden activation: ReLU (default) · GELU (transformers) · tanh (RNNs) — never sigmoid in hidden layers
  4. Output activation: none/logits for CrossEntropyLoss · sigmoid for binary · softmax only if you need explicit probabilities
  5. Depth vs width: hierarchical data → go deep; flat tabular → moderate width; hold the parameter budget fixed to compare fairly
  6. Watch the failure modes: sigmoid → vanishing gradient across depth; large lr / bad init → dying ReLU (permanently zero neurons)

76. Where does it stop working: The MLP architecture recipe

Edge cases

Discussion prompt

The MLP architecture recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Trace forward once: z = Wh + b, then h = σ(z), layer by layer — confirm shapes and one numeric output against torch
  2. Count params per layer: d_in × d_out + d_out; sum across all nn.Linear (activations add zero)
  3. Hidden activation: ReLU (default) · GELU (transformers) · tanh (RNNs) — never sigmoid in hidden layers
  4. Output activation: none/logits for CrossEntropyLoss · sigmoid for binary · softmax only if you need explicit probabilities
  5. Depth vs width: hierarchical data → go deep; flat tabular → moderate width; hold the parameter budget fixed to compare fairly
  6. Watch the failure modes: sigmoid → vanishing gradient across depth; large lr / bad init → dying ReLU (permanently zero neurons)

77. Rule out three: Check yourself — the forward pass

Elimination

Eliminate the wrong options

With W⁽²⁾ = [1, −1, 2], b⁽²⁾ = 0.5, and hidden output h⁽¹⁾ = [0, 0.5, 0.5], what is the network's output?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. 1.0
  • B. 0.5
  • C. 2.0
  • D. 1.5

Survives elimination: A

Why: z⁽²⁾ = (1)(0) + (−1)(0.5) + (2)(0.5) + 0.5 = 0 − 0.5 + 1.0 + 0.5 = 1.0. The dead neuron 0 drops out; neuron 1 subtracts 0.5, neuron 2 adds 1.0, and the bias adds 0.5.

78. Check yourself — the forward pass

Check

Use the running network: x = [1, 2], ReLU hidden, and h⁽¹⁾ = [0, 0.5, 0.5].

Check your understanding

With W⁽²⁾ = [1, −1, 2], b⁽²⁾ = 0.5, and hidden output h⁽¹⁾ = [0, 0.5, 0.5], what is the network's output?

  • A. 1.0 (correct)
  • B. 0.5
  • C. 2.0
  • D. 1.5

Answer: A

Why: z⁽²⁾ = (1)(0) + (−1)(0.5) + (2)(0.5) + 0.5 = 0 − 0.5 + 1.0 + 0.5 = 1.0. The dead neuron 0 drops out; neuron 1 subtracts 0.5, neuron 2 adds 1.0, and the bias adds 0.5.

Why B tempts people
This is just the bias b⁽²⁾ = 0.5 — it forgets the two firing hidden neurons' contributions (−0.5 and +1.0), which net to +0.5 on top of the bias.
Why C tempts people
Uses the raw pre-activations z⁽¹⁾ instead of the ReLU outputs, or forgets neuron 0 was clamped to 0 — double-counting the silenced neuron.
Why D tempts people
Adds the two firing outputs as +0.5 and +1.0 (both positive), ignoring that neuron 1's committee weight is −1, so its contribution is −0.5, not +0.5.

79. Answer it before you see the options: Check yourself — parameter count

Prediction

Predict first

How many trainable parameters does nn.Linear(128, 64) have?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: 8 256 (128×64 + 64)

Why: d_in × d_out + d_out = 128×64 + 64 = 8192 + 64 = 8256. The +64 is the bias, one scalar per output neuron. Verified: sum(p.numel() for p in nn.Linear(128,64).parameters()) = 8256.

80. Check yourself — parameter count

Check

Apply d_in × d_out + d_out before looking.

Check your understanding

How many trainable parameters does nn.Linear(128, 64) have?

  • A. 8 256 (128×64 + 64) (correct)
  • B. 8 192 (128×64 only)
  • C. 8 320 (128×64 + 128)
  • D. 192 (128 + 64)

Answer: A

Why: d_in × d_out + d_out = 128×64 + 64 = 8192 + 64 = 8256. The +64 is the bias, one scalar per output neuron. Verified: sum(p.numel() for p in nn.Linear(128,64).parameters()) = 8256.

Why B tempts people
Counts only the weight matrix (128×64 = 8192) and forgets the 64-element bias vector — the classic off-by-bias error.
Why C tempts people
Adds d_in (128) instead of d_out (64) as the bias count. The bias length matches the OUTPUT dimension, not the input.
Why D tempts people
Adds the two sizes as if params were additive. The weight matrix is multiplicative (128×64); only the bias is additive (+64).

81. Answer it before you see the options: Check yourself — activation choice

Prediction

Predict first

You build a 12-layer MLP. Which hidden-layer activation is most likely to stall training because gradients vanish?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: sigmoid — derivative ≤ 0.25 compounds to near-zero across layers

Why: σ'(x) ≤ 0.25 everywhere. Across 12 layers the gradient into the first layer is ≤ 0.25¹² ≈ 6×10⁻⁸ — effectively zero. ReLU's gradient is 1 for active neurons, so it does not compound.

82. Check yourself — activation choice

Check

Think about what each derivative does across depth.

Check your understanding

You build a 12-layer MLP. Which hidden-layer activation is most likely to stall training because gradients vanish?

  • A. sigmoid — derivative ≤ 0.25 compounds to near-zero across layers (correct)
  • B. ReLU — outputs zero for negative inputs
  • C. GELU — its smooth curve attenuates the signal
  • D. tanh — its range (−1, 1) is too small

Answer: A

Why: σ'(x) ≤ 0.25 everywhere. Across 12 layers the gradient into the first layer is ≤ 0.25¹² ≈ 6×10⁻⁸ — effectively zero. ReLU's gradient is 1 for active neurons, so it does not compound.

Why B tempts people
ReLU does zero out negatives (and dead neurons are real), but its gradient for ACTIVE neurons is exactly 1 — no decay compounds across depth.
Why C tempts people
GELU is smooth and never hard-zeros; its derivative is near 1 for positive inputs, so it does not vanish across depth the way sigmoid does.
Why D tempts people
tanh's range is (−1,1) but its derivative at 0 is 1 and it decays less steeply than sigmoid; range size is not what causes vanishing gradients — the small derivative is.

83. Rule out three: Check yourself — depth vs width

Elimination

Eliminate the wrong options

On load_digits at a ~38k parameter budget, the wide model (1 hidden, width 512) hit 97.5% and the deep model (10 hidden, width 64) hit 94.7%. Best interpretation?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Width won here because digits is flat/tabular; depth helps mainly on hierarchical data
  • B. Depth always beats width at equal parameter count
  • C. The deep model overfit because it had more parameters
  • D. It proves sigmoid is the better activation for digits

Survives elimination: A

Why: load_digits is a flat 64-feature problem with no spatial hierarchy, so a single wide layer interpolates efficiently and the deep ReLU stack is slightly harder to optimize. Depth shines on hierarchical data (images, sequences) — the result is dataset-dependent.

84. Check yourself — depth vs width

Check

Use this lesson's experiment result.

Check your understanding

On load_digits at a ~38k parameter budget, the wide model (1 hidden, width 512) hit 97.5% and the deep model (10 hidden, width 64) hit 94.7%. Best interpretation?

  • A. Width won here because digits is flat/tabular; depth helps mainly on hierarchical data (correct)
  • B. Depth always beats width at equal parameter count
  • C. The deep model overfit because it had more parameters
  • D. It proves sigmoid is the better activation for digits

Answer: A

Why: load_digits is a flat 64-feature problem with no spatial hierarchy, so a single wide layer interpolates efficiently and the deep ReLU stack is slightly harder to optimize. Depth shines on hierarchical data (images, sequences) — the result is dataset-dependent.

Why B tempts people
This experiment is a direct counterexample. Depth wins when the target has compositional/hierarchical structure — not universally.
Why C tempts people
The deep model actually had FEWER parameters (38 090 vs 38 410), and overfitting would show low train loss with high test loss; both converged to near-zero train loss.
Why D tempts people
Both models used ReLU; sigmoid was never tested here and is irrelevant to the accuracy comparison.

85. Your turn: build & count

Section

Part 8 of 8 — the project

86. Project: architect, count, and verify

Concept

Build MLPs of different shapes on load_digits, count their parameters exactly, and reimplement GELU from the formula — proving every number against torch and scipy.

#requirementtool / formula
1count params for any Linear stackd_in×d_out + d_out, summed
2implement GELU from scratchscipy.special.erf
3train the [64→128→64→10] MLP, score test accAdam, CrossEntropyLoss, load_digits

Build rules: type every line yourself, predict each number before printing, and verify every count with sum(p.numel() for p in model.parameters()).

87. Break it if you can: Project: architect, count, and verify

Counterexample

Discussion prompt

Build MLPs of different shapes on load_digits, count their parameters exactly, and reimplement GELU from the formula — proving every number against torch and scipy.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Build rules: type every line yourself, predict each number before printing, and verify every count with sum(p.numel() for p in model.parameters()).

88. Milestone 1 — parameter count by hand

Worked example

Your turn: predict the total for [64 → 128 → 64 → 10] before running. Write each layer's count first.

Hint: apply d_in × d_out + d_out to each Linear and sum the three. The activations contribute nothing.

import torch.nn as nn
fc1 = nn.Linear(64, 128); fc2 = nn.Linear(128, 64); fc3 = nn.Linear(64, 10)
for name, lyr in [('fc1', fc1), ('fc2', fc2), ('fc3', fc3)]:
    print(name, sum(p.numel() for p in lyr.parameters()))
total = sum(p.numel() for lyr in (fc1, fc2, fc3) for p in lyr.parameters())
print('total:', total)
layerformulacount
fc1 Linear(64,128)64×128 + 1288 320
fc2 Linear(128,64)128×64 + 648 256
fc3 Linear(64,10)64×10 + 10650
total17 226

89. Milestone 2 — GELU from scratch

Worked example

Your turn: implement gelu(x) with scipy.special.erf. Predict the value at x = 1 before you run it.

Hint: GELU(x) = x · 0.5 · (1 + erf(x / √2)). At x = 1: 1 · 0.5 · (1 + erf(0.707)) ≈ 0.841.

import numpy as np, torch, torch.nn as nn
from scipy.special import erf
def gelu(x):
    return x * 0.5 * (1.0 + erf(x / np.sqrt(2)))
pts = np.array([-1., 0., 1., 2.])
manual = gelu(pts)
torch_v = nn.GELU()(torch.tensor(pts, dtype=torch.float32)).numpy()
for x, m, t in zip(pts, manual, torch_v):
    print(f'x={x:+.1f} scratch={m:.6f} nn.GELU={t:.6f}')
xgelu (scratch)nn.GELU()match
−1.0−0.158655−0.158655yes
0.00.0000000.000000yes
1.00.8413450.841345yes
2.01.9545001.954500yes

90. What stays fixed: Milestone 2 — GELU from scratch

Invariant

Step through it

Step through Milestone 2 — GELU from scratch one row at a time. One of these columns never changes — find it, and say why it cannot.

  1. Step 1: x is −1.0
  2. Step 2: x is 0.0
  3. Step 3: x is 1.0
  4. Step 4: x is 2.0

91. Milestone 3 — train and score

Worked example

Your turn: train the [64→128→64→10] MLP on scaled load_digits and read off the test accuracy. Predict whether it clears 95%.

Hint: standardize features, use Adam(lr=0.005), CrossEntropyLoss, 400 epochs, then compare argmax(1) to the labels.

import torch, torch.nn as nn, numpy as np
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
X, y = load_digits(return_X_y=True)
Xtr, Xte, ytr, yte = train_test_split(X.astype(np.float32), y, test_size=0.2, random_state=42)
sc = StandardScaler(); Xtr = sc.fit_transform(Xtr); Xte = sc.transform(Xte)
torch.manual_seed(0)
model = nn.Sequential(nn.Linear(64,128), nn.ReLU(), nn.Linear(128,64), nn.ReLU(), nn.Linear(64,10))
opt = torch.optim.Adam(model.parameters(), lr=0.005)
for _ in range(400):
    opt.zero_grad(); nn.CrossEntropyLoss()(model(torch.tensor(Xtr)), torch.tensor(ytr, dtype=torch.long)).backward(); opt.step()
with torch.no_grad():
    acc = (model(torch.tensor(Xte)).argmax(1) == torch.tensor(yte)).float().mean()
print('test acc:', round(acc.item(), 4))
quantityvalue (verified)
params17 226
test acc0.9694 (≈ 97%)
clears 95%?yes

92. What each one costs: Milestone 3 — train and score

Trade off

Comparison matrix

From Milestone 3 — train and score: every row here is a choice with a cost. Fill the value (verified) column, then say which row you would actually pick and what you give up for it.

quantityvalue (verified)
params17 226
test acc0.9694 (≈ 97%)
clears 95%?yes

93. The full program

Concept

import torch, torch.nn as nn, numpy as np
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from scipy.special import erf

torch.manual_seed(0)
model = nn.Sequential(nn.Linear(64,128), nn.ReLU(),
                      nn.Linear(128,64), nn.ReLU(), nn.Linear(64,10))
print('params:', sum(p.numel() for p in model.parameters()))   # 17226

gelu = lambda x: x * 0.5 * (1 + erf(x / np.sqrt(2)))
print('GELU(1):', round(float(gelu(1.0)), 6))                   # 0.841345

X, y = load_digits(return_X_y=True)
Xtr, Xte, ytr, yte = train_test_split(X.astype(np.float32), y, test_size=0.2, random_state=42)
sc = StandardScaler(); Xtr = sc.fit_transform(Xtr); Xte = sc.transform(Xte)
opt = torch.optim.Adam(model.parameters(), lr=0.005)
for _ in range(400):
    opt.zero_grad(); nn.CrossEntropyLoss()(model(torch.tensor(Xtr)), torch.tensor(ytr, dtype=torch.long)).backward(); opt.step()
with torch.no_grad():
    acc = (model(torch.tensor(Xte)).argmax(1) == torch.tensor(yte)).float().mean()
print('test acc:', round(acc.item(), 4))                        # 0.9694
printed linevalue
params:17226
GELU(1):0.841345
test acc:0.9694

If your count reads 17226, GELU(1) is 0.841345, and the net clears 96% — you can audit any MLP's size, activation math, and capacity from first principles.

94. Fill in: value for The full program

Comparison

Comparison matrix

From The full program: refill the value column from what you know. The rest of the table is as it appeared.

printed linevalue
params:17226
GELU(1):0.841345
test acc:0.9694

95. Show it off

Concept

Slides closed, out loud: (1) forward-propagate x = [1, 2] through the running net to the output 1.0, (2) state what universal approximation guarantees and what it does not, (3) count parameters for [64→128→64→10] without code, (4) explain why 12 stacked sigmoids stall training.

Stretch: rerun the wide-vs-deep experiment with GELU instead of ReLU and with a residual connection every two layers — see whether depth catches up once optimization is easier. Next: Lesson 45 — convolutional layers (spatial inductive bias).

96. Connect it up: Lesson 44: MLP Architectures — Depth, Width & Activations

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — What an MLP actually computes · Universal approximation — and its catch · Depth vs width — why depth wins · Parameter counting — exact formula · Activation functions — outputs & gradients · Deep vs wide — a fair experiment. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

97. What you can do now

Recap

conceptthe one thing to remember
forward passz = Wh + b, then h = σ(z); ReLU silences negatives
universal approx.one wide layer CAN — but depth is exponentially more efficient
param countd_in × d_out + d_out; bias adds d_out per layer
sigmoid hiddenσ' ≤ 0.25 per layer — never stack sigmoid deep
GELUx · 0.5 · (1 + erf(x/√2)); GELU(1) = 0.841345
wide vs deepdigits: wide 97.5% > deep 94.7% — match shape to structure

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 44 (Week 15 — MLP Architectures) — Barron · USAAIO Round 2 Preparation, 2026
  2. Cybenko (1989) / Hornik (1991), Universal Approximation Theorem — Approximation by superpositions of a sigmoidal function; Multilayer feedforward networks are universal approximators
  3. Hendrycks & Gimpel (2016), Gaussian Error Linear Units (GELUs)
  4. PyTorch nn.Linear / nn.GELU documentation
  5. Every forward pass, parameter count, activation value, XOR fit, gradient product, and digits accuracy produced by real execution — torch 2.7.1 + numpy 2.2.6 + scikit-learn 1.9.0, verification run July 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108