USAAIO Lesson 44, from Week 15 of Phase 2, fully worked. One running 2→3→1 MLP is hand-traced neuron by neuron and confirmed against torch. It states the universal approximation theorem and demonstrates it on XOR, showing a single Linear layer fail, then covers depth against width as expressivity by counting linear regions. It derives the exact nn.Linear parameter formula and sums it layer by layer, computes all four activations - sigmoid, tanh, ReLU, and GELU - at five inputs with their gradient consequences, and rebuilds GELU from erf. It quantifies the vanishing-gradient and dying-ReLU failure modes, runs a fair wide-against-deep experiment at about 38k parameters on load_digits, and ends with a from-scratch build-it project. Every snippet runs standalone in a fresh interpreter, and every number came from real execution with torch 2.7.1 and numpy 2.2.6. The lesson runs to 63 slides.
Subject: Machine Learning · 97 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 44 · Week 15 (Phase 2)
We trace one tiny network by hand — every neuron, every weight — then scale the same ideas up: universal approximation, why depth beats width, exact parameter counting, and the activation math that decides whether a deep net learns at all.
Objectives
torchd_in × d_out + d_out and sum any nn.Linear stackWarm-up
Discussion prompt
Before we open Lesson 44: MLP Architectures — Depth, Width & Activations: without looking back, what was the main idea of Support Vector Machines, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
hard-margin SVM geometric derivation, soft-margin slack variables and the role of C, kernel trick (RBF, polynomial), hyperparameter grid search with cross-validation, multi-class OvO vs OvR strategies, and SVM vs logistic regression trade-offs. Build and evaluate SVC on Iris and moons; verified with sklearn and real execution numbers.
Section
Part 1 of 8 — one running example
Concept
A multilayer perceptron alternates two moves: an affine map z = Wh + b (a matrix multiply plus a bias), then a nonlinear activation h' = σ(z) applied element-wise. Repeat for each layer.
\[ h^{(0)} = x, \qquad z^{(\ell)} = W^{(\ell)} h^{(\ell-1)} + b^{(\ell)}, \qquad h^{(\ell)} = \sigma\!\big(z^{(\ell)}\big) \]
The activation σ is the only nonlinear part. Strip it out and the whole stack collapses to a single affine map — which is exactly the trap we'll prove on XOR.
Concept
We fix one concrete network for the whole lesson: input dimension 2, a hidden layer of 3 ReLU neurons, and a single linear output. Every part of this deck comes back to it.
\[ W^{(1)} = \begin{bmatrix} 1 & -1 \\ 0.5 & 0.5 \\ -2 & 1 \end{bmatrix},\; b^{(1)} = \begin{bmatrix} 0 \\ -1 \\ 0.5 \end{bmatrix},\; W^{(2)} = \begin{bmatrix} 1 & -1 & 2 \end{bmatrix},\; b^{(2)} = \begin{bmatrix} 0.5 \end{bmatrix} \]
We feed it the input x = [1, 2]. Row j of W⁽¹⁾ holds the two weights of hidden neuron j. Keep this table of numbers in view — the next four slides trace it end to end.
Counterexample
Discussion prompt
We fix one concrete network for the whole lesson: input dimension 2, a hidden layer of 3 ReLU neurons, and a single linear output. Every part of this deck comes back to it.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
We feed it the input x = [1, 2]. Row j of W⁽¹⁾ holds the two weights of hidden neuron j. Keep this table of numbers in view — the next four slides trace it end to end.
Intuition
Think of every hidden neuron as a feature detector. It takes a weighted vote of the inputs (Wⱼ·x + bⱼ), and the activation decides how loudly it fires — ReLU stays silent (0) until its vote goes positive, then passes the vote straight through.
The output layer is a weighted committee of those detectors. Learning means tuning who votes, how strongly each detector responds, and how the committee combines them.
That is the whole mental model. Now let's turn the crank on real numbers — no hand-waving, one arithmetic step per neuron.
Analogy
Discussion prompt
Explain Each hidden neuron is a little detector by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
The output layer is a weighted committee of those detectors. Learning means tuning who votes, how strongly each detector responds, and how the committee combines them.
Ranking
Put in order
Put the moves of Forward pass, step 1 — the pre-activations z⁽¹⁾ into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Row 0 of W⁽¹⁾ is [1, −1], bias 0.
Worked example
First layer: z⁽¹⁾ = W⁽¹⁾x + b⁽¹⁾. Each hidden neuron dots its weight row with x = [1, 2] and adds its bias. Do all three, one at a time:
Neuron 0: (1)(1) + (−1)(2) + 0 = 1 − 2 = −1
Why: Row 0 of W⁽¹⁾ is [1, −1], bias 0. This vote came out negative — ReLU will silence it.
import numpy as np
x = np.array([1.0, 2.0])
W1 = np.array([[ 1.0, -1.0],
[ 0.5, 0.5],
[-2.0, 1.0]])
b1 = np.array([0.0, -1.0, 0.5])
z1 = W1 @ x + b1
print(z1) # [-1. 0.5 0.5]Neuron 1: (0.5)(1) + (0.5)(2) − 1 = 0.5 + 1 − 1 = 0.5
Why: Row 1 is [0.5, 0.5], bias −1. Positive vote — survives ReLU.
Neuron 2: (−2)(1) + (1)(2) + 0.5 = −2 + 2 + 0.5 = 0.5
Why: Row 2 is [−2, 1], bias 0.5. Also positive. So z⁽¹⁾ = [−1, 0.5, 0.5], matching the printed array exactly.
| neuron j | row Wⱼ | Wⱼ·x | + bⱼ | z⁽¹⁾ⱼ |
|---|---|---|---|---|
| 0 | [1, −1] | 1 − 2 = −1 | +0 | −1.0 |
| 1 | [0.5, 0.5] | 0.5 + 1 = 1.5 | −1 | 0.5 |
| 2 | [−2, 1] | −2 + 2 = 0 | +0.5 | 0.5 |
Comparison
Comparison matrix
From Forward pass, step 1 — the pre-activations z⁽¹⁾: refill the row Wⱼ column from what you know. The rest of the table is as it appeared.
| neuron j | row Wⱼ | Wⱼ·x | + bⱼ | z⁽¹⁾ⱼ |
|---|---|---|---|---|
| 0 | [1, −1] | 1 − 2 = −1 | +0 | −1.0 |
| 1 | [0.5, 0.5] | 0.5 + 1 = 1.5 | −1 | 0.5 |
| 2 | [−2, 1] | −2 + 2 = 0 | +0.5 | 0.5 |
Estimation
Predict first
Apply ReLU(z) = max(0, z) element-wise to z⁽¹⁾ = [−1, 0.5, 0.5]. ReLU clamps negatives to zero and leaves positives untouched:
Commit before you compute: what does Forward pass, step 2 — ReLU gives h⁽¹⁾ come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Neurons 1, 2: max(0, 0.5) = 0.5 each → pass through
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Positive pre-activations are unchanged, so h⁽¹⁾ = [0, 0.5, 0.5].
Worked example
Apply ReLU(z) = max(0, z) element-wise to z⁽¹⁾ = [−1, 0.5, 0.5]. ReLU clamps negatives to zero and leaves positives untouched:
import numpy as np
z1 = np.array([-1.0, 0.5, 0.5])
h = np.maximum(0.0, z1) # ReLU
print(h) # [0. 0.5 0.5]Neuron 0: max(0, −1) = 0 → dead on this input
Why: Its negative vote is silenced. This neuron contributes nothing to the output for x = [1, 2] — a live illustration of ReLU gating.
Neurons 1, 2: max(0, 0.5) = 0.5 each → pass through
Why: Positive pre-activations are unchanged, so h⁽¹⁾ = [0, 0.5, 0.5].
| neuron j | z⁽¹⁾ⱼ | ReLU(z⁽¹⁾ⱼ) | state |
|---|---|---|---|
| 0 | −1.0 | 0.0 | silenced |
| 1 | 0.5 | 0.5 | firing |
| 2 | 0.5 | 0.5 | firing |
Trade off
Comparison matrix
From Forward pass, step 2 — ReLU gives h⁽¹⁾: every row here is a choice with a cost. Fill the state column, then say which row you would actually pick and what you give up for it.
| neuron j | z⁽¹⁾ⱼ | ReLU(z⁽¹⁾ⱼ) | state |
|---|---|---|---|
| 0 | −1.0 | 0.0 | silenced |
| 1 | 0.5 | 0.5 | firing |
| 2 | 0.5 | 0.5 | firing |
Missing information
Discussion prompt
Output layer is linear: z⁽²⁾ = W⁽²⁾h⁽¹⁾ + b⁽²⁾ with W⁽²⁾ = [1, −1, 2], b⁽²⁾ = 0.5, and h⁽¹⁾ = [0, 0.5, 0.5]. One dot product:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Multiply each hidden output by its committee weight, sum, add the output bias.
Worked example
Output layer is linear: z⁽²⁾ = W⁽²⁾h⁽¹⁾ + b⁽²⁾ with W⁽²⁾ = [1, −1, 2], b⁽²⁾ = 0.5, and h⁽¹⁾ = [0, 0.5, 0.5]. One dot product:
(1)(0) + (−1)(0.5) + (2)(0.5) + 0.5
Why: Multiply each hidden output by its committee weight, sum, add the output bias.
import numpy as np
h = np.array([0.0, 0.5, 0.5])
W2 = np.array([1.0, -1.0, 2.0])
b2 = 0.5
z2 = W2 @ h + b2
print(z2) # 1.0= 0 − 0.5 + 1.0 + 0.5 = 1.0
Why: The dead neuron drops out, neuron 1 subtracts 0.5, neuron 2 adds 1.0, bias adds 0.5. Network output for x = [1, 2] is exactly 1.0.
| term | wⱼ · hⱼ | value |
|---|---|---|
| neuron 0 | 1 × 0 | 0.0 |
| neuron 1 | −1 × 0.5 | −0.5 |
| neuron 2 | 2 × 0.5 | +1.0 |
| bias | — | +0.5 |
| output z⁽²⁾ | 1.0 |
Error analysis
Annotate
Walk the callouts on Forward pass, step 3 — the output z⁽²⁾. Each one is a place this is easy to get subtly wrong.
Pattern
Predict first
The table runs: hand trace (numpy) | 1.0 · nn.Sequential(torch) | 1.0
In Confirm the hand trace against PyTorch, given the rows so far: what is the next one — the row where source is match??
Correct: match? | yes
| source | output for x = [1, 2] |
|---|---|
| hand trace (numpy) | 1.0 |
| nn.Sequential(torch) | 1.0 |
| match? | yes |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. nn.Linear stores weight as (out, in) and computes xWᵀ + b, which is exactly W⁽¹⁾x + b⁽¹⁾ for our row layout.
Worked example
Load the exact same weights into nn.Linear layers and run the network. It must print our hand-computed 1.0:
import torch, torch.nn as nn
x = torch.tensor([1.0, 2.0])
lin1 = nn.Linear(2, 3)
lin1.weight.data = torch.tensor([[1.,-1.],[0.5,0.5],[-2.,1.]])
lin1.bias.data = torch.tensor([0., -1., 0.5])
lin2 = nn.Linear(3, 1)
lin2.weight.data = torch.tensor([[1., -1., 2.]])
lin2.bias.data = torch.tensor([0.5])
net = nn.Sequential(lin1, nn.ReLU(), lin2)
with torch.no_grad():
print(net(x).numpy()) # [1.]torch prints [1.] — identical to the hand trace
Why: nn.Linear stores weight as (out, in) and computes xWᵀ + b, which is exactly W⁽¹⁾x + b⁽¹⁾ for our row layout. The math and the library agree to the last digit.
| source | output for x = [1, 2] |
|---|---|
| hand trace (numpy) | 1.0 |
| nn.Sequential(torch) | 1.0 |
| match? | yes |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
torch prints [1.] — identical to the hand trace
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Load the exact same weights into nn.Linear layers and run the network. It must print our hand-computed 1.0:
Section
Part 2 of 8 — what one hidden layer can do
Concept
A feedforward net with one hidden layer and a non-polynomial activation can approximate any continuous function on a compact domain, to any accuracy ε — given enough hidden neurons N.
\[ \forall\, f \in C([0,1]^d),\; \varepsilon > 0 \;\exists\; N,\, w,\, a,\, b :\; \sup_{x} \Big|\, f(x) - \sum_{i=1}^{N} w_i\, \sigma\!\big(a_i^{\top} x + b_i\big) \Big| < \varepsilon \]
It is a statement of existence (Cybenko 1989, Hornik 1991): the weights exist. It says nothing about how many neurons you need or whether gradient descent can find them.
Explain it
Discussion prompt
Explain The universal approximation theorem to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
A feedforward net with one hidden layer and a non-polynomial activation can approximate any continuous function on a compact domain, to any accuracy ε — given enough hidden neurons N.
Intuition
'Enough neurons' can mean exponentially many. A single hidden layer approximates a wiggly function by stacking up many bump-shaped pieces — one small region at a time — so a function with lots of structure needs a huge, flat layer.
Depth is the escape hatch: by composing layers, a deep net reuses earlier features and represents the same function with far fewer parameters. Existence (one wide layer) and efficiency (depth) are different questions.
So the theorem is reassuring, not prescriptive. In practice you reach for depth. But first let's watch a single hidden layer actually solve a problem a linear model cannot.
Worked example
XOR is the classic not linearly separable problem: no straight line separates the 1s from the 0s. One hidden layer (width 4, tanh) fits it. Train and read off the predictions:
import torch, torch.nn as nn
torch.manual_seed(42)
X = torch.tensor([[0.,0.],[0.,1.],[1.,0.],[1.,1.]])
y = torch.tensor([[0.],[1.],[1.],[0.]])
model = nn.Sequential(nn.Linear(2,4), nn.Tanh(),
nn.Linear(4,1), nn.Sigmoid())
opt = torch.optim.Adam(model.parameters(), lr=0.1)
for _ in range(3000):
opt.zero_grad()
nn.MSELoss()(model(X), y).backward()
opt.step()
with torch.no_grad():
print(model(X).numpy().round(3).flatten())Predictions [0.0, 0.998, 0.998, 0.003] round to [0, 1, 1, 0]
Why: The four hidden tanh units bend the input space so the output layer can separate the classes. Universal approximation made concrete: a hidden layer represents a function a single Linear cannot.
| input | target | prediction | rounds to |
|---|---|---|---|
| (0,0) | 0 | 0.000 | 0 ✓ |
| (0,1) | 1 | 0.998 | 1 ✓ |
| (1,0) | 1 | 0.998 | 1 ✓ |
| (1,1) | 0 | 0.003 | 0 ✓ |
Discrimination
Sort into buckets
Sort these by target, from memory, without looking back at XOR: one hidden layer solves it. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Anomaly
Predict first
A student writes this, and it looks reasonable:
XOR just needs more capacity — drop the hidden layer and train a single Linear(2,1) harder.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: A Linear + Sigmoid is one linear boundary.
Add a nonlinear hidden layer. Nonlinearity — not size — is what defeats non-separability.
Why: A Linear + Sigmoid is one linear boundary. No line separates XOR, so the best it can do is sit on the fence at 0.5 for all four points. More epochs never fix a representational limit.
Trap
XOR just needs more capacity — drop the hidden layer and train a single Linear(2,1) harder.
import torch, torch.nn as nn
torch.manual_seed(0)
X = torch.tensor([[0.,0.],[0.,1.],[1.,0.],[1.,1.]])
y = torch.tensor([[0.],[1.],[1.],[0.]])
model = nn.Sequential(nn.Linear(2,1), nn.Sigmoid()) # no hidden layer
opt = torch.optim.Adam(model.parameters(), lr=0.1)
for _ in range(3000):
opt.zero_grad(); nn.BCELoss()(model(X), y).backward(); opt.step()
with torch.no_grad():
print(model(X).numpy().round(3).flatten()) # [0.5 0.5 0.5 0.5]Output [0.5, 0.5, 0.5, 0.5] — pure guessing
Why: A Linear + Sigmoid is one linear boundary. No line separates XOR, so the best it can do is sit on the fence at 0.5 for all four points. More epochs never fix a representational limit.
Add a nonlinear hidden layer. Nonlinearity — not size — is what defeats non-separability.
model = nn.Sequential(nn.Linear(2,4), nn.Tanh(),
nn.Linear(4,1), nn.Sigmoid())
# ... same training loop ...
# predictions -> [0.0, 0.998, 0.998, 0.003]Hidden layer → [0, 1, 1, 0]
Why: Four tanh units warp the space until the classes are linearly separable in hidden coordinates. The fix is a nonlinearity between two Linear layers, exactly what the universal approximation theorem promises.
Break the constraint
Discussion prompt
The rule this trap just fixed:
Add a nonlinear hidden layer. Nonlinearity — not size — is what defeats non-separability.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
A Linear + Sigmoid is one linear boundary. No line separates XOR, so the best it can do is sit on the fence at 0.5 for all four points. More epochs never fix a representational limit.
Section
Part 3 of 8 — expressivity
Concept
A depth-L network composes L nonlinear maps: f = σ∘W_L ∘ … ∘ σ∘W_1. Each layer refines a feature hierarchy — edges from pixels, textures from edges, objects from textures.
| strategy | mechanism | practical effect |
|---|---|---|
| one wide layer | interpolation over many bumps | may need exponentially many neurons |
| many narrow layers | function composition | expressivity grows fast with depth |
| deep + moderate width | both | most competition architectures |
Rule of thumb: hierarchical targets (images, sequences) reward depth; a smooth scalar target with no hierarchy is fine with a moderately wide shallow net.
Worked example
A ReLU network is piecewise linear — it carves input space into flat regions. A rough upper bound: one width-w layer makes ~w regions; composing L such layers can multiply them, giving ~wᴸ. Count it:
# ReLU regions grow ~w per layer when stacked -> ~w**L
w = 4
for L in [1, 2, 3, 4]:
print(f'L={L} width {w}: up to ~{w**L} linear regions')Width 4: 1 layer ≈ 4 regions, 4 layers ≈ 256 regions
Why: Adding a layer MULTIPLIES the region count; adding width only ADDS to it. That multiplicative-vs-additive gap is the formal reason depth is more parameter-efficient than width.
| layers L | wᴸ (width 4) | growth vs +width |
|---|---|---|
| 1 | 4 | baseline |
| 2 | 16 | ×4 |
| 3 | 64 | ×4 |
| 4 | 256 | ×4 (exponential) |
Intuition
A wide shallow layer is like describing a face by listing every pixel pattern separately — no shared vocabulary. A deep net first learns 'edge', then reuses 'edge' everywhere to build 'eye', then reuses 'eye' to build 'face'.
Reuse is the whole advantage. The same features feed many higher-level features, so a deep net says far more with far fewer parameters — as long as the target really has that layered structure.
Sorting
Sort into buckets
These are the pieces of Lesson 44: MLP Architectures — Depth, Width & Activations, out of order. Put each one back under the part of the lesson it belongs to.
Section
Part 4 of 8 — audit any architecture
Concept
nn.Linear(d_in, d_out) holds a weight matrix of shape (d_out, d_in) and a bias vector of length d_out. Count both:
\[ \operatorname{params}\big(\text{Linear}(d_{\text{in}}, d_{\text{out}})\big) = \underbrace{d_{\text{in}}\, d_{\text{out}}}_{\text{weights}} + \underbrace{d_{\text{out}}}_{\text{bias}} \]
The + d_out is the bias — one scalar per output neuron. Activations (ReLU, Tanh, GELU, …) have zero parameters. A stack's total is the sum over its Linear layers.
Ranking
Put in order
Put the moves of Params of our 2 → 3 → 1 network into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. 6 weights (the 3×2 matrix) plus 3 biases (one per hidden neuron).
Worked example
Apply the formula to the running network before trusting any library. Two Linear layers:
Layer 1 Linear(2, 3): 2×3 + 3 = 6 + 3 = 9
Why: 6 weights (the 3×2 matrix) plus 3 biases (one per hidden neuron).
Layer 2 Linear(3, 1): 3×1 + 1 = 3 + 1 = 4
Why: 3 weights (the committee) plus 1 output bias. Total = 9 + 4 = 13.
import torch.nn as nn
net = nn.Sequential(nn.Linear(2, 3), nn.ReLU(), nn.Linear(3, 1))
for name, p in net.named_parameters():
print(name, p.numel())
print('total:', sum(p.numel() for p in net.parameters()))torch confirms total = 13
Why: 0.weight=6, 0.bias=3, 2.weight=3, 2.bias=1 → 13. The ReLU at index 1 has no parameters, so it never appears in named_parameters().
| tensor | shape | numel |
|---|---|---|
| 0.weight | (3, 2) | 6 |
| 0.bias | (3,) | 3 |
| 2.weight | (1, 3) | 3 |
| 2.bias | (1,) | 1 |
| total | 13 |
Blank canvas
Draw it
Draw what Params of our 2 → 3 → 1 network just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Worked example
Scale up to a digits-sized classifier (64 input features, 10 classes). Sum the three Linear layers:
import torch.nn as nn
model = nn.Sequential(nn.Linear(64,128), nn.ReLU(),
nn.Linear(128,64), nn.ReLU(),
nn.Linear(64,10))
for m in model:
if isinstance(m, nn.Linear):
print(m, sum(p.numel() for p in m.parameters()))
print('total:', sum(p.numel() for p in model.parameters()))L1: 64×128+128 = 8320; L2: 128×64+64 = 8256; L3: 64×10+10 = 650
Why: Each layer by the formula. The activations add nothing; only the three Linear layers carry parameters.
Total = 8320 + 8256 + 650 = 17226
Why: torch prints 17226 — matching the hand sum exactly. You can now audit any architecture's size on paper.
| layer | formula | params |
|---|---|---|
| Linear(64,128) | 64×128 + 128 | 8 320 |
| Linear(128,64) | 128×64 + 64 | 8 256 |
| Linear(64,10) | 64×10 + 10 | 650 |
| TOTAL | 17 226 |
Blank canvas
Draw it
Draw what Params of a real MLP: [64 → 128 → 64 → 10] just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Linear(64, 128) is a 64×128 weight matrix, so 64 × 128 = 8192 parameters.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Misses the 128-element bias vector.
Add the bias: d_in × d_out + d_out = 64×128 + 128 = 8320.
Why: Misses the 128-element bias vector. That is an off-by-d_out error on EVERY layer; across a deep net it compounds into thousands of uncounted parameters and a wrong size budget.
Trap
Linear(64, 128) is a 64×128 weight matrix, so 64 × 128 = 8192 parameters.
import torch.nn as nn
lin = nn.Linear(64, 128)
print(lin.weight.numel()) # 8192 (weights ONLY)Count only the weight matrix → 8192
Why: Misses the 128-element bias vector. That is an off-by-d_out error on EVERY layer; across a deep net it compounds into thousands of uncounted parameters and a wrong size budget.
Add the bias: d_in × d_out + d_out = 64×128 + 128 = 8320.
import torch.nn as nn
lin = nn.Linear(64, 128)
w = lin.weight.numel(); b = lin.bias.numel()
print(w, b, w + b) # 8192 128 8320Weights + bias → 8320
Why: The bias adds one scalar per output neuron (128 of them). Verified: sum(p.numel() for p in Linear(64,128).parameters()) = 8320. Always add d_out.
Error analysis
Annotate
Walk the callouts on Trap: forgetting the bias in the count. Each one is a place this is easy to get subtly wrong.
Section
Part 5 of 8 — the nonlinear core
Concept
All four map a scalar pre-activation to a scalar output. Their differences in range, zero-centering, and gradient behaviour decide which layer they belong in.
| activation | range | zero-centered? | key issue |
|---|---|---|---|
| sigmoid σ(x) | (0, 1) | no (σ(0)=0.5) | saturates → vanishing gradient |
| tanh(x) | (−1, 1) | yes (tanh(0)=0) | saturates, but less than sigmoid |
| ReLU max(0,x) | [0, ∞) | no | dead neurons if preact stays < 0 |
| GELU x·Φ(x) | ≈(−0.17, ∞) | ≈ yes (GELU(0)=0) | smooth; default in transformers |
Estimation
Predict first
Evaluate all four at x ∈ {−2, −1, 0, 1, 2}, and also the sigmoid derivative σ'(x) = σ(x)(1−σ(x)) — the quantity that governs gradient flow:
Commit before you compute: what does Exact activation values — computed come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: σ'(±2) = 0.1050 already; the max of σ' is only 0.25 (at x=0)
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The sigmoid derivative never exceeds 0.25 anywhere, and shrinks fast toward the tails.
Worked example
Evaluate all four at x ∈ {−2, −1, 0, 1, 2}, and also the sigmoid derivative σ'(x) = σ(x)(1−σ(x)) — the quantity that governs gradient flow:
import torch, torch.nn as nn
x = torch.tensor([-2., -1., 0., 1., 2.])
sig = torch.sigmoid(x)
tanh = torch.tanh(x)
relu = torch.relu(x)
gelu = nn.GELU()(x)
sdv = sig * (1 - sig) # sigmoid derivative
for xi, s, ta, r, g, sd in zip(x.tolist(), sig, tanh, relu, gelu, sdv):
print(f"x={xi:+.0f} sig={s:.4f} tanh={ta:.4f} relu={r:.4f} gelu={g:.4f} sig'={sd:.4f}")σ'(±2) = 0.1050 already; the max of σ' is only 0.25 (at x=0)
Why: The sigmoid derivative never exceeds 0.25 anywhere, and shrinks fast toward the tails. In a deep stack these sub-0.25 factors multiply — the seed of the vanishing gradient.
| x | sigmoid | tanh | relu | gelu | σ'(x) |
|---|---|---|---|---|---|
| −2.0 | 0.1192 | −0.9640 | 0.0000 | −0.0455 | 0.1050 |
| −1.0 | 0.2689 | −0.7616 | 0.0000 | −0.1587 | 0.1966 |
| 0.0 | 0.5000 | 0.0000 | 0.0000 | 0.0000 | 0.2500 |
| 1.0 | 0.7311 | 0.7616 | 1.0000 | 0.8413 | 0.1966 |
| 2.0 | 0.8808 | 0.9640 | 2.0000 | 1.9545 | 0.1050 |
Discrimination
Sort into buckets
Sort these by relu, from memory, without looking back at Exact activation values — computed. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Concept
GELU (used in GPT, BERT, ViT) multiplies the input by the probability that a standard normal is below it — softly gating small values while passing large positives almost unchanged.
\[ \operatorname{GELU}(x) = x\cdot \Phi(x) = x\cdot \tfrac{1}{2}\Big[\,1 + \operatorname{erf}\!\big(\tfrac{x}{\sqrt{2}}\big)\Big] \]
Unlike ReLU it is smooth everywhere and produces small negative outputs for slightly negative inputs (never a hard zero), so a GELU unit never fully 'dies'.
Fill the middle
Fill in the blanks
From GELU from scratch — verified against nn.GELU — one line has had its right-hand side removed. Put it back.
import numpy as np, torch, torch.nn as nn
from scipy.special import erf
def gelu_scratch(x):
return x * 0.5 * (1.0 + erf(x / np.sqrt(2)))
pts = np.array([-1., 0., 1., 2.])
manual = gelu_scratch(pts)
torch_v = nn.GELU()(torch.tensor(pts, dtype=torch.float32)).numpy()
for x, m, t in zip(pts, manual, torch_v):
print(f'x=___ manual=___ torch=___ match=___')
Why: manual is what everything below it consumes, so the wrong expression here fails later and somewhere else. GELU(−1)=−0.158655, GELU(0)=0, GELU(1)=0.841345, GELU(2)=1.954500.
Worked example
Implement GELU straight from the erf formula and confirm it matches PyTorch's exact (non-tanh) implementation:
import numpy as np, torch, torch.nn as nn
from scipy.special import erf
def gelu_scratch(x):
return x * 0.5 * (1.0 + erf(x / np.sqrt(2)))
pts = np.array([-1., 0., 1., 2.])
manual = gelu_scratch(pts)
torch_v = nn.GELU()(torch.tensor(pts, dtype=torch.float32)).numpy()
for x, m, t in zip(pts, manual, torch_v):
print(f'x={x:+.1f} manual={m:.6f} torch={t:.6f} match={abs(m-t)<1e-5}')All four agree to < 1e-5
Why: GELU(−1)=−0.158655, GELU(0)=0, GELU(1)=0.841345, GELU(2)=1.954500. nn.GELU defaults to the exact erf form (not the tanh approximation), so a from-scratch scipy.special.erf build matches it.
| x | GELU (formula) | nn.GELU() | match |
|---|---|---|---|
| −1.0 | −0.158655 | −0.158655 | yes |
| 0.0 | 0.000000 | 0.000000 | yes |
| 1.0 | 0.841345 | 0.841345 | yes |
| 2.0 | 1.954500 | 1.954500 | yes |
Invariant
Step through it
Step through GELU from scratch — verified against nn.GELU one row at a time. One of these columns never changes — find it, and say why it cannot.
Concept
Backprop multiplies one activation derivative per layer. With sigmoid, each factor is ≤ 0.25, so across depth the gradient reaching the early layers shrinks geometrically — they stop learning.
This is the whole reason ReLU replaced sigmoid in hidden layers: ReLU'(z) = 1 for active neurons, so there is no shrinking factor to compound.
Explain it
Discussion prompt
Explain Failure mode 1: vanishing gradients to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Backprop multiplies one activation derivative per layer. With sigmoid, each factor is ≤ 0.25, so across depth the gradient reaching the early layers shrinks geometrically — they stop learning.
Worked example
Take the best case σ' = 0.25 and a realistic case σ' ≈ 0.1966 (at activation ≈ 1), and raise each to the depth. Watch it collapse:
for L in [1, 4, 8, 12]:
print(f'depth {L:2d}: 0.25**{L} = {0.25**L:.3e}')
print('---')
for L in [1, 4, 8, 12]:
print(f'depth {L:2d}: 0.1966**{L} = {0.1966**L:.3e}')At depth 12 the gradient factor is ~5.96e−08 (best case)
Why: Even the most generous σ' = 0.25 gives 0.25¹² ≈ 6×10⁻⁸ — the first layers receive essentially zero signal and training stalls. The realistic case is worse (~3.3×10⁻⁹).
| depth L | 0.25ᴸ (best) | 0.1966ᴸ (typical) |
|---|---|---|
| 1 | 2.500e−01 | 1.966e−01 |
| 4 | 3.906e−03 | 1.494e−03 |
| 8 | 1.526e−05 | 2.232e−06 |
| 12 | 5.960e−08 | 3.334e−09 |
Comparison
Comparison matrix
From Quantify the vanishing gradient: refill the 0.25ᴸ (best) column from what you know. The rest of the table is as it appeared.
| depth L | 0.25ᴸ (best) | 0.1966ᴸ (typical) |
|---|---|---|
| 1 | 2.500e−01 | 1.966e−01 |
| 4 | 3.906e−03 | 1.494e−03 |
| 8 | 1.526e−05 | 2.232e−06 |
| 12 | 5.960e−08 | 3.334e−09 |
Pattern
Predict first
The table runs: 0 | −5.164 | 0.0 | 0 · 1 | −6.220 | 0.0 | 0 · 2 | −3.890 | 0.0 | 0
In Failure mode 2: the dying ReLU, given the rows so far: what is the next one — the row where sample is … (all 6)?
Correct: … (all 6) | all < 0 | 0.0 | 0 (dead)
| sample | pre-activation | ReLU out | grad gate |
|---|---|---|---|
| 0 | −5.164 | 0.0 | 0 |
| 1 | −6.220 | 0.0 | 0 |
| 2 | −3.890 | 0.0 | 0 |
| … (all 6) | all < 0 | 0.0 | 0 (dead) |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. preacts ≈ [−5.16, −6.22, −3.89, −3.80, −3.31, −4.60], all < 0, so the gradient gate is 0 on every sample.
Worked example
A ReLU neuron dies when its pre-activation is negative for every input: output 0, and ReLU'(z) = 0 too, so no gradient ever flows to revive it. Construct one and check its gate on all data:
import numpy as np
np.random.seed(0)
data = np.random.randn(6, 2) # 6 samples, 2 features
w = np.array([-1.0, -1.0]); b = -3.0 # forces preact very negative
preact = data @ w + b
gate = (preact > 0).astype(float) # ReLU'(z): 1 if alive else 0
print('preacts:', preact.round(3))
print('gate :', gate)
print('alive on any sample?', bool(gate.any()))Every pre-activation is negative → gate is all zeros
Why: preacts ≈ [−5.16, −6.22, −3.89, −3.80, −3.31, −4.60], all < 0, so the gradient gate is 0 on every sample. The neuron outputs 0 forever and receives no update — permanently dead.
| sample | pre-activation | ReLU out | grad gate |
|---|---|---|---|
| 0 | −5.164 | 0.0 | 0 |
| 1 | −6.220 | 0.0 | 0 |
| 2 | −3.890 | 0.0 | 0 |
| … (all 6) | all < 0 | 0.0 | 0 (dead) |
Invariant
Step through it
Step through Failure mode 2: the dying ReLU one row at a time. One of these columns never changes — find it, and say why it cannot.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Sigmoid is a classic activation — use it in every hidden layer so activations stay in (0, 1).
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Each backprop step multiplies by σ'(x) ≤ 0.25.
Use ReLU or GELU in hidden layers; reserve sigmoid for the output of binary classification, where saturation is the desired probability clamp.
Why: Each backprop step multiplies by σ'(x) ≤ 0.25. After 12 layers the gradient is ≤ 0.25¹² ≈ 6×10⁻⁸ — the early layers get no signal and training freezes.
Trap
Sigmoid is a classic activation — use it in every hidden layer so activations stay in (0, 1).
Stack 12 sigmoid hidden layers, train with SGD
Why: Each backprop step multiplies by σ'(x) ≤ 0.25. After 12 layers the gradient is ≤ 0.25¹² ≈ 6×10⁻⁸ — the early layers get no signal and training freezes.
# effective gradient scale into the first layer, 12 sigmoid layers
print(0.25**12) # 5.96e-08 -> vanishedUse ReLU or GELU in hidden layers; reserve sigmoid for the output of binary classification, where saturation is the desired probability clamp.
Hidden: ReLU / GELU / tanh · Output: sigmoid (binary) or logits
Why: ReLU's gradient is 1 for active neurons — no compounding decay. Sigmoid is safe only in a single output layer where you WANT (0,1) squashing.
# ReLU gradient does not compound across depth
print(1.0**12) # 1.0 -> signal survivesTwo truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
x = [1, 2]. Row j of W⁽¹⁾ holds the two weights of hidden neuron j. Keep this table of numbers in view — the next four slides trace it end to end.; The output layer is a weighted committee of those detectors. Learning means tuning who votes, how strongly each detector responds, and how the committee combines them.; Rule of thumb: hierarchical targets (images, sequences) reward depth; a smooth scalar target with no hierarchy is fine with a moderately wide shallow net.Linear(2,1) harder.; Linear(64, 128) is a 64×128 weight matrix, so 64 × 128 = 8192 parameters.Section
Part 6 of 8 — equal parameter budgets
Missing information
Discussion prompt
A fair comparison holds the parameter count fixed and varies only the shape. Compute the budgets by hand: a wide net 64→512→10 vs a deep net (9 hidden layers of width 64, then 64→10).
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Within 1% of each other — so any accuracy gap is due to shape, not size. The wide net spends its budget on one fat layer; the deep net spreads it across ten thin ones.
Worked example
A fair comparison holds the parameter count fixed and varies only the shape. Compute the budgets by hand: a wide net 64→512→10 vs a deep net (9 hidden layers of width 64, then 64→10).
# wide: Linear(64,512) + Linear(512,10)
w1 = 64*512 + 512
w2 = 512*10 + 10
print('wide:', w1, '+', w2, '=', w1 + w2)
# deep: 9 x Linear(64,64) + Linear(64,10)
d_hidden = 9 * (64*64 + 64)
d_out = 64*10 + 10
print('deep:', d_hidden, '+', d_out, '=', d_hidden + d_out)Wide = 33280 + 5130 = 38410; Deep = 37440 + 650 = 38090
Why: Within 1% of each other — so any accuracy gap is due to shape, not size. The wide net spends its budget on one fat layer; the deep net spreads it across ten thin ones.
| model | shape | params |
|---|---|---|
| wide | 64 → 512 → 10 | 38 410 |
| deep | 64 → (64)×9 → 10 | 38 090 |
| gap | < 1% |
Estimation
Predict first
Train both on load_digits (1 797 samples, 64 features, 10 classes), same optimizer, same seed, same 400 epochs. This block is self-contained:
Commit before you compute: what does Wide vs deep on load_digits come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: wide → 97.50% test acc; deep → 94.72% test acc
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. On flat tabular digits the wide net WINS.
Worked example
Train both on load_digits (1 797 samples, 64 features, 10 classes), same optimizer, same seed, same 400 epochs. This block is self-contained:
import torch, torch.nn as nn, numpy as np
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
X, y = load_digits(return_X_y=True)
Xtr, Xte, ytr, yte = train_test_split(X.astype(np.float32), y, test_size=0.2, random_state=42)
sc = StandardScaler(); Xtr = sc.fit_transform(Xtr); Xte = sc.transform(Xte)
Xtr_t, ytr_t = torch.tensor(Xtr), torch.tensor(ytr, dtype=torch.long)
Xte_t, yte_t = torch.tensor(Xte), torch.tensor(yte, dtype=torch.long)
torch.manual_seed(0)
wide = nn.Sequential(nn.Linear(64,512), nn.ReLU(), nn.Linear(512,10))
torch.manual_seed(0)
layers = []
for i in range(10):
layers += [nn.Linear(64,64), nn.ReLU()] if i < 9 else [nn.Linear(64,10)]
deep = nn.Sequential(*layers)
for name, m in [('wide', wide), ('deep', deep)]:
opt = torch.optim.Adam(m.parameters(), lr=0.005)
for _ in range(400):
opt.zero_grad(); nn.CrossEntropyLoss()(m(Xtr_t), ytr_t).backward(); opt.step()
acc = (m(Xte_t).argmax(1) == yte_t).float().mean().item()
print(f'{name}: params={sum(p.numel() for p in m.parameters())} acc={acc*100:.2f}%')wide → 97.50% test acc; deep → 94.72% test acc
Why: On flat tabular digits the wide net WINS. There is no spatial hierarchy for depth to exploit, and the deep ReLU stack is slightly harder to optimize. The result is task-dependent, not a universal law.
| model | params | init loss | final loss | test acc |
|---|---|---|---|---|
| wide (1 hidden, w=512) | 38 410 | 2.3481 | 0.0002 | 97.50% |
| deep (10 hidden, w=64) | 38 090 | 2.3047 | 0.0001 | 94.72% |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
wide → 97.50% test acc; deep → 94.72% test acc
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Train both on load_digits (1 797 samples, 64 features, 10 classes), same optimizer, same seed, same 400 epochs. This block is self-contained:
Concept
| scenario | recommended choice | reason |
|---|---|---|
| flat tabular, < 1 000 features | 2–3 layers, moderate width | no hierarchy; a wide-ish shallow net interpolates well |
| image / sequence (hierarchical) | deep, narrow (ResNet-style) | each layer learns finer structure |
| tight parameter budget | prefer depth | depth is more parameter-efficient on structured data |
| deep sigmoid stalls | swap to ReLU / add BatchNorm / residuals | avoid compounding sub-0.25 gradients |
| very deep net diverges | reduce depth first, then retune lr | depth × lr interact — start shallow, grow |
The digits result is the guideline in action: match the architecture to the data's structure, not to a slogan like 'deeper is always better'.
Analogy
Discussion prompt
Explain USAAIO architecture-design guidelines by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
The digits result is the guideline in action: match the architecture to the data's structure, not to a slogan like 'deeper is always better'.
Section
Part 7 of 8
Constraint
Discussion prompt
Run The MLP architecture recipe with this step confiscated:
Output activation: none/logits for CrossEntropyLoss · sigmoid for binary · softmax only if you need explicit probabilities
Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.
Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.
Answer:
z = Wh + b, then h = σ(z), layer by layer — confirm shapes and one numeric output against torchd_in × d_out + d_out; sum across all nn.Linear (activations add zero)CrossEntropyLoss · sigmoid for binary · softmax only if you need explicit probabilitiesPattern
z = Wh + b, then h = σ(z), layer by layer — confirm shapes and one numeric output against torchd_in × d_out + d_out; sum across all nn.Linear (activations add zero)CrossEntropyLoss · sigmoid for binary · softmax only if you need explicit probabilitiesEdge cases
Discussion prompt
The MLP architecture recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
z = Wh + b, then h = σ(z), layer by layer — confirm shapes and one numeric output against torchd_in × d_out + d_out; sum across all nn.Linear (activations add zero)CrossEntropyLoss · sigmoid for binary · softmax only if you need explicit probabilitiesElimination
Eliminate the wrong options
With W⁽²⁾ = [1, −1, 2], b⁽²⁾ = 0.5, and hidden output h⁽¹⁾ = [0, 0.5, 0.5], what is the network's output?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: z⁽²⁾ = (1)(0) + (−1)(0.5) + (2)(0.5) + 0.5 = 0 − 0.5 + 1.0 + 0.5 = 1.0. The dead neuron 0 drops out; neuron 1 subtracts 0.5, neuron 2 adds 1.0, and the bias adds 0.5.
Check
Use the running network: x = [1, 2], ReLU hidden, and h⁽¹⁾ = [0, 0.5, 0.5].
Check your understanding
With W⁽²⁾ = [1, −1, 2], b⁽²⁾ = 0.5, and hidden output h⁽¹⁾ = [0, 0.5, 0.5], what is the network's output?
Answer: A
Why: z⁽²⁾ = (1)(0) + (−1)(0.5) + (2)(0.5) + 0.5 = 0 − 0.5 + 1.0 + 0.5 = 1.0. The dead neuron 0 drops out; neuron 1 subtracts 0.5, neuron 2 adds 1.0, and the bias adds 0.5.
Prediction
Predict first
How many trainable parameters does nn.Linear(128, 64) have?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: 8 256 (128×64 + 64)
Why: d_in × d_out + d_out = 128×64 + 64 = 8192 + 64 = 8256. The +64 is the bias, one scalar per output neuron. Verified: sum(p.numel() for p in nn.Linear(128,64).parameters()) = 8256.
Check
Apply d_in × d_out + d_out before looking.
Check your understanding
How many trainable parameters does nn.Linear(128, 64) have?
Answer: A
Why: d_in × d_out + d_out = 128×64 + 64 = 8192 + 64 = 8256. The +64 is the bias, one scalar per output neuron. Verified: sum(p.numel() for p in nn.Linear(128,64).parameters()) = 8256.
Prediction
Predict first
You build a 12-layer MLP. Which hidden-layer activation is most likely to stall training because gradients vanish?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: sigmoid — derivative ≤ 0.25 compounds to near-zero across layers
Why: σ'(x) ≤ 0.25 everywhere. Across 12 layers the gradient into the first layer is ≤ 0.25¹² ≈ 6×10⁻⁸ — effectively zero. ReLU's gradient is 1 for active neurons, so it does not compound.
Check
Think about what each derivative does across depth.
Check your understanding
You build a 12-layer MLP. Which hidden-layer activation is most likely to stall training because gradients vanish?
Answer: A
Why: σ'(x) ≤ 0.25 everywhere. Across 12 layers the gradient into the first layer is ≤ 0.25¹² ≈ 6×10⁻⁸ — effectively zero. ReLU's gradient is 1 for active neurons, so it does not compound.
Elimination
Eliminate the wrong options
On load_digits at a ~38k parameter budget, the wide model (1 hidden, width 512) hit 97.5% and the deep model (10 hidden, width 64) hit 94.7%. Best interpretation?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: load_digits is a flat 64-feature problem with no spatial hierarchy, so a single wide layer interpolates efficiently and the deep ReLU stack is slightly harder to optimize. Depth shines on hierarchical data (images, sequences) — the result is dataset-dependent.
Check
Use this lesson's experiment result.
Check your understanding
On load_digits at a ~38k parameter budget, the wide model (1 hidden, width 512) hit 97.5% and the deep model (10 hidden, width 64) hit 94.7%. Best interpretation?
Answer: A
Why: load_digits is a flat 64-feature problem with no spatial hierarchy, so a single wide layer interpolates efficiently and the deep ReLU stack is slightly harder to optimize. Depth shines on hierarchical data (images, sequences) — the result is dataset-dependent.
Section
Part 8 of 8 — the project
Concept
Build MLPs of different shapes on load_digits, count their parameters exactly, and reimplement GELU from the formula — proving every number against torch and scipy.
| # | requirement | tool / formula |
|---|---|---|
| 1 | count params for any Linear stack | d_in×d_out + d_out, summed |
| 2 | implement GELU from scratch | scipy.special.erf |
| 3 | train the [64→128→64→10] MLP, score test acc | Adam, CrossEntropyLoss, load_digits |
Build rules: type every line yourself, predict each number before printing, and verify every count with sum(p.numel() for p in model.parameters()).
Counterexample
Discussion prompt
Build MLPs of different shapes on load_digits, count their parameters exactly, and reimplement GELU from the formula — proving every number against torch and scipy.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Build rules: type every line yourself, predict each number before printing, and verify every count with sum(p.numel() for p in model.parameters()).
Worked example
Your turn: predict the total for [64 → 128 → 64 → 10] before running. Write each layer's count first.
Hint: apply d_in × d_out + d_out to each Linear and sum the three. The activations contribute nothing.
import torch.nn as nn
fc1 = nn.Linear(64, 128); fc2 = nn.Linear(128, 64); fc3 = nn.Linear(64, 10)
for name, lyr in [('fc1', fc1), ('fc2', fc2), ('fc3', fc3)]:
print(name, sum(p.numel() for p in lyr.parameters()))
total = sum(p.numel() for lyr in (fc1, fc2, fc3) for p in lyr.parameters())
print('total:', total)| layer | formula | count |
|---|---|---|
| fc1 Linear(64,128) | 64×128 + 128 | 8 320 |
| fc2 Linear(128,64) | 128×64 + 64 | 8 256 |
| fc3 Linear(64,10) | 64×10 + 10 | 650 |
| total | 17 226 |
Worked example
Your turn: implement gelu(x) with scipy.special.erf. Predict the value at x = 1 before you run it.
Hint: GELU(x) = x · 0.5 · (1 + erf(x / √2)). At x = 1: 1 · 0.5 · (1 + erf(0.707)) ≈ 0.841.
import numpy as np, torch, torch.nn as nn
from scipy.special import erf
def gelu(x):
return x * 0.5 * (1.0 + erf(x / np.sqrt(2)))
pts = np.array([-1., 0., 1., 2.])
manual = gelu(pts)
torch_v = nn.GELU()(torch.tensor(pts, dtype=torch.float32)).numpy()
for x, m, t in zip(pts, manual, torch_v):
print(f'x={x:+.1f} scratch={m:.6f} nn.GELU={t:.6f}')| x | gelu (scratch) | nn.GELU() | match |
|---|---|---|---|
| −1.0 | −0.158655 | −0.158655 | yes |
| 0.0 | 0.000000 | 0.000000 | yes |
| 1.0 | 0.841345 | 0.841345 | yes |
| 2.0 | 1.954500 | 1.954500 | yes |
Invariant
Step through it
Step through Milestone 2 — GELU from scratch one row at a time. One of these columns never changes — find it, and say why it cannot.
Worked example
Your turn: train the [64→128→64→10] MLP on scaled load_digits and read off the test accuracy. Predict whether it clears 95%.
Hint: standardize features, use Adam(lr=0.005), CrossEntropyLoss, 400 epochs, then compare argmax(1) to the labels.
import torch, torch.nn as nn, numpy as np
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
X, y = load_digits(return_X_y=True)
Xtr, Xte, ytr, yte = train_test_split(X.astype(np.float32), y, test_size=0.2, random_state=42)
sc = StandardScaler(); Xtr = sc.fit_transform(Xtr); Xte = sc.transform(Xte)
torch.manual_seed(0)
model = nn.Sequential(nn.Linear(64,128), nn.ReLU(), nn.Linear(128,64), nn.ReLU(), nn.Linear(64,10))
opt = torch.optim.Adam(model.parameters(), lr=0.005)
for _ in range(400):
opt.zero_grad(); nn.CrossEntropyLoss()(model(torch.tensor(Xtr)), torch.tensor(ytr, dtype=torch.long)).backward(); opt.step()
with torch.no_grad():
acc = (model(torch.tensor(Xte)).argmax(1) == torch.tensor(yte)).float().mean()
print('test acc:', round(acc.item(), 4))| quantity | value (verified) |
|---|---|
| params | 17 226 |
| test acc | 0.9694 (≈ 97%) |
| clears 95%? | yes |
Trade off
Comparison matrix
From Milestone 3 — train and score: every row here is a choice with a cost. Fill the value (verified) column, then say which row you would actually pick and what you give up for it.
| quantity | value (verified) |
|---|---|
| params | 17 226 |
| test acc | 0.9694 (≈ 97%) |
| clears 95%? | yes |
Concept
import torch, torch.nn as nn, numpy as np
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from scipy.special import erf
torch.manual_seed(0)
model = nn.Sequential(nn.Linear(64,128), nn.ReLU(),
nn.Linear(128,64), nn.ReLU(), nn.Linear(64,10))
print('params:', sum(p.numel() for p in model.parameters())) # 17226
gelu = lambda x: x * 0.5 * (1 + erf(x / np.sqrt(2)))
print('GELU(1):', round(float(gelu(1.0)), 6)) # 0.841345
X, y = load_digits(return_X_y=True)
Xtr, Xte, ytr, yte = train_test_split(X.astype(np.float32), y, test_size=0.2, random_state=42)
sc = StandardScaler(); Xtr = sc.fit_transform(Xtr); Xte = sc.transform(Xte)
opt = torch.optim.Adam(model.parameters(), lr=0.005)
for _ in range(400):
opt.zero_grad(); nn.CrossEntropyLoss()(model(torch.tensor(Xtr)), torch.tensor(ytr, dtype=torch.long)).backward(); opt.step()
with torch.no_grad():
acc = (model(torch.tensor(Xte)).argmax(1) == torch.tensor(yte)).float().mean()
print('test acc:', round(acc.item(), 4)) # 0.9694| printed line | value |
|---|---|
| params: | 17226 |
| GELU(1): | 0.841345 |
| test acc: | 0.9694 |
If your count reads 17226, GELU(1) is 0.841345, and the net clears 96% — you can audit any MLP's size, activation math, and capacity from first principles.
Comparison
Comparison matrix
From The full program: refill the value column from what you know. The rest of the table is as it appeared.
| printed line | value |
|---|---|
| params: | 17226 |
| GELU(1): | 0.841345 |
| test acc: | 0.9694 |
Concept
Slides closed, out loud: (1) forward-propagate x = [1, 2] through the running net to the output 1.0, (2) state what universal approximation guarantees and what it does not, (3) count parameters for [64→128→64→10] without code, (4) explain why 12 stacked sigmoids stall training.
Stretch: rerun the wide-vs-deep experiment with GELU instead of ReLU and with a residual connection every two layers — see whether depth catches up once optimization is easier. Next: Lesson 45 — convolutional layers (spatial inductive bias).
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — What an MLP actually computes · Universal approximation — and its catch · Depth vs width — why depth wins · Parameter counting — exact formula · Activation functions — outputs & gradients · Deep vs wide — a fair experiment. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
z = Wh + b, then σ — and match it to torch (our net: x=[1,2] → 1.0)~wᴸ) while width only adds themd_in × d_out + d_out (our nets: 13, 17 226, 38 410, 38 090)97.5% > deep 94.7%) and pick a shape from the data's structure| concept | the one thing to remember |
|---|---|
| forward pass | z = Wh + b, then h = σ(z); ReLU silences negatives |
| universal approx. | one wide layer CAN — but depth is exponentially more efficient |
| param count | d_in × d_out + d_out; bias adds d_out per layer |
| sigmoid hidden | σ' ≤ 0.25 per layer — never stack sigmoid deep |
| GELU | x · 0.5 · (1 + erf(x/√2)); GELU(1) = 0.841345 |
| wide vs deep | digits: wide 97.5% > deep 94.7% — match shape to structure |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.