USAAIO Lesson 82, from Phase 3 on transformers. It builds LayerNorm from scratch, normalizing over the feature dimension and verifying against nn.LayerNorm, then compares pre-norm with post-norm for stability. It covers the feed-forward sublayer - two Linear layers with a GELU between them, at d_ff = 4·d_model - and residual connections for gradient flow, then assembles the complete pre-norm EncoderBlock as an nn.Module. All the numbers were verified with torch 2.7.1+cpu in June 2026. The lesson runs to 30 slides.
Subject: Machine Learning · 60 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 82 · Phase 3 (Transformers)
The two sublayers that wrap multi-head attention: LayerNorm (normalize over features, not batch), GELU feed-forward network, residual connections, and the complete pre-norm EncoderBlock as a reusable nn.Module.
Objectives
nn.LayerNormNorm → sublayer) and post-norm and state which is more stableLinear layers + GELU, d_ff = 4·d_model)nn.Module with residual connections and both sublayersWarm-up
Discussion prompt
Before we open Lesson 82: LayerNorm, FFN with GELU, and the Full Encoder Block: without looking back, what was the main idea of Multi-Head Attention, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
multi-head attention from scratch — projecting Q/K/V into h parallel subspaces, scaled dot-product per head, concat + W_O projection, parameter count proof (4·d_model² regardless of h), and h=1 equivalence to single-head attention. Implement MultiHeadAttention as nn.Module and verify every shape.
Section
Part 1 of 3
Concept
LayerNorm normalizes each token's feature vector independently: compute the mean and variance across the d_model features for that one token, then rescale.
\[ \hat{x}_i = \frac{x_i - \mu}{\sqrt{\sigma^2 + \varepsilon}}, \quad \mu = \frac{1}{d}\sum_{i=1}^d x_i, \quad \sigma^2 = \frac{1}{d}\sum_{i=1}^d (x_i - \mu)^2 \]
After normalization, learnable parameters γ (scale) and β (shift) — both of shape d_model — let the network undo the normalization if needed. nn.LayerNorm trains γ and β automatically.
Counterexample
Discussion prompt
LayerNorm normalizes each token's feature vector independently: compute the mean and variance across the d_model features for that one token, then rescale.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
After normalization, learnable parameters γ (scale) and β (shift) — both of shape d_model — let the network undo the normalization if needed. nn.LayerNorm trains γ and β automatically.
Concept
BatchNorm normalizes across the batch (per feature channel). LayerNorm normalizes across features (per token). For sequences, batch sizes vary and sequence lengths vary — BatchNorm statistics become unreliable.
| norm | mean/var computed over | suitable for |
|---|---|---|
| BatchNorm | batch dimension (N samples) | fixed-size CNN feature maps |
| LayerNorm | feature dimension (d_model) | variable-length sequences, NLP |
For a token embedding x ∈ ℝ^d, LN's statistics are entirely self-contained — independent of what other tokens or other sequences are in the batch (Lesson 17 CE context).
Comparison
Comparison matrix
From LayerNorm vs BatchNorm: the axis difference: refill the mean/var computed over column from what you know. The rest of the table is as it appeared.
| norm | mean/var computed over | suitable for |
|---|---|---|
| BatchNorm | batch dimension (N samples) | fixed-size CNN feature maps |
| LayerNorm | feature dimension (d_model) | variable-length sequences, NLP |
Estimation
Predict first
Normalize x = [1, 3, 5, 7] manually, then verify against nn.LayerNorm.
Commit before you compute: what does LayerNorm from scratch come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: mean = 4.0; var = 5.0; std = 2.2361
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. mean = (1+3+5+7)/4; var = [(1-4)² + (3-4)² + (5-4)² + (7-4)²] / 4 = (9+1+1+9)/4 = 5.0; std = √5 ≈ 2.2361
Worked example
Normalize x = [1, 3, 5, 7] manually, then verify against nn.LayerNorm.
import torch, torch.nn as nn
x = torch.tensor([[1.0, 3.0, 5.0, 7.0]])
mean = x.mean(dim=-1, keepdim=True) # 4.0
var = x.var(dim=-1, keepdim=True, unbiased=False) # 5.0
xhat = (x - mean) / (var + 1e-5).sqrt() # no gamma/beta yet
ln = nn.LayerNorm(4, elementwise_affine=False, eps=1e-5)
print('scratch:', xhat.tolist())
print('nn.LN: ', ln(x).tolist())
print('match: ', torch.allclose(xhat, ln(x), atol=1e-4))mean = 4.0; var = 5.0; std = 2.2361
Why: mean = (1+3+5+7)/4; var = [(1-4)² + (3-4)² + (5-4)² + (7-4)²] / 4 = (9+1+1+9)/4 = 5.0; std = √5 ≈ 2.2361
| x_i | x_i − μ | (x_i − μ) / std | xhat_i |
|---|---|---|---|
| 1.0 | −3.0 | −3.0 / 2.2361 | −1.3416 |
| 3.0 | −1.0 | −1.0 / 2.2361 | −0.4472 |
| 5.0 | 1.0 | 1.0 / 2.2361 | 0.4472 |
| 7.0 | 3.0 | 3.0 / 2.2361 | 1.3416 |
Trade off
Comparison matrix
From LayerNorm from scratch: every row here is a choice with a cost. Fill the xhat_i column, then say which row you would actually pick and what you give up for it.
| x_i | x_i − μ | (x_i − μ) / std | xhat_i |
|---|---|---|---|
| 1.0 | −3.0 | −3.0 / 2.2361 | −1.3416 |
| 3.0 | −1.0 | −1.0 / 2.2361 | −0.4472 |
| 5.0 | 1.0 | 1.0 / 2.2361 | 0.4472 |
| 7.0 | 3.0 | 3.0 / 2.2361 | 1.3416 |
Anomaly
Predict first
A student writes this, and it looks reasonable:
LayerNorm over a batch of two sequences [[1,2,3,4],[2,4,6,8]] is computed per feature column — same as BatchNorm.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: This is BatchNorm's statistics — averaging down the batch axis per feature.
LayerNorm normalizes each row (each token vector) independently: dim=-1, not dim=0.
Why: This is BatchNorm's statistics — averaging down the batch axis per feature. LayerNorm never touches the batch axis.
Trap
LayerNorm over a batch of two sequences [[1,2,3,4],[2,4,6,8]] is computed per feature column — same as BatchNorm.
mean_col = [(1+2)/2, (2+4)/2, (3+6)/2, (4+8)/2] = [1.5, 3.0, 4.5, 6.0]
Why: This is BatchNorm's statistics — averaging down the batch axis per feature. LayerNorm never touches the batch axis.
LayerNorm normalizes each row (each token vector) independently: dim=-1, not dim=0.
Row 0: mean = 2.5, std = 1.118 → xhat = [−1.34, −0.45, 0.45, 1.34]; Row 1: same ratios, identical xhat
Why: Each token is normalized by its own statistics. Both rows produce the same xhat because they have the same shape — [1,2,3,4] scaled by 2. BatchNorm would give different results because it mixes the two samples.
Break the constraint
Discussion prompt
The rule this trap just fixed:
LayerNorm normalizes each row (each token vector) independently: dim=-1, not dim=0.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
This is BatchNorm's statistics — averaging down the batch axis per feature. LayerNorm never touches the batch axis.
Section
Part 2 of 3
Concept
Two orderings exist for inserting LayerNorm relative to a sublayer (MHA or FFN):
| variant | formula | training stability |
|---|---|---|
| Post-norm (original) | x = LayerNorm(x + Sublayer(x)) | harder to train deep; norms after residual |
| Pre-norm (modern) | x = x + Sublayer(LayerNorm(x)) | more stable; norm before sublayer, residual path clean |
Pre-norm keeps the residual stream unnormalized, so gradients flow back through the + x path without touching the norm. This is why GPT-2, GPT-3, LLaMA, and most modern transformers use pre-norm.
Analogy
Discussion prompt
Explain Pre-norm: the modern default by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Two orderings exist for inserting LayerNorm relative to a sublayer (MHA or FFN):
Concept
The residual formula output = x + Sublayer(Norm(x)) means the network learns a correction (the residual) on top of the identity path.
\[ \frac{\partial \mathcal{L}}{\partial x} = \frac{\partial \mathcal{L}}{\partial(x + F(x))} \cdot \left(1 + \frac{\partial F}{\partial x}\right) \]
The 1 in the gradient ensures that, even if ∂F/∂x ≈ 0 in early training, a full gradient signal still flows to every earlier layer (Lesson 9 backprop context). Without residuals, deep networks suffer vanishing gradients.
Explain it
Discussion prompt
Explain Residual connections: the gradient highway to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
The residual formula output = x + Sublayer(Norm(x)) means the network learns a correction (the residual) on top of the identity path.
Missing information
Discussion prompt
Train an 8-layer network with and without residual connections. Compare the gradient norm reaching layer 1 (deepest backprop path).
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Without residuals the gradient degrades through tanh saturation. With residuals the + identity path keeps the gradient magnitude healthy in all layers — the foundation for training deep transformers.
Worked example
Train an 8-layer network with and without residual connections. Compare the gradient norm reaching layer 1 (deepest backprop path).
import torch, torch.nn as nn
torch.manual_seed(7)
depth, d = 8, 16
class NoRes(nn.Module):
def __init__(self):
super().__init__()
self.layers = nn.ModuleList([nn.Linear(d, d) for _ in range(depth)])
for l in self.layers: nn.init.normal_(l.weight, std=0.3)
def forward(self, x):
for l in self.layers: x = torch.tanh(l(x))
return x
class WithRes(nn.Module):
def __init__(self):
super().__init__()
self.layers = nn.ModuleList([nn.Linear(d, d) for _ in range(depth)])
def forward(self, x):
for l in self.layers: x = x + torch.tanh(l(x))
return x
x = torch.randn(4, d)
for Cls, name in [(NoRes,'no-res'), (WithRes,'residual')]:
model = Cls(); model(x).sum().backward()
gnorms = [l.weight.grad.norm().item() for l in model.layers]
print(f'{name}: L1={gnorms[0]:.3f}, L8={gnorms[-1]:.3f}')no-res: L1=15.07, L8=9.79; residual: L1=29.27, L8=31.16
Why: Without residuals the gradient degrades through tanh saturation. With residuals the + identity path keeps the gradient magnitude healthy in all layers — the foundation for training deep transformers.
| layer | no-residual ‖∇‖ | with-residual ‖∇‖ |
|---|---|---|
| L1 (earliest) | 15.0687 | 29.2718 |
| L2 | 13.1109 | 22.8678 |
| L4 | 11.9878 | 23.5143 |
| L6 | 13.6133 | 23.4945 |
| L8 (latest) | 9.7869 | 31.1626 |
Pattern
Step through it
Step through Residual gradient flow: with vs without one row at a time. What is driving the change, and what would the row after the last one be?
Section
Part 3 of 3
Concept
Each encoder position independently passes through a two-layer MLP: expand to d_ff = 4·d_model, apply GELU, project back. The expansion and contraction happen at every token in parallel.
\[ \text{FFN}(x) = W_2\, \text{GELU}(W_1 x + b_1) + b_2, \quad W_1 \in \mathbb{R}^{d_{ff} \times d}, \; W_2 \in \mathbb{R}^{d \times d_{ff}} \]
Why 4×? Empirical finding from the original transformer paper. The extra width gives the network capacity to store factual associations (Lesson 82 onwards). Why GELU over ReLU? Smoother gradient near zero; small negative values not completely zeroed out.
Socratic
Discussion prompt
Each encoder position independently passes through a two-layer MLP: expand to d_ff = 4·d_model, apply GELU, project back. The expansion and contraction happen at every token in parallel.
Suppose that were not true. What is the first thing in Lesson 82: LayerNorm, FFN with GELU, and the Full Encoder Block that would stop working?
Hint: Follow it one step downstream. The answer is whatever was quietly relying on it.
Concept
GELU (Gaussian Error Linear Unit) weights the input by the probability it is positive: GELU(x) ≈ x · Φ(x) where Φ is the standard normal CDF. It smoothly suppresses small or negative inputs.
| x | GELU(x) | ReLU(x) | difference |
|---|---|---|---|
| −2.0 | −0.0455 | 0.0000 | GELU leaks small negative |
| −1.0 | −0.1587 | 0.0000 | GELU not exactly zero |
| −0.5 | −0.1543 | 0.0000 | smooth suppression |
| 0.0 | 0.0000 | 0.0000 | both zero at origin |
| 0.5 | 0.3457 | 0.5000 | GELU slightly below ReLU |
| 1.0 | 0.8413 | 1.0000 | approaches identity |
| 2.0 | 1.9545 | 2.0000 | ≈ identity for large x |
The smooth gradient at x < 0 avoids the dead neuron problem of ReLU (Lesson 44 activation context). BERT, GPT-2, and GPT-3 all use GELU.
Explain it
Discussion prompt
Explain GELU vs ReLU: the smooth gate to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
GELU (Gaussian Error Linear Unit) weights the input by the probability it is positive: GELU(x) ≈ x · Φ(x) where Φ is the standard normal CDF. It smoothly suppresses small or negative inputs.
Estimation
Predict first
Build the FFN as an nn.Module. Use d_model=4, d_ff=16 (4×). Verify parameter count and output shape.
Commit before you compute: what does FFN sublayer in PyTorch come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: output shape (1, 4) — same as input; params = 4·16+16 + 16·4+4 = 80+68 = 148
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. W1: d_model×d_ff + d_ff = 4·16+16 = 80.
Worked example
Build the FFN as an nn.Module. Use d_model=4, d_ff=16 (4×). Verify parameter count and output shape.
import torch, torch.nn as nn, torch.nn.functional as F
torch.manual_seed(0)
class FFN(nn.Module):
def __init__(self, d_model, d_ff):
super().__init__()
self.w1 = nn.Linear(d_model, d_ff)
self.w2 = nn.Linear(d_ff, d_model)
def forward(self, x):
return self.w2(F.gelu(self.w1(x)))
ffn = FFN(d_model=4, d_ff=16)
x = torch.tensor([[1.0, 0.0, -1.0, 0.5]])
hid = F.gelu(ffn.w1(x))
out = ffn(x)
params = sum(p.numel() for p in ffn.parameters())
print(f'hidden shape: {hid.shape}') # (1, 16)
print(f'output shape: {out.shape}') # (1, 4)
print(f'output: {out.squeeze().tolist()}') # [-0.3974, 0.2011, 0.0934, -0.1582]
print(f'params: {params}') # 148output shape (1, 4) — same as input; params = 4·16+16 + 16·4+4 = 80+68 = 148
Why: W1: d_model×d_ff + d_ff = 4·16+16 = 80. W2: d_ff×d_model + d_model = 16·4+4 = 68. Total = 148. The FFN expands then contracts, restoring d_model for the residual add.
| step | tensor shape | note |
|---|---|---|
| input x | (1, 4) | one token, d_model=4 |
| W1(x) | (1, 16) | expand to d_ff=4×d_model |
| GELU(W1(x)) | (1, 16) | smooth gating in-place |
| W2(GELU(W1(x))) | (1, 4) | contract back to d_model |
| FFN params | — | 80 + 68 = 148 total |
Discrimination
Sort into buckets
Sort these by tensor shape, from memory, without looking back at FFN sublayer in PyTorch. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Anomaly
Predict first
A student writes this, and it looks reasonable:
The FFN processes the full sequence at once with shared position-sensitive weights — similar to a recurrent layer connecting adjacent positions.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: This is wrong — the FFN weight matrix W1 is (d_ff, d_model), applied identically and independently at every position.
The FFN is position-wise: the same W1 and W2 are applied independently at every token position.
Why: This is wrong — the FFN weight matrix W1 is (d_ff, d_model), applied identically and independently at every position. There is no cross-position communication in the FFN.
Trap
The FFN processes the full sequence at once with shared position-sensitive weights — similar to a recurrent layer connecting adjacent positions.
Feed W1 a matrix of shape (T, d_model) and assume different positions share information through the weight
Why: This is wrong — the FFN weight matrix W1 is (d_ff, d_model), applied identically and independently at every position. There is no cross-position communication in the FFN.
The FFN is position-wise: the same W1 and W2 are applied independently at every token position.
Input: (B, T, d_model) → W1 applied as (B·T, d_model) @ W1.T → (B, T, d_ff) → GELU → W2 → (B, T, d_model)
Why: nn.Linear broadcasts over leading dimensions, so passing (B, T, d_model) is valid — it processes each of the B·T tokens independently through the same weights. Cross-token interactions happen only in MHA, not the FFN.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
d_model features for that one token, then rescale.; For a token embedding x ∈ ℝ^d, LN's statistics are entirely self-contained — independent of what other tokens or other sequences are in the batch (Lesson 17 CE context).; Two orderings exist for inserting LayerNorm relative to a sublayer (MHA or FFN):[[1,2,3,4],[2,4,6,8]] is computed per feature column — same as BatchNorm.; The FFN processes the full sequence at once with shared position-sensitive weights — similar to a recurrent layer connecting adjacent positions.Concept
The complete pre-norm encoder block chains two sublayers, each guarded by a LayerNorm and a residual connection:
x = x + MHA(LayerNorm(x)) — multi-head self-attention (Lesson 81)x = x + FFN(LayerNorm(x)) — position-wise feed-forward networkx has the same shape as input — it goes into the next encoder blockThe two LayerNorms (norm1, norm2) are distinct: norm1 guards the attention sublayer, norm2 guards the FFN sublayer. Each has its own learned γ and β.
Analogy
Discussion prompt
Explain Assembling the full EncoderBlock by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
The complete pre-norm encoder block chains two sublayers, each guarded by a LayerNorm and a residual connection:
Estimation
Predict first
Implement the full pre-norm EncoderBlock. Use d_model=8, num_heads=2, d_ff=32. Verify shape and parameter count.
Commit before you compute: what does EncoderBlock as nn.Module come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: output shape (2, 5, 8) = input shape — the encoder block is shape-preserving
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Every sublayer returns the same (B, T, d_model) tensor.
Worked example
Implement the full pre-norm EncoderBlock. Use d_model=8, num_heads=2, d_ff=32. Verify shape and parameter count.
import torch, torch.nn as nn, torch.nn.functional as F
torch.manual_seed(42)
class MultiHeadSelfAttn(nn.Module):
def __init__(self, d, h):
super().__init__()
self.h = h; self.dk = d // h
self.Wq = nn.Linear(d, d, bias=False)
self.Wk = nn.Linear(d, d, bias=False)
self.Wv = nn.Linear(d, d, bias=False)
self.Wo = nn.Linear(d, d, bias=False)
def forward(self, x):
B, T, D = x.shape
def split(W): return W(x).view(B,T,self.h,self.dk).transpose(1,2)
Q, K, V = split(self.Wq), split(self.Wk), split(self.Wv)
attn = (Q @ K.transpose(-2,-1) / self.dk**0.5).softmax(-1)
return self.Wo((attn @ V).transpose(1,2).reshape(B,T,D))
class EncoderBlock(nn.Module):
def __init__(self, d, h, d_ff):
super().__init__()
self.norm1 = nn.LayerNorm(d)
self.attn = MultiHeadSelfAttn(d, h)
self.norm2 = nn.LayerNorm(d)
self.ffn = nn.Sequential(
nn.Linear(d, d_ff), nn.GELU(), nn.Linear(d_ff, d))
def forward(self, x):
x = x + self.attn(self.norm1(x)) # sublayer 1
x = x + self.ffn(self.norm2(x)) # sublayer 2
return x
block = EncoderBlock(d=8, h=2, d_ff=32)
x = torch.randn(2, 5, 8)
out = block(x)
params = sum(p.numel() for p in block.parameters())
print(f'input: {x.shape}') # (2, 5, 8)
print(f'output: {out.shape}') # (2, 5, 8)
print(f'params: {params}') # 840output shape (2, 5, 8) = input shape — the encoder block is shape-preserving
Why: Every sublayer returns the same (B, T, d_model) tensor. The residual x + ... requires shapes to match. 840 parameters = 4·8² (MHA no-bias) + 8·32+32 + 32·8+8 (FFN) + 2·2·8 (two LN).
| component | params (d=8) | formula |
|---|---|---|
| MHA (Wq,Wk,Wv,Wo, no bias) | 256 | 4 × d² = 4 × 64 |
| FFN W1 + bias | 288 | d × d_ff + d_ff = 8×32+32 |
| FFN W2 + bias | 264 | d_ff × d + d = 32×8+8 |
| LayerNorm ×2 (γ + β) | 32 | 2 × 2d = 4×8 |
| Total | 840 | — |
Comparison
Comparison matrix
From EncoderBlock as nn.Module: refill the params (d=8) column from what you know. The rest of the table is as it appeared.
| component | params (d=8) | formula |
|---|---|---|
| MHA (Wq,Wk,Wv,Wo, no bias) | 256 | 4 × d² = 4 × 64 |
| FFN W1 + bias | 288 | d × d_ff + d_ff = 8×32+32 |
| FFN W2 + bias | 264 | d_ff × d + d = 32×8+8 |
| LayerNorm ×2 (γ + β) | 32 | 2 × 2d = 4×8 |
| Total | 840 | — |
Constraint
Discussion prompt
Run The Encoder Block recipe with this step confiscated:
Residual connection: x = x + Sublayer(...) — the +x term keeps gradient magnitude healthy through depth (the 1 + in the chain rule)
Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.
Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.
Answer:
mean = x.mean(-1); var = x.var(-1); xhat = (x-mean)/(var+ε).sqrt() — normalize each token's feature vector independentlyx + Sub(Norm(x))); post-norm wraps the residual sum (Norm(x + Sub(x))). Use pre-norm for stable…x = x + Sublayer(...) — the +x term keeps gradient magnitude healthy through depth (the 1 + in the chain rule)Linear(d → 4d) → GELU → Linear(4d → d); applied position-wise (same weights every token, no cross-token mixing)norm1 → MHA → residual then norm2 → FFN → residual; output shape = input shape (B, T, d_model)Pattern
mean = x.mean(-1); var = x.var(-1); xhat = (x-mean)/(var+ε).sqrt() — normalize each token's feature vector independentlyx + Sub(Norm(x))); post-norm wraps the residual sum (Norm(x + Sub(x))). Use pre-norm for stable deep trainingx = x + Sublayer(...) — the +x term keeps gradient magnitude healthy through depth (the 1 + in the chain rule)Linear(d → 4d) → GELU → Linear(4d → d); applied position-wise (same weights every token, no cross-token mixing)norm1 → MHA → residual then norm2 → FFN → residual; output shape = input shape (B, T, d_model)Edge cases
Discussion prompt
The Encoder Block recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
mean = x.mean(-1); var = x.var(-1); xhat = (x-mean)/(var+ε).sqrt() — normalize each token's feature vector independentlyx + Sub(Norm(x))); post-norm wraps the residual sum (Norm(x + Sub(x))). Use pre-norm for stable…x = x + Sublayer(...) — the +x term keeps gradient magnitude healthy through depth (the 1 + in the chain rule)Linear(d → 4d) → GELU → Linear(4d → d); applied position-wise (same weights every token, no cross-token mixing)norm1 → MHA → residual then norm2 → FFN → residual; output shape = input shape (B, T, d_model)Elimination
Eliminate the wrong options
A token embedding x = [2.0, 4.0, 6.0, 8.0] is passed through LayerNorm (no γ/β, eps=0). What is xhat[0]?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: mean = (2+4+6+8)/4 = 5.0; var = [(2-5)²+(4-5)²+(6-5)²+(8-5)²]/4 = (9+1+1+9)/4 = 5.0; std = √5 ≈ 2.2361. xhat[0] = (2.0−5.0)/2.2361 = −3.0/2.2361 ≈ −1.3416. Verified in the trace table.
Check
Compute before clicking.
Check your understanding
A token embedding x = [2.0, 4.0, 6.0, 8.0] is passed through LayerNorm (no γ/β, eps=0). What is xhat[0]?
Answer: A
Why: mean = (2+4+6+8)/4 = 5.0; var = [(2-5)²+(4-5)²+(6-5)²+(8-5)²]/4 = (9+1+1+9)/4 = 5.0; std = √5 ≈ 2.2361. xhat[0] = (2.0−5.0)/2.2361 = −3.0/2.2361 ≈ −1.3416. Verified in the trace table.
Prediction
Predict first
GELU(−0.5) = −0.1543 while ReLU(−0.5) = 0.0. Why does this matter for the FFN sublayer?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: GELU provides a smooth, non-zero gradient for negative inputs, avoiding dead neurons
Why: ReLU hard-zeros all negative inputs, so any neuron whose pre-activation is always negative gets zero gradient and never updates (dead neuron). GELU smoothly suppresses near-zero negatives while keeping a small gradient, so neurons can recover. This is why BERT, GPT-2, and GPT-3 prefer GELU.
Check
What distinguishes them?
Check your understanding
GELU(−0.5) = −0.1543 while ReLU(−0.5) = 0.0. Why does this matter for the FFN sublayer?
Answer: A
Why: ReLU hard-zeros all negative inputs, so any neuron whose pre-activation is always negative gets zero gradient and never updates (dead neuron). GELU smoothly suppresses near-zero negatives while keeping a small gradient, so neurons can recover. This is why BERT, GPT-2, and GPT-3 prefer GELU.
Elimination
Eliminate the wrong options
Which is the correct pre-norm encoder sublayer 1 (the MHA sublayer)?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Pre-norm: normalize the input before the sublayer, then add the original (un-normalized) x via the residual. This keeps a clean residual path and is what GPT-2/LLaMA use. The formula is x = x + Sub(Norm(x)).
Check
Identify the correct data flow.
Check your understanding
Which is the correct pre-norm encoder sublayer 1 (the MHA sublayer)?
Answer: A
Why: Pre-norm: normalize the input before the sublayer, then add the original (un-normalized) x via the residual. This keeps a clean residual path and is what GPT-2/LLaMA use. The formula is x = x + Sub(Norm(x)).
+ x term, gradients vanish in deep stacks and the network cannot learn an identity mapping as a base case.Section
Project
Concept
Build all three components from scratch: LayerNorm (verify vs nn.LayerNorm), FFN sublayer (Linear+GELU+Linear), full EncoderBlock nn.Module with pre-norm residual wiring.
| # | milestone | key check |
|---|---|---|
| 1 | LayerNorm from scratch on x=[1,3,5,7] | match nn.LayerNorm to atol=1e-4 |
| 2 | FFN(d_model=4, d_ff=16) — shape + param count | output (1,4), params=148 |
| 3 | Full EncoderBlock(d=8, h=2, d_ff=32) | output (B,T,8), params=840 |
Build rules: type every line, implement LayerNorm manually before reaching for nn.LayerNorm, and verify shapes after each sublayer with print(x.shape).
Counterexample
Discussion prompt
Build all three components from scratch: LayerNorm (verify vs nn.LayerNorm), FFN sublayer (Linear+GELU+Linear), full EncoderBlock nn.Module with pre-norm residual wiring.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Build rules: type every line, implement LayerNorm manually before reaching for nn.LayerNorm, and verify shapes after each sublayer with print(x.shape).
Worked example
Your turn: implement LayerNorm manually on x = [1, 3, 5, 7]. Predict xhat before running.
Hint: mean = x.mean(dim=-1, keepdim=True); var = x.var(dim=-1, keepdim=True, unbiased=False); xhat = (x - mean) / (var + 1e-5).sqrt().
import torch, torch.nn as nn
x = torch.tensor([[1.0, 3.0, 5.0, 7.0]])
mean = x.mean(dim=-1, keepdim=True)
var = x.var(dim=-1, keepdim=True, unbiased=False)
xhat = (x - mean) / (var + 1e-5).sqrt()
ln = nn.LayerNorm(4, elementwise_affine=False, eps=1e-5)
print('scratch:', [round(v,4) for v in xhat.squeeze().tolist()])
print('nn.LN: ', [round(v,4) for v in ln(x).squeeze().tolist()])
print('match: ', torch.allclose(xhat, ln(x), atol=1e-4))| quantity | value |
|---|---|
| mean | 4.0 |
| var | 5.0000 |
| std | 2.2361 |
| xhat | [−1.3416, −0.4472, 0.4472, 1.3416] |
| match | True |
Worked example
Your turn: implement FFN(d_model=4, d_ff=16), pass x = [[1, 0, −1, 0.5]], and count the parameters. Predict the param count before running.
Hint: nn.Linear(d_model, d_ff) → F.gelu(...) → nn.Linear(d_ff, d_model). Params = d_model*d_ff + d_ff + d_ff*d_model + d_model.
import torch, torch.nn as nn, torch.nn.functional as F
torch.manual_seed(0)
class FFN(nn.Module):
def __init__(self, d_model, d_ff):
super().__init__()
self.w1 = nn.Linear(d_model, d_ff)
self.w2 = nn.Linear(d_ff, d_model)
def forward(self, x): return self.w2(F.gelu(self.w1(x)))
ffn = FFN(4, 16)
x = torch.tensor([[1.0, 0.0, -1.0, 0.5]])
out = ffn(x)
params = sum(p.numel() for p in ffn.parameters())
print('output:', [round(v,4) for v in out.squeeze().tolist()])
print('params:', params)| output | value |
|---|---|
| FFN(x)[0] | −0.3974 |
| FFN(x)[1] | 0.2011 |
| FFN(x)[2] | 0.0934 |
| FFN(x)[3] | −0.1582 |
| params | 148 |
Trade off
Comparison matrix
From Milestone 2 — FFN sublayer: every row here is a choice with a cost. Fill the value column, then say which row you would actually pick and what you give up for it.
| output | value |
|---|---|
| FFN(x)[0] | −0.3974 |
| FFN(x)[1] | 0.2011 |
| FFN(x)[2] | 0.0934 |
| FFN(x)[3] | −0.1582 |
| params | 148 |
Worked example
Your turn: assemble EncoderBlock(d=8, h=2, d_ff=32) with pre-norm residuals. Verify that output.shape == input.shape and params == 840.
Hint: forward: x = x + attn(norm1(x)); x = x + ffn(norm2(x)). Use nn.Sequential(nn.Linear(d,d_ff), nn.GELU(), nn.Linear(d_ff,d)) for the FFN.
import torch, torch.nn as nn
torch.manual_seed(42)
# (reuse MultiHeadSelfAttn from Milestone 2 shell)
class EncoderBlock(nn.Module):
def __init__(self, d, h, d_ff):
super().__init__()
self.norm1 = nn.LayerNorm(d)
self.attn = MultiHeadSelfAttn(d, h) # from prev. milestone
self.norm2 = nn.LayerNorm(d)
self.ffn = nn.Sequential(
nn.Linear(d, d_ff), nn.GELU(), nn.Linear(d_ff, d))
def forward(self, x):
x = x + self.attn(self.norm1(x))
x = x + self.ffn(self.norm2(x))
return x
block = EncoderBlock(d=8, h=2, d_ff=32)
x = torch.randn(2, 5, 8)
out = block(x)
params = sum(p.numel() for p in block.parameters())
print(f'input: {x.shape}') # torch.Size([2, 5, 8])
print(f'output: {out.shape}') # torch.Size([2, 5, 8])
print(f'params: {params}') # 840| check | expected | note |
|---|---|---|
| output.shape | (2, 5, 8) | same as input — shape-preserving |
| params | 840 | MHA 256 + FFN 552 + LN 32 |
| sublayer 1 | x + attn(norm1(x)) | pre-norm, residual |
| sublayer 2 | x + ffn(norm2(x)) | pre-norm, residual |
Discrimination
Sort into buckets
Sort these by note, from memory, without looking back at Milestone 3 — full EncoderBlock. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Concept
import torch, torch.nn as nn, torch.nn.functional as F
torch.manual_seed(42)
# LayerNorm from scratch
def layernorm_scratch(x, eps=1e-5):
mu = x.mean(-1, keepdim=True)
var = x.var(-1, keepdim=True, unbiased=False)
return (x - mu) / (var + eps).sqrt()
# Multi-head self-attention
class MHA(nn.Module):
def __init__(self, d, h):
super().__init__()
self.h = h; self.dk = d // h
self.Wq = nn.Linear(d,d,bias=False); self.Wk = nn.Linear(d,d,bias=False)
self.Wv = nn.Linear(d,d,bias=False); self.Wo = nn.Linear(d,d,bias=False)
def forward(self, x):
B,T,D = x.shape
def sp(W): return W(x).view(B,T,self.h,self.dk).transpose(1,2)
Q,K,V = sp(self.Wq), sp(self.Wk), sp(self.Wv)
return self.Wo(((Q@K.transpose(-2,-1)/self.dk**0.5).softmax(-1)@V)
.transpose(1,2).reshape(B,T,D))
# EncoderBlock (pre-norm)
class EncoderBlock(nn.Module):
def __init__(self, d, h, d_ff):
super().__init__()
self.norm1=nn.LayerNorm(d); self.attn=MHA(d,h)
self.norm2=nn.LayerNorm(d)
self.ffn=nn.Sequential(nn.Linear(d,d_ff),nn.GELU(),nn.Linear(d_ff,d))
def forward(self, x):
x = x + self.attn(self.norm1(x))
x = x + self.ffn(self.norm2(x))
return x
# verify
xv = torch.tensor([[1.0,3.0,5.0,7.0]])
print('LN match:', torch.allclose(layernorm_scratch(xv), nn.LayerNorm(4,elementwise_affine=False)(xv),atol=1e-4))
enc = EncoderBlock(8,2,32)
xe = torch.randn(2,5,8)
print('shape:', enc(xe).shape) # torch.Size([2, 5, 8])
print('params:', sum(p.numel() for p in enc.parameters())) # 840| output | value | what it confirms |
|---|---|---|
| LN match | True | scratch matches nn.LayerNorm |
| shape | (2, 5, 8) | shape-preserving block |
| params | 840 | matches hand-computed budget |
Stack N of these EncoderBlock modules — that's the BERT/GPT encoder. Every block: normalize → attend → normalize → feed-forward, with clean residual paths throughout.
Comparison
Comparison matrix
From The full program: refill the what it confirms column from what you know. The rest of the table is as it appeared.
| output | value | what it confirms |
|---|---|---|
| LN match | True | scratch matches nn.LayerNorm |
| shape | (2, 5, 8) | shape-preserving block |
| params | 840 | matches hand-computed budget |
Concept
Out loud, slides closed: explain (1) why LayerNorm normalizes over the feature dimension rather than the batch dimension; (2) the pre-norm formula and why it stabilizes deep training; (3) what the FFN's 4× expansion achieves and why GELU is preferred over ReLU.
Stretch (homework from L82): implement LayerNorm as a proper nn.Module with trainable γ and β; run the pre-norm vs post-norm training experiment on a 6-layer model and log the loss curves; experiment with d_ff = 2·d_model vs 4·d_model and observe quality difference. Next: Lesson 83 — positional encodings and full encoder stack.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — LayerNorm — normalize over features · Pre-norm vs Post-norm & residual connections · FFN sublayer — GELU and the 4× rule · Your turn: build EncoderBlock. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
mean/var over last dim) and verify against nn.LayerNormx = x + Sub(Norm(x))) and why it stabilizes deep transformersLinear → GELU → Linear, d_ff = 4d), position-wisenn.Module with two pre-norm sublayers and residual connections| concept | the one thing to remember |
|---|---|
| LayerNorm | normalize each token's feature vector; dim=-1 |
| LN vs BN | LN: per-token (feature axis); BN: per-feature (batch axis) |
| pre-norm | Sub(Norm(x)) then add x; residual path is clean |
| FFN | two Linear layers + GELU; 4× expansion; position-wise |
| EncoderBlock | norm→MHA→residual, then norm→FFN→residual; shape-preserving |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.