Lesson 82: LayerNorm, FFN with GELU, and the Full Encoder Block

USAAIO Lesson 82, from Phase 3 on transformers. It builds LayerNorm from scratch, normalizing over the feature dimension and verifying against nn.LayerNorm, then compares pre-norm with post-norm for stability. It covers the feed-forward sublayer - two Linear layers with a GELU between them, at d_ff = 4·d_model - and residual connections for gradient flow, then assembles the complete pre-norm EncoderBlock as an nn.Module. All the numbers were verified with torch 2.7.1+cpu in June 2026. The lesson runs to 30 slides.

Subject: Machine Learning · 60 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. LayerNorm, FFN with GELU & the Full Encoder Block

Title

USAAIO · Lesson 82 · Phase 3 (Transformers)

The two sublayers that wrap multi-head attention: LayerNorm (normalize over features, not batch), GELU feed-forward network, residual connections, and the complete pre-norm EncoderBlock as a reusable nn.Module.

2. By the end of this lesson you can

Objectives

  1. Implement LayerNorm from scratch and verify it against nn.LayerNorm
  2. Explain why layer norm (over features) beats batch norm for variable-length sequences
  3. Contrast pre-norm (Norm → sublayer) and post-norm and state which is more stable
  4. Build the FFN sublayer (two Linear layers + GELU, d_ff = 4·d_model)
  5. Assemble the full EncoderBlock nn.Module with residual connections and both sublayers

3. What survived from Multi-Head Attention?

Warm-up

Discussion prompt

Before we open Lesson 82: LayerNorm, FFN with GELU, and the Full Encoder Block: without looking back, what was the main idea of Multi-Head Attention, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

multi-head attention from scratch — projecting Q/K/V into h parallel subspaces, scaled dot-product per head, concat + W_O projection, parameter count proof (4·d_model² regardless of h), and h=1 equivalence to single-head attention. Implement MultiHeadAttention as nn.Module and verify every shape.

4. LayerNorm — normalize over features

Section

Part 1 of 3

5. What LayerNorm does

Concept

LayerNorm normalizes each token's feature vector independently: compute the mean and variance across the d_model features for that one token, then rescale.

\[ \hat{x}_i = \frac{x_i - \mu}{\sqrt{\sigma^2 + \varepsilon}}, \quad \mu = \frac{1}{d}\sum_{i=1}^d x_i, \quad \sigma^2 = \frac{1}{d}\sum_{i=1}^d (x_i - \mu)^2 \]

After normalization, learnable parameters γ (scale) and β (shift) — both of shape d_model — let the network undo the normalization if needed. nn.LayerNorm trains γ and β automatically.

6. Break it if you can: What LayerNorm does

Counterexample

Discussion prompt

LayerNorm normalizes each token's feature vector independently: compute the mean and variance across the d_model features for that one token, then rescale.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

After normalization, learnable parameters γ (scale) and β (shift) — both of shape d_model — let the network undo the normalization if needed. nn.LayerNorm trains γ and β automatically.

7. LayerNorm vs BatchNorm: the axis difference

Concept

BatchNorm normalizes across the batch (per feature channel). LayerNorm normalizes across features (per token). For sequences, batch sizes vary and sequence lengths vary — BatchNorm statistics become unreliable.

normmean/var computed oversuitable for
BatchNormbatch dimension (N samples)fixed-size CNN feature maps
LayerNormfeature dimension (d_model)variable-length sequences, NLP

For a token embedding x ∈ ℝ^d, LN's statistics are entirely self-contained — independent of what other tokens or other sequences are in the batch (Lesson 17 CE context).

8. Fill in: mean/var computed over for LayerNorm vs BatchNorm: the axis difference

Comparison

Comparison matrix

From LayerNorm vs BatchNorm: the axis difference: refill the mean/var computed over column from what you know. The rest of the table is as it appeared.

normmean/var computed oversuitable for
BatchNormbatch dimension (N samples)fixed-size CNN feature maps
LayerNormfeature dimension (d_model)variable-length sequences, NLP

9. Guess the shape of the answer: LayerNorm from scratch

Estimation

Predict first

Normalize x = [1, 3, 5, 7] manually, then verify against nn.LayerNorm.

Commit before you compute: what does LayerNorm from scratch come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: mean = 4.0; var = 5.0; std = 2.2361

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. mean = (1+3+5+7)/4; var = [(1-4)² + (3-4)² + (5-4)² + (7-4)²] / 4 = (9+1+1+9)/4 = 5.0; std = √5 ≈ 2.2361

10. LayerNorm from scratch

Worked example

Normalize x = [1, 3, 5, 7] manually, then verify against nn.LayerNorm.

import torch, torch.nn as nn

x = torch.tensor([[1.0, 3.0, 5.0, 7.0]])
mean = x.mean(dim=-1, keepdim=True)          # 4.0
var  = x.var(dim=-1, keepdim=True, unbiased=False)  # 5.0
xhat = (x - mean) / (var + 1e-5).sqrt()     # no gamma/beta yet

ln = nn.LayerNorm(4, elementwise_affine=False, eps=1e-5)
print('scratch:', xhat.tolist())
print('nn.LN:  ', ln(x).tolist())
print('match:  ', torch.allclose(xhat, ln(x), atol=1e-4))

mean = 4.0; var = 5.0; std = 2.2361

Why: mean = (1+3+5+7)/4; var = [(1-4)² + (3-4)² + (5-4)² + (7-4)²] / 4 = (9+1+1+9)/4 = 5.0; std = √5 ≈ 2.2361

x_ix_i − μ(x_i − μ) / stdxhat_i
1.0−3.0−3.0 / 2.2361−1.3416
3.0−1.0−1.0 / 2.2361−0.4472
5.0 1.0 1.0 / 2.2361 0.4472
7.0 3.0 3.0 / 2.2361 1.3416

11. What each one costs: LayerNorm from scratch

Trade off

Comparison matrix

From LayerNorm from scratch: every row here is a choice with a cost. Fill the xhat_i column, then say which row you would actually pick and what you give up for it.

x_ix_i − μ(x_i − μ) / stdxhat_i
1.0−3.0−3.0 / 2.2361−1.3416
3.0−1.0−1.0 / 2.2361−0.4472
5.01.01.0 / 2.23610.4472
7.03.03.0 / 2.23611.3416

12. Something is wrong here: normalizing over the wrong dimension

Anomaly

Predict first

A student writes this, and it looks reasonable:

LayerNorm over a batch of two sequences [[1,2,3,4],[2,4,6,8]] is computed per feature column — same as BatchNorm.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: This is BatchNorm's statistics — averaging down the batch axis per feature.

LayerNorm normalizes each row (each token vector) independently: dim=-1, not dim=0.

Why: This is BatchNorm's statistics — averaging down the batch axis per feature. LayerNorm never touches the batch axis.

13. Trap: normalizing over the wrong dimension

Trap

The trap

LayerNorm over a batch of two sequences [[1,2,3,4],[2,4,6,8]] is computed per feature column — same as BatchNorm.

mean_col = [(1+2)/2, (2+4)/2, (3+6)/2, (4+8)/2] = [1.5, 3.0, 4.5, 6.0]

Why: This is BatchNorm's statistics — averaging down the batch axis per feature. LayerNorm never touches the batch axis.

The fix

LayerNorm normalizes each row (each token vector) independently: dim=-1, not dim=0.

Row 0: mean = 2.5, std = 1.118 → xhat = [−1.34, −0.45, 0.45, 1.34]; Row 1: same ratios, identical xhat

Why: Each token is normalized by its own statistics. Both rows produce the same xhat because they have the same shape — [1,2,3,4] scaled by 2. BatchNorm would give different results because it mixes the two samples.

14. Break it on purpose: normalizing over the wrong dimension

Break the constraint

Discussion prompt

The rule this trap just fixed:

LayerNorm normalizes each row (each token vector) independently: dim=-1, not dim=0.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

This is BatchNorm's statistics — averaging down the batch axis per feature. LayerNorm never touches the batch axis.

15. Pre-norm vs Post-norm & residual connections

Section

Part 2 of 3

16. Pre-norm: the modern default

Concept

Two orderings exist for inserting LayerNorm relative to a sublayer (MHA or FFN):

variantformulatraining stability
Post-norm (original)x = LayerNorm(x + Sublayer(x))harder to train deep; norms after residual
Pre-norm (modern)x = x + Sublayer(LayerNorm(x))more stable; norm before sublayer, residual path clean

Pre-norm keeps the residual stream unnormalized, so gradients flow back through the + x path without touching the norm. This is why GPT-2, GPT-3, LLaMA, and most modern transformers use pre-norm.

17. By analogy: Pre-norm: the modern default

Analogy

Discussion prompt

Explain Pre-norm: the modern default by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Two orderings exist for inserting LayerNorm relative to a sublayer (MHA or FFN):

18. Residual connections: the gradient highway

Concept

The residual formula output = x + Sublayer(Norm(x)) means the network learns a correction (the residual) on top of the identity path.

\[ \frac{\partial \mathcal{L}}{\partial x} = \frac{\partial \mathcal{L}}{\partial(x + F(x))} \cdot \left(1 + \frac{\partial F}{\partial x}\right) \]

The 1 in the gradient ensures that, even if ∂F/∂x ≈ 0 in early training, a full gradient signal still flows to every earlier layer (Lesson 9 backprop context). Without residuals, deep networks suffer vanishing gradients.

19. Teach it back: Residual connections: the gradient highway

Explain it

Discussion prompt

Explain Residual connections: the gradient highway to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

The residual formula output = x + Sublayer(Norm(x)) means the network learns a correction (the residual) on top of the identity path.

20. What has to be given first: Residual gradient flow: with vs without

Missing information

Discussion prompt

Train an 8-layer network with and without residual connections. Compare the gradient norm reaching layer 1 (deepest backprop path).

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

Without residuals the gradient degrades through tanh saturation. With residuals the + identity path keeps the gradient magnitude healthy in all layers — the foundation for training deep transformers.

21. Residual gradient flow: with vs without

Worked example

Train an 8-layer network with and without residual connections. Compare the gradient norm reaching layer 1 (deepest backprop path).

import torch, torch.nn as nn
torch.manual_seed(7)
depth, d = 8, 16

class NoRes(nn.Module):
    def __init__(self):
        super().__init__()
        self.layers = nn.ModuleList([nn.Linear(d, d) for _ in range(depth)])
        for l in self.layers: nn.init.normal_(l.weight, std=0.3)
    def forward(self, x):
        for l in self.layers: x = torch.tanh(l(x))
        return x

class WithRes(nn.Module):
    def __init__(self):
        super().__init__()
        self.layers = nn.ModuleList([nn.Linear(d, d) for _ in range(depth)])
    def forward(self, x):
        for l in self.layers: x = x + torch.tanh(l(x))
        return x

x = torch.randn(4, d)
for Cls, name in [(NoRes,'no-res'), (WithRes,'residual')]:
    model = Cls(); model(x).sum().backward()
    gnorms = [l.weight.grad.norm().item() for l in model.layers]
    print(f'{name}: L1={gnorms[0]:.3f}, L8={gnorms[-1]:.3f}')

no-res: L1=15.07, L8=9.79; residual: L1=29.27, L8=31.16

Why: Without residuals the gradient degrades through tanh saturation. With residuals the + identity path keeps the gradient magnitude healthy in all layers — the foundation for training deep transformers.

layerno-residual ‖∇‖with-residual ‖∇‖
L1 (earliest)15.068729.2718
L213.110922.8678
L411.987823.5143
L613.613323.4945
L8 (latest) 9.786931.1626

22. Watch it run: Residual gradient flow: with vs without

Pattern

Step through it

Step through Residual gradient flow: with vs without one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: layer is L1 (earliest)
  2. Step 2: layer is L2
  3. Step 3: layer is L4
  4. Step 4: layer is L6
  5. Step 5: layer is L8 (latest)

23. FFN sublayer — GELU and the 4× rule

Section

Part 3 of 3

24. The FFN sublayer

Concept

Each encoder position independently passes through a two-layer MLP: expand to d_ff = 4·d_model, apply GELU, project back. The expansion and contraction happen at every token in parallel.

\[ \text{FFN}(x) = W_2\, \text{GELU}(W_1 x + b_1) + b_2, \quad W_1 \in \mathbb{R}^{d_{ff} \times d}, \; W_2 \in \mathbb{R}^{d \times d_{ff}} \]

Why 4×? Empirical finding from the original transformer paper. The extra width gives the network capacity to store factual associations (Lesson 82 onwards). Why GELU over ReLU? Smoother gradient near zero; small negative values not completely zeroed out.

25. What rests on this: The FFN sublayer

Socratic

Discussion prompt

Each encoder position independently passes through a two-layer MLP: expand to d_ff = 4·d_model, apply GELU, project back. The expansion and contraction happen at every token in parallel.

Suppose that were not true. What is the first thing in Lesson 82: LayerNorm, FFN with GELU, and the Full Encoder Block that would stop working?

Hint: Follow it one step downstream. The answer is whatever was quietly relying on it.

26. GELU vs ReLU: the smooth gate

Concept

GELU (Gaussian Error Linear Unit) weights the input by the probability it is positive: GELU(x) ≈ x · Φ(x) where Φ is the standard normal CDF. It smoothly suppresses small or negative inputs.

xGELU(x)ReLU(x)difference
−2.0−0.04550.0000GELU leaks small negative
−1.0−0.15870.0000GELU not exactly zero
−0.5−0.15430.0000smooth suppression
0.0 0.00000.0000both zero at origin
0.5 0.34570.5000GELU slightly below ReLU
1.0 0.84131.0000approaches identity
2.0 1.95452.0000≈ identity for large x

The smooth gradient at x < 0 avoids the dead neuron problem of ReLU (Lesson 44 activation context). BERT, GPT-2, and GPT-3 all use GELU.

27. Teach it back: GELU vs ReLU: the smooth gate

Explain it

Discussion prompt

Explain GELU vs ReLU: the smooth gate to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

GELU (Gaussian Error Linear Unit) weights the input by the probability it is positive: GELU(x) ≈ x · Φ(x) where Φ is the standard normal CDF. It smoothly suppresses small or negative inputs.

28. Guess the shape of the answer: FFN sublayer in PyTorch

Estimation

Predict first

Build the FFN as an nn.Module. Use d_model=4, d_ff=16 (4×). Verify parameter count and output shape.

Commit before you compute: what does FFN sublayer in PyTorch come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: output shape (1, 4) — same as input; params = 4·16+16 + 16·4+4 = 80+68 = 148

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. W1: d_model×d_ff + d_ff = 4·16+16 = 80.

29. FFN sublayer in PyTorch

Worked example

Build the FFN as an nn.Module. Use d_model=4, d_ff=16 (4×). Verify parameter count and output shape.

import torch, torch.nn as nn, torch.nn.functional as F
torch.manual_seed(0)

class FFN(nn.Module):
    def __init__(self, d_model, d_ff):
        super().__init__()
        self.w1 = nn.Linear(d_model, d_ff)
        self.w2 = nn.Linear(d_ff, d_model)
    def forward(self, x):
        return self.w2(F.gelu(self.w1(x)))

ffn = FFN(d_model=4, d_ff=16)
x   = torch.tensor([[1.0, 0.0, -1.0, 0.5]])
hid = F.gelu(ffn.w1(x))
out = ffn(x)
params = sum(p.numel() for p in ffn.parameters())
print(f'hidden shape: {hid.shape}')   # (1, 16)
print(f'output shape: {out.shape}')   # (1,  4)
print(f'output:       {out.squeeze().tolist()}')  # [-0.3974, 0.2011, 0.0934, -0.1582]
print(f'params:       {params}')      # 148

output shape (1, 4) — same as input; params = 4·16+16 + 16·4+4 = 80+68 = 148

Why: W1: d_model×d_ff + d_ff = 4·16+16 = 80. W2: d_ff×d_model + d_model = 16·4+4 = 68. Total = 148. The FFN expands then contracts, restoring d_model for the residual add.

steptensor shapenote
input x(1, 4)one token, d_model=4
W1(x)(1, 16)expand to d_ff=4×d_model
GELU(W1(x))(1, 16)smooth gating in-place
W2(GELU(W1(x)))(1, 4)contract back to d_model
FFN params—80 + 68 = 148 total

30. Which is which, by tensor shape

Discrimination

Sort into buckets

Sort these by tensor shape, from memory, without looking back at FFN sublayer in PyTorch. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

(1, 4)
input x; W2(GELU(W1(x)))
(1, 16)
W1(x); GELU(W1(x))
—
FFN params
g1
tensor shape is "(1, 4)" for input x, W2(GELU(W1(x))) — that is what the table on "FFN sublayer in PyTorch" records, and it is the single property separating this group from the rest.
g2
tensor shape is "(1, 16)" for W1(x), GELU(W1(x)) — that is what the table on "FFN sublayer in PyTorch" records, and it is the single property separating this group from the rest.
g3
tensor shape is "—" for FFN params — that is what the table on "FFN sublayer in PyTorch" records, and it is the single property separating this group from the rest.

31. Something is wrong here: applying the FFN across the sequence instead of…

Anomaly

Predict first

A student writes this, and it looks reasonable:

The FFN processes the full sequence at once with shared position-sensitive weights — similar to a recurrent layer connecting adjacent positions.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: This is wrong — the FFN weight matrix W1 is (d_ff, d_model), applied identically and independently at every position.

The FFN is position-wise: the same W1 and W2 are applied independently at every token position.

Why: This is wrong — the FFN weight matrix W1 is (d_ff, d_model), applied identically and independently at every position. There is no cross-position communication in the FFN.

32. Trap: applying the FFN across the sequence instead of per-token

Trap

The trap

The FFN processes the full sequence at once with shared position-sensitive weights — similar to a recurrent layer connecting adjacent positions.

Feed W1 a matrix of shape (T, d_model) and assume different positions share information through the weight

Why: This is wrong — the FFN weight matrix W1 is (d_ff, d_model), applied identically and independently at every position. There is no cross-position communication in the FFN.

The fix

The FFN is position-wise: the same W1 and W2 are applied independently at every token position.

Input: (B, T, d_model) → W1 applied as (B·T, d_model) @ W1.T → (B, T, d_ff) → GELU → W2 → (B, T, d_model)

Why: nn.Linear broadcasts over leading dimensions, so passing (B, T, d_model) is valid — it processes each of the B·T tokens independently through the same weights. Cross-token interactions happen only in MHA, not the FFN.

33. Which of these survive contact with Lesson 82: LayerNorm, FFN with GELU, and the…?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
LayerNorm normalizes each token's feature vector independently: compute the mean and variance across the d_model features for that one token, then rescale.; For a token embedding x ∈ ℝ^d, LN's statistics are entirely self-contained — independent of what other tokens or other sequences are in the batch (Lesson 17 CE context).; Two orderings exist for inserting LayerNorm relative to a sublayer (MHA or FFN):
Breaks
LayerNorm over a batch of two sequences [[1,2,3,4],[2,4,6,8]] is computed per feature column — same as BatchNorm.; The FFN processes the full sequence at once with shared position-sensitive weights — similar to a recurrent layer connecting adjacent positions.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 82: LayerNorm, FFN with GELU, and the Full Encoder Block puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

34. Assembling the full EncoderBlock

Concept

The complete pre-norm encoder block chains two sublayers, each guarded by a LayerNorm and a residual connection:

  1. Sublayer 1: x = x + MHA(LayerNorm(x)) — multi-head self-attention (Lesson 81)
  2. Sublayer 2: x = x + FFN(LayerNorm(x)) — position-wise feed-forward network
  3. Output x has the same shape as input — it goes into the next encoder block

The two LayerNorms (norm1, norm2) are distinct: norm1 guards the attention sublayer, norm2 guards the FFN sublayer. Each has its own learned γ and β.

35. By analogy: Assembling the full EncoderBlock

Analogy

Discussion prompt

Explain Assembling the full EncoderBlock by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

The complete pre-norm encoder block chains two sublayers, each guarded by a LayerNorm and a residual connection:

36. Guess the shape of the answer: EncoderBlock as nn.Module

Estimation

Predict first

Implement the full pre-norm EncoderBlock. Use d_model=8, num_heads=2, d_ff=32. Verify shape and parameter count.

Commit before you compute: what does EncoderBlock as nn.Module come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: output shape (2, 5, 8) = input shape — the encoder block is shape-preserving

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Every sublayer returns the same (B, T, d_model) tensor.

37. EncoderBlock as nn.Module

Worked example

Implement the full pre-norm EncoderBlock. Use d_model=8, num_heads=2, d_ff=32. Verify shape and parameter count.

import torch, torch.nn as nn, torch.nn.functional as F
torch.manual_seed(42)

class MultiHeadSelfAttn(nn.Module):
    def __init__(self, d, h):
        super().__init__()
        self.h = h; self.dk = d // h
        self.Wq = nn.Linear(d, d, bias=False)
        self.Wk = nn.Linear(d, d, bias=False)
        self.Wv = nn.Linear(d, d, bias=False)
        self.Wo = nn.Linear(d, d, bias=False)
    def forward(self, x):
        B, T, D = x.shape
        def split(W): return W(x).view(B,T,self.h,self.dk).transpose(1,2)
        Q, K, V = split(self.Wq), split(self.Wk), split(self.Wv)
        attn = (Q @ K.transpose(-2,-1) / self.dk**0.5).softmax(-1)
        return self.Wo((attn @ V).transpose(1,2).reshape(B,T,D))

class EncoderBlock(nn.Module):
    def __init__(self, d, h, d_ff):
        super().__init__()
        self.norm1 = nn.LayerNorm(d)
        self.attn  = MultiHeadSelfAttn(d, h)
        self.norm2 = nn.LayerNorm(d)
        self.ffn   = nn.Sequential(
            nn.Linear(d, d_ff), nn.GELU(), nn.Linear(d_ff, d))
    def forward(self, x):
        x = x + self.attn(self.norm1(x))  # sublayer 1
        x = x + self.ffn(self.norm2(x))   # sublayer 2
        return x

block = EncoderBlock(d=8, h=2, d_ff=32)
x = torch.randn(2, 5, 8)
out = block(x)
params = sum(p.numel() for p in block.parameters())
print(f'input:  {x.shape}')    # (2, 5, 8)
print(f'output: {out.shape}')  # (2, 5, 8)
print(f'params: {params}')     # 840

output shape (2, 5, 8) = input shape — the encoder block is shape-preserving

Why: Every sublayer returns the same (B, T, d_model) tensor. The residual x + ... requires shapes to match. 840 parameters = 4·8² (MHA no-bias) + 8·32+32 + 32·8+8 (FFN) + 2·2·8 (two LN).

componentparams (d=8)formula
MHA (Wq,Wk,Wv,Wo, no bias)2564 × d² = 4 × 64
FFN W1 + bias288d × d_ff + d_ff = 8×32+32
FFN W2 + bias264d_ff × d + d = 32×8+8
LayerNorm ×2 (γ + β) 322 × 2d = 4×8
Total840—

38. Fill in: params (d=8) for EncoderBlock as nn.Module

Comparison

Comparison matrix

From EncoderBlock as nn.Module: refill the params (d=8) column from what you know. The rest of the table is as it appeared.

componentparams (d=8)formula
MHA (Wq,Wk,Wv,Wo, no bias)2564 × d² = 4 × 64
FFN W1 + bias288d × d_ff + d_ff = 8×32+32
FFN W2 + bias264d_ff × d + d = 32×8+8
LayerNorm ×2 (γ + β)322 × 2d = 4×8
Total840—

39. Without one step: The Encoder Block recipe

Constraint

Discussion prompt

Run The Encoder Block recipe with this step confiscated:

Residual connection: x = x + Sublayer(...) — the +x term keeps gradient magnitude healthy through depth (the 1 + in the chain rule)

Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.

Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.

Answer:

  1. LayerNorm: mean = x.mean(-1); var = x.var(-1); xhat = (x-mean)/(var+ε).sqrt() — normalize each token's feature vector independently
  2. Pre-norm vs post-norm: pre-norm wraps the sublayer (x + Sub(Norm(x))); post-norm wraps the residual sum (Norm(x + Sub(x))). Use pre-norm for stable…
  3. Residual connection: x = x + Sublayer(...) — the +x term keeps gradient magnitude healthy through depth (the 1 + in the chain rule)
  4. FFN sublayer: Linear(d → 4d) → GELU → Linear(4d → d); applied position-wise (same weights every token, no cross-token mixing)
  5. Full EncoderBlock: norm1 → MHA → residual then norm2 → FFN → residual; output shape = input shape (B, T, d_model)

40. The Encoder Block recipe

Pattern

  1. LayerNorm: mean = x.mean(-1); var = x.var(-1); xhat = (x-mean)/(var+ε).sqrt() — normalize each token's feature vector independently
  2. Pre-norm vs post-norm: pre-norm wraps the sublayer (x + Sub(Norm(x))); post-norm wraps the residual sum (Norm(x + Sub(x))). Use pre-norm for stable deep training
  3. Residual connection: x = x + Sublayer(...) — the +x term keeps gradient magnitude healthy through depth (the 1 + in the chain rule)
  4. FFN sublayer: Linear(d → 4d) → GELU → Linear(4d → d); applied position-wise (same weights every token, no cross-token mixing)
  5. Full EncoderBlock: norm1 → MHA → residual then norm2 → FFN → residual; output shape = input shape (B, T, d_model)

41. Where does it stop working: The Encoder Block recipe

Edge cases

Discussion prompt

The Encoder Block recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. LayerNorm: mean = x.mean(-1); var = x.var(-1); xhat = (x-mean)/(var+ε).sqrt() — normalize each token's feature vector independently
  2. Pre-norm vs post-norm: pre-norm wraps the sublayer (x + Sub(Norm(x))); post-norm wraps the residual sum (Norm(x + Sub(x))). Use pre-norm for stable…
  3. Residual connection: x = x + Sublayer(...) — the +x term keeps gradient magnitude healthy through depth (the 1 + in the chain rule)
  4. FFN sublayer: Linear(d → 4d) → GELU → Linear(4d → d); applied position-wise (same weights every token, no cross-token mixing)
  5. Full EncoderBlock: norm1 → MHA → residual then norm2 → FFN → residual; output shape = input shape (B, T, d_model)

42. Rule out three: Check yourself — LayerNorm axis

Elimination

Eliminate the wrong options

A token embedding x = [2.0, 4.0, 6.0, 8.0] is passed through LayerNorm (no γ/β, eps=0). What is xhat[0]?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. −1.3416
  • B. 0.0
  • C. −0.5
  • D. −0.4472

Survives elimination: A

Why: mean = (2+4+6+8)/4 = 5.0; var = [(2-5)²+(4-5)²+(6-5)²+(8-5)²]/4 = (9+1+1+9)/4 = 5.0; std = √5 ≈ 2.2361. xhat[0] = (2.0−5.0)/2.2361 = −3.0/2.2361 ≈ −1.3416. Verified in the trace table.

43. Check yourself — LayerNorm axis

Check

Compute before clicking.

Check your understanding

A token embedding x = [2.0, 4.0, 6.0, 8.0] is passed through LayerNorm (no γ/β, eps=0). What is xhat[0]?

  • A. −1.3416 (correct)
  • B. 0.0
  • C. −0.5
  • D. −0.4472

Answer: A

Why: mean = (2+4+6+8)/4 = 5.0; var = [(2-5)²+(4-5)²+(6-5)²+(8-5)²]/4 = (9+1+1+9)/4 = 5.0; std = √5 ≈ 2.2361. xhat[0] = (2.0−5.0)/2.2361 = −3.0/2.2361 ≈ −1.3416. Verified in the trace table.

Why B tempts people
0.0 would be the normalized value for the mean element (x=5.0), not for x=2.0. The mean is 5.0 here, not 2.0.
Why C tempts people
−0.5 is x_scaled if you simply divided by the range (8−2=6), not by the standard deviation. LayerNorm divides by std, not range.
Why D tempts people
−0.4472 is xhat[1] for x[1]=4.0 (one unit below the mean of 5.0 divided by std 2.2361 = −0.4472), not for x[0]=2.0 (three units below the mean).

44. Answer it before you see the options: Check yourself — GELU vs ReLU

Prediction

Predict first

GELU(−0.5) = −0.1543 while ReLU(−0.5) = 0.0. Why does this matter for the FFN sublayer?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: GELU provides a smooth, non-zero gradient for negative inputs, avoiding dead neurons

Why: ReLU hard-zeros all negative inputs, so any neuron whose pre-activation is always negative gets zero gradient and never updates (dead neuron). GELU smoothly suppresses near-zero negatives while keeping a small gradient, so neurons can recover. This is why BERT, GPT-2, and GPT-3 prefer GELU.

45. Check yourself — GELU vs ReLU

Check

What distinguishes them?

Check your understanding

GELU(−0.5) = −0.1543 while ReLU(−0.5) = 0.0. Why does this matter for the FFN sublayer?

  • A. GELU provides a smooth, non-zero gradient for negative inputs, avoiding dead neurons (correct)
  • B. GELU is faster to compute than ReLU, reducing training time
  • C. GELU ensures the FFN output sums to zero, which helps LayerNorm
  • D. GELU sets negative values to zero exactly like ReLU but with a random threshold

Answer: A

Why: ReLU hard-zeros all negative inputs, so any neuron whose pre-activation is always negative gets zero gradient and never updates (dead neuron). GELU smoothly suppresses near-zero negatives while keeping a small gradient, so neurons can recover. This is why BERT, GPT-2, and GPT-3 prefer GELU.

Why B tempts people
GELU is actually more expensive than ReLU (it involves the Gaussian CDF or an approximation). Speed is not the reason for choosing it.
Why C tempts people
GELU does not enforce zero-sum output. LayerNorm handles normalization independently; GELU is just a nonlinearity inside the FFN.
Why D tempts people
GELU is a deterministic function — not a random threshold. It approaches 0 smoothly as x → −∞ based on the Gaussian CDF, not a stochastic gate.

46. Rule out three: Check yourself — pre-norm formula

Elimination

Eliminate the wrong options

Which is the correct pre-norm encoder sublayer 1 (the MHA sublayer)?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. x = x + MHA(LayerNorm(x))
  • B. x = LayerNorm(x + MHA(x))
  • C. x = MHA(LayerNorm(x))
  • D. x = LayerNorm(MHA(x) + LayerNorm(x))

Survives elimination: A

Why: Pre-norm: normalize the input before the sublayer, then add the original (un-normalized) x via the residual. This keeps a clean residual path and is what GPT-2/LLaMA use. The formula is x = x + Sub(Norm(x)).

47. Check yourself — pre-norm formula

Check

Identify the correct data flow.

Check your understanding

Which is the correct pre-norm encoder sublayer 1 (the MHA sublayer)?

  • A. x = x + MHA(LayerNorm(x)) (correct)
  • B. x = LayerNorm(x + MHA(x))
  • C. x = MHA(LayerNorm(x))
  • D. x = LayerNorm(MHA(x) + LayerNorm(x))

Answer: A

Why: Pre-norm: normalize the input before the sublayer, then add the original (un-normalized) x via the residual. This keeps a clean residual path and is what GPT-2/LLaMA use. The formula is x = x + Sub(Norm(x)).

Why B tempts people
x = LayerNorm(x + MHA(x)) is post-norm — the norm wraps the residual sum, not the sublayer input. Harder to train deep networks; the original 2017 transformer used this.
Why C tempts people
x = MHA(LayerNorm(x)) discards the residual connection entirely. Without the + x term, gradients vanish in deep stacks and the network cannot learn an identity mapping as a base case.
Why D tempts people
Double LayerNorm is not a standard pattern and normalizing the residual path before the add destroys the gradient highway that makes residuals useful.

48. Your turn: build EncoderBlock

Section

Project

49. Project: LayerNorm + FFN + EncoderBlock

Concept

Build all three components from scratch: LayerNorm (verify vs nn.LayerNorm), FFN sublayer (Linear+GELU+Linear), full EncoderBlock nn.Module with pre-norm residual wiring.

#milestonekey check
1LayerNorm from scratch on x=[1,3,5,7]match nn.LayerNorm to atol=1e-4
2FFN(d_model=4, d_ff=16) — shape + param countoutput (1,4), params=148
3Full EncoderBlock(d=8, h=2, d_ff=32)output (B,T,8), params=840

Build rules: type every line, implement LayerNorm manually before reaching for nn.LayerNorm, and verify shapes after each sublayer with print(x.shape).

50. Break it if you can: Project: LayerNorm + FFN + EncoderBlock

Counterexample

Discussion prompt

Build all three components from scratch: LayerNorm (verify vs nn.LayerNorm), FFN sublayer (Linear+GELU+Linear), full EncoderBlock nn.Module with pre-norm residual wiring.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Build rules: type every line, implement LayerNorm manually before reaching for nn.LayerNorm, and verify shapes after each sublayer with print(x.shape).

51. Milestone 1 — LayerNorm from scratch

Worked example

Your turn: implement LayerNorm manually on x = [1, 3, 5, 7]. Predict xhat before running.

Hint: mean = x.mean(dim=-1, keepdim=True); var = x.var(dim=-1, keepdim=True, unbiased=False); xhat = (x - mean) / (var + 1e-5).sqrt().

import torch, torch.nn as nn
x = torch.tensor([[1.0, 3.0, 5.0, 7.0]])
mean = x.mean(dim=-1, keepdim=True)
var  = x.var(dim=-1, keepdim=True, unbiased=False)
xhat = (x - mean) / (var + 1e-5).sqrt()
ln   = nn.LayerNorm(4, elementwise_affine=False, eps=1e-5)
print('scratch:', [round(v,4) for v in xhat.squeeze().tolist()])
print('nn.LN:  ', [round(v,4) for v in ln(x).squeeze().tolist()])
print('match:  ', torch.allclose(xhat, ln(x), atol=1e-4))
quantityvalue
mean4.0
var5.0000
std2.2361
xhat[−1.3416, −0.4472, 0.4472, 1.3416]
matchTrue

52. Milestone 2 — FFN sublayer

Worked example

Your turn: implement FFN(d_model=4, d_ff=16), pass x = [[1, 0, −1, 0.5]], and count the parameters. Predict the param count before running.

Hint: nn.Linear(d_model, d_ff) → F.gelu(...) → nn.Linear(d_ff, d_model). Params = d_model*d_ff + d_ff + d_ff*d_model + d_model.

import torch, torch.nn as nn, torch.nn.functional as F
torch.manual_seed(0)
class FFN(nn.Module):
    def __init__(self, d_model, d_ff):
        super().__init__()
        self.w1 = nn.Linear(d_model, d_ff)
        self.w2 = nn.Linear(d_ff, d_model)
    def forward(self, x): return self.w2(F.gelu(self.w1(x)))

ffn = FFN(4, 16)
x   = torch.tensor([[1.0, 0.0, -1.0, 0.5]])
out = ffn(x)
params = sum(p.numel() for p in ffn.parameters())
print('output:', [round(v,4) for v in out.squeeze().tolist()])
print('params:', params)
outputvalue
FFN(x)[0]−0.3974
FFN(x)[1] 0.2011
FFN(x)[2] 0.0934
FFN(x)[3]−0.1582
params148

53. What each one costs: Milestone 2 — FFN sublayer

Trade off

Comparison matrix

From Milestone 2 — FFN sublayer: every row here is a choice with a cost. Fill the value column, then say which row you would actually pick and what you give up for it.

outputvalue
FFN(x)[0]−0.3974
FFN(x)[1]0.2011
FFN(x)[2]0.0934
FFN(x)[3]−0.1582
params148

54. Milestone 3 — full EncoderBlock

Worked example

Your turn: assemble EncoderBlock(d=8, h=2, d_ff=32) with pre-norm residuals. Verify that output.shape == input.shape and params == 840.

Hint: forward: x = x + attn(norm1(x)); x = x + ffn(norm2(x)). Use nn.Sequential(nn.Linear(d,d_ff), nn.GELU(), nn.Linear(d_ff,d)) for the FFN.

import torch, torch.nn as nn
torch.manual_seed(42)
# (reuse MultiHeadSelfAttn from Milestone 2 shell)
class EncoderBlock(nn.Module):
    def __init__(self, d, h, d_ff):
        super().__init__()
        self.norm1 = nn.LayerNorm(d)
        self.attn  = MultiHeadSelfAttn(d, h)   # from prev. milestone
        self.norm2 = nn.LayerNorm(d)
        self.ffn   = nn.Sequential(
            nn.Linear(d, d_ff), nn.GELU(), nn.Linear(d_ff, d))
    def forward(self, x):
        x = x + self.attn(self.norm1(x))
        x = x + self.ffn(self.norm2(x))
        return x

block = EncoderBlock(d=8, h=2, d_ff=32)
x = torch.randn(2, 5, 8)
out = block(x)
params = sum(p.numel() for p in block.parameters())
print(f'input:  {x.shape}')    # torch.Size([2, 5, 8])
print(f'output: {out.shape}')  # torch.Size([2, 5, 8])
print(f'params: {params}')     # 840
checkexpectednote
output.shape(2, 5, 8)same as input — shape-preserving
params840MHA 256 + FFN 552 + LN 32
sublayer 1x + attn(norm1(x))pre-norm, residual
sublayer 2x + ffn(norm2(x))pre-norm, residual

55. Which is which, by note

Discrimination

Sort into buckets

Sort these by note, from memory, without looking back at Milestone 3 — full EncoderBlock. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

same as input — shape-preserving
output.shape
MHA 256 + FFN 552 + LN 32
params
pre-norm, residual
sublayer 1; sublayer 2
g1
note is "same as input — shape-preserving" for output.shape — that is what the table on "Milestone 3 — full EncoderBlock" records, and it is the single property separating this group from the rest.
g2
note is "MHA 256 + FFN 552 + LN 32" for params — that is what the table on "Milestone 3 — full EncoderBlock" records, and it is the single property separating this group from the rest.
g3
note is "pre-norm, residual" for sublayer 1, sublayer 2 — that is what the table on "Milestone 3 — full EncoderBlock" records, and it is the single property separating this group from the rest.

56. The full program

Concept

import torch, torch.nn as nn, torch.nn.functional as F
torch.manual_seed(42)

# LayerNorm from scratch
def layernorm_scratch(x, eps=1e-5):
    mu  = x.mean(-1, keepdim=True)
    var = x.var(-1, keepdim=True, unbiased=False)
    return (x - mu) / (var + eps).sqrt()

# Multi-head self-attention
class MHA(nn.Module):
    def __init__(self, d, h):
        super().__init__()
        self.h = h; self.dk = d // h
        self.Wq = nn.Linear(d,d,bias=False); self.Wk = nn.Linear(d,d,bias=False)
        self.Wv = nn.Linear(d,d,bias=False); self.Wo = nn.Linear(d,d,bias=False)
    def forward(self, x):
        B,T,D = x.shape
        def sp(W): return W(x).view(B,T,self.h,self.dk).transpose(1,2)
        Q,K,V = sp(self.Wq), sp(self.Wk), sp(self.Wv)
        return self.Wo(((Q@K.transpose(-2,-1)/self.dk**0.5).softmax(-1)@V)
                       .transpose(1,2).reshape(B,T,D))

# EncoderBlock (pre-norm)
class EncoderBlock(nn.Module):
    def __init__(self, d, h, d_ff):
        super().__init__()
        self.norm1=nn.LayerNorm(d); self.attn=MHA(d,h)
        self.norm2=nn.LayerNorm(d)
        self.ffn=nn.Sequential(nn.Linear(d,d_ff),nn.GELU(),nn.Linear(d_ff,d))
    def forward(self, x):
        x = x + self.attn(self.norm1(x))
        x = x + self.ffn(self.norm2(x))
        return x

# verify
xv = torch.tensor([[1.0,3.0,5.0,7.0]])
print('LN match:', torch.allclose(layernorm_scratch(xv), nn.LayerNorm(4,elementwise_affine=False)(xv),atol=1e-4))
enc = EncoderBlock(8,2,32)
xe  = torch.randn(2,5,8)
print('shape:', enc(xe).shape)    # torch.Size([2, 5, 8])
print('params:', sum(p.numel() for p in enc.parameters()))  # 840
outputvaluewhat it confirms
LN matchTruescratch matches nn.LayerNorm
shape(2, 5, 8)shape-preserving block
params840matches hand-computed budget

Stack N of these EncoderBlock modules — that's the BERT/GPT encoder. Every block: normalize → attend → normalize → feed-forward, with clean residual paths throughout.

57. Fill in: what it confirms for The full program

Comparison

Comparison matrix

From The full program: refill the what it confirms column from what you know. The rest of the table is as it appeared.

outputvaluewhat it confirms
LN matchTruescratch matches nn.LayerNorm
shape(2, 5, 8)shape-preserving block
params840matches hand-computed budget

58. Show it off

Concept

Out loud, slides closed: explain (1) why LayerNorm normalizes over the feature dimension rather than the batch dimension; (2) the pre-norm formula and why it stabilizes deep training; (3) what the FFN's 4× expansion achieves and why GELU is preferred over ReLU.

Stretch (homework from L82): implement LayerNorm as a proper nn.Module with trainable γ and β; run the pre-norm vs post-norm training experiment on a 6-layer model and log the loss curves; experiment with d_ff = 2·d_model vs 4·d_model and observe quality difference. Next: Lesson 83 — positional encodings and full encoder stack.

59. Connect it up: Lesson 82: LayerNorm, FFN with GELU, and the Full Encoder Block

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — LayerNorm — normalize over features · Pre-norm vs Post-norm & residual connections · FFN sublayer — GELU and the 4× rule · Your turn: build EncoderBlock. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

60. What you can do now

Recap

conceptthe one thing to remember
LayerNormnormalize each token's feature vector; dim=-1
LN vs BNLN: per-token (feature axis); BN: per-feature (batch axis)
pre-normSub(Norm(x)) then add x; residual path is clean
FFNtwo Linear layers + GELU; 4× expansion; position-wise
EncoderBlocknorm→MHA→residual, then norm→FFN→residual; shape-preserving

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 82 — LayerNorm, FFN, Encoder Block — Barron · USAAIO Round 2 Preparation, 2026
  2. LayerNorm from scratch vs nn.LayerNorm, pre-norm/post-norm gradient norms, GELU trace, EncoderBlock forward pass — all verified with torch 2.7.1+cpu and numpy 2.2.6, June 2026 — Real execution, verified

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108