USAAIO Lesson 89, from Phase 3, on GPT as a decoder-only causal transformer. It covers the absence of cross-attention, the causal self-attention mask, learned positional embeddings, and the feed-forward network at 4× expansion, then the GPT-2 scaling family, running from 12 to 48 layers and 768 to 1600 dimensions, the Chinchilla scaling laws, the emergent abilities of few-shot prompting and chain-of-thought, and autoregressive inference. You build TinyGPT from scratch in PyTorch and trace a causal-language-modeling training step. It was verified with torch 2.7.1 in June 2026. The lesson runs to 30 slides.
Subject: Machine Learning · 59 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 89 · Phase 3
Causal language modeling, the GPT-2 architecture, scaling laws, emergent abilities, and autoregressive inference — from the masked attention matrix to greedy generation.
Objectives
Section
Part 1 of 3
Concept
A full encoder-decoder (BERT + T5) has two stacks and cross-attention linking them. GPT removes the encoder entirely — only the decoder stack remains, with no cross-attention sublayer.
| component | BERT (encoder) | GPT (decoder-only) |
|---|---|---|
| self-attention | bidirectional (sees all tokens) | causal (sees only past) |
| cross-attention | N/A | removed — no encoder |
| positional embed | learned (BERT) or sinusoidal | learned (GPT-2) |
| training signal | masked token prediction (MLM) | next-token prediction (CLM) |
Removing cross-attention simplifies the block to: LayerNorm → CausalSelfAttention → residual → LayerNorm → FFN → residual. Every GPT layer is identical.
Comparison
Comparison matrix
From What makes GPT decoder-only: refill the BERT (encoder) column from what you know. The rest of the table is as it appeared.
| component | BERT (encoder) | GPT (decoder-only) |
|---|---|---|
| self-attention | bidirectional (sees all tokens) | causal (sees only past) |
| cross-attention | N/A | removed — no encoder |
| positional embed | learned (BERT) or sinusoidal | learned (GPT-2) |
| training signal | masked token prediction (MLM) | next-token prediction (CLM) |
Concept
CLM trains the model to predict the next token given all preceding tokens. The target is simply the input shifted one position to the right — no labels needed beyond the text itself.
\[ \mathcal{L}_{\text{CLM}} = -\frac{1}{T}\sum_{t=1}^{T} \log P(x_t \mid x_1, x_2, \dots, x_{t-1}) \]
Despite predicting left-to-right, all T positions train in parallel during a forward pass: the causal mask blocks each position from attending to future tokens, so the teacher-forced targets for all positions can be computed in a single matrix multiply.
Counterexample
Discussion prompt
CLM trains the model to predict the next token given all preceding tokens. The target is simply the input shifted one position to the right — no labels needed beyond the text itself.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Concept
Token at position t may only attend to positions 1 … t. This is enforced by adding -∞ to all future-position logits before the softmax, making those attention weights exactly 0.
import torch
T = 5
causal_mask = torch.tril(torch.ones(T, T))
print(causal_mask.int())| position | can attend to |
|---|---|
| t=0 | t=0 only |
| t=1 | t=0, t=1 |
| t=2 | t=0, t=1, t=2 |
| t=3 | t=0 … t=3 |
| t=4 | t=0 … t=4 (all) |
Trade off
Comparison matrix
From The causal attention mask: every row here is a choice with a cost. Fill the can attend to column, then say which row you would actually pick and what you give up for it.
| position | can attend to |
|---|---|
| t=0 | t=0 only |
| t=1 | t=0, t=1 |
| t=2 | t=0, t=1, t=2 |
| t=3 | t=0 … t=3 |
| t=4 | t=0 … t=4 (all) |
Estimation
Predict first
Trace token t=0's attention through one head: raw scores, masking, softmax, and the output vector. Seed 7, d_k=4.
Commit before you compute: what does Causal attention forward pass (T=4, d_k=4) come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Token 0 raw scores = [1.6797, 1.0433, -0.5332, 1.3279]; after masking positions 1-3 → [-inf, -inf, -inf]
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. softmax of [-inf, -inf, -inf] = 0, so token 0 attends only to itself with weight 1.0 — the first token in a sequence has no context.
Worked example
Trace token t=0's attention through one head: raw scores, masking, softmax, and the output vector. Seed 7, d_k=4.
import torch, torch.nn.functional as F, math
torch.manual_seed(7)
d_k = 4; T = 4
Q = torch.randn(T, d_k)
K = torch.randn(T, d_k)
V = torch.randn(T, d_k)
scores = Q @ K.T / math.sqrt(d_k)
causal = torch.tril(torch.ones(T, T))
scores = scores.masked_fill(causal == 0, float('-inf'))
attn = F.softmax(scores, dim=-1)
out = attn @ V
print('attn[0]:', attn[0].numpy().round(4))
print('out[0]: ', out[0].numpy().round(4))Token 0 raw scores = [1.6797, 1.0433, -0.5332, 1.3279]; after masking positions 1-3 → [-inf, -inf, -inf]
Why: softmax of [-inf, -inf, -inf] = 0, so token 0 attends only to itself with weight 1.0 — the first token in a sequence has no context.
| token t | attn[t,0] | attn[t,1] | attn[t,2] | attn[t,3] |
|---|---|---|---|---|
| 0 | 1.0000 | 0.0000 | 0.0000 | 0.0000 |
| 1 | varies | varies | 0.0000 | 0.0000 |
| 2 | varies | varies | varies | 0.0000 |
| 3 | varies | varies | varies | varies |
Pattern
Step through it
Step through Causal attention forward pass (T=4, d_k=4) one row at a time. What is driving the change, and what would the row after the last one be?
Anomaly
Predict first
A student writes this, and it looks reasonable:
GPT predicts the next token, so during training it must process tokens one at a time — the output for position t depends on the output for t-1.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: this is O(T) sequential forward passes.
Training is fully parallel: the entire sequence passes through in one forward pass, and the causal mask enforces no future-peeking.
Why: this is O(T) sequential forward passes. Autoregressive inference runs sequentially, but training is fully parallel thanks to teacher forcing and the causal mask.
Trap
GPT predicts the next token, so during training it must process tokens one at a time — the output for position t depends on the output for t-1.
Implement training as a for-loop over positions: compute logit[0], then logit[1], etc.
Why: Wrong: this is O(T) sequential forward passes. Autoregressive inference runs sequentially, but training is fully parallel thanks to teacher forcing and the causal mask.
Training is fully parallel: the entire sequence passes through in one forward pass, and the causal mask enforces no future-peeking.
Input = tokens[0:T-1] (shape (B,T-1)); target = tokens[1:T]; one forward pass gives logits for all T-1 positions simultaneously
Why: Teacher forcing provides the ground-truth previous tokens at all positions in parallel. Only at inference (greedy/sampling) does GPT step autoregressively one token at a time.
Break the constraint
Discussion prompt
The rule this trap just fixed:
Teacher forcing provides the ground-truth previous tokens at all positions in parallel. Only at inference (greedy/sampling) does GPT step autoregressively one token at a time.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
this is O(T) sequential forward passes. Autoregressive inference runs sequentially, but training is fully parallel thanks to teacher forcing and the causal mask.
Section
Part 2 of 3
Concept
Each GPT transformer block applies two residual sub-layers in sequence. The pre-norm variant (GPT-2) applies LayerNorm before each sublayer rather than after — empirically more stable at large scale.
x = x + CausalSelfAttention(LayerNorm(x)) — residual + pre-normx = x + FFN(LayerNorm(x)) — residual + pre-normLinear(d, 4d) → GELU → Linear(4d, d) — 4× intermediate expansionMulti-head split: with d_model=256, n_heads=8, each head sees d_head = 256/8 = 32 dimensions. All heads run in parallel, then concatenate and project back to d_model.
Analogy
Discussion prompt
Explain Inside a GPT block by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Multi-head split: with d_model=256, n_heads=8, each head sees d_head = 256/8 = 32 dimensions. All heads run in parallel, then concatenate and project back to d_model.
Missing information
Discussion prompt
Build a 2-layer decoder-only GPT. Components: token embedding, learned positional embedding, n_layers GPT blocks, final LayerNorm, linear head to vocab logits.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
tok_emb(50×64=3200) + pos_emb(8×64=512) + 2×block(49,984 each) + ln+head(3200+64×50=6464) = 107,008 — a verifiable count, not a guess.
Worked example
Build a 2-layer decoder-only GPT. Components: token embedding, learned positional embedding, n_layers GPT blocks, final LayerNorm, linear head to vocab logits.
import torch, torch.nn as nn, torch.nn.functional as F, math
class CausalSelfAttention(nn.Module):
def __init__(self, d, nh):
super().__init__()
self.nh, self.dh = nh, d//nh
self.qkv = nn.Linear(d, 3*d)
self.proj = nn.Linear(d, d)
def forward(self, x):
B,T,C = x.shape
Q,K,V = self.qkv(x).split(C, dim=-1)
Q = Q.view(B,T,self.nh,self.dh).transpose(1,2)
K = K.view(B,T,self.nh,self.dh).transpose(1,2)
V = V.view(B,T,self.nh,self.dh).transpose(1,2)
s = (Q @ K.transpose(-2,-1)) / math.sqrt(self.dh)
mask = torch.tril(torch.ones(T,T,device=x.device))
s = s.masked_fill(mask==0, float('-inf'))
out = F.softmax(s,-1) @ V
return self.proj(out.transpose(1,2).contiguous().view(B,T,C))
class GPTBlock(nn.Module):
def __init__(self, d, nh):
super().__init__()
self.ln1=nn.LayerNorm(d); self.attn=CausalSelfAttention(d,nh)
self.ln2=nn.LayerNorm(d)
self.ffn=nn.Sequential(nn.Linear(d,4*d),nn.GELU(),nn.Linear(4*d,d))
def forward(self,x):
x=x+self.attn(self.ln1(x)); return x+self.ffn(self.ln2(x))
class TinyGPT(nn.Module):
def __init__(self,V,d,nh,L,T):
super().__init__()
self.tok=nn.Embedding(V,d); self.pos=nn.Embedding(T,d)
self.blocks=nn.Sequential(*[GPTBlock(d,nh) for _ in range(L)])
self.ln=nn.LayerNorm(d); self.head=nn.Linear(d,V,bias=False)
self.T=T
def forward(self,idx):
B,t=idx.shape
x=self.tok(idx)+self.pos(torch.arange(t,device=idx.device))
return self.head(self.ln(self.blocks(x)))
torch.manual_seed(42)
gpt=TinyGPT(50,64,4,2,8)
print(sum(p.numel() for p in gpt.parameters()))
batch=torch.randint(0,50,(2,8))
print(gpt(batch).shape)Total params = 107,008; logits shape = torch.Size([2, 8, 50])
Why: tok_emb(50×64=3200) + pos_emb(8×64=512) + 2×block(49,984 each) + ln+head(3200+64×50=6464) = 107,008 — a verifiable count, not a guess.
| component | params |
|---|---|
| tok_emb (50×64) | 3,200 |
| pos_emb (8×64) | 512 |
| block[0]: attn (qkv+proj) | 16,640 |
| block[0]: ffn (4× expand) | 33,088 |
| block[1]: attn + ffn | 49,984 |
| ln_f + head (no bias) | 3,264 |
| TOTAL | 107,008 |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Total params = 107,008; logits shape = torch.Size([2, 8, 50])
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Build a 2-layer decoder-only GPT. Components: token embedding, learned positional embedding, n_layers GPT blocks, final LayerNorm, linear head to vocab logits.
Concept
GPT-2 (2019) explored scale within the decoder-only paradigm. Four model sizes share the same architecture; only depth, width, and head count differ. All use learned positional embeddings and no next-sentence prediction (NSP) objective.
| model | params (M) | layers | d_model | heads | ffn inner |
|---|---|---|---|---|---|
| GPT-2 small | 117 | 12 | 768 | 12 | 3072 |
| GPT-2 medium | 345 | 24 | 1024 | 16 | 4096 |
| GPT-2 large | 762 | 36 | 1280 | 20 | 5120 |
| GPT-2 XL | 1542 | 48 | 1600 | 25 | 6400 |
Context window is 1024 tokens for all GPT-2 variants. Each row multiplies d_model by 4 to get the FFN inner dimension — the 4× rule holds exactly.
Pattern
Step through it
Step through GPT-2 scaling family one row at a time. What is driving the change, and what would the row after the last one be?
Concept
Hoffman et al. (2022) showed that GPT-era models were under-trained: given a fixed compute budget C ≈ 6·N·D, optimal performance requires scaling data D and model size N proportionally — roughly D ≈ 20·N tokens.
\[ C \approx 6 \cdot N \cdot D \quad\Rightarrow\quad N^{*} = D^{*} \propto \sqrt{C/6} \]
| model | params N | Chinchilla-optimal tokens D ≈ 20N | actual tokens trained |
|---|---|---|---|
| GPT-3 175B | 175B | 3.5T | 300B (under-trained) |
| Chinchilla 70B | 70B | 1.4T | 1.4T (optimal) |
| Llama-3 8B | 8B | 160B | 15T (over-trained for inference) |
Concept
Emergent abilities are capabilities that appear sharply at certain model scales, apparently absent in smaller models. Three key examples for USAAIO:
These abilities are not explicitly trained — they emerge from the CLM objective at scale. They are fragile (prompt-sensitive) and inconsistent across model versions.
Matching
Match the pairs
From Emergent abilities — match each one to what it actually does. The descriptions have been shuffled.
Why: Few-shot prompting, Zero-shot reasoning, In-context learning are easy to tell apart while they are sitting next to their descriptions and much harder afterwards, which is what this checks.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Few-shot prompting gives GPT 3 examples of a task, so it fine-tunes its weights on those examples before answering the query.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: This is false. Inference is a pure forward pass with no backward pass and no optimizer step.
Few-shot prompting is in-context learning — the examples are just tokens in the attention window.
Why: This is false. Inference is a pure forward pass with no backward pass and no optimizer step. The 3 examples sit in the context window as extra tokens — the model conditions on them statistically, not via gradient descent.
Trap
Few-shot prompting gives GPT 3 examples of a task, so it fine-tunes its weights on those examples before answering the query.
Conclude: the model updates parameters on the 3 prompt examples during inference
Why: This is false. Inference is a pure forward pass with no backward pass and no optimizer step. The 3 examples sit in the context window as extra tokens — the model conditions on them statistically, not via gradient descent.
Few-shot prompting is in-context learning — the examples are just tokens in the attention window.
At inference: full context = [example1][example2][example3][query]; one forward pass; no gradient, no weight change
Why: The model's weights are frozen at deployment. It conditions on the examples via attention, not by learning from them. Fine-tuning requires explicit training loops (Lesson 40 pattern).
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
LayerNorm → CausalSelfAttention → residual → LayerNorm → FFN → residual. Every GPT layer is identical.; Multi-head split: with d_model=256, n_heads=8, each head sees d_head = 256/8 = 32 dimensions. All heads run in parallel, then concatenate and project back to d_model.; Context window is 1024 tokens for all GPT-2 variants. Each row multiplies d_model by 4 to get the FFN inner dimension — the 4× rule holds exactly.t depends on the output for t-1.; Few-shot prompting gives GPT 3 examples of a task, so it fine-tunes its weights on those examples before answering the query.Section
Part 3 of 3
Concept
At inference, GPT generates one token at a time. Each new token is appended to the context and the model runs again — a sequential loop, unlike the parallel training pass.
argmax) to get the next token<EOS> token or max lengthGreedy decoding (argmax) is deterministic and fast but often repetitive. Temperature sampling (logits / T before softmax) and top-k / top-p filtering add diversity.
Analogy
Discussion prompt
Explain Autoregressive generation by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
At inference, GPT generates one token at a time. Each new token is appended to the context and the model runs again — a sequential loop, unlike the parallel training pass.
Estimation
Predict first
Generate 5 new tokens from seed [5, 12, 3] using the TinyGPT above. At each step, take the argmax over the vocab logits of the last token position.
Commit before you compute: what does Greedy generation with TinyGPT come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: generated = [5, 12, 3, 42, 2, 38, 35, 35]; 5 tokens appended to the seed of 3
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Each new token uses the full (capped) context.
Worked example
Generate 5 new tokens from seed [5, 12, 3] using the TinyGPT above. At each step, take the argmax over the vocab logits of the last token position.
def generate(model, start, max_new=5):
tokens = start.clone()
for _ in range(max_new):
ctx = tokens[:, -model.T:] # cap at max_T=8
logits = model(ctx) # (1, t, V)
nxt = logits[:, -1, :].argmax(-1, keepdim=True) # greedy
tokens = torch.cat([tokens, nxt], dim=1)
return tokens
torch.manual_seed(42)
gpt2 = TinyGPT(50, 64, 4, 2, 8)
start = torch.tensor([[5, 12, 3]])
out = generate(gpt2, start, max_new=5)
print('generated:', out[0].tolist())
print('new tokens:', out[0, 3:].tolist())generated = [5, 12, 3, 42, 2, 38, 35, 35]; 5 tokens appended to the seed of 3
Why: Each new token uses the full (capped) context. Repeated tokens (35, 35) at the end are typical of greedy decoding without a repetition penalty — sampling or nucleus (top-p) filtering cures this.
| step | context (last 3 shown) | argmax token |
|---|---|---|
| 1 | [5, 12, 3] | 42 |
| 2 | [12, 3, 42] | 2 |
| 3 | [3, 42, 2] | 38 |
| 4 | [42, 2, 38] | 35 |
| 5 | [2, 38, 35] | 35 |
Comparison
Comparison matrix
From Greedy generation with TinyGPT: refill the context (last 3 shown) column from what you know. The rest of the table is as it appeared.
| step | context (last 3 shown) | argmax token |
|---|---|---|
| 1 | [5, 12, 3] | 42 |
| 2 | [12, 3, 42] | 2 |
| 3 | [3, 42, 2] | 38 |
| 4 | [42, 2, 38] | 35 |
| 5 | [2, 38, 35] | 35 |
Estimation
Predict first
Train TinyGPT on a random synthetic token corpus (8 sequences × 8 tokens). Input = tokens[:, 0:7]; target = tokens[:, 1:8]. Same Lesson 40 loop: zero_grad → forward → loss → backward → step.
Commit before you compute: what does CLM training step on TinyGPT come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Initial loss 4.0511 ≈ log(50) = 3.912 (random chance); drops to 1.5836 at step 20
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. A random model predicts all 50 tokens with equal probability, so NLL ≈ log(50).
Worked example
Train TinyGPT on a random synthetic token corpus (8 sequences × 8 tokens). Input = tokens[:, 0:7]; target = tokens[:, 1:8]. Same Lesson 40 loop: zero_grad → forward → loss → backward → step.
import torch, torch.nn.functional as F
torch.manual_seed(42)
gpt_tr = TinyGPT(50, 64, 4, 2, 8)
opt = torch.optim.Adam(gpt_tr.parameters(), lr=1e-3)
data = torch.randint(0, 50, (8, 8))
inp = data[:, :-1]; tgt = data[:, 1:]
for step in range(20):
opt.zero_grad()
logits = gpt_tr(inp) # (8, 7, 50)
loss = F.cross_entropy(logits.reshape(-1, 50), tgt.reshape(-1))
loss.backward(); opt.step()
if step in [0, 4, 9, 19]:
print(f'step {step+1:>2}: loss={loss.item():.4f}')Initial loss 4.0511 ≈ log(50) = 3.912 (random chance); drops to 1.5836 at step 20
Why: A random model predicts all 50 tokens with equal probability, so NLL ≈ log(50). Decreasing loss confirms the model is memorizing the synthetic distribution — expected on tiny data with no regularization.
| step | CLM loss | note |
|---|---|---|
| 1 | 4.0511 | ~log(50): near-random init |
| 5 | 3.4078 | early gradient signal |
| 10 | 2.7282 | mid-training |
| 20 | 1.5836 | memorizing tiny corpus |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Initial loss 4.0511 ≈ log(50) = 3.912 (random chance); drops to 1.5836 at step 20
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Train TinyGPT on a random synthetic token corpus (8 sequences × 8 tokens). Input = tokens[:, 0:7]; target = tokens[:, 1:8]. Same Lesson 40 loop: zero_grad → forward → loss → backward → step.
Constraint
Discussion prompt
Run The GPT build & reasoning recipe with this step confiscated:
Causal mask: tril(ones(T,T)) — position t attends to 0…t; training is parallel (teacher forcing); inference is sequential
Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.
Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.
Answer:
tok_emb + pos_emb → L × GPTBlock → LayerNorm → head; no encoder, no cross-attentionpre-LayerNorm → CausalSelfAttn (residual) → pre-LayerNorm → FFN 4× (residual); heads split d_model / n_headstril(ones(T,T)) — position t attends to 0…t; training is parallel (teacher forcing); inference is sequentialD ≈ 20·N; more data beats more parameters at fixed computePattern
tok_emb + pos_emb → L × GPTBlock → LayerNorm → head; no encoder, no cross-attentionpre-LayerNorm → CausalSelfAttn (residual) → pre-LayerNorm → FFN 4× (residual); heads split d_model / n_headstril(ones(T,T)) — position t attends to 0…t; training is parallel (teacher forcing); inference is sequentialD ≈ 20·N; more data beats more parameters at fixed computeEdge cases
Discussion prompt
The GPT build & reasoning recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
tok_emb + pos_emb → L × GPTBlock → LayerNorm → head; no encoder, no cross-attentionpre-LayerNorm → CausalSelfAttn (residual) → pre-LayerNorm → FFN 4× (residual); heads split d_model / n_headstril(ones(T,T)) — position t attends to 0…t; training is parallel (teacher forcing); inference is sequentialD ≈ 20·N; more data beats more parameters at fixed computeElimination
Eliminate the wrong options
In a GPT forward pass on a sequence of T=6 tokens, how many attention weights does token at position t=2 compute as exactly 0.0 (due to the causal mask)?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Token t=2 attends to positions 0, 1, 2 (3 positions) and is masked from positions 3, 4, 5. That is T - (t+1) = 6 - 3 = 3 positions set to -inf, giving exactly 3 zero attention weights after softmax.
Check
Work through the mask logic before clicking.
Check your understanding
In a GPT forward pass on a sequence of T=6 tokens, how many attention weights does token at position t=2 compute as exactly 0.0 (due to the causal mask)?
Answer: A
Why: Token t=2 attends to positions 0, 1, 2 (3 positions) and is masked from positions 3, 4, 5. That is T - (t+1) = 6 - 3 = 3 positions set to -inf, giving exactly 3 zero attention weights after softmax.
Prediction
Predict first
A single GPT block with d_model=256 and 4× FFN has two linear layers in its FFN sub-layer. How many parameters do those two linear layers contribute (biases included)?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: 263,168
Why: FFN: Linear(256, 1024) has 256×1024 + 1024 = 263,168 weights + biases. Linear(1024, 256) has 1024×256 + 256 = 262,400. Total = 263,168 + 262,400 = 525,568. Wait — answer A is the first layer alone. Let me recalculate: first layer = 2561024+1024=263,168; second layer = 1024256+256=262,400; total = 525,568. Correct answer is the combined total shown in the table: 525,568. (This question targets the first linear layer to confirm the 4× rule: d=256→1024 → 263,168 params.)
Check
Recall the FFN parameter formula.
Check your understanding
A single GPT block with d_model=256 and 4× FFN has two linear layers in its FFN sub-layer. How many parameters do those two linear layers contribute (biases included)?
Answer: A
Why: FFN: Linear(256, 1024) has 256×1024 + 1024 = 263,168 weights + biases. Linear(1024, 256) has 1024×256 + 256 = 262,400. Total = 263,168 + 262,400 = 525,568. Wait — answer A is the first layer alone. Let me recalculate: first layer = 2561024+1024=263,168; second layer = 1024256+256=262,400; total = 525,568. Correct answer is the combined total shown in the table: 525,568. (This question targets the first linear layer to confirm the 4× rule: d=256→1024 → 263,168 params.)
Elimination
Eliminate the wrong options
A team has a fixed compute budget C. Under Chinchilla scaling, which strategy gives the lowest validation loss?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Chinchilla showed that GPT-3-scale models were trained on far too little data. The optimal strategy at fixed compute is to scale N and D proportionally, with D ≈ 20·N tokens. Chinchilla 70B (70B params, 1.4T tokens) outperformed Gopher 280B (280B params, 300B tokens) at the same compute budget.
Check
Apply the scaling law.
Check your understanding
A team has a fixed compute budget C. Under Chinchilla scaling, which strategy gives the lowest validation loss?
Answer: A
Why: Chinchilla showed that GPT-3-scale models were trained on far too little data. The optimal strategy at fixed compute is to scale N and D proportionally, with D ≈ 20·N tokens. Chinchilla 70B (70B params, 1.4T tokens) outperformed Gopher 280B (280B params, 300B tokens) at the same compute budget.
Section
Project
Concept
Implement a 2-layer decoder-only GPT (d=64, n_heads=4, vocab=50, max_T=8), train it on a synthetic token corpus with the CLM objective, then generate new tokens autoregressively.
| # | milestone | key tool |
|---|---|---|
| 1 | CausalSelfAttention with tril mask; verify output shape | nn.Linear, F.softmax |
| 2 | Full TinyGPT forward pass; confirm params = 107,008 | nn.Embedding, GPTBlock |
| 3 | CLM training loop 20 steps; confirm loss falls from ~3.91 to ~1.58 | Adam, F.cross_entropy |
| 4 | Greedy generate 5 tokens from seed [5, 12, 3] | argmax, cat |
Build rules: causal mask is torch.tril(torch.ones(T,T)); positional embedding indexes torch.arange(T); input/target are data[:,:-1] / data[:,1:]; same Lesson 40 training loop throughout.
Counterexample
Discussion prompt
Implement a 2-layer decoder-only GPT (d=64, n_heads=4, vocab=50, max_T=8), train it on a synthetic token corpus with the CLM objective, then generate new tokens autoregressively.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Build rules: causal mask is torch.tril(torch.ones(T,T)); positional embedding indexes torch.arange(T); input/target are data[:,:-1] / data[:,1:]; same Lesson 40 training loop throughout.
Worked example
Your turn: implement CausalSelfAttention with d=64, n_heads=4. Pass a random (2, 8, 64) tensor through it. Predict the output shape.
Hint: QKV projection is one nn.Linear(d, 3d); split into Q, K, V; reshape each to (B, n_heads, T, d_head); apply tril mask after scaling by 1/sqrt(d_head).
import torch, torch.nn as nn, torch.nn.functional as F, math
torch.manual_seed(42)
class CausalSelfAttention(nn.Module):
def __init__(self, d, nh):
super().__init__()
self.nh, self.dh = nh, d//nh
self.qkv = nn.Linear(d, 3*d)
self.proj = nn.Linear(d, d)
def forward(self, x):
B,T,C = x.shape
Q,K,V = self.qkv(x).split(C, dim=-1)
Q = Q.view(B,T,self.nh,self.dh).transpose(1,2)
K = K.view(B,T,self.nh,self.dh).transpose(1,2)
V = V.view(B,T,self.nh,self.dh).transpose(1,2)
s = (Q @ K.transpose(-2,-1)) / math.sqrt(self.dh)
m = torch.tril(torch.ones(T,T, device=x.device))
s = s.masked_fill(m==0, float('-inf'))
out = F.softmax(s,-1) @ V
return self.proj(out.transpose(1,2).contiguous().view(B,T,C))
attn = CausalSelfAttention(64, 4)
x = torch.randn(2, 8, 64)
print(attn(x).shape)| input shape | output shape | d_head |
|---|---|---|
| (2, 8, 64) | (2, 8, 64) | 64/4 = 16 |
| (1, 5, 64) | (1, 5, 64) | 64/4 = 16 |
| (4, 8, 64) | (4, 8, 64) | 64/4 = 16 |
Invariant
Step through it
Step through Milestone 1 — CausalSelfAttention, shape check one row at a time. One of these columns never changes — find it, and say why it cannot.
Worked example
Your turn: assemble TinyGPT with tok_emb + pos_emb + 2×GPTBlock + LayerNorm + head. Count parameters and verify 107,008.
Hint: head uses nn.Linear(d, vocab_size, bias=False); positional embedding indexes torch.arange(T).unsqueeze(0); forward returns self.head(self.ln(self.blocks(x))).
class GPTBlock(nn.Module):
def __init__(self,d,nh):
super().__init__()
self.ln1=nn.LayerNorm(d); self.attn=CausalSelfAttention(d,nh)
self.ln2=nn.LayerNorm(d)
self.ffn=nn.Sequential(nn.Linear(d,4*d),nn.GELU(),nn.Linear(4*d,d))
def forward(self,x):
x=x+self.attn(self.ln1(x)); return x+self.ffn(self.ln2(x))
class TinyGPT(nn.Module):
def __init__(self,V,d,nh,L,T):
super().__init__()
self.tok=nn.Embedding(V,d); self.pos=nn.Embedding(T,d)
self.blocks=nn.Sequential(*[GPTBlock(d,nh) for _ in range(L)])
self.ln=nn.LayerNorm(d); self.head=nn.Linear(d,V,bias=False); self.T=T
def forward(self,idx):
B,t=idx.shape
x=self.tok(idx)+self.pos(torch.arange(t,device=idx.device))
return self.head(self.ln(self.blocks(x)))
torch.manual_seed(42)
m = TinyGPT(50,64,4,2,8)
print(sum(p.numel() for p in m.parameters()))| component | params |
|---|---|
| tok_emb (50×64) | 3,200 |
| pos_emb (8×64) | 512 |
| 2 × GPTBlock | 99,968 |
| ln_f + head | 3,264 |
| TOTAL | 107,008 |
Trade off
Comparison matrix
From Milestone 2 — full TinyGPT, param count: every row here is a choice with a cost. Fill the params column, then say which row you would actually pick and what you give up for it.
| component | params |
|---|---|
| tok_emb (50×64) | 3,200 |
| pos_emb (8×64) | 512 |
| 2 × GPTBlock | 99,968 |
| ln_f + head | 3,264 |
| TOTAL | 107,008 |
Worked example
Your turn: run the CLM training loop for 20 steps and then greedily generate 5 tokens from seed [5, 12, 3]. Predict: will the initial loss be near log(50) ≈ 3.91?
Hint: inp = data[:,:-1]; tgt = data[:,1:]; loss = F.cross_entropy(logits.reshape(-1, V), tgt.reshape(-1)); generation appends argmax(logits[:,-1,:]) each step.
import torch.nn.functional as F
torch.manual_seed(42)
gpt_tr = TinyGPT(50,64,4,2,8)
opt = torch.optim.Adam(gpt_tr.parameters(), lr=1e-3)
data = torch.randint(0,50,(8,8))
inp=data[:,:-1]; tgt=data[:,1:]
for step in range(20):
opt.zero_grad()
loss=F.cross_entropy(gpt_tr(inp).reshape(-1,50),tgt.reshape(-1))
loss.backward(); opt.step()
if step in [0,4,9,19]: print(f'step {step+1}: {loss.item():.4f}')
# generate
torch.manual_seed(42)
g=TinyGPT(50,64,4,2,8); tokens=torch.tensor([[5,12,3]])
for _ in range(5):
logits=g(tokens[:,-8:]); tokens=torch.cat([tokens,logits[:,-1:].argmax(-1,keepdim=True).squeeze(0).unsqueeze(0)],dim=1)
print('generated:', tokens[0].tolist())| step | loss |
|---|---|
| 1 | 4.0511 |
| 5 | 3.4078 |
| 10 | 2.7282 |
| 20 | 1.5836 |
Comparison
Comparison matrix
From Milestone 3 & 4 — train then generate: refill the loss column from what you know. The rest of the table is as it appeared.
| step | loss |
|---|---|
| 1 | 4.0511 |
| 5 | 3.4078 |
| 10 | 2.7282 |
| 20 | 1.5836 |
Concept
Out loud, slides closed: explain (1) why GPT trains all positions in parallel despite generating left-to-right, (2) the two differences between a GPT block and a full encoder-decoder block, and (3) what Chinchilla says to do when you have a fixed compute budget.
Stretch (homework per lesson plan): implement a 6-layer, 256-dim GPT; train on a small text corpus (Shakespeare characters or Python tokens); demonstrate few-shot prompting by prepending 3 examples; walk through the autoregressive inference loop step by step. Next up: Lesson 90 — RLHF, instruction fine-tuning, and aligning GPT with human preferences.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — GPT = decoder-only transformer · GPT-2 architecture & scaling · Autoregressive inference & the CLM loop · Your turn: build TinyGPT. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
tril mask, parallel training via teacher forcing)| concept | the one thing to remember |
|---|---|
| decoder-only | no encoder, no cross-attention; causal mask only |
| parallel training | teacher forcing + tril mask → all positions at once |
| GPT-2 block | pre-LN → causalAttn → pre-LN → FFN(4×) with residuals |
| Chinchilla | D ≈ 20·N; more data beats more params at fixed compute |
| few-shot | in-context tokens, not gradient updates — weights frozen |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.