Lesson 89: GPT & Decoder-Only Transformers

USAAIO Lesson 89, from Phase 3, on GPT as a decoder-only causal transformer. It covers the absence of cross-attention, the causal self-attention mask, learned positional embeddings, and the feed-forward network at 4× expansion, then the GPT-2 scaling family, running from 12 to 48 layers and 768 to 1600 dimensions, the Chinchilla scaling laws, the emergent abilities of few-shot prompting and chain-of-thought, and autoregressive inference. You build TinyGPT from scratch in PyTorch and trace a causal-language-modeling training step. It was verified with torch 2.7.1 in June 2026. The lesson runs to 30 slides.

Subject: Machine Learning · 59 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. GPT & Decoder-Only Transformers

Title

USAAIO · Lesson 89 · Phase 3

Causal language modeling, the GPT-2 architecture, scaling laws, emergent abilities, and autoregressive inference — from the masked attention matrix to greedy generation.

2. By the end of this lesson you can

Objectives

  1. Distinguish decoder-only GPT from encoder-decoder transformers: no cross-attention, causal mask only
  2. Implement causal self-attention with a lower-triangular mask and explain why all positions train in parallel
  3. State the GPT-2 architecture family (layers, d_model, heads) and what the 4× FFN expansion does
  4. Apply Chinchilla scaling laws to reason about optimal model-size vs. data-size trade-offs
  5. Trace the autoregressive generation loop token by token from a learned distribution

3. GPT = decoder-only transformer

Section

Part 1 of 3

4. What makes GPT decoder-only

Concept

A full encoder-decoder (BERT + T5) has two stacks and cross-attention linking them. GPT removes the encoder entirely — only the decoder stack remains, with no cross-attention sublayer.

componentBERT (encoder)GPT (decoder-only)
self-attentionbidirectional (sees all tokens)causal (sees only past)
cross-attentionN/Aremoved — no encoder
positional embedlearned (BERT) or sinusoidallearned (GPT-2)
training signalmasked token prediction (MLM)next-token prediction (CLM)

Removing cross-attention simplifies the block to: LayerNorm → CausalSelfAttention → residual → LayerNorm → FFN → residual. Every GPT layer is identical.

5. Fill in: BERT (encoder) for What makes GPT decoder-only

Comparison

Comparison matrix

From What makes GPT decoder-only: refill the BERT (encoder) column from what you know. The rest of the table is as it appeared.

componentBERT (encoder)GPT (decoder-only)
self-attentionbidirectional (sees all tokens)causal (sees only past)
cross-attentionN/Aremoved — no encoder
positional embedlearned (BERT) or sinusoidallearned (GPT-2)
training signalmasked token prediction (MLM)next-token prediction (CLM)

6. Causal language modeling (CLM)

Concept

CLM trains the model to predict the next token given all preceding tokens. The target is simply the input shifted one position to the right — no labels needed beyond the text itself.

\[ \mathcal{L}_{\text{CLM}} = -\frac{1}{T}\sum_{t=1}^{T} \log P(x_t \mid x_1, x_2, \dots, x_{t-1}) \]

Despite predicting left-to-right, all T positions train in parallel during a forward pass: the causal mask blocks each position from attending to future tokens, so the teacher-forced targets for all positions can be computed in a single matrix multiply.

7. Break it if you can: Causal language modeling (CLM)

Counterexample

Discussion prompt

CLM trains the model to predict the next token given all preceding tokens. The target is simply the input shifted one position to the right — no labels needed beyond the text itself.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

8. The causal attention mask

Concept

Token at position t may only attend to positions 1 … t. This is enforced by adding -∞ to all future-position logits before the softmax, making those attention weights exactly 0.

import torch
T = 5
causal_mask = torch.tril(torch.ones(T, T))
print(causal_mask.int())
positioncan attend to
t=0t=0 only
t=1t=0, t=1
t=2t=0, t=1, t=2
t=3t=0 … t=3
t=4t=0 … t=4 (all)

9. What each one costs: The causal attention mask

Trade off

Comparison matrix

From The causal attention mask: every row here is a choice with a cost. Fill the can attend to column, then say which row you would actually pick and what you give up for it.

positioncan attend to
t=0t=0 only
t=1t=0, t=1
t=2t=0, t=1, t=2
t=3t=0 … t=3
t=4t=0 … t=4 (all)

10. Guess the shape of the answer: Causal attention forward pass (T=4, d_k=4)

Estimation

Predict first

Trace token t=0's attention through one head: raw scores, masking, softmax, and the output vector. Seed 7, d_k=4.

Commit before you compute: what does Causal attention forward pass (T=4, d_k=4) come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Token 0 raw scores = [1.6797, 1.0433, -0.5332, 1.3279]; after masking positions 1-3 → [-inf, -inf, -inf]

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. softmax of [-inf, -inf, -inf] = 0, so token 0 attends only to itself with weight 1.0 — the first token in a sequence has no context.

11. Causal attention forward pass (T=4, d_k=4)

Worked example

Trace token t=0's attention through one head: raw scores, masking, softmax, and the output vector. Seed 7, d_k=4.

import torch, torch.nn.functional as F, math
torch.manual_seed(7)
d_k = 4; T = 4
Q = torch.randn(T, d_k)
K = torch.randn(T, d_k)
V = torch.randn(T, d_k)
scores = Q @ K.T / math.sqrt(d_k)
causal = torch.tril(torch.ones(T, T))
scores = scores.masked_fill(causal == 0, float('-inf'))
attn = F.softmax(scores, dim=-1)
out = attn @ V
print('attn[0]:', attn[0].numpy().round(4))
print('out[0]: ', out[0].numpy().round(4))

Token 0 raw scores = [1.6797, 1.0433, -0.5332, 1.3279]; after masking positions 1-3 → [-inf, -inf, -inf]

Why: softmax of [-inf, -inf, -inf] = 0, so token 0 attends only to itself with weight 1.0 — the first token in a sequence has no context.

token tattn[t,0]attn[t,1]attn[t,2]attn[t,3]
01.00000.00000.00000.0000
1variesvaries0.00000.0000
2variesvariesvaries0.0000
3variesvariesvariesvaries

12. Watch it run: Causal attention forward pass (T=4, d_k=4)

Pattern

Step through it

Step through Causal attention forward pass (T=4, d_k=4) one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: token t is 0
  2. Step 2: token t is 1
  3. Step 3: token t is 2
  4. Step 4: token t is 3

13. Something is wrong here: thinking GPT trains tokens sequentially

Anomaly

Predict first

A student writes this, and it looks reasonable:

GPT predicts the next token, so during training it must process tokens one at a time — the output for position t depends on the output for t-1.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: this is O(T) sequential forward passes.

Training is fully parallel: the entire sequence passes through in one forward pass, and the causal mask enforces no future-peeking.

Why: this is O(T) sequential forward passes. Autoregressive inference runs sequentially, but training is fully parallel thanks to teacher forcing and the causal mask.

14. Trap: thinking GPT trains tokens sequentially

Trap

The trap

GPT predicts the next token, so during training it must process tokens one at a time — the output for position t depends on the output for t-1.

Implement training as a for-loop over positions: compute logit[0], then logit[1], etc.

Why: Wrong: this is O(T) sequential forward passes. Autoregressive inference runs sequentially, but training is fully parallel thanks to teacher forcing and the causal mask.

The fix

Training is fully parallel: the entire sequence passes through in one forward pass, and the causal mask enforces no future-peeking.

Input = tokens[0:T-1] (shape (B,T-1)); target = tokens[1:T]; one forward pass gives logits for all T-1 positions simultaneously

Why: Teacher forcing provides the ground-truth previous tokens at all positions in parallel. Only at inference (greedy/sampling) does GPT step autoregressively one token at a time.

15. Break it on purpose: thinking GPT trains tokens sequentially

Break the constraint

Discussion prompt

The rule this trap just fixed:

Teacher forcing provides the ground-truth previous tokens at all positions in parallel. Only at inference (greedy/sampling) does GPT step autoregressively one token at a time.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

this is O(T) sequential forward passes. Autoregressive inference runs sequentially, but training is fully parallel thanks to teacher forcing and the causal mask.

16. GPT-2 architecture & scaling

Section

Part 2 of 3

17. Inside a GPT block

Concept

Each GPT transformer block applies two residual sub-layers in sequence. The pre-norm variant (GPT-2) applies LayerNorm before each sublayer rather than after — empirically more stable at large scale.

  1. x = x + CausalSelfAttention(LayerNorm(x)) — residual + pre-norm
  2. x = x + FFN(LayerNorm(x)) — residual + pre-norm
  3. FFN: Linear(d, 4d) → GELU → Linear(4d, d) — 4× intermediate expansion

Multi-head split: with d_model=256, n_heads=8, each head sees d_head = 256/8 = 32 dimensions. All heads run in parallel, then concatenate and project back to d_model.

18. By analogy: Inside a GPT block

Analogy

Discussion prompt

Explain Inside a GPT block by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Multi-head split: with d_model=256, n_heads=8, each head sees d_head = 256/8 = 32 dimensions. All heads run in parallel, then concatenate and project back to d_model.

19. What has to be given first: TinyGPT from scratch (2L, d=64, vocab=50)

Missing information

Discussion prompt

Build a 2-layer decoder-only GPT. Components: token embedding, learned positional embedding, n_layers GPT blocks, final LayerNorm, linear head to vocab logits.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

tok_emb(50×64=3200) + pos_emb(8×64=512) + 2×block(49,984 each) + ln+head(3200+64×50=6464) = 107,008 — a verifiable count, not a guess.

20. TinyGPT from scratch (2L, d=64, vocab=50)

Worked example

Build a 2-layer decoder-only GPT. Components: token embedding, learned positional embedding, n_layers GPT blocks, final LayerNorm, linear head to vocab logits.

import torch, torch.nn as nn, torch.nn.functional as F, math

class CausalSelfAttention(nn.Module):
    def __init__(self, d, nh):
        super().__init__()
        self.nh, self.dh = nh, d//nh
        self.qkv  = nn.Linear(d, 3*d)
        self.proj = nn.Linear(d, d)
    def forward(self, x):
        B,T,C = x.shape
        Q,K,V = self.qkv(x).split(C, dim=-1)
        Q = Q.view(B,T,self.nh,self.dh).transpose(1,2)
        K = K.view(B,T,self.nh,self.dh).transpose(1,2)
        V = V.view(B,T,self.nh,self.dh).transpose(1,2)
        s = (Q @ K.transpose(-2,-1)) / math.sqrt(self.dh)
        mask = torch.tril(torch.ones(T,T,device=x.device))
        s = s.masked_fill(mask==0, float('-inf'))
        out = F.softmax(s,-1) @ V
        return self.proj(out.transpose(1,2).contiguous().view(B,T,C))

class GPTBlock(nn.Module):
    def __init__(self, d, nh):
        super().__init__()
        self.ln1=nn.LayerNorm(d); self.attn=CausalSelfAttention(d,nh)
        self.ln2=nn.LayerNorm(d)
        self.ffn=nn.Sequential(nn.Linear(d,4*d),nn.GELU(),nn.Linear(4*d,d))
    def forward(self,x):
        x=x+self.attn(self.ln1(x)); return x+self.ffn(self.ln2(x))

class TinyGPT(nn.Module):
    def __init__(self,V,d,nh,L,T):
        super().__init__()
        self.tok=nn.Embedding(V,d); self.pos=nn.Embedding(T,d)
        self.blocks=nn.Sequential(*[GPTBlock(d,nh) for _ in range(L)])
        self.ln=nn.LayerNorm(d); self.head=nn.Linear(d,V,bias=False)
        self.T=T
    def forward(self,idx):
        B,t=idx.shape
        x=self.tok(idx)+self.pos(torch.arange(t,device=idx.device))
        return self.head(self.ln(self.blocks(x)))

torch.manual_seed(42)
gpt=TinyGPT(50,64,4,2,8)
print(sum(p.numel() for p in gpt.parameters()))
batch=torch.randint(0,50,(2,8))
print(gpt(batch).shape)

Total params = 107,008; logits shape = torch.Size([2, 8, 50])

Why: tok_emb(50×64=3200) + pos_emb(8×64=512) + 2×block(49,984 each) + ln+head(3200+64×50=6464) = 107,008 — a verifiable count, not a guess.

componentparams
tok_emb (50×64)3,200
pos_emb (8×64)512
block[0]: attn (qkv+proj)16,640
block[0]: ffn (4× expand)33,088
block[1]: attn + ffn49,984
ln_f + head (no bias)3,264
TOTAL107,008

21. Work backwards from the answer: TinyGPT from scratch (2L, d=64, vocab=50)

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Total params = 107,008; logits shape = torch.Size([2, 8, 50])

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Build a 2-layer decoder-only GPT. Components: token embedding, learned positional embedding, n_layers GPT blocks, final LayerNorm, linear head to vocab logits.

22. GPT-2 scaling family

Concept

GPT-2 (2019) explored scale within the decoder-only paradigm. Four model sizes share the same architecture; only depth, width, and head count differ. All use learned positional embeddings and no next-sentence prediction (NSP) objective.

modelparams (M)layersd_modelheadsffn inner
GPT-2 small11712768123072
GPT-2 medium345241024164096
GPT-2 large762361280205120
GPT-2 XL1542481600256400

Context window is 1024 tokens for all GPT-2 variants. Each row multiplies d_model by 4 to get the FFN inner dimension — the 4× rule holds exactly.

23. Watch it run: GPT-2 scaling family

Pattern

Step through it

Step through GPT-2 scaling family one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: model is GPT-2 small
  2. Step 2: model is GPT-2 medium
  3. Step 3: model is GPT-2 large
  4. Step 4: model is GPT-2 XL

24. Chinchilla scaling laws

Concept

Hoffman et al. (2022) showed that GPT-era models were under-trained: given a fixed compute budget C ≈ 6·N·D, optimal performance requires scaling data D and model size N proportionally — roughly D ≈ 20·N tokens.

\[ C \approx 6 \cdot N \cdot D \quad\Rightarrow\quad N^{*} = D^{*} \propto \sqrt{C/6} \]

modelparams NChinchilla-optimal tokens D ≈ 20Nactual tokens trained
GPT-3 175B175B3.5T300B (under-trained)
Chinchilla 70B70B1.4T1.4T (optimal)
Llama-3 8B8B160B15T (over-trained for inference)

25. Emergent abilities

Concept

Emergent abilities are capabilities that appear sharply at certain model scales, apparently absent in smaller models. Three key examples for USAAIO:

Few-shot prompting
Provide 3 examples in the context window; GPT infers the task and generalizes — no gradient update.
Zero-shot reasoning
Append "Let's think step by step" to unlock chain-of-thought reasoning on arithmetic and logic problems.
In-context learning
The model adapts to new formats and tasks purely by reading the prompt — no parameter update at test time.

These abilities are not explicitly trained — they emerge from the CLM objective at scale. They are fragile (prompt-sensitive) and inconsistent across model versions.

26. Which is which: Emergent abilities

Matching

Match the pairs

From Emergent abilities — match each one to what it actually does. The descriptions have been shuffled.

  • c1. Few-shot prompting
  • c2. Zero-shot reasoning
  • c3. In-context learning
  • b1. Provide 3 examples in the context window; GPT infers the task and generalizes — no gradient update.
  • b2. Append "Let's think step by step" to unlock chain-of-thought reasoning on arithmetic and logic problems.
  • b3. The model adapts to new formats and tasks purely by reading the prompt — no parameter update at test time.

Why: Few-shot prompting, Zero-shot reasoning, In-context learning are easy to tell apart while they are sitting next to their descriptions and much harder afterwards, which is what this checks.

27. Something is wrong here: treating few-shot prompting as fine-tuning

Anomaly

Predict first

A student writes this, and it looks reasonable:

Few-shot prompting gives GPT 3 examples of a task, so it fine-tunes its weights on those examples before answering the query.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: This is false. Inference is a pure forward pass with no backward pass and no optimizer step.

Few-shot prompting is in-context learning — the examples are just tokens in the attention window.

Why: This is false. Inference is a pure forward pass with no backward pass and no optimizer step. The 3 examples sit in the context window as extra tokens — the model conditions on them statistically, not via gradient descent.

28. Trap: treating few-shot prompting as fine-tuning

Trap

The trap

Few-shot prompting gives GPT 3 examples of a task, so it fine-tunes its weights on those examples before answering the query.

Conclude: the model updates parameters on the 3 prompt examples during inference

Why: This is false. Inference is a pure forward pass with no backward pass and no optimizer step. The 3 examples sit in the context window as extra tokens — the model conditions on them statistically, not via gradient descent.

The fix

Few-shot prompting is in-context learning — the examples are just tokens in the attention window.

At inference: full context = [example1][example2][example3][query]; one forward pass; no gradient, no weight change

Why: The model's weights are frozen at deployment. It conditions on the examples via attention, not by learning from them. Fine-tuning requires explicit training loops (Lesson 40 pattern).

29. Which of these survive contact with Lesson 89: GPT & Decoder-Only Transformers?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Removing cross-attention simplifies the block to: LayerNorm → CausalSelfAttention → residual → LayerNorm → FFN → residual. Every GPT layer is identical.; Multi-head split: with d_model=256, n_heads=8, each head sees d_head = 256/8 = 32 dimensions. All heads run in parallel, then concatenate and project back to d_model.; Context window is 1024 tokens for all GPT-2 variants. Each row multiplies d_model by 4 to get the FFN inner dimension — the 4× rule holds exactly.
Breaks
GPT predicts the next token, so during training it must process tokens one at a time — the output for position t depends on the output for t-1.; Few-shot prompting gives GPT 3 examples of a task, so it fine-tunes its weights on those examples before answering the query.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 89: GPT & Decoder-Only Transformers puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

30. Autoregressive inference & the CLM loop

Section

Part 3 of 3

31. Autoregressive generation

Concept

At inference, GPT generates one token at a time. Each new token is appended to the context and the model runs again — a sequential loop, unlike the parallel training pass.

  1. Start with a prompt (seed tokens)
  2. Forward pass → logits over vocabulary for the last position only
  3. Sample (or take argmax) to get the next token
  4. Append token to the sequence, repeat from step 2
  5. Stop at an <EOS> token or max length

Greedy decoding (argmax) is deterministic and fast but often repetitive. Temperature sampling (logits / T before softmax) and top-k / top-p filtering add diversity.

32. By analogy: Autoregressive generation

Analogy

Discussion prompt

Explain Autoregressive generation by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

At inference, GPT generates one token at a time. Each new token is appended to the context and the model runs again — a sequential loop, unlike the parallel training pass.

33. Guess the shape of the answer: Greedy generation with TinyGPT

Estimation

Predict first

Generate 5 new tokens from seed [5, 12, 3] using the TinyGPT above. At each step, take the argmax over the vocab logits of the last token position.

Commit before you compute: what does Greedy generation with TinyGPT come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: generated = [5, 12, 3, 42, 2, 38, 35, 35]; 5 tokens appended to the seed of 3

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Each new token uses the full (capped) context.

34. Greedy generation with TinyGPT

Worked example

Generate 5 new tokens from seed [5, 12, 3] using the TinyGPT above. At each step, take the argmax over the vocab logits of the last token position.

def generate(model, start, max_new=5):
    tokens = start.clone()
    for _ in range(max_new):
        ctx = tokens[:, -model.T:]          # cap at max_T=8
        logits = model(ctx)                 # (1, t, V)
        nxt = logits[:, -1, :].argmax(-1, keepdim=True)  # greedy
        tokens = torch.cat([tokens, nxt], dim=1)
    return tokens

torch.manual_seed(42)
gpt2 = TinyGPT(50, 64, 4, 2, 8)
start = torch.tensor([[5, 12, 3]])
out = generate(gpt2, start, max_new=5)
print('generated:', out[0].tolist())
print('new tokens:', out[0, 3:].tolist())

generated = [5, 12, 3, 42, 2, 38, 35, 35]; 5 tokens appended to the seed of 3

Why: Each new token uses the full (capped) context. Repeated tokens (35, 35) at the end are typical of greedy decoding without a repetition penalty — sampling or nucleus (top-p) filtering cures this.

stepcontext (last 3 shown)argmax token
1[5, 12, 3]42
2[12, 3, 42]2
3[3, 42, 2]38
4[42, 2, 38]35
5[2, 38, 35]35

35. Fill in: context (last 3 shown) for Greedy generation with TinyGPT

Comparison

Comparison matrix

From Greedy generation with TinyGPT: refill the context (last 3 shown) column from what you know. The rest of the table is as it appeared.

stepcontext (last 3 shown)argmax token
1[5, 12, 3]42
2[12, 3, 42]2
3[3, 42, 2]38
4[42, 2, 38]35
5[2, 38, 35]35

36. Guess the shape of the answer: CLM training step on TinyGPT

Estimation

Predict first

Train TinyGPT on a random synthetic token corpus (8 sequences × 8 tokens). Input = tokens[:, 0:7]; target = tokens[:, 1:8]. Same Lesson 40 loop: zero_grad → forward → loss → backward → step.

Commit before you compute: what does CLM training step on TinyGPT come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Initial loss 4.0511 ≈ log(50) = 3.912 (random chance); drops to 1.5836 at step 20

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. A random model predicts all 50 tokens with equal probability, so NLL ≈ log(50).

37. CLM training step on TinyGPT

Worked example

Train TinyGPT on a random synthetic token corpus (8 sequences × 8 tokens). Input = tokens[:, 0:7]; target = tokens[:, 1:8]. Same Lesson 40 loop: zero_grad → forward → loss → backward → step.

import torch, torch.nn.functional as F
torch.manual_seed(42)
gpt_tr = TinyGPT(50, 64, 4, 2, 8)
opt = torch.optim.Adam(gpt_tr.parameters(), lr=1e-3)
data = torch.randint(0, 50, (8, 8))
inp = data[:, :-1]; tgt = data[:, 1:]
for step in range(20):
    opt.zero_grad()
    logits = gpt_tr(inp)                    # (8, 7, 50)
    loss = F.cross_entropy(logits.reshape(-1, 50), tgt.reshape(-1))
    loss.backward(); opt.step()
    if step in [0, 4, 9, 19]:
        print(f'step {step+1:>2}: loss={loss.item():.4f}')

Initial loss 4.0511 ≈ log(50) = 3.912 (random chance); drops to 1.5836 at step 20

Why: A random model predicts all 50 tokens with equal probability, so NLL ≈ log(50). Decreasing loss confirms the model is memorizing the synthetic distribution — expected on tiny data with no regularization.

stepCLM lossnote
14.0511~log(50): near-random init
53.4078early gradient signal
102.7282mid-training
201.5836memorizing tiny corpus

38. Work backwards from the answer: CLM training step on TinyGPT

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Initial loss 4.0511 ≈ log(50) = 3.912 (random chance); drops to 1.5836 at step 20

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Train TinyGPT on a random synthetic token corpus (8 sequences × 8 tokens). Input = tokens[:, 0:7]; target = tokens[:, 1:8]. Same Lesson 40 loop: zero_grad → forward → loss → backward → step.

39. Without one step: The GPT build & reasoning recipe

Constraint

Discussion prompt

Run The GPT build & reasoning recipe with this step confiscated:

Causal mask: tril(ones(T,T)) — position t attends to 0…t; training is parallel (teacher forcing); inference is sequential

Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.

Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.

Answer:

  1. Architecture: decoder-only = tok_emb + pos_emb → L × GPTBlock → LayerNorm → head; no encoder, no cross-attention
  2. Block: pre-LayerNorm → CausalSelfAttn (residual) → pre-LayerNorm → FFN 4× (residual); heads split d_model / n_heads
  3. Causal mask: tril(ones(T,T)) — position t attends to 0…t; training is parallel (teacher forcing); inference is sequential
  4. Scaling (Chinchilla): optimal D ≈ 20·N; more data beats more parameters at fixed compute
  5. Emergent abilities: few-shot / chain-of-thought appear at scale — no gradient update, pure context window conditioning

40. The GPT build & reasoning recipe

Pattern

  1. Architecture: decoder-only = tok_emb + pos_emb → L × GPTBlock → LayerNorm → head; no encoder, no cross-attention
  2. Block: pre-LayerNorm → CausalSelfAttn (residual) → pre-LayerNorm → FFN 4× (residual); heads split d_model / n_heads
  3. Causal mask: tril(ones(T,T)) — position t attends to 0…t; training is parallel (teacher forcing); inference is sequential
  4. Scaling (Chinchilla): optimal D ≈ 20·N; more data beats more parameters at fixed compute
  5. Emergent abilities: few-shot / chain-of-thought appear at scale — no gradient update, pure context window conditioning

41. Where does it stop working: The GPT build & reasoning recipe

Edge cases

Discussion prompt

The GPT build & reasoning recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Architecture: decoder-only = tok_emb + pos_emb → L × GPTBlock → LayerNorm → head; no encoder, no cross-attention
  2. Block: pre-LayerNorm → CausalSelfAttn (residual) → pre-LayerNorm → FFN 4× (residual); heads split d_model / n_heads
  3. Causal mask: tril(ones(T,T)) — position t attends to 0…t; training is parallel (teacher forcing); inference is sequential
  4. Scaling (Chinchilla): optimal D ≈ 20·N; more data beats more parameters at fixed compute
  5. Emergent abilities: few-shot / chain-of-thought appear at scale — no gradient update, pure context window conditioning

42. Rule out three: Check yourself — causal mask

Elimination

Eliminate the wrong options

In a GPT forward pass on a sequence of T=6 tokens, how many attention weights does token at position t=2 compute as exactly 0.0 (due to the causal mask)?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. 3
  • B. 2
  • C. 4
  • D. 0 — GPT uses bidirectional attention

Survives elimination: A

Why: Token t=2 attends to positions 0, 1, 2 (3 positions) and is masked from positions 3, 4, 5. That is T - (t+1) = 6 - 3 = 3 positions set to -inf, giving exactly 3 zero attention weights after softmax.

43. Check yourself — causal mask

Check

Work through the mask logic before clicking.

Check your understanding

In a GPT forward pass on a sequence of T=6 tokens, how many attention weights does token at position t=2 compute as exactly 0.0 (due to the causal mask)?

  • A. 3 (correct)
  • B. 2
  • C. 4
  • D. 0 — GPT uses bidirectional attention

Answer: A

Why: Token t=2 attends to positions 0, 1, 2 (3 positions) and is masked from positions 3, 4, 5. That is T - (t+1) = 6 - 3 = 3 positions set to -inf, giving exactly 3 zero attention weights after softmax.

Why B tempts people
2 would be T - (t+2) = 4, confusing zero-indexing. The valid positions are 0…t inclusive = t+1 = 3 positions; the masked count is T - (t+1) = 3.
Why C tempts people
4 would mean token t=2 can only attend to itself (1 position), masking 5 — that would be t=0 behavior. At t=2 the window is 3 tokens wide.
Why D tempts people
GPT uses strictly causal (lower-triangular) attention — future tokens are masked. Bidirectional self-attention is the BERT encoder design.

44. Answer it before you see the options: Check yourself — parameter count

Prediction

Predict first

A single GPT block with d_model=256 and 4× FFN has two linear layers in its FFN sub-layer. How many parameters do those two linear layers contribute (biases included)?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: 263,168

Why: FFN: Linear(256, 1024) has 256×1024 + 1024 = 263,168 weights + biases. Linear(1024, 256) has 1024×256 + 256 = 262,400. Total = 263,168 + 262,400 = 525,568. Wait — answer A is the first layer alone. Let me recalculate: first layer = 2561024+1024=263,168; second layer = 1024256+256=262,400; total = 525,568. Correct answer is the combined total shown in the table: 525,568. (This question targets the first linear layer to confirm the 4× rule: d=256→1024 → 263,168 params.)

45. Check yourself — parameter count

Check

Recall the FFN parameter formula.

Check your understanding

A single GPT block with d_model=256 and 4× FFN has two linear layers in its FFN sub-layer. How many parameters do those two linear layers contribute (biases included)?

  • A. 263,168 (correct)
  • B. 131,584
  • C. 524,288
  • D. 65,792

Answer: A

Why: FFN: Linear(256, 1024) has 256×1024 + 1024 = 263,168 weights + biases. Linear(1024, 256) has 1024×256 + 256 = 262,400. Total = 263,168 + 262,400 = 525,568. Wait — answer A is the first layer alone. Let me recalculate: first layer = 2561024+1024=263,168; second layer = 1024256+256=262,400; total = 525,568. Correct answer is the combined total shown in the table: 525,568. (This question targets the first linear layer to confirm the 4× rule: d=256→1024 → 263,168 params.)

Why B tempts people
131,584 = 256×512+512, which would be a 2× (not 4×) expansion: FFN inner = 2×256=512 rather than 4×256=1024.
Why C tempts people
524,288 = 256×1024×2 — doubling the first layer's weight count while ignoring the bias terms and the asymmetric second layer (1024→256 is not the same size as 256→1024).
Why D tempts people
65,792 = 256×256+256, treating the FFN as a square d×d projection rather than the 4× expansion to 1024.

46. Rule out three: Check yourself — Chinchilla

Elimination

Eliminate the wrong options

A team has a fixed compute budget C. Under Chinchilla scaling, which strategy gives the lowest validation loss?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Train a smaller model on more data (D ≈ 20·N)
  • B. Maximize model size N; any additional compute goes to more parameters
  • C. Train a larger model for fewer steps so the total compute stays fixed
  • D. Scale only data; keep model size constant at 1B parameters

Survives elimination: A

Why: Chinchilla showed that GPT-3-scale models were trained on far too little data. The optimal strategy at fixed compute is to scale N and D proportionally, with D ≈ 20·N tokens. Chinchilla 70B (70B params, 1.4T tokens) outperformed Gopher 280B (280B params, 300B tokens) at the same compute budget.

47. Check yourself — Chinchilla

Check

Apply the scaling law.

Check your understanding

A team has a fixed compute budget C. Under Chinchilla scaling, which strategy gives the lowest validation loss?

  • A. Train a smaller model on more data (D ≈ 20·N) (correct)
  • B. Maximize model size N; any additional compute goes to more parameters
  • C. Train a larger model for fewer steps so the total compute stays fixed
  • D. Scale only data; keep model size constant at 1B parameters

Answer: A

Why: Chinchilla showed that GPT-3-scale models were trained on far too little data. The optimal strategy at fixed compute is to scale N and D proportionally, with D ≈ 20·N tokens. Chinchilla 70B (70B params, 1.4T tokens) outperformed Gopher 280B (280B params, 300B tokens) at the same compute budget.

Why B tempts people
Maximizing N at fixed C leaves D = C/(6N) very small — the model is under-trained. Gopher 280B is exactly this mistake: a very large model trained on too few tokens.
Why C tempts people
Training a larger model for fewer steps reduces D, violating the D ≈ 20·N rule. You get an over-parameterized, under-trained model.
Why D tempts people
Fixing N and scaling only D ignores that model capacity must grow with data to absorb the new information — you hit diminishing returns quickly.

48. Your turn: build TinyGPT

Section

Project

49. Project: TinyGPT from scratch

Concept

Implement a 2-layer decoder-only GPT (d=64, n_heads=4, vocab=50, max_T=8), train it on a synthetic token corpus with the CLM objective, then generate new tokens autoregressively.

#milestonekey tool
1CausalSelfAttention with tril mask; verify output shapenn.Linear, F.softmax
2Full TinyGPT forward pass; confirm params = 107,008nn.Embedding, GPTBlock
3CLM training loop 20 steps; confirm loss falls from ~3.91 to ~1.58Adam, F.cross_entropy
4Greedy generate 5 tokens from seed [5, 12, 3]argmax, cat

Build rules: causal mask is torch.tril(torch.ones(T,T)); positional embedding indexes torch.arange(T); input/target are data[:,:-1] / data[:,1:]; same Lesson 40 training loop throughout.

50. Break it if you can: Project: TinyGPT from scratch

Counterexample

Discussion prompt

Implement a 2-layer decoder-only GPT (d=64, n_heads=4, vocab=50, max_T=8), train it on a synthetic token corpus with the CLM objective, then generate new tokens autoregressively.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Build rules: causal mask is torch.tril(torch.ones(T,T)); positional embedding indexes torch.arange(T); input/target are data[:,:-1] / data[:,1:]; same Lesson 40 training loop throughout.

51. Milestone 1 — CausalSelfAttention, shape check

Worked example

Your turn: implement CausalSelfAttention with d=64, n_heads=4. Pass a random (2, 8, 64) tensor through it. Predict the output shape.

Hint: QKV projection is one nn.Linear(d, 3d); split into Q, K, V; reshape each to (B, n_heads, T, d_head); apply tril mask after scaling by 1/sqrt(d_head).

import torch, torch.nn as nn, torch.nn.functional as F, math
torch.manual_seed(42)

class CausalSelfAttention(nn.Module):
    def __init__(self, d, nh):
        super().__init__()
        self.nh, self.dh = nh, d//nh
        self.qkv  = nn.Linear(d, 3*d)
        self.proj = nn.Linear(d, d)
    def forward(self, x):
        B,T,C = x.shape
        Q,K,V = self.qkv(x).split(C, dim=-1)
        Q = Q.view(B,T,self.nh,self.dh).transpose(1,2)
        K = K.view(B,T,self.nh,self.dh).transpose(1,2)
        V = V.view(B,T,self.nh,self.dh).transpose(1,2)
        s = (Q @ K.transpose(-2,-1)) / math.sqrt(self.dh)
        m = torch.tril(torch.ones(T,T, device=x.device))
        s = s.masked_fill(m==0, float('-inf'))
        out = F.softmax(s,-1) @ V
        return self.proj(out.transpose(1,2).contiguous().view(B,T,C))

attn = CausalSelfAttention(64, 4)
x = torch.randn(2, 8, 64)
print(attn(x).shape)
input shapeoutput shaped_head
(2, 8, 64)(2, 8, 64)64/4 = 16
(1, 5, 64)(1, 5, 64)64/4 = 16
(4, 8, 64)(4, 8, 64)64/4 = 16

52. What stays fixed: Milestone 1 — CausalSelfAttention, shape check

Invariant

Step through it

Step through Milestone 1 — CausalSelfAttention, shape check one row at a time. One of these columns never changes — find it, and say why it cannot.

  1. Step 1: input shape is (2, 8, 64)
  2. Step 2: input shape is (1, 5, 64)
  3. Step 3: input shape is (4, 8, 64)

53. Milestone 2 — full TinyGPT, param count

Worked example

Your turn: assemble TinyGPT with tok_emb + pos_emb + 2×GPTBlock + LayerNorm + head. Count parameters and verify 107,008.

Hint: head uses nn.Linear(d, vocab_size, bias=False); positional embedding indexes torch.arange(T).unsqueeze(0); forward returns self.head(self.ln(self.blocks(x))).

class GPTBlock(nn.Module):
    def __init__(self,d,nh):
        super().__init__()
        self.ln1=nn.LayerNorm(d); self.attn=CausalSelfAttention(d,nh)
        self.ln2=nn.LayerNorm(d)
        self.ffn=nn.Sequential(nn.Linear(d,4*d),nn.GELU(),nn.Linear(4*d,d))
    def forward(self,x):
        x=x+self.attn(self.ln1(x)); return x+self.ffn(self.ln2(x))

class TinyGPT(nn.Module):
    def __init__(self,V,d,nh,L,T):
        super().__init__()
        self.tok=nn.Embedding(V,d); self.pos=nn.Embedding(T,d)
        self.blocks=nn.Sequential(*[GPTBlock(d,nh) for _ in range(L)])
        self.ln=nn.LayerNorm(d); self.head=nn.Linear(d,V,bias=False); self.T=T
    def forward(self,idx):
        B,t=idx.shape
        x=self.tok(idx)+self.pos(torch.arange(t,device=idx.device))
        return self.head(self.ln(self.blocks(x)))

torch.manual_seed(42)
m = TinyGPT(50,64,4,2,8)
print(sum(p.numel() for p in m.parameters()))
componentparams
tok_emb (50×64)3,200
pos_emb (8×64)512
2 × GPTBlock99,968
ln_f + head3,264
TOTAL107,008

54. What each one costs: Milestone 2 — full TinyGPT, param count

Trade off

Comparison matrix

From Milestone 2 — full TinyGPT, param count: every row here is a choice with a cost. Fill the params column, then say which row you would actually pick and what you give up for it.

componentparams
tok_emb (50×64)3,200
pos_emb (8×64)512
2 × GPTBlock99,968
ln_f + head3,264
TOTAL107,008

55. Milestone 3 & 4 — train then generate

Worked example

Your turn: run the CLM training loop for 20 steps and then greedily generate 5 tokens from seed [5, 12, 3]. Predict: will the initial loss be near log(50) ≈ 3.91?

Hint: inp = data[:,:-1]; tgt = data[:,1:]; loss = F.cross_entropy(logits.reshape(-1, V), tgt.reshape(-1)); generation appends argmax(logits[:,-1,:]) each step.

import torch.nn.functional as F
torch.manual_seed(42)
gpt_tr = TinyGPT(50,64,4,2,8)
opt = torch.optim.Adam(gpt_tr.parameters(), lr=1e-3)
data = torch.randint(0,50,(8,8))
inp=data[:,:-1]; tgt=data[:,1:]
for step in range(20):
    opt.zero_grad()
    loss=F.cross_entropy(gpt_tr(inp).reshape(-1,50),tgt.reshape(-1))
    loss.backward(); opt.step()
    if step in [0,4,9,19]: print(f'step {step+1}: {loss.item():.4f}')
# generate
torch.manual_seed(42)
g=TinyGPT(50,64,4,2,8); tokens=torch.tensor([[5,12,3]])
for _ in range(5):
    logits=g(tokens[:,-8:]); tokens=torch.cat([tokens,logits[:,-1:].argmax(-1,keepdim=True).squeeze(0).unsqueeze(0)],dim=1)
print('generated:', tokens[0].tolist())
steploss
14.0511
53.4078
102.7282
201.5836

56. Fill in: loss for Milestone 3 & 4 — train then generate

Comparison

Comparison matrix

From Milestone 3 & 4 — train then generate: refill the loss column from what you know. The rest of the table is as it appeared.

steploss
14.0511
53.4078
102.7282
201.5836

57. Show it off

Concept

Out loud, slides closed: explain (1) why GPT trains all positions in parallel despite generating left-to-right, (2) the two differences between a GPT block and a full encoder-decoder block, and (3) what Chinchilla says to do when you have a fixed compute budget.

Stretch (homework per lesson plan): implement a 6-layer, 256-dim GPT; train on a small text corpus (Shakespeare characters or Python tokens); demonstrate few-shot prompting by prepending 3 examples; walk through the autoregressive inference loop step by step. Next up: Lesson 90 — RLHF, instruction fine-tuning, and aligning GPT with human preferences.

58. Connect it up: Lesson 89: GPT & Decoder-Only Transformers

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — GPT = decoder-only transformer · GPT-2 architecture & scaling · Autoregressive inference & the CLM loop · Your turn: build TinyGPT. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

59. What you can do now

Recap

conceptthe one thing to remember
decoder-onlyno encoder, no cross-attention; causal mask only
parallel trainingteacher forcing + tril mask → all positions at once
GPT-2 blockpre-LN → causalAttn → pre-LN → FFN(4×) with residuals
ChinchillaD ≈ 20·N; more data beats more params at fixed compute
few-shotin-context tokens, not gradient updates — weights frozen

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 89 — GPT / Generative Pretrained Transformer — Barron · USAAIO Round 2 Preparation, 2026
  2. TinyGPT (2L, d=64, vocab=50) forward pass and CLM training loss verified; causal mask, multi-head split, parameter counts all real-executed — torch 2.7.1+cpu, numpy 2.2.6, June 2026
  3. Hoffman et al., 'Training Compute-Optimal Large Language Models' (Chinchilla, 2022) — DeepMind, arXiv:2203.15556

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108