Lesson 114: Phase 3 Mock Exam — Transformers, NLP, and Computer Vision

USAAIO Lesson 114, a Phase 3 review. It is a timed mock-exam simulation covering all the Phase 3 topics: deriving scaled dot-product attention, counting the parameters of multi-head attention, causal masking, sinusoidal positional encoding, the BERT architecture and how to fine-tune it, LoRA adapters, greedy against beam-search decoding, computing BLEU-4 and ROUGE-N, non-maximum suppression with IoU, and the cross-entropy and perplexity of a full transformer training run. Every value in a trace table was computed with torch 2.7.1+cpu, and each section comes with a worked answer-key walkthrough. The lesson runs to 27 slides.

Subject: Machine Learning · 53 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Phase 3 Mock Exam

Title

USAAIO · Lesson 114 · Phase 3 Capstone

Timed simulation: Section 1 (transformers), Section 2 (NLP), Section 3 (CV + coding). Full answer-key walkthrough. Target: >80% theory, one coding problem complete.

2. What this session tests

Objectives

  1. Derive scaled dot-product attention from scratch and trace numeric outputs (Lessons 88-90)
  2. Count MHA parameters from d_model and n_heads; apply causal masking (Lesson 89)
  3. Compute BLEU-4 and ROUGE-N by hand on a toy sentence pair (Lesson 101)
  4. Explain BERT fine-tuning vs LoRA adapters and give parameter counts (Lesson 98)
  5. Implement multi-head attention in PyTorch and verify output shapes (Lesson 89)
  6. Apply NMS with IoU threshold to suppress redundant boxes (Lesson 107)

3. What survived from CLIP — Contrastive Language-Image Pretraining?

Warm-up

Discussion prompt

Before we open Lesson 114: Phase 3 Mock Exam — Transformers, NLP, and Computer Vision: without looking back, what was the main idea of CLIP — Contrastive Language-Image Pretraining, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

symmetric InfoNCE contrastive loss, dual-encoder CLIP architecture (image + text tower), learned temperature parameter, zero-shot transfer via text-based class descriptors, and CLIP extensions (ALIGN, BLIP-2).

4. Section 1 — Transformer Theory

Section

60 minutes

5. Attention derivation — the four steps

Concept

Every exam question on attention reduces to four operations. Write them from memory before reading the formula.

\[ \text{Attention}(Q,K,V) = \text{softmax}\!\left(\frac{QK^\top}{\sqrt{d_k}}\right)V \]

stepoperationshape out (seq=T, d_k)
1raw scores S = QK^T(T, T)
2scale S / sqrt(d_k)(T, T)
3softmax row-wise(T, T) rows sum to 1
4weighted sum A · V(T, d_v)

6. Fill in: operation for Attention derivation — the four steps

Comparison

Comparison matrix

From Attention derivation — the four steps: refill the operation column from what you know. The rest of the table is as it appeared.

stepoperationshape out (seq=T, d_k)
1raw scores S = QK^T(T, T)
2scale S / sqrt(d_k)(T, T)
3softmax row-wise(T, T) rows sum to 1
4weighted sum A · V(T, d_v)

7. What has to happen first: Attention trace — Q, K, V ∈ R^{3×4}, d_k = 4

Ranking

Put in order

Put the moves of Attention trace — Q, K, V ∈ R^{3×4}, d_k = 4 into the order they have to happen.

  1. Compute raw score matrix S = QK^T
  2. Scale by 1/√d_k = 1/2 → S_scaled[0] = [1.000, 0.000, 0.500]
  3. Softmax row 0: exp([1.0, 0.0, 0.5]) / sum → [0.5065, 0.1863, 0.3072]
  4. Output row 0 = 0.5065·V[0] + 0.1863·V[1] + 0.3072·V[2] = [4.2029, 5.2029, 6.2029, 7.2029]

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Dot product of each query row with every key row.

8. Attention trace — Q, K, V ∈ R^{3×4}, d_k = 4

Worked example

Compute raw score matrix S = QK^T

Why: Dot product of each query row with every key row. With Q[0]=[1,0,1,0] and K[0]=[1,0,1,0]: Q[0]·K[0]=2, Q[0]·K[1]=0, Q[0]·K[2]=1 → row 0 = [2, 0, 1].

Scale by 1/√d_k = 1/2 → S_scaled[0] = [1.000, 0.000, 0.500]

Why: Without scaling, large d_k pushes dot products into saturated softmax regions, collapsing gradients. Dividing by sqrt(d_k) stabilises variance.

Softmax row 0: exp([1.0, 0.0, 0.5]) / sum → [0.5065, 0.1863, 0.3072]

Why: exp(1.0)=2.718, exp(0.0)=1.000, exp(0.5)=1.649; sum=5.366; divide each. These are the attention weights for query 0 over the three keys.

Output row 0 = 0.5065·V[0] + 0.1863·V[1] + 0.3072·V[2] = [4.2029, 5.2029, 6.2029, 7.2029]

Why: V rows are [1,2,3,4], [5,6,7,8], [9,10,11,12]. Weighted sum: 0.5065·1 + 0.1863·5 + 0.3072·9 = 4.2029; pattern holds for all 4 columns.

rowattn weightsoutput (first 2 cols)
Q[0][0.5065, 0.1863, 0.3072][4.2029, 5.2029]
Q[1][0.1863, 0.5065, 0.3072][5.4835, 6.4835]
Q[2][0.3837, 0.3837, 0.2327][4.3962, 5.3962]

9. What each one costs: Attention trace — Q, K, V ∈ R^{3×4}, d_k = 4

Trade off

Comparison matrix

From Attention trace — Q, K, V ∈ R^{3×4}, d_k = 4: every row here is a choice with a cost. Fill the attn weights column, then say which row you would actually pick and what you give up for it.

rowattn weightsoutput (first 2 cols)
Q[0][0.5065, 0.1863, 0.3072][4.2029, 5.2029]
Q[1][0.1863, 0.5065, 0.3072][5.4835, 6.4835]
Q[2][0.3837, 0.3837, 0.2327][4.3962, 5.3962]

10. Causal masking in decoder self-attention

Concept

A causal (autoregressive) mask sets future positions to -inf before softmax, forcing each position to attend only to itself and earlier tokens. Applied in GPT-style decoders; absent in BERT encoders.

import torch, math, torch.nn.functional as F

# seq=3, d_k=4 (same Q,K,V from worked example)
scores_scaled = torch.tensor(
    [[1.0000, 0.0000, 0.5000],
     [0.0000, 1.0000, 0.5000],
     [0.5000, 0.5000, 0.0000]])

causal_mask = torch.triu(torch.ones(3, 3), diagonal=1).bool()
masked = scores_scaled.masked_fill(causal_mask, float('-inf'))
weights = F.softmax(masked, dim=-1)
print(masked)
print(weights)
positionmasked scores rowattn weights (causal)
0[1.000, -inf, -inf][1.000, 0.000, 0.000]
1[0.000, 1.000, -inf][0.269, 0.731, 0.000]
2[0.500, 0.500, 0.000][0.384, 0.384, 0.233]

11. MHA parameter count — from d_model and n_heads

Concept

Multi-head attention (Lesson 89) splits d_model into n_heads heads of size d_k = d_model / n_heads. Four weight matrices: W_Q, W_K, W_V (each d_model × d_k per head) and W_O (d_model × d_model).

\[ \text{params} = n_h \cdot 3 d_{\text{model}} d_k + d_{\text{model}}^2 = 4 d_{\text{model}}^2 \quad\text{(no bias)} \]

d_modeln_headsd_k3×d×dk×hW_O (d×d)total
644163×64×16×4 = 12,28864×64 = 4,09616,384
76812643×768×64×12 = 1,769,472768×768 = 589,8242,359,296
102416643×1024×64×16 = 3,145,7281024×1024 = 1,048,5764,194,304

Verified: nn.MultiheadAttention(embed_dim=64, num_heads=4, bias=False) reports 16,384 parameters — matches the formula exactly.

12. Break it if you can: MHA parameter count — from d_model and n_heads

Counterexample

Discussion prompt

Verified: nn.MultiheadAttention(embed_dim=64, num_heads=4, bias=False) reports 16,384 parameters — matches the formula exactly.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

13. Something is wrong here: forgetting W_O when counting MHA params

Anomaly

Predict first

A student writes this, and it looks reasonable:

Exam question: total MHA params for d_model=64, n_heads=4, bias=False?

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Stops after computing W_Q, W_K, W_V for all heads.

Exam question: total MHA params for d_model=64, n_heads=4, bias=False?

Why: Stops after computing W_Q, W_K, W_V for all heads. Misses the output projection W_O that recombines head outputs back to d_model.

14. Trap: forgetting W_O when counting MHA params

Trap

The trap

Exam question: total MHA params for d_model=64, n_heads=4, bias=False?

Count only the QKV projections: 4 heads × 3 × 64 × 16 = 12,288

Why: Stops after computing W_Q, W_K, W_V for all heads. Misses the output projection W_O that recombines head outputs back to d_model.

Answer: 12,288 — WRONG

Why: Off by 4,096 (the W_O matrix). In PyTorch nn.MultiheadAttention stores W_O as out_proj, which has d_model×d_model weights.

The fix

Exam question: total MHA params for d_model=64, n_heads=4, bias=False?

QKV projections: 4 heads × 3 × (64 × 16) = 12,288; W_O: 64 × 64 = 4,096

Why: W_O is a d_model × d_model matrix that merges the concatenated head outputs back into d_model dimensions — always present, always d^2.

Answer: 12,288 + 4,096 = 16,384 — CORRECT

Why: Matches sum(p.numel() for p in nn.MultiheadAttention(64, 4, bias=False).parameters()) = 16384.

15. Break it on purpose: forgetting W_O when counting MHA params

Break the constraint

Discussion prompt

The rule this trap just fixed:

W_O is a d_model × d_model matrix that merges the concatenated head outputs back into d_model dimensions — always present, always d^2.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

Stops after computing W_Q, W_K, W_V for all heads. Misses the output projection W_O that recombines head outputs back to d_model.

16. Section 2 — NLP: BERT, Generation, Metrics

Section

60 minutes

17. Sinusoidal positional encoding — what the values look like

Concept

\[ PE_{(pos,\,2i)} = \sin\!\left(\frac{pos}{10000^{2i/d}}\right), \quad PE_{(pos,\,2i+1)} = \cos\!\left(\frac{pos}{10000^{2i/d}}\right) \]

Key exam facts: PE(0) is all-zeros for sin dims and all-ones for cos dims. High-frequency dimensions (small i) change rapidly with position; low-frequency dims (large i) barely change. These are added to the token embedding, not concatenated.

posdim 0 (sin)dim 1 (cos)dim 2 (sin)dim 3 (cos)
00.00001.00000.00001.0000
10.84150.54030.09980.9950
20.9093-0.41610.19870.9801
30.1411-0.99000.29550.9553
4-0.7568-0.65360.38940.9211

18. Watch it run: Sinusoidal positional encoding — what the values look…

Pattern

Step through it

Step through Sinusoidal positional encoding — what the values look like one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: pos is 0
  2. Step 2: pos is 1
  3. Step 3: pos is 2
  4. Step 4: pos is 3
  5. Step 5: pos is 4

19. BERT fine-tuning vs LoRA — parameter efficiency

Concept

Full fine-tuning updates all 110M BERT-base parameters. Classification fine-tuning adds only a d_model × n_classes head and freezes the encoder. LoRA (Lesson 98) instead injects low-rank adapters into the frozen W_Q and W_V matrices.

\[ \Delta W = BA, \quad B \in \mathbb{R}^{d \times r},\; A \in \mathbb{R}^{r \times d}, \quad r \ll d \]

strategytrainable params% of BERT-base
Full fine-tune110,000,000100.0%
Head-only (2-class)1,5380.0014%
LoRA r=8 (Q+V, 12 layers)294,9120.27%

LoRA per adapter: B is (768 × 8) = 6,144 params, A is (8 × 768) = 6,144 params. Two adapters (Q and V) per layer × 12 layers × 2 matrices × 6,144 = 294,912 trainable params.

20. By analogy: BERT fine-tuning vs LoRA — parameter efficiency

Analogy

Discussion prompt

Explain BERT fine-tuning vs LoRA — parameter efficiency by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

LoRA per adapter: B is (768 × 8) = 6,144 params, A is (8 × 768) = 6,144 params. Two adapters (Q and V) per layer × 12 layers × 2 matrices × 6,144 = 294,912 trainable params.

21. Greedy decoding vs top-k sampling — trace

Concept

Greedy decoding picks argmax of the softmax at each step — deterministic, fast, but repetitive. Top-k sampling restricts to the k highest-prob tokens then resamples — stochastic, higher diversity. Beam search keeps the top-b partial sequences at each step.

import torch, torch.nn.functional as F

vocab = ['<end>', 'A', 'B', 'C']
logits_t1 = torch.tensor([0.1, 2.0, 1.5, 0.4])

# Greedy: argmax
probs_t1 = F.softmax(logits_t1, dim=0)
greedy_tok = vocab[int(probs_t1.argmax())]

# Top-k (k=2): renormalise over top-2
topk_vals, topk_idx = torch.topk(probs_t1, k=2)
topk_probs = topk_vals / topk_vals.sum()

print(f'probs: {[f"{p:.3f}" for p in probs_t1.tolist()]}')
print(f'greedy: {greedy_tok}')
print(f'top-2 probs: {topk_probs.tolist()}')
print(f'top-2 tokens: {[vocab[i] for i in topk_idx.tolist()]}')
steplogitssoftmax probsgreedy pick
1[0.1, 2.0, 1.5, 0.4][0.076, 0.511, 0.310, 0.103]A
2[0.2, 0.5, 3.0, 0.3][0.050, 0.068, 0.826, 0.056]B
3[5.0, 0.1, 0.2, 0.1][0.977, 0.007, 0.008, 0.007]<end>

22. What has to happen first: BLEU-4 by hand — ref vs hypothesis

Ranking

Put in order

Put the moves of BLEU-4 by hand — ref vs hypothesis into the order they have to happen.

  1. Set up: ref = 'the cat sat on the mat', hyp = 'the cat sat on a mat'
  2. Compute clipped n-gram precisions: P_1=5/6, P_2=3/5, P_3=2/4, P_4=1/3
  3. Brevity penalty: BP = 1 since |hyp|=|ref|=6
  4. BLEU-4 = BP · exp((1/4) · Σ log P_n) = exp((log(5/6)+log(3/5)+log(2/4)+log(1/3))/4) = 0.5373

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. One word differs: 'a' replaces 'the'.

23. BLEU-4 by hand — ref vs hypothesis

Worked example

Set up: ref = 'the cat sat on the mat', hyp = 'the cat sat on a mat'

Why: One word differs: 'a' replaces 'the'. BLEU clips counts to the reference count to prevent gaming by repetition.

Compute clipped n-gram precisions: P_1=5/6, P_2=3/5, P_3=2/4, P_4=1/3

Why: 1-grams: 5 of 6 hyp unigrams appear in ref (all but 'a'). 2-grams: 3 of 5 match. 3-grams: 2 of 4 match. 4-grams: 1 of 3 match ('the cat sat on' is not in hyp but 'cat sat on a' isn't in ref; 'the cat sat on' is not a 4-gram of hyp — only 'cat sat on a' and 'sat on a mat' and 'the cat sat on'. ref has 'the cat sat on' not 'cat sat on a'. Verify manually: hyp 4-grams = {'the cat sat on','cat sat on a','sat on a mat'}; ref 4-grams = {'the cat sat on','cat sat on the','sat on the mat','on the mat '}. Match: 1. Correct.)

Brevity penalty: BP = 1 since |hyp|=|ref|=6

Why: BP = 1 when the hypothesis is not shorter than the reference. BP = exp(1 - r/c) when c < r. Here c = r = 6 so BP = 1.

BLEU-4 = BP · exp((1/4) · Σ log P_n) = exp((log(5/6)+log(3/5)+log(2/4)+log(1/3))/4) = 0.5373

Why: log(0.8333)+log(0.6000)+log(0.5000)+log(0.3333) = -0.1823-0.5108-0.6931-1.0986 = -2.4848; /4 = -0.6212; exp(-0.6212) = 0.5373.

nmatcheshyp n-gramsP_n
1560.8333
2350.6000
3240.5000
4130.3333
BLEU-4BP=1.000exp(mean log)0.5373

24. Fill in: hyp n-grams for BLEU-4 by hand — ref vs hypothesis

Comparison

Comparison matrix

From BLEU-4 by hand — ref vs hypothesis: refill the hyp n-grams column from what you know. The rest of the table is as it appeared.

nmatcheshyp n-gramsP_n
1560.8333
2350.6000
3240.5000
4130.3333
BLEU-4BP=1.000exp(mean log)0.5373

25. ROUGE-N — recall-oriented complement to BLEU

Concept

BLEU is precision-oriented (how much of the hypothesis is in the reference). ROUGE is recall-oriented (how much of the reference is covered by the hypothesis). Both compute n-gram overlap; ROUGE is the dominant metric for summarisation.

\[ \text{ROUGE-N Recall} = \frac{\sum_{\text{ngrams}\in\text{ref}} \min(\text{count}_{\text{hyp}}, \text{count}_{\text{ref}})}{\sum_{\text{ngrams}\in\text{ref}} \text{count}_{\text{ref}}} \]

metricprecisionrecallF1
ROUGE-10.83330.83330.8333
ROUGE-20.60000.60000.6000

Same ref/hyp pair as BLEU. P=R=F1 here because |hyp|=|ref|. In general they differ: a verbose hypothesis has high recall but low precision; a terse one has high precision but low recall.

26. By analogy: ROUGE-N — recall-oriented complement to BLEU

Analogy

Discussion prompt

Explain ROUGE-N — recall-oriented complement to BLEU by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Same ref/hyp pair as BLEU. P=R=F1 here because |hyp|=|ref|. In general they differ: a verbose hypothesis has high recall but low precision; a terse one has high precision but low recall.

27. Section 3 — Coding & Computer Vision

Section

60 minutes

28. NMS — non-maximum suppression with IoU

Concept

Object detectors produce many overlapping candidate boxes. NMS keeps the highest-score box, suppresses all others with IoU above a threshold, then repeats on the survivors. IoU = intersection / union.

\[ \text{IoU}(A, B) = \frac{|A \cap B|}{|A \cup B|} = \frac{\text{inter\_area}}{\text{area}_A + \text{area}_B - \text{inter\_area}} \]

box pairintersection areaunion areaIoUNMS action (thresh=0.5)
box0 vs box1 ([10,10,60,60] vs [12,12,62,62])48×48=230450×50+50×50-2304=26960.8546SUPPRESS box1
box0 vs box2 ([10,10,60,60] vs [80,80,120,120])02500+1600-0=41000.0000keep box2
box1 vs box2 (already suppressed)0—0.0000—

NMS result: boxes 0 (score 0.9) and 2 (score 0.7) kept; box 1 suppressed because IoU(box0, box1) = 0.8546 > 0.50.

29. Break it if you can: NMS — non-maximum suppression with IoU

Counterexample

Discussion prompt

NMS result: boxes 0 (score 0.9) and 2 (score 0.7) kept; box 1 suppressed because IoU(box0, box1) = 0.8546 > 0.50.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

30. What has to happen first: Implement multi-head attention from scratch

Ranking

Put in order

Put the moves of Implement multi-head attention from scratch into the order they have to happen.

  1. Project X into Q, K, V using linear layers, then reshape to (B, n_heads, T, d_k)
  2. Compute scaled dot-product attention per head: attn = softmax(Q @ K^T / sqrt(d_k)) @ V
  3. Merge heads: transpose back to (B, T, heads, d_k), reshape to (B, T, d_model), apply W_O

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Each of the n_heads heads attends to a d_k-dimensional subspace.

31. Implement multi-head attention from scratch

Worked example

Project X into Q, K, V using linear layers, then reshape to (B, n_heads, T, d_k)

Why: Each of the n_heads heads attends to a d_k-dimensional subspace. Reshape from (B, T, d_model) → (B, T, n_heads, d_k) then transpose dims 1 and 2.

Compute scaled dot-product attention per head: attn = softmax(Q @ K^T / sqrt(d_k)) @ V

Why: Broadcasting: Q and K are (B, heads, T, d_k); Q @ K.transpose(-2,-1) gives (B, heads, T, T). Softmax on dim=-1 (over keys). Result (B, heads, T, d_k).

Merge heads: transpose back to (B, T, heads, d_k), reshape to (B, T, d_model), apply W_O

Why: .transpose(1,2).contiguous().view(B, T, C) — contiguous() is required before view because transpose creates a non-contiguous tensor.

import torch, torch.nn as nn, torch.nn.functional as F, math

class TinyMHA(nn.Module):
    def __init__(self, d_model=16, n_heads=2):
        super().__init__()
        self.d_k = d_model // n_heads
        self.n_heads = n_heads
        self.Wq = nn.Linear(d_model, d_model, bias=False)
        self.Wk = nn.Linear(d_model, d_model, bias=False)
        self.Wv = nn.Linear(d_model, d_model, bias=False)
        self.Wo = nn.Linear(d_model, d_model, bias=False)

    def forward(self, x):
        B, T, C = x.shape
        # Project + split heads
        Q = self.Wq(x).view(B, T, self.n_heads, self.d_k).transpose(1, 2)
        K = self.Wk(x).view(B, T, self.n_heads, self.d_k).transpose(1, 2)
        V = self.Wv(x).view(B, T, self.n_heads, self.d_k).transpose(1, 2)
        # Attention
        attn = (Q @ K.transpose(-2, -1)) / math.sqrt(self.d_k)
        attn = F.softmax(attn, dim=-1)
        out = (attn @ V).transpose(1, 2).contiguous().view(B, T, C)
        return self.Wo(out)

torch.manual_seed(42)
model = TinyMHA(d_model=16, n_heads=2)
x = torch.randn(2, 5, 16)
print(model(x).shape)     # (2, 5, 16)
params = sum(p.numel() for p in model.parameters())
print(f'params: {params}')  # 4 * 16^2 / bias=False = 1024
tensorshape after stepnote
x input(2, 5, 16)B=2, T=5, d=16
Q (after Linear + view + transpose)(2, 2, 5, 8)B, heads=2, T, d_k=8
attn weights(2, 2, 5, 5)B, heads, T, T
attn @ V(2, 2, 5, 8)B, heads, T, d_k
after merge + Wo(2, 5, 16)back to input shape
param count4 × 16²= 1024 (bias=False)

32. Inspect it line by line: Implement multi-head attention from scratch

Error analysis

Annotate

Walk the callouts on Implement multi-head attention from scratch. Each one is a place this is easy to get subtly wrong.

  • Each of the n_heads heads attends to a d_k-dimensional subspace. Reshape from (B, T, d_model) → (B, T, n_heads, d_k) then transpose dims 1 and 2.
  • Broadcasting: Q and K are (B, heads, T, d_k); Q @ K.transpose(-2,-1) gives (B, heads, T, T). Softmax on dim=-1 (over keys). Result (B, heads, T, d_k).
  • .transpose(1,2).contiguous().view(B, T, C) — contiguous() is required before view because transpose creates a non-contiguous tensor.

33. Cross-entropy loss and perplexity in language modelling

Concept

A language model's training loss is CrossEntropyLoss over the vocabulary at each position. Perplexity = exp(CE loss) — the geometric mean number of equally-likely tokens the model is considering. Lower is better; perfect model has PP=1.

\[ \mathcal{L}_{\text{CE}} = -\frac{1}{N}\sum_i \log p_{\theta}(y_i \mid x_i), \qquad \text{PP} = e^{\mathcal{L}_{\text{CE}}} \]

import torch, torch.nn as nn, math

# batch=2, seq=4, vocab=3
logits = torch.tensor([
    [[2.0, 0.5, 0.1], [0.1, 3.0, 0.2], [0.3, 0.2, 2.8], [1.0, 0.5, 0.1]],
    [[0.5, 2.2, 0.3], [0.2, 0.1, 3.5], [2.0, 0.5, 0.2], [0.5, 1.0, 0.5]],
])
targets = torch.tensor([[0, 1, 2, 0], [1, 2, 0, 1]])

loss = nn.CrossEntropyLoss()(logits.view(-1, 3), targets.view(-1))
print(f'CE loss: {loss.item():.4f}')    # 0.3436
print(f'Perplexity: {math.exp(loss.item()):.4f}')  # 1.4100
CE lossperplexityinterpretation
0.34361.4100model is almost always right (3-class)
1.09863.0000uniform over 3 tokens — no signal
0.01.0perfect model — 100% prob on correct token

34. Where does each piece belong: Lesson 114: Phase 3 Mock Exam —…

Sorting

Sort into buckets

These are the pieces of Lesson 114: Phase 3 Mock Exam — Transformers, NLP, and Computer Vision, out of order. Put each one back under the part of the lesson it belongs to.

Section 1 — Transformer Theory
Attention derivation — the four steps; Attention trace — Q, K, V ∈ R^{3×4}, d_k = 4; Causal masking in decoder self-attention
Section 2 — NLP: BERT, Generation, Metrics
Sinusoidal positional encoding — what the values look like; BERT fine-tuning vs LoRA — parameter efficiency; Greedy decoding vs top-k sampling — trace
Section 3 — Coding & Computer Vision
NMS — non-maximum suppression with IoU; Implement multi-head attention from scratch; Cross-entropy loss and perplexity in language modelling
s1
Section 1 — Transformer Theory is where Lesson 114: Phase 3 Mock Exam — Transformers, NLP, and Computer Vision puts Attention derivation — the four steps, Attention trace — Q, K, V ∈ R^{3×4}, d_k = 4, Causal masking in decoder self-attention. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
Section 2 — NLP: BERT, Generation, Metrics is where Lesson 114: Phase 3 Mock Exam — Transformers, NLP, and Computer Vision puts Sinusoidal positional encoding — what the values look like, BERT fine-tuning vs LoRA — parameter efficiency, Greedy decoding vs top-k sampling — trace. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
Section 3 — Coding & Computer Vision is where Lesson 114: Phase 3 Mock Exam — Transformers, NLP, and Computer Vision puts NMS — non-maximum suppression with IoU, Implement multi-head attention from scratch, Cross-entropy loss and perplexity in language modelling. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

35. Without one step: Phase 3 exam strategy — 6-step recipe

Constraint

Discussion prompt

Run Phase 3 exam strategy — 6-step recipe with this step confiscated:

ROUGE: compute n-gram overlap; recall = overlap / ref_ngrams; precision = overlap / hyp_ngrams; F1 = harmonic mean.

Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.

Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.

Answer:

  1. Attention derivation: write S=QK^T, scale by 1/√d_k, softmax row-wise, multiply V. For causal mask: triu(..., diagonal=1) filled with -inf before softmax.
  2. MHA param count: 4×d_model² (no bias). Add bias counts if specified: 3×d_model×d_k×heads + d_model² + 4×d_model bias terms.
  3. BLEU-4: clipped n-gram precision for n=1..4, geometric mean, multiply by brevity penalty exp(1 - r/c) when c < r.
  4. ROUGE: compute n-gram overlap; recall = overlap / ref_ngrams; precision = overlap / hyp_ngrams; F1 = harmonic mean.
  5. NMS: sort by score descending; greedily keep top box; suppress all remaining boxes with IoU > threshold; repeat.
  6. LoRA param count: 2×(d×r + r×d) per adapter × layers × (Q+V). Always much less than d².

36. Phase 3 exam strategy — 6-step recipe

Pattern

  1. Attention derivation: write S=QK^T, scale by 1/√d_k, softmax row-wise, multiply V. For causal mask: triu(..., diagonal=1) filled with -inf before softmax.
  2. MHA param count: 4×d_model² (no bias). Add bias counts if specified: 3×d_model×d_k×heads + d_model² + 4×d_model bias terms.
  3. BLEU-4: clipped n-gram precision for n=1..4, geometric mean, multiply by brevity penalty exp(1 - r/c) when c < r.
  4. ROUGE: compute n-gram overlap; recall = overlap / ref_ngrams; precision = overlap / hyp_ngrams; F1 = harmonic mean.
  5. NMS: sort by score descending; greedily keep top box; suppress all remaining boxes with IoU > threshold; repeat.
  6. LoRA param count: 2×(d×r + r×d) per adapter × layers × (Q+V). Always much less than d².

37. Where does it stop working: Phase 3 exam strategy — 6-step recipe

Edge cases

Discussion prompt

Phase 3 exam strategy — 6-step recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Attention derivation: write S=QK^T, scale by 1/√d_k, softmax row-wise, multiply V. For causal mask: triu(..., diagonal=1) filled with -inf before softmax.
  2. MHA param count: 4×d_model² (no bias). Add bias counts if specified: 3×d_model×d_k×heads + d_model² + 4×d_model bias terms.
  3. BLEU-4: clipped n-gram precision for n=1..4, geometric mean, multiply by brevity penalty exp(1 - r/c) when c < r.
  4. ROUGE: compute n-gram overlap; recall = overlap / ref_ngrams; precision = overlap / hyp_ngrams; F1 = harmonic mean.
  5. NMS: sort by score descending; greedily keep top box; suppress all remaining boxes with IoU > threshold; repeat.
  6. LoRA param count: 2×(d×r + r×d) per adapter × layers × (Q+V). Always much less than d².

38. Rule out three: Check 1 — attention output shape

Elimination

Eliminate the wrong options

A single-head attention module has Q ∈ R^{8×64}, K ∈ R^{6×64}, V ∈ R^{6×32}. What is the shape of the attention output?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. (8, 32)
  • B. (8, 64)
  • C. (6, 32)
  • D. (8, 6)

Survives elimination: A

Why: QK^T gives (8,64)×(64,6) = (8,6) score matrix. After softmax (8,6). Multiply by V ∈ (6,32): (8,6)×(6,32) = (8,32). Output shape matches V's last dimension, rows match Q's sequence length.

39. Check 1 — attention output shape

Check

Solve on paper before selecting.

Check your understanding

A single-head attention module has Q ∈ R^{8×64}, K ∈ R^{6×64}, V ∈ R^{6×32}. What is the shape of the attention output?

  • A. (8, 32) (correct)
  • B. (8, 64)
  • C. (6, 32)
  • D. (8, 6)

Answer: A

Why: QK^T gives (8,64)×(64,6) = (8,6) score matrix. After softmax (8,6). Multiply by V ∈ (6,32): (8,6)×(6,32) = (8,32). Output shape matches V's last dimension, rows match Q's sequence length.

Why B tempts people
Confuses output with the query key-dimension d_k=64. The output dimension comes from d_v (V's last dim=32), not d_k.
Why C tempts people
Uses K/V sequence length (6) instead of Q sequence length (8). Output sequence length always matches Q's row count.
Why D tempts people
Returns the score matrix shape (8,6) before the weighted sum with V. This is an intermediate step, not the final output.

40. Answer it before you see the options: Check 2 — BLEU brevity penalty

Prediction

Predict first

Reference length r=10, hypothesis length c=7. What is the BLEU brevity penalty?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: exp(1 - 10/7) ≈ 0.654

Why: BP = exp(1 - r/c) when c < r. Here c=7 < r=10 so BP = exp(1 - 10/7) = exp(-0.4286) ≈ 0.6514. This penalises short hypotheses — a one-word hypothesis with a perfect 1-gram match would still get a near-zero BLEU.

41. Check 2 — BLEU brevity penalty

Check

Work out BP before selecting.

Check your understanding

Reference length r=10, hypothesis length c=7. What is the BLEU brevity penalty?

  • A. exp(1 - 10/7) ≈ 0.654 (correct)
  • B. 1.000
  • C. 7/10 = 0.700
  • D. exp(1 - 7/10) ≈ 1.350

Answer: A

Why: BP = exp(1 - r/c) when c < r. Here c=7 < r=10 so BP = exp(1 - 10/7) = exp(-0.4286) ≈ 0.6514. This penalises short hypotheses — a one-word hypothesis with a perfect 1-gram match would still get a near-zero BLEU.

Why B tempts people
BP=1 only when c ≥ r. Here c=7 < r=10, so the hypothesis is too short and must be penalised.
Why C tempts people
Simple ratio c/r is not the formula. The exponential form ensures BP is always ≤ 1 and approaches 0 as the hypothesis shrinks.
Why D tempts people
Inverted the ratio: used exp(1 - c/r) = exp(0.3) ≈ 1.35, which is > 1. BP is capped at 1 and can never exceed it.

42. Rule out three: Check 3 — LoRA parameter count

Elimination

Eliminate the wrong options

A 12-layer transformer has d_model=768. LoRA adapters (rank r=8, no bias) are inserted into W_Q and W_V of every layer. How many trainable parameters does LoRA add?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. 294,912
  • B. 147,456
  • C. 589,824
  • D. 12,288

Survives elimination: A

Why: Each LoRA adapter = A (r×d) + B (d×r) = 2×768×8 = 12,288 params. Two adapters per layer (Q and V): 24,576. Times 12 layers: 294,912. Verified with the formula 12 × 2 × 2 × 768 × 8.

43. Check 3 — LoRA parameter count

Check

Compute the answer on paper first.

Check your understanding

A 12-layer transformer has d_model=768. LoRA adapters (rank r=8, no bias) are inserted into W_Q and W_V of every layer. How many trainable parameters does LoRA add?

  • A. 294,912 (correct)
  • B. 147,456
  • C. 589,824
  • D. 12,288

Answer: A

Why: Each LoRA adapter = A (r×d) + B (d×r) = 2×768×8 = 12,288 params. Two adapters per layer (Q and V): 24,576. Times 12 layers: 294,912. Verified with the formula 12 × 2 × 2 × 768 × 8.

Why B tempts people
Counts only Q adapters (12 layers × 1 adapter × 2 matrices × 768×8 = 147,456). Misses the V adapters, which are identical.
Why C tempts people
Doubles the correct answer — perhaps by counting A and B as separate matrices twice each, or mistakenly adding a third adapter per layer.
Why D tempts people
Computes only a single adapter's parameters (2×768×8=12,288) without multiplying by the 12 layers or 2 adapter locations per layer.

44. Your Turn — Coding Problem

Section

Timed coding challenge

45. Coding challenge brief

Concept

Goal: Implement scaled dot-product attention from scratch in PyTorch (no nn.MultiheadAttention), add optional causal masking, and verify output shapes and parameter counts.

  1. Milestone 1: Write scaled_dot_product_attention(Q, K, V, causal=False) that returns the attended output and the attention weight matrix.
  2. Milestone 2: Verify on Q=(3,4), K=(3,4), V=(3,4) inputs that output shape is (3,4) and weights sum to 1 per row.
  3. Milestone 3: Add causal masking — pass causal=True and confirm position 0 attends only to itself (weight row 0 = [1.0, 0.0, 0.0]).
  4. Show-it-off: Wrap in a single-head SelfAttention(nn.Module) with learnable W_Q, W_K, W_V, W_O.

46. Predict the next row: Coding challenge — solution walkthrough

Pattern

Predict first

The table runs: output shape | (3, 4) | (3, 4) · weight row sums | [1.000, 1.000, 1.000] | [1.0000, 1.0000, 1.0000] · causal row 0 | [1.0, 0.0, 0.0] | [1.0000, 0.0000, 0.0000] · causal row 1 | [0.269, 0.731, 0.0] | [0.2689, 0.7311, 0.0000]

In Coding challenge — solution walkthrough, given the rows so far: what is the next one — the row where check is output[0,0]?

Correct: output[0,0] | 4.2029 | 4.2029

checkexpectedactual
output shape(3, 4)(3, 4)
weight row sums[1.000, 1.000, 1.000][1.0000, 1.0000, 1.0000]
causal row 0[1.0, 0.0, 0.0][1.0000, 0.0000, 0.0000]
causal row 1[0.269, 0.731, 0.0][0.2689, 0.7311, 0.0000]
output[0,0]4.20294.2029

Why: The relationship between the columns, not the individual numbers, is what generates the next row. Upper-triangular mask (diagonal=1) zeros out future positions before softmax.

47. Coding challenge — solution walkthrough

Worked example

Implement the attention function with optional causal mask

Why: Upper-triangular mask (diagonal=1) zeros out future positions before softmax. masked_fill(-inf) ensures those positions become 0 after softmax (since exp(-inf)=0).

import torch, torch.nn as nn, torch.nn.functional as F, math

def scaled_dot_product_attention(Q, K, V, causal=False):
    d_k = Q.shape[-1]
    scores = Q @ K.transpose(-2, -1) / math.sqrt(d_k)
    if causal:
        T = scores.shape[-1]
        mask = torch.triu(torch.ones(T, T), diagonal=1).bool()
        scores = scores.masked_fill(mask, float('-inf'))
    weights = F.softmax(scores, dim=-1)
    return weights @ V, weights

torch.manual_seed(42)
Q = torch.tensor([[1.,0.,1.,0.],[0.,1.,0.,1.],[1.,1.,0.,0.]])
K = torch.tensor([[1.,0.,1.,0.],[0.,1.,0.,1.],[0.,0.,1.,1.]])
V = torch.tensor([[1.,2.,3.,4.],[5.,6.,7.,8.],[9.,10.,11.,12.]])

out, w = scaled_dot_product_attention(Q, K, V, causal=False)
out_c, w_c = scaled_dot_product_attention(Q, K, V, causal=True)
print('output shape:', out.shape)
print('weights row sums:', w.sum(dim=-1))
print('causal row 0:', w_c[0].tolist())
checkexpectedactual
output shape(3, 4)(3, 4)
weight row sums[1.000, 1.000, 1.000][1.0000, 1.0000, 1.0000]
causal row 0[1.0, 0.0, 0.0][1.0000, 0.0000, 0.0000]
causal row 1[0.269, 0.731, 0.0][0.2689, 0.7311, 0.0000]
output[0,0]4.20294.2029

48. What each one costs: Coding challenge — solution walkthrough

Trade off

Comparison matrix

From Coding challenge — solution walkthrough: every row here is a choice with a cost. Fill the actual column, then say which row you would actually pick and what you give up for it.

checkexpectedactual
output shape(3, 4)(3, 4)
weight row sums[1.000, 1.000, 1.000][1.0000, 1.0000, 1.0000]
causal row 0[1.0, 0.0, 0.0][1.0000, 0.0000, 0.0000]
causal row 1[0.269, 0.731, 0.0][0.2689, 0.7311, 0.0000]
output[0,0]4.20294.2029

49. Restore the missing line: Show-it-off — SelfAttention module

Fill the middle

Fill in the blanks

From Show-it-off — SelfAttention module — one line has had its right-hand side removed. Put it back.

class SelfAttention(nn.Module):
def __init__(self, d_model):
super().__init__()
self.Wq = nn.Linear(d_model, d_model, bias=False)
self.Wk = nn.Linear(d_model, d_model, bias=False)
self.Wv = nn.Linear(d_model, d_model, bias=False)
self.Wo = nn.Linear(d_model, d_model, bias=False)

def forward(self, x, causal=False):
Q, K, V = self.Wq(x), self.Wk(x), self.Wv(x)
out, weights = scaled_dot_product_attention(Q, K, V, causal)
return self.Wo(out), weights

sa = SelfAttention(d_model=16)
x = torch.randn(2, 5, 16) # batch=2, seq=5, d=16
out, w = sa(x, causal=True)
print('output:', out.shape) # (2, 5, 16)
print('weights:', w.shape) # (2, 5, 5)
params = sum(p.numel() for p in sa.parameters())
print('params:', params) # 1024

Why: x is what everything below it consumes, so the wrong expression here fails later and somewhere else. The four projection matrices are the only trainable parameters.

50. Show-it-off — SelfAttention module

Worked example

Wrap scaled_dot_product_attention in a single-head SelfAttention nn.Module with W_Q, W_K, W_V, W_O

Why: The four projection matrices are the only trainable parameters. For d_model=16, bias=False: 4 × 16² = 1024 parameters.

class SelfAttention(nn.Module):
    def __init__(self, d_model):
        super().__init__()
        self.Wq = nn.Linear(d_model, d_model, bias=False)
        self.Wk = nn.Linear(d_model, d_model, bias=False)
        self.Wv = nn.Linear(d_model, d_model, bias=False)
        self.Wo = nn.Linear(d_model, d_model, bias=False)

    def forward(self, x, causal=False):
        Q, K, V = self.Wq(x), self.Wk(x), self.Wv(x)
        out, weights = scaled_dot_product_attention(Q, K, V, causal)
        return self.Wo(out), weights

sa = SelfAttention(d_model=16)
x = torch.randn(2, 5, 16)   # batch=2, seq=5, d=16
out, w = sa(x, causal=True)
print('output:', out.shape)      # (2, 5, 16)
print('weights:', w.shape)        # (2, 5, 5)
params = sum(p.numel() for p in sa.parameters())
print('params:', params)          # 1024
attributevaluederivation
output shape(2, 5, 16)same as input — standard for self-attention
weight shape(2, 5, 5)(batch, T_query, T_key)
param count10244 × 16² = 1024 (no bias)
causal guaranteew[b,t,t'] = 0 for t' > tupper-tri mask before softmax

51. Fill in: derivation for Show-it-off — SelfAttention module

Comparison

Comparison matrix

From Show-it-off — SelfAttention module: refill the derivation column from what you know. The rest of the table is as it appeared.

attributevaluederivation
output shape(2, 5, 16)same as input — standard for self-attention
weight shape(2, 5, 5)(batch, T_query, T_key)
param count10244 × 16² = 1024 (no bias)
causal guaranteew[b,t,t'] = 0 for t' > tupper-tri mask before softmax

52. Connect it up: Lesson 114: Phase 3 Mock Exam — Transformers, NLP, and Computer Vision

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Section 1 — Transformer Theory · Section 2 — NLP: BERT, Generation, Metrics · Section 3 — Coding & Computer Vision · Your Turn — Coding Problem. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

53. Phase 3 complete — what you can now do

Recap

topickey formula / factlesson ref
Scaled dot-product attentionsoftmax(QK^T / sqrt(d_k)) VLesson 88
MHA params4 × d_model² (no bias)Lesson 89
Causal masktriu(ones, diagonal=1) → -infLesson 89
Sinusoidal PEsin/cos with freq 10000^(2i/d)Lesson 88
BLEU-4BP · exp(mean log P_1..4)Lesson 101
LoRA2drl × n_layers × n_adaptersLesson 98
NMSsort → greedily keep → suppress IoU > threshLesson 107

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 114 — Phase 3 Timed Mock Exam — Barron · USAAIO Round 2 Preparation, 2026
  2. Vaswani et al. 'Attention Is All You Need' (NeurIPS 2017) — arXiv:1706.03762
  3. Devlin et al. 'BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding' (NAACL 2019) — arXiv:1810.04805
  4. Hu et al. 'LoRA: Low-Rank Adaptation of Large Language Models' (ICLR 2022) — arXiv:2106.09685
  5. Papineni et al. 'BLEU: a Method for Automatic Evaluation of Machine Translation' (ACL 2002) — Association for Computational Linguistics
  6. Attention math, BLEU-4, ROUGE-N, NMS/IoU, LoRA param counts, and CE/perplexity verified with torch 2.7.1+cpu and numpy 2.2.6, June 2026 — Real execution, verified

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108