USAAIO Lesson 114, a Phase 3 review. It is a timed mock-exam simulation covering all the Phase 3 topics: deriving scaled dot-product attention, counting the parameters of multi-head attention, causal masking, sinusoidal positional encoding, the BERT architecture and how to fine-tune it, LoRA adapters, greedy against beam-search decoding, computing BLEU-4 and ROUGE-N, non-maximum suppression with IoU, and the cross-entropy and perplexity of a full transformer training run. Every value in a trace table was computed with torch 2.7.1+cpu, and each section comes with a worked answer-key walkthrough. The lesson runs to 27 slides.
Subject: Machine Learning · 53 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 114 · Phase 3 Capstone
Timed simulation: Section 1 (transformers), Section 2 (NLP), Section 3 (CV + coding). Full answer-key walkthrough. Target: >80% theory, one coding problem complete.
Objectives
Warm-up
Discussion prompt
Before we open Lesson 114: Phase 3 Mock Exam — Transformers, NLP, and Computer Vision: without looking back, what was the main idea of CLIP — Contrastive Language-Image Pretraining, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
symmetric InfoNCE contrastive loss, dual-encoder CLIP architecture (image + text tower), learned temperature parameter, zero-shot transfer via text-based class descriptors, and CLIP extensions (ALIGN, BLIP-2).
Section
60 minutes
Concept
Every exam question on attention reduces to four operations. Write them from memory before reading the formula.
\[ \text{Attention}(Q,K,V) = \text{softmax}\!\left(\frac{QK^\top}{\sqrt{d_k}}\right)V \]
| step | operation | shape out (seq=T, d_k) |
|---|---|---|
| 1 | raw scores S = QK^T | (T, T) |
| 2 | scale S / sqrt(d_k) | (T, T) |
| 3 | softmax row-wise | (T, T) rows sum to 1 |
| 4 | weighted sum A · V | (T, d_v) |
Comparison
Comparison matrix
From Attention derivation — the four steps: refill the operation column from what you know. The rest of the table is as it appeared.
| step | operation | shape out (seq=T, d_k) |
|---|---|---|
| 1 | raw scores S = QK^T | (T, T) |
| 2 | scale S / sqrt(d_k) | (T, T) |
| 3 | softmax row-wise | (T, T) rows sum to 1 |
| 4 | weighted sum A · V | (T, d_v) |
Ranking
Put in order
Put the moves of Attention trace — Q, K, V ∈ R^{3×4}, d_k = 4 into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Dot product of each query row with every key row.
Worked example
Compute raw score matrix S = QK^T
Why: Dot product of each query row with every key row. With Q[0]=[1,0,1,0] and K[0]=[1,0,1,0]: Q[0]·K[0]=2, Q[0]·K[1]=0, Q[0]·K[2]=1 → row 0 = [2, 0, 1].
Scale by 1/√d_k = 1/2 → S_scaled[0] = [1.000, 0.000, 0.500]
Why: Without scaling, large d_k pushes dot products into saturated softmax regions, collapsing gradients. Dividing by sqrt(d_k) stabilises variance.
Softmax row 0: exp([1.0, 0.0, 0.5]) / sum → [0.5065, 0.1863, 0.3072]
Why: exp(1.0)=2.718, exp(0.0)=1.000, exp(0.5)=1.649; sum=5.366; divide each. These are the attention weights for query 0 over the three keys.
Output row 0 = 0.5065·V[0] + 0.1863·V[1] + 0.3072·V[2] = [4.2029, 5.2029, 6.2029, 7.2029]
Why: V rows are [1,2,3,4], [5,6,7,8], [9,10,11,12]. Weighted sum: 0.5065·1 + 0.1863·5 + 0.3072·9 = 4.2029; pattern holds for all 4 columns.
| row | attn weights | output (first 2 cols) |
|---|---|---|
| Q[0] | [0.5065, 0.1863, 0.3072] | [4.2029, 5.2029] |
| Q[1] | [0.1863, 0.5065, 0.3072] | [5.4835, 6.4835] |
| Q[2] | [0.3837, 0.3837, 0.2327] | [4.3962, 5.3962] |
Trade off
Comparison matrix
From Attention trace — Q, K, V ∈ R^{3×4}, d_k = 4: every row here is a choice with a cost. Fill the attn weights column, then say which row you would actually pick and what you give up for it.
| row | attn weights | output (first 2 cols) |
|---|---|---|
| Q[0] | [0.5065, 0.1863, 0.3072] | [4.2029, 5.2029] |
| Q[1] | [0.1863, 0.5065, 0.3072] | [5.4835, 6.4835] |
| Q[2] | [0.3837, 0.3837, 0.2327] | [4.3962, 5.3962] |
Concept
A causal (autoregressive) mask sets future positions to -inf before softmax, forcing each position to attend only to itself and earlier tokens. Applied in GPT-style decoders; absent in BERT encoders.
import torch, math, torch.nn.functional as F
# seq=3, d_k=4 (same Q,K,V from worked example)
scores_scaled = torch.tensor(
[[1.0000, 0.0000, 0.5000],
[0.0000, 1.0000, 0.5000],
[0.5000, 0.5000, 0.0000]])
causal_mask = torch.triu(torch.ones(3, 3), diagonal=1).bool()
masked = scores_scaled.masked_fill(causal_mask, float('-inf'))
weights = F.softmax(masked, dim=-1)
print(masked)
print(weights)| position | masked scores row | attn weights (causal) |
|---|---|---|
| 0 | [1.000, -inf, -inf] | [1.000, 0.000, 0.000] |
| 1 | [0.000, 1.000, -inf] | [0.269, 0.731, 0.000] |
| 2 | [0.500, 0.500, 0.000] | [0.384, 0.384, 0.233] |
Concept
Multi-head attention (Lesson 89) splits d_model into n_heads heads of size d_k = d_model / n_heads. Four weight matrices: W_Q, W_K, W_V (each d_model × d_k per head) and W_O (d_model × d_model).
\[ \text{params} = n_h \cdot 3 d_{\text{model}} d_k + d_{\text{model}}^2 = 4 d_{\text{model}}^2 \quad\text{(no bias)} \]
| d_model | n_heads | d_k | 3×d×dk×h | W_O (d×d) | total |
|---|---|---|---|---|---|
| 64 | 4 | 16 | 3×64×16×4 = 12,288 | 64×64 = 4,096 | 16,384 |
| 768 | 12 | 64 | 3×768×64×12 = 1,769,472 | 768×768 = 589,824 | 2,359,296 |
| 1024 | 16 | 64 | 3×1024×64×16 = 3,145,728 | 1024×1024 = 1,048,576 | 4,194,304 |
Verified: nn.MultiheadAttention(embed_dim=64, num_heads=4, bias=False) reports 16,384 parameters — matches the formula exactly.
Counterexample
Discussion prompt
Verified: nn.MultiheadAttention(embed_dim=64, num_heads=4, bias=False) reports 16,384 parameters — matches the formula exactly.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Exam question: total MHA params for d_model=64, n_heads=4, bias=False?
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Stops after computing W_Q, W_K, W_V for all heads.
Exam question: total MHA params for d_model=64, n_heads=4, bias=False?
Why: Stops after computing W_Q, W_K, W_V for all heads. Misses the output projection W_O that recombines head outputs back to d_model.
Trap
Exam question: total MHA params for d_model=64, n_heads=4, bias=False?
Count only the QKV projections: 4 heads × 3 × 64 × 16 = 12,288
Why: Stops after computing W_Q, W_K, W_V for all heads. Misses the output projection W_O that recombines head outputs back to d_model.
Answer: 12,288 — WRONG
Why: Off by 4,096 (the W_O matrix). In PyTorch nn.MultiheadAttention stores W_O as out_proj, which has d_model×d_model weights.
Exam question: total MHA params for d_model=64, n_heads=4, bias=False?
QKV projections: 4 heads × 3 × (64 × 16) = 12,288; W_O: 64 × 64 = 4,096
Why: W_O is a d_model × d_model matrix that merges the concatenated head outputs back into d_model dimensions — always present, always d^2.
Answer: 12,288 + 4,096 = 16,384 — CORRECT
Why: Matches sum(p.numel() for p in nn.MultiheadAttention(64, 4, bias=False).parameters()) = 16384.
Break the constraint
Discussion prompt
The rule this trap just fixed:
W_O is a d_model × d_model matrix that merges the concatenated head outputs back into d_model dimensions — always present, always d^2.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
Stops after computing W_Q, W_K, W_V for all heads. Misses the output projection W_O that recombines head outputs back to d_model.
Section
60 minutes
Concept
\[ PE_{(pos,\,2i)} = \sin\!\left(\frac{pos}{10000^{2i/d}}\right), \quad PE_{(pos,\,2i+1)} = \cos\!\left(\frac{pos}{10000^{2i/d}}\right) \]
Key exam facts: PE(0) is all-zeros for sin dims and all-ones for cos dims. High-frequency dimensions (small i) change rapidly with position; low-frequency dims (large i) barely change. These are added to the token embedding, not concatenated.
| pos | dim 0 (sin) | dim 1 (cos) | dim 2 (sin) | dim 3 (cos) |
|---|---|---|---|---|
| 0 | 0.0000 | 1.0000 | 0.0000 | 1.0000 |
| 1 | 0.8415 | 0.5403 | 0.0998 | 0.9950 |
| 2 | 0.9093 | -0.4161 | 0.1987 | 0.9801 |
| 3 | 0.1411 | -0.9900 | 0.2955 | 0.9553 |
| 4 | -0.7568 | -0.6536 | 0.3894 | 0.9211 |
Pattern
Step through it
Step through Sinusoidal positional encoding — what the values look like one row at a time. What is driving the change, and what would the row after the last one be?
Concept
Full fine-tuning updates all 110M BERT-base parameters. Classification fine-tuning adds only a d_model × n_classes head and freezes the encoder. LoRA (Lesson 98) instead injects low-rank adapters into the frozen W_Q and W_V matrices.
\[ \Delta W = BA, \quad B \in \mathbb{R}^{d \times r},\; A \in \mathbb{R}^{r \times d}, \quad r \ll d \]
| strategy | trainable params | % of BERT-base |
|---|---|---|
| Full fine-tune | 110,000,000 | 100.0% |
| Head-only (2-class) | 1,538 | 0.0014% |
| LoRA r=8 (Q+V, 12 layers) | 294,912 | 0.27% |
LoRA per adapter: B is (768 × 8) = 6,144 params, A is (8 × 768) = 6,144 params. Two adapters (Q and V) per layer × 12 layers × 2 matrices × 6,144 = 294,912 trainable params.
Analogy
Discussion prompt
Explain BERT fine-tuning vs LoRA — parameter efficiency by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
LoRA per adapter: B is (768 × 8) = 6,144 params, A is (8 × 768) = 6,144 params. Two adapters (Q and V) per layer × 12 layers × 2 matrices × 6,144 = 294,912 trainable params.
Concept
Greedy decoding picks argmax of the softmax at each step — deterministic, fast, but repetitive. Top-k sampling restricts to the k highest-prob tokens then resamples — stochastic, higher diversity. Beam search keeps the top-b partial sequences at each step.
import torch, torch.nn.functional as F
vocab = ['<end>', 'A', 'B', 'C']
logits_t1 = torch.tensor([0.1, 2.0, 1.5, 0.4])
# Greedy: argmax
probs_t1 = F.softmax(logits_t1, dim=0)
greedy_tok = vocab[int(probs_t1.argmax())]
# Top-k (k=2): renormalise over top-2
topk_vals, topk_idx = torch.topk(probs_t1, k=2)
topk_probs = topk_vals / topk_vals.sum()
print(f'probs: {[f"{p:.3f}" for p in probs_t1.tolist()]}')
print(f'greedy: {greedy_tok}')
print(f'top-2 probs: {topk_probs.tolist()}')
print(f'top-2 tokens: {[vocab[i] for i in topk_idx.tolist()]}')| step | logits | softmax probs | greedy pick |
|---|---|---|---|
| 1 | [0.1, 2.0, 1.5, 0.4] | [0.076, 0.511, 0.310, 0.103] | A |
| 2 | [0.2, 0.5, 3.0, 0.3] | [0.050, 0.068, 0.826, 0.056] | B |
| 3 | [5.0, 0.1, 0.2, 0.1] | [0.977, 0.007, 0.008, 0.007] | <end> |
Ranking
Put in order
Put the moves of BLEU-4 by hand — ref vs hypothesis into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. One word differs: 'a' replaces 'the'.
Worked example
Set up: ref = 'the cat sat on the mat', hyp = 'the cat sat on a mat'
Why: One word differs: 'a' replaces 'the'. BLEU clips counts to the reference count to prevent gaming by repetition.
Compute clipped n-gram precisions: P_1=5/6, P_2=3/5, P_3=2/4, P_4=1/3
Why: 1-grams: 5 of 6 hyp unigrams appear in ref (all but 'a'). 2-grams: 3 of 5 match. 3-grams: 2 of 4 match. 4-grams: 1 of 3 match ('the cat sat on' is not in hyp but 'cat sat on a' isn't in ref; 'the cat sat on' is not a 4-gram of hyp — only 'cat sat on a' and 'sat on a mat' and 'the cat sat on'. ref has 'the cat sat on' not 'cat sat on a'. Verify manually: hyp 4-grams = {'the cat sat on','cat sat on a','sat on a mat'}; ref 4-grams = {'the cat sat on','cat sat on the','sat on the mat','on the mat '}. Match: 1. Correct.)
Brevity penalty: BP = 1 since |hyp|=|ref|=6
Why: BP = 1 when the hypothesis is not shorter than the reference. BP = exp(1 - r/c) when c < r. Here c = r = 6 so BP = 1.
BLEU-4 = BP · exp((1/4) · Σ log P_n) = exp((log(5/6)+log(3/5)+log(2/4)+log(1/3))/4) = 0.5373
Why: log(0.8333)+log(0.6000)+log(0.5000)+log(0.3333) = -0.1823-0.5108-0.6931-1.0986 = -2.4848; /4 = -0.6212; exp(-0.6212) = 0.5373.
| n | matches | hyp n-grams | P_n |
|---|---|---|---|
| 1 | 5 | 6 | 0.8333 |
| 2 | 3 | 5 | 0.6000 |
| 3 | 2 | 4 | 0.5000 |
| 4 | 1 | 3 | 0.3333 |
| BLEU-4 | BP=1.000 | exp(mean log) | 0.5373 |
Comparison
Comparison matrix
From BLEU-4 by hand — ref vs hypothesis: refill the hyp n-grams column from what you know. The rest of the table is as it appeared.
| n | matches | hyp n-grams | P_n |
|---|---|---|---|
| 1 | 5 | 6 | 0.8333 |
| 2 | 3 | 5 | 0.6000 |
| 3 | 2 | 4 | 0.5000 |
| 4 | 1 | 3 | 0.3333 |
| BLEU-4 | BP=1.000 | exp(mean log) | 0.5373 |
Concept
BLEU is precision-oriented (how much of the hypothesis is in the reference). ROUGE is recall-oriented (how much of the reference is covered by the hypothesis). Both compute n-gram overlap; ROUGE is the dominant metric for summarisation.
\[ \text{ROUGE-N Recall} = \frac{\sum_{\text{ngrams}\in\text{ref}} \min(\text{count}_{\text{hyp}}, \text{count}_{\text{ref}})}{\sum_{\text{ngrams}\in\text{ref}} \text{count}_{\text{ref}}} \]
| metric | precision | recall | F1 |
|---|---|---|---|
| ROUGE-1 | 0.8333 | 0.8333 | 0.8333 |
| ROUGE-2 | 0.6000 | 0.6000 | 0.6000 |
Same ref/hyp pair as BLEU. P=R=F1 here because |hyp|=|ref|. In general they differ: a verbose hypothesis has high recall but low precision; a terse one has high precision but low recall.
Analogy
Discussion prompt
Explain ROUGE-N — recall-oriented complement to BLEU by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Same ref/hyp pair as BLEU. P=R=F1 here because |hyp|=|ref|. In general they differ: a verbose hypothesis has high recall but low precision; a terse one has high precision but low recall.
Section
60 minutes
Concept
Object detectors produce many overlapping candidate boxes. NMS keeps the highest-score box, suppresses all others with IoU above a threshold, then repeats on the survivors. IoU = intersection / union.
\[ \text{IoU}(A, B) = \frac{|A \cap B|}{|A \cup B|} = \frac{\text{inter\_area}}{\text{area}_A + \text{area}_B - \text{inter\_area}} \]
| box pair | intersection area | union area | IoU | NMS action (thresh=0.5) |
|---|---|---|---|---|
| box0 vs box1 ([10,10,60,60] vs [12,12,62,62]) | 48×48=2304 | 50×50+50×50-2304=2696 | 0.8546 | SUPPRESS box1 |
| box0 vs box2 ([10,10,60,60] vs [80,80,120,120]) | 0 | 2500+1600-0=4100 | 0.0000 | keep box2 |
| box1 vs box2 (already suppressed) | 0 | — | 0.0000 | — |
NMS result: boxes 0 (score 0.9) and 2 (score 0.7) kept; box 1 suppressed because IoU(box0, box1) = 0.8546 > 0.50.
Counterexample
Discussion prompt
NMS result: boxes 0 (score 0.9) and 2 (score 0.7) kept; box 1 suppressed because IoU(box0, box1) = 0.8546 > 0.50.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Ranking
Put in order
Put the moves of Implement multi-head attention from scratch into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Each of the n_heads heads attends to a d_k-dimensional subspace.
Worked example
Project X into Q, K, V using linear layers, then reshape to (B, n_heads, T, d_k)
Why: Each of the n_heads heads attends to a d_k-dimensional subspace. Reshape from (B, T, d_model) → (B, T, n_heads, d_k) then transpose dims 1 and 2.
Compute scaled dot-product attention per head: attn = softmax(Q @ K^T / sqrt(d_k)) @ V
Why: Broadcasting: Q and K are (B, heads, T, d_k); Q @ K.transpose(-2,-1) gives (B, heads, T, T). Softmax on dim=-1 (over keys). Result (B, heads, T, d_k).
Merge heads: transpose back to (B, T, heads, d_k), reshape to (B, T, d_model), apply W_O
Why: .transpose(1,2).contiguous().view(B, T, C) — contiguous() is required before view because transpose creates a non-contiguous tensor.
import torch, torch.nn as nn, torch.nn.functional as F, math
class TinyMHA(nn.Module):
def __init__(self, d_model=16, n_heads=2):
super().__init__()
self.d_k = d_model // n_heads
self.n_heads = n_heads
self.Wq = nn.Linear(d_model, d_model, bias=False)
self.Wk = nn.Linear(d_model, d_model, bias=False)
self.Wv = nn.Linear(d_model, d_model, bias=False)
self.Wo = nn.Linear(d_model, d_model, bias=False)
def forward(self, x):
B, T, C = x.shape
# Project + split heads
Q = self.Wq(x).view(B, T, self.n_heads, self.d_k).transpose(1, 2)
K = self.Wk(x).view(B, T, self.n_heads, self.d_k).transpose(1, 2)
V = self.Wv(x).view(B, T, self.n_heads, self.d_k).transpose(1, 2)
# Attention
attn = (Q @ K.transpose(-2, -1)) / math.sqrt(self.d_k)
attn = F.softmax(attn, dim=-1)
out = (attn @ V).transpose(1, 2).contiguous().view(B, T, C)
return self.Wo(out)
torch.manual_seed(42)
model = TinyMHA(d_model=16, n_heads=2)
x = torch.randn(2, 5, 16)
print(model(x).shape) # (2, 5, 16)
params = sum(p.numel() for p in model.parameters())
print(f'params: {params}') # 4 * 16^2 / bias=False = 1024| tensor | shape after step | note |
|---|---|---|
| x input | (2, 5, 16) | B=2, T=5, d=16 |
| Q (after Linear + view + transpose) | (2, 2, 5, 8) | B, heads=2, T, d_k=8 |
| attn weights | (2, 2, 5, 5) | B, heads, T, T |
| attn @ V | (2, 2, 5, 8) | B, heads, T, d_k |
| after merge + Wo | (2, 5, 16) | back to input shape |
| param count | 4 × 16² | = 1024 (bias=False) |
Error analysis
Annotate
Walk the callouts on Implement multi-head attention from scratch. Each one is a place this is easy to get subtly wrong.
.transpose(1,2).contiguous().view(B, T, C) — contiguous() is required before view because transpose creates a non-contiguous tensor.Concept
A language model's training loss is CrossEntropyLoss over the vocabulary at each position. Perplexity = exp(CE loss) — the geometric mean number of equally-likely tokens the model is considering. Lower is better; perfect model has PP=1.
\[ \mathcal{L}_{\text{CE}} = -\frac{1}{N}\sum_i \log p_{\theta}(y_i \mid x_i), \qquad \text{PP} = e^{\mathcal{L}_{\text{CE}}} \]
import torch, torch.nn as nn, math
# batch=2, seq=4, vocab=3
logits = torch.tensor([
[[2.0, 0.5, 0.1], [0.1, 3.0, 0.2], [0.3, 0.2, 2.8], [1.0, 0.5, 0.1]],
[[0.5, 2.2, 0.3], [0.2, 0.1, 3.5], [2.0, 0.5, 0.2], [0.5, 1.0, 0.5]],
])
targets = torch.tensor([[0, 1, 2, 0], [1, 2, 0, 1]])
loss = nn.CrossEntropyLoss()(logits.view(-1, 3), targets.view(-1))
print(f'CE loss: {loss.item():.4f}') # 0.3436
print(f'Perplexity: {math.exp(loss.item()):.4f}') # 1.4100| CE loss | perplexity | interpretation |
|---|---|---|
| 0.3436 | 1.4100 | model is almost always right (3-class) |
| 1.0986 | 3.0000 | uniform over 3 tokens — no signal |
| 0.0 | 1.0 | perfect model — 100% prob on correct token |
Sorting
Sort into buckets
These are the pieces of Lesson 114: Phase 3 Mock Exam — Transformers, NLP, and Computer Vision, out of order. Put each one back under the part of the lesson it belongs to.
Constraint
Discussion prompt
Run Phase 3 exam strategy — 6-step recipe with this step confiscated:
ROUGE: compute n-gram overlap; recall = overlap / ref_ngrams; precision = overlap / hyp_ngrams; F1 = harmonic mean.
Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.
Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.
Answer:
triu(..., diagonal=1) filled with -inf before softmax.Pattern
triu(..., diagonal=1) filled with -inf before softmax.Edge cases
Discussion prompt
Phase 3 exam strategy — 6-step recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
triu(..., diagonal=1) filled with -inf before softmax.Elimination
Eliminate the wrong options
A single-head attention module has Q ∈ R^{8×64}, K ∈ R^{6×64}, V ∈ R^{6×32}. What is the shape of the attention output?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: QK^T gives (8,64)×(64,6) = (8,6) score matrix. After softmax (8,6). Multiply by V ∈ (6,32): (8,6)×(6,32) = (8,32). Output shape matches V's last dimension, rows match Q's sequence length.
Check
Solve on paper before selecting.
Check your understanding
A single-head attention module has Q ∈ R^{8×64}, K ∈ R^{6×64}, V ∈ R^{6×32}. What is the shape of the attention output?
Answer: A
Why: QK^T gives (8,64)×(64,6) = (8,6) score matrix. After softmax (8,6). Multiply by V ∈ (6,32): (8,6)×(6,32) = (8,32). Output shape matches V's last dimension, rows match Q's sequence length.
Prediction
Predict first
Reference length r=10, hypothesis length c=7. What is the BLEU brevity penalty?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: exp(1 - 10/7) ≈ 0.654
Why: BP = exp(1 - r/c) when c < r. Here c=7 < r=10 so BP = exp(1 - 10/7) = exp(-0.4286) ≈ 0.6514. This penalises short hypotheses — a one-word hypothesis with a perfect 1-gram match would still get a near-zero BLEU.
Check
Work out BP before selecting.
Check your understanding
Reference length r=10, hypothesis length c=7. What is the BLEU brevity penalty?
Answer: A
Why: BP = exp(1 - r/c) when c < r. Here c=7 < r=10 so BP = exp(1 - 10/7) = exp(-0.4286) ≈ 0.6514. This penalises short hypotheses — a one-word hypothesis with a perfect 1-gram match would still get a near-zero BLEU.
Elimination
Eliminate the wrong options
A 12-layer transformer has d_model=768. LoRA adapters (rank r=8, no bias) are inserted into W_Q and W_V of every layer. How many trainable parameters does LoRA add?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Each LoRA adapter = A (r×d) + B (d×r) = 2×768×8 = 12,288 params. Two adapters per layer (Q and V): 24,576. Times 12 layers: 294,912. Verified with the formula 12 × 2 × 2 × 768 × 8.
Check
Compute the answer on paper first.
Check your understanding
A 12-layer transformer has d_model=768. LoRA adapters (rank r=8, no bias) are inserted into W_Q and W_V of every layer. How many trainable parameters does LoRA add?
Answer: A
Why: Each LoRA adapter = A (r×d) + B (d×r) = 2×768×8 = 12,288 params. Two adapters per layer (Q and V): 24,576. Times 12 layers: 294,912. Verified with the formula 12 × 2 × 2 × 768 × 8.
Section
Timed coding challenge
Concept
Goal: Implement scaled dot-product attention from scratch in PyTorch (no nn.MultiheadAttention), add optional causal masking, and verify output shapes and parameter counts.
scaled_dot_product_attention(Q, K, V, causal=False) that returns the attended output and the attention weight matrix.causal=True and confirm position 0 attends only to itself (weight row 0 = [1.0, 0.0, 0.0]).SelfAttention(nn.Module) with learnable W_Q, W_K, W_V, W_O.Pattern
Predict first
The table runs: output shape | (3, 4) | (3, 4) · weight row sums | [1.000, 1.000, 1.000] | [1.0000, 1.0000, 1.0000] · causal row 0 | [1.0, 0.0, 0.0] | [1.0000, 0.0000, 0.0000] · causal row 1 | [0.269, 0.731, 0.0] | [0.2689, 0.7311, 0.0000]
In Coding challenge — solution walkthrough, given the rows so far: what is the next one — the row where check is output[0,0]?
Correct: output[0,0] | 4.2029 | 4.2029
| check | expected | actual |
|---|---|---|
| output shape | (3, 4) | (3, 4) |
| weight row sums | [1.000, 1.000, 1.000] | [1.0000, 1.0000, 1.0000] |
| causal row 0 | [1.0, 0.0, 0.0] | [1.0000, 0.0000, 0.0000] |
| causal row 1 | [0.269, 0.731, 0.0] | [0.2689, 0.7311, 0.0000] |
| output[0,0] | 4.2029 | 4.2029 |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. Upper-triangular mask (diagonal=1) zeros out future positions before softmax.
Worked example
Implement the attention function with optional causal mask
Why: Upper-triangular mask (diagonal=1) zeros out future positions before softmax. masked_fill(-inf) ensures those positions become 0 after softmax (since exp(-inf)=0).
import torch, torch.nn as nn, torch.nn.functional as F, math
def scaled_dot_product_attention(Q, K, V, causal=False):
d_k = Q.shape[-1]
scores = Q @ K.transpose(-2, -1) / math.sqrt(d_k)
if causal:
T = scores.shape[-1]
mask = torch.triu(torch.ones(T, T), diagonal=1).bool()
scores = scores.masked_fill(mask, float('-inf'))
weights = F.softmax(scores, dim=-1)
return weights @ V, weights
torch.manual_seed(42)
Q = torch.tensor([[1.,0.,1.,0.],[0.,1.,0.,1.],[1.,1.,0.,0.]])
K = torch.tensor([[1.,0.,1.,0.],[0.,1.,0.,1.],[0.,0.,1.,1.]])
V = torch.tensor([[1.,2.,3.,4.],[5.,6.,7.,8.],[9.,10.,11.,12.]])
out, w = scaled_dot_product_attention(Q, K, V, causal=False)
out_c, w_c = scaled_dot_product_attention(Q, K, V, causal=True)
print('output shape:', out.shape)
print('weights row sums:', w.sum(dim=-1))
print('causal row 0:', w_c[0].tolist())| check | expected | actual |
|---|---|---|
| output shape | (3, 4) | (3, 4) |
| weight row sums | [1.000, 1.000, 1.000] | [1.0000, 1.0000, 1.0000] |
| causal row 0 | [1.0, 0.0, 0.0] | [1.0000, 0.0000, 0.0000] |
| causal row 1 | [0.269, 0.731, 0.0] | [0.2689, 0.7311, 0.0000] |
| output[0,0] | 4.2029 | 4.2029 |
Trade off
Comparison matrix
From Coding challenge — solution walkthrough: every row here is a choice with a cost. Fill the actual column, then say which row you would actually pick and what you give up for it.
| check | expected | actual |
|---|---|---|
| output shape | (3, 4) | (3, 4) |
| weight row sums | [1.000, 1.000, 1.000] | [1.0000, 1.0000, 1.0000] |
| causal row 0 | [1.0, 0.0, 0.0] | [1.0000, 0.0000, 0.0000] |
| causal row 1 | [0.269, 0.731, 0.0] | [0.2689, 0.7311, 0.0000] |
| output[0,0] | 4.2029 | 4.2029 |
Fill the middle
Fill in the blanks
From Show-it-off — SelfAttention module — one line has had its right-hand side removed. Put it back.
class SelfAttention(nn.Module):
def __init__(self, d_model):
super().__init__()
self.Wq = nn.Linear(d_model, d_model, bias=False)
self.Wk = nn.Linear(d_model, d_model, bias=False)
self.Wv = nn.Linear(d_model, d_model, bias=False)
self.Wo = nn.Linear(d_model, d_model, bias=False)
def forward(self, x, causal=False):
Q, K, V = self.Wq(x), self.Wk(x), self.Wv(x)
out, weights = scaled_dot_product_attention(Q, K, V, causal)
return self.Wo(out), weights
sa = SelfAttention(d_model=16)
x = torch.randn(2, 5, 16) # batch=2, seq=5, d=16
out, w = sa(x, causal=True)
print('output:', out.shape) # (2, 5, 16)
print('weights:', w.shape) # (2, 5, 5)
params = sum(p.numel() for p in sa.parameters())
print('params:', params) # 1024
Why: x is what everything below it consumes, so the wrong expression here fails later and somewhere else. The four projection matrices are the only trainable parameters.
Worked example
Wrap scaled_dot_product_attention in a single-head SelfAttention nn.Module with W_Q, W_K, W_V, W_O
Why: The four projection matrices are the only trainable parameters. For d_model=16, bias=False: 4 × 16² = 1024 parameters.
class SelfAttention(nn.Module):
def __init__(self, d_model):
super().__init__()
self.Wq = nn.Linear(d_model, d_model, bias=False)
self.Wk = nn.Linear(d_model, d_model, bias=False)
self.Wv = nn.Linear(d_model, d_model, bias=False)
self.Wo = nn.Linear(d_model, d_model, bias=False)
def forward(self, x, causal=False):
Q, K, V = self.Wq(x), self.Wk(x), self.Wv(x)
out, weights = scaled_dot_product_attention(Q, K, V, causal)
return self.Wo(out), weights
sa = SelfAttention(d_model=16)
x = torch.randn(2, 5, 16) # batch=2, seq=5, d=16
out, w = sa(x, causal=True)
print('output:', out.shape) # (2, 5, 16)
print('weights:', w.shape) # (2, 5, 5)
params = sum(p.numel() for p in sa.parameters())
print('params:', params) # 1024| attribute | value | derivation |
|---|---|---|
| output shape | (2, 5, 16) | same as input — standard for self-attention |
| weight shape | (2, 5, 5) | (batch, T_query, T_key) |
| param count | 1024 | 4 × 16² = 1024 (no bias) |
| causal guarantee | w[b,t,t'] = 0 for t' > t | upper-tri mask before softmax |
Comparison
Comparison matrix
From Show-it-off — SelfAttention module: refill the derivation column from what you know. The rest of the table is as it appeared.
| attribute | value | derivation |
|---|---|---|
| output shape | (2, 5, 16) | same as input — standard for self-attention |
| weight shape | (2, 5, 5) | (batch, T_query, T_key) |
| param count | 1024 | 4 × 16² = 1024 (no bias) |
| causal guarantee | w[b,t,t'] = 0 for t' > t | upper-tri mask before softmax |
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Section 1 — Transformer Theory · Section 2 — NLP: BERT, Generation, Metrics · Section 3 — Coding & Computer Vision · Your Turn — Coding Problem. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
| topic | key formula / fact | lesson ref |
|---|---|---|
| Scaled dot-product attention | softmax(QK^T / sqrt(d_k)) V | Lesson 88 |
| MHA params | 4 × d_model² (no bias) | Lesson 89 |
| Causal mask | triu(ones, diagonal=1) → -inf | Lesson 89 |
| Sinusoidal PE | sin/cos with freq 10000^(2i/d) | Lesson 88 |
| BLEU-4 | BP · exp(mean log P_1..4) | Lesson 101 |
| LoRA | 2drl × n_layers × n_adapters | Lesson 98 |
| NMS | sort → greedily keep → suppress IoU > thresh | Lesson 107 |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.