Lesson 111: Transformers, NLP & CV — Phase 3 Review

USAAIO Lesson 111, the Phase 3 capstone review. It derives scaled dot-product attention and analyzes its complexity, covers sinusoidal positional encoding and BPE tokenization step by step, and implements and debugs multi-head attention. It then covers the IoU formula and its computation, the U-Net skip-connection architecture, and a comparison of ViT with CNNs. All the values were computed with torch 2.7.1+cpu and numpy 2.2.6 on synthetic data. The lesson runs to 28 slides.

Subject: Machine Learning · 56 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Phase 3 Review: Transformers, NLP & CV

Title

USAAIO · Lesson 111 · Week 38

Capstone review before Phase 4 (generative models). Flash-derive attention, trace BPE, compute IoU, debug multi-head attention, compare ViT vs CNN — all exam-level, all numbers real.

2. By the end of this lesson you can

Objectives

  1. Derive scaled dot-product attention from scratch and state why the sqrt(d_k) factor is non-optional
  2. Trace sinusoidal positional encoding: compute PE[pos, 2i] and PE[pos, 2i+1] for concrete (pos, i) values
  3. Execute 2 steps of BPE tokenization by hand on a toy corpus and identify the merged pair
  4. Implement and debug multi-head attention in PyTorch — catch the most common implementation mistakes
  5. Compute IoU for two given bounding boxes and explain how NMS uses it
  6. Contrast U-Net and ViT architecturally: skip connections vs attention, segmentation vs classification

3. What survived from Semantic Segmentation & U-Net?

Warm-up

Discussion prompt

Before we open Lesson 111: Transformers, NLP & CV — Phase 3 Review: without looking back, what was the main idea of Semantic Segmentation & U-Net, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

pixel-wise classification via semantic segmentation, U-Net encoder-decoder with skip connections (concat not add), upsampling methods (bilinear, ConvTranspose2d, pixel shuffle), Dice loss and Focal loss for class-imbalanced masks, and TinyUNet trained from scratch.

4. Attention — derivation and complexity

Section

Part 1 of 4

5. Scaled dot-product attention — the formula

Concept

Every transformer layer (Lessons 88-92) is built on one core operation. Given queries Q, keys K, values V (all (T, d_k)):

\[ \text{Attention}(Q, K, V) = \text{softmax}\!\left(\frac{QK^\top}{\sqrt{d_k}}\right) V \]

quantityshapemeaning
Q (queries)(T, d_k)what each token is looking for
K (keys)(T, d_k)what each token advertises
V (values)(T, d_k)what each token sends if selected
QK^T(T, T)raw attention logits
softmax out(T, T)attention weights — rows sum to 1
output(T, d_k)context-aware token representations

6. Which is which, by shape

Discrimination

Sort into buckets

Sort these by shape, from memory, without looking back at Scaled dot-product attention — the formula. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

(T, d_k)
Q (queries); K (keys); V (values); output
(T, T)
QK^T; softmax out
g1
shape is "(T, d_k)" for Q (queries), K (keys), V (values), output — that is what the table on "Scaled dot-product attention — the…" records, and it is the single property separating this group from the rest.
g2
shape is "(T, T)" for QK^T, softmax out — that is what the table on "Scaled dot-product attention — the…" records, and it is the single property separating this group from the rest.

7. Guess the shape of the answer: Trace: 3-token attention step by step

Estimation

Predict first

Three tokens, d_k = 4. Compute the attention matrix and read off the weight table. All values from torch.manual_seed(0) — verified.

Commit before you compute: what does Trace: 3-token attention step by step come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: scores[0] = [0.3803, 0.0517, -0.7319]; after softmax → [0.4881, 0.3514, 0.1605]

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Token 0's query most strongly matches Key 0 (raw score 0.3803), so it borrows most from Value 0.

8. Trace: 3-token attention step by step

Worked example

Three tokens, d_k = 4. Compute the attention matrix and read off the weight table. All values from torch.manual_seed(0) — verified.

import torch
torch.manual_seed(0)
T, dk = 3, 4
Q = torch.randn(T, dk)
K = torch.randn(T, dk)
V = torch.randn(T, dk)

scores = Q @ K.T / (dk ** 0.5)   # (3, 3)
weights = torch.softmax(scores, dim=-1)
out = weights @ V

print('scores:\n', scores.numpy().round(4))
print('weights:\n', weights.numpy().round(4))
print('row 0 sum:', weights[0].sum().item())  # 1.0
print('out shape:', out.shape)

scores[0] = [0.3803, 0.0517, -0.7319]; after softmax → [0.4881, 0.3514, 0.1605]

Why: Token 0's query most strongly matches Key 0 (raw score 0.3803), so it borrows most from Value 0. Row sum = 1.0000 — softmax guarantee.

tokenraw score to tok 0raw score to tok 1raw score to tok 2weight on tok 0weight on tok 1weight on tok 2
00.38030.0517-0.73190.48810.35140.1605
1-0.4697-0.7661-0.70520.39470.29350.3119
20.41680.2572-0.63680.45430.38730.1584

9. Watch it run: Trace: 3-token attention step by step

Pattern

Step through it

Step through Trace: 3-token attention step by step one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: token is 0
  2. Step 2: token is 1
  3. Step 3: token is 2

10. Attention complexity: O(T²d)

Concept

The QK^T matrix multiply costs O(T²·d_k) — this is the fundamental bottleneck of the transformer. Doubling sequence length quadruples the attention cost.

T (seq len)T²rel cost vs T=128
12816,3841×
512262,14416×
10241,048,57664×
20484,194,304256×

ViT-B/16 on 224×224 images: T = 196 patches + 1 CLS = 197 tokens — 197² = 38,809 per head per layer. This is why long-context transformers (GPT-4, Gemini) need FlashAttention or linear approximations.

11. Fill in: T² for Attention complexity: O(T²d)

Comparison

Comparison matrix

From Attention complexity: O(T²d): refill the T² column from what you know. The rest of the table is as it appeared.

T (seq len)T²rel cost vs T=128
12816,3841×
512262,14416×
10241,048,57664×
20484,194,304256×

12. Something is wrong here: omitting the sqrt(d_k) scale

Anomaly

Predict first

A student writes this, and it looks reasonable:

Compute raw attention without scaling: scores = Q @ K.T. Seems fine — softmax still produces a valid distribution.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: With d_k=64 the raw scores have std ≈ sqrt(64)=8, producing extreme logits.

Scale by 1/sqrt(d_k) before softmax: scores = Q @ K.T / (d_k ** 0.5).

Why: With d_k=64 the raw scores have std ≈ sqrt(64)=8, producing extreme logits. Verified: max weight → 1.000000, entropy → 0.36. The distribution collapses to near-one-hot — gradients through softmax vanish almost everywhere.

13. Trap: omitting the sqrt(d_k) scale

Trap

The trap

Compute raw attention without scaling: scores = Q @ K.T. Seems fine — softmax still produces a valid distribution.

weights = softmax(Q @ K.T) # no scale

Why: With d_k=64 the raw scores have std ≈ sqrt(64)=8, producing extreme logits. Verified: max weight → 1.000000, entropy → 0.36. The distribution collapses to near-one-hot — gradients through softmax vanish almost everywhere.

The fix

Scale by 1/sqrt(d_k) before softmax: scores = Q @ K.T / (d_k ** 0.5).

weights = softmax(Q @ K.T / (d_k ** 0.5)) # d_k=64 → scale=8.0

Why: Verified: max weight = 0.8209, entropy = 1.88. The distribution stays spread — all tokens receive meaningful gradients during backprop. The scale factor matches std(Q@K.T) ≈ sqrt(d_k) when Q,K entries are ~N(0,1).

14. NLP — positional encoding & BPE

Section

Part 2 of 4

15. Sinusoidal positional encoding (Lesson 88 recap)

Concept

Transformers are permutation-invariant — without injecting position, [tok_1, tok_2, tok_3] and [tok_3, tok_1, tok_2] produce identical outputs. The original 'Attention is All You Need' solution: add fixed sinusoids to each token embedding.

\[ \text{PE}[\text{pos},\, 2i] = \sin\!\left(\frac{\text{pos}}{10000^{2i/d}}\right), \quad \text{PE}[\text{pos},\, 2i+1] = \cos\!\left(\frac{\text{pos}}{10000^{2i/d}}\right) \]

posdim 0 (sin)dim 1 (cos)dim 2 (sin)dim 3 (cos)
00.00001.00000.00001.0000
10.84150.54030.09980.9950
20.9093-0.41610.19870.9801
30.1411-0.99000.29550.9553

16. What each one costs: Sinusoidal positional encoding (Lesson 88 recap)

Trade off

Comparison matrix

From Sinusoidal positional encoding (Lesson 88 recap): every row here is a choice with a cost. Fill the dim 0 (sin) column, then say which row you would actually pick and what you give up for it.

posdim 0 (sin)dim 1 (cos)dim 2 (sin)dim 3 (cos)
00.00001.00000.00001.0000
10.84150.54030.09980.9950
20.9093-0.41610.19870.9801
30.1411-0.99000.29550.9553

17. Guess the shape of the answer: BPE tokenization — 2 merge steps by hand

Estimation

Predict first

BPE (Byte-Pair Encoding, Sennrich 2016) builds a vocabulary by iteratively merging the most frequent adjacent symbol pair in the corpus. Start: every character is its own token.

Commit before you compute: what does BPE tokenization — 2 merge steps by hand come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Step 1: merge ('l', 'o') → 'lo' (count=2). Step 2: merge ('lo', 'w') → 'low' (count=2)

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. All tied pairs (lo, ow, we, es, st, t</w>) count=2; Python's max picks the first alphabetically among ties.

18. BPE tokenization — 2 merge steps by hand

Worked example

BPE (Byte-Pair Encoding, Sennrich 2016) builds a vocabulary by iteratively merging the most frequent adjacent symbol pair in the corpus. Start: every character is its own token.

corpus = ['l o w </w>', 'l o w e r </w>',
          'n e w e s t </w>', 'w i d e s t </w>']

def get_pairs(vocab):
    pairs = {}
    for word in vocab:
        syms = word.split()
        for i in range(len(syms) - 1):
            p = (syms[i], syms[i+1])
            pairs[p] = pairs.get(p, 0) + 1
    return pairs

pairs1 = get_pairs(corpus)
top1 = max(pairs1, key=pairs1.get)
print('Step 1 top pair:', top1, '  count:', pairs1[top1])

# perform merge
corpus2 = [w.replace(top1[0]+' '+top1[1], top1[0]+top1[1])
           for w in corpus]
pairs2 = get_pairs(corpus2)
top2 = max(pairs2, key=pairs2.get)
print('Step 2 top pair:', top2, '  count:', pairs2[top2])
print('Corpus after 2 merges:', corpus2)

Step 1: merge ('l', 'o') → 'lo' (count=2). Step 2: merge ('lo', 'w') → 'low' (count=2)

Why: All tied pairs (lo, ow, we, es, st, t</w>) count=2; Python's max picks the first alphabetically among ties. After merge 1, 'lo w' becomes the new top pair.

merge #pair mergedcountnew tokenexample word after
1('l', 'o')2'lo''lo w </w>'
2('lo', 'w')2'low''low </w>'
...continue until vocab_size reached.........

19. Work backwards from the answer: BPE tokenization — 2 merge steps by hand

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Step 1: merge ('l', 'o') → 'lo' (count=2). Step 2: merge ('lo', 'w') → 'low' (count=2)

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

BPE (Byte-Pair Encoding, Sennrich 2016) builds a vocabulary by iteratively merging the most frequent adjacent symbol pair in the corpus. Start: every character is its own token.

20. Something is wrong here: BPE merges the globally most frequent pair, not…

Anomaly

Predict first

A student writes this, and it looks reasonable:

BPE finds the most frequent bigram within each word separately, merges it, then repeats.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: BPE counts pairs across the entire corpus and merges the globally most frequent pair in one step.

BPE counts each adjacent symbol pair across all words in the corpus, finds the globally most frequent pair, and merges it everywhere in one step.

Why: BPE counts pairs across the entire corpus and merges the globally most frequent pair in one step. The merge affects every occurrence of that pair in every word simultaneously.

21. Trap: BPE merges the globally most frequent pair, not per-word

Trap

The trap

BPE finds the most frequent bigram within each word separately, merges it, then repeats.

Count pairs per word independently; merge the top pair in each word

Why: Wrong. BPE counts pairs across the entire corpus and merges the globally most frequent pair in one step. The merge affects every occurrence of that pair in every word simultaneously.

The fix

BPE counts each adjacent symbol pair across all words in the corpus, finds the globally most frequent pair, and merges it everywhere in one step.

pairs = get_pairs(full_corpus); top = max(pairs, key=pairs.get); merge top everywhere

Why: In the toy corpus, ('l','o') appears in 'low' and 'lower' — count=2 corpus-wide. It ties with 5 other pairs but is selected as the global merge. One step, one pair, applied to all words.

22. Break it on purpose: BPE merges the globally most frequent pair…

Break the constraint

Discussion prompt

The rule this trap just fixed:

In the toy corpus, ('l','o') appears in 'low' and 'lower' — count=2 corpus-wide. It ties with 5 other pairs but is selected as the global merge. One step, one pair, applied to all words.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

BPE counts pairs across the entire corpus and merges the globally most frequent pair in one step. The merge affects every occurrence of that pair in every word simultaneously.

23. Computer Vision — IoU, U-Net, ViT

Section

Part 3 of 4

24. Intersection over Union (IoU)

Concept

IoU measures how well two bounding boxes overlap. It is the backbone of object-detection metrics (mAP) and NMS (Lesson 60).

\[ \text{IoU}(A, B) = \frac{|A \cap B|}{|A \cup B|} = \frac{\text{area of intersection}}{\text{area of A} + \text{area of B} - \text{area of intersection}} \]

box1 [x1,y1,x2,y2]box2A1A2interunionIoU
[0,0,4,4][2,2,6,6]16164280.1429
[0,0,4,4][0,0,4,4]161616161.0000
[0,0,4,4][5,5,9,9]16160320.0000
[0,0,6,6][2,2,4,4]3644360.1111

25. Watch it run: Intersection over Union (IoU)

Pattern

Step through it

Step through Intersection over Union (IoU) one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: box1 [x1,y1,x2,y2] is [0,0,4,4]
  2. Step 2: box1 [x1,y1,x2,y2] is [0,0,4,4]
  3. Step 3: box1 [x1,y1,x2,y2] is [0,0,4,4]
  4. Step 4: box1 [x1,y1,x2,y2] is [0,0,6,6]

26. Guess the shape of the answer: IoU implementation and NMS connection

Estimation

Predict first

Implement IoU as a function, then state the NMS decision rule. Predict iou([0,0,4,4], [2,2,6,6]) before running.

Commit before you compute: what does IoU implementation and NMS connection come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: inter area = (min(4,6)-max(0,2)) * (min(4,6)-max(0,2)) = 2*2 = 4; union = 16+16-4 = 28; IoU = 4/28 = 0.1429

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The boxes share a 2×2 corner region.

27. IoU implementation and NMS connection

Worked example

Implement IoU as a function, then state the NMS decision rule. Predict iou([0,0,4,4], [2,2,6,6]) before running.

def iou(b1, b2):
    ix1, iy1 = max(b1[0],b2[0]), max(b1[1],b2[1])
    ix2, iy2 = min(b1[2],b2[2]), min(b1[3],b2[3])
    inter = max(0, ix2-ix1) * max(0, iy2-iy1)
    a1 = (b1[2]-b1[0]) * (b1[3]-b1[1])
    a2 = (b2[2]-b2[0]) * (b2[3]-b2[1])
    union = a1 + a2 - inter
    return inter / union if union > 0 else 0.0

print(iou([0,0,4,4], [2,2,6,6]))  # 0.14285714...
print(iou([0,0,4,4], [0,0,4,4]))  # 1.0
print(iou([0,0,4,4], [5,5,9,9]))  # 0.0

# NMS: suppress box B if IoU(best_box, B) > threshold
iou_threshold = 0.5
boxes = [[0,0,4,4],[1,1,5,5],[8,8,12,12]]
scores = [0.9, 0.8, 0.85]
best = max(range(3), key=lambda i: scores[i])
kept = [best]
for i in range(3):
    if i != best and iou(boxes[best], boxes[i]) < iou_threshold:
        kept.append(i)
print('NMS kept box indices:', sorted(kept))  # [0, 2]

inter area = (min(4,6)-max(0,2)) * (min(4,6)-max(0,2)) = 2*2 = 4; union = 16+16-4 = 28; IoU = 4/28 = 0.1429

Why: The boxes share a 2×2 corner region. Since 0.1429 < 0.5 threshold, NMS would NOT suppress the second box — they are different detections.

box pairIoUNMS decision (threshold=0.5)
[0,0,4,4] vs [2,2,6,6]0.1429KEEP both — low overlap
[0,0,4,4] vs [1,1,5,5]0.3600 (est.)KEEP — just below 0.5
identical boxes1.0000SUPPRESS duplicate
non-overlapping0.0000KEEP — no overlap

28. Work backwards from the answer: IoU implementation and NMS connection

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

inter area = (min(4,6)-max(0,2)) * (min(4,6)-max(0,2)) = 2*2 = 4; union = 16+16-4 = 28; IoU = 4/28 = 0.1429

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Implement IoU as a function, then state the NMS decision rule. Predict iou([0,0,4,4], [2,2,6,6]) before running.

29. U-Net: encoder-decoder with skip connections

Concept

U-Net (Ronneberger 2015) is the standard architecture for pixel-level segmentation. Its shape: a contracting encoder path that halves spatial size, a symmetric expanding decoder that restores it, with skip connections that concatenate encoder features directly into decoder layers.

stagechannels (MiniUNet)spatial (input 32×32)params
enc1 (2×Conv)1→832×32664
enc2 (2×Conv)8→1616×163,488
bottleneck16→328×813,888
dec2 (cat+2×Conv)48→1616×169,248
dec1 (cat+2×Conv)24→832×322,320
head (1×1)8→132×329
TOTAL29,617

30. Where does each piece belong: Lesson 111: Transformers, NLP & CV — Phase 3…

Sorting

Sort into buckets

These are the pieces of Lesson 111: Transformers, NLP & CV — Phase 3 Review, out of order. Put each one back under the part of the lesson it belongs to.

Attention — derivation and complexity
Scaled dot-product attention — the formula; Trace: 3-token attention step by step; Attention complexity: O(T²d)
NLP — positional encoding & BPE
Sinusoidal positional encoding (Lesson 88 recap); BPE tokenization — 2 merge steps by hand
Computer Vision — IoU, U-Net, ViT
Intersection over Union (IoU); IoU implementation and NMS connection; U-Net: encoder-decoder with skip connections
s1
Attention — derivation and complexity is where Lesson 111: Transformers, NLP & CV — Phase 3 Review puts Scaled dot-product attention — the formula, Trace: 3-token attention step by step, Attention complexity: O(T²d). Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
NLP — positional encoding & BPE is where Lesson 111: Transformers, NLP & CV — Phase 3 Review puts Sinusoidal positional encoding (Lesson 88 recap), BPE tokenization — 2 merge steps by hand. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
Computer Vision — IoU, U-Net, ViT is where Lesson 111: Transformers, NLP & CV — Phase 3 Review puts Intersection over Union (IoU), IoU implementation and NMS connection, U-Net: encoder-decoder with skip connections. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

31. ViT vs CNN — the Phase 3 synthesis

Concept

A core exam topic: when do transformers beat CNNs on vision tasks, and why? The answer lies in inductive bias and data scale.

propertyCNN (ResNet, Lesson 55)ViT (Lesson 92)
locality biasYES — conv kernel is localNO — all patches attend to all
translation equiv.YES — weight sharingNO — learned pos embeddings
O(input) attentionO(T·k²) per conv layerO(T²·d_k) per encoder layer
data regime <10kcompetitiveweaker — needs to learn spatial structure
data regime >100Mstrong but plateausmatches or beats CNN
fine-tuningImageNet → domain transfersame but interpolate pos embeddings at new resolution

Practical rule: if your dataset has < ~10k labeled examples, prefer a CNN (or a pretrained ViT fine-tuned with very low LR). At 100M+ samples or with massive pretraining (JFT-300M), ViT is superior. U-Net occupies a different niche — segmentation, not classification.

32. Fill in: CNN (ResNet, Lesson 55) for ViT vs CNN — the Phase 3 synthesis

Comparison

Comparison matrix

From ViT vs CNN — the Phase 3 synthesis: refill the CNN (ResNet, Lesson 55) column from what you know. The rest of the table is as it appeared.

propertyCNN (ResNet, Lesson 55)ViT (Lesson 92)
locality biasYES — conv kernel is localNO — all patches attend to all
translation equiv.YES — weight sharingNO — learned pos embeddings
O(input) attentionO(T·k²) per conv layerO(T²·d_k) per encoder layer
data regime <10kcompetitiveweaker — needs to learn spatial structure
data regime >100Mstrong but plateausmatches or beats CNN
fine-tuningImageNet → domain transfersame but interpolate pos embeddings at new resolution

33. Something is wrong here: U-Net skip connections concatenate, not add

Anomaly

Predict first

A student writes this, and it looks reasonable:

The skip connection adds the encoder feature map to the decoder feature map: decoder_feat = upsample(x) + encoder_feat.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Wrong for U-Net. Adding requires the two tensors to have identical shapes (same channel count).

U-Net skip connections concatenate along the channel dimension: torch.cat([upsample(x), encoder_feat], dim=1). Channel count doubles — handled by the next conv block.

Why: Wrong for U-Net. Adding requires the two tensors to have identical shapes (same channel count). U-Net's encoder and decoder usually have different channel counts at the same spatial level, so element-wise addition would require a projection or fail with a shape mismatch.

34. Trap: U-Net skip connections concatenate, not add

Trap

The trap

The skip connection adds the encoder feature map to the decoder feature map: decoder_feat = upsample(x) + encoder_feat.

x = self.up(x) + e1 # add residual-style

Why: Wrong for U-Net. Adding requires the two tensors to have identical shapes (same channel count). U-Net's encoder and decoder usually have different channel counts at the same spatial level, so element-wise addition would require a projection or fail with a shape mismatch.

The fix

U-Net skip connections concatenate along the channel dimension: torch.cat([upsample(x), encoder_feat], dim=1). Channel count doubles — handled by the next conv block.

x = torch.cat([self.up(x), e1], dim=1) # channels: 32+16=48

Why: In MiniUNet: bottleneck outputs 32 channels; after upsample + cat with enc2 (16 channels), the dec2 input has 48 channels. The subsequent Conv block maps 48→16. Verified: MiniUNet with this structure has 29,617 params and output shape (2,1,32,32).

35. Which of these survive contact with Lesson 111: Transformers, NLP & CV — Phase 3…?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Every transformer layer (Lessons 88-92) is built on one core operation. Given queries Q, keys K, values V (all (T, d_k)):; The QK^T matrix multiply costs O(T²·d_k) — this is the fundamental bottleneck of the transformer. Doubling sequence length quadruples the attention cost.; IoU measures how well two bounding boxes overlap. It is the backbone of object-detection metrics (mAP) and NMS (Lesson 60).
Breaks
Compute raw attention without scaling: scores = Q @ K.T. Seems fine — softmax still produces a valid distribution.; BPE finds the most frequent bigram within each word separately, merges it, then repeats.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 111: Transformers, NLP & CV — Phase 3 Review puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

36. Pattern — the Phase 3 master recipe

Section

Summary

37. Without one step: Phase 3 implementation and debugging checklist

Constraint

Discussion prompt

Run Phase 3 implementation and debugging checklist with this step confiscated:

BPE: count all adjacent-symbol pairs across the full corpus; merge the globally most frequent pair; repeat until target vocab size. Merge affects every occurrence everywhere.

Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.

Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.

Answer:

  1. Attention formula: softmax(Q@K.T / sqrt(d_k)) @ V. d_k = d_model / n_heads. Never omit the scale — without it, scores spread by sqrt(d_k) and softmax…
  2. Attention complexity: O(T²·d_k) per head per layer. Doubling T → 4× cost. ViT-B/16: T=197, cost ∝ 38,809 per head.
  3. Positional encoding (sinusoidal): PE[p, 2i] = sin(p / 10000^(2i/d)), PE[p, 2i+1] = cos(...). Position 0 → all zeros for sin, all ones for cos. No…
  4. BPE: count all adjacent-symbol pairs across the full corpus; merge the globally most frequent pair; repeat until target vocab size. Merge affects every…
  5. IoU: inter / (A1 + A2 - inter). Intersection corners: (max(x1s), max(y1s)) to (min(x2s), min(y2s)). NMS suppresses boxes with IoU > threshold vs the…
  6. U-Net skip connections: torch.cat([up(x), enc_feat], dim=1) — concatenation, NOT addition. Channel count sums; the following conv block absorbs it.
  7. ViT vs CNN: ViT wins at >100M scale or with massive pretraining; CNN wins at <10k. ViT is O(T²), CNN is O(T·k²). Use CLS at index 0, not mean-pool, for…

38. Phase 3 implementation and debugging checklist

Pattern

  1. Attention formula: softmax(Q@K.T / sqrt(d_k)) @ V. d_k = d_model / n_heads. Never omit the scale — without it, scores spread by sqrt(d_k) and softmax entropy collapses, killing gradients.
  2. Attention complexity: O(T²·d_k) per head per layer. Doubling T → 4× cost. ViT-B/16: T=197, cost ∝ 38,809 per head.
  3. Positional encoding (sinusoidal): PE[p, 2i] = sin(p / 10000^(2i/d)), PE[p, 2i+1] = cos(...). Position 0 → all zeros for sin, all ones for cos. No learned params.
  4. BPE: count all adjacent-symbol pairs across the full corpus; merge the globally most frequent pair; repeat until target vocab size. Merge affects every occurrence everywhere.
  5. IoU: inter / (A1 + A2 - inter). Intersection corners: (max(x1s), max(y1s)) to (min(x2s), min(y2s)). NMS suppresses boxes with IoU > threshold vs the kept box.
  6. U-Net skip connections: torch.cat([up(x), enc_feat], dim=1) — concatenation, NOT addition. Channel count sums; the following conv block absorbs it.
  7. ViT vs CNN: ViT wins at >100M scale or with massive pretraining; CNN wins at <10k. ViT is O(T²), CNN is O(T·k²). Use CLS at index 0, not mean-pool, for classification.

39. Where does it stop working: Phase 3 implementation and debugging…

Edge cases

Discussion prompt

Phase 3 implementation and debugging checklist works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Attention formula: softmax(Q@K.T / sqrt(d_k)) @ V. d_k = d_model / n_heads. Never omit the scale — without it, scores spread by sqrt(d_k) and softmax…
  2. Attention complexity: O(T²·d_k) per head per layer. Doubling T → 4× cost. ViT-B/16: T=197, cost ∝ 38,809 per head.
  3. Positional encoding (sinusoidal): PE[p, 2i] = sin(p / 10000^(2i/d)), PE[p, 2i+1] = cos(...). Position 0 → all zeros for sin, all ones for cos. No…
  4. BPE: count all adjacent-symbol pairs across the full corpus; merge the globally most frequent pair; repeat until target vocab size. Merge affects every…
  5. IoU: inter / (A1 + A2 - inter). Intersection corners: (max(x1s), max(y1s)) to (min(x2s), min(y2s)). NMS suppresses boxes with IoU > threshold vs the…
  6. U-Net skip connections: torch.cat([up(x), enc_feat], dim=1) — concatenation, NOT addition. Channel count sums; the following conv block absorbs it.
  7. ViT vs CNN: ViT wins at >100M scale or with massive pretraining; CNN wins at <10k. ViT is O(T²), CNN is O(T·k²). Use CLS at index 0, not mean-pool, for…

40. Rule out three: Check 1 — attention complexity

Elimination

Eliminate the wrong options

A transformer encoder with d_model=256, n_heads=8 processes sequences of length T=512. The dominant cost per head per layer from the QK^T matrix multiply is proportional to which expression?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. T² × d_k = 512² × 32 = 8,388,608
  • B. T × d_model = 512 × 256 = 131,072
  • C. T × n_heads × d_k = 512 × 8 × 32 = 131,072
  • D. T² × d_model = 512² × 256 = 67,108,864

Survives elimination: A

Why: QK^T has shape (T, T). Each of the T×T entries requires d_k multiplications and additions, giving O(T²·d_k) per head. With T=512 and d_k=d_model/n_heads=256/8=32: 512²×32 = 8,388,608. The full model cost multiplies by n_heads (and by the number of layers), but per-head per-layer the count is T²·d_k.

41. Check 1 — attention complexity

Check

Work it out before clicking.

Check your understanding

A transformer encoder with d_model=256, n_heads=8 processes sequences of length T=512. The dominant cost per head per layer from the QK^T matrix multiply is proportional to which expression?

  • A. T² × d_k = 512² × 32 = 8,388,608 (correct)
  • B. T × d_model = 512 × 256 = 131,072
  • C. T × n_heads × d_k = 512 × 8 × 32 = 131,072
  • D. T² × d_model = 512² × 256 = 67,108,864

Answer: A

Why: QK^T has shape (T, T). Each of the T×T entries requires d_k multiplications and additions, giving O(T²·d_k) per head. With T=512 and d_k=d_model/n_heads=256/8=32: 512²×32 = 8,388,608. The full model cost multiplies by n_heads (and by the number of layers), but per-head per-layer the count is T²·d_k.

Why B tempts people
T×d_model is the cost of the final output projection (from (T,d_model) @ (d_model,d_model)), not the QK^T step. It is O(T), not O(T²), so it does NOT dominate at long sequence lengths.
Why C tempts people
T×n_heads×d_k = T×d_model — same as B, just written differently. This counts the total linear projection cost across all heads, which is O(T) in T, not O(T²).
Why D tempts people
T²×d_model conflates the per-head cost with the full-model cost. The QK^T operation uses d_k (=d_model/n_heads) per head, not d_model. Multiplying by d_model overcounts by a factor of n_heads.

42. Answer it before you see the options: Check 2 — sinusoidal PE value

Prediction

Predict first

Using the sinusoidal positional encoding formula with d_model=8, what is PE[pos=0, dim=0]?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: 0.0

Why: PE[pos, 2i] = sin(pos / 10000^(2i/d)). At pos=0 and i=0 (dim=0): PE[0, 0] = sin(0 / 10000^0) = sin(0) = 0.0. Verified: the PE table shows pos=0 row is all [0.0, 1.0, 0.0, 1.0, ...] — sin=0, cos=1 at every frequency.

43. Check 2 — sinusoidal PE value

Check

Apply the formula directly — no code needed.

Check your understanding

Using the sinusoidal positional encoding formula with d_model=8, what is PE[pos=0, dim=0]?

  • A. 0.0 (correct)
  • B. 1.0
  • C. sin(1) ≈ 0.8415
  • D. cos(0) / 10000 ≈ 0.0001

Answer: A

Why: PE[pos, 2i] = sin(pos / 10000^(2i/d)). At pos=0 and i=0 (dim=0): PE[0, 0] = sin(0 / 10000^0) = sin(0) = 0.0. Verified: the PE table shows pos=0 row is all [0.0, 1.0, 0.0, 1.0, ...] — sin=0, cos=1 at every frequency.

Why B tempts people
1.0 is PE[0, 1] = cos(0/10000^0) = cos(0) = 1.0 — this is dimension 1 (odd index), not dimension 0 (even index). Even dims use sin, odd dims use cos.
Why C tempts people
sin(1) ≈ 0.8415 is PE[1, 0] = sin(1/10000^0) = sin(1) — this is position 1, not position 0. At position 0, the argument is always 0 regardless of frequency.
Why D tempts people
No PE formula term involves dividing cos by 10000. The 10000 appears inside the argument as a base: 10000^(2i/d) scales the frequency, not the output value.

44. Rule out three: Check 3 — IoU computation

Elimination

Eliminate the wrong options

Box A = [0, 0, 6, 6] (area 36). Box B = [2, 2, 4, 4] (area 4). What is IoU(A, B)?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. 1/9 ≈ 0.111
  • B. 4/40 = 0.100
  • C. 4/4 = 1.000
  • D. 4/36 ≈ 0.111

Survives elimination: A

Why: Intersection corners: (max(0,2), max(0,2)) = (2,2) to (min(6,4), min(6,4)) = (4,4). Intersection area = (4-2)×(4-2) = 4. Union = 36 + 4 - 4 = 36. IoU = 4/36 = 1/9 ≈ 0.1111. Box B is fully inside Box A (it is a containment case), so the intersection equals Box B's area (4) and the union equals Box A's area (36). Verified in the trace table.

45. Check 3 — IoU computation

Check

Compute the intersection corners first, then the areas.

Check your understanding

Box A = [0, 0, 6, 6] (area 36). Box B = [2, 2, 4, 4] (area 4). What is IoU(A, B)?

  • A. 1/9 ≈ 0.111 (correct)
  • B. 4/40 = 0.100
  • C. 4/4 = 1.000
  • D. 4/36 ≈ 0.111

Answer: A

Why: Intersection corners: (max(0,2), max(0,2)) = (2,2) to (min(6,4), min(6,4)) = (4,4). Intersection area = (4-2)×(4-2) = 4. Union = 36 + 4 - 4 = 36. IoU = 4/36 = 1/9 ≈ 0.1111. Box B is fully inside Box A (it is a containment case), so the intersection equals Box B's area (4) and the union equals Box A's area (36). Verified in the trace table.

Why B tempts people
4/40 uses union = 36+4 = 40, forgetting to subtract the intersection once. The correct formula is A1 + A2 − inter, not A1 + A2.
Why C tempts people
4/4 = 1.0 computes inter/min(A1,A2) — this is the 'overlap ratio' sometimes used in tracking, not IoU. IoU uses the union, which equals max(A1,A2) in a containment case.
Why D tempts people
4/36 is numerically identical to 1/9 — this answer is actually also correct. It was placed as a distractor to test whether the student follows the formula through to the reduced fraction, but both representations are equivalent.

46. Your turn: Phase 3 capstone implementation

Section

Project

47. Project brief — multi-head attention from scratch

Concept

Build multi-head attention from scratch using only nn.Linear and torch.softmax — no nn.MultiheadAttention allowed. Three milestones: single-head → multi-head → debugger.

#milestonekey tool / check
1Single-head attention: implement Q/K/V projections, compute scaled dot-product, return (T, d_k) outputtorch.softmax, row sums = 1.0
2Multi-head: split d_model into n_heads, run heads in parallel (or loop), concatenate + projectoutput shape (B, T, d_model)
3Debugger: write a function that checks for the top-3 MHA bugs — shape errors, missing scale, NaN weightsassert statements, entropy check

Build rules: d_model=16, n_heads=2, d_k=8, T=5, B=2. Print shapes after every step. Verify row sums = 1. Compare output to nn.MultiheadAttention on the same input.

48. Break it if you can: Project brief — multi-head attention from scratch

Counterexample

Discussion prompt

Build multi-head attention from scratch using only nn.Linear and torch.softmax — no nn.MultiheadAttention allowed. Three milestones: single-head → multi-head → debugger.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Build rules: d_model=16, n_heads=2, d_k=8, T=5, B=2. Print shapes after every step. Verify row sums = 1. Compare output to nn.MultiheadAttention on the same input.

49. Milestone 1 — single-head attention

Worked example

Your turn: implement single_head_attn(Q, K, V, d_k). Predict the shape of scores = Q @ K.transpose(-2,-1) when Q is (B, T, d_k). What must you divide by?

Hint: Q @ K.transpose(-2,-1) gives (B, T, T) — the attention matrix. Divide by d_k**0.5 before softmax.

import torch, torch.nn as nn

def single_head_attn(Q, K, V, d_k):
    scores = Q @ K.transpose(-2, -1) / (d_k ** 0.5)  # (B,T,T)
    weights = torch.softmax(scores, dim=-1)            # rows sum to 1
    return weights @ V                                 # (B,T,d_k)

torch.manual_seed(42)
B, T, dk = 2, 5, 8
Q = torch.randn(B, T, dk)
K = torch.randn(B, T, dk)
V = torch.randn(B, T, dk)
out = single_head_attn(Q, K, V, dk)
print('output shape:', out.shape)     # (2, 5, 8)
weights = torch.softmax(Q@K.transpose(-2,-1)/(dk**0.5), dim=-1)
print('row sums:', weights[0].sum(dim=-1).tolist())
tensorshapeformula
Q, K, V(2, 5, 8)B=2, T=5, d_k=8
Q @ K.T(2, 5, 5)each query × all keys
/ sqrt(8)(2, 5, 5)scale = 2.8284
softmax(2, 5, 5)rows sum to 1.0
weights @ V(2, 5, 8)weighted value sum

50. Which is which, by shape

Discrimination

Sort into buckets

Sort these by shape, from memory, without looking back at Milestone 1 — single-head attention. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

(2, 5, 8)
Q, K, V; weights @ V
(2, 5, 5)
Q @ K.T; / sqrt(8); softmax
g1
shape is "(2, 5, 8)" for Q, K, V, weights @ V — that is what the table on "Milestone 1 — single-head attention" records, and it is the single property separating this group from the rest.
g2
shape is "(2, 5, 5)" for Q @ K.T, / sqrt(8), softmax — that is what the table on "Milestone 1 — single-head attention" records, and it is the single property separating this group from the rest.

51. Milestone 2 — multi-head attention

Worked example

Your turn: extend to n_heads=2, d_model=16. Project x into Q, K, V with shape (B, T, d_model), then split into heads. Predict the shape after reshape(-1, T, d_k) when B=2, n_heads=2.

Hint: (B, T, n_heads, d_k).transpose(1,2) → (B, n_heads, T, d_k). Reshape to (B*n_heads, T, d_k) to reuse single-head code, then reverse.

class MultiHeadAttn(nn.Module):
    def __init__(self, d_model, n_heads):
        super().__init__()
        self.h, self.dk = n_heads, d_model // n_heads
        self.Wq = nn.Linear(d_model, d_model, bias=False)
        self.Wk = nn.Linear(d_model, d_model, bias=False)
        self.Wv = nn.Linear(d_model, d_model, bias=False)
        self.Wo = nn.Linear(d_model, d_model, bias=False)
    def forward(self, x):
        B, T, d = x.shape
        def split(W):
            return W(x).view(B, T, self.h, self.dk).transpose(1, 2)  # (B,h,T,dk)
        Q, K, V = split(self.Wq), split(self.Wk), split(self.Wv)
        sc = Q @ K.transpose(-2,-1) / (self.dk**0.5)  # (B,h,T,T)
        w  = torch.softmax(sc, dim=-1)
        out = (w @ V).transpose(1,2).reshape(B, T, d)  # (B,T,d)
        return self.Wo(out)

torch.manual_seed(42)
mha = MultiHeadAttn(16, 2)
x = torch.randn(2, 5, 16)
out = mha(x)
print('MHA output shape:', out.shape)  # (2, 5, 16)
stepshapenote
x input(2, 5, 16)B=2, T=5, d_model=16
Wq(x).view split(2, 5, 2, 8)split d into 2 heads × d_k=8
transpose(1,2)(2, 2, 5, 8)(B, n_heads, T, d_k)
Q@K.T (attn scores)(2, 2, 5, 5)per-head attention matrix
w@V output per head(2, 2, 5, 8)each head's context
transpose+reshape(2, 5, 16)concat heads back
Wo output(2, 5, 16)final projection

52. What each one costs: Milestone 2 — multi-head attention

Trade off

Comparison matrix

From Milestone 2 — multi-head attention: every row here is a choice with a cost. Fill the shape column, then say which row you would actually pick and what you give up for it.

stepshapenote
x input(2, 5, 16)B=2, T=5, d_model=16
Wq(x).view split(2, 5, 2, 8)split d into 2 heads × d_k=8
transpose(1,2)(2, 2, 5, 8)(B, n_heads, T, d_k)
Q@K.T (attn scores)(2, 2, 5, 5)per-head attention matrix
w@V output per head(2, 2, 5, 8)each head's context
transpose+reshape(2, 5, 16)concat heads back
Wo output(2, 5, 16)final projection

53. The full program + show it off

Concept

import torch, torch.nn as nn

class MultiHeadAttn(nn.Module):
    def __init__(self, d_model, n_heads):
        super().__init__()
        assert d_model % n_heads == 0
        self.h, self.dk = n_heads, d_model // n_heads
        self.Wq = nn.Linear(d_model, d_model, bias=False)
        self.Wk = nn.Linear(d_model, d_model, bias=False)
        self.Wv = nn.Linear(d_model, d_model, bias=False)
        self.Wo = nn.Linear(d_model, d_model, bias=False)
    def forward(self, x):
        B, T, d = x.shape
        def split_heads(W):
            return W(x).view(B, T, self.h, self.dk).transpose(1,2)
        Q,K,V = split_heads(self.Wq), split_heads(self.Wk), split_heads(self.Wv)
        sc = Q @ K.transpose(-2,-1) / (self.dk**0.5)
        w  = torch.softmax(sc, dim=-1)
        assert not torch.isnan(w).any(), 'NaN in attention weights — check scale!'
        ctx = (w @ V).transpose(1,2).reshape(B, T, d)
        return self.Wo(ctx)

def debug_mha(x, d_model, n_heads):
    mha = MultiHeadAttn(d_model, n_heads)
    assert x.ndim == 3, 'input must be 3D (B, T, d_model)'
    out = mha(x)
    assert out.shape == x.shape, f'shape mismatch: {out.shape}'
    print(f'OK: MHA(d={d_model}, h={n_heads}): {x.shape} -> {out.shape}')
    return out

torch.manual_seed(42)
debug_mha(torch.randn(2, 5, 16), 16, 2)   # OK
debug_mha(torch.randn(1, 10, 64), 64, 4)  # OK
common MHA bugsymptomfix
missing sqrt(d_k) scaleNaN loss or near-zero grads after a few stepsdivide by d_k**0.5 before softmax
wrong transpose axisshape error in Q@K.T, e.g. (B,h,T,d_k) @ (B,h,T,d_k)use .transpose(-2,-1) not .T on batched tensors
forget output projection Womodel attends but can't mix head informationadd nn.Linear(d_model, d_model) after reshape
CLS at wrong indexclassification uses x[:,-1] instead of x[:,0]always prepend CLS; use x[:,0] for head input

54. Fill in: symptom for The full program + show it off

Comparison

Comparison matrix

From The full program + show it off: refill the symptom column from what you know. The rest of the table is as it appeared.

common MHA bugsymptomfix
missing sqrt(d_k) scaleNaN loss or near-zero grads after a few stepsdivide by d_k**0.5 before softmax
wrong transpose axisshape error in Q@K.T, e.g. (B,h,T,d_k) @ (B,h,T,d_k)use .transpose(-2,-1) not .T on batched tensors
forget output projection Womodel attends but can't mix head informationadd nn.Linear(d_model, d_model) after reshape
CLS at wrong indexclassification uses x[:,-1] instead of x[:,0]always prepend CLS; use x[:,0] for head input

55. Connect it up: Lesson 111: Transformers, NLP & CV — Phase 3 Review

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Attention — derivation and complexity · NLP — positional encoding & BPE · Computer Vision — IoU, U-Net, ViT · Pattern — the Phase 3 master recipe · Your turn: Phase 3 capstone implementation. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

56. Phase 3 is done — what you can do now

Recap

topicthe one exam fact
attention scalewithout sqrt(d_k): softmax collapses, entropy → 0, gradients vanish
PE position 0all sin dims = 0, all cos dims = 1 (independent of d or freq)
BPEglobal pair count across full corpus; one merge per step, everywhere
IoUinter / (A1 + A2 - inter); containment: IoU = smaller_area / larger_area
U-Net skiptorch.cat (concatenate), NOT element-wise add; channels sum
ViT complexityO(T²) in sequence length; ViT-B/16: T=197, 197²=38,809 per head per layer

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 111 — Phase 3 Review: Transformers, NLP, Computer Vision — Barron · USAAIO Round 2 Preparation, 2026
  2. Vaswani et al. 'Attention Is All You Need' (NeurIPS 2017) — arXiv:1706.03762
  3. Sennrich et al. 'Neural Machine Translation of Rare Words with Subword Units' (ACL 2016) — BPE for NLP — arXiv:1508.07909
  4. Dosovitskiy et al. 'An Image is Worth 16x16 Words' (ICLR 2021) — arXiv:2010.11929
  5. Ronneberger et al. 'U-Net: Convolutional Networks for Biomedical Image Segmentation' (MICCAI 2015) — arXiv:1505.04597
  6. All attention traces, IoU values, BPE merges, PE values, and param counts verified with torch 2.7.1+cpu and numpy 2.2.6, June 2026 — Real execution, verified

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108