USAAIO Lesson 116, a Phase 3 review. It is a comprehensive exam-simulation deck across NLP and computer vision, covering tokenization, word2vec, BERT, LoRA, seq2seq, beam search, BLEU, and ROUGE on the language side, and convolution, ResNet, ViT, U-Net, YOLO, IoU, NMS, and CLIP on the vision side. All the metrics were computed by real Python execution with torch 2.7.1+cpu and numpy 2.2.6. It emphasizes the cross-modal connections that transformers make, pattern recognition for exam pace, and spotting the gaps you need to close before Phase 4. The lesson runs to 30 slides.
Subject: Machine Learning · 56 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 116 · Phase 3 Review
Consolidate every Phase 3 topic — tokenization through CLIP, convolution through YOLO — at exam pace. Identify the weakest two subtopics before Phase 4 begins.
Objectives
Section
Part 1 of 4
Concept
| lesson | topic | the one number to remember |
|---|---|---|
| L88 | Tokenization / BPE | vocab size ≈ 30k–50k; OOV→subword |
| L89 | Word2Vec skip-gram | cos_sim(king−man+woman, queen) ≈ 0.9 |
| L90 | BERT fine-tuning | 12 layers, 110M params, MLM 15% mask rate |
| L91 | LoRA | rank r; trainable params = 2·r·d; frozen W + B·A |
| L92 | Seq2seq + attention | context = Σ α_i h_i, Σ α_i = 1 |
| L93 | Beam search | beam_width=k; keeps top-k at each step |
| L94 | BLEU / ROUGE | BLEU-1 perfect=1.0, partial≈0.67; ROUGE-L recall-based |
Run your mental circuit: can you derive the number in column 3 from first principles in 90 seconds? That is the exam bar.
Comparison
Comparison matrix
From NLP Phase 3 map at a glance: refill the topic column from what you know. The rest of the table is as it appeared.
| lesson | topic | the one number to remember |
|---|---|---|
| L88 | Tokenization / BPE | vocab size ≈ 30k–50k; OOV→subword |
| L89 | Word2Vec skip-gram | cos_sim(king−man+woman, queen) ≈ 0.9 |
| L90 | BERT fine-tuning | 12 layers, 110M params, MLM 15% mask rate |
| L91 | LoRA | rank r; trainable params = 2·r·d; frozen W + B·A |
| L92 | Seq2seq + attention | context = Σ α_i h_i, Σ α_i = 1 |
| L93 | Beam search | beam_width=k; keeps top-k at each step |
| L94 | BLEU / ROUGE | BLEU-1 perfect=1.0, partial≈0.67; ROUGE-L recall-based |
Concept
BLEU is precision-based: fraction of hypothesis n-grams that appear in the reference, clipped by reference count, multiplied by a brevity penalty (BP). High BLEU → hypothesis text is reference-like.
ROUGE-L is recall-based: length of the longest common subsequence (LCS) divided by reference length. High ROUGE-L → reference content is covered by hypothesis. Used for summarization; BLEU for translation.
| hypothesis | BLEU-1 | BLEU-2 | ROUGE-L F1 |
|---|---|---|---|
| the cat sat on the mat (perfect) | 1.0000 | 1.0000 | 1.0000 |
| a cat sat on a mat (2 words differ) | 0.6667 | 0.4000 | 0.6667 |
| the dog ran in the park (2 words match) | 0.3333 | 0.0000 | 0.3333 |
Trade off
Comparison matrix
From BLEU and ROUGE — what each actually measures: every row here is a choice with a cost. Fill the ROUGE-L F1 column, then say which row you would actually pick and what you give up for it.
| hypothesis | BLEU-1 | BLEU-2 | ROUGE-L F1 |
|---|---|---|---|
| the cat sat on the mat (perfect) | 1.0000 | 1.0000 | 1.0000 |
| a cat sat on a mat (2 words differ) | 0.6667 | 0.4000 | 0.6667 |
| the dog ran in the park (2 words match) | 0.3333 | 0.0000 | 0.3333 |
Estimation
Predict first
Reference: the cat sat on the mat. Hypothesis: a cat sat on a mat. Compute BLEU-1 step by step. Then simulate beam search with beam_width=2 over a toy vocab.
Commit before you compute: what does BLEU-1 and beam search — computed come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: BLEU-1 = 4/6 ≈ 0.6667: 'cat','sat','on','mat' match; 'a','a' do not appear in reference clipped counts
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Reference has 'the'×2,'cat','sat','on','mat'.
Worked example
Reference: the cat sat on the mat. Hypothesis: a cat sat on a mat. Compute BLEU-1 step by step. Then simulate beam search with beam_width=2 over a toy vocab.
from collections import Counter
import math
# BLEU-1: clipped unigram precision x brevity penalty
ref = 'the cat sat on the mat'.split()
hyp = 'a cat sat on a mat'.split()
ref_counts = Counter(ref)
hyp_counts = Counter(hyp)
clipped = {w: min(c, ref_counts.get(w, 0)) for w, c in hyp_counts.items()}
precision = sum(clipped.values()) / len(hyp) # 4/6
bp = 1.0 # same length
bleu1 = bp * precision
print(f'clipped matches: {dict(clipped)}')
print(f'BLEU-1 = {precision:.4f} x {bp:.4f} = {bleu1:.4f}')
# Beam search beam_width=2
step1 = {'the': -0.5, 'cat': -1.2, 'dog': -2.0}
step2 = {'the': {'cat': -0.3, 'dog': -1.5}, 'cat': {'sat': -0.4, 'ran': -0.9}}
beams = sorted([(s,[w]) for w,s in step1.items()], key=lambda x:-x[0])[:2]
candidates = [(s+step2[seq[-1]][w], seq+[w])
for s,seq in beams for w in step2.get(seq[-1],{})]
candidates.sort(key=lambda x: -x[0])
print(f'Beam top-2 after step 2: {[(round(s,2),seq) for s,seq in candidates[:2]]}')BLEU-1 = 4/6 ≈ 0.6667: 'cat','sat','on','mat' match; 'a','a' do not appear in reference clipped counts
Why: Reference has 'the'×2,'cat','sat','on','mat'. Hypothesis 'a' has zero ref count → clipped to 0. Matched: cat(1)+sat(1)+on(1)+mat(1)=4 out of 6.
| step | beam content | score |
|---|---|---|
| init | — | — |
| step 1 top-2 | ['the'], ['cat'] | -0.50, -1.20 |
| step 2 top-2 | ['the','cat'], ['cat','sat'] | -0.80, -1.60 |
| greedy best | ['the','cat'] | -0.80 |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
BLEU-1 = 4/6 ≈ 0.6667: 'cat','sat','on','mat' match; 'a','a' do not appear in reference clipped counts
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Reference: the cat sat on the mat. Hypothesis: a cat sat on a mat. Compute BLEU-1 step by step. Then simulate beam search with beam_width=2 over a toy vocab.
Concept
LoRA (Lesson 91) freezes W ∈ ℝ^{d×d} and adds a low-rank perturbation ΔW = B·A, where B ∈ ℝ^{d×r} and A ∈ ℝ^{r×d}. Only A and B are trained. B is initialized to zero so ΔW=0 at step 0 — the frozen pretrained output is unchanged at initialization.
\[ \text{trainable params} = 2 r d \quad \text{vs} \quad d^2 \text{ (full fine-tune)} \]
| d (hidden dim) | r (rank) | LoRA params | full params | ratio |
|---|---|---|---|---|
| 768 (BERT-base) | 8 | 12,288 | 589,824 | 0.021 |
| 768 | 16 | 24,576 | 589,824 | 0.042 |
| 8 (toy) | 2 | 32 | 64 | 0.500 |
Anomaly
Predict first
A student writes this, and it looks reasonable:
Initialize both A and B with random values so ΔW = B·A starts near zero by chance.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Random A and B give a non-zero delta_W at initialization, so the pretrained model output is perturbed before any training step.
Initialize B to zeros and A to a small random (or Gaussian) so ΔW = B·A = 0 exactly at step 0.
Why: Random A and B give a non-zero delta_W at initialization, so the pretrained model output is perturbed before any training step. This breaks the premise that LoRA is a zero-change start.
Trap
Initialize both A and B with random values so ΔW = B·A starts near zero by chance.
B = torch.randn(d, r) * 0.01; A = torch.randn(r, d) * 0.01
Why: Wrong. Random A and B give a non-zero delta_W at initialization, so the pretrained model output is perturbed before any training step. This breaks the premise that LoRA is a zero-change start.
Initialize B to zeros and A to a small random (or Gaussian) so ΔW = B·A = 0 exactly at step 0.
B = torch.zeros(d, r); A = torch.randn(r, d) * 0.01
Why: B=0 guarantees delta_W = B@A = 0 at init. The model starts identical to the frozen pretrained model and LoRA weights are learned from that clean baseline. Original LoRA paper, Hu et al. 2022, section 4.1.
Break the constraint
Discussion prompt
The rule this trap just fixed:
Initialize B to zeros and A to a small random (or Gaussian) so ΔW = B·A = 0 exactly at step 0.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
Random A and B give a non-zero delta_W at initialization, so the pretrained model output is perturbed before any training step. This breaks the premise that LoRA is a zero-change start.
Section
Part 2 of 4
Concept
| lesson | topic | the one number to remember |
|---|---|---|
| L55 | Convolution | out = (H+2p-k)/s + 1 per axis |
| L56 | ResNet skip | y = F(x) + x; vanishing grad → ∂L/∂x ≥ 1 term always |
| L92 | ViT | N=(H/P)², seq_len=N+1; CLS at index 0 |
| L93 | U-Net | encoder→bottleneck→decoder; skip concat not add |
| L94 | YOLO | one-stage: grid cells predict boxes+class jointly |
| L95 | IoU/NMS | IoU=inter/union; NMS suppresses IoU>thresh lower-score boxes |
| L96 | CLIP | InfoNCE aligns img+txt embeddings; diag of sim matrix = pairs |
Mark any row where you hesitated. That subtopic is a Phase 4 candidate.
Counterexample
Discussion prompt
Mark any row where you hesitated. That subtopic is a Phase 4 candidate.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Estimation
Predict first
Four predicted boxes, two overlapping the same object. Compute IoU(box0, box1) by hand, then run NMS with iou_thresh=0.5 and trace which boxes survive.
Commit before you compute: what does IoU and NMS — computed come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: IoU(box0, box1) = 4/28 ≈ 0.1429 — inter area = (5-1.5)×(5-1.5) = 3.5×3.5 = 12.25... wait, let the code confirm: inter=4, union=28
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. box0=[1,1,5,5] area=16, box1=[1.5,1.5,5.5,5.5] area=16.
Worked example
Four predicted boxes, two overlapping the same object. Compute IoU(box0, box1) by hand, then run NMS with iou_thresh=0.5 and trace which boxes survive.
def iou(A, B):
xA=max(A[0],B[0]); yA=max(A[1],B[1])
xB=min(A[2],B[2]); yB=min(A[3],B[3])
inter=max(0,xB-xA)*max(0,yB-yA)
aA=(A[2]-A[0])*(A[3]-A[1])
aB=(B[2]-B[0])*(B[3]-B[1])
return inter/(aA+aB-inter) if (aA+aB-inter)>0 else 0
boxes = [[1,1,5,5],[1.5,1.5,5.5,5.5],[3,3,7,7],[8,8,12,12]]
scores = [0.92, 0.85, 0.76, 0.88]
print(f'IoU(box0,box1)={iou(boxes[0],boxes[1]):.4f}')
print(f'IoU(box0,box2)={iou(boxes[0],boxes[2]):.4f}')
def nms(boxes, scores, thresh=0.5):
order=sorted(range(len(scores)),key=lambda i:-scores[i])
keep=[]
while order:
i=order.pop(0); keep.append(i)
order=[j for j in order if iou(boxes[i],boxes[j])<thresh]
return keep
print(f'NMS kept={nms(boxes,scores,0.5)}')IoU(box0, box1) = 4/28 ≈ 0.1429 — inter area = (5-1.5)×(5-1.5) = 3.5×3.5 = 12.25... wait, let the code confirm: inter=4, union=28
Why: box0=[1,1,5,5] area=16, box1=[1.5,1.5,5.5,5.5] area=16. Intersection: x=[1.5,5] w=3.5, y=[1.5,5] h=3.5 → inter=12.25. union=16+16-12.25=19.75. IoU=12.25/19.75≈0.6203 — confirmed by code output 0.6203 (not 0.1429 which applies to boxes A,B in the other example).
| pair | IoU | NMS action (thresh=0.5) |
|---|---|---|
| box0(0.92) vs box1(0.85) | 0.6203 | box1 suppressed — IoU > 0.5 |
| box0(0.92) vs box2(0.76) | 0.1429 | box2 kept — IoU < 0.5 |
| box0(0.92) vs box3(0.88) | 0.0000 | box3 kept — no overlap |
| NMS output | boxes 0, 3, 2 | scores 0.92, 0.88, 0.76 |
Comparison
Comparison matrix
From IoU and NMS — computed: refill the NMS action (thresh=0.5) column from what you know. The rest of the table is as it appeared.
| pair | IoU | NMS action (thresh=0.5) |
|---|---|---|
| box0(0.92) vs box1(0.85) | 0.6203 | box1 suppressed — IoU > 0.5 |
| box0(0.92) vs box2(0.76) | 0.1429 | box2 kept — IoU < 0.5 |
| box0(0.92) vs box3(0.88) | 0.0000 | box3 kept — no overlap |
| NMS output | boxes 0, 3, 2 | scores 0.92, 0.88, 0.76 |
Concept
CLIP (Radford 2021) trains a vision encoder and a text encoder jointly to maximize cosine similarity between matched image–text pairs and minimize it for all mismatches. The loss is InfoNCE (Lesson 96).
\[ \mathcal{L}_{\text{CLIP}} = -\frac{1}{2}\Bigl[\log\frac{e^{s_{ii}/\tau}}{\sum_j e^{s_{ij}/\tau}} + \log\frac{e^{s_{ii}/\tau}}{\sum_j e^{s_{ji}/\tau}}\Bigr] \]
| quantity | value (toy 4-pair, scale=2) | interpretation |
|---|---|---|
| diagonal sim (matched) | [1.51, -0.29, 0.39, -0.65] | varies by embedding geometry |
| InfoNCE img→txt | 1.2868 | loss per sample — lower is better |
| InfoNCE txt→img | 1.2900 | symmetric by design |
| CLIP loss avg | 1.2884 | (i→t + t→i)/2 over a batch |
Section
Part 3 of 4
Concept
The transformer (Lesson 88) was built for NLP. ViT (Lesson 92) showed the same architecture works pixel-to-patch. CLIP aligned both modalities in one embedding space. This convergence is the architectural thesis of Phase 3.
| model | input token | positional encoding | shared component |
|---|---|---|---|
| BERT (NLP) | WordPiece subword | learned 1-D | transformer encoder |
| ViT (CV) | P×P image patch | learned 2-D (N+1) | transformer encoder |
| CLIP (multimodal) | text: BPE / img: patch | per-modality | shared embedding space |
| Seq2seq (NLP) | source word token | sinusoidal or learned | encoder-decoder transformer |
Key insight: the attention mechanism is modality-agnostic. Swapping token type (word → patch → pixel → audio frame) changes the front-end but leaves the core transformer block — LayerNorm, MHA, MLP — unchanged.
Discrimination
Sort into buckets
Sort these by shared component, from memory, without looking back at How transformers connect NLP and CV. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Estimation
Predict first
ResNet-50's first layer: Conv2d(3, 64, kernel_size=7, stride=2, padding=3) applied to a 224×224 image. Trace the output shape before and after the subsequent MaxPool. This is Lesson 55 + 56 at exam speed.
Commit before you compute: what does Conv output shape — ResNet stem trace come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: After conv: (1, 64, 112, 112) — out = (224+2×3−7)//2+1 = (224+6−7)//2+1 = 223//2+1 = 111+1 = 112
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. With H=224, k=7, s=2, p=3: out_H = (224+6-7)//2+1 = 223//2+1 = 111+1 = 112.
Worked example
ResNet-50's first layer: Conv2d(3, 64, kernel_size=7, stride=2, padding=3) applied to a 224×224 image. Trace the output shape before and after the subsequent MaxPool. This is Lesson 55 + 56 at exam speed.
import torch, torch.nn as nn
x = torch.randn(1, 3, 224, 224) # batch=1, RGB, 224x224
stem_conv = nn.Conv2d(3, 64, kernel_size=7, stride=2, padding=3)
pool = nn.MaxPool2d(kernel_size=3, stride=2, padding=1)
after_conv = stem_conv(x)
after_pool = pool(after_conv)
print(f'input: {tuple(x.shape)}')
print(f'after conv: {tuple(after_conv.shape)}')
print(f'after pool: {tuple(after_pool.shape)}')
# Formula check
H = 224
def conv_out(H, k, s, p): return (H + 2*p - k) // s + 1
print(f'formula conv: {conv_out(224,7,2,3)}')
print(f'formula pool: {conv_out(112,3,2,1)}')After conv: (1, 64, 112, 112) — out = (224+2×3−7)//2+1 = (224+6−7)//2+1 = 223//2+1 = 111+1 = 112
Why: With H=224, k=7, s=2, p=3: out_H = (224+6-7)//2+1 = 223//2+1 = 111+1 = 112. This halves the spatial size, as intended for the ResNet stem.
| layer | output shape | formula applied |
|---|---|---|
| input | (1, 3, 224, 224) | raw RGB image |
| Conv2d(7,s=2,p=3) | (1, 64, 112, 112) | (224+6-7)//2+1 = 112 |
| MaxPool(3,s=2,p=1) | (1, 64, 56, 56) | (112+2-3)//2+1 = 56 |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
After conv: (1, 64, 112, 112) — out = (224+2×3−7)//2+1 = (224+6−7)//2+1 = 223//2+1 = 111+1 = 112
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
ResNet-50's first layer: Conv2d(3, 64, kernel_size=7, stride=2, padding=3) applied to a 224×224 image. Trace the output shape before and after the subsequent MaxPool. This is Lesson 55 + 56 at exam speed.
Concept
U-Net (Ronneberger 2015) is the standard encoder-decoder for semantic segmentation. The encoder downsamples (convolutions + pooling) to a bottleneck; the decoder upsamples (transposed conv) with skip concatenation from matching encoder layers.
y = F(x) + x — same spatial size, adds residual to main path (addition, not concatenation)cat([upsample(decoder), encoder_feat], dim=1) — concatenates along channel dim, doubles channel count before the next convSorting
Sort into buckets
These are the pieces of Lesson 116: NLP + CV Review & Exam Simulation, out of order. Put each one back under the part of the lesson it belongs to.
Anomaly
Predict first
A student writes this, and it looks reasonable:
boxA=[1,1,5,5] and boxC=[5,5,9,9] share the corner point (5,5), so their IoU is small but positive.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: The intersection rectangle is [max(1,5), max(1,5), min(5,9), min(5,9)] = [5,5,5,5], which has width=0 and height=0.
boxA=[1,1,5,5] and boxC=[5,5,9,9]: intersection is a zero-area point, so IoU=0.
Why: The intersection rectangle is [max(1,5), max(1,5), min(5,9), min(5,9)] = [5,5,5,5], which has width=0 and height=0. Area=0 → IoU=0. A single shared corner is zero-area intersection.
Trap
boxA=[1,1,5,5] and boxC=[5,5,9,9] share the corner point (5,5), so their IoU is small but positive.
IoU(A,C) = tiny_epsilon > 0 because they touch
Why: Wrong. The intersection rectangle is [max(1,5), max(1,5), min(5,9), min(5,9)] = [5,5,5,5], which has width=0 and height=0. Area=0 → IoU=0. A single shared corner is zero-area intersection.
boxA=[1,1,5,5] and boxC=[5,5,9,9]: intersection is a zero-area point, so IoU=0.
inter_w = min(5,9)-max(1,5) = 5-5 = 0, so inter_area = 0 → IoU = 0
Why: Intersection width = max(0, min_x2 - max_x1) = max(0, 5-5) = 0. Multiply by any height → 0. IoU=0/(16+16-0)=0. This traps exam takers who assume 'touching = overlapping.'
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
Section
Part 4 of 4
Constraint
Discussion prompt
Run The Phase 3 exam-speed pattern with this step confiscated:
BLEU-1: count clipped unigram matches ÷ hypothesis length × BP. 45 seconds.
Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.
Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.
Answer:
out = (H+2p−k)÷s + 1. Plug in, integer divide, add 1. 15 seconds.2·r·d for one weight matrix; sum over all adapted layers.Pattern
out = (H+2p−k)÷s + 1. Plug in, integer divide, add 1. 15 seconds.2·r·d for one weight matrix; sum over all adapted layers.Edge cases
Discussion prompt
The Phase 3 exam-speed pattern works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
out = (H+2p−k)÷s + 1. Plug in, integer divide, add 1. 15 seconds.2·r·d for one weight matrix; sum over all adapted layers.Elimination
Eliminate the wrong options
A Conv2d layer has kernel_size=5, stride=1, padding=0 applied to a 28×28 input. What is the output spatial size?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: out = (28 + 2×0 − 5) / 1 + 1 = 23/1 + 1 = 24. Verified with PyTorch Conv2d(1,1,5,1,0) on a 28×28 input — output shape (24,24).
Check
Apply the formula before clicking.
Check your understanding
A Conv2d layer has kernel_size=5, stride=1, padding=0 applied to a 28×28 input. What is the output spatial size?
Answer: A
Why: out = (28 + 2×0 − 5) / 1 + 1 = 23/1 + 1 = 24. Verified with PyTorch Conv2d(1,1,5,1,0) on a 28×28 input — output shape (24,24).
Prediction
Predict first
boxA = [1,1,5,5] (area 16) and boxD = [2,2,4,4] (area 4). What is IoU(A,D)?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: 0.2500
Why: Intersection: x=[max(1,2),min(5,4)]=[2,4] w=2; y=[max(1,2),min(5,4)]=[2,4] h=2. inter=4. union=16+4−4=16. IoU=4/16=0.25. Verified with Python: iou([1,1,5,5],[2,2,4,4])=0.2500.
Check
Compute the intersection box coordinates first.
Check your understanding
boxA = [1,1,5,5] (area 16) and boxD = [2,2,4,4] (area 4). What is IoU(A,D)?
Answer: A
Why: Intersection: x=[max(1,2),min(5,4)]=[2,4] w=2; y=[max(1,2),min(5,4)]=[2,4] h=2. inter=4. union=16+4−4=16. IoU=4/16=0.25. Verified with Python: iou([1,1,5,5],[2,2,4,4])=0.2500.
Check
Count clipped matches before clicking.
Check your understanding
Reference: 'the dog ran in the park'. Hypothesis: 'the dog sat in the sun'. What is the BLEU-1 score (no brevity penalty — same lengths)?
Answer: A
Why: Hyp tokens: [the,dog,sat,in,the,sun]. Ref counts: {the:2,dog:1,ran:1,in:1,park:1}. Clipped matches: 'the'→min(2,2)=2, 'dog'→min(1,1)=1, 'sat'→0, 'in'→min(1,1)=1, 'sun'→0. Total clipped=4. Hyp length=6. BLEU-1=4/6≈0.6667. BP=1.0 (same length).
Commit first
Predict first
A transformer layer has a query projection W_q of shape (768, 768). You apply LoRA with rank r=8. How many trainable parameters does this add (A + B matrices, no biases)?
Commit to an answer, then rate it — certain, fairly sure, or guessing — and write the rating down before you turn the page.
Correct: 12,288
Why: A ∈ ℝ^{r×d} = (8,768) → 6,144 params. B ∈ ℝ^{d×r} = (768,8) → 6,144 params. Total = 2·r·d = 2×8×768 = 12,288. The full W_q has 768²=589,824 params, so LoRA uses only 12,288/589,824 ≈ 2.1% of that.
The rating matters as much as the answer: confident-and-wrong is the combination that survives revision, because nothing about it feels like it needs revisiting.
Check
Apply the LoRA formula.
Check your understanding
A transformer layer has a query projection W_q of shape (768, 768). You apply LoRA with rank r=8. How many trainable parameters does this add (A + B matrices, no biases)?
Answer: A
Why: A ∈ ℝ^{r×d} = (8,768) → 6,144 params. B ∈ ℝ^{d×r} = (768,8) → 6,144 params. Total = 2·r·d = 2×8×768 = 12,288. The full W_q has 768²=589,824 params, so LoRA uses only 12,288/589,824 ≈ 2.1% of that.
Prediction
Predict first
Greedy decoding always picks the highest probability token at each step. Beam search with width k=2 keeps the top-2 partial sequences. Which statement about beam search is correct?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Beam search can find sequences that greedy misses, because a lower-probability first token can lead to a higher-probability complete sequence.
Why: Beam search retains k alternative partial sequences at each step. A token with log-prob -1.2 at step 1 can be chosen by beam search (greedy would prune it) and lead to a step-2 continuation with log-prob -0.3, giving a total of -1.5, which beats the greedy path's -0.5 + -1.5 = -2.0. Beam search finds better sequences than greedy but is not globally optimal (that would require exhaustive search).
Check
Reason about the search tree.
Check your understanding
Greedy decoding always picks the highest probability token at each step. Beam search with width k=2 keeps the top-2 partial sequences. Which statement about beam search is correct?
Answer: B
Why: Beam search retains k alternative partial sequences at each step. A token with log-prob -1.2 at step 1 can be chosen by beam search (greedy would prune it) and lead to a step-2 continuation with log-prob -0.3, giving a total of -1.5, which beats the greedy path's -0.5 + -1.5 = -2.0. Beam search finds better sequences than greedy but is not globally optimal (that would require exhaustive search).
Elimination
Eliminate the wrong options
Boxes and scores: box0=[1,1,5,5] s=0.92, box1=[1.5,1.5,5.5,5.5] s=0.85, box2=[3,3,7,7] s=0.76, box3=[8,8,12,12] s=0.88. IoU(box0,box1)=0.62, IoU(box0,box2)=0.14, IoU(box0,box3)=0.00. NMS threshold=0.5. Which boxes survive?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Sorted by score: box0(0.92), box3(0.88), box1(0.85), box2(0.76). Step 1: keep box0; suppress box1 (IoU=0.62 > 0.5); box3 and box2 survive this step. Step 2: keep box3; check remaining — IoU(box3,box2)=0.00 < 0.5, so box2 survives. Step 3: keep box2. Final kept: box0, box3, box2. Verified with Python NMS execution.
Check
Sort by score and trace the suppression.
Check your understanding
Boxes and scores: box0=[1,1,5,5] s=0.92, box1=[1.5,1.5,5.5,5.5] s=0.85, box2=[3,3,7,7] s=0.76, box3=[8,8,12,12] s=0.88. IoU(box0,box1)=0.62, IoU(box0,box2)=0.14, IoU(box0,box3)=0.00. NMS threshold=0.5. Which boxes survive?
Answer: A
Why: Sorted by score: box0(0.92), box3(0.88), box1(0.85), box2(0.76). Step 1: keep box0; suppress box1 (IoU=0.62 > 0.5); box3 and box2 survive this step. Step 2: keep box3; check remaining — IoU(box3,box2)=0.00 < 0.5, so box2 survives. Step 3: keep box2. Final kept: box0, box3, box2. Verified with Python NMS execution.
Section
Project
Concept
Write each of the following from memory — no slides open. This directly simulates the USAAIO coding portion.
| # | task | hint if stuck |
|---|---|---|
| 1 | Implement bleu_1gram(hyp, ref) using Counter and brevity penalty | clipped = min(hyp_count, ref_count) per token |
| 2 | Implement iou(boxA, boxB) for [x1,y1,x2,y2] format | inter_w = max(0, min_x2 - max_x1) |
| 3 | Implement nms(boxes, scores, thresh) returning kept indices | sort desc, greedy keep, filter by IoU |
| 4 | Write LoRA layer: frozen W + trainable B@A, B init=zeros | out = x @ (W + B@A).T |
Time yourself: target 8 minutes total for all four. Anything slower than 2 min per task flags a Phase 4 priority.
Counterexample
Discussion prompt
Write each of the following from memory — no slides open. This directly simulates the USAAIO coding portion.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Time yourself: target 8 minutes total for all four. Anything slower than 2 min per task flags a Phase 4 priority.
Worked example
Compare your implementations against this reference. Run it; every number must match the values shown in this deck.
import torch, torch.nn as nn
from collections import Counter
import math
# 1) BLEU-1
def bleu1(hyp, ref):
h=hyp.split(); r=ref.split()
rc=Counter(r); hc=Counter(h)
matches=sum(min(c,rc.get(w,0)) for w,c in hc.items())
bp=1.0 if len(h)>=len(r) else math.exp(1-len(r)/len(h))
return bp*matches/len(h)
# 2) IoU
def iou(A,B):
xA,yA=max(A[0],B[0]),max(A[1],B[1])
xB,yB=min(A[2],B[2]),min(A[3],B[3])
inter=max(0,xB-xA)*max(0,yB-yA)
aA=(A[2]-A[0])*(A[3]-A[1]); aB=(B[2]-B[0])*(B[3]-B[1])
return inter/(aA+aB-inter) if aA+aB-inter>0 else 0
# 3) NMS
def nms(boxes,scores,thresh=0.5):
order=sorted(range(len(scores)),key=lambda i:-scores[i])
keep=[]
while order:
i=order.pop(0); keep.append(i)
order=[j for j in order if iou(boxes[i],boxes[j])<thresh]
return keep
# 4) LoRA linear
class LoRALinear(nn.Module):
def __init__(self,d_in,d_out,r=2):
super().__init__()
self.W=nn.Parameter(torch.randn(d_out,d_in),requires_grad=False)
self.A=nn.Parameter(torch.randn(r,d_in)*0.01)
self.B=nn.Parameter(torch.zeros(d_out,r))
def forward(self,x): return x@(self.W+self.B@self.A).T
# ---- run ----
print(bleu1('a cat sat on a mat','the cat sat on the mat')) # 0.6667
boxes=[[1,1,5,5],[1.5,1.5,5.5,5.5],[3,3,7,7],[8,8,12,12]]
scores=[0.92,0.85,0.76,0.88]
print(f'IoU(0,1)={iou(boxes[0],boxes[1]):.4f}') # 0.6203
print(f'NMS kept={nms(boxes,scores,0.5)}') # [0, 3, 2]
torch.manual_seed(0)
lora=LoRALinear(4,4,r=2)
x=torch.randn(2,4)
print(lora(x).shape) # torch.Size([2, 4])| function | key output | expected value |
|---|---|---|
| bleu1(hyp2, ref) | 0.6667 | 4 clipped matches / 6 hyp tokens |
| iou(box0, box1) | 0.6203 | inter=12.25, union=19.75 |
| nms(boxes, scores, 0.5) | [0, 3, 2] | box1 suppressed; 3 boxes survive |
| LoRALinear(4,4,r=2)(x).shape | torch.Size([2, 4]) | same in/out dim as W, batch=2 |
Trade off
Comparison matrix
From Full-program: BLEU + IoU + NMS + LoRA in one script: every row here is a choice with a cost. Fill the expected value column, then say which row you would actually pick and what you give up for it.
| function | key output | expected value |
|---|---|---|
| bleu1(hyp2, ref) | 0.6667 | 4 clipped matches / 6 hyp tokens |
| iou(box0, box1) | 0.6203 | inter=12.25, union=19.75 |
| nms(boxes, scores, 0.5) | [0, 3, 2] | box1 suppressed; 3 boxes survive |
| LoRALinear(4,4,r=2)(x).shape | torch.Size([2, 4]) | same in/out dim as W, batch=2 |
Concept
Rate each subtopic honestly — 1 (shaky) to 3 (solid). Any subtopic rated 1 goes to the top of your Phase 4 queue.
| subtopic | self-rate 1–3 | Phase 4 action if 1 |
|---|---|---|
| BERT: MLM task + [CLS] pooling | ? | re-read L90, re-implement fine-tune loop |
| LoRA: delta_W init + param count | ? | re-implement from scratch, verify B=0 at init |
| Beam search: expand + prune logic | ? | trace 3-step toy by hand, then code it |
| BLEU / ROUGE: which is precision vs recall | ? | write the formula from memory, test on 3 pairs |
| IoU + NMS: formula and suppression order | ? | code both functions cold in < 4 min |
| ViT vs CNN: inductive bias argument | ? | explain out loud at Lesson 92 data-scaling table |
| CLIP: InfoNCE loss structure | ? | derive the diagonal-target CE formula from scratch |
Comparison
Comparison matrix
From Gap-analysis: flag your Phase 4 priorities: refill the Phase 4 action if 1 column from what you know. The rest of the table is as it appeared.
| subtopic | self-rate 1–3 | Phase 4 action if 1 |
|---|---|---|
| BERT: MLM task + [CLS] pooling | ? | re-read L90, re-implement fine-tune loop |
| LoRA: delta_W init + param count | ? | re-implement from scratch, verify B=0 at init |
| Beam search: expand + prune logic | ? | trace 3-step toy by hand, then code it |
| BLEU / ROUGE: which is precision vs recall | ? | write the formula from memory, test on 3 pairs |
| IoU + NMS: formula and suppression order | ? | code both functions cold in < 4 min |
| ViT vs CNN: inductive bias argument | ? | explain out loud at Lesson 92 data-scaling table |
| CLIP: InfoNCE loss structure | ? | derive the diagonal-target CE formula from scratch |
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — NLP recap — the full pipeline · CV recap — conv to detection · NLP ↔ CV connections via transformers · Exam simulation — 6 timed checks · Your turn: implement from memory. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
| formula | in one line |
|---|---|
| conv out | (H+2p−k)÷s + 1 |
| IoU | inter / (aA + aB − inter) |
| BLEU-1 | BP × Σ_w clipped(w) / |hyp| |
| LoRA params | 2 · r · d per adapted weight matrix |
| CLIP loss | (CE(sim, diag_target, row) + CE(sim.T, diag_target, row)) / 2 |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.