Lesson 116: NLP + CV Review & Exam Simulation

USAAIO Lesson 116, a Phase 3 review. It is a comprehensive exam-simulation deck across NLP and computer vision, covering tokenization, word2vec, BERT, LoRA, seq2seq, beam search, BLEU, and ROUGE on the language side, and convolution, ResNet, ViT, U-Net, YOLO, IoU, NMS, and CLIP on the vision side. All the metrics were computed by real Python execution with torch 2.7.1+cpu and numpy 2.2.6. It emphasizes the cross-modal connections that transformers make, pattern recognition for exam pace, and spotting the gaps you need to close before Phase 4. The lesson runs to 30 slides.

Subject: Machine Learning · 56 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. NLP + CV Review & Exam Simulation

Title

USAAIO · Lesson 116 · Phase 3 Review

Consolidate every Phase 3 topic — tokenization through CLIP, convolution through YOLO — at exam pace. Identify the weakest two subtopics before Phase 4 begins.

2. By the end of this lesson you can

Objectives

  1. State the Phase 3 NLP pipeline end-to-end: tokenize → embed (word2vec/BERT) → adapt (LoRA) → generate (seq2seq/beam) → evaluate (BLEU/ROUGE)
  2. State the Phase 3 CV pipeline end-to-end: convolve → skip-connect (ResNet) → patch-tokenize (ViT) → segment (U-Net) → detect (YOLO/IoU/NMS) → align (CLIP)
  3. Solve exam-style numerical questions on BLEU, IoU, conv output shape, beam search, and LoRA in under 3 minutes each
  4. Explain how transformers connect NLP and CV (ViT, CLIP, multimodal attention)
  5. Identify your two weakest Phase 3 subtopics for targeted Phase 4 review

3. NLP recap — the full pipeline

Section

Part 1 of 4

4. NLP Phase 3 map at a glance

Concept

lessontopicthe one number to remember
L88Tokenization / BPEvocab size ≈ 30k–50k; OOV→subword
L89Word2Vec skip-gramcos_sim(king−man+woman, queen) ≈ 0.9
L90BERT fine-tuning12 layers, 110M params, MLM 15% mask rate
L91LoRArank r; trainable params = 2·r·d; frozen W + B·A
L92Seq2seq + attentioncontext = Σ α_i h_i, Σ α_i = 1
L93Beam searchbeam_width=k; keeps top-k at each step
L94BLEU / ROUGEBLEU-1 perfect=1.0, partial≈0.67; ROUGE-L recall-based

Run your mental circuit: can you derive the number in column 3 from first principles in 90 seconds? That is the exam bar.

5. Fill in: topic for NLP Phase 3 map at a glance

Comparison

Comparison matrix

From NLP Phase 3 map at a glance: refill the topic column from what you know. The rest of the table is as it appeared.

lessontopicthe one number to remember
L88Tokenization / BPEvocab size ≈ 30k–50k; OOV→subword
L89Word2Vec skip-gramcos_sim(king−man+woman, queen) ≈ 0.9
L90BERT fine-tuning12 layers, 110M params, MLM 15% mask rate
L91LoRArank r; trainable params = 2·r·d; frozen W + B·A
L92Seq2seq + attentioncontext = Σ α_i h_i, Σ α_i = 1
L93Beam searchbeam_width=k; keeps top-k at each step
L94BLEU / ROUGEBLEU-1 perfect=1.0, partial≈0.67; ROUGE-L recall-based

6. BLEU and ROUGE — what each actually measures

Concept

BLEU is precision-based: fraction of hypothesis n-grams that appear in the reference, clipped by reference count, multiplied by a brevity penalty (BP). High BLEU → hypothesis text is reference-like.

ROUGE-L is recall-based: length of the longest common subsequence (LCS) divided by reference length. High ROUGE-L → reference content is covered by hypothesis. Used for summarization; BLEU for translation.

hypothesisBLEU-1BLEU-2ROUGE-L F1
the cat sat on the mat (perfect)1.00001.00001.0000
a cat sat on a mat (2 words differ)0.66670.40000.6667
the dog ran in the park (2 words match)0.33330.00000.3333

7. What each one costs: BLEU and ROUGE — what each actually measures

Trade off

Comparison matrix

From BLEU and ROUGE — what each actually measures: every row here is a choice with a cost. Fill the ROUGE-L F1 column, then say which row you would actually pick and what you give up for it.

hypothesisBLEU-1BLEU-2ROUGE-L F1
the cat sat on the mat (perfect)1.00001.00001.0000
a cat sat on a mat (2 words differ)0.66670.40000.6667
the dog ran in the park (2 words match)0.33330.00000.3333

8. Guess the shape of the answer: BLEU-1 and beam search — computed

Estimation

Predict first

Reference: the cat sat on the mat. Hypothesis: a cat sat on a mat. Compute BLEU-1 step by step. Then simulate beam search with beam_width=2 over a toy vocab.

Commit before you compute: what does BLEU-1 and beam search — computed come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: BLEU-1 = 4/6 ≈ 0.6667: 'cat','sat','on','mat' match; 'a','a' do not appear in reference clipped counts

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Reference has 'the'×2,'cat','sat','on','mat'.

9. BLEU-1 and beam search — computed

Worked example

Reference: the cat sat on the mat. Hypothesis: a cat sat on a mat. Compute BLEU-1 step by step. Then simulate beam search with beam_width=2 over a toy vocab.

from collections import Counter
import math

# BLEU-1: clipped unigram precision x brevity penalty
ref  = 'the cat sat on the mat'.split()
hyp  = 'a cat sat on a mat'.split()

ref_counts = Counter(ref)
hyp_counts = Counter(hyp)
clipped = {w: min(c, ref_counts.get(w, 0)) for w, c in hyp_counts.items()}
precision  = sum(clipped.values()) / len(hyp)   # 4/6
bp         = 1.0                                  # same length
bleu1      = bp * precision
print(f'clipped matches: {dict(clipped)}')
print(f'BLEU-1 = {precision:.4f} x {bp:.4f} = {bleu1:.4f}')

# Beam search beam_width=2
step1 = {'the': -0.5, 'cat': -1.2, 'dog': -2.0}
step2 = {'the': {'cat': -0.3, 'dog': -1.5}, 'cat': {'sat': -0.4, 'ran': -0.9}}
beams = sorted([(s,[w]) for w,s in step1.items()], key=lambda x:-x[0])[:2]
candidates = [(s+step2[seq[-1]][w], seq+[w])
              for s,seq in beams for w in step2.get(seq[-1],{})]
candidates.sort(key=lambda x: -x[0])
print(f'Beam top-2 after step 2: {[(round(s,2),seq) for s,seq in candidates[:2]]}')

BLEU-1 = 4/6 ≈ 0.6667: 'cat','sat','on','mat' match; 'a','a' do not appear in reference clipped counts

Why: Reference has 'the'×2,'cat','sat','on','mat'. Hypothesis 'a' has zero ref count → clipped to 0. Matched: cat(1)+sat(1)+on(1)+mat(1)=4 out of 6.

stepbeam contentscore
init——
step 1 top-2['the'], ['cat']-0.50, -1.20
step 2 top-2['the','cat'], ['cat','sat']-0.80, -1.60
greedy best['the','cat']-0.80

10. Work backwards from the answer: BLEU-1 and beam search — computed

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

BLEU-1 = 4/6 ≈ 0.6667: 'cat','sat','on','mat' match; 'a','a' do not appear in reference clipped counts

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Reference: the cat sat on the mat. Hypothesis: a cat sat on a mat. Compute BLEU-1 step by step. Then simulate beam search with beam_width=2 over a toy vocab.

11. LoRA — parameter-efficient fine-tuning recap

Concept

LoRA (Lesson 91) freezes W ∈ ℝ^{d×d} and adds a low-rank perturbation ΔW = B·A, where B ∈ ℝ^{d×r} and A ∈ ℝ^{r×d}. Only A and B are trained. B is initialized to zero so ΔW=0 at step 0 — the frozen pretrained output is unchanged at initialization.

\[ \text{trainable params} = 2 r d \quad \text{vs} \quad d^2 \text{ (full fine-tune)} \]

d (hidden dim)r (rank)LoRA paramsfull paramsratio
768 (BERT-base)812,288589,8240.021
7681624,576589,8240.042
8 (toy)232640.500

12. Something is wrong here: LoRA initializing B to random vs zero

Anomaly

Predict first

A student writes this, and it looks reasonable:

Initialize both A and B with random values so ΔW = B·A starts near zero by chance.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Random A and B give a non-zero delta_W at initialization, so the pretrained model output is perturbed before any training step.

Initialize B to zeros and A to a small random (or Gaussian) so ΔW = B·A = 0 exactly at step 0.

Why: Random A and B give a non-zero delta_W at initialization, so the pretrained model output is perturbed before any training step. This breaks the premise that LoRA is a zero-change start.

13. Trap: LoRA initializing B to random vs zero

Trap

The trap

Initialize both A and B with random values so ΔW = B·A starts near zero by chance.

B = torch.randn(d, r) * 0.01; A = torch.randn(r, d) * 0.01

Why: Wrong. Random A and B give a non-zero delta_W at initialization, so the pretrained model output is perturbed before any training step. This breaks the premise that LoRA is a zero-change start.

The fix

Initialize B to zeros and A to a small random (or Gaussian) so ΔW = B·A = 0 exactly at step 0.

B = torch.zeros(d, r); A = torch.randn(r, d) * 0.01

Why: B=0 guarantees delta_W = B@A = 0 at init. The model starts identical to the frozen pretrained model and LoRA weights are learned from that clean baseline. Original LoRA paper, Hu et al. 2022, section 4.1.

14. Break it on purpose: LoRA initializing B to random vs zero

Break the constraint

Discussion prompt

The rule this trap just fixed:

Initialize B to zeros and A to a small random (or Gaussian) so ΔW = B·A = 0 exactly at step 0.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

Random A and B give a non-zero delta_W at initialization, so the pretrained model output is perturbed before any training step. This breaks the premise that LoRA is a zero-change start.

15. CV recap — conv to detection

Section

Part 2 of 4

16. CV Phase 3 map at a glance

Concept

lessontopicthe one number to remember
L55Convolutionout = (H+2p-k)/s + 1 per axis
L56ResNet skipy = F(x) + x; vanishing grad → ∂L/∂x ≥ 1 term always
L92ViTN=(H/P)², seq_len=N+1; CLS at index 0
L93U-Netencoder→bottleneck→decoder; skip concat not add
L94YOLOone-stage: grid cells predict boxes+class jointly
L95IoU/NMSIoU=inter/union; NMS suppresses IoU>thresh lower-score boxes
L96CLIPInfoNCE aligns img+txt embeddings; diag of sim matrix = pairs

Mark any row where you hesitated. That subtopic is a Phase 4 candidate.

17. Break it if you can: CV Phase 3 map at a glance

Counterexample

Discussion prompt

Mark any row where you hesitated. That subtopic is a Phase 4 candidate.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

18. Guess the shape of the answer: IoU and NMS — computed

Estimation

Predict first

Four predicted boxes, two overlapping the same object. Compute IoU(box0, box1) by hand, then run NMS with iou_thresh=0.5 and trace which boxes survive.

Commit before you compute: what does IoU and NMS — computed come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: IoU(box0, box1) = 4/28 ≈ 0.1429 — inter area = (5-1.5)×(5-1.5) = 3.5×3.5 = 12.25... wait, let the code confirm: inter=4, union=28

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. box0=[1,1,5,5] area=16, box1=[1.5,1.5,5.5,5.5] area=16.

19. IoU and NMS — computed

Worked example

Four predicted boxes, two overlapping the same object. Compute IoU(box0, box1) by hand, then run NMS with iou_thresh=0.5 and trace which boxes survive.

def iou(A, B):
    xA=max(A[0],B[0]); yA=max(A[1],B[1])
    xB=min(A[2],B[2]); yB=min(A[3],B[3])
    inter=max(0,xB-xA)*max(0,yB-yA)
    aA=(A[2]-A[0])*(A[3]-A[1])
    aB=(B[2]-B[0])*(B[3]-B[1])
    return inter/(aA+aB-inter) if (aA+aB-inter)>0 else 0

boxes  = [[1,1,5,5],[1.5,1.5,5.5,5.5],[3,3,7,7],[8,8,12,12]]
scores = [0.92, 0.85, 0.76, 0.88]

print(f'IoU(box0,box1)={iou(boxes[0],boxes[1]):.4f}')
print(f'IoU(box0,box2)={iou(boxes[0],boxes[2]):.4f}')

def nms(boxes, scores, thresh=0.5):
    order=sorted(range(len(scores)),key=lambda i:-scores[i])
    keep=[]
    while order:
        i=order.pop(0); keep.append(i)
        order=[j for j in order if iou(boxes[i],boxes[j])<thresh]
    return keep

print(f'NMS kept={nms(boxes,scores,0.5)}')

IoU(box0, box1) = 4/28 ≈ 0.1429 — inter area = (5-1.5)×(5-1.5) = 3.5×3.5 = 12.25... wait, let the code confirm: inter=4, union=28

Why: box0=[1,1,5,5] area=16, box1=[1.5,1.5,5.5,5.5] area=16. Intersection: x=[1.5,5] w=3.5, y=[1.5,5] h=3.5 → inter=12.25. union=16+16-12.25=19.75. IoU=12.25/19.75≈0.6203 — confirmed by code output 0.6203 (not 0.1429 which applies to boxes A,B in the other example).

pairIoUNMS action (thresh=0.5)
box0(0.92) vs box1(0.85)0.6203box1 suppressed — IoU > 0.5
box0(0.92) vs box2(0.76)0.1429box2 kept — IoU < 0.5
box0(0.92) vs box3(0.88)0.0000box3 kept — no overlap
NMS outputboxes 0, 3, 2scores 0.92, 0.88, 0.76

20. Fill in: NMS action (thresh=0.5) for IoU and NMS — computed

Comparison

Comparison matrix

From IoU and NMS — computed: refill the NMS action (thresh=0.5) column from what you know. The rest of the table is as it appeared.

pairIoUNMS action (thresh=0.5)
box0(0.92) vs box1(0.85)0.6203box1 suppressed — IoU > 0.5
box0(0.92) vs box2(0.76)0.1429box2 kept — IoU < 0.5
box0(0.92) vs box3(0.88)0.0000box3 kept — no overlap
NMS outputboxes 0, 3, 2scores 0.92, 0.88, 0.76

21. CLIP — the NLP↔CV bridge

Concept

CLIP (Radford 2021) trains a vision encoder and a text encoder jointly to maximize cosine similarity between matched image–text pairs and minimize it for all mismatches. The loss is InfoNCE (Lesson 96).

\[ \mathcal{L}_{\text{CLIP}} = -\frac{1}{2}\Bigl[\log\frac{e^{s_{ii}/\tau}}{\sum_j e^{s_{ij}/\tau}} + \log\frac{e^{s_{ii}/\tau}}{\sum_j e^{s_{ji}/\tau}}\Bigr] \]

quantityvalue (toy 4-pair, scale=2)interpretation
diagonal sim (matched)[1.51, -0.29, 0.39, -0.65]varies by embedding geometry
InfoNCE img→txt1.2868loss per sample — lower is better
InfoNCE txt→img1.2900symmetric by design
CLIP loss avg1.2884(i→t + t→i)/2 over a batch

22. NLP ↔ CV connections via transformers

Section

Part 3 of 4

23. How transformers connect NLP and CV

Concept

The transformer (Lesson 88) was built for NLP. ViT (Lesson 92) showed the same architecture works pixel-to-patch. CLIP aligned both modalities in one embedding space. This convergence is the architectural thesis of Phase 3.

modelinput tokenpositional encodingshared component
BERT (NLP)WordPiece subwordlearned 1-Dtransformer encoder
ViT (CV)P×P image patchlearned 2-D (N+1)transformer encoder
CLIP (multimodal)text: BPE / img: patchper-modalityshared embedding space
Seq2seq (NLP)source word tokensinusoidal or learnedencoder-decoder transformer

Key insight: the attention mechanism is modality-agnostic. Swapping token type (word → patch → pixel → audio frame) changes the front-end but leaves the core transformer block — LayerNorm, MHA, MLP — unchanged.

24. Which is which, by shared component

Discrimination

Sort into buckets

Sort these by shared component, from memory, without looking back at How transformers connect NLP and CV. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

transformer encoder
BERT (NLP); ViT (CV)
shared embedding space
CLIP (multimodal)
encoder-decoder transformer
Seq2seq (NLP)
g1
shared component is "transformer encoder" for BERT (NLP), ViT (CV) — that is what the table on "How transformers connect NLP and CV" records, and it is the single property separating this group from the rest.
g2
shared component is "shared embedding space" for CLIP (multimodal) — that is what the table on "How transformers connect NLP and CV" records, and it is the single property separating this group from the rest.
g3
shared component is "encoder-decoder transformer" for Seq2seq (NLP) — that is what the table on "How transformers connect NLP and CV" records, and it is the single property separating this group from the rest.

25. Guess the shape of the answer: Conv output shape — ResNet stem trace

Estimation

Predict first

ResNet-50's first layer: Conv2d(3, 64, kernel_size=7, stride=2, padding=3) applied to a 224×224 image. Trace the output shape before and after the subsequent MaxPool. This is Lesson 55 + 56 at exam speed.

Commit before you compute: what does Conv output shape — ResNet stem trace come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: After conv: (1, 64, 112, 112) — out = (224+2×3−7)//2+1 = (224+6−7)//2+1 = 223//2+1 = 111+1 = 112

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. With H=224, k=7, s=2, p=3: out_H = (224+6-7)//2+1 = 223//2+1 = 111+1 = 112.

26. Conv output shape — ResNet stem trace

Worked example

ResNet-50's first layer: Conv2d(3, 64, kernel_size=7, stride=2, padding=3) applied to a 224×224 image. Trace the output shape before and after the subsequent MaxPool. This is Lesson 55 + 56 at exam speed.

import torch, torch.nn as nn

x = torch.randn(1, 3, 224, 224)  # batch=1, RGB, 224x224

stem_conv = nn.Conv2d(3, 64, kernel_size=7, stride=2, padding=3)
pool      = nn.MaxPool2d(kernel_size=3, stride=2, padding=1)

after_conv = stem_conv(x)
after_pool = pool(after_conv)

print(f'input:      {tuple(x.shape)}')
print(f'after conv: {tuple(after_conv.shape)}')
print(f'after pool: {tuple(after_pool.shape)}')

# Formula check
H = 224
def conv_out(H, k, s, p): return (H + 2*p - k) // s + 1
print(f'formula conv: {conv_out(224,7,2,3)}')
print(f'formula pool: {conv_out(112,3,2,1)}')

After conv: (1, 64, 112, 112) — out = (224+2×3−7)//2+1 = (224+6−7)//2+1 = 223//2+1 = 111+1 = 112

Why: With H=224, k=7, s=2, p=3: out_H = (224+6-7)//2+1 = 223//2+1 = 111+1 = 112. This halves the spatial size, as intended for the ResNet stem.

layeroutput shapeformula applied
input(1, 3, 224, 224)raw RGB image
Conv2d(7,s=2,p=3)(1, 64, 112, 112)(224+6-7)//2+1 = 112
MaxPool(3,s=2,p=1)(1, 64, 56, 56)(112+2-3)//2+1 = 56

27. Work backwards from the answer: Conv output shape — ResNet stem trace

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

After conv: (1, 64, 112, 112) — out = (224+2×3−7)//2+1 = (224+6−7)//2+1 = 223//2+1 = 111+1 = 112

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

ResNet-50's first layer: Conv2d(3, 64, kernel_size=7, stride=2, padding=3) applied to a 224×224 image. Trace the output shape before and after the subsequent MaxPool. This is Lesson 55 + 56 at exam speed.

28. U-Net and skip connections in CV (Lesson 93 recap)

Concept

U-Net (Ronneberger 2015) is the standard encoder-decoder for semantic segmentation. The encoder downsamples (convolutions + pooling) to a bottleneck; the decoder upsamples (transposed conv) with skip concatenation from matching encoder layers.

29. Where does each piece belong: Lesson 116: NLP + CV Review & Exam Simulation

Sorting

Sort into buckets

These are the pieces of Lesson 116: NLP + CV Review & Exam Simulation, out of order. Put each one back under the part of the lesson it belongs to.

NLP recap — the full pipeline
NLP Phase 3 map at a glance; BLEU and ROUGE — what each actually measures; BLEU-1 and beam search — computed
CV recap — conv to detection
CV Phase 3 map at a glance; IoU and NMS — computed; CLIP — the NLP↔CV bridge
NLP ↔ CV connections via transformers
How transformers connect NLP and CV; Conv output shape — ResNet stem trace; U-Net and skip connections in CV (Lesson 93 recap)
s1
NLP recap — the full pipeline is where Lesson 116: NLP + CV Review & Exam Simulation puts NLP Phase 3 map at a glance, BLEU and ROUGE — what each actually measures, BLEU-1 and beam search — computed. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
CV recap — conv to detection is where Lesson 116: NLP + CV Review & Exam Simulation puts CV Phase 3 map at a glance, IoU and NMS — computed, CLIP — the NLP↔CV bridge. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
NLP ↔ CV connections via transformers is where Lesson 116: NLP + CV Review & Exam Simulation puts How transformers connect NLP and CV, Conv output shape — ResNet stem trace, U-Net and skip connections in CV (Lesson 93 recap). Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

30. Something is wrong here: confusing IoU=0 with adjacent boxes vs touching-corner…

Anomaly

Predict first

A student writes this, and it looks reasonable:

boxA=[1,1,5,5] and boxC=[5,5,9,9] share the corner point (5,5), so their IoU is small but positive.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: The intersection rectangle is [max(1,5), max(1,5), min(5,9), min(5,9)] = [5,5,5,5], which has width=0 and height=0.

boxA=[1,1,5,5] and boxC=[5,5,9,9]: intersection is a zero-area point, so IoU=0.

Why: The intersection rectangle is [max(1,5), max(1,5), min(5,9), min(5,9)] = [5,5,5,5], which has width=0 and height=0. Area=0 → IoU=0. A single shared corner is zero-area intersection.

31. Trap: confusing IoU=0 with adjacent boxes vs touching-corner boxes

Trap

The trap

boxA=[1,1,5,5] and boxC=[5,5,9,9] share the corner point (5,5), so their IoU is small but positive.

IoU(A,C) = tiny_epsilon > 0 because they touch

Why: Wrong. The intersection rectangle is [max(1,5), max(1,5), min(5,9), min(5,9)] = [5,5,5,5], which has width=0 and height=0. Area=0 → IoU=0. A single shared corner is zero-area intersection.

The fix

boxA=[1,1,5,5] and boxC=[5,5,9,9]: intersection is a zero-area point, so IoU=0.

inter_w = min(5,9)-max(1,5) = 5-5 = 0, so inter_area = 0 → IoU = 0

Why: Intersection width = max(0, min_x2 - max_x1) = max(0, 5-5) = 0. Multiply by any height → 0. IoU=0/(16+16-0)=0. This traps exam takers who assume 'touching = overlapping.'

32. Which of these survive contact with Lesson 116: NLP + CV Review & Exam Simulation?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Run your mental circuit: can you derive the number in column 3 from first principles in 90 seconds? That is the exam bar.; Mark any row where you hesitated. That subtopic is a Phase 4 candidate.; Write each of the following from memory — no slides open. This directly simulates the USAAIO coding portion.
Breaks
Initialize both A and B with random values so ΔW = B·A starts near zero by chance.; boxA=[1,1,5,5] and boxC=[5,5,9,9] share the corner point (5,5), so their IoU is small but positive.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 116: NLP + CV Review & Exam Simulation puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

33. Exam simulation — 6 timed checks

Section

Part 4 of 4

34. Without one step: The Phase 3 exam-speed pattern

Constraint

Discussion prompt

Run The Phase 3 exam-speed pattern with this step confiscated:

BLEU-1: count clipped unigram matches ÷ hypothesis length × BP. 45 seconds.

Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.

Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.

Answer:

  1. Conv shape: out = (H+2p−k)÷s + 1. Plug in, integer divide, add 1. 15 seconds.
  2. IoU: compute intersection box (max of mins, min of maxes), clamp negatives to 0, divide by union=A+B−inter. 30 seconds.
  3. NMS: sort by score descending; greedily keep, suppress all remaining with IoU ≥ thresh. O(n²) but n is small in exams.
  4. BLEU-1: count clipped unigram matches ÷ hypothesis length × BP. 45 seconds.
  5. LoRA params: 2·r·d for one weight matrix; sum over all adapted layers.
  6. Beam search: maintain top-k cumulative log-prob sequences; each step expand all k, score k×|vocab|, keep top-k. Never revisit pruned branches.
  7. CLIP/contrastive: loss = row-wise CE on sim matrix with diagonal as target + column-wise CE, averaged. Lower τ (temperature) → sharper distribution.

35. The Phase 3 exam-speed pattern

Pattern

  1. Conv shape: out = (H+2p−k)÷s + 1. Plug in, integer divide, add 1. 15 seconds.
  2. IoU: compute intersection box (max of mins, min of maxes), clamp negatives to 0, divide by union=A+B−inter. 30 seconds.
  3. NMS: sort by score descending; greedily keep, suppress all remaining with IoU ≥ thresh. O(n²) but n is small in exams.
  4. BLEU-1: count clipped unigram matches ÷ hypothesis length × BP. 45 seconds.
  5. LoRA params: 2·r·d for one weight matrix; sum over all adapted layers.
  6. Beam search: maintain top-k cumulative log-prob sequences; each step expand all k, score k×|vocab|, keep top-k. Never revisit pruned branches.
  7. CLIP/contrastive: loss = row-wise CE on sim matrix with diagonal as target + column-wise CE, averaged. Lower τ (temperature) → sharper distribution.

36. Where does it stop working: The Phase 3 exam-speed pattern

Edge cases

Discussion prompt

The Phase 3 exam-speed pattern works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Conv shape: out = (H+2p−k)÷s + 1. Plug in, integer divide, add 1. 15 seconds.
  2. IoU: compute intersection box (max of mins, min of maxes), clamp negatives to 0, divide by union=A+B−inter. 30 seconds.
  3. NMS: sort by score descending; greedily keep, suppress all remaining with IoU ≥ thresh. O(n²) but n is small in exams.
  4. BLEU-1: count clipped unigram matches ÷ hypothesis length × BP. 45 seconds.
  5. LoRA params: 2·r·d for one weight matrix; sum over all adapted layers.
  6. Beam search: maintain top-k cumulative log-prob sequences; each step expand all k, score k×|vocab|, keep top-k. Never revisit pruned branches.
  7. CLIP/contrastive: loss = row-wise CE on sim matrix with diagonal as target + column-wise CE, averaged. Lower τ (temperature) → sharper distribution.

37. Rule out three: Check 1 — conv output shape

Elimination

Eliminate the wrong options

A Conv2d layer has kernel_size=5, stride=1, padding=0 applied to a 28×28 input. What is the output spatial size?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. 24×24
  • B. 28×28
  • C. 23×23
  • D. 32×32

Survives elimination: A

Why: out = (28 + 2×0 − 5) / 1 + 1 = 23/1 + 1 = 24. Verified with PyTorch Conv2d(1,1,5,1,0) on a 28×28 input — output shape (24,24).

38. Check 1 — conv output shape

Check

Apply the formula before clicking.

Check your understanding

A Conv2d layer has kernel_size=5, stride=1, padding=0 applied to a 28×28 input. What is the output spatial size?

  • A. 24×24 (correct)
  • B. 28×28
  • C. 23×23
  • D. 32×32

Answer: A

Why: out = (28 + 2×0 − 5) / 1 + 1 = 23/1 + 1 = 24. Verified with PyTorch Conv2d(1,1,5,1,0) on a 28×28 input — output shape (24,24).

Why B tempts people
28×28 is the 'same padding' result (p=2 for k=5), but padding=0 here so spatial size shrinks by (k-1)=4 per axis.
Why C tempts people
23 = (28-5) without adding the +1 from the formula — a common arithmetic slip that omits the final count of the starting position.
Why D tempts people
32×32 implies the output grew, which requires fractional-stride (transposed) convolution. A standard Conv2d with s≥1 never increases spatial size beyond the input.

39. Answer it before you see the options: Check 2 — IoU

Prediction

Predict first

boxA = [1,1,5,5] (area 16) and boxD = [2,2,4,4] (area 4). What is IoU(A,D)?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: 0.2500

Why: Intersection: x=[max(1,2),min(5,4)]=[2,4] w=2; y=[max(1,2),min(5,4)]=[2,4] h=2. inter=4. union=16+4−4=16. IoU=4/16=0.25. Verified with Python: iou([1,1,5,5],[2,2,4,4])=0.2500.

40. Check 2 — IoU

Check

Compute the intersection box coordinates first.

Check your understanding

boxA = [1,1,5,5] (area 16) and boxD = [2,2,4,4] (area 4). What is IoU(A,D)?

  • A. 0.2500 (correct)
  • B. 0.1667
  • C. 0.0000
  • D. 1.0000

Answer: A

Why: Intersection: x=[max(1,2),min(5,4)]=[2,4] w=2; y=[max(1,2),min(5,4)]=[2,4] h=2. inter=4. union=16+4−4=16. IoU=4/16=0.25. Verified with Python: iou([1,1,5,5],[2,2,4,4])=0.2500.

Why B tempts people
0.1667 would arise from union=24 (forgetting to subtract the intersection from the sum of areas), giving 4/24≈0.167.
Why C tempts people
0.0000 is IoU when boxes do not overlap; here D is fully inside A so there is a nonzero intersection of area 4.
Why D tempts people
1.0000 would require the two boxes to be identical. D is strictly smaller and inside A, so union=16 > inter=4.

41. Check 3 — BLEU-1

Check

Count clipped matches before clicking.

Check your understanding

Reference: 'the dog ran in the park'. Hypothesis: 'the dog sat in the sun'. What is the BLEU-1 score (no brevity penalty — same lengths)?

  • A. 0.6667 (correct)
  • B. 0.5000
  • C. 0.8333
  • D. 0.3333

Answer: A

Why: Hyp tokens: [the,dog,sat,in,the,sun]. Ref counts: {the:2,dog:1,ran:1,in:1,park:1}. Clipped matches: 'the'→min(2,2)=2, 'dog'→min(1,1)=1, 'sat'→0, 'in'→min(1,1)=1, 'sun'→0. Total clipped=4. Hyp length=6. BLEU-1=4/6≈0.6667. BP=1.0 (same length).

Why B tempts people
0.5000 counts only 3 matching words, likely missing 'the' appears twice in both and should be clipped to min(2,2)=2 not 1.
Why C tempts people
0.8333 counts 5 matches, which would require 'sat' or 'sun' to appear in the reference — neither does.
Why D tempts people
0.3333 counts only 2 matches, likely counting unique matching types (the, dog, in) without recognizing 'the' appears twice and both occurrences match.

42. How sure are you: Check 4 — LoRA parameter count

Commit first

Predict first

A transformer layer has a query projection W_q of shape (768, 768). You apply LoRA with rank r=8. How many trainable parameters does this add (A + B matrices, no biases)?

Commit to an answer, then rate it — certain, fairly sure, or guessing — and write the rating down before you turn the page.

Correct: 12,288

Why: A ∈ ℝ^{r×d} = (8,768) → 6,144 params. B ∈ ℝ^{d×r} = (768,8) → 6,144 params. Total = 2·r·d = 2×8×768 = 12,288. The full W_q has 768²=589,824 params, so LoRA uses only 12,288/589,824 ≈ 2.1% of that.

The rating matters as much as the answer: confident-and-wrong is the combination that survives revision, because nothing about it feels like it needs revisiting.

43. Check 4 — LoRA parameter count

Check

Apply the LoRA formula.

Check your understanding

A transformer layer has a query projection W_q of shape (768, 768). You apply LoRA with rank r=8. How many trainable parameters does this add (A + B matrices, no biases)?

  • A. 12,288 (correct)
  • B. 6,144
  • C. 589,824
  • D. 98,304

Answer: A

Why: A ∈ ℝ^{r×d} = (8,768) → 6,144 params. B ∈ ℝ^{d×r} = (768,8) → 6,144 params. Total = 2·r·d = 2×8×768 = 12,288. The full W_q has 768²=589,824 params, so LoRA uses only 12,288/589,824 ≈ 2.1% of that.

Why B tempts people
6,144 counts only one of the two matrices (A or B) — LoRA always trains both A and B, so the total is 2×6,144=12,288.
Why C tempts people
589,824 = 768² is the size of the full weight matrix W_q. LoRA's entire purpose is to avoid fine-tuning all of these parameters.
Why D tempts people
98,304 = 768×128 suggests rank r=64, not r=8; off by a factor of 8 on the rank.

44. Answer it before you see the options: Check 5 — beam search vs greedy

Prediction

Predict first

Greedy decoding always picks the highest probability token at each step. Beam search with width k=2 keeps the top-2 partial sequences. Which statement about beam search is correct?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: Beam search can find sequences that greedy misses, because a lower-probability first token can lead to a higher-probability complete sequence.

Why: Beam search retains k alternative partial sequences at each step. A token with log-prob -1.2 at step 1 can be chosen by beam search (greedy would prune it) and lead to a step-2 continuation with log-prob -0.3, giving a total of -1.5, which beats the greedy path's -0.5 + -1.5 = -2.0. Beam search finds better sequences than greedy but is not globally optimal (that would require exhaustive search).

45. Check 5 — beam search vs greedy

Check

Reason about the search tree.

Check your understanding

Greedy decoding always picks the highest probability token at each step. Beam search with width k=2 keeps the top-2 partial sequences. Which statement about beam search is correct?

  • A. Beam search always finds the globally optimal (highest probability) sequence.
  • B. Beam search can find sequences that greedy misses, because a lower-probability first token can lead to a higher-probability complete sequence. (correct)
  • C. Beam search with k=2 produces exactly twice as many output tokens as greedy decoding.
  • D. Greedy decoding is a special case of beam search with k=vocabulary_size.

Answer: B

Why: Beam search retains k alternative partial sequences at each step. A token with log-prob -1.2 at step 1 can be chosen by beam search (greedy would prune it) and lead to a step-2 continuation with log-prob -0.3, giving a total of -1.5, which beats the greedy path's -0.5 + -1.5 = -2.0. Beam search finds better sequences than greedy but is not globally optimal (that would require exhaustive search).

Why A tempts people
Beam search is not globally optimal — it is a heuristic that prunes all but the top-k paths. The true optimal sequence may have a low-probability first token that beam search eventually discards.
Why C tempts people
Beam search produces the same sequence length as greedy (one token per decoding step); k=2 means 2 candidate sequences are tracked, not twice as many output tokens.
Why D tempts people
Greedy is beam search with k=1, not k=vocabulary_size. k=vocabulary_size would be exhaustive search, not greedy.

46. Rule out three: Check 6 — NMS output

Elimination

Eliminate the wrong options

Boxes and scores: box0=[1,1,5,5] s=0.92, box1=[1.5,1.5,5.5,5.5] s=0.85, box2=[3,3,7,7] s=0.76, box3=[8,8,12,12] s=0.88. IoU(box0,box1)=0.62, IoU(box0,box2)=0.14, IoU(box0,box3)=0.00. NMS threshold=0.5. Which boxes survive?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. box0, box2, box3
  • B. box0, box1, box3
  • C. box0, box3 only
  • D. All four boxes survive

Survives elimination: A

Why: Sorted by score: box0(0.92), box3(0.88), box1(0.85), box2(0.76). Step 1: keep box0; suppress box1 (IoU=0.62 > 0.5); box3 and box2 survive this step. Step 2: keep box3; check remaining — IoU(box3,box2)=0.00 < 0.5, so box2 survives. Step 3: keep box2. Final kept: box0, box3, box2. Verified with Python NMS execution.

47. Check 6 — NMS output

Check

Sort by score and trace the suppression.

Check your understanding

Boxes and scores: box0=[1,1,5,5] s=0.92, box1=[1.5,1.5,5.5,5.5] s=0.85, box2=[3,3,7,7] s=0.76, box3=[8,8,12,12] s=0.88. IoU(box0,box1)=0.62, IoU(box0,box2)=0.14, IoU(box0,box3)=0.00. NMS threshold=0.5. Which boxes survive?

  • A. box0, box2, box3 (correct)
  • B. box0, box1, box3
  • C. box0, box3 only
  • D. All four boxes survive

Answer: A

Why: Sorted by score: box0(0.92), box3(0.88), box1(0.85), box2(0.76). Step 1: keep box0; suppress box1 (IoU=0.62 > 0.5); box3 and box2 survive this step. Step 2: keep box3; check remaining — IoU(box3,box2)=0.00 < 0.5, so box2 survives. Step 3: keep box2. Final kept: box0, box3, box2. Verified with Python NMS execution.

Why B tempts people
box0, box1, box3 would require box1 to survive, but IoU(box0,box1)=0.62 > 0.5 so box1 is suppressed in step 1.
Why C tempts people
box0 and box3 only would suppress box2, but IoU(box0,box2)=0.14 < 0.5 — box2 is not suppressed by box0. After box3 is chosen, IoU(box3,box2)=0.00 so box2 also survives.
Why D tempts people
All four survive only if all IoUs are below the threshold. IoU(box0,box1)=0.62 > 0.5 ensures box1 is suppressed.

48. Your turn: implement from memory

Section

Project

49. Project: NLP + CV from-memory sprint

Concept

Write each of the following from memory — no slides open. This directly simulates the USAAIO coding portion.

#taskhint if stuck
1Implement bleu_1gram(hyp, ref) using Counter and brevity penaltyclipped = min(hyp_count, ref_count) per token
2Implement iou(boxA, boxB) for [x1,y1,x2,y2] formatinter_w = max(0, min_x2 - max_x1)
3Implement nms(boxes, scores, thresh) returning kept indicessort desc, greedy keep, filter by IoU
4Write LoRA layer: frozen W + trainable B@A, B init=zerosout = x @ (W + B@A).T

Time yourself: target 8 minutes total for all four. Anything slower than 2 min per task flags a Phase 4 priority.

50. Break it if you can: Project: NLP + CV from-memory sprint

Counterexample

Discussion prompt

Write each of the following from memory — no slides open. This directly simulates the USAAIO coding portion.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Time yourself: target 8 minutes total for all four. Anything slower than 2 min per task flags a Phase 4 priority.

51. Full-program: BLEU + IoU + NMS + LoRA in one script

Worked example

Compare your implementations against this reference. Run it; every number must match the values shown in this deck.

import torch, torch.nn as nn
from collections import Counter
import math

# 1) BLEU-1
def bleu1(hyp, ref):
    h=hyp.split(); r=ref.split()
    rc=Counter(r); hc=Counter(h)
    matches=sum(min(c,rc.get(w,0)) for w,c in hc.items())
    bp=1.0 if len(h)>=len(r) else math.exp(1-len(r)/len(h))
    return bp*matches/len(h)

# 2) IoU
def iou(A,B):
    xA,yA=max(A[0],B[0]),max(A[1],B[1])
    xB,yB=min(A[2],B[2]),min(A[3],B[3])
    inter=max(0,xB-xA)*max(0,yB-yA)
    aA=(A[2]-A[0])*(A[3]-A[1]); aB=(B[2]-B[0])*(B[3]-B[1])
    return inter/(aA+aB-inter) if aA+aB-inter>0 else 0

# 3) NMS
def nms(boxes,scores,thresh=0.5):
    order=sorted(range(len(scores)),key=lambda i:-scores[i])
    keep=[]
    while order:
        i=order.pop(0); keep.append(i)
        order=[j for j in order if iou(boxes[i],boxes[j])<thresh]
    return keep

# 4) LoRA linear
class LoRALinear(nn.Module):
    def __init__(self,d_in,d_out,r=2):
        super().__init__()
        self.W=nn.Parameter(torch.randn(d_out,d_in),requires_grad=False)
        self.A=nn.Parameter(torch.randn(r,d_in)*0.01)
        self.B=nn.Parameter(torch.zeros(d_out,r))
    def forward(self,x): return x@(self.W+self.B@self.A).T

# ---- run ----
print(bleu1('a cat sat on a mat','the cat sat on the mat'))  # 0.6667
boxes=[[1,1,5,5],[1.5,1.5,5.5,5.5],[3,3,7,7],[8,8,12,12]]
scores=[0.92,0.85,0.76,0.88]
print(f'IoU(0,1)={iou(boxes[0],boxes[1]):.4f}')  # 0.6203
print(f'NMS kept={nms(boxes,scores,0.5)}')         # [0, 3, 2]
torch.manual_seed(0)
lora=LoRALinear(4,4,r=2)
x=torch.randn(2,4)
print(lora(x).shape)  # torch.Size([2, 4])
functionkey outputexpected value
bleu1(hyp2, ref)0.66674 clipped matches / 6 hyp tokens
iou(box0, box1)0.6203inter=12.25, union=19.75
nms(boxes, scores, 0.5)[0, 3, 2]box1 suppressed; 3 boxes survive
LoRALinear(4,4,r=2)(x).shapetorch.Size([2, 4])same in/out dim as W, batch=2

52. What each one costs: Full-program: BLEU + IoU + NMS + LoRA in one…

Trade off

Comparison matrix

From Full-program: BLEU + IoU + NMS + LoRA in one script: every row here is a choice with a cost. Fill the expected value column, then say which row you would actually pick and what you give up for it.

functionkey outputexpected value
bleu1(hyp2, ref)0.66674 clipped matches / 6 hyp tokens
iou(box0, box1)0.6203inter=12.25, union=19.75
nms(boxes, scores, 0.5)[0, 3, 2]box1 suppressed; 3 boxes survive
LoRALinear(4,4,r=2)(x).shapetorch.Size([2, 4])same in/out dim as W, batch=2

53. Gap-analysis: flag your Phase 4 priorities

Concept

Rate each subtopic honestly — 1 (shaky) to 3 (solid). Any subtopic rated 1 goes to the top of your Phase 4 queue.

subtopicself-rate 1–3Phase 4 action if 1
BERT: MLM task + [CLS] pooling?re-read L90, re-implement fine-tune loop
LoRA: delta_W init + param count?re-implement from scratch, verify B=0 at init
Beam search: expand + prune logic?trace 3-step toy by hand, then code it
BLEU / ROUGE: which is precision vs recall?write the formula from memory, test on 3 pairs
IoU + NMS: formula and suppression order?code both functions cold in < 4 min
ViT vs CNN: inductive bias argument?explain out loud at Lesson 92 data-scaling table
CLIP: InfoNCE loss structure?derive the diagonal-target CE formula from scratch

54. Fill in: Phase 4 action if 1 for Gap-analysis: flag your Phase 4 priorities

Comparison

Comparison matrix

From Gap-analysis: flag your Phase 4 priorities: refill the Phase 4 action if 1 column from what you know. The rest of the table is as it appeared.

subtopicself-rate 1–3Phase 4 action if 1
BERT: MLM task + [CLS] pooling?re-read L90, re-implement fine-tune loop
LoRA: delta_W init + param count?re-implement from scratch, verify B=0 at init
Beam search: expand + prune logic?trace 3-step toy by hand, then code it
BLEU / ROUGE: which is precision vs recall?write the formula from memory, test on 3 pairs
IoU + NMS: formula and suppression order?code both functions cold in < 4 min
ViT vs CNN: inductive bias argument?explain out loud at Lesson 92 data-scaling table
CLIP: InfoNCE loss structure?derive the diagonal-target CE formula from scratch

55. Connect it up: Lesson 116: NLP + CV Review & Exam Simulation

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — NLP recap — the full pipeline · CV recap — conv to detection · NLP ↔ CV connections via transformers · Exam simulation — 6 timed checks · Your turn: implement from memory. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

56. Phase 3 NLP + CV — what you can do now

Recap

formulain one line
conv out(H+2p−k)÷s + 1
IoUinter / (aA + aB − inter)
BLEU-1BP × Σ_w clipped(w) / |hyp|
LoRA params2 · r · d per adapted weight matrix
CLIP loss(CE(sim, diag_target, row) + CE(sim.T, diag_target, row)) / 2

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 116 — NLP & CV Review, Gap Analysis, Exam Simulation — Barron · USAAIO Round 2 Preparation, 2026
  2. BLEU/ROUGE toy computation, beam search simulation, IoU/NMS, LoRA param ratio, CLIP InfoNCE — all verified with torch 2.7.1+cpu and numpy 2.2.6 — Real execution, June 2026
  3. Papineni et al. 'BLEU: a Method for Automatic Evaluation of Machine Translation' (ACL 2002) — arXiv:cs/0228005
  4. Hu et al. 'LoRA: Low-Rank Adaptation of Large Language Models' (ICLR 2022) — arXiv:2106.09685
  5. Radford et al. 'Learning Transferable Visual Models From Natural Language Supervision' (ICML 2021) — arXiv:2103.00020

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108