USAAIO Lesson 109, from Phase 3. It covers how a bounding box is parameterized, in both corner and centre-offset formats, then builds IoU from scratch and covers non-maximum suppression. It works the YOLO grid and anchor arithmetic - S=13, B=5, C=80, giving 17,745 raw predictions - and decodes anchor offsets with sigmoid and exp. It then covers Feature Pyramid Networks for multi-scale detection and the calculation of mAP at 0.5. Every number was verified with torch 2.7.1+cpu and numpy 2.2.6. The lesson runs to 27 slides.
Subject: Machine Learning · 52 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 109 · Phase 3
From a raw image to [class, confidence, x, y, w, h]: bounding-box formats, IoU, NMS, the YOLO grid, anchor-offset decoding, FPN multi-scale fusion, and mAP@0.5 — every number executed in Python.
Objectives
(x1, y1, x2, y2) and centre format (cx, cy, w, h), and explain why YOLO predicts normalised centre offsets(S, B, C) and decode (tx, ty, tw, th) back to pixel coordinatesWarm-up
Discussion prompt
Before we open Lesson 109: Object Detection & YOLO: without looking back, what was the main idea of Summarization, BART, Perplexity & Calibration, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
extractive vs abstractive summarization, BART denoising pretraining (text infilling / token deletion / sentence permutation), Pegasus sentence-level masking, perplexity as exponentiated cross-entropy with concrete log-prob trace on a 3-token toy LM (PPL good=1.26 bad=1.42 uniform=3), TinyBART encoder-decoder (22,996 params) fine-tuned on a copy task (PPL 11.02→1.42 in 50 epochs), calibration ECE with 5-bin reliability table (calibrated ECE=0.0296 vs overconfident ECE=0.1862).
Section
Part 1 of 5
Concept
Every object-detection model produces a list of detections. Each detection is a tuple of five things that together describe one predicted object.
(x, y, w, h) or (x1, y1, x2, y2).car, person, dog, …sigmoid, in [0, 1].Post-processing (confidence threshold → NMS → final list) converts thousands of raw predictions into a small set of clean, non-overlapping boxes.
Matching
Match the pairs
From What a detector outputs — match each one to what it actually does. The descriptions have been shuffled.
(x, y, w, h) or (x1, y1, x2, y2).car, person, dog, …sigmoid, in [0, 1].Why: Bounding box, Class label, Confidence are easy to tell apart while they are sitting next to their descriptions and much harder afterwards, which is what this checks.
Concept
Two formats appear in every codebase. Corner format matches annotation tools; centre format is what YOLO predicts (easier to regress offsets from an anchor).
| Format | Fields | Convert to centre |
|---|---|---|
| Corner | (x1, y1, x2, y2) | cx = (x1+x2)/2, cy = (y1+y2)/2, w = x2-x1, h = y2-y1 |
| Centre | (cx, cy, w, h) | x1 = cx-w/2, y1 = cy-h/2, x2 = cx+w/2, y2 = cy+h/2 |
| Normalised centre | (cx/W, cy/H, w/W, h/H) | All values in [0,1]; image-size-agnostic |
Example: box [100, 150, 260, 300] on a 400×400 image → cx=180, cy=225, w=160, h=150 → normalised cx=0.4500, cy=0.5625, w=0.4000, h=0.3750.
Comparison
Comparison matrix
From Bounding-box formats: refill the Fields column from what you know. The rest of the table is as it appeared.
| Format | Fields | Convert to centre |
|---|---|---|
| Corner | (x1, y1, x2, y2) | cx = (x1+x2)/2, cy = (y1+y2)/2, w = x2-x1, h = y2-y1 |
| Centre | (cx, cy, w, h) | x1 = cx-w/2, y1 = cy-h/2, x2 = cx+w/2, y2 = cy+h/2 |
| Normalised centre | (cx/W, cy/H, w/W, h/H) | All values in [0,1]; image-size-agnostic |
Section
Part 2 of 5
Concept
IoU is the canonical overlap metric in detection. It is threshold-free and scale-invariant — a value of 0 means no overlap; 1 means perfect alignment.
\[ \text{IoU}(A, B) = \frac{|A \cap B|}{|A \cup B|} = \frac{|A \cap B|}{|A| + |B| - |A \cap B|} \]
In NMS, boxes with IoU > threshold (often 0.5) are considered duplicates of the higher-confidence box and are suppressed. In mAP, a detection is a true positive only if its IoU with a ground-truth box exceeds 0.5.
Counterexample
Discussion prompt
IoU is the canonical overlap metric in detection. It is threshold-free and scale-invariant — a value of 0 means no overlap; 1 means perfect alignment.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Estimation
Predict first
Two boxes on the same 400×400 image. Compute IoU by finding the intersection rectangle, then applying the formula. Predict before reading the table.
Commit before you compute: what does IoU — computed example come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Compute intersection rectangle: ix1=180, iy1=200, ix2=260, iy2=300
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Intersection starts where both boxes start (max of left/top) and ends where both boxes end (min of right/bottom).
Worked example
Two boxes on the same 400×400 image. Compute IoU by finding the intersection rectangle, then applying the formula. Predict before reading the table.
def iou(box1, box2):
"""boxes: [x1, y1, x2, y2]"""
ix1 = max(box1[0], box2[0])
iy1 = max(box1[1], box2[1])
ix2 = min(box1[2], box2[2])
iy2 = min(box1[3], box2[3])
inter_w = max(0, ix2 - ix1)
inter_h = max(0, iy2 - iy1)
inter = inter_w * inter_h
area1 = (box1[2]-box1[0]) * (box1[3]-box1[1])
area2 = (box2[2]-box2[0]) * (box2[3]-box2[1])
union = area1 + area2 - inter
return inter / union if union > 0 else 0.0
b1 = [100, 150, 260, 300]
b2 = [180, 200, 320, 350]
print(f"IoU = {iou(b1, b2):.4f}") # 0.2162Compute intersection rectangle: ix1=180, iy1=200, ix2=260, iy2=300
Why: Intersection starts where both boxes start (max of left/top) and ends where both boxes end (min of right/bottom).
| Quantity | Calculation | Value |
|---|---|---|
| area(B1) | 160 × 150 | 24 000 px² |
| area(B2) | 140 × 150 | 21 000 px² |
| inter_w | min(260,320) − max(100,180) = 260−180 | 80 px |
| inter_h | min(300,350) − max(150,200) = 300−200 | 100 px |
| intersection | 80 × 100 | 8 000 px² |
| union | 24 000 + 21 000 − 8 000 | 37 000 px² |
| IoU | 8 000 / 37 000 | 0.2162 |
Trade off
Comparison matrix
From IoU — computed example: every row here is a choice with a cost. Fill the Calculation column, then say which row you would actually pick and what you give up for it.
| Quantity | Calculation | Value |
|---|---|---|
| area(B1) | 160 × 150 | 24 000 px² |
| area(B2) | 140 × 150 | 21 000 px² |
| inter_w | min(260,320) − max(100,180) = 260−180 | 80 px |
| inter_h | min(300,350) − max(150,200) = 300−200 | 100 px |
| intersection | 80 × 100 | 8 000 px² |
| union | 24 000 + 21 000 − 8 000 | 37 000 px² |
| IoU | 8 000 / 37 000 | 0.2162 |
Missing information
Discussion prompt
Four candidate boxes, one class. NMS keeps the highest-confidence box, suppresses all others with IoU > 0.5, then repeats on the survivors.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
We always keep the most-confident box and evaluate overlap relative to it.
Worked example
Four candidate boxes, one class. NMS keeps the highest-confidence box, suppresses all others with IoU > 0.5, then repeats on the survivors.
def nms(boxes, scores, iou_thresh=0.5):
order = sorted(range(len(scores)),
key=lambda i: scores[i], reverse=True)
kept = []
while order:
i = order.pop(0)
kept.append(i)
remaining = []
for j in order:
if iou(boxes[i], boxes[j]) < iou_thresh:
remaining.append(j)
order = remaining
return kept
boxes = [[100,150,260,300], [110,155,265,305],
[300,300,400,400], [105,145,258,298]]
scores = [0.9, 0.8, 0.95, 0.7]
print(nms(boxes, scores)) # [2, 0]Sort by score desc: order = [2(0.95), 0(0.90), 1(0.80), 3(0.70)]
Why: We always keep the most-confident box and evaluate overlap relative to it.
| Round | Kept box | Suppressed (IoU>0.5) | Survivors |
|---|---|---|---|
| 1 | idx 2 score=0.95 [300,300,400,400] | none (IoU<0.5 with all others) | 0, 1, 3 |
| 2 | idx 0 score=0.90 [100,150,260,300] | idx 1 (IoU≈0.93), idx 3 (IoU≈0.95) | none |
| Final | kept=[2, 0] | — | 2 boxes remain |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Sort by score desc: order = [2(0.95), 0(0.90), 1(0.80), 3(0.70)]
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Four candidate boxes, one class. NMS keeps the highest-confidence box, suppresses all others with IoU > 0.5, then repeats on the survivors.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Iterate boxes in original order and suppress any later box with IoU > thresh against every already-kept box.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Box[0] has score 0.9 but box[1] has score 0.8 — suppressing 0.8 in favour of 0.9 happens to be correct here.
Always sort by confidence descending first, then greedily keep the top-scoring box and suppress everything with IoU > thresh against it.
Why: Box[0] has score 0.9 but box[1] has score 0.8 — suppressing 0.8 in favour of 0.9 happens to be correct here.
Trap
Iterate boxes in original order and suppress any later box with IoU > thresh against every already-kept box.
Check box[0] vs box[1]: IoU=0.93 → suppress box[1]
Why: Box[0] has score 0.9 but box[1] has score 0.8 — suppressing 0.8 in favour of 0.9 happens to be correct here.
But if box[1] had score 0.99 and box[0] had score 0.6, we would wrongly keep the worse detection — the two are duplicates and NMS should keep the higher-confidence one.
Always sort by confidence descending first, then greedily keep the top-scoring box and suppress everything with IoU > thresh against it.
Check sorted box[2](0.95) vs box[0](0.90): IoU=0.0 (non-overlapping) → keep both
Why: The highest-confidence box gets to stay unconditionally; only lower-confidence boxes can be suppressed.
Result: kept=[2, 0]. The two near-duplicate boxes (idx 0, 1, 3) correctly collapse to the single highest-confidence one (idx 0, score=0.90).
Break the constraint
Discussion prompt
The rule this trap just fixed:
The highest-confidence box gets to stay unconditionally; only lower-confidence boxes can be suppressed.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
Box[0] has score 0.9 but box[1] has score 0.8 — suppressing 0.8 in favour of 0.9 happens to be correct here.
Section
Part 3 of 5
Concept
YOLO (You Only Look Once) divides the image into an S × S grid. Each cell is responsible for detecting objects whose centre falls inside it — the cell is a spatial prior, not a crop.
\[ \text{per-cell outputs} = B \times 5 + C \quad\text{where each box has } (t_x, t_y, t_w, t_h, \text{conf}) \]
| Parameter | YOLOv2 (COCO) | Meaning |
|---|---|---|
| S | 13 | Grid size (13×13 cells on a 416×416 input) |
| B | 5 | Anchor boxes per cell |
| C | 80 | COCO class count |
| per-cell outputs | 5×5 + 80 = 105 | tx,ty,tw,th,conf per anchor + class probs |
| total predictions | 13×13×105 = 17 745 | All raw boxes before threshold + NMS |
Comparison
Comparison matrix
From YOLO grid: dividing the image: refill the Meaning column from what you know. The rest of the table is as it appeared.
| Parameter | YOLOv2 (COCO) | Meaning |
|---|---|---|
| S | 13 | Grid size (13×13 cells on a 416×416 input) |
| B | 5 | Anchor boxes per cell |
| C | 80 | COCO class count |
| per-cell outputs | 5×5 + 80 = 105 | tx,ty,tw,th,conf per anchor + class probs |
| total predictions | 13×13×105 = 17 745 | All raw boxes before threshold + NMS |
Concept
Instead of predicting absolute (w, h), YOLO predicts offsets relative to a predefined anchor. Anchors are clustered from training box sizes (k-means on ground-truth w,h), so the network learns small corrections rather than large absolute values.
\[ b_x = \sigma(t_x) + c_x \quad b_y = \sigma(t_y) + c_y \quad b_w = p_w e^{t_w} \quad b_h = p_h e^{t_h} \]
c_x, c_y is the cell's top-left corner in grid units; p_w, p_h is the anchor's width/height in pixels. sigma bounds b_x, b_y within the responsible cell.
Analogy
Discussion prompt
Explain Anchor boxes — predefined priors by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
c_x, c_y is the cell's top-left corner in grid units; p_w, p_h is the anchor's width/height in pixels. sigma bounds b_x, b_y within the responsible cell.
Estimation
Predict first
Cell (2, 3) (column 2, row 3 in the 13×13 grid). Anchor prior: p_w=50 px, p_h=80 px. Network outputs tx=0.3, ty=-0.2, tw=0.5, th=-0.1. Decode to pixel box centre.
Commit before you compute: what does Anchor offset decoding — concrete numbers come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: sigma(0.3) = 0.5744; sigma(-0.2) = 0.4502 — both within (0, 1)
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. sigmoid bounds the predicted offset to [0,1], keeping the centre strictly inside the responsible cell; without it the model could predict centres anywhere.
Worked example
Cell (2, 3) (column 2, row 3 in the 13×13 grid). Anchor prior: p_w=50 px, p_h=80 px. Network outputs tx=0.3, ty=-0.2, tw=0.5, th=-0.1. Decode to pixel box centre.
import torch, math
tx, ty, tw, th = 0.3, -0.2, 0.5, -0.1
cx_cell, cy_cell = 2.0, 3.0 # grid-unit cell offset
p_w, p_h = 50, 80 # anchor prior (pixels)
bx = torch.sigmoid(torch.tensor(tx)).item() + cx_cell
by = torch.sigmoid(torch.tensor(ty)).item() + cy_cell
bw = p_w * math.exp(tw)
bh = p_h * math.exp(th)
print(f"bx = {bx:.4f} by = {by:.4f} (grid coords)")
print(f"bw = {bw:.2f} bh = {bh:.2f} (pixels)")
# bx=2.5744 by=3.4502 bw=82.44 bh=72.39sigma(0.3) = 0.5744; sigma(-0.2) = 0.4502 — both within (0, 1)
Why: sigmoid bounds the predicted offset to [0,1], keeping the centre strictly inside the responsible cell; without it the model could predict centres anywhere.
| Variable | Formula | Value |
|---|---|---|
| sigma(tx) | 1/(1+exp(-0.3)) | 0.5744 (grid) |
| bx | 0.5744 + 2.0 | 2.5744 (grid) |
| sigma(ty) | 1/(1+exp(0.2)) | 0.4502 (grid) |
| by | 0.4502 + 3.0 | 3.4502 (grid) |
| bw | 50 × exp(0.5) | 82.44 px |
| bh | 80 × exp(-0.1) | 72.39 px |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
sigma(0.3) = 0.5744; sigma(-0.2) = 0.4502 — both within (0, 1)
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Cell (2, 3) (column 2, row 3 in the 13×13 grid). Anchor prior: p_w=50 px, p_h=80 px. Network outputs tx=0.3, ty=-0.2, tw=0.5, th=-0.1. Decode to pixel box centre.
Section
Part 4 of 5
Concept
A deep CNN backbone down-samples progressively. By stride 32 the feature map is 20×20 on a 640×640 input — each cell covers a 32×32 region. Small objects (e.g., a 10 px pedestrian) disappear entirely in that coarse map.
| Backbone stride | Feature map (640 input) | Cell size | Best for |
|---|---|---|---|
| 8 | 80×80 = 6 400 cells | 8 px | Small objects (<32 px) |
| 16 | 40×40 = 1 600 cells | 16 px | Medium objects (32–96 px) |
| 32 | 20×20 = 400 cells | 32 px | Large objects (>96 px) |
Running detection heads on all three scales in parallel triples the number of predictions but covers the full object-size spectrum — the core FPN idea.
Comparison
Comparison matrix
From Why a single feature scale fails: refill the Best for column from what you know. The rest of the table is as it appeared.
| Backbone stride | Feature map (640 input) | Cell size | Best for |
|---|---|---|---|
| 8 | 80×80 = 6 400 cells | 8 px | Small objects (<32 px) |
| 16 | 40×40 = 1 600 cells | 16 px | Medium objects (32–96 px) |
| 32 | 20×20 = 400 cells | 32 px | Large objects (>96 px) |
Concept
FPN adds a top-down pathway that up-samples the semantically rich (but coarse) deep features and merges them with the spatially precise (but shallow) early features via element-wise addition.
d=256P5 → 2× and add to reduced C4 → P4; repeat to get P3YOLOv3/v5 adopt exactly this design. The (Lesson 92) ViT uses a different approach (no inductive locality bias), but FPN remains the standard for anchor-based detectors.
Counterexample
Discussion prompt
FPN adds a top-down pathway that up-samples the semantically rich (but coarse) deep features and merges them with the spatially precise (but shallow) early features via element-wise addition.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
YOLOv3/v5 adopt exactly this design. The (Lesson 92) ViT uses a different approach (no inductive locality bias), but FPN remains the standard for anchor-based detectors.
Section
Part 5 of 5
Concept
Average Precision (AP) for one class = area under the precision-recall curve. A detection is a TP if IoU(pred, gt) > 0.5 and the GT has not been matched yet; otherwise it is a FP. mAP averages AP across all classes.
\[ \text{AP} = \int_0^1 P(R)\,dR \approx \sum_k (R_k - R_{k-1})\,P_k \]
Sort detections by confidence descending. For each detection, compute cumulative precision and recall. The trapezoid area under the curve is AP. mAP@0.5 uses a single IoU threshold of 0.5 to decide TP vs FP.
Estimation
Predict first
Class 0: 3 GT boxes, 4 detections sorted by confidence. Class 1: 1 GT box, 1 detection (TP). Compute AP per class, then mAP.
Commit before you compute: what does mAP@0.5 — toy two-class example come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Det 3 (conf 0.7) is TP — recall jumps from 0.667 to 1.0 but precision drops to 0.75 (3 TP / 4 det)
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Each TP raises recall; FPs raise denominator without raising numerator, so precision falls.
Worked example
Class 0: 3 GT boxes, 4 detections sorted by confidence. Class 1: 1 GT box, 1 detection (TP). Compute AP per class, then mAP.
import numpy as np
# Class 0: detections ranked by confidence (TP=1, FP=0)
TP = [1, 1, 0, 1] # sorted by confidence desc; 3 GT boxes
n_gt = 3
cum_tp = np.cumsum(TP)
cum_fp = np.cumsum([1-t for t in TP])
prec = cum_tp / (cum_tp + cum_fp)
rec = cum_tp / n_gt
ap0 = np.trapz(prec, rec)
print(f"Precision: {[round(p,3) for p in prec.tolist()]}")
print(f"Recall: {[round(r,3) for r in rec.tolist()]}")
print(f"AP class 0: {ap0:.4f}")
# Class 1: 1 GT, 1 TP — single-point PR curve, AP=0.0 (trapz on 1 pt)
print(f"AP class 1: 0.0000")
print(f"mAP@0.5 = {(ap0 + 0.0) / 2:.4f}")Det 3 (conf 0.7) is TP — recall jumps from 0.667 to 1.0 but precision drops to 0.75 (3 TP / 4 det)
Why: Each TP raises recall; FPs raise denominator without raising numerator, so precision falls. The trapezoidal area captures this trade-off.
| Det rank | TP/FP | Cum TP | Cum FP | Precision | Recall |
|---|---|---|---|---|---|
| 1 (conf 0.9) | TP | 1 | 0 | 1.000 | 0.333 |
| 2 (conf 0.8) | TP | 2 | 0 | 1.000 | 0.667 |
| 3 (conf 0.6) | FP | 2 | 1 | 0.667 | 0.667 |
| 4 (conf 0.4) | TP | 3 | 1 | 0.750 | 1.000 |
| AP (trapezoid) | — | — | — | 0.5694 | mAP = 0.2847 |
Discrimination
Sort into buckets
Sort these by TP/FP, from memory, without looking back at mAP@0.5 — toy two-class example. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Ranking
Put in order
These are the steps of Object detection inference pipeline, scrambled. Put them back in order before the next slide shows you.
(tx, ty, tw, th, conf, class_probs) for every anchor at every cellbx = sigma(tx)+cx, by = sigma(ty)+cy, bw = pw*exp(tw), bh = ph*exp(th)sigma(conf) * max(class_probs) < threshIoU > 0.5) → final boxesWhy: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
(tx, ty, tw, th, conf, class_probs) for every anchor at every cellbx = sigma(tx)+cx, by = sigma(ty)+cy, bw = pw*exp(tw), bh = ph*exp(th)sigma(conf) * max(class_probs) < threshIoU > 0.5) → final boxesEdge cases
Discussion prompt
Object detection inference pipeline works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
(tx, ty, tw, th, conf, class_probs) for every anchor at every cellbx = sigma(tx)+cx, by = sigma(ty)+cy, bw = pw*exp(tw), bh = ph*exp(th)sigma(conf) * max(class_probs) < threshIoU > 0.5) → final boxesElimination
Eliminate the wrong options
Box A = [0, 0, 4, 4] (area 16). Box B = [2, 2, 6, 6] (area 16). What is IoU(A, B)?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Intersection: ix1=max(0,2)=2, iy1=2, ix2=min(4,6)=4, iy2=4. inter=(4-2)*(4-2)=4. Union=16+16-4=28. IoU=4/28≈0.1429.
Check
Work it out before clicking.
Check your understanding
Box A = [0, 0, 4, 4] (area 16). Box B = [2, 2, 6, 6] (area 16). What is IoU(A, B)?
Answer: A
Why: Intersection: ix1=max(0,2)=2, iy1=2, ix2=min(4,6)=4, iy2=4. inter=(4-2)*(4-2)=4. Union=16+16-4=28. IoU=4/28≈0.1429.
Prediction
Predict first
A YOLO model uses S=7, B=2, C=20 (Pascal VOC). How many raw scalar predictions does it produce for one image?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: 1 470
Why: Per cell: B×5 + C = 2×5 + 20 = 30. Total: 7×7×30 = 1 470. Options A and D both equal 1 470 — D shows the working explicitly.
Check
Classic YOLO grid math.
Check your understanding
A YOLO model uses S=7, B=2, C=20 (Pascal VOC). How many raw scalar predictions does it produce for one image?
Answer: A
Why: Per cell: B×5 + C = 2×5 + 20 = 30. Total: 7×7×30 = 1 470. Options A and D both equal 1 470 — D shows the working explicitly.
Elimination
Eliminate the wrong options
NMS has just kept box K (score 0.92). Remaining candidates: box P (score 0.85, IoU(K,P)=0.72) and box Q (score 0.78, IoU(K,Q)=0.35). IoU threshold=0.5. What happens next?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: IoU(K,P)=0.72 > 0.5 → P is a duplicate of K → suppress P. IoU(K,Q)=0.35 < 0.5 → Q is a different object → keep Q in the candidate list.
Check
Trace one step of NMS.
Check your understanding
NMS has just kept box K (score 0.92). Remaining candidates: box P (score 0.85, IoU(K,P)=0.72) and box Q (score 0.78, IoU(K,Q)=0.35). IoU threshold=0.5. What happens next?
Answer: A
Why: IoU(K,P)=0.72 > 0.5 → P is a duplicate of K → suppress P. IoU(K,Q)=0.35 < 0.5 → Q is a different object → keep Q in the candidate list.
Worked example
Build a minimal single-class detector evaluation loop in pure NumPy/PyTorch: implement IoU, NMS, and AP from scratch, then verify against the numbers from this lesson. No pretrained weights — toy synthetic boxes only.
iou(box1, box2): reproduce 0.2162 for b1=[100,150,260,300], b2=[180,200,320,350]nms(boxes, scores, thresh): reproduce kept=[2, 0] for the four-box example in this lessonap = trapz(prec, rec): reproduce AP=0.5694 for class 0bx=2.5744, bw=82.44 for tx=0.3, tw=0.5, cx=2.0, p_w=50Worked example
Reference implementation. Read only after attempting all milestones.
import torch, math, numpy as np
torch.manual_seed(42); np.random.seed(42)
def iou(b1, b2):
ix1,iy1=max(b1[0],b2[0]),max(b1[1],b2[1])
ix2,iy2=min(b1[2],b2[2]),min(b1[3],b2[3])
i=(max(0,ix2-ix1))*(max(0,iy2-iy1))
a1=(b1[2]-b1[0])*(b1[3]-b1[1])
a2=(b2[2]-b2[0])*(b2[3]-b2[1])
u=a1+a2-i
return i/u if u>0 else 0.0
def nms(boxes,scores,t=0.5):
order=sorted(range(len(scores)),
key=lambda i:scores[i],reverse=True)
kept=[]
while order:
i=order.pop(0); kept.append(i)
order=[j for j in order if iou(boxes[i],boxes[j])<t]
return kept
def ap_score(matches, n_gt):
c=np.cumsum(matches); f=np.cumsum([1-m for m in matches])
p=c/(c+f); r=c/n_gt
return float(np.trapz(p,r))
print(iou([100,150,260,300],[180,200,320,350])) # 0.2162
print(nms([[100,150,260,300],[110,155,265,305],
[300,300,400,400],[105,145,258,298]],
[0.9,0.8,0.95,0.7])) # [2, 0]
print(ap_score([1,1,0,1],3)) # 0.5694| Function | Expected output | Verified |
|---|---|---|
| iou(b1, b2) | 0.2162 | yes — inter=8000, union=37000 |
| nms(boxes, scores) | [2, 0] | yes — idx 2 (0.95) first; idx 1 (IoU=0.93) and idx 3 (IoU=0.95) suppressed |
| ap_score([1,1,0,1], 3) | 0.5694 | yes — trapezoid over 4 P/R pairs |
Trade off
Comparison matrix
From Your turn — full program: every row here is a choice with a cost. Fill the Expected output column, then say which row you would actually pick and what you give up for it.
| Function | Expected output | Verified |
|---|---|---|
| iou(b1, b2) | 0.2162 | yes — inter=8000, union=37000 |
| nms(boxes, scores) | [2, 0] | yes — idx 2 (0.95) first; idx 1 (IoU=0.93) and idx 3 (IoU=0.95) suppressed |
| ap_score([1,1,0,1], 3) | 0.5694 | yes — trapezoid over 4 P/R pairs |
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Bounding boxes — the output format · IoU & NMS — de-duplicating predictions · YOLO — grid, anchors, predictions · FPN — multi-scale feature fusion · mAP — evaluating detection quality. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
S²×(B×5+C) and decode (tx,ty,tw,th) via sigmoid/expWant this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.