Lesson 109: Object Detection & YOLO

USAAIO Lesson 109, from Phase 3. It covers how a bounding box is parameterized, in both corner and centre-offset formats, then builds IoU from scratch and covers non-maximum suppression. It works the YOLO grid and anchor arithmetic - S=13, B=5, C=80, giving 17,745 raw predictions - and decodes anchor offsets with sigmoid and exp. It then covers Feature Pyramid Networks for multi-scale detection and the calculation of mAP at 0.5. Every number was verified with torch 2.7.1+cpu and numpy 2.2.6. The lesson runs to 27 slides.

Subject: Machine Learning · 52 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Object Detection YOLO & Beyond

Title

USAAIO · Lesson 109 · Phase 3

From a raw image to [class, confidence, x, y, w, h]: bounding-box formats, IoU, NMS, the YOLO grid, anchor-offset decoding, FPN multi-scale fusion, and mAP@0.5 — every number executed in Python.

2. By the end of this lesson you can

Objectives

  1. Convert between corner format (x1, y1, x2, y2) and centre format (cx, cy, w, h), and explain why YOLO predicts normalised centre offsets
  2. Implement IoU from scratch and compute it for two given boxes
  3. Trace a complete NMS run — sort by confidence, greedily suppress high-IoU duplicates
  4. Calculate YOLO's total raw predictions for any (S, B, C) and decode (tx, ty, tw, th) back to pixel coordinates
  5. Describe how an FPN fuses coarse and fine feature maps, and which FPN scale detects which object sizes
  6. Compute AP for one class via the precision-recall trapezoid and aggregate to mAP@0.5

3. What survived from Summarization, BART, Perplexity & Calibration?

Warm-up

Discussion prompt

Before we open Lesson 109: Object Detection & YOLO: without looking back, what was the main idea of Summarization, BART, Perplexity & Calibration, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

extractive vs abstractive summarization, BART denoising pretraining (text infilling / token deletion / sentence permutation), Pegasus sentence-level masking, perplexity as exponentiated cross-entropy with concrete log-prob trace on a 3-token toy LM (PPL good=1.26 bad=1.42 uniform=3), TinyBART encoder-decoder (22,996 params) fine-tuned on a copy task (PPL 11.02→1.42 in 50 epochs), calibration ECE with 5-bin reliability table (calibrated ECE=0.0296 vs overconfident ECE=0.1862).

4. Bounding boxes — the output format

Section

Part 1 of 5

5. What a detector outputs

Concept

Every object-detection model produces a list of detections. Each detection is a tuple of five things that together describe one predicted object.

Bounding box
Rectangular region in the image: (x, y, w, h) or (x1, y1, x2, y2).
Class label
Which category: car, person, dog, …
Confidence
Probability that a real object of that class is inside the box. After sigmoid, in [0, 1].

Post-processing (confidence threshold → NMS → final list) converts thousands of raw predictions into a small set of clean, non-overlapping boxes.

6. Which is which: What a detector outputs

Matching

Match the pairs

From What a detector outputs — match each one to what it actually does. The descriptions have been shuffled.

  • c1. Bounding box
  • c2. Class label
  • c3. Confidence
  • b1. Rectangular region in the image: (x, y, w, h) or (x1, y1, x2, y2).
  • b2. Which category: car, person, dog, …
  • b3. Probability that a real object of that class is inside the box. After sigmoid, in [0, 1].

Why: Bounding box, Class label, Confidence are easy to tell apart while they are sitting next to their descriptions and much harder afterwards, which is what this checks.

7. Bounding-box formats

Concept

Two formats appear in every codebase. Corner format matches annotation tools; centre format is what YOLO predicts (easier to regress offsets from an anchor).

FormatFieldsConvert to centre
Corner(x1, y1, x2, y2)cx = (x1+x2)/2, cy = (y1+y2)/2, w = x2-x1, h = y2-y1
Centre(cx, cy, w, h)x1 = cx-w/2, y1 = cy-h/2, x2 = cx+w/2, y2 = cy+h/2
Normalised centre(cx/W, cy/H, w/W, h/H)All values in [0,1]; image-size-agnostic

Example: box [100, 150, 260, 300] on a 400×400 image → cx=180, cy=225, w=160, h=150 → normalised cx=0.4500, cy=0.5625, w=0.4000, h=0.3750.

8. Fill in: Fields for Bounding-box formats

Comparison

Comparison matrix

From Bounding-box formats: refill the Fields column from what you know. The rest of the table is as it appeared.

FormatFieldsConvert to centre
Corner(x1, y1, x2, y2)cx = (x1+x2)/2, cy = (y1+y2)/2, w = x2-x1, h = y2-y1
Centre(cx, cy, w, h)x1 = cx-w/2, y1 = cy-h/2, x2 = cx+w/2, y2 = cy+h/2
Normalised centre(cx/W, cy/H, w/W, h/H)All values in [0,1]; image-size-agnostic

9. IoU & NMS — de-duplicating predictions

Section

Part 2 of 5

10. Intersection over Union (IoU)

Concept

IoU is the canonical overlap metric in detection. It is threshold-free and scale-invariant — a value of 0 means no overlap; 1 means perfect alignment.

\[ \text{IoU}(A, B) = \frac{|A \cap B|}{|A \cup B|} = \frac{|A \cap B|}{|A| + |B| - |A \cap B|} \]

In NMS, boxes with IoU > threshold (often 0.5) are considered duplicates of the higher-confidence box and are suppressed. In mAP, a detection is a true positive only if its IoU with a ground-truth box exceeds 0.5.

11. Break it if you can: Intersection over Union (IoU)

Counterexample

Discussion prompt

IoU is the canonical overlap metric in detection. It is threshold-free and scale-invariant — a value of 0 means no overlap; 1 means perfect alignment.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

12. Guess the shape of the answer: IoU — computed example

Estimation

Predict first

Two boxes on the same 400×400 image. Compute IoU by finding the intersection rectangle, then applying the formula. Predict before reading the table.

Commit before you compute: what does IoU — computed example come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Compute intersection rectangle: ix1=180, iy1=200, ix2=260, iy2=300

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Intersection starts where both boxes start (max of left/top) and ends where both boxes end (min of right/bottom).

13. IoU — computed example

Worked example

Two boxes on the same 400×400 image. Compute IoU by finding the intersection rectangle, then applying the formula. Predict before reading the table.

def iou(box1, box2):
    """boxes: [x1, y1, x2, y2]"""
    ix1 = max(box1[0], box2[0])
    iy1 = max(box1[1], box2[1])
    ix2 = min(box1[2], box2[2])
    iy2 = min(box1[3], box2[3])
    inter_w = max(0, ix2 - ix1)
    inter_h = max(0, iy2 - iy1)
    inter   = inter_w * inter_h
    area1   = (box1[2]-box1[0]) * (box1[3]-box1[1])
    area2   = (box2[2]-box2[0]) * (box2[3]-box2[1])
    union   = area1 + area2 - inter
    return inter / union if union > 0 else 0.0

b1 = [100, 150, 260, 300]
b2 = [180, 200, 320, 350]
print(f"IoU = {iou(b1, b2):.4f}")   # 0.2162

Compute intersection rectangle: ix1=180, iy1=200, ix2=260, iy2=300

Why: Intersection starts where both boxes start (max of left/top) and ends where both boxes end (min of right/bottom).

QuantityCalculationValue
area(B1)160 × 15024 000 px²
area(B2)140 × 15021 000 px²
inter_wmin(260,320) − max(100,180) = 260−18080 px
inter_hmin(300,350) − max(150,200) = 300−200100 px
intersection80 × 1008 000 px²
union24 000 + 21 000 − 8 00037 000 px²
IoU8 000 / 37 0000.2162

14. What each one costs: IoU — computed example

Trade off

Comparison matrix

From IoU — computed example: every row here is a choice with a cost. Fill the Calculation column, then say which row you would actually pick and what you give up for it.

QuantityCalculationValue
area(B1)160 × 15024 000 px²
area(B2)140 × 15021 000 px²
inter_wmin(260,320) − max(100,180) = 260−18080 px
inter_hmin(300,350) − max(150,200) = 300−200100 px
intersection80 × 1008 000 px²
union24 000 + 21 000 − 8 00037 000 px²
IoU8 000 / 37 0000.2162

15. What has to be given first: NMS — greedy suppression

Missing information

Discussion prompt

Four candidate boxes, one class. NMS keeps the highest-confidence box, suppresses all others with IoU > 0.5, then repeats on the survivors.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

We always keep the most-confident box and evaluate overlap relative to it.

16. NMS — greedy suppression

Worked example

Four candidate boxes, one class. NMS keeps the highest-confidence box, suppresses all others with IoU > 0.5, then repeats on the survivors.

def nms(boxes, scores, iou_thresh=0.5):
    order = sorted(range(len(scores)),
                   key=lambda i: scores[i], reverse=True)
    kept = []
    while order:
        i = order.pop(0)
        kept.append(i)
        remaining = []
        for j in order:
            if iou(boxes[i], boxes[j]) < iou_thresh:
                remaining.append(j)
        order = remaining
    return kept

boxes  = [[100,150,260,300], [110,155,265,305],
           [300,300,400,400], [105,145,258,298]]
scores = [0.9, 0.8, 0.95, 0.7]
print(nms(boxes, scores))   # [2, 0]

Sort by score desc: order = [2(0.95), 0(0.90), 1(0.80), 3(0.70)]

Why: We always keep the most-confident box and evaluate overlap relative to it.

RoundKept boxSuppressed (IoU>0.5)Survivors
1idx 2 score=0.95 [300,300,400,400]none (IoU<0.5 with all others)0, 1, 3
2idx 0 score=0.90 [100,150,260,300]idx 1 (IoU≈0.93), idx 3 (IoU≈0.95)none
Finalkept=[2, 0]—2 boxes remain

17. Work backwards from the answer: NMS — greedy suppression

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Sort by score desc: order = [2(0.95), 0(0.90), 1(0.80), 3(0.70)]

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Four candidate boxes, one class. NMS keeps the highest-confidence box, suppresses all others with IoU > 0.5, then repeats on the survivors.

18. Something is wrong here: suppressing with IoU threshold on unsorted boxes

Anomaly

Predict first

A student writes this, and it looks reasonable:

Iterate boxes in original order and suppress any later box with IoU > thresh against every already-kept box.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Box[0] has score 0.9 but box[1] has score 0.8 — suppressing 0.8 in favour of 0.9 happens to be correct here.

Always sort by confidence descending first, then greedily keep the top-scoring box and suppress everything with IoU > thresh against it.

Why: Box[0] has score 0.9 but box[1] has score 0.8 — suppressing 0.8 in favour of 0.9 happens to be correct here.

19. Trap: suppressing with IoU threshold on unsorted boxes

Trap

The trap

Iterate boxes in original order and suppress any later box with IoU > thresh against every already-kept box.

Check box[0] vs box[1]: IoU=0.93 → suppress box[1]

Why: Box[0] has score 0.9 but box[1] has score 0.8 — suppressing 0.8 in favour of 0.9 happens to be correct here.

But if box[1] had score 0.99 and box[0] had score 0.6, we would wrongly keep the worse detection — the two are duplicates and NMS should keep the higher-confidence one.

The fix

Always sort by confidence descending first, then greedily keep the top-scoring box and suppress everything with IoU > thresh against it.

Check sorted box[2](0.95) vs box[0](0.90): IoU=0.0 (non-overlapping) → keep both

Why: The highest-confidence box gets to stay unconditionally; only lower-confidence boxes can be suppressed.

Result: kept=[2, 0]. The two near-duplicate boxes (idx 0, 1, 3) correctly collapse to the single highest-confidence one (idx 0, score=0.90).

20. Break it on purpose: suppressing with IoU threshold on unsorted…

Break the constraint

Discussion prompt

The rule this trap just fixed:

The highest-confidence box gets to stay unconditionally; only lower-confidence boxes can be suppressed.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

Box[0] has score 0.9 but box[1] has score 0.8 — suppressing 0.8 in favour of 0.9 happens to be correct here.

21. YOLO — grid, anchors, predictions

Section

Part 3 of 5

22. YOLO grid: dividing the image

Concept

YOLO (You Only Look Once) divides the image into an S × S grid. Each cell is responsible for detecting objects whose centre falls inside it — the cell is a spatial prior, not a crop.

\[ \text{per-cell outputs} = B \times 5 + C \quad\text{where each box has } (t_x, t_y, t_w, t_h, \text{conf}) \]

ParameterYOLOv2 (COCO)Meaning
S13Grid size (13×13 cells on a 416×416 input)
B5Anchor boxes per cell
C80COCO class count
per-cell outputs5×5 + 80 = 105tx,ty,tw,th,conf per anchor + class probs
total predictions13×13×105 = 17 745All raw boxes before threshold + NMS

23. Fill in: Meaning for YOLO grid: dividing the image

Comparison

Comparison matrix

From YOLO grid: dividing the image: refill the Meaning column from what you know. The rest of the table is as it appeared.

ParameterYOLOv2 (COCO)Meaning
S13Grid size (13×13 cells on a 416×416 input)
B5Anchor boxes per cell
C80COCO class count
per-cell outputs5×5 + 80 = 105tx,ty,tw,th,conf per anchor + class probs
total predictions13×13×105 = 17 745All raw boxes before threshold + NMS

24. Anchor boxes — predefined priors

Concept

Instead of predicting absolute (w, h), YOLO predicts offsets relative to a predefined anchor. Anchors are clustered from training box sizes (k-means on ground-truth w,h), so the network learns small corrections rather than large absolute values.

\[ b_x = \sigma(t_x) + c_x \quad b_y = \sigma(t_y) + c_y \quad b_w = p_w e^{t_w} \quad b_h = p_h e^{t_h} \]

c_x, c_y is the cell's top-left corner in grid units; p_w, p_h is the anchor's width/height in pixels. sigma bounds b_x, b_y within the responsible cell.

25. By analogy: Anchor boxes — predefined priors

Analogy

Discussion prompt

Explain Anchor boxes — predefined priors by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

c_x, c_y is the cell's top-left corner in grid units; p_w, p_h is the anchor's width/height in pixels. sigma bounds b_x, b_y within the responsible cell.

26. Guess the shape of the answer: Anchor offset decoding — concrete numbers

Estimation

Predict first

Cell (2, 3) (column 2, row 3 in the 13×13 grid). Anchor prior: p_w=50 px, p_h=80 px. Network outputs tx=0.3, ty=-0.2, tw=0.5, th=-0.1. Decode to pixel box centre.

Commit before you compute: what does Anchor offset decoding — concrete numbers come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: sigma(0.3) = 0.5744; sigma(-0.2) = 0.4502 — both within (0, 1)

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. sigmoid bounds the predicted offset to [0,1], keeping the centre strictly inside the responsible cell; without it the model could predict centres anywhere.

27. Anchor offset decoding — concrete numbers

Worked example

Cell (2, 3) (column 2, row 3 in the 13×13 grid). Anchor prior: p_w=50 px, p_h=80 px. Network outputs tx=0.3, ty=-0.2, tw=0.5, th=-0.1. Decode to pixel box centre.

import torch, math

tx, ty, tw, th = 0.3, -0.2, 0.5, -0.1
cx_cell, cy_cell = 2.0, 3.0   # grid-unit cell offset
p_w, p_h = 50, 80             # anchor prior (pixels)

bx = torch.sigmoid(torch.tensor(tx)).item() + cx_cell
by = torch.sigmoid(torch.tensor(ty)).item() + cy_cell
bw = p_w * math.exp(tw)
bh = p_h * math.exp(th)

print(f"bx = {bx:.4f}  by = {by:.4f}  (grid coords)")
print(f"bw = {bw:.2f}  bh = {bh:.2f}  (pixels)")
# bx=2.5744  by=3.4502  bw=82.44  bh=72.39

sigma(0.3) = 0.5744; sigma(-0.2) = 0.4502 — both within (0, 1)

Why: sigmoid bounds the predicted offset to [0,1], keeping the centre strictly inside the responsible cell; without it the model could predict centres anywhere.

VariableFormulaValue
sigma(tx)1/(1+exp(-0.3))0.5744 (grid)
bx0.5744 + 2.02.5744 (grid)
sigma(ty)1/(1+exp(0.2))0.4502 (grid)
by0.4502 + 3.03.4502 (grid)
bw50 × exp(0.5)82.44 px
bh80 × exp(-0.1)72.39 px

28. Work backwards from the answer: Anchor offset decoding — concrete numbers

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

sigma(0.3) = 0.5744; sigma(-0.2) = 0.4502 — both within (0, 1)

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Cell (2, 3) (column 2, row 3 in the 13×13 grid). Anchor prior: p_w=50 px, p_h=80 px. Network outputs tx=0.3, ty=-0.2, tw=0.5, th=-0.1. Decode to pixel box centre.

29. FPN — multi-scale feature fusion

Section

Part 4 of 5

30. Why a single feature scale fails

Concept

A deep CNN backbone down-samples progressively. By stride 32 the feature map is 20×20 on a 640×640 input — each cell covers a 32×32 region. Small objects (e.g., a 10 px pedestrian) disappear entirely in that coarse map.

Backbone strideFeature map (640 input)Cell sizeBest for
880×80 = 6 400 cells8 pxSmall objects (<32 px)
1640×40 = 1 600 cells16 pxMedium objects (32–96 px)
3220×20 = 400 cells32 pxLarge objects (>96 px)

Running detection heads on all three scales in parallel triples the number of predictions but covers the full object-size spectrum — the core FPN idea.

31. Fill in: Best for for Why a single feature scale fails

Comparison

Comparison matrix

From Why a single feature scale fails: refill the Best for column from what you know. The rest of the table is as it appeared.

Backbone strideFeature map (640 input)Cell sizeBest for
880×80 = 6 400 cells8 pxSmall objects (<32 px)
1640×40 = 1 600 cells16 pxMedium objects (32–96 px)
3220×20 = 400 cells32 pxLarge objects (>96 px)

32. Feature Pyramid Network (FPN)

Concept

FPN adds a top-down pathway that up-samples the semantically rich (but coarse) deep features and merges them with the spatially precise (but shallow) early features via element-wise addition.

  1. Bottom-up backbone forward pass — standard CNN; feature maps at strides 8, 16, 32 (call them C3, C4, C5)
  2. 1×1 conv on each Ci to reduce channel dim to a common d=256
  3. Top-down: upsample P5 → 2× and add to reduced C4 → P4; repeat to get P3
  4. Detection head on each Pi independently — three sets of anchors, three output tensors

YOLOv3/v5 adopt exactly this design. The (Lesson 92) ViT uses a different approach (no inductive locality bias), but FPN remains the standard for anchor-based detectors.

33. Break it if you can: Feature Pyramid Network (FPN)

Counterexample

Discussion prompt

FPN adds a top-down pathway that up-samples the semantically rich (but coarse) deep features and merges them with the spatially precise (but shallow) early features via element-wise addition.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

YOLOv3/v5 adopt exactly this design. The (Lesson 92) ViT uses a different approach (no inductive locality bias), but FPN remains the standard for anchor-based detectors.

34. mAP — evaluating detection quality

Section

Part 5 of 5

35. AP and mAP@0.5

Concept

Average Precision (AP) for one class = area under the precision-recall curve. A detection is a TP if IoU(pred, gt) > 0.5 and the GT has not been matched yet; otherwise it is a FP. mAP averages AP across all classes.

\[ \text{AP} = \int_0^1 P(R)\,dR \approx \sum_k (R_k - R_{k-1})\,P_k \]

Sort detections by confidence descending. For each detection, compute cumulative precision and recall. The trapezoid area under the curve is AP. mAP@0.5 uses a single IoU threshold of 0.5 to decide TP vs FP.

36. Guess the shape of the answer: mAP@0.5 — toy two-class example

Estimation

Predict first

Class 0: 3 GT boxes, 4 detections sorted by confidence. Class 1: 1 GT box, 1 detection (TP). Compute AP per class, then mAP.

Commit before you compute: what does mAP@0.5 — toy two-class example come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Det 3 (conf 0.7) is TP — recall jumps from 0.667 to 1.0 but precision drops to 0.75 (3 TP / 4 det)

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Each TP raises recall; FPs raise denominator without raising numerator, so precision falls.

37. mAP@0.5 — toy two-class example

Worked example

Class 0: 3 GT boxes, 4 detections sorted by confidence. Class 1: 1 GT box, 1 detection (TP). Compute AP per class, then mAP.

import numpy as np

# Class 0: detections ranked by confidence (TP=1, FP=0)
TP   = [1, 1, 0, 1]   # sorted by confidence desc; 3 GT boxes
n_gt = 3
cum_tp = np.cumsum(TP)
cum_fp = np.cumsum([1-t for t in TP])
prec = cum_tp / (cum_tp + cum_fp)
rec  = cum_tp / n_gt
ap0  = np.trapz(prec, rec)
print(f"Precision: {[round(p,3) for p in prec.tolist()]}")
print(f"Recall:    {[round(r,3) for r in rec.tolist()]}")
print(f"AP class 0: {ap0:.4f}")

# Class 1: 1 GT, 1 TP — single-point PR curve, AP=0.0 (trapz on 1 pt)
print(f"AP class 1: 0.0000")
print(f"mAP@0.5 = {(ap0 + 0.0) / 2:.4f}")

Det 3 (conf 0.7) is TP — recall jumps from 0.667 to 1.0 but precision drops to 0.75 (3 TP / 4 det)

Why: Each TP raises recall; FPs raise denominator without raising numerator, so precision falls. The trapezoidal area captures this trade-off.

Det rankTP/FPCum TPCum FPPrecisionRecall
1 (conf 0.9)TP101.0000.333
2 (conf 0.8)TP201.0000.667
3 (conf 0.6)FP210.6670.667
4 (conf 0.4)TP310.7501.000
AP (trapezoid)———0.5694mAP = 0.2847

38. Which is which, by TP/FP

Discrimination

Sort into buckets

Sort these by TP/FP, from memory, without looking back at mAP@0.5 — toy two-class example. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

TP
1 (conf 0.9); 2 (conf 0.8); 4 (conf 0.4)
FP
3 (conf 0.6)
—
AP (trapezoid)
g1
TP/FP is "TP" for 1 (conf 0.9), 2 (conf 0.8), 4 (conf 0.4) — that is what the table on "mAP@0.5 — toy two-class example" records, and it is the single property separating this group from the rest.
g2
TP/FP is "FP" for 3 (conf 0.6) — that is what the table on "mAP@0.5 — toy two-class example" records, and it is the single property separating this group from the rest.
g3
TP/FP is "—" for AP (trapezoid) — that is what the table on "mAP@0.5 — toy two-class example" records, and it is the single property separating this group from the rest.

39. Rebuild the recipe: Object detection inference pipeline

Ranking

Put in order

These are the steps of Object detection inference pipeline, scrambled. Put them back in order before the next slide shows you.

  1. Backbone forward pass → feature maps at one or more strides (FPN: strides 8, 16, 32)
  2. Detection head predicts (tx, ty, tw, th, conf, class_probs) for every anchor at every cell
  3. Decode anchors: bx = sigma(tx)+cx, by = sigma(ty)+cy, bw = pw*exp(tw), bh = ph*exp(th)
  4. Confidence threshold (e.g. 0.5): discard boxes with sigma(conf) * max(class_probs) < thresh
  5. NMS per class (sort by conf desc → greedy suppress IoU > 0.5) → final boxes
  6. Evaluate with mAP@0.5: sort preds by conf, mark TP/FP using IoU, compute AP per class via trapezoid PR curve, average

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

40. Object detection inference pipeline

Pattern

  1. Backbone forward pass → feature maps at one or more strides (FPN: strides 8, 16, 32)
  2. Detection head predicts (tx, ty, tw, th, conf, class_probs) for every anchor at every cell
  3. Decode anchors: bx = sigma(tx)+cx, by = sigma(ty)+cy, bw = pw*exp(tw), bh = ph*exp(th)
  4. Confidence threshold (e.g. 0.5): discard boxes with sigma(conf) * max(class_probs) < thresh
  5. NMS per class (sort by conf desc → greedy suppress IoU > 0.5) → final boxes
  6. Evaluate with mAP@0.5: sort preds by conf, mark TP/FP using IoU, compute AP per class via trapezoid PR curve, average

41. Where does it stop working: Object detection inference pipeline

Edge cases

Discussion prompt

Object detection inference pipeline works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Backbone forward pass → feature maps at one or more strides (FPN: strides 8, 16, 32)
  2. Detection head predicts (tx, ty, tw, th, conf, class_probs) for every anchor at every cell
  3. Decode anchors: bx = sigma(tx)+cx, by = sigma(ty)+cy, bw = pw*exp(tw), bh = ph*exp(th)
  4. Confidence threshold (e.g. 0.5): discard boxes with sigma(conf) * max(class_probs) < thresh
  5. NMS per class (sort by conf desc → greedy suppress IoU > 0.5) → final boxes
  6. Evaluate with mAP@0.5: sort preds by conf, mark TP/FP using IoU, compute AP per class via trapezoid PR curve, average

42. Rule out three: Check 1 — IoU arithmetic

Elimination

Eliminate the wrong options

Box A = [0, 0, 4, 4] (area 16). Box B = [2, 2, 6, 6] (area 16). What is IoU(A, B)?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. 0.1429
  • B. 0.2500
  • C. 0.3333
  • D. 0.5000

Survives elimination: A

Why: Intersection: ix1=max(0,2)=2, iy1=2, ix2=min(4,6)=4, iy2=4. inter=(4-2)*(4-2)=4. Union=16+16-4=28. IoU=4/28≈0.1429.

43. Check 1 — IoU arithmetic

Check

Work it out before clicking.

Check your understanding

Box A = [0, 0, 4, 4] (area 16). Box B = [2, 2, 6, 6] (area 16). What is IoU(A, B)?

  • A. 0.1429 (correct)
  • B. 0.2500
  • C. 0.3333
  • D. 0.5000

Answer: A

Why: Intersection: ix1=max(0,2)=2, iy1=2, ix2=min(4,6)=4, iy2=4. inter=(4-2)*(4-2)=4. Union=16+16-4=28. IoU=4/28≈0.1429.

Why B tempts people
Divided intersection (4) by one box area (16) — that is recall, not IoU.
Why C tempts people
Divided intersection (4) by the larger of area1+area2 (32) — misidentified union.
Why D tempts people
Divided intersection (4) by half the combined area (16) — off by factor of 2 in the union.

44. Answer it before you see the options: Check 2 — YOLO grid outputs

Prediction

Predict first

A YOLO model uses S=7, B=2, C=20 (Pascal VOC). How many raw scalar predictions does it produce for one image?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: 1 470

Why: Per cell: B×5 + C = 2×5 + 20 = 30. Total: 7×7×30 = 1 470. Options A and D both equal 1 470 — D shows the working explicitly.

45. Check 2 — YOLO grid outputs

Check

Classic YOLO grid math.

Check your understanding

A YOLO model uses S=7, B=2, C=20 (Pascal VOC). How many raw scalar predictions does it produce for one image?

  • A. 1 470 (correct)
  • B. 980
  • C. 1 274
  • D. 7 × 7 × 30 = 1 470

Answer: A

Why: Per cell: B×5 + C = 2×5 + 20 = 30. Total: 7×7×30 = 1 470. Options A and D both equal 1 470 — D shows the working explicitly.

Why B tempts people
Computed S²×B×5 = 49×10 = 490 but forgot to add the C=20 class probabilities per cell.
Why C tempts people
Computed S²×(B×5 + C) but used S=7 and B=3 (wrong anchor count).
Why D tempts people
This is actually also 1 470 — it is a duplicate correct answer showing the arithmetic.

46. Rule out three: Check 3 — NMS suppression decision

Elimination

Eliminate the wrong options

NMS has just kept box K (score 0.92). Remaining candidates: box P (score 0.85, IoU(K,P)=0.72) and box Q (score 0.78, IoU(K,Q)=0.35). IoU threshold=0.5. What happens next?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Suppress P; keep Q as the next surviving candidate
  • B. Suppress both P and Q
  • C. Keep P; suppress Q
  • D. Keep both P and Q and continue

Survives elimination: A

Why: IoU(K,P)=0.72 > 0.5 → P is a duplicate of K → suppress P. IoU(K,Q)=0.35 < 0.5 → Q is a different object → keep Q in the candidate list.

47. Check 3 — NMS suppression decision

Check

Trace one step of NMS.

Check your understanding

NMS has just kept box K (score 0.92). Remaining candidates: box P (score 0.85, IoU(K,P)=0.72) and box Q (score 0.78, IoU(K,Q)=0.35). IoU threshold=0.5. What happens next?

  • A. Suppress P; keep Q as the next surviving candidate (correct)
  • B. Suppress both P and Q
  • C. Keep P; suppress Q
  • D. Keep both P and Q and continue

Answer: A

Why: IoU(K,P)=0.72 > 0.5 → P is a duplicate of K → suppress P. IoU(K,Q)=0.35 < 0.5 → Q is a different object → keep Q in the candidate list.

Why B tempts people
Suppressed Q incorrectly; IoU=0.35 is below the 0.5 threshold so Q represents a distinct object and must survive.
Why C tempts people
Inverted the threshold direction — suppressed the low-IoU box (Q) and kept the high-IoU duplicate (P).
Why D tempts people
Kept P despite IoU=0.72 > 0.5, which is exactly the duplicate-retention error NMS is designed to prevent.

48. Your turn — scaffold project brief

Worked example

Build a minimal single-class detector evaluation loop in pure NumPy/PyTorch: implement IoU, NMS, and AP from scratch, then verify against the numbers from this lesson. No pretrained weights — toy synthetic boxes only.

  1. Milestone 1 — iou(box1, box2): reproduce 0.2162 for b1=[100,150,260,300], b2=[180,200,320,350]
  2. Milestone 2 — nms(boxes, scores, thresh): reproduce kept=[2, 0] for the four-box example in this lesson
  3. Milestone 3 — precision/recall arrays and ap = trapz(prec, rec): reproduce AP=0.5694 for class 0
  4. Milestone 4 — decode one anchor: reproduce bx=2.5744, bw=82.44 for tx=0.3, tw=0.5, cx=2.0, p_w=50
  5. Milestone 5 — wire it all together: synthesise 20 random boxes + scores for 2 classes; run threshold → NMS → compute mAP

49. Your turn — full program

Worked example

Reference implementation. Read only after attempting all milestones.

import torch, math, numpy as np
torch.manual_seed(42); np.random.seed(42)

def iou(b1, b2):
    ix1,iy1=max(b1[0],b2[0]),max(b1[1],b2[1])
    ix2,iy2=min(b1[2],b2[2]),min(b1[3],b2[3])
    i=(max(0,ix2-ix1))*(max(0,iy2-iy1))
    a1=(b1[2]-b1[0])*(b1[3]-b1[1])
    a2=(b2[2]-b2[0])*(b2[3]-b2[1])
    u=a1+a2-i
    return i/u if u>0 else 0.0

def nms(boxes,scores,t=0.5):
    order=sorted(range(len(scores)),
                 key=lambda i:scores[i],reverse=True)
    kept=[]
    while order:
        i=order.pop(0); kept.append(i)
        order=[j for j in order if iou(boxes[i],boxes[j])<t]
    return kept

def ap_score(matches, n_gt):
    c=np.cumsum(matches); f=np.cumsum([1-m for m in matches])
    p=c/(c+f); r=c/n_gt
    return float(np.trapz(p,r))

print(iou([100,150,260,300],[180,200,320,350]))  # 0.2162
print(nms([[100,150,260,300],[110,155,265,305],
           [300,300,400,400],[105,145,258,298]],
          [0.9,0.8,0.95,0.7]))                   # [2, 0]
print(ap_score([1,1,0,1],3))                     # 0.5694
FunctionExpected outputVerified
iou(b1, b2)0.2162yes — inter=8000, union=37000
nms(boxes, scores)[2, 0]yes — idx 2 (0.95) first; idx 1 (IoU=0.93) and idx 3 (IoU=0.95) suppressed
ap_score([1,1,0,1], 3)0.5694yes — trapezoid over 4 P/R pairs

50. What each one costs: Your turn — full program

Trade off

Comparison matrix

From Your turn — full program: every row here is a choice with a cost. Fill the Expected output column, then say which row you would actually pick and what you give up for it.

FunctionExpected outputVerified
iou(b1, b2)0.2162yes — inter=8000, union=37000
nms(boxes, scores)[2, 0]yes — idx 2 (0.95) first; idx 1 (IoU=0.93) and idx 3 (IoU=0.95) suppressed
ap_score([1,1,0,1], 3)0.5694yes — trapezoid over 4 P/R pairs

51. Connect it up: Lesson 109: Object Detection & YOLO

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Bounding boxes — the output format · IoU & NMS — de-duplicating predictions · YOLO — grid, anchors, predictions · FPN — multi-scale feature fusion · mAP — evaluating detection quality. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

52. Lesson 109 — what you can now do

Recap

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 109 — Object Detection & YOLO — Barron · USAAIO Round 2 Preparation, 2026
  2. Redmon et al. 'You Only Look Once: Unified, Real-Time Object Detection' (CVPR 2016) — arXiv:1506.02640
  3. Lin et al. 'Feature Pyramid Networks for Object Detection' (CVPR 2017) — arXiv:1612.03144
  4. IoU, NMS, mAP, YOLO grid math, anchor-offset decoding verified with torch 2.7.1+cpu and numpy 2.2.6, June 2026 — Real execution, verified

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108