USAAIO Lesson 26, from Week 9, fully worked on one running 12-email spam classifier and one six-house regression set. It builds the confusion matrix count by count, derives precision, recall, and F1 and computes them by hand, and explains why accuracy lies on imbalanced data. It sweeps the precision-recall trade-off across thresholds, derives ROC and AUC as a ranking probability - 31 of 35 pairs, by hand - and confirms it against sklearn, then covers AUC-PR and log-loss with a confident-but-wrong blow-up. It then computes MSE, RMSE, MAE, R2, and MAPE entry by entry, splits RMSE from MAE with an outlier, and gives each generation metric - perplexity, BLEU, ROUGE, and FID - a runnable toy example. It ends with a from-scratch metrics library verified against sklearn. Every number was produced by real execution. The lesson runs to 61 slides.
Subject: Machine Learning · 119 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 26 · Week 9
The number you optimize decides which model you ship. We build every metric on one running spam classifier and one house-price set — confusion counts, precision/recall/F1, the ROC curve and AUC as a ranking probability, log-loss, then MSE/RMSE/MAE/R² and the generation metrics — each derived by hand and confirmed against sklearn.
Objectives
P(random positive ranked above random negative)sklearnWarm-up
Discussion prompt
Before we open Lesson 26: Evaluation Metrics: without looking back, what was the main idea of Projections & the Hat Matrix, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
orthogonal projection matrices (P²=P, Pᵀ=P), the hat matrix H=X(XᵀX)⁻¹Xᵀ that projects y onto the column space, the residual projector I−H, leverage scores, and Cook's distance. Build the hat matrix and leverage scores from scratch.
Section
Part 1 of 6 — the four counts
Concept
One dataset carries the whole classification half of this lesson: 12 emails. Each has a true label y (1 = spam, 0 = ham) and the model's score p = its estimated P(spam). Five are truly spam, seven are ham.
| true y | score p | |
|---|---|---|
| e1 | 1 (spam) | 0.95 |
| e2 | 1 (spam) | 0.90 |
| e3 | 1 (spam) | 0.80 |
| e4 | 1 (spam) | 0.60 |
| e5 | 1 (spam) | 0.30 |
| e6 | 0 (ham) | 0.70 |
| e7 | 0 (ham) | 0.55 |
| e8..e12 | 0 (ham) | 0.40, 0.20, 0.15, 0.10, 0.05 |
A score is not yet a decision. To turn p into a predicted label we pick a threshold — flag as spam when p ≥ t. We start with the default t = 0.5.
Comparison
Comparison matrix
From The running example: a spam classifier: refill the true y column from what you know. The rest of the table is as it appeared.
| true y | score p | |
|---|---|---|
| e1 | 1 (spam) | 0.95 |
| e2 | 1 (spam) | 0.90 |
| e3 | 1 (spam) | 0.80 |
| e4 | 1 (spam) | 0.60 |
| e5 | 1 (spam) | 0.30 |
| e6 | 0 (ham) | 0.70 |
| e7 | 0 (ham) | 0.55 |
| e8..e12 | 0 (ham) | 0.40, 0.20, 0.15, 0.10, 0.05 |
Estimation
Predict first
Apply pred = (p ≥ 0.5). Two of the five spam emails clear 0.5 easily (0.95, 0.90, 0.80, 0.60) but e5 at 0.30 does not; two ham emails (e6 0.70, e7 0.55) wrongly clear it. This is our prediction vector.
Commit before you compute: what does Threshold the scores at t = 0.5 come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: pred = [1 1 1 1 0 1 1 0 0 0 0 0]
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Four of five spam correctly flagged (e1–e4), e5 missed; two ham wrongly flagged (e6, e7).
Worked example
Apply pred = (p ≥ 0.5). Two of the five spam emails clear 0.5 easily (0.95, 0.90, 0.80, 0.60) but e5 at 0.30 does not; two ham emails (e6 0.70, e7 0.55) wrongly clear it. This is our prediction vector.
import numpy as np
y = np.array([1,1,1,1,1,0,0,0,0,0,0,0])
p = np.array([0.95,0.90,0.80,0.60,0.30,
0.70,0.55,0.40,0.20,0.15,0.10,0.05])
pred = (p >= 0.5).astype(int)
print(pred)pred = [1 1 1 1 0 1 1 0 0 0 0 0]
Why: Four of five spam correctly flagged (e1–e4), e5 missed; two ham wrongly flagged (e6, e7). Every metric below is a different summary of exactly these agreements and disagreements.
| y | p | p ≥ 0.5 → pred | |
|---|---|---|---|
| e1 | 1 | 0.95 | 1 ✓ |
| e4 | 1 | 0.60 | 1 ✓ |
| e5 | 1 | 0.30 | 0 ✗ missed spam |
| e6 | 0 | 0.70 | 1 ✗ flagged ham |
| e12 | 0 | 0.05 | 0 ✓ |
Trade off
Comparison matrix
From Threshold the scores at t = 0.5: every row here is a choice with a cost. Fill the p ≥ 0.5 → pred column, then say which row you would actually pick and what you give up for it.
| y | p | p ≥ 0.5 → pred | |
|---|---|---|---|
| e1 | 1 | 0.95 | 1 ✓ |
| e4 | 1 | 0.60 | 1 ✓ |
| e5 | 1 | 0.30 | 0 ✗ missed spam |
| e6 | 0 | 0.70 | 1 ✗ flagged ham |
| e12 | 0 | 0.05 | 0 ✓ |
Concept
Every (prediction, truth) pair falls into exactly one of four boxes. Positive/negative names what the model said; true/false names whether it was right.
| actual spam (+) | actual ham (−) | |
|---|---|---|
| predict spam (+) | TP — caught spam | FP — false alarm |
| predict ham (−) | FN — missed spam | TN — correct pass |
FP is a false alarm (real mail trashed); FN is a miss (spam let through). These two error types have different real-world costs — that asymmetry drives every choice in this lesson.
Counterexample
Discussion prompt
Every (prediction, truth) pair falls into exactly one of four boxes. Positive/negative names what the model said; true/false names whether it was right.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Missing information
Discussion prompt
A boolean mask like (pred==1)&(y==1) is True exactly on the true positives; .sum() counts the Trues. Build all four counts from pred and y:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
e1–e4 are true positives; e6, e7 are false positives; e5 is the lone false negative; the other five ham are true negatives. The four counts sum to 12 — every email lands in exactly one box.
Worked example
A boolean mask like (pred==1)&(y==1) is True exactly on the true positives; .sum() counts the Trues. Build all four counts from pred and y:
import numpy as np
y = np.array([1,1,1,1,1,0,0,0,0,0,0,0])
pred = np.array([1,1,1,1,0,1,1,0,0,0,0,0])
TP = ((pred==1)&(y==1)).sum(); FP = ((pred==1)&(y==0)).sum()
FN = ((pred==0)&(y==1)).sum(); TN = ((pred==0)&(y==0)).sum()
print(int(TP), int(FP), int(FN), int(TN))TP=4, FP=2, FN=1, TN=5
Why: e1–e4 are true positives; e6, e7 are false positives; e5 is the lone false negative; the other five ham are true negatives. The four counts sum to 12 — every email lands in exactly one box.
| count | which emails | value |
|---|---|---|
| TP | e1, e2, e3, e4 | 4 |
| FP | e6, e7 | 2 |
| FN | e5 | 1 |
| TN | e8, e9, e10, e11, e12 | 5 |
| total | all | 12 |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
TP=4, FP=2, FN=1, TN=5
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
A boolean mask like (pred==1)&(y==1) is True exactly on the true positives; .sum() counts the Trues. Build all four counts from pred and y:
Intuition
The four names decode cleanly once you split them. The second word is what the model predicted — 'positive' means it flagged spam, 'negative' means it passed the email through.
The first word says whether that prediction was correct: 'true' means the model was right, 'false' means it was wrong. So a false positive is a wrong 'spam' flag, and a false negative is a wrong 'ham' pass.
Whenever a metric confuses you, expand its counts this way. Precision, recall, and every rate below are just ratios of these four boxes — nothing more.
Analogy
Discussion prompt
Explain Reading the names: true/false, positive/negative by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
The four names decode cleanly once you split them. The second word is what the model predicted — 'positive' means it flagged spam, 'negative' means it passed the email through.
Section
Part 2 of 6 — from counts to scores
Concept
Precision looks only at what you flagged: of the emails I called spam, how many really were? Recall looks only at the real spam: of all actual spam, how many did I catch?
\[ \text{precision} = \frac{TP}{TP + FP}, \qquad \text{recall} = \frac{TP}{TP + FN} \]
Precision's enemy is the false alarm (FP); recall's enemy is the miss (FN). Same TP on top — they differ only in which error type sits in the denominator.
Explain it
Discussion prompt
Explain Precision and recall: two different questions to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Precision's enemy is the false alarm (FP); recall's enemy is the miss (FN). Same TP on top — they differ only in which error type sits in the denominator.
Intuition
Either metric alone is trivial to game. Flag every email as spam and you catch all real spam — recall = 1.0 — but almost everything you flag is wrong, so precision collapses.
Flag only the single email you are most certain about and precision is likely 1.0 — but you miss all the rest, so recall collapses. Any honest score has to look at both at once.
A spam filter leans toward precision (never trash real mail); a cancer screen leans toward recall (never miss a case). The cost of FP vs FN decides the balance — the metric just makes that choice explicit.
Pattern
Predict first
The table runs: precision | TP/(TP+FP) | 4/6 | 0.6667 · recall | TP/(TP+FN) | 4/5 | 0.8000
In Precision and recall on our counts, given the rows so far: what is the next one — the row where metric is accuracy?
Correct: accuracy | (TP+TN)/n | 9/12 | 0.7500
| metric | formula | arithmetic | value |
|---|---|---|---|
| precision | TP/(TP+FP) | 4/6 | 0.6667 |
| recall | TP/(TP+FN) | 4/5 | 0.8000 |
| accuracy | (TP+TN)/n | 9/12 | 0.7500 |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. Of the 6 emails we flagged as spam, 4 truly were — the 2 false alarms (e6, e7) drag precision down to two-thirds.
Worked example
Plug TP=4, FP=2, FN=1 straight into the formulas — no code needed yet, just the two divisions.
precision = TP/(TP+FP) = 4/(4+2) = 4/6 = 0.6667
Why: Of the 6 emails we flagged as spam, 4 truly were — the 2 false alarms (e6, e7) drag precision down to two-thirds.
recall = TP/(TP+FN) = 4/(4+1) = 4/5 = 0.8
Why: Of the 5 real spam emails, we caught 4 — only e5 (score 0.30) slipped through, so recall is 0.8.
| metric | formula | arithmetic | value |
|---|---|---|---|
| precision | TP/(TP+FP) | 4/6 | 0.6667 |
| recall | TP/(TP+FN) | 4/5 | 0.8000 |
| accuracy | (TP+TN)/n | 9/12 | 0.7500 |
Pattern
Step through it
Step through Precision and recall on our counts one row at a time. What is driving the change, and what would the row after the last one be?
Concept
F1 is the harmonic mean of precision P and recall R — it collapses the pair into a single score that is high only when both are high.
\[ F_1 = \frac{2}{\tfrac{1}{P} + \tfrac{1}{R}} = \frac{2PR}{P + R} \]
The harmonic mean is dominated by the smaller term: if either P or R is near zero, F1 is near zero. That is exactly why the 'flag everything' cheat (recall 1, precision tiny) scores a terrible F1.
Ranking
Put in order
Put the moves of F1 by hand, then verified in code into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Multiply the two rates, double it.
Worked example
Numerator: 2PR = 2·(4/6)·(4/5) = 2·(16/30) = 32/30
Why: Multiply the two rates, double it. Keep fractions to stay exact.
Denominator: P + R = 4/6 + 4/5 = 20/30 + 24/30 = 44/30
Why: Common denominator 30. The 30s cancel in the ratio.
F1 = (32/30)/(44/30) = 32/44 = 0.7273
Why: F1 sits between recall (0.8) and precision (0.667), pulled toward the smaller — the signature of a harmonic mean.
import numpy as np
y = np.array([1,1,1,1,1,0,0,0,0,0,0,0])
pred = np.array([1,1,1,1,0,1,1,0,0,0,0,0])
TP = ((pred==1)&(y==1)).sum(); FP = ((pred==1)&(y==0)).sum()
FN = ((pred==0)&(y==1)).sum(); TN = ((pred==0)&(y==0)).sum()
prec = TP/(TP+FP); rec = TP/(TP+FN)
f1 = 2*prec*rec/(prec+rec)
print(int(TP), int(FP), int(FN), int(TN))
print(round(prec,4), round(rec,4), round(f1,4))| printed | value |
|---|---|
| TP FP FN TN | 4 2 1 5 |
| precision | 0.6667 |
| recall | 0.8 |
| F1 | 0.7273 |
Error analysis
Annotate
Walk the callouts on F1 by hand, then verified in code. Each one is a place this is easy to get subtly wrong.
Ranking
Put in order
These are the steps of From confusion matrix to metrics, scrambled. Put them back in order before the next slide shows you.
pred = (p ≥ t)TP, FP, FN, TN (they must sum to n)TP/(TP+FP) — punishes false alarmsTP/(TP+FN) — punishes misses2PR/(P+R) — the harmonic mean, high only when both areWhy: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
pred = (p ≥ t)TP, FP, FN, TN (they must sum to n)TP/(TP+FP) — punishes false alarmsTP/(TP+FN) — punishes misses2PR/(P+R) — the harmonic mean, high only when both areAnomaly
Predict first
A student writes this, and it looks reasonable:
A fraud detector reports 99% accuracy on a stream that is 99% legitimate. Ship it — it's almost never wrong.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Predicting 'negative' for EVERYONE scores 99% accuracy while catching zero fraud.
On imbalanced data report recall, F1, or AUC-PR — metrics that stare at the minority class instead of being drowned out by the majority.
Why: Predicting 'negative' for EVERYONE scores 99% accuracy while catching zero fraud. Accuracy rewards a model that has learned nothing about the rare class you actually care about.
Trap
A fraud detector reports 99% accuracy on a stream that is 99% legitimate. Ship it — it's almost never wrong.
Judge a 1%-positive dataset by accuracy
Why: Predicting 'negative' for EVERYONE scores 99% accuracy while catching zero fraud. Accuracy rewards a model that has learned nothing about the rare class you actually care about.
import numpy as np
y = np.array([1]*10 + [0]*990) # 1% fraud
pred = np.zeros(1000, dtype=int) # predict legit for all
acc = (pred == y).mean()
TP = ((pred==1)&(y==1)).sum()
recall = TP / (y==1).sum()
print(round(acc,4), int(TP), round(recall,4))| metric | value |
|---|---|
| accuracy | 0.99 (looks great) |
| fraud caught (TP) | 0 |
| recall | 0.0 (catches nothing) |
On imbalanced data report recall, F1, or AUC-PR — metrics that stare at the minority class instead of being drowned out by the majority.
Report recall / F1 / AUC-PR for the positive class
Why: The very same all-negative model that scored 0.99 accuracy scores recall = 0.0 and F1 = 0.0. The right metric instantly exposes the useless model that accuracy praised.
import numpy as np
y = np.array([1]*10 + [0]*990) # 1% fraud
pred = np.zeros(1000, dtype=int) # predict legit for all
acc = (pred == y).mean()
TP = ((pred==1)&(y==1)).sum()
recall = TP / (y==1).sum()
print(round(acc,4), int(TP), round(recall,4))| metric | value |
|---|---|
| accuracy | 0.99 (misleading) |
| recall | 0.0 (the honest signal) |
| F1 | 0.0 |
Concept
With k > 2 classes you compute precision/recall/F1 per class (that class vs the rest), then average. How you average changes what the number means.
Pick macro when you care about performance on rare classes; micro when overall correctness is the goal. The gap between them is itself a diagnostic.
Fill the middle
Fill in the blanks
From Macro and micro F1 diverge on a rare class — one line has had its right-hand side removed. Put it back.
import numpy as np
from sklearn.metrics import f1_score
y = np.array([0,0,0,0,1,1,2]) # class 2 is rare (1 example)
pred = np.array([0,0,0,1,1,1,0])
print(round(f1_score(y,pred,average="macro"),4),
round(f1_score(y,pred,average="micro"),4))
Why: pred is what everything below it consumes, so the wrong expression here fails later and somewhere else. Class 2's F1 is 0, and macro averages it in with equal weight, dragging the score down.
Worked example
Seven examples, three classes; class 2 appears just once and the model gets it wrong. Watch macro punish that while micro barely notices.
import numpy as np
from sklearn.metrics import f1_score
y = np.array([0,0,0,0,1,1,2]) # class 2 is rare (1 example)
pred = np.array([0,0,0,1,1,1,0])
print(round(f1_score(y,pred,average="macro"),4),
round(f1_score(y,pred,average="micro"),4))macro-F1 = 0.5167 vs micro-F1 = 0.7143
Why: Class 2's F1 is 0, and macro averages it in with equal weight, dragging the score down. Micro pools counts, so the four correct class-0 predictions dominate and the rare miss is diluted.
| averaging | value | reading |
|---|---|---|
| macro-F1 | 0.5167 | rare class counts equally — flags the failure |
| micro-F1 | 0.7143 | frequent classes dominate — hides it |
Section
Part 3 of 6 — precision vs recall, ROC, AUC
Intuition
Nothing forces the threshold to be 0.5. Lower it and you flag more emails — you catch more spam (recall up) but raise more false alarms (precision down). Raise it and the reverse.
So precision and recall are not fixed properties of the model — they are a curve the single threshold slides you along. To judge the model itself, we look across all thresholds at once.
Faded example
Fill in the blanks
Sweep the threshold: watch P and R trade, with the scaffolding fading: two lines are gone now — fill both.
import numpy as np
y = np.array([1,1,1,1,1,0,0,0,0,0,0,0])
p = np.array([0.95,0.90,0.80,0.60,0.30,0.70,0.55,0.40,0.20,0.15,0.10,0.05])
for t in [0.25, 0.50, 0.75]:
pred = (p >= t).astype(int)
TP = ((pred==1)&(y==1)).sum(); FP = ((pred==1)&(y==0)).sum()
FN = ((pred==0)&(y==1)).sum()
prec = TP/(TP+FP); rec = TP/(TP+FN)
print(t, int(TP), int(FP), int(FN), round(prec,3), round(rec,3))
Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. Lowering the bar admits more true spam (recall up) but also more ham false alarms (precision down).
Worked example
Recompute precision and recall at three thresholds on the running data. Same scores p, three different cut lines.
import numpy as np
y = np.array([1,1,1,1,1,0,0,0,0,0,0,0])
p = np.array([0.95,0.90,0.80,0.60,0.30,0.70,0.55,0.40,0.20,0.15,0.10,0.05])
for t in [0.25, 0.50, 0.75]:
pred = (p >= t).astype(int)
TP = ((pred==1)&(y==1)).sum(); FP = ((pred==1)&(y==0)).sum()
FN = ((pred==0)&(y==1)).sum()
prec = TP/(TP+FP); rec = TP/(TP+FN)
print(t, int(TP), int(FP), int(FN), round(prec,3), round(rec,3))As t drops 0.75 → 0.50 → 0.25, recall rises 0.6 → 0.8 → 1.0 while precision falls 1.0 → 0.667 → 0.625
Why: Lowering the bar admits more true spam (recall up) but also more ham false alarms (precision down). The two move in opposite directions — the core trade-off.
| t | TP | FP | FN | precision | recall |
|---|---|---|---|---|---|
| 0.75 | 3 | 0 | 2 | 1.000 | 0.600 |
| 0.50 | 4 | 2 | 1 | 0.667 | 0.800 |
| 0.25 | 5 | 3 | 0 | 0.625 | 1.000 |
Pattern
Step through it
Step through Sweep the threshold: watch P and R trade one row at a time. What is driving the change, and what would the row after the last one be?
Concept
The ROC curve plots two rates as the threshold sweeps from high to low. TPR (= recall) is the fraction of positives caught; FPR is the fraction of negatives wrongly flagged.
\[ \text{TPR} = \frac{TP}{TP + FN} = \frac{TP}{P}, \qquad \text{FPR} = \frac{FP}{FP + TN} = \frac{FP}{N} \]
At t = 1 nothing is flagged, so we sit at (0, 0). At t = 0 everything is flagged, so we reach (1, 1). A good model bows toward the top-left — high TPR at low FPR.
Sorting
Sort into buckets
These are the pieces of Lesson 26: Evaluation Metrics, out of order. Put each one back under the part of the lesson it belongs to.
Worked example
With P = 5 positives and N = 7 negatives, compute (FPR, TPR) at a ladder of thresholds. These are the corners the curve steps through.
import numpy as np
y = np.array([1,1,1,1,1,0,0,0,0,0,0,0])
p = np.array([0.95,0.90,0.80,0.60,0.30,0.70,0.55,0.40,0.20,0.15,0.10,0.05])
P = (y==1).sum(); N = (y==0).sum()
for t in [1.0, 0.75, 0.65, 0.50, 0.35, 0.0]:
pred = (p >= t).astype(int)
TP = ((pred==1)&(y==1)).sum(); FP = ((pred==1)&(y==0)).sum()
print(t, round(FP/N,3), round(TP/P,3))The curve runs (0,0) → (0,0.6) → (0.143,0.6) → (0.286,0.8) → (0.429,0.8) → (1,1)
Why: It climbs straight up first (catching the three most-confident spam at zero false alarms), then steps right each time a threshold crosses a ham score. Bowed toward top-left = good.
| threshold t | FPR | TPR |
|---|---|---|
| 1.00 | 0.000 | 0.000 |
| 0.75 | 0.000 | 0.600 |
| 0.65 | 0.143 | 0.600 |
| 0.50 | 0.286 | 0.800 |
| 0.35 | 0.429 | 0.800 |
| 0.00 | 1.000 | 1.000 |
Scale up
Step through it
Step through Trace the ROC curve points and watch the numbers move. Now imagine the input ten times bigger: which column is the one that stops this being practical?
Concept
AUC is the area under that ROC curve. A perfect ranker hugs the top-left and encloses area 1.0; a coin-flip model sits on the diagonal with area 0.5.
But AUC has a second, more useful reading — a plain probability — that lets us compute it by pure counting, no curve needed. That is the next slide, and it is the fact the exam asks for.
Concept
AUC equals the probability that a randomly chosen positive scores higher than a randomly chosen negative. It measures pure ranking — does spam get higher scores than ham? — and is completely threshold-free.
\[ \text{AUC} = \Pr\big(p_{+} > p_{-}\big) = \frac{1}{|\text{pos}|\,|\text{neg}|}\sum_{i \in \text{pos}}\sum_{j \in \text{neg}} \Big[\,\mathbb{1}(p_i > p_j) + \tfrac{1}{2}\mathbb{1}(p_i = p_j)\Big] \]
Count every positive–negative pair: score 1 if the positive outranks the negative, 0.5 for a tie, 0 if the negative wins. Average over all pairs. With 5 positives and 7 negatives there are 5·7 = 35 pairs to judge.
Step zero
Discussion prompt
AUC by hand: count the losing pairs — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: 0.60 loses to 0.70 → 1 losing pair
Answer:
Worked example
Positive scores are {0.95, 0.90, 0.80, 0.60, 0.30}; negative scores are {0.70, 0.55, 0.40, 0.20, 0.15, 0.10, 0.05}. It is easiest to count the pairs the model gets wrong (a positive scoring below a negative).
0.60 loses to 0.70 → 1 losing pair
Why: The positive e4 (0.60) is outranked by the negative e6 (0.70).
0.30 loses to 0.70, 0.55, 0.40 → 3 losing pairs
Why: The weak positive e5 (0.30) is outranked by three negatives. No other positive loses to any negative.
wins = 35 − 4 = 31 → AUC = 31/35 = 0.8857
Why: 31 of 35 positive–negative pairs are correctly ordered, and there are no ties. AUC = 0.8857.
import numpy as np
from sklearn.metrics import roc_auc_score
y = np.array([1,1,1,1,1,0,0,0,0,0,0,0])
p = np.array([0.95,0.90,0.80,0.60,0.30,0.70,0.55,0.40,0.20,0.15,0.10,0.05])
pos, neg = p[y==1], p[y==0]
wins = sum((a > b) + 0.5*(a == b) for a in pos for b in neg)
auc = wins / (len(pos)*len(neg))
print(int(len(pos)*len(neg)), wins, round(auc,4), round(roc_auc_score(y, p),4))| source | pairs | wins | AUC |
|---|---|---|---|
| hand count | 35 | 31 | 0.8857 |
| pairwise code | 35 | 31.0 | 0.8857 |
| sklearn roc_auc_score | — | — | 0.8857 |
Blank canvas
Draw it
Draw what AUC by hand: count the losing pairs just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Anomaly
Predict first
A student writes this, and it looks reasonable:
"AUC = 0.8857, so the model is about 89% accurate."
It is wrong. Say what breaks — and say it before you turn the page.
Correct: AUC never looks at a threshold, so it says nothing about how many LABELS are right.
Read AUC as a ranking probability: P(random spam scores above random ham) = 0.8857.
Why: AUC never looks at a threshold, so it says nothing about how many LABELS are right. Our model's accuracy at t=0.5 is 0.75, not 0.8857 — different numbers measuring different things.
Trap
"AUC = 0.8857, so the model is about 89% accurate."
Treat AUC as a fraction of correct labels
Why: AUC never looks at a threshold, so it says nothing about how many LABELS are right. Our model's accuracy at t=0.5 is 0.75, not 0.8857 — different numbers measuring different things.
Read AUC as a ranking probability: P(random spam scores above random ham) = 0.8857.
AUC = ranking quality across all thresholds; accuracy = correct labels at one threshold
Why: AUC is threshold-free (0.8857 here); accuracy depends entirely on the cut (0.75 at t=0.5, different at every other t). A model can have high AUC yet poor accuracy if you threshold it badly.
Break the constraint
Discussion prompt
The rule this trap just fixed:
AUC is threshold-free (0.8857 here); accuracy depends entirely on the cut (0.75 at t=0.5, different at every other t). A model can have high AUC yet poor accuracy if you threshold it badly.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
AUC never looks at a threshold, so it says nothing about how many LABELS are right. Our model's accuracy at t=0.5 is 0.75, not 0.8857 — different numbers measuring different things.
Fill the middle
Fill in the blanks
From AUC-PR: when negatives swamp positives — one line has had its right-hand side removed. Put it back.
import numpy as np
from sklearn.metrics import average_precision_score, roc_auc_score
rng = np.random.default_rng(0)
y = *np.array([1,1,1] + [0]27)** # 3 positives, 27 negatives (10%)
p = np.concatenate([[0.80,0.60,0.40], rng.uniform(0,0.7,27)])
print(round(roc_auc_score(y,p),4), round(average_precision_score(y,p),4))
Why: y is what everything below it consumes, so the wrong expression here fails later and somewhere else. ROC-AUC looks decent; AUC-PR reveals the model is barely better than the 0.1 base rate at surfacing the rare positives.
Concept
ROC's FPR has the huge negative count N in its denominator, so under heavy imbalance a flood of false positives barely nudges it — ROC-AUC can look rosy on a bad detector. AUC-PR (area under precision–recall, a.k.a. average precision) puts FP against TP, so it stays honest.
import numpy as np
from sklearn.metrics import average_precision_score, roc_auc_score
rng = np.random.default_rng(0)
y = np.array([1,1,1] + [0]*27) # 3 positives, 27 negatives (10%)
p = np.concatenate([[0.80,0.60,0.40], rng.uniform(0,0.7,27)])
print(round(roc_auc_score(y,p),4), round(average_precision_score(y,p),4))Same predictions: ROC-AUC = 0.7654 but AUC-PR = 0.4874
Why: ROC-AUC looks decent; AUC-PR reveals the model is barely better than the 0.1 base rate at surfacing the rare positives. On imbalanced problems, prefer AUC-PR.
| metric | value | reading |
|---|---|---|
| ROC-AUC | 0.7654 | flattered by the 27 negatives |
| AUC-PR (avg precision) | 0.4874 | the honest, imbalance-aware score |
| base rate (positives/total) | 0.100 | AUC-PR floor for a random model |
Intuition
Both curves sweep the same threshold; they differ in what a good score requires. ROC rewards separating positives from negatives; PR additionally demands that your positive predictions are pure.
The tell is the base rate. A random model's ROC-AUC is always 0.5 regardless of imbalance, but its AUC-PR floor is the positive base rate — 0.10 in the example above. So on rare-positive problems, PR has room to expose a weak model that ROC hides.
Concept
Everything so far judged labels or rankings. Log-loss (binary cross-entropy) judges the probabilities themselves — it rewards being confident and right, and savagely punishes being confident and wrong.
\[ \text{LogLoss} = -\frac{1}{n}\sum_{i=1}^{n}\Big[\,y_i \ln p_i + (1 - y_i)\ln(1 - p_i)\,\Big] \]
This is exactly the cross-entropy from Lesson 14. Lower is better; a perfect, perfectly-confident model scores 0. It is the loss you train on, and a good final metric when calibrated probabilities matter.
Estimation
Predict first
Two models, both right on every label on four examples, but one is confident (0.9/0.1) and one is timid (0.6/0.4). Log-loss separates them; accuracy would call them identical.
Commit before you compute: what does Log-loss rewards calibration come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Timid: every term is −ln(0.6) = 0.5108 → log-loss 0.5108
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Same correct labels, but only 0.6 confidence, so each term is −ln(0.6).
Worked example
Two models, both right on every label on four examples, but one is confident (0.9/0.1) and one is timid (0.6/0.4). Log-loss separates them; accuracy would call them identical.
import numpy as np
from sklearn.metrics import log_loss
y = np.array([1, 0, 1, 0])
p_conf = np.array([0.9, 0.1, 0.9, 0.1]) # confident & right
p_timid= np.array([0.6, 0.4, 0.6, 0.4]) # right label, unsure
def ll(y, p):
p = np.clip(p, 1e-15, 1-1e-15)
return -np.mean(y*np.log(p) + (1-y)*np.log(1-p))
print(round(ll(y,p_conf),4), round(log_loss(y,p_conf),4))
print(round(ll(y,p_timid),4), round(log_loss(y,p_timid),4))Confident: every term is −ln(0.9) = 0.1054 → log-loss 0.1054
Why: All four probabilities put 0.9 on the correct class, so each term is −ln(0.9); the mean is −ln(0.9).
Timid: every term is −ln(0.6) = 0.5108 → log-loss 0.5108
Why: Same correct labels, but only 0.6 confidence, so each term is −ln(0.6). Log-loss is ~5× worse — it penalizes hedging even when the label is right.
| model | accuracy | log-loss | from-scratch = sklearn |
|---|---|---|---|
| confident 0.9 | 1.00 | 0.1054 | 0.1054 = 0.1054 ✓ |
| timid 0.6 | 1.00 | 0.5108 | 0.5108 = 0.5108 ✓ |
Anomaly
Predict first
A student writes this, and it looks reasonable:
"Being 0.99 sure costs the same whether I'm right or wrong — it's just one prediction."
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Log-loss on that one wrong-but-certain email is −ln(1−0.99) = −ln(0.01) ≈ 4.6 — one confidently-wrong prediction can outweigh dozens of correct ones.
Calibrate: only claim 0.99 when you would be right 99 times in 100. Log-loss forces this discipline — it is unbounded above as p → 0 on a true positive.
Why: Log-loss on that one wrong-but-certain email is −ln(1−0.99) = −ln(0.01) ≈ 4.6 — one confidently-wrong prediction can outweigh dozens of correct ones.
Trap
"Being 0.99 sure costs the same whether I'm right or wrong — it's just one prediction."
Assume high confidence is free
Why: Log-loss on that one wrong-but-certain email is −ln(1−0.99) = −ln(0.01) ≈ 4.6 — one confidently-wrong prediction can outweigh dozens of correct ones.
import numpy as np
from sklearn.metrics import log_loss
y = np.array([1, 1, 1, 0])
p_ok = np.array([0.7, 0.7, 0.7, 0.3])
p_wrong = np.array([0.7, 0.7, 0.7, 0.99]) # sure the ham is spam
print(round(log_loss(y,p_ok),4), round(log_loss(y,p_wrong),4))| 4th prediction | log-loss |
|---|---|
| 0.3 (unsure, correct-ish) | 0.3567 |
| 0.99 (confident, wrong) | 1.4188 |
Calibrate: only claim 0.99 when you would be right 99 times in 100. Log-loss forces this discipline — it is unbounded above as p → 0 on a true positive.
Confident-wrong nearly 4× the log-loss of unsure
Why: Flipping the last prediction from 0.3 to a confident-but-wrong 0.99 explodes log-loss from 0.3567 to 1.4188. Confidence must be earned, or the metric bites.
import numpy as np
from sklearn.metrics import log_loss
y = np.array([1, 1, 1, 0])
p_ok = np.array([0.7, 0.7, 0.7, 0.3])
p_wrong = np.array([0.7, 0.7, 0.7, 0.99]) # sure the ham is spam
print(round(log_loss(y,p_ok),4), round(log_loss(y,p_wrong),4))| 4th prediction | log-loss |
|---|---|
| 0.3 (calibrated) | 0.3567 |
| 0.99 (over-confident, wrong) | 1.4188 |
Comparison
Comparison matrix
From Trap: a confident wrong answer is cheap: refill the log-loss column from what you know. The rest of the table is as it appeared.
| 4th prediction | log-loss |
|---|---|
| 0.3 (unsure, correct-ish) | 0.3567 |
| 0.99 (confident, wrong) | 1.4188 |
Section
Part 4 of 6 — error, in target units
Concept
Switch domains: 6 houses, each with a true price yt and a model prediction yp, in units of $100k. The residual e = yt − yp is how far off we are on each house.
| house | true yt | pred yp | e = yt − yp |
|---|---|---|---|
| h1 | 2.0 | 2.5 | −0.5 |
| h2 | 3.0 | 2.0 | +1.0 |
| h3 | 5.0 | 5.0 | 0.0 |
| h4 | 4.0 | 5.0 | −1.0 |
| h5 | 6.0 | 5.0 | +1.0 |
| h6 | 8.0 | 9.0 | −1.0 |
Every regression metric is a different way to summarize this error column into one number. The residuals are small and mixed in sign — a clean case with no outliers (yet).
Pattern
Step through it
Step through The running regression example: house prices one row at a time. What is driving the change, and what would the row after the last one be?
Ranking
Put in order
Put the moves of MSE and RMSE, entry by entry into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. (−0.5)²=0.25, (+1)²=1, 0²=0, (−1)²=1, (+1)²=1, (−1)²=1.
Worked example
MSE averages the squared errors; RMSE takes its square root to return to the target's units ($100k). Square each residual first:
Squared errors: 0.25, 1.00, 0.00, 1.00, 1.00, 1.00 → sum 4.25
Why: (−0.5)²=0.25, (+1)²=1, 0²=0, (−1)²=1, (+1)²=1, (−1)²=1. Squaring erases sign and magnifies the larger misses.
MSE = 4.25 / 6 = 0.7083
Why: Average of the six squared errors. Units are ($100k)² — awkward, which is why we take a root next.
RMSE = √0.7083 = 0.8416
Why: Back in $100k. Read it as a typical error size of about $84k per house.
| house | e | e² |
|---|---|---|
| h1 | −0.5 | 0.25 |
| h2 | +1.0 | 1.00 |
| h3 | 0.0 | 0.00 |
| h4..h6 | −1, +1, −1 | 1, 1, 1 |
| Σ / MSE / RMSE | — | 4.25 / 0.7083 / 0.8416 |
Blank canvas
Draw it
Draw what MSE and RMSE, entry by entry just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Ranking
Put in order
Put the moves of MAE and R², entry by entry into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Absolute values, then average. Every miss counts in proportion to its size — no extra weight on the big ones.
Worked example
MAE averages the absolute errors (no squaring). R² compares our error to the error of always guessing the mean ȳ.
MAE = (0.5+1+0+1+1+1)/6 = 4.5/6 = 0.75
Why: Absolute values, then average. Every miss counts in proportion to its size — no extra weight on the big ones.
ȳ = 28/6 = 4.6667; SS_tot = Σ(yt − ȳ)² = 23.333
Why: Deviations of yt=[2,3,5,4,6,8] from 4.6667 squared and summed give 23.333 — the error of the dumb mean-only baseline.
R² = 1 − SS_res/SS_tot = 1 − 4.25/23.333 = 0.8179
Why: The model removes about 82% of the baseline's squared error — it explains 82% of the variance in price.
import numpy as np
yt = np.array([2.0, 3.0, 5.0, 4.0, 6.0, 8.0])
yp = np.array([2.5, 2.0, 5.0, 5.0, 5.0, 9.0])
err = yt - yp
mse = np.mean(err**2); rmse = mse**0.5
mae = np.mean(np.abs(err))
r2 = 1 - np.sum(err**2)/np.sum((yt-yt.mean())**2)
print(round(mse,4), round(rmse,4), round(mae,4), round(r2,4))| metric | value |
|---|---|
| MSE | 0.7083 |
| RMSE | 0.8416 |
| MAE | 0.7500 |
| R² | 0.8179 |
Error analysis
Annotate
Walk the callouts on MAE and R², entry by entry. Each one is a place this is easy to get subtly wrong.
Intuition
Squaring inside MSE gives extra weight to big residuals, so the root-mean-square is pulled up toward the largest errors. The mean-absolute treats every dollar of error the same.
By Jensen's inequality RMSE ≥ MAE, with equality only when every error has the same magnitude. The gap between them is a direct readout of how spread out your errors are — a built-in outlier detector.
Explain it
Discussion prompt
Explain Why RMSE ≥ MAE, always to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Squaring inside MSE gives extra weight to big residuals, so the root-mean-square is pulled up toward the largest errors. The mean-absolute treats every dollar of error the same.
Anomaly
Predict first
A student writes this, and it looks reasonable:
"RMSE and MAE both average error, so report whichever — they tell the same story."
It is wrong. Say what breaks — and say it before you turn the page.
Correct: One big miss barely moves MAE but explodes RMSE, because RMSE squares it.
Choose by how you want large errors treated: RMSE to punish big misses, MAE to stay robust to outliers.
Why: One big miss barely moves MAE but explodes RMSE, because RMSE squares it. They diverge exactly when outliers exist — which is precisely when the choice of metric matters most.
Trap
"RMSE and MAE both average error, so report whichever — they tell the same story."
Assume RMSE ≈ MAE regardless of the data
Why: One big miss barely moves MAE but explodes RMSE, because RMSE squares it. They diverge exactly when outliers exist — which is precisely when the choice of metric matters most.
import numpy as np
err_clean = np.array([0.2,-0.1,0.15,-0.2,0.1,0.05])
err_out = np.array([0.2,-0.1,0.15,-0.2,0.1,3.0]) # one outlier
for e in (err_clean, err_out):
rmse = np.sqrt(np.mean(e**2)); mae = np.mean(np.abs(e))
print(round(rmse,4), round(mae,4), round(rmse/mae,3))| errors | RMSE | MAE | RMSE/MAE |
|---|---|---|---|
| all small | 0.1443 | 0.1333 | 1.083 |
| one 3.0 outlier | 1.2331 | 0.6250 | 1.973 |
Choose by how you want large errors treated: RMSE to punish big misses, MAE to stay robust to outliers.
The ratio RMSE/MAE nearly doubles when one outlier appears
Why: Same five tiny errors; swapping the last for a 3.0 outlier pushes RMSE/MAE from 1.08 to 1.97. The gap IS the outlier signal — reporting only one metric hides it.
import numpy as np
err_clean = np.array([0.2,-0.1,0.15,-0.2,0.1,0.05])
err_out = np.array([0.2,-0.1,0.15,-0.2,0.1,3.0]) # one outlier
for e in (err_clean, err_out):
rmse = np.sqrt(np.mean(e**2)); mae = np.mean(np.abs(e))
print(round(rmse,4), round(mae,4), round(rmse/mae,3))| case | RMSE/MAE | reading |
|---|---|---|
| clean | 1.083 | errors are uniform |
| outlier | 1.973 | one error dominates |
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
p into a predicted label we pick a threshold — flag as spam when p ≥ t. We start with the default t = 0.5.; Every (prediction, truth) pair falls into exactly one of four boxes. Positive/negative names what the model said; true/false names whether it was right.; Precision's enemy is the false alarm (FP); recall's enemy is the miss (FN). Same TP on top — they differ only in which error type sits in the denominator.Faded example
Fill in the blanks
MAPE and its blow-up near zero, with the scaffolding fading: two lines are gone now — fill both.
import numpy as np
yt = np.array([100.0, 50.0, 0.5]) # third true value is tiny
yp = np.array([110.0, 45.0, 1.0])
per = *np.abs((yt-yp)/yt)100**
print([float(round(v,1)) for v in per], round(float(per.mean()),2))
Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. The first two houses are off by a sane 10%.
Worked example
MAPE reports error as a percentage of the true value — intuitive, but it divides by yt, so a true value near zero detonates it.
import numpy as np
yt = np.array([100.0, 50.0, 0.5]) # third true value is tiny
yp = np.array([110.0, 45.0, 1.0])
per = np.abs((yt-yp)/yt)*100
print([float(round(v,1)) for v in per], round(float(per.mean()),2))Per-row percent error: 10%, 10%, 100% → MAPE = 40%
Why: The first two houses are off by a sane 10%. The third has true value 0.5; a 0.5 absolute miss is 100% of it, dragging the mean to 40%.
| yt | yp | |e|/yt · 100 |
|---|---|---|
| 100.0 | 110.0 | 10.0% |
| 50.0 | 45.0 | 10.0% |
| 0.5 | 1.0 | 100.0% ← blow-up |
| MAPE (mean) | — | 40.0% |
Discrimination
Sort into buckets
Sort these by |e|/yt · 100, from memory, without looking back at MAPE and its blow-up near zero. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Constraint
Discussion prompt
Run Metric-selection cheat sheet with this step confiscated:
Calibrated probabilities: log-loss (punishes confident-wrong)
Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.
Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.
Answer:
P(pos ranked above neg)y = 0)Pattern
P(pos ranked above neg)y = 0)Edge cases
Discussion prompt
Metric-selection cheat sheet works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
P(pos ranked above neg)y = 0)Section
Part 5 of 6 — text and images
Intuition
There is no single 'correct label' for a translation, a summary, or a generated image — many outputs are good. So generation metrics compare against reference examples or against a distribution, not one gold answer.
Four workhorses: perplexity (how surprised is the language model?), BLEU (n-gram precision vs references), ROUGE (n-gram recall vs references), and FID (distance between real and generated image feature distributions).
Intuition
Picture the model at each token facing a fork in the road. If it were perfectly torn between exactly k equally-likely next words, its perplexity would be exactly k — the effective number of choices it is juggling.
So perplexity 3.3 means the model is about as uncertain as flipping between roughly three words at each step. A perfect model that always assigns probability 1 to the true token has perplexity 1 (no branching); a uniform model over a vocabulary of size V has perplexity V.
That is why lower is better and why perplexity is comparable across models on the same text — it is a unit-free surprise, not a raw loss.
Fill the middle
Fill in the blanks
From Perplexity: how surprised is the model? — one line has had its right-hand side removed. Put it back.
import numpy as np
# model's probability on each true next token, 5-token sentence
q = np.array([0.5, 0.25, 0.5, 0.1, 0.4])
nll = -np.mean(np.log(q)) # average negative log-likelihood
ppl = np.exp(nll)
print(round(nll,4), round(ppl,4))
Why: ppl is what everything below it consumes, so the wrong expression here fails later and somewhere else. The model is, on average, about as uncertain as a fair choice among ~3.3 tokens.
Worked example
Perplexity is exp of the average negative log-likelihood the model assigned to the true next tokens. Lower is better; perplexity k means the model is as unsure as if choosing uniformly among k options.
\[ \text{PPL} = \exp\!\Big(-\frac{1}{T}\sum_{t=1}^{T}\ln q_t\Big), \qquad q_t = \text{model's prob of the true token} \]
import numpy as np
# model's probability on each true next token, 5-token sentence
q = np.array([0.5, 0.25, 0.5, 0.1, 0.4])
nll = -np.mean(np.log(q)) # average negative log-likelihood
ppl = np.exp(nll)
print(round(nll,4), round(ppl,4))avg NLL = 1.1983 → perplexity = exp(1.1983) = 3.3145
Why: The model is, on average, about as uncertain as a fair choice among ~3.3 tokens. Ties straight back to cross-entropy from Lesson 14 — perplexity is just its exponential.
| quantity | value |
|---|---|
| token probs q | 0.5, 0.25, 0.5, 0.1, 0.4 |
| avg NLL = −mean(ln q) | 1.1983 |
| perplexity = exp(NLL) | 3.3145 |
Translation
\( \text{PPL} = \exp\!\Big(-\frac{1}{T}\sum_{t=1}^{T}\ln q_t\Big), \qquad q_t = \text{model's prob of the true token} \)
Draw it
Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.
Pattern
Predict first
The table runs: the | 3 | 2 | 2 · cat | 1 | 1 | 1 · mat | 1 | 1 | 1
In BLEU idea: clipped n-gram precision, given the rows so far: what is the next one — the row where word is total / p1?
Correct: total / p1 | 5 | — | 4 → 0.8
| word | cand count | ref count | clipped |
|---|---|---|---|
| the | 3 | 2 | 2 |
| cat | 1 | 1 | 1 |
| mat | 1 | 1 | 1 |
| total / p1 | 5 | — | 4 → 0.8 |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. Without clipping, the three 'the's would each count and precision would be inflated.
Worked example
BLEU rewards candidate n-grams that appear in a reference — but clips each count at how often it appears in the reference, so spamming a word can't inflate the score. Here is the unigram (1-gram) precision, the heart of BLEU.
from collections import Counter
ref = "the cat sat on the mat".split()
cand = "the the the cat mat".split()
ref_c, cand_c = Counter(ref), Counter(cand)
clipped = sum(min(cand_c[w], ref_c[w]) for w in cand_c)
p1 = clipped / len(cand)
print(clipped, len(cand), round(p1,4))'the' appears 3× in candidate but only 2× in reference → clipped to 2
Why: Without clipping, the three 'the's would each count and precision would be inflated. min(3, 2) caps it at 2. 'cat' and 'mat' each match once.
clipped matches = 2 + 1 + 1 = 4, over 5 candidate words → p1 = 0.8
Why: Unigram precision is 4/5 = 0.8. Full BLEU multiplies precisions across n-gram sizes and adds a brevity penalty, but this clipped-precision core is the exam-relevant idea.
| word | cand count | ref count | clipped |
|---|---|---|---|
| the | 3 | 2 | 2 |
| cat | 1 | 1 | 1 |
| mat | 1 | 1 | 1 |
| total / p1 | 5 | — | 4 → 0.8 |
Discrimination
Sort into buckets
Sort these by cand count, from memory, without looking back at BLEU idea: clipped n-gram precision. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Missing information
Discussion prompt
Where BLEU asks 'how much of my output is in the reference?' (precision), ROUGE asks 'how much of the reference did I cover?' (recall) — the natural fit for summarization, where you want to capture the key content.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
The candidate 'the cat mat' covers 'the' (once), 'cat', and 'mat' — 3 of the reference's 6 tokens. Recall = 3/6 = 0.5. Note the same overlap divided by candidate length would be precision.
Worked example
Where BLEU asks 'how much of my output is in the reference?' (precision), ROUGE asks 'how much of the reference did I cover?' (recall) — the natural fit for summarization, where you want to capture the key content.
from collections import Counter
ref = "the cat sat on the mat".split()
cand = "the cat mat".split()
ref_c, cand_c = Counter(ref), Counter(cand)
overlap = sum(min(cand_c[w], ref_c[w]) for w in ref_c)
recall = overlap / len(ref)
print(overlap, len(ref), round(recall,4))overlap = 3 covered words out of 6 reference words → ROUGE-1 = 0.5
Why: The candidate 'the cat mat' covers 'the' (once), 'cat', and 'mat' — 3 of the reference's 6 tokens. Recall = 3/6 = 0.5. Note the same overlap divided by candidate length would be precision.
| measure | denominator | value |
|---|---|---|
| overlap (clipped) | — | 3 |
| ROUGE-1 recall | len(ref) = 6 | 0.5 |
| (precision would use len(cand) = 3) | 3 | 1.0 |
Estimation
Predict first
For image generation there is no reference per image. FID embeds real and generated images into a feature space and measures the distance between the two Gaussians (means μ, covariances Σ). Lower = generated set looks more like real. A 1-D toy shows the formula:
Commit before you compute: what does FID idea: distance between distributions come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: FID = 0.25 + (1 + 1.44 − 2·1.2) = 0.25 + 0.04 = 0.29
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The mean gap contributes (0.5)²=0.25; the variance mismatch contributes 1+1.44−2√(1.44)=0.04.
Worked example
For image generation there is no reference per image. FID embeds real and generated images into a feature space and measures the distance between the two Gaussians (means μ, covariances Σ). Lower = generated set looks more like real. A 1-D toy shows the formula:
\[ \text{FID} = \lVert \mu_r - \mu_f \rVert^2 + \operatorname{Tr}\!\big(\Sigma_r + \Sigma_f - 2(\Sigma_r \Sigma_f)^{1/2}\big) \]
import numpy as np
mu_r, var_r = 0.0, 1.0 # real features ~ N(0, 1)
mu_f, var_f = 0.5, 1.44 # fake features ~ N(0.5, 1.2^2)
fid = (mu_r-mu_f)**2 + var_r + var_f - 2*np.sqrt(var_r*var_f)
print(round(fid,4))FID = 0.25 + (1 + 1.44 − 2·1.2) = 0.25 + 0.04 = 0.29
Why: The mean gap contributes (0.5)²=0.25; the variance mismatch contributes 1+1.44−2√(1.44)=0.04. In 1-D the trace and matrix sqrt reduce to plain variance and √. Match real distribution → FID → 0.
| term | value |
|---|---|
| mean gap ‖μr−μf‖² | 0.25 |
| variance term σr²+σf²−2σrσf | 0.04 |
| FID | 0.29 |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
FID = 0.25 + (1 + 1.44 − 2·1.2) = 0.25 + 0.04 = 0.29
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
For image generation there is no reference per image. FID embeds real and generated images into a feature space and measures the distance between the two Gaussians (means μ, covariances Σ). Lower = generated set looks more like real. A 1-D toy shows the formula:
Section
Part 6 of 6 — then build it
Elimination
Eliminate the wrong options
A model has precision = 0.8 and recall = 0.6. Its F1 score is:
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: F1 = 2PR/(P+R) = 2(0.8)(0.6)/(0.8+0.6) = 0.96/1.4 ≈ 0.686. The harmonic mean sits below the arithmetic mean, pulled toward the smaller value (recall).
Check
Compute it on paper before clicking.
Check your understanding
A model has precision = 0.8 and recall = 0.6. Its F1 score is:
Answer: A
Why: F1 = 2PR/(P+R) = 2(0.8)(0.6)/(0.8+0.6) = 0.96/1.4 ≈ 0.686. The harmonic mean sits below the arithmetic mean, pulled toward the smaller value (recall).
Prediction
Predict first
A dataset is 99% negative. A model predicts 'negative' for every example. Which is true?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Accuracy is 0.99 but recall for the positive class is 0
Why: Predicting all-negative is right on every true negative (99% of the data) so accuracy = 0.99, but it catches zero positives, so recall = TP/(TP+FN) = 0 and F1 = 0. Accuracy hides a useless model.
Check
Recall the all-negative baseline.
Check your understanding
A dataset is 99% negative. A model predicts 'negative' for every example. Which is true?
Answer: A
Why: Predicting all-negative is right on every true negative (99% of the data) so accuracy = 0.99, but it catches zero positives, so recall = TP/(TP+FN) = 0 and F1 = 0. Accuracy hides a useless model.
Commit first
Predict first
A classifier has AUC-ROC = 0.90. The most precise interpretation is:
Commit to an answer, then rate it — certain, fairly sure, or guessing — and write the rating down before you turn the page.
Correct: A random positive is scored higher than a random negative 90% of the time
Why: AUC-ROC equals P(random positive scored above random negative) — a threshold-free measure of ranking quality. Equivalently it is the area under the ROC curve.
The rating matters as much as the answer: confident-and-wrong is the combination that survives revision, because nothing about it feels like it needs revisiting.
Check
Ranking, not labels.
Check your understanding
A classifier has AUC-ROC = 0.90. The most precise interpretation is:
Answer: A
Why: AUC-ROC equals P(random positive scored above random negative) — a threshold-free measure of ranking quality. Equivalently it is the area under the ROC curve.
Prediction
Predict first
You compute RMSE = 1.23 and MAE = 0.63 on the same predictions. What does the large gap most likely indicate?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: A few large errors (outliers) are inflating RMSE relative to MAE
Why: RMSE squares errors, so a few big misses pull it well above MAE. A wide RMSE/MAE ratio (here ~1.95) is a signature of outliers or high error variance.
Check
Think about what squaring does.
Check your understanding
You compute RMSE = 1.23 and MAE = 0.63 on the same predictions. What does the large gap most likely indicate?
Answer: A
Why: RMSE squares errors, so a few big misses pull it well above MAE. A wide RMSE/MAE ratio (here ~1.95) is a signature of outliers or high error variance.
Elimination
Eliminate the wrong options
Model A outputs 0.9 on the correct class every time; Model B outputs 0.6 on the correct class every time. Both are 100% accurate. Which has the lower (better) log-loss?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Log-loss per example is −ln(p) on the correct class: −ln(0.9)=0.105 for A vs −ln(0.6)=0.511 for B. A is ~5× lower, so A is better. Log-loss rewards calibrated confidence, which accuracy cannot see.
Check
Two models agree on every label.
Check your understanding
Model A outputs 0.9 on the correct class every time; Model B outputs 0.6 on the correct class every time. Both are 100% accurate. Which has the lower (better) log-loss?
Answer: A
Why: Log-loss per example is −ln(p) on the correct class: −ln(0.9)=0.105 for A vs −ln(0.6)=0.511 for B. A is ~5× lower, so A is better. Log-loss rewards calibrated confidence, which accuracy cannot see.
Section
The project
Concept
Implement the core classification metrics on the running spam data using only NumPy, then prove each one matches sklearn. You derived every piece — now assemble the library.
| # | requirement | tool |
|---|---|---|
| 1 | confusion counts → precision, recall | boolean masks |
| 2 | F1 = 2PR/(P+R) | harmonic mean |
| 3 | AUC via pairwise ranking vs sklearn | roc_auc_score |
Build rules: type every line yourself, run after each milestone, build the four counts with boolean masks, and compare AUC to roc_auc_score with np.isclose. If a count looks wrong, print the masks — don't guess.
Analogy
Discussion prompt
Explain Project: a metrics library from scratch by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Implement the core classification metrics on the running spam data using only NumPy, then prove each one matches sklearn. You derived every piece — now assemble the library.
Worked example
Your turn: threshold p at 0.5, then build TP, FP, FN, TN with masks. Predict the counts from the by-hand result before you print.
Hint: pred = (p >= 0.5).astype(int), then TP = ((pred==1)&(y==1)).sum(), and so on. The four should sum to 12.
import numpy as np
y = np.array([1,1,1,1,1,0,0,0,0,0,0,0])
p = np.array([0.95,0.90,0.80,0.60,0.30,0.70,0.55,0.40,0.20,0.15,0.10,0.05])
pred = (p >= 0.5).astype(int)
TP = ((pred==1)&(y==1)).sum(); FP = ((pred==1)&(y==0)).sum()
FN = ((pred==0)&(y==1)).sum(); TN = ((pred==0)&(y==0)).sum()
print(int(TP), int(FP), int(FN), int(TN))
print(round(TP/(TP+FP),4), round(TP/(TP+FN),4))| quantity | value |
|---|---|
| TP, FP, FN, TN | 4, 2, 1, 5 |
| precision | 0.6667 |
| recall | 0.8000 |
Pattern
Step through it
Step through Milestone 1 — the four counts one row at a time. What is driving the change, and what would the row after the last one be?
Worked example
Your turn: combine precision and recall into F1. Predict whether F1 lands above or below their arithmetic mean of 0.733.
Hint: f1 = 2*prec*rec/(prec+rec) — the harmonic mean, which sits below the arithmetic mean.
prec, rec = 4/6, 4/5
f1 = 2*prec*rec/(prec+rec)
print(round(prec,4), round(rec,4), round(f1,4))| quantity | value |
|---|---|
| precision | 0.6667 |
| recall | 0.8000 |
| F1 (harmonic mean) | 0.7273 |
| (arithmetic mean, for contrast) | 0.7333 |
Worked example
Your turn: compute AUC as the fraction of correctly-ordered positive/negative pairs and verify against sklearn with np.isclose. Predict the value from the 31/35 hand count.
Hint: pos, neg = p[y==1], p[y==0], then average (a>b)+0.5*(a==b) over all pos/neg pairs; compare to roc_auc_score.
import numpy as np
from sklearn.metrics import roc_auc_score
y = np.array([1,1,1,1,1,0,0,0,0,0,0,0])
p = np.array([0.95,0.90,0.80,0.60,0.30,0.70,0.55,0.40,0.20,0.15,0.10,0.05])
pos, neg = p[y==1], p[y==0]
auc = np.mean([(a>b)+0.5*(a==b) for a in pos for b in neg])
print(round(auc,4), np.isclose(auc, roc_auc_score(y,p)))| source | AUC | match |
|---|---|---|
| pairwise ranking (31/35) | 0.8857 | — |
| sklearn roc_auc_score | 0.8857 | True |
Trade off
Comparison matrix
From Milestone 3 — AUC vs sklearn: every row here is a choice with a cost. Fill the AUC column, then say which row you would actually pick and what you give up for it.
| source | AUC | match |
|---|---|---|
| pairwise ranking (31/35) | 0.8857 | — |
| sklearn roc_auc_score | 0.8857 | True |
Concept
import numpy as np
from sklearn.metrics import roc_auc_score, average_precision_score
def confusion(y, pred):
TP=((pred==1)&(y==1)).sum(); FP=((pred==1)&(y==0)).sum()
FN=((pred==0)&(y==1)).sum(); TN=((pred==0)&(y==0)).sum()
return TP, FP, FN, TN
def prf(y, pred):
TP, FP, FN, TN = confusion(y, pred)
P = TP/(TP+FP); R = TP/(TP+FN)
return P, R, 2*P*R/(P+R)
def auc_pairwise(y, p):
pos, neg = p[y==1], p[y==0]
wins = sum((a>b)+0.5*(a==b) for a in pos for b in neg)
return wins / (len(pos)*len(neg))
y = np.array([1,1,1,1,1,0,0,0,0,0,0,0])
p = np.array([0.95,0.90,0.80,0.60,0.30,0.70,0.55,0.40,0.20,0.15,0.10,0.05])
pred = (p >= 0.5).astype(int)
P, R, F1 = prf(y, pred)
auc = auc_pairwise(y, p)
print('P,R,F1 =', tuple(round(v,4) for v in (P,R,F1)))
print('AUC =', round(auc,4), '== sklearn', np.isclose(auc, roc_auc_score(y,p)))| printed line | value |
|---|---|
| P,R,F1 = | (0.6667, 0.8, 0.7273) |
| AUC = | 0.8857 == sklearn True |
If your P, R, F1 read 0.6667, 0.8, 0.7273 and AUC matches sklearn, you can score a classifier from the counts up — and you know exactly what each number sees.
Comparison
Comparison matrix
From The full program: refill the value column from what you know. The rest of the table is as it appeared.
| printed line | value |
|---|---|
| P,R,F1 = | (0.6667, 0.8, 0.7273) |
| AUC = | 0.8857 == sklearn True |
Concept
Slides closed, out loud: explain (1) why accuracy fails on 99%-negative data, (2) what AUC-ROC means as a probability and how the 31/35 pairs produced 0.8857, and (3) when RMSE diverges from MAE.
Stretch: add average_precision_score to the library and show it drops below ROC-AUC on the imbalanced set; then reproduce roc_curve by sweeping thresholds and confirm the (FPR, TPR) ladder you traced by hand. Metrics drive model selection (Week 22) and GAN/diffusion FID (Week 48).
Counterexample
Discussion prompt
Slides closed, out loud: explain (1) why accuracy fails on 99%-negative data, (2) what AUC-ROC means as a probability and how the 31/35 pairs produced 0.8857, and (3) when RMSE diverges from MAE.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — The confusion matrix · Precision, recall, F1 · The threshold trade-off · Regression metrics · Generation metrics · Check yourself. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
P(random pos ranked above random neg) — 31/35 = 0.8857 heresklearn| metric | the one thing to remember |
|---|---|
| F1 | 2PR/(P+R) — harmonic mean, pulled to the smaller |
| accuracy | lies on imbalanced data — use F1 / recall / AUC-PR |
| AUC-ROC | P(random pos > random neg); threshold-free ranking |
| log-loss | calibrated probs; confident-wrong explodes it |
| RMSE vs MAE | RMSE ≥ MAE; the gap is your outlier detector |
| generation | perplexity=exp(NLL), BLEU=precision, ROUGE=recall, FID=distribution distance |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.