Lesson 26: Evaluation Metrics

USAAIO Lesson 26, from Week 9, fully worked on one running 12-email spam classifier and one six-house regression set. It builds the confusion matrix count by count, derives precision, recall, and F1 and computes them by hand, and explains why accuracy lies on imbalanced data. It sweeps the precision-recall trade-off across thresholds, derives ROC and AUC as a ranking probability - 31 of 35 pairs, by hand - and confirms it against sklearn, then covers AUC-PR and log-loss with a confident-but-wrong blow-up. It then computes MSE, RMSE, MAE, R2, and MAPE entry by entry, splits RMSE from MAE with an outlier, and gives each generation metric - perplexity, BLEU, ROUGE, and FID - a runnable toy example. It ends with a from-scratch metrics library verified against sklearn. Every number was produced by real execution. The lesson runs to 61 slides.

Subject: Machine Learning · 119 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Evaluation Metrics

Title

USAAIO · Lesson 26 · Week 9

The number you optimize decides which model you ship. We build every metric on one running spam classifier and one house-price set — confusion counts, precision/recall/F1, the ROC curve and AUC as a ranking probability, log-loss, then MSE/RMSE/MAE/R² and the generation metrics — each derived by hand and confirmed against sklearn.

2. By the end of this lesson you can

Objectives

  1. Build a confusion matrix count by count and derive precision, recall, F1 from it — including why F1 is a harmonic mean
  2. Show exactly why accuracy misleads on imbalanced data, and pick the metric that exposes the useless model
  3. Sweep the threshold to trace the precision–recall trade-off, then plot the ROC curve and compute AUC as P(random positive ranked above random negative)
  4. Compute AUC-PR and log-loss, and show log-loss punishing a confident-but-wrong prediction
  5. Compute MSE, RMSE, MAE, R², MAPE entry by entry, and know when RMSE and MAE diverge
  6. Define the generation metrics — perplexity, BLEU, ROUGE, FID — and implement a runnable toy of each
  7. Implement a from-scratch metrics library and prove it matches sklearn

3. What survived from Projections & the Hat Matrix?

Warm-up

Discussion prompt

Before we open Lesson 26: Evaluation Metrics: without looking back, what was the main idea of Projections & the Hat Matrix, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

orthogonal projection matrices (P²=P, Pᵀ=P), the hat matrix H=X(XᵀX)⁻¹Xᵀ that projects y onto the column space, the residual projector I−H, leverage scores, and Cook's distance. Build the hat matrix and leverage scores from scratch.

4. The confusion matrix

Section

Part 1 of 6 — the four counts

5. The running example: a spam classifier

Concept

One dataset carries the whole classification half of this lesson: 12 emails. Each has a true label y (1 = spam, 0 = ham) and the model's score p = its estimated P(spam). Five are truly spam, seven are ham.

emailtrue yscore p
e11 (spam)0.95
e21 (spam)0.90
e31 (spam)0.80
e41 (spam)0.60
e51 (spam)0.30
e60 (ham)0.70
e70 (ham)0.55
e8..e120 (ham)0.40, 0.20, 0.15, 0.10, 0.05

A score is not yet a decision. To turn p into a predicted label we pick a threshold — flag as spam when p ≥ t. We start with the default t = 0.5.

6. Fill in: true y for The running example: a spam classifier

Comparison

Comparison matrix

From The running example: a spam classifier: refill the true y column from what you know. The rest of the table is as it appeared.

emailtrue yscore p
e11 (spam)0.95
e21 (spam)0.90
e31 (spam)0.80
e41 (spam)0.60
e51 (spam)0.30
e60 (ham)0.70
e70 (ham)0.55
e8..e120 (ham)0.40, 0.20, 0.15, 0.10, 0.05

7. Guess the shape of the answer: Threshold the scores at t = 0.5

Estimation

Predict first

Apply pred = (p ≥ 0.5). Two of the five spam emails clear 0.5 easily (0.95, 0.90, 0.80, 0.60) but e5 at 0.30 does not; two ham emails (e6 0.70, e7 0.55) wrongly clear it. This is our prediction vector.

Commit before you compute: what does Threshold the scores at t = 0.5 come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: pred = [1 1 1 1 0 1 1 0 0 0 0 0]

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Four of five spam correctly flagged (e1–e4), e5 missed; two ham wrongly flagged (e6, e7).

8. Threshold the scores at t = 0.5

Worked example

Apply pred = (p ≥ 0.5). Two of the five spam emails clear 0.5 easily (0.95, 0.90, 0.80, 0.60) but e5 at 0.30 does not; two ham emails (e6 0.70, e7 0.55) wrongly clear it. This is our prediction vector.

import numpy as np
y = np.array([1,1,1,1,1,0,0,0,0,0,0,0])
p = np.array([0.95,0.90,0.80,0.60,0.30,
              0.70,0.55,0.40,0.20,0.15,0.10,0.05])
pred = (p >= 0.5).astype(int)
print(pred)

pred = [1 1 1 1 0 1 1 0 0 0 0 0]

Why: Four of five spam correctly flagged (e1–e4), e5 missed; two ham wrongly flagged (e6, e7). Every metric below is a different summary of exactly these agreements and disagreements.

emailypp ≥ 0.5 → pred
e110.951 ✓
e410.601 ✓
e510.300 ✗ missed spam
e600.701 ✗ flagged ham
e1200.050 ✓

9. What each one costs: Threshold the scores at t = 0.5

Trade off

Comparison matrix

From Threshold the scores at t = 0.5: every row here is a choice with a cost. Fill the p ≥ 0.5 → pred column, then say which row you would actually pick and what you give up for it.

emailypp ≥ 0.5 → pred
e110.951 ✓
e410.601 ✓
e510.300 ✗ missed spam
e600.701 ✗ flagged ham
e1200.050 ✓

10. Four outcomes, four names

Concept

Every (prediction, truth) pair falls into exactly one of four boxes. Positive/negative names what the model said; true/false names whether it was right.

actual spam (+)actual ham (−)
predict spam (+)TP — caught spamFP — false alarm
predict ham (−)FN — missed spamTN — correct pass

FP is a false alarm (real mail trashed); FN is a miss (spam let through). These two error types have different real-world costs — that asymmetry drives every choice in this lesson.

11. Break it if you can: Four outcomes, four names

Counterexample

Discussion prompt

Every (prediction, truth) pair falls into exactly one of four boxes. Positive/negative names what the model said; true/false names whether it was right.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

12. What has to be given first: Count the four boxes with boolean masks

Missing information

Discussion prompt

A boolean mask like (pred==1)&(y==1) is True exactly on the true positives; .sum() counts the Trues. Build all four counts from pred and y:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

e1–e4 are true positives; e6, e7 are false positives; e5 is the lone false negative; the other five ham are true negatives. The four counts sum to 12 — every email lands in exactly one box.

13. Count the four boxes with boolean masks

Worked example

A boolean mask like (pred==1)&(y==1) is True exactly on the true positives; .sum() counts the Trues. Build all four counts from pred and y:

import numpy as np
y    = np.array([1,1,1,1,1,0,0,0,0,0,0,0])
pred = np.array([1,1,1,1,0,1,1,0,0,0,0,0])
TP = ((pred==1)&(y==1)).sum(); FP = ((pred==1)&(y==0)).sum()
FN = ((pred==0)&(y==1)).sum(); TN = ((pred==0)&(y==0)).sum()
print(int(TP), int(FP), int(FN), int(TN))

TP=4, FP=2, FN=1, TN=5

Why: e1–e4 are true positives; e6, e7 are false positives; e5 is the lone false negative; the other five ham are true negatives. The four counts sum to 12 — every email lands in exactly one box.

countwhich emailsvalue
TPe1, e2, e3, e44
FPe6, e72
FNe51
TNe8, e9, e10, e11, e125
totalall12

14. Work backwards from the answer: Count the four boxes with boolean masks

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

TP=4, FP=2, FN=1, TN=5

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

A boolean mask like (pred==1)&(y==1) is True exactly on the true positives; .sum() counts the Trues. Build all four counts from pred and y:

15. Reading the names: true/false, positive/negative

Intuition

The four names decode cleanly once you split them. The second word is what the model predicted — 'positive' means it flagged spam, 'negative' means it passed the email through.

The first word says whether that prediction was correct: 'true' means the model was right, 'false' means it was wrong. So a false positive is a wrong 'spam' flag, and a false negative is a wrong 'ham' pass.

Whenever a metric confuses you, expand its counts this way. Precision, recall, and every rate below are just ratios of these four boxes — nothing more.

16. By analogy: Reading the names: true/false, positive/negative

Analogy

Discussion prompt

Explain Reading the names: true/false, positive/negative by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

The four names decode cleanly once you split them. The second word is what the model predicted — 'positive' means it flagged spam, 'negative' means it passed the email through.

17. Precision, recall, F1

Section

Part 2 of 6 — from counts to scores

18. Precision and recall: two different questions

Concept

Precision looks only at what you flagged: of the emails I called spam, how many really were? Recall looks only at the real spam: of all actual spam, how many did I catch?

\[ \text{precision} = \frac{TP}{TP + FP}, \qquad \text{recall} = \frac{TP}{TP + FN} \]

Precision's enemy is the false alarm (FP); recall's enemy is the miss (FN). Same TP on top — they differ only in which error type sits in the denominator.

19. Teach it back: Precision and recall: two different questions

Explain it

Discussion prompt

Explain Precision and recall: two different questions to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Precision's enemy is the false alarm (FP); recall's enemy is the miss (FN). Same TP on top — they differ only in which error type sits in the denominator.

20. The knob you can always cheat

Intuition

Either metric alone is trivial to game. Flag every email as spam and you catch all real spam — recall = 1.0 — but almost everything you flag is wrong, so precision collapses.

Flag only the single email you are most certain about and precision is likely 1.0 — but you miss all the rest, so recall collapses. Any honest score has to look at both at once.

A spam filter leans toward precision (never trash real mail); a cancer screen leans toward recall (never miss a case). The cost of FP vs FN decides the balance — the metric just makes that choice explicit.

21. Predict the next row: Precision and recall on our counts

Pattern

Predict first

The table runs: precision | TP/(TP+FP) | 4/6 | 0.6667 · recall | TP/(TP+FN) | 4/5 | 0.8000

In Precision and recall on our counts, given the rows so far: what is the next one — the row where metric is accuracy?

Correct: accuracy | (TP+TN)/n | 9/12 | 0.7500

metricformulaarithmeticvalue
precisionTP/(TP+FP)4/60.6667
recallTP/(TP+FN)4/50.8000
accuracy(TP+TN)/n9/120.7500

Why: The relationship between the columns, not the individual numbers, is what generates the next row. Of the 6 emails we flagged as spam, 4 truly were — the 2 false alarms (e6, e7) drag precision down to two-thirds.

22. Precision and recall on our counts

Worked example

Plug TP=4, FP=2, FN=1 straight into the formulas — no code needed yet, just the two divisions.

precision = TP/(TP+FP) = 4/(4+2) = 4/6 = 0.6667

Why: Of the 6 emails we flagged as spam, 4 truly were — the 2 false alarms (e6, e7) drag precision down to two-thirds.

recall = TP/(TP+FN) = 4/(4+1) = 4/5 = 0.8

Why: Of the 5 real spam emails, we caught 4 — only e5 (score 0.30) slipped through, so recall is 0.8.

metricformulaarithmeticvalue
precisionTP/(TP+FP)4/60.6667
recallTP/(TP+FN)4/50.8000
accuracy(TP+TN)/n9/120.7500

23. Watch it run: Precision and recall on our counts

Pattern

Step through it

Step through Precision and recall on our counts one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: metric is precision
  2. Step 2: metric is recall
  3. Step 3: metric is accuracy

24. F1: one number that needs both

Concept

F1 is the harmonic mean of precision P and recall R — it collapses the pair into a single score that is high only when both are high.

\[ F_1 = \frac{2}{\tfrac{1}{P} + \tfrac{1}{R}} = \frac{2PR}{P + R} \]

The harmonic mean is dominated by the smaller term: if either P or R is near zero, F1 is near zero. That is exactly why the 'flag everything' cheat (recall 1, precision tiny) scores a terrible F1.

25. What has to happen first: F1 by hand, then verified in code

Ranking

Put in order

Put the moves of F1 by hand, then verified in code into the order they have to happen.

  1. Numerator: 2PR = 2·(4/6)·(4/5) = 2·(16/30) = 32/30
  2. Denominator: P + R = 4/6 + 4/5 = 20/30 + 24/30 = 44/30
  3. F1 = (32/30)/(44/30) = 32/44 = 0.7273

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Multiply the two rates, double it.

26. F1 by hand, then verified in code

Worked example

Numerator: 2PR = 2·(4/6)·(4/5) = 2·(16/30) = 32/30

Why: Multiply the two rates, double it. Keep fractions to stay exact.

Denominator: P + R = 4/6 + 4/5 = 20/30 + 24/30 = 44/30

Why: Common denominator 30. The 30s cancel in the ratio.

F1 = (32/30)/(44/30) = 32/44 = 0.7273

Why: F1 sits between recall (0.8) and precision (0.667), pulled toward the smaller — the signature of a harmonic mean.

import numpy as np
y    = np.array([1,1,1,1,1,0,0,0,0,0,0,0])
pred = np.array([1,1,1,1,0,1,1,0,0,0,0,0])
TP = ((pred==1)&(y==1)).sum(); FP = ((pred==1)&(y==0)).sum()
FN = ((pred==0)&(y==1)).sum(); TN = ((pred==0)&(y==0)).sum()
prec = TP/(TP+FP); rec = TP/(TP+FN)
f1 = 2*prec*rec/(prec+rec)
print(int(TP), int(FP), int(FN), int(TN))
print(round(prec,4), round(rec,4), round(f1,4))
printedvalue
TP FP FN TN4 2 1 5
precision0.6667
recall0.8
F10.7273

27. Inspect it line by line: F1 by hand, then verified in code

Error analysis

Annotate

Walk the callouts on F1 by hand, then verified in code. Each one is a place this is easy to get subtly wrong.

  • Multiply the two rates, double it. Keep fractions to stay exact.
  • Common denominator 30. The 30s cancel in the ratio.
  • F1 sits between recall (0.8) and precision (0.667), pulled toward the smaller — the signature of a harmonic mean.

28. Rebuild the recipe: From confusion matrix to metrics

Ranking

Put in order

These are the steps of From confusion matrix to metrics, scrambled. Put them back in order before the next slide shows you.

  1. Threshold: turn scores into labels with pred = (p ≥ t)
  2. Count: four boolean masks → TP, FP, FN, TN (they must sum to n)
  3. Precision = TP/(TP+FP) — punishes false alarms
  4. Recall = TP/(TP+FN) — punishes misses
  5. F1 = 2PR/(P+R) — the harmonic mean, high only when both are

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

29. From confusion matrix to metrics

Pattern

  1. Threshold: turn scores into labels with pred = (p ≥ t)
  2. Count: four boolean masks → TP, FP, FN, TN (they must sum to n)
  3. Precision = TP/(TP+FP) — punishes false alarms
  4. Recall = TP/(TP+FN) — punishes misses
  5. F1 = 2PR/(P+R) — the harmonic mean, high only when both are

30. Something is wrong here: accuracy on imbalanced data

Anomaly

Predict first

A student writes this, and it looks reasonable:

A fraud detector reports 99% accuracy on a stream that is 99% legitimate. Ship it — it's almost never wrong.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Predicting 'negative' for EVERYONE scores 99% accuracy while catching zero fraud.

On imbalanced data report recall, F1, or AUC-PR — metrics that stare at the minority class instead of being drowned out by the majority.

Why: Predicting 'negative' for EVERYONE scores 99% accuracy while catching zero fraud. Accuracy rewards a model that has learned nothing about the rare class you actually care about.

31. Trap: accuracy on imbalanced data

Trap

The trap

A fraud detector reports 99% accuracy on a stream that is 99% legitimate. Ship it — it's almost never wrong.

Judge a 1%-positive dataset by accuracy

Why: Predicting 'negative' for EVERYONE scores 99% accuracy while catching zero fraud. Accuracy rewards a model that has learned nothing about the rare class you actually care about.

import numpy as np
y = np.array([1]*10 + [0]*990)      # 1% fraud
pred = np.zeros(1000, dtype=int)    # predict legit for all
acc = (pred == y).mean()
TP = ((pred==1)&(y==1)).sum()
recall = TP / (y==1).sum()
print(round(acc,4), int(TP), round(recall,4))
metricvalue
accuracy0.99 (looks great)
fraud caught (TP)0
recall0.0 (catches nothing)

The fix

On imbalanced data report recall, F1, or AUC-PR — metrics that stare at the minority class instead of being drowned out by the majority.

Report recall / F1 / AUC-PR for the positive class

Why: The very same all-negative model that scored 0.99 accuracy scores recall = 0.0 and F1 = 0.0. The right metric instantly exposes the useless model that accuracy praised.

import numpy as np
y = np.array([1]*10 + [0]*990)      # 1% fraud
pred = np.zeros(1000, dtype=int)    # predict legit for all
acc = (pred == y).mean()
TP = ((pred==1)&(y==1)).sum()
recall = TP / (y==1).sum()
print(round(acc,4), int(TP), round(recall,4))
metricvalue
accuracy0.99 (misleading)
recall0.0 (the honest signal)
F10.0

32. More than two classes: macro vs micro

Concept

With k > 2 classes you compute precision/recall/F1 per class (that class vs the rest), then average. How you average changes what the number means.

Pick macro when you care about performance on rare classes; micro when overall correctness is the goal. The gap between them is itself a diagnostic.

33. Restore the missing line: Macro and micro F1 diverge on a rare class

Fill the middle

Fill in the blanks

From Macro and micro F1 diverge on a rare class — one line has had its right-hand side removed. Put it back.

import numpy as np
from sklearn.metrics import f1_score
y = np.array([0,0,0,0,1,1,2]) # class 2 is rare (1 example)
pred = np.array([0,0,0,1,1,1,0])
print(round(f1_score(y,pred,average="macro"),4),
round(f1_score(y,pred,average="micro"),4))

Why: pred is what everything below it consumes, so the wrong expression here fails later and somewhere else. Class 2's F1 is 0, and macro averages it in with equal weight, dragging the score down.

34. Macro and micro F1 diverge on a rare class

Worked example

Seven examples, three classes; class 2 appears just once and the model gets it wrong. Watch macro punish that while micro barely notices.

import numpy as np
from sklearn.metrics import f1_score
y    = np.array([0,0,0,0,1,1,2])   # class 2 is rare (1 example)
pred = np.array([0,0,0,1,1,1,0])
print(round(f1_score(y,pred,average="macro"),4),
      round(f1_score(y,pred,average="micro"),4))

macro-F1 = 0.5167 vs micro-F1 = 0.7143

Why: Class 2's F1 is 0, and macro averages it in with equal weight, dragging the score down. Micro pools counts, so the four correct class-0 predictions dominate and the rare miss is diluted.

averagingvaluereading
macro-F10.5167rare class counts equally — flags the failure
micro-F10.7143frequent classes dominate — hides it

35. The threshold trade-off

Section

Part 3 of 6 — precision vs recall, ROC, AUC

36. 0.5 is a choice, not a law

Intuition

Nothing forces the threshold to be 0.5. Lower it and you flag more emails — you catch more spam (recall up) but raise more false alarms (precision down). Raise it and the reverse.

So precision and recall are not fixed properties of the model — they are a curve the single threshold slides you along. To judge the model itself, we look across all thresholds at once.

37. Finish it with less help: Sweep the threshold: watch P and R trade

Faded example

Fill in the blanks

Sweep the threshold: watch P and R trade, with the scaffolding fading: two lines are gone now — fill both.

import numpy as np
y = np.array([1,1,1,1,1,0,0,0,0,0,0,0])
p = np.array([0.95,0.90,0.80,0.60,0.30,0.70,0.55,0.40,0.20,0.15,0.10,0.05])
for t in [0.25, 0.50, 0.75]:
pred = (p >= t).astype(int)
TP = ((pred==1)&(y==1)).sum(); FP = ((pred==1)&(y==0)).sum()
FN = ((pred==0)&(y==1)).sum()
prec = TP/(TP+FP); rec = TP/(TP+FN)
print(t, int(TP), int(FP), int(FN), round(prec,3), round(rec,3))

Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. Lowering the bar admits more true spam (recall up) but also more ham false alarms (precision down).

38. Sweep the threshold: watch P and R trade

Worked example

Recompute precision and recall at three thresholds on the running data. Same scores p, three different cut lines.

import numpy as np
y = np.array([1,1,1,1,1,0,0,0,0,0,0,0])
p = np.array([0.95,0.90,0.80,0.60,0.30,0.70,0.55,0.40,0.20,0.15,0.10,0.05])
for t in [0.25, 0.50, 0.75]:
    pred = (p >= t).astype(int)
    TP = ((pred==1)&(y==1)).sum(); FP = ((pred==1)&(y==0)).sum()
    FN = ((pred==0)&(y==1)).sum()
    prec = TP/(TP+FP); rec = TP/(TP+FN)
    print(t, int(TP), int(FP), int(FN), round(prec,3), round(rec,3))

As t drops 0.75 → 0.50 → 0.25, recall rises 0.6 → 0.8 → 1.0 while precision falls 1.0 → 0.667 → 0.625

Why: Lowering the bar admits more true spam (recall up) but also more ham false alarms (precision down). The two move in opposite directions — the core trade-off.

tTPFPFNprecisionrecall
0.753021.0000.600
0.504210.6670.800
0.255300.6251.000

39. Watch it run: Sweep the threshold: watch P and R trade

Pattern

Step through it

Step through Sweep the threshold: watch P and R trade one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: t is 0.75
  2. Step 2: t is 0.50
  3. Step 3: t is 0.25

40. The ROC curve: TPR vs FPR

Concept

The ROC curve plots two rates as the threshold sweeps from high to low. TPR (= recall) is the fraction of positives caught; FPR is the fraction of negatives wrongly flagged.

\[ \text{TPR} = \frac{TP}{TP + FN} = \frac{TP}{P}, \qquad \text{FPR} = \frac{FP}{FP + TN} = \frac{FP}{N} \]

At t = 1 nothing is flagged, so we sit at (0, 0). At t = 0 everything is flagged, so we reach (1, 1). A good model bows toward the top-left — high TPR at low FPR.

41. Where does each piece belong: Lesson 26: Evaluation Metrics

Sorting

Sort into buckets

These are the pieces of Lesson 26: Evaluation Metrics, out of order. Put each one back under the part of the lesson it belongs to.

The confusion matrix
The running example: a spam classifier; Threshold the scores at t = 0.5; Four outcomes, four names
Precision, recall, F1
Precision and recall: two different questions; The knob you can always cheat; Precision and recall on our counts
The threshold trade-off
0.5 is a choice, not a law; Sweep the threshold: watch P and R trade; The ROC curve: TPR vs FPR
s1
The confusion matrix is where Lesson 26: Evaluation Metrics puts The running example: a spam classifier, Threshold the scores at t = 0.5, Four outcomes, four names. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
Precision, recall, F1 is where Lesson 26: Evaluation Metrics puts Precision and recall: two different questions, The knob you can always cheat, Precision and recall on our counts. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
The threshold trade-off is where Lesson 26: Evaluation Metrics puts 0.5 is a choice, not a law, Sweep the threshold: watch P and R trade, The ROC curve: TPR vs FPR. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

42. Trace the ROC curve points

Worked example

With P = 5 positives and N = 7 negatives, compute (FPR, TPR) at a ladder of thresholds. These are the corners the curve steps through.

import numpy as np
y = np.array([1,1,1,1,1,0,0,0,0,0,0,0])
p = np.array([0.95,0.90,0.80,0.60,0.30,0.70,0.55,0.40,0.20,0.15,0.10,0.05])
P = (y==1).sum(); N = (y==0).sum()
for t in [1.0, 0.75, 0.65, 0.50, 0.35, 0.0]:
    pred = (p >= t).astype(int)
    TP = ((pred==1)&(y==1)).sum(); FP = ((pred==1)&(y==0)).sum()
    print(t, round(FP/N,3), round(TP/P,3))

The curve runs (0,0) → (0,0.6) → (0.143,0.6) → (0.286,0.8) → (0.429,0.8) → (1,1)

Why: It climbs straight up first (catching the three most-confident spam at zero false alarms), then steps right each time a threshold crosses a ham score. Bowed toward top-left = good.

threshold tFPRTPR
1.000.0000.000
0.750.0000.600
0.650.1430.600
0.500.2860.800
0.350.4290.800
0.001.0001.000

43. What happens as it grows: Trace the ROC curve points

Scale up

Step through it

Step through Trace the ROC curve points and watch the numbers move. Now imagine the input ten times bigger: which column is the one that stops this being practical?

  1. Step 1: threshold t is 1.00
  2. Step 2: threshold t is 0.75
  3. Step 3: threshold t is 0.65
  4. Step 4: threshold t is 0.50
  5. Step 5: threshold t is 0.35
  6. Step 6: threshold t is 0.00

44. AUC = area under the ROC curve

Concept

AUC is the area under that ROC curve. A perfect ranker hugs the top-left and encloses area 1.0; a coin-flip model sits on the diagonal with area 0.5.

But AUC has a second, more useful reading — a plain probability — that lets us compute it by pure counting, no curve needed. That is the next slide, and it is the fact the exam asks for.

45. AUC as a ranking probability

Concept

AUC equals the probability that a randomly chosen positive scores higher than a randomly chosen negative. It measures pure ranking — does spam get higher scores than ham? — and is completely threshold-free.

\[ \text{AUC} = \Pr\big(p_{+} > p_{-}\big) = \frac{1}{|\text{pos}|\,|\text{neg}|}\sum_{i \in \text{pos}}\sum_{j \in \text{neg}} \Big[\,\mathbb{1}(p_i > p_j) + \tfrac{1}{2}\mathbb{1}(p_i = p_j)\Big] \]

Count every positive–negative pair: score 1 if the positive outranks the negative, 0.5 for a tie, 0 if the negative wins. Average over all pairs. With 5 positives and 7 negatives there are 5·7 = 35 pairs to judge.

46. Plan first: AUC by hand: count the losing pairs

Step zero

Discussion prompt

AUC by hand: count the losing pairs — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: 0.60 loses to 0.70 → 1 losing pair

Answer:

  1. 0.60 loses to 0.70 → 1 losing pair
  2. 0.30 loses to 0.70, 0.55, 0.40 → 3 losing pairs
  3. wins = 35 − 4 = 31 → AUC = 31/35 = 0.8857

47. AUC by hand: count the losing pairs

Worked example

Positive scores are {0.95, 0.90, 0.80, 0.60, 0.30}; negative scores are {0.70, 0.55, 0.40, 0.20, 0.15, 0.10, 0.05}. It is easiest to count the pairs the model gets wrong (a positive scoring below a negative).

0.60 loses to 0.70 → 1 losing pair

Why: The positive e4 (0.60) is outranked by the negative e6 (0.70).

0.30 loses to 0.70, 0.55, 0.40 → 3 losing pairs

Why: The weak positive e5 (0.30) is outranked by three negatives. No other positive loses to any negative.

wins = 35 − 4 = 31 → AUC = 31/35 = 0.8857

Why: 31 of 35 positive–negative pairs are correctly ordered, and there are no ties. AUC = 0.8857.

import numpy as np
from sklearn.metrics import roc_auc_score
y = np.array([1,1,1,1,1,0,0,0,0,0,0,0])
p = np.array([0.95,0.90,0.80,0.60,0.30,0.70,0.55,0.40,0.20,0.15,0.10,0.05])
pos, neg = p[y==1], p[y==0]
wins = sum((a > b) + 0.5*(a == b) for a in pos for b in neg)
auc = wins / (len(pos)*len(neg))
print(int(len(pos)*len(neg)), wins, round(auc,4), round(roc_auc_score(y, p),4))
sourcepairswinsAUC
hand count35310.8857
pairwise code3531.00.8857
sklearn roc_auc_score——0.8857

48. Draw the shape of it: AUC by hand: count the losing pairs

Blank canvas

Draw it

Draw what AUC by hand: count the losing pairs just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.

49. Something is wrong here: reading AUC as accuracy

Anomaly

Predict first

A student writes this, and it looks reasonable:

"AUC = 0.8857, so the model is about 89% accurate."

It is wrong. Say what breaks — and say it before you turn the page.

Correct: AUC never looks at a threshold, so it says nothing about how many LABELS are right.

Read AUC as a ranking probability: P(random spam scores above random ham) = 0.8857.

Why: AUC never looks at a threshold, so it says nothing about how many LABELS are right. Our model's accuracy at t=0.5 is 0.75, not 0.8857 — different numbers measuring different things.

50. Trap: reading AUC as accuracy

Trap

The trap

"AUC = 0.8857, so the model is about 89% accurate."

Treat AUC as a fraction of correct labels

Why: AUC never looks at a threshold, so it says nothing about how many LABELS are right. Our model's accuracy at t=0.5 is 0.75, not 0.8857 — different numbers measuring different things.

The fix

Read AUC as a ranking probability: P(random spam scores above random ham) = 0.8857.

AUC = ranking quality across all thresholds; accuracy = correct labels at one threshold

Why: AUC is threshold-free (0.8857 here); accuracy depends entirely on the cut (0.75 at t=0.5, different at every other t). A model can have high AUC yet poor accuracy if you threshold it badly.

51. Break it on purpose: reading AUC as accuracy

Break the constraint

Discussion prompt

The rule this trap just fixed:

AUC is threshold-free (0.8857 here); accuracy depends entirely on the cut (0.75 at t=0.5, different at every other t). A model can have high AUC yet poor accuracy if you threshold it badly.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

AUC never looks at a threshold, so it says nothing about how many LABELS are right. Our model's accuracy at t=0.5 is 0.75, not 0.8857 — different numbers measuring different things.

52. Restore the missing line: AUC-PR: when negatives swamp positives

Fill the middle

Fill in the blanks

From AUC-PR: when negatives swamp positives — one line has had its right-hand side removed. Put it back.

import numpy as np
from sklearn.metrics import average_precision_score, roc_auc_score
rng = np.random.default_rng(0)
y = *np.array([1,1,1] + [0]27)** # 3 positives, 27 negatives (10%)
p = np.concatenate([[0.80,0.60,0.40], rng.uniform(0,0.7,27)])
print(round(roc_auc_score(y,p),4), round(average_precision_score(y,p),4))

Why: y is what everything below it consumes, so the wrong expression here fails later and somewhere else. ROC-AUC looks decent; AUC-PR reveals the model is barely better than the 0.1 base rate at surfacing the rare positives.

53. AUC-PR: when negatives swamp positives

Concept

ROC's FPR has the huge negative count N in its denominator, so under heavy imbalance a flood of false positives barely nudges it — ROC-AUC can look rosy on a bad detector. AUC-PR (area under precision–recall, a.k.a. average precision) puts FP against TP, so it stays honest.

import numpy as np
from sklearn.metrics import average_precision_score, roc_auc_score
rng = np.random.default_rng(0)
y = np.array([1,1,1] + [0]*27)          # 3 positives, 27 negatives (10%)
p = np.concatenate([[0.80,0.60,0.40], rng.uniform(0,0.7,27)])
print(round(roc_auc_score(y,p),4), round(average_precision_score(y,p),4))

Same predictions: ROC-AUC = 0.7654 but AUC-PR = 0.4874

Why: ROC-AUC looks decent; AUC-PR reveals the model is barely better than the 0.1 base rate at surfacing the rare positives. On imbalanced problems, prefer AUC-PR.

metricvaluereading
ROC-AUC0.7654flattered by the 27 negatives
AUC-PR (avg precision)0.4874the honest, imbalance-aware score
base rate (positives/total)0.100AUC-PR floor for a random model

54. ROC or PR — which curve to trust

Intuition

Both curves sweep the same threshold; they differ in what a good score requires. ROC rewards separating positives from negatives; PR additionally demands that your positive predictions are pure.

The tell is the base rate. A random model's ROC-AUC is always 0.5 regardless of imbalance, but its AUC-PR floor is the positive base rate — 0.10 in the example above. So on rare-positive problems, PR has room to expose a weak model that ROC hides.

55. Log-loss: scoring the probabilities

Concept

Everything so far judged labels or rankings. Log-loss (binary cross-entropy) judges the probabilities themselves — it rewards being confident and right, and savagely punishes being confident and wrong.

\[ \text{LogLoss} = -\frac{1}{n}\sum_{i=1}^{n}\Big[\,y_i \ln p_i + (1 - y_i)\ln(1 - p_i)\,\Big] \]

This is exactly the cross-entropy from Lesson 14. Lower is better; a perfect, perfectly-confident model scores 0. It is the loss you train on, and a good final metric when calibrated probabilities matter.

56. Guess the shape of the answer: Log-loss rewards calibration

Estimation

Predict first

Two models, both right on every label on four examples, but one is confident (0.9/0.1) and one is timid (0.6/0.4). Log-loss separates them; accuracy would call them identical.

Commit before you compute: what does Log-loss rewards calibration come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Timid: every term is −ln(0.6) = 0.5108 → log-loss 0.5108

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Same correct labels, but only 0.6 confidence, so each term is −ln(0.6).

57. Log-loss rewards calibration

Worked example

Two models, both right on every label on four examples, but one is confident (0.9/0.1) and one is timid (0.6/0.4). Log-loss separates them; accuracy would call them identical.

import numpy as np
from sklearn.metrics import log_loss
y      = np.array([1, 0, 1, 0])
p_conf = np.array([0.9, 0.1, 0.9, 0.1])   # confident & right
p_timid= np.array([0.6, 0.4, 0.6, 0.4])   # right label, unsure
def ll(y, p):
    p = np.clip(p, 1e-15, 1-1e-15)
    return -np.mean(y*np.log(p) + (1-y)*np.log(1-p))
print(round(ll(y,p_conf),4), round(log_loss(y,p_conf),4))
print(round(ll(y,p_timid),4), round(log_loss(y,p_timid),4))

Confident: every term is −ln(0.9) = 0.1054 → log-loss 0.1054

Why: All four probabilities put 0.9 on the correct class, so each term is −ln(0.9); the mean is −ln(0.9).

Timid: every term is −ln(0.6) = 0.5108 → log-loss 0.5108

Why: Same correct labels, but only 0.6 confidence, so each term is −ln(0.6). Log-loss is ~5× worse — it penalizes hedging even when the label is right.

modelaccuracylog-lossfrom-scratch = sklearn
confident 0.91.000.10540.1054 = 0.1054 ✓
timid 0.61.000.51080.5108 = 0.5108 ✓

58. Something is wrong here: a confident wrong answer is cheap

Anomaly

Predict first

A student writes this, and it looks reasonable:

"Being 0.99 sure costs the same whether I'm right or wrong — it's just one prediction."

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Log-loss on that one wrong-but-certain email is −ln(1−0.99) = −ln(0.01) ≈ 4.6 — one confidently-wrong prediction can outweigh dozens of correct ones.

Calibrate: only claim 0.99 when you would be right 99 times in 100. Log-loss forces this discipline — it is unbounded above as p → 0 on a true positive.

Why: Log-loss on that one wrong-but-certain email is −ln(1−0.99) = −ln(0.01) ≈ 4.6 — one confidently-wrong prediction can outweigh dozens of correct ones.

59. Trap: a confident wrong answer is cheap

Trap

The trap

"Being 0.99 sure costs the same whether I'm right or wrong — it's just one prediction."

Assume high confidence is free

Why: Log-loss on that one wrong-but-certain email is −ln(1−0.99) = −ln(0.01) ≈ 4.6 — one confidently-wrong prediction can outweigh dozens of correct ones.

import numpy as np
from sklearn.metrics import log_loss
y       = np.array([1, 1, 1, 0])
p_ok    = np.array([0.7, 0.7, 0.7, 0.3])
p_wrong = np.array([0.7, 0.7, 0.7, 0.99])  # sure the ham is spam
print(round(log_loss(y,p_ok),4), round(log_loss(y,p_wrong),4))
4th predictionlog-loss
0.3 (unsure, correct-ish)0.3567
0.99 (confident, wrong)1.4188

The fix

Calibrate: only claim 0.99 when you would be right 99 times in 100. Log-loss forces this discipline — it is unbounded above as p → 0 on a true positive.

Confident-wrong nearly 4× the log-loss of unsure

Why: Flipping the last prediction from 0.3 to a confident-but-wrong 0.99 explodes log-loss from 0.3567 to 1.4188. Confidence must be earned, or the metric bites.

import numpy as np
from sklearn.metrics import log_loss
y       = np.array([1, 1, 1, 0])
p_ok    = np.array([0.7, 0.7, 0.7, 0.3])
p_wrong = np.array([0.7, 0.7, 0.7, 0.99])  # sure the ham is spam
print(round(log_loss(y,p_ok),4), round(log_loss(y,p_wrong),4))
4th predictionlog-loss
0.3 (calibrated)0.3567
0.99 (over-confident, wrong)1.4188

60. Fill in: log-loss for Trap: a confident wrong answer is cheap

Comparison

Comparison matrix

From Trap: a confident wrong answer is cheap: refill the log-loss column from what you know. The rest of the table is as it appeared.

4th predictionlog-loss
0.3 (unsure, correct-ish)0.3567
0.99 (confident, wrong)1.4188

61. Regression metrics

Section

Part 4 of 6 — error, in target units

62. The running regression example: house prices

Concept

Switch domains: 6 houses, each with a true price yt and a model prediction yp, in units of $100k. The residual e = yt − yp is how far off we are on each house.

housetrue ytpred ype = yt − yp
h12.02.5−0.5
h23.02.0+1.0
h35.05.00.0
h44.05.0−1.0
h56.05.0+1.0
h68.09.0−1.0

Every regression metric is a different way to summarize this error column into one number. The residuals are small and mixed in sign — a clean case with no outliers (yet).

63. Watch it run: The running regression example: house prices

Pattern

Step through it

Step through The running regression example: house prices one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: house is h1
  2. Step 2: house is h2
  3. Step 3: house is h3
  4. Step 4: house is h4
  5. Step 5: house is h5
  6. Step 6: house is h6

64. What has to happen first: MSE and RMSE, entry by entry

Ranking

Put in order

Put the moves of MSE and RMSE, entry by entry into the order they have to happen.

  1. Squared errors: 0.25, 1.00, 0.00, 1.00, 1.00, 1.00 → sum 4.25
  2. MSE = 4.25 / 6 = 0.7083
  3. RMSE = √0.7083 = 0.8416

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. (−0.5)²=0.25, (+1)²=1, 0²=0, (−1)²=1, (+1)²=1, (−1)²=1.

65. MSE and RMSE, entry by entry

Worked example

MSE averages the squared errors; RMSE takes its square root to return to the target's units ($100k). Square each residual first:

Squared errors: 0.25, 1.00, 0.00, 1.00, 1.00, 1.00 → sum 4.25

Why: (−0.5)²=0.25, (+1)²=1, 0²=0, (−1)²=1, (+1)²=1, (−1)²=1. Squaring erases sign and magnifies the larger misses.

MSE = 4.25 / 6 = 0.7083

Why: Average of the six squared errors. Units are ($100k)² — awkward, which is why we take a root next.

RMSE = √0.7083 = 0.8416

Why: Back in $100k. Read it as a typical error size of about $84k per house.

houseee²
h1−0.50.25
h2+1.01.00
h30.00.00
h4..h6−1, +1, −11, 1, 1
Σ / MSE / RMSE—4.25 / 0.7083 / 0.8416

66. Draw the shape of it: MSE and RMSE, entry by entry

Blank canvas

Draw it

Draw what MSE and RMSE, entry by entry just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.

67. What has to happen first: MAE and R², entry by entry

Ranking

Put in order

Put the moves of MAE and R², entry by entry into the order they have to happen.

  1. MAE = (0.5+1+0+1+1+1)/6 = 4.5/6 = 0.75
  2. ȳ = 28/6 = 4.6667; SS_tot = Σ(yt − ȳ)² = 23.333
  3. R² = 1 − SS_res/SS_tot = 1 − 4.25/23.333 = 0.8179

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Absolute values, then average. Every miss counts in proportion to its size — no extra weight on the big ones.

68. MAE and R², entry by entry

Worked example

MAE averages the absolute errors (no squaring). R² compares our error to the error of always guessing the mean ȳ.

MAE = (0.5+1+0+1+1+1)/6 = 4.5/6 = 0.75

Why: Absolute values, then average. Every miss counts in proportion to its size — no extra weight on the big ones.

ȳ = 28/6 = 4.6667; SS_tot = Σ(yt − ȳ)² = 23.333

Why: Deviations of yt=[2,3,5,4,6,8] from 4.6667 squared and summed give 23.333 — the error of the dumb mean-only baseline.

R² = 1 − SS_res/SS_tot = 1 − 4.25/23.333 = 0.8179

Why: The model removes about 82% of the baseline's squared error — it explains 82% of the variance in price.

import numpy as np
yt = np.array([2.0, 3.0, 5.0, 4.0, 6.0, 8.0])
yp = np.array([2.5, 2.0, 5.0, 5.0, 5.0, 9.0])
err = yt - yp
mse  = np.mean(err**2); rmse = mse**0.5
mae  = np.mean(np.abs(err))
r2   = 1 - np.sum(err**2)/np.sum((yt-yt.mean())**2)
print(round(mse,4), round(rmse,4), round(mae,4), round(r2,4))
metricvalue
MSE0.7083
RMSE0.8416
MAE0.7500
R²0.8179

69. Inspect it line by line: MAE and R², entry by entry

Error analysis

Annotate

Walk the callouts on MAE and R², entry by entry. Each one is a place this is easy to get subtly wrong.

  • Absolute values, then average. Every miss counts in proportion to its size — no extra weight on the big ones.
  • Deviations of yt=[2,3,5,4,6,8] from 4.6667 squared and summed give 23.333 — the error of the dumb mean-only baseline.
  • The model removes about 82% of the baseline's squared error — it explains 82% of the variance in price.

70. Why RMSE ≥ MAE, always

Intuition

Squaring inside MSE gives extra weight to big residuals, so the root-mean-square is pulled up toward the largest errors. The mean-absolute treats every dollar of error the same.

By Jensen's inequality RMSE ≥ MAE, with equality only when every error has the same magnitude. The gap between them is a direct readout of how spread out your errors are — a built-in outlier detector.

71. Teach it back: Why RMSE ≥ MAE, always

Explain it

Discussion prompt

Explain Why RMSE ≥ MAE, always to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Squaring inside MSE gives extra weight to big residuals, so the root-mean-square is pulled up toward the largest errors. The mean-absolute treats every dollar of error the same.

72. Something is wrong here: RMSE and MAE are interchangeable

Anomaly

Predict first

A student writes this, and it looks reasonable:

"RMSE and MAE both average error, so report whichever — they tell the same story."

It is wrong. Say what breaks — and say it before you turn the page.

Correct: One big miss barely moves MAE but explodes RMSE, because RMSE squares it.

Choose by how you want large errors treated: RMSE to punish big misses, MAE to stay robust to outliers.

Why: One big miss barely moves MAE but explodes RMSE, because RMSE squares it. They diverge exactly when outliers exist — which is precisely when the choice of metric matters most.

73. Trap: RMSE and MAE are interchangeable

Trap

The trap

"RMSE and MAE both average error, so report whichever — they tell the same story."

Assume RMSE ≈ MAE regardless of the data

Why: One big miss barely moves MAE but explodes RMSE, because RMSE squares it. They diverge exactly when outliers exist — which is precisely when the choice of metric matters most.

import numpy as np
err_clean = np.array([0.2,-0.1,0.15,-0.2,0.1,0.05])
err_out   = np.array([0.2,-0.1,0.15,-0.2,0.1,3.0])   # one outlier
for e in (err_clean, err_out):
    rmse = np.sqrt(np.mean(e**2)); mae = np.mean(np.abs(e))
    print(round(rmse,4), round(mae,4), round(rmse/mae,3))
errorsRMSEMAERMSE/MAE
all small0.14430.13331.083
one 3.0 outlier1.23310.62501.973

The fix

Choose by how you want large errors treated: RMSE to punish big misses, MAE to stay robust to outliers.

The ratio RMSE/MAE nearly doubles when one outlier appears

Why: Same five tiny errors; swapping the last for a 3.0 outlier pushes RMSE/MAE from 1.08 to 1.97. The gap IS the outlier signal — reporting only one metric hides it.

import numpy as np
err_clean = np.array([0.2,-0.1,0.15,-0.2,0.1,0.05])
err_out   = np.array([0.2,-0.1,0.15,-0.2,0.1,3.0])   # one outlier
for e in (err_clean, err_out):
    rmse = np.sqrt(np.mean(e**2)); mae = np.mean(np.abs(e))
    print(round(rmse,4), round(mae,4), round(rmse/mae,3))
caseRMSE/MAEreading
clean1.083errors are uniform
outlier1.973one error dominates

74. Which of these survive contact with Lesson 26: Evaluation Metrics?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
A score is not yet a decision. To turn p into a predicted label we pick a threshold — flag as spam when p ≥ t. We start with the default t = 0.5.; Every (prediction, truth) pair falls into exactly one of four boxes. Positive/negative names what the model said; true/false names whether it was right.; Precision's enemy is the false alarm (FP); recall's enemy is the miss (FN). Same TP on top — they differ only in which error type sits in the denominator.
Breaks
A fraud detector reports 99% accuracy on a stream that is 99% legitimate. Ship it — it's almost never wrong.; "AUC = 0.8857, so the model is about 89% accurate."
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 26: Evaluation Metrics puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

75. Finish it with less help: MAPE and its blow-up near zero

Faded example

Fill in the blanks

MAPE and its blow-up near zero, with the scaffolding fading: two lines are gone now — fill both.

import numpy as np
yt = np.array([100.0, 50.0, 0.5]) # third true value is tiny
yp = np.array([110.0, 45.0, 1.0])
per = *np.abs((yt-yp)/yt)100**
print([float(round(v,1)) for v in per], round(float(per.mean()),2))

Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. The first two houses are off by a sane 10%.

76. MAPE and its blow-up near zero

Worked example

MAPE reports error as a percentage of the true value — intuitive, but it divides by yt, so a true value near zero detonates it.

import numpy as np
yt = np.array([100.0, 50.0, 0.5])   # third true value is tiny
yp = np.array([110.0, 45.0, 1.0])
per = np.abs((yt-yp)/yt)*100
print([float(round(v,1)) for v in per], round(float(per.mean()),2))

Per-row percent error: 10%, 10%, 100% → MAPE = 40%

Why: The first two houses are off by a sane 10%. The third has true value 0.5; a 0.5 absolute miss is 100% of it, dragging the mean to 40%.

ytyp|e|/yt · 100
100.0110.010.0%
50.045.010.0%
0.51.0100.0% ← blow-up
MAPE (mean)—40.0%

77. Which is which, by |e|/yt · 100

Discrimination

Sort into buckets

Sort these by |e|/yt · 100, from memory, without looking back at MAPE and its blow-up near zero. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

10.0%
100.0; 50.0
100.0% ← blow-up
0.5
40.0%
MAPE (mean)
g1
|e|/yt · 100 is "10.0%" for 100.0, 50.0 — that is what the table on "MAPE and its blow-up near zero" records, and it is the single property separating this group from the rest.
g2
|e|/yt · 100 is "100.0% ← blow-up" for 0.5 — that is what the table on "MAPE and its blow-up near zero" records, and it is the single property separating this group from the rest.
g3
|e|/yt · 100 is "40.0%" for MAPE (mean) — that is what the table on "MAPE and its blow-up near zero" records, and it is the single property separating this group from the rest.

78. Without one step: Metric-selection cheat sheet

Constraint

Discussion prompt

Run Metric-selection cheat sheet with this step confiscated:

Calibrated probabilities: log-loss (punishes confident-wrong)

Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.

Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.

Answer:

  1. Balanced classes, labels only: accuracy is fine
  2. Imbalanced classes: F1, recall, or AUC-PR — never accuracy
  3. Ranking quality, threshold-free: AUC-ROC = P(pos ranked above neg)
  4. Calibrated probabilities: log-loss (punishes confident-wrong)
  5. Regression, punish big misses: RMSE; robust to outliers: MAE; variance explained: R²; percent error: MAPE (never near y = 0)
  6. Generation: perplexity (LMs), BLEU (translation), ROUGE (summaries), FID (images)

79. Metric-selection cheat sheet

Pattern

  1. Balanced classes, labels only: accuracy is fine
  2. Imbalanced classes: F1, recall, or AUC-PR — never accuracy
  3. Ranking quality, threshold-free: AUC-ROC = P(pos ranked above neg)
  4. Calibrated probabilities: log-loss (punishes confident-wrong)
  5. Regression, punish big misses: RMSE; robust to outliers: MAE; variance explained: R²; percent error: MAPE (never near y = 0)
  6. Generation: perplexity (LMs), BLEU (translation), ROUGE (summaries), FID (images)

80. Where does it stop working: Metric-selection cheat sheet

Edge cases

Discussion prompt

Metric-selection cheat sheet works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Balanced classes, labels only: accuracy is fine
  2. Imbalanced classes: F1, recall, or AUC-PR — never accuracy
  3. Ranking quality, threshold-free: AUC-ROC = P(pos ranked above neg)
  4. Calibrated probabilities: log-loss (punishes confident-wrong)
  5. Regression, punish big misses: RMSE; robust to outliers: MAE; variance explained: R²; percent error: MAPE (never near y = 0)
  6. Generation: perplexity (LMs), BLEU (translation), ROUGE (summaries), FID (images)

81. Generation metrics

Section

Part 5 of 6 — text and images

82. Why generation needs different metrics

Intuition

There is no single 'correct label' for a translation, a summary, or a generated image — many outputs are good. So generation metrics compare against reference examples or against a distribution, not one gold answer.

Four workhorses: perplexity (how surprised is the language model?), BLEU (n-gram precision vs references), ROUGE (n-gram recall vs references), and FID (distance between real and generated image feature distributions).

83. Perplexity is an effective branching factor

Intuition

Picture the model at each token facing a fork in the road. If it were perfectly torn between exactly k equally-likely next words, its perplexity would be exactly k — the effective number of choices it is juggling.

So perplexity 3.3 means the model is about as uncertain as flipping between roughly three words at each step. A perfect model that always assigns probability 1 to the true token has perplexity 1 (no branching); a uniform model over a vocabulary of size V has perplexity V.

That is why lower is better and why perplexity is comparable across models on the same text — it is a unit-free surprise, not a raw loss.

84. Restore the missing line: Perplexity: how surprised is the model?

Fill the middle

Fill in the blanks

From Perplexity: how surprised is the model? — one line has had its right-hand side removed. Put it back.

import numpy as np
# model's probability on each true next token, 5-token sentence
q = np.array([0.5, 0.25, 0.5, 0.1, 0.4])
nll = -np.mean(np.log(q)) # average negative log-likelihood
ppl = np.exp(nll)
print(round(nll,4), round(ppl,4))

Why: ppl is what everything below it consumes, so the wrong expression here fails later and somewhere else. The model is, on average, about as uncertain as a fair choice among ~3.3 tokens.

85. Perplexity: how surprised is the model?

Worked example

Perplexity is exp of the average negative log-likelihood the model assigned to the true next tokens. Lower is better; perplexity k means the model is as unsure as if choosing uniformly among k options.

\[ \text{PPL} = \exp\!\Big(-\frac{1}{T}\sum_{t=1}^{T}\ln q_t\Big), \qquad q_t = \text{model's prob of the true token} \]

import numpy as np
# model's probability on each true next token, 5-token sentence
q = np.array([0.5, 0.25, 0.5, 0.1, 0.4])
nll = -np.mean(np.log(q))          # average negative log-likelihood
ppl = np.exp(nll)
print(round(nll,4), round(ppl,4))

avg NLL = 1.1983 → perplexity = exp(1.1983) = 3.3145

Why: The model is, on average, about as uncertain as a fair choice among ~3.3 tokens. Ties straight back to cross-entropy from Lesson 14 — perplexity is just its exponential.

quantityvalue
token probs q0.5, 0.25, 0.5, 0.1, 0.4
avg NLL = −mean(ln q)1.1983
perplexity = exp(NLL)3.3145

86. Say it in words: Perplexity: how surprised is the model?

Translation

\( \text{PPL} = \exp\!\Big(-\frac{1}{T}\sum_{t=1}^{T}\ln q_t\Big), \qquad q_t = \text{model's prob of the true token} \)

Draw it

Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.

87. Predict the next row: BLEU idea: clipped n-gram precision

Pattern

Predict first

The table runs: the | 3 | 2 | 2 · cat | 1 | 1 | 1 · mat | 1 | 1 | 1

In BLEU idea: clipped n-gram precision, given the rows so far: what is the next one — the row where word is total / p1?

Correct: total / p1 | 5 | — | 4 → 0.8

wordcand countref countclipped
the322
cat111
mat111
total / p15—4 → 0.8

Why: The relationship between the columns, not the individual numbers, is what generates the next row. Without clipping, the three 'the's would each count and precision would be inflated.

88. BLEU idea: clipped n-gram precision

Worked example

BLEU rewards candidate n-grams that appear in a reference — but clips each count at how often it appears in the reference, so spamming a word can't inflate the score. Here is the unigram (1-gram) precision, the heart of BLEU.

from collections import Counter
ref  = "the cat sat on the mat".split()
cand = "the the the cat mat".split()
ref_c, cand_c = Counter(ref), Counter(cand)
clipped = sum(min(cand_c[w], ref_c[w]) for w in cand_c)
p1 = clipped / len(cand)
print(clipped, len(cand), round(p1,4))

'the' appears 3× in candidate but only 2× in reference → clipped to 2

Why: Without clipping, the three 'the's would each count and precision would be inflated. min(3, 2) caps it at 2. 'cat' and 'mat' each match once.

clipped matches = 2 + 1 + 1 = 4, over 5 candidate words → p1 = 0.8

Why: Unigram precision is 4/5 = 0.8. Full BLEU multiplies precisions across n-gram sizes and adds a brevity penalty, but this clipped-precision core is the exam-relevant idea.

wordcand countref countclipped
the322
cat111
mat111
total / p15—4 → 0.8

89. Which is which, by cand count

Discrimination

Sort into buckets

Sort these by cand count, from memory, without looking back at BLEU idea: clipped n-gram precision. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

3
the
1
cat; mat
5
total / p1
g1
cand count is "3" for the — that is what the table on "BLEU idea: clipped n-gram precision" records, and it is the single property separating this group from the rest.
g2
cand count is "1" for cat, mat — that is what the table on "BLEU idea: clipped n-gram precision" records, and it is the single property separating this group from the rest.
g3
cand count is "5" for total / p1 — that is what the table on "BLEU idea: clipped n-gram precision" records, and it is the single property separating this group from the rest.

90. What has to be given first: ROUGE idea: n-gram recall

Missing information

Discussion prompt

Where BLEU asks 'how much of my output is in the reference?' (precision), ROUGE asks 'how much of the reference did I cover?' (recall) — the natural fit for summarization, where you want to capture the key content.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

The candidate 'the cat mat' covers 'the' (once), 'cat', and 'mat' — 3 of the reference's 6 tokens. Recall = 3/6 = 0.5. Note the same overlap divided by candidate length would be precision.

91. ROUGE idea: n-gram recall

Worked example

Where BLEU asks 'how much of my output is in the reference?' (precision), ROUGE asks 'how much of the reference did I cover?' (recall) — the natural fit for summarization, where you want to capture the key content.

from collections import Counter
ref  = "the cat sat on the mat".split()
cand = "the cat mat".split()
ref_c, cand_c = Counter(ref), Counter(cand)
overlap = sum(min(cand_c[w], ref_c[w]) for w in ref_c)
recall = overlap / len(ref)
print(overlap, len(ref), round(recall,4))

overlap = 3 covered words out of 6 reference words → ROUGE-1 = 0.5

Why: The candidate 'the cat mat' covers 'the' (once), 'cat', and 'mat' — 3 of the reference's 6 tokens. Recall = 3/6 = 0.5. Note the same overlap divided by candidate length would be precision.

measuredenominatorvalue
overlap (clipped)—3
ROUGE-1 recalllen(ref) = 60.5
(precision would use len(cand) = 3)31.0

92. Guess the shape of the answer: FID idea: distance between distributions

Estimation

Predict first

For image generation there is no reference per image. FID embeds real and generated images into a feature space and measures the distance between the two Gaussians (means μ, covariances Σ). Lower = generated set looks more like real. A 1-D toy shows the formula:

Commit before you compute: what does FID idea: distance between distributions come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: FID = 0.25 + (1 + 1.44 − 2·1.2) = 0.25 + 0.04 = 0.29

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The mean gap contributes (0.5)²=0.25; the variance mismatch contributes 1+1.44−2√(1.44)=0.04.

93. FID idea: distance between distributions

Worked example

For image generation there is no reference per image. FID embeds real and generated images into a feature space and measures the distance between the two Gaussians (means μ, covariances Σ). Lower = generated set looks more like real. A 1-D toy shows the formula:

\[ \text{FID} = \lVert \mu_r - \mu_f \rVert^2 + \operatorname{Tr}\!\big(\Sigma_r + \Sigma_f - 2(\Sigma_r \Sigma_f)^{1/2}\big) \]

import numpy as np
mu_r, var_r = 0.0, 1.0     # real features  ~ N(0, 1)
mu_f, var_f = 0.5, 1.44    # fake features  ~ N(0.5, 1.2^2)
fid = (mu_r-mu_f)**2 + var_r + var_f - 2*np.sqrt(var_r*var_f)
print(round(fid,4))

FID = 0.25 + (1 + 1.44 − 2·1.2) = 0.25 + 0.04 = 0.29

Why: The mean gap contributes (0.5)²=0.25; the variance mismatch contributes 1+1.44−2√(1.44)=0.04. In 1-D the trace and matrix sqrt reduce to plain variance and √. Match real distribution → FID → 0.

termvalue
mean gap ‖μr−μf‖²0.25
variance term σr²+σf²−2σrσf0.04
FID0.29

94. Work backwards from the answer: FID idea: distance between distributions

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

FID = 0.25 + (1 + 1.44 − 2·1.2) = 0.25 + 0.04 = 0.29

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

For image generation there is no reference per image. FID embeds real and generated images into a feature space and measures the distance between the two Gaussians (means μ, covariances Σ). Lower = generated set looks more like real. A 1-D toy shows the formula:

95. Check yourself

Section

Part 6 of 6 — then build it

96. Rule out three: Check yourself — F1 is a harmonic mean

Elimination

Eliminate the wrong options

A model has precision = 0.8 and recall = 0.6. Its F1 score is:

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. ≈ 0.686
  • B. 0.70 (the average)
  • C. 0.48 (the product)
  • D. 1.40 (the sum)

Survives elimination: A

Why: F1 = 2PR/(P+R) = 2(0.8)(0.6)/(0.8+0.6) = 0.96/1.4 ≈ 0.686. The harmonic mean sits below the arithmetic mean, pulled toward the smaller value (recall).

97. Check yourself — F1 is a harmonic mean

Check

Compute it on paper before clicking.

Check your understanding

A model has precision = 0.8 and recall = 0.6. Its F1 score is:

  • A. ≈ 0.686 (correct)
  • B. 0.70 (the average)
  • C. 0.48 (the product)
  • D. 1.40 (the sum)

Answer: A

Why: F1 = 2PR/(P+R) = 2(0.8)(0.6)/(0.8+0.6) = 0.96/1.4 ≈ 0.686. The harmonic mean sits below the arithmetic mean, pulled toward the smaller value (recall).

Why B tempts people
0.70 is the ARITHMETIC mean (0.8+0.6)/2. F1 is the HARMONIC mean, always ≤ the arithmetic mean, and it penalizes imbalance between P and R.
Why C tempts people
0.48 = P·R is just the product. F1 doubles the product and divides by the sum P+R.
Why D tempts people
1.40 = P+R is the denominator of the F1 formula, not the score itself — you stopped one step early.

98. Answer it before you see the options: Check yourself — accuracy on imbalance

Prediction

Predict first

A dataset is 99% negative. A model predicts 'negative' for every example. Which is true?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: Accuracy is 0.99 but recall for the positive class is 0

Why: Predicting all-negative is right on every true negative (99% of the data) so accuracy = 0.99, but it catches zero positives, so recall = TP/(TP+FN) = 0 and F1 = 0. Accuracy hides a useless model.

99. Check yourself — accuracy on imbalance

Check

Recall the all-negative baseline.

Check your understanding

A dataset is 99% negative. A model predicts 'negative' for every example. Which is true?

  • A. Accuracy is 0.99 but recall for the positive class is 0 (correct)
  • B. Both accuracy and recall are 0.99
  • C. Accuracy is 0.50
  • D. F1 is 0.99

Answer: A

Why: Predicting all-negative is right on every true negative (99% of the data) so accuracy = 0.99, but it catches zero positives, so recall = TP/(TP+FN) = 0 and F1 = 0. Accuracy hides a useless model.

Why B tempts people
Recall = TP/(TP+FN) = 0/(all positives) = 0, not 0.99 — no positive is ever caught, so recall cannot be high.
Why C tempts people
Accuracy is the fraction correct = 990/1000 = 0.99, not 0.50. The model is right on every negative.
Why D tempts people
F1 = 2PR/(P+R) needs nonzero recall; with recall 0 (and precision undefined/0), F1 = 0, not 0.99.

100. How sure are you: Check yourself — what AUC means

Commit first

Predict first

A classifier has AUC-ROC = 0.90. The most precise interpretation is:

Commit to an answer, then rate it — certain, fairly sure, or guessing — and write the rating down before you turn the page.

Correct: A random positive is scored higher than a random negative 90% of the time

Why: AUC-ROC equals P(random positive scored above random negative) — a threshold-free measure of ranking quality. Equivalently it is the area under the ROC curve.

The rating matters as much as the answer: confident-and-wrong is the combination that survives revision, because nothing about it feels like it needs revisiting.

101. Check yourself — what AUC means

Check

Ranking, not labels.

Check your understanding

A classifier has AUC-ROC = 0.90. The most precise interpretation is:

  • A. A random positive is scored higher than a random negative 90% of the time (correct)
  • B. The model is 90% accurate
  • C. 90% of predicted probabilities exceed 0.9
  • D. Precision equals 0.90 at the 0.5 threshold

Answer: A

Why: AUC-ROC equals P(random positive scored above random negative) — a threshold-free measure of ranking quality. Equivalently it is the area under the ROC curve.

Why B tempts people
Accuracy depends on a chosen threshold; AUC is threshold-free and measures ranking, not the fraction of correct labels. Our running model has AUC 0.8857 but accuracy 0.75.
Why C tempts people
AUC ignores the absolute probability values entirely — only their relative ORDERING between the two classes matters.
Why D tempts people
Precision is a single-threshold quantity; AUC aggregates ranking across ALL thresholds and never fixes one.

102. Answer it before you see the options: Check yourself — RMSE vs MAE

Prediction

Predict first

You compute RMSE = 1.23 and MAE = 0.63 on the same predictions. What does the large gap most likely indicate?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: A few large errors (outliers) are inflating RMSE relative to MAE

Why: RMSE squares errors, so a few big misses pull it well above MAE. A wide RMSE/MAE ratio (here ~1.95) is a signature of outliers or high error variance.

103. Check yourself — RMSE vs MAE

Check

Think about what squaring does.

Check your understanding

You compute RMSE = 1.23 and MAE = 0.63 on the same predictions. What does the large gap most likely indicate?

  • A. A few large errors (outliers) are inflating RMSE relative to MAE (correct)
  • B. A bug — RMSE can never exceed MAE
  • C. The errors are all identical in size
  • D. R² must be negative

Answer: A

Why: RMSE squares errors, so a few big misses pull it well above MAE. A wide RMSE/MAE ratio (here ~1.95) is a signature of outliers or high error variance.

Why B tempts people
RMSE ≥ MAE always (Jensen's inequality); a gap is expected, not a bug. RMSE below MAE would be the impossible case.
Why C tempts people
Identical error sizes give RMSE = MAE exactly (ratio 1). A large gap means the opposite — errors vary a lot.
Why D tempts people
R² measures variance explained and is unrelated to the RMSE/MAE gap; a model can have a large gap and still a high positive R².

104. Rule out three: Check yourself — log-loss

Elimination

Eliminate the wrong options

Model A outputs 0.9 on the correct class every time; Model B outputs 0.6 on the correct class every time. Both are 100% accurate. Which has the lower (better) log-loss?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Model A — it is more confident on the correct class, so −ln(p) is smaller
  • B. Model B — hedging is safer
  • C. They tie, because accuracy is identical
  • D. Neither — log-loss ignores probabilities

Survives elimination: A

Why: Log-loss per example is −ln(p) on the correct class: −ln(0.9)=0.105 for A vs −ln(0.6)=0.511 for B. A is ~5× lower, so A is better. Log-loss rewards calibrated confidence, which accuracy cannot see.

105. Check yourself — log-loss

Check

Two models agree on every label.

Check your understanding

Model A outputs 0.9 on the correct class every time; Model B outputs 0.6 on the correct class every time. Both are 100% accurate. Which has the lower (better) log-loss?

  • A. Model A — it is more confident on the correct class, so −ln(p) is smaller (correct)
  • B. Model B — hedging is safer
  • C. They tie, because accuracy is identical
  • D. Neither — log-loss ignores probabilities

Answer: A

Why: Log-loss per example is −ln(p) on the correct class: −ln(0.9)=0.105 for A vs −ln(0.6)=0.511 for B. A is ~5× lower, so A is better. Log-loss rewards calibrated confidence, which accuracy cannot see.

Why B tempts people
Hedging raises log-loss when you are right: −ln(0.6) > −ln(0.9). Confidence on the correct class is rewarded, not punished, as long as you are right.
Why C tempts people
Accuracy is blind to probabilities; log-loss is exactly the metric that separates two equally-accurate models by their confidence.
Why D tempts people
The opposite — log-loss is computed entirely from the predicted probabilities via −[y ln p + (1−y) ln(1−p)].

106. Your turn: build the metrics

Section

The project

107. Project: a metrics library from scratch

Concept

Implement the core classification metrics on the running spam data using only NumPy, then prove each one matches sklearn. You derived every piece — now assemble the library.

#requirementtool
1confusion counts → precision, recallboolean masks
2F1 = 2PR/(P+R)harmonic mean
3AUC via pairwise ranking vs sklearnroc_auc_score

Build rules: type every line yourself, run after each milestone, build the four counts with boolean masks, and compare AUC to roc_auc_score with np.isclose. If a count looks wrong, print the masks — don't guess.

108. By analogy: Project: a metrics library from scratch

Analogy

Discussion prompt

Explain Project: a metrics library from scratch by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Implement the core classification metrics on the running spam data using only NumPy, then prove each one matches sklearn. You derived every piece — now assemble the library.

109. Milestone 1 — the four counts

Worked example

Your turn: threshold p at 0.5, then build TP, FP, FN, TN with masks. Predict the counts from the by-hand result before you print.

Hint: pred = (p >= 0.5).astype(int), then TP = ((pred==1)&(y==1)).sum(), and so on. The four should sum to 12.

import numpy as np
y    = np.array([1,1,1,1,1,0,0,0,0,0,0,0])
p    = np.array([0.95,0.90,0.80,0.60,0.30,0.70,0.55,0.40,0.20,0.15,0.10,0.05])
pred = (p >= 0.5).astype(int)
TP = ((pred==1)&(y==1)).sum(); FP = ((pred==1)&(y==0)).sum()
FN = ((pred==0)&(y==1)).sum(); TN = ((pred==0)&(y==0)).sum()
print(int(TP), int(FP), int(FN), int(TN))
print(round(TP/(TP+FP),4), round(TP/(TP+FN),4))
quantityvalue
TP, FP, FN, TN4, 2, 1, 5
precision0.6667
recall0.8000

110. Watch it run: Milestone 1 — the four counts

Pattern

Step through it

Step through Milestone 1 — the four counts one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: quantity is TP, FP, FN, TN
  2. Step 2: quantity is precision
  3. Step 3: quantity is recall

111. Milestone 2 — F1 from P and R

Worked example

Your turn: combine precision and recall into F1. Predict whether F1 lands above or below their arithmetic mean of 0.733.

Hint: f1 = 2*prec*rec/(prec+rec) — the harmonic mean, which sits below the arithmetic mean.

prec, rec = 4/6, 4/5
f1 = 2*prec*rec/(prec+rec)
print(round(prec,4), round(rec,4), round(f1,4))
quantityvalue
precision0.6667
recall0.8000
F1 (harmonic mean)0.7273
(arithmetic mean, for contrast)0.7333

112. Milestone 3 — AUC vs sklearn

Worked example

Your turn: compute AUC as the fraction of correctly-ordered positive/negative pairs and verify against sklearn with np.isclose. Predict the value from the 31/35 hand count.

Hint: pos, neg = p[y==1], p[y==0], then average (a>b)+0.5*(a==b) over all pos/neg pairs; compare to roc_auc_score.

import numpy as np
from sklearn.metrics import roc_auc_score
y = np.array([1,1,1,1,1,0,0,0,0,0,0,0])
p = np.array([0.95,0.90,0.80,0.60,0.30,0.70,0.55,0.40,0.20,0.15,0.10,0.05])
pos, neg = p[y==1], p[y==0]
auc = np.mean([(a>b)+0.5*(a==b) for a in pos for b in neg])
print(round(auc,4), np.isclose(auc, roc_auc_score(y,p)))
sourceAUCmatch
pairwise ranking (31/35)0.8857—
sklearn roc_auc_score0.8857True

113. What each one costs: Milestone 3 — AUC vs sklearn

Trade off

Comparison matrix

From Milestone 3 — AUC vs sklearn: every row here is a choice with a cost. Fill the AUC column, then say which row you would actually pick and what you give up for it.

sourceAUCmatch
pairwise ranking (31/35)0.8857—
sklearn roc_auc_score0.8857True

114. The full program

Concept

import numpy as np
from sklearn.metrics import roc_auc_score, average_precision_score

def confusion(y, pred):
    TP=((pred==1)&(y==1)).sum(); FP=((pred==1)&(y==0)).sum()
    FN=((pred==0)&(y==1)).sum(); TN=((pred==0)&(y==0)).sum()
    return TP, FP, FN, TN

def prf(y, pred):
    TP, FP, FN, TN = confusion(y, pred)
    P = TP/(TP+FP); R = TP/(TP+FN)
    return P, R, 2*P*R/(P+R)

def auc_pairwise(y, p):
    pos, neg = p[y==1], p[y==0]
    wins = sum((a>b)+0.5*(a==b) for a in pos for b in neg)
    return wins / (len(pos)*len(neg))

y    = np.array([1,1,1,1,1,0,0,0,0,0,0,0])
p    = np.array([0.95,0.90,0.80,0.60,0.30,0.70,0.55,0.40,0.20,0.15,0.10,0.05])
pred = (p >= 0.5).astype(int)
P, R, F1 = prf(y, pred)
auc = auc_pairwise(y, p)
print('P,R,F1 =', tuple(round(v,4) for v in (P,R,F1)))
print('AUC =', round(auc,4), '== sklearn', np.isclose(auc, roc_auc_score(y,p)))
printed linevalue
P,R,F1 =(0.6667, 0.8, 0.7273)
AUC =0.8857 == sklearn True

If your P, R, F1 read 0.6667, 0.8, 0.7273 and AUC matches sklearn, you can score a classifier from the counts up — and you know exactly what each number sees.

115. Fill in: value for The full program

Comparison

Comparison matrix

From The full program: refill the value column from what you know. The rest of the table is as it appeared.

printed linevalue
P,R,F1 =(0.6667, 0.8, 0.7273)
AUC =0.8857 == sklearn True

116. Show it off

Concept

Slides closed, out loud: explain (1) why accuracy fails on 99%-negative data, (2) what AUC-ROC means as a probability and how the 31/35 pairs produced 0.8857, and (3) when RMSE diverges from MAE.

Stretch: add average_precision_score to the library and show it drops below ROC-AUC on the imbalanced set; then reproduce roc_curve by sweeping thresholds and confirm the (FPR, TPR) ladder you traced by hand. Metrics drive model selection (Week 22) and GAN/diffusion FID (Week 48).

117. Break it if you can: Show it off

Counterexample

Discussion prompt

Slides closed, out loud: explain (1) why accuracy fails on 99%-negative data, (2) what AUC-ROC means as a probability and how the 31/35 pairs produced 0.8857, and (3) when RMSE diverges from MAE.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

118. Connect it up: Lesson 26: Evaluation Metrics

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — The confusion matrix · Precision, recall, F1 · The threshold trade-off · Regression metrics · Generation metrics · Check yourself. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

119. What you can do now

Recap

metricthe one thing to remember
F12PR/(P+R) — harmonic mean, pulled to the smaller
accuracylies on imbalanced data — use F1 / recall / AUC-PR
AUC-ROCP(random pos > random neg); threshold-free ranking
log-losscalibrated probs; confident-wrong explodes it
RMSE vs MAERMSE ≥ MAE; the gap is your outlier detector
generationperplexity=exp(NLL), BLEU=precision, ROUGE=recall, FID=distribution distance

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 26 (Week 9 — Evaluation Metrics) — Barron · USAAIO Round 2 Preparation, 2026
  2. scikit-learn metrics: precision_recall_fscore, roc_auc_score, average_precision_score, log_loss, mean_squared_error, r2_score
  3. Fawcett, An introduction to ROC analysis — Pattern Recognition Letters 27 (2006) 861–874
  4. Papineni et al., BLEU; Lin, ROUGE; Heusel et al., FID (GANs / two-time-scale update rule) — ACL 2002 · WAS 2004 · NeurIPS 2017
  5. Every count, curve point, AUC, and metric produced by real execution — numpy 2.2.6 + scikit-learn 1.9.0, verification run July 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108