USAAIO Lesson 35, from Week 12, fully worked. It defines the confusion matrix and every rate derived from it, then builds the ROC curve threshold by threshold on one running 10-example dataset. AUC is computed three independent ways - by the trapezoid rule, by sklearn, and by counting concordant pairs - and all three land on 0.875. The precision-recall curve and average precision are derived the same way, and the lesson shows why ROC-AUC misleads under 2% imbalance while PR-AUC, at 0.111, tells the truth. It then defines calibration and the Expected Calibration Error bin by bin, and sweeps temperature scaling to its ECE minimum at T=3. Every snippet runs standalone in a fresh interpreter, and every number was produced by real execution. The lesson runs to 61 slides.
Subject: Machine Learning · 110 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 35 · Week 12 (Evaluation)
Threshold-free evaluation, and probabilities you can trust. We build the ROC curve one threshold at a time, compute its AUC three independent ways, expose why it lies under imbalance, then fix an overconfident model with temperature scaling — every number produced by real execution.
Objectives
TPR, FPR, and precision from it by handAUCAUC = P(random positive scores above random negative) by counting concordant pairs — the same 0.875 three waysROC-AUC flatters an imbalanced model while PR-AUC doesn'tECE to its minimum with temperature scaling TWarm-up
Discussion prompt
Before we open Lesson 35: ROC, PR Curves & Calibration: without looking back, what was the main idea of Concentration Inequalities & Generalization, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
Markov, Chebyshev, and Hoeffding inequalities, VC dimension as a complexity measure, the VC generalization bound, and why overparameterized networks still generalize (double descent). Verify the inequalities empirically and compute a VC generalization bound.
Section
Part 1 of 6 — the foundation
Concept
A classifier outputs a score s ∈ [0,1] — how positive it thinks each example is. To make an actual decision you pick a threshold t and call everything with s ≥ t positive.
Change t and the decisions change. Evaluation metrics like accuracy depend on that one arbitrary knob — the whole point of ROC and PR curves is to escape it by sweeping t across all its values at once.
Counterexample
Discussion prompt
A classifier outputs a score s ∈ [0,1] — how positive it thinks each example is. To make an actual decision you pick a threshold t and call everything with s ≥ t positive.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Concept
One running example for the whole lesson. Ten examples, each with a model score s and a true label y (1 = positive). Sorted high score to low:
| # | score s | true y |
|---|---|---|
| 1 | 0.95 | 1 |
| 2 | 0.90 | 1 |
| 3 | 0.80 | 0 |
| 4 | 0.70 | 1 |
| 5 | 0.55 | 0 |
| 6 | 0.45 | 1 |
| 7 | 0.40 | 0 |
| 8 | 0.30 | 0 |
| 9 | 0.20 | 0 |
| 10 | 0.10 | 0 |
There are P = 4 positives and N = 6 negatives. Notice the ranking is good but not perfect: example 3 (a negative) scores 0.80, above example 4 (a positive) at 0.70. That one swap is what keeps the AUC below 1.
Pattern
Step through it
Step through The data: ten scored examples one row at a time. What is driving the change, and what would the row after the last one be?
Concept
Fix a threshold. Every example falls into exactly one of four boxes, comparing what the model predicted against the truth.
| predicted + | predicted − | |
|---|---|---|
| actually + | TP (true positive) | FN (false negative) |
| actually − | FP (false positive) | TN (true negative) |
TP / FP / FN / TN — TP: called positive, was positive (a hit). FP: called positive, was negative (a false alarm). FN: called negative, was positive (a miss). TN: called negative, was negative (a correct rejection). Every metric today is built from just these four numbers.
Comparison
Comparison matrix
From The four confusion-matrix counts: refill the predicted − column from what you know. The rest of the table is as it appeared.
| predicted + | predicted − | |
|---|---|---|
| actually + | TP (true positive) | FN (false negative) |
| actually − | FP (false positive) | TN (true negative) |
Estimation
Predict first
Threshold t = 0.5: predict positive when s ≥ 0.5. That flags examples 1–5 (scores 0.95, 0.90, 0.80, 0.70, 0.55). Compare each flag to its true label:
Commit before you compute: what does Count the four boxes at t = 0.5 come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Predicted negative: #6(y=1)✗ #7(y=0)✓ #8(y=0)✓ #9(y=0)✓ #10(y=0)✓
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Example 6 is a positive we missed (FN = 1); the other four are correctly rejected negatives (TN = 4).
Worked example
Threshold t = 0.5: predict positive when s ≥ 0.5. That flags examples 1–5 (scores 0.95, 0.90, 0.80, 0.70, 0.55). Compare each flag to its true label:
Predicted positive: #1(y=1)✓ #2(y=1)✓ #3(y=0)✗ #4(y=1)✓ #5(y=0)✗
Why: Three of the five flagged examples are truly positive (TP = 3); two are truly negative false alarms (FP = 2).
Predicted negative: #6(y=1)✗ #7(y=0)✓ #8(y=0)✓ #9(y=0)✓ #10(y=0)✓
Why: Example 6 is a positive we missed (FN = 1); the other four are correctly rejected negatives (TN = 4).
\[ TP = 3, \quad FP = 2, \quad FN = 1, \quad TN = 4 \qquad (TP+FP+FN+TN = 10\;\checkmark) \]
Notation
Annotate
From Count the four boxes at t = 0.5 — read this one piece at a time. What is each part doing?
On: \( TP = 3, \quad FP = 2, \quad FN = 1, \quad TN = 4 \qquad (TP+FP+FN+TN = 10\;\checkmark) \)
Concept
The two axes of the ROC curve, plus precision, are ratios of those four counts. Learn which denominator each one uses — that choice is the whole story of imbalance.
\[ \text{TPR (recall)} = \frac{TP}{TP+FN}, \qquad \text{FPR} = \frac{FP}{FP+TN}, \qquad \text{precision} = \frac{TP}{TP+FP} \]
TPR (recall) is the fraction of real positives caught. FPR is the fraction of real negatives falsely flagged. Precision is the fraction of flags that were right. TPR and FPR split by the true class; precision splits by the predicted class.
Analogy
Discussion prompt
Explain Three rates read off the boxes by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
The two axes of the ROC curve, plus precision, are ratios of those four counts. Learn which denominator each one uses — that choice is the whole story of imbalance.
Missing information
Discussion prompt
Plug TP=3, FP=2, FN=1, TN=4 into the three formulas, then confirm with a runnable snippet that rebuilds the whole thing from the raw scores:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
The threshold t = 0.5 gives a single (FPR, TPR) = (0.333, 0.75) point. The ROC curve is what you get by sweeping t and collecting all such points.
Worked example
Plug TP=3, FP=2, FN=1, TN=4 into the three formulas, then confirm with a runnable snippet that rebuilds the whole thing from the raw scores:
import numpy as np
scores = np.array([0.95,0.90,0.80,0.70,0.55,0.45,0.40,0.30,0.20,0.10])
y = np.array([1, 1, 0, 1, 0, 1, 0, 0, 0, 0 ])
thr = 0.5
pred = (scores >= thr).astype(int)
TP = int(((pred==1)&(y==1)).sum())
FP = int(((pred==1)&(y==0)).sum())
FN = int(((pred==0)&(y==1)).sum())
TN = int(((pred==0)&(y==0)).sum())
print("TP FP FN TN =", TP, FP, FN, TN)
print("TPR =", TP/(TP+FN))
print("FPR =", FP/(FP+TN))
print("precision =", TP/(TP+FP))TPR = 3/4 = 0.75, FPR = 2/6 = 0.333, precision = 3/5 = 0.60
Why: The threshold t = 0.5 gives a single (FPR, TPR) = (0.333, 0.75) point. The ROC curve is what you get by sweeping t and collecting all such points.
| quantity | formula | value (verified) |
|---|---|---|
| TP, FP, FN, TN | count boxes | 3, 2, 1, 4 |
| TPR = TP/(TP+FN) | 3/(3+1) | 0.750 |
| FPR = FP/(FP+TN) | 2/(2+4) | 0.333 |
| precision = TP/(TP+FP) | 3/(3+2) | 0.600 |
Trade off
Comparison matrix
From The rates at t = 0.5, verified in code: every row here is a choice with a cost. Fill the value (verified) column, then say which row you would actually pick and what you give up for it.
| quantity | formula | value (verified) |
|---|---|---|
| TP, FP, FN, TN | count boxes | 3, 2, 1, 4 |
| TPR = TP/(TP+FN) | 3/(3+1) | 0.750 |
| FPR = FP/(FP+TN) | 2/(2+4) | 0.333 |
| precision = TP/(TP+FP) | 3/(3+2) | 0.600 |
Section
Part 2 of 6 — sweep every threshold
Intuition
The (0.333, 0.75) point above depended entirely on our arbitrary t = 0.5. A different threshold gives a different trade-off — lower t catches more positives (higher TPR) but raises false alarms (higher FPR).
Rather than defend one threshold, plot all of them. Slide t from +∞ (nothing flagged) down to −∞ (everything flagged) and trace the (FPR, TPR) point as it moves.
That trace is the ROC curve — Receiver Operating Characteristic. It shows the model's entire achievable frontier of trade-offs in one picture, independent of any single threshold choice.
Explain it
Discussion prompt
Explain Why one point is not enough to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Rather than defend one threshold, plot all of them. Slide t from +∞ (nothing flagged) down to −∞ (everything flagged) and trace the (FPR, TPR) point as it moves.
Concept
Start with t above every score: nothing is flagged, so TP = FP = 0 and the point sits at the origin (0, 0). Now lower t past one example's score at a time.
Each example you cross adds one flag. If that example is a real positive, TP grows → the point steps up. If it is a real negative, FP grows → the point steps right.
So the ROC curve is a staircase from (0,0) to (1,1): up on a positive, right on a negative. A good model front-loads all its ups (climbs the left edge early) before the rights.
Pattern
Predict first
The table runs: (start) | — | 0 | 0 | 0.000 | 0.000 · 0.95 | 1 | 1 | 0 | 0.250 | 0.000 · 0.90 | 1 | 2 | 0 | 0.500 | 0.000 · 0.80 | 0 | 2 | 1 | 0.500 | 0.167 · 0.70 | 1 | 3 | 1 | 0.750 | 0.167 · 0.55 | 0 | 3 | 2 | 0.750 | 0.333 · 0.45 | 1 | 4 | 2 | 1.000 | 0.333 · 0.40 | 0 | 4 | 3 | 1.000 | 0.500 · 0.30 | 0 | 4 | 4 | 1.000 | 0.667 · 0.20 | 0 | 4 | 5 | 1.000 | 0.833
In Build the ROC curve from scratch, given the rows so far: what is the next one — the row where crossed score is 0.10?
Correct: 0.10 | 0 | 4 | 6 | 1.000 | 1.000
| crossed score | label | tp | fp | TPR | FPR |
|---|---|---|---|---|---|
| (start) | — | 0 | 0 | 0.000 | 0.000 |
| 0.95 | 1 | 1 | 0 | 0.250 | 0.000 |
| 0.90 | 1 | 2 | 0 | 0.500 | 0.000 |
| 0.80 | 0 | 2 | 1 | 0.500 | 0.167 |
| 0.70 | 1 | 3 | 1 | 0.750 | 0.167 |
| 0.55 | 0 | 3 | 2 | 0.750 | 0.333 |
| 0.45 | 1 | 4 | 2 | 1.000 | 0.333 |
| 0.40 | 0 | 4 | 3 | 1.000 | 0.500 |
| 0.30 | 0 | 4 | 4 | 1.000 | 0.667 |
| 0.20 | 0 | 4 | 5 | 1.000 | 0.833 |
| 0.10 | 0 | 4 | 6 | 1.000 | 1.000 |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. Descending score is exactly the order thresholds cross the examples as t drops.
Worked example
Sort by descending score, walk the examples, and after each one record the running (FPR, TPR). This snippet is complete and self-contained — it prints every point on the curve:
import numpy as np
from sklearn.metrics import roc_auc_score
scores = np.array([0.95,0.90,0.80,0.70,0.55,0.45,0.40,0.30,0.20,0.10])
y = np.array([1, 1, 0, 1, 0, 1, 0, 0, 0, 0 ])
P, N = y.sum(), (1-y).sum()
tp = fp = 0
tpr = [0.0]; fpr = [0.0]
for i in np.argsort(-scores):
tp += y[i]; fp += 1 - y[i]
tpr.append(tp / P); fpr.append(fp / N)
print(round(np.trapz(tpr, fpr), 4), round(roc_auc_score(y, scores), 4))np.argsort(-scores) walks examples high score → low
Why: Descending score is exactly the order thresholds cross the examples as t drops. Each loop pass lowers the effective threshold past one more example.
tp += y[i] steps up; fp += 1 - y[i] steps right
Why: If y[i]=1 the example is positive so TP increments (up); if y[i]=0 it is negative so FP increments (right). Exactly the staircase rule.
| crossed score | label | tp | fp | TPR | FPR |
|---|---|---|---|---|---|
| (start) | — | 0 | 0 | 0.000 | 0.000 |
| 0.95 | 1 | 1 | 0 | 0.250 | 0.000 |
| 0.90 | 1 | 2 | 0 | 0.500 | 0.000 |
| 0.80 | 0 | 2 | 1 | 0.500 | 0.167 |
| 0.70 | 1 | 3 | 1 | 0.750 | 0.167 |
| 0.55 | 0 | 3 | 2 | 0.750 | 0.333 |
| 0.45 | 1 | 4 | 2 | 1.000 | 0.333 |
| 0.40 | 0 | 4 | 3 | 1.000 | 0.500 |
| 0.30 | 0 | 4 | 4 | 1.000 | 0.667 |
| 0.20 | 0 | 4 | 5 | 1.000 | 0.833 |
| 0.10 | 0 | 4 | 6 | 1.000 | 1.000 |
Concept
Trace the table: the first two crossings are positives, so the point climbs straight up the left edge to TPR = 0.5 with FPR still 0. Then example 3 (score 0.80, a negative) forces the first step right.
After example 6 (the last positive, score 0.45), TPR hits 1.0 — every positive is caught. The remaining four crossings are all negatives, so the curve runs flat across the top from FPR = 0.333 to 1.0.
Figure (svg): An ROC staircase rising from the bottom-left origin up the left edge to TPR 0.5, stepping up and right through the middle, reaching TPR 1.0 and then running flat across the top to the top-right corner, well above the dashed diagonal.
Section
Part 2 continued — 0.875 every time
Concept
One number summarizes the whole staircase: the area under the ROC curve, AUC ∈ [0,1]. A curve that hugs the top-left corner encloses nearly the full unit square (AUC → 1); the random diagonal encloses half (AUC = 0.5).
For a staircase, the area is a sum of rectangles: at each step right (width ΔFPR), the height is the current TPR. That is exactly what the trapezoid rule np.trapz(tpr, fpr) computes.
Ranking
Put in order
Put the moves of Way 1 — trapezoid area by hand into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. First negative crossed; the height of the rectangle is the TPR standing at that moment, 0.5.
Worked example
The curve only advances rightward at the four negative crossings. Sum TPR × ΔFPR over just those steps (each ΔFPR = 1/6). Read the TPR at each right-step from the table:
Right-step at 0.80: TPR = 0.5, width 1/6 → 0.5·(1/6)
Why: First negative crossed; the height of the rectangle is the TPR standing at that moment, 0.5.
Right-steps at 0.55, 0.40, 0.30, 0.20, 0.10
Why: Heights 0.75 (at 0.55) then 1.0, 1.0, 1.0, 1.0 for the last four negatives. With the 0.80 step above, that is all 6 negatives — 6 right-steps, each width 1/6.
\[ \text{AUC} = \tfrac{1}{6}\big(0.5 + 0.75 + 1.0 + 1.0 + 1.0 + 1.0\big) = \tfrac{1}{6}(5.25) = 0.875 \]
Sum of heights 5.25, ÷ 6 = 0.875
Why: Heights are the TPR at each of the 6 negative crossings: 0.5 (at 0.80), 0.75 (at 0.55), then 1.0 four times (at 0.40, 0.30, 0.20, 0.10). This is the exact area under the staircase.
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Sum of heights 5.25, ÷ 6 = 0.875
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
The curve only advances rightward at the four negative crossings. Sum TPR × ΔFPR over just those steps (each ΔFPR = 1/6). Read the TPR at each right-step from the table:
Fill the middle
Fill in the blanks
From Way 2 — np.trapz matches sklearn — one line has had its right-hand side removed. Put it back.
import numpy as np
from sklearn.metrics import roc_auc_score
scores = np.array([0.95,0.90,0.80,0.70,0.55,0.45,0.40,0.30,0.20,0.10])
y = np.array([1, 1, 0, 1, 0, 1, 0, 0, 0, 0 ])
P, N = y.sum(), (1-y).sum()
tp = fp = 0; tpr = [0.0]; fpr = [0.0]
for i in np.argsort(-scores):
tp += y[i]; fp += 1 - y[i]
tpr.append(tp / P); fpr.append(fp / N)
auc_manual = np.trapz(tpr, fpr)
print(round(auc_manual, 4), round(roc_auc_score(y, scores), 4))
Why: auc_manual is what everything below it consumes, so the wrong expression here fails later and somewhere else. trapz sums trapezoids under the (fpr, tpr) points.
Worked example
The same block from before ends by integrating the curve and printing it beside sklearn's roc_auc_score. Both must agree with our hand value 0.875:
import numpy as np
from sklearn.metrics import roc_auc_score
scores = np.array([0.95,0.90,0.80,0.70,0.55,0.45,0.40,0.30,0.20,0.10])
y = np.array([1, 1, 0, 1, 0, 1, 0, 0, 0, 0 ])
P, N = y.sum(), (1-y).sum()
tp = fp = 0; tpr = [0.0]; fpr = [0.0]
for i in np.argsort(-scores):
tp += y[i]; fp += 1 - y[i]
tpr.append(tp / P); fpr.append(fp / N)
auc_manual = np.trapz(tpr, fpr)
print(round(auc_manual, 4), round(roc_auc_score(y, scores), 4))np.trapz(tpr, fpr) integrates y=tpr over x=fpr
Why: trapz sums trapezoids under the (fpr, tpr) points. On a staircase the trapezoids are rectangles, giving the same 0.875 as the hand sum.
| method | AUC (verified) |
|---|---|
| hand trapezoid | 0.8750 |
| np.trapz(tpr, fpr) | 0.8750 |
| sklearn roc_auc_score | 0.8750 |
Sorting
Sort into buckets
These are the pieces of Lesson 35: ROC, PR Curves & Calibration, out of order. Put each one back under the part of the lesson it belongs to.
Intuition
Here is the interpretation worth memorizing: AUC is the probability that a randomly chosen positive scores higher than a randomly chosen negative.
That is why AUC is threshold-free — it never fixes a t, it only asks whether the model ranks positives above negatives. A perfect ranker always does (AUC = 1); a coin flip does it half the time (0.5).
\[ \text{AUC} = \Pr\big(\,s^{+} > s^{-}\,\big) \quad\text{over random }(s^{+},s^{-})\text{ pairs} \]
Worked example
Make that probability literal: form every (positive, negative) pair, count how many the positive outranks, and divide by the total P·N = 4·6 = 24:
import numpy as np
scores = np.array([0.95,0.90,0.80,0.70,0.55,0.45,0.40,0.30,0.20,0.10])
y = np.array([1, 1, 0, 1, 0, 1, 0, 0, 0, 0 ])
pos = scores[y==1]; neg = scores[y==0]
concordant = 0.0
for sp in pos:
for sn in neg:
if sp > sn: concordant += 1.0
elif sp == sn: concordant += 0.5
print(concordant, len(pos)*len(neg))
print(round(concordant/(len(pos)*len(neg)), 4))Only 3 of the 24 pairs are discordant
Why: A pair is discordant when a negative outranks a positive. Here that happens exactly three times: (pos 0.70 < neg 0.80), (pos 0.45 < neg 0.80), and (pos 0.45 < neg 0.55). Every other pair has the positive on top.
concordant = 24 − 3 = 21 → 21/24 = 0.875
Why: There are no ties, so concordant + discordant = 24. The same 0.875 as the trapezoid and sklearn — the ranking view and the area view are one number.
| method | value (verified) |
|---|---|
| positive scores | [0.95, 0.90, 0.70, 0.45] |
| negative scores | [0.80, 0.55, 0.40, 0.30, 0.20, 0.10] |
| concordant pairs | 21 of 24 |
| AUC = 21/24 | 0.8750 |
Error analysis
Annotate
Walk the callouts on Way 3 — count concordant pairs. Each one is a place this is easy to get subtly wrong.
Concept
Memorize the three landmarks so you can read any AUC instantly:
| classifier | ROC shape | AUC |
|---|---|---|
| perfect | up the left edge, across the top | 1.0 |
| random | the diagonal | 0.5 |
| worse than random | below the diagonal | < 0.5 |
A worse-than-random model still has signal — it just points backwards. Flip its predictions and the AUC becomes 1 − AUC > 0.5. Never discard an AUC of 0.3.
Fill the middle
Fill in the blanks
From Flipping a backwards classifier — one line has had its right-hand side removed. Put it back.
import numpy as np
from sklearn.metrics import roc_auc_score
scores = np.array([0.95,0.90,0.80,0.70,0.55,0.45,0.40,0.30,0.20,0.10])
y = np.array([0, 0, 1, 0, 1, 0, 1, 1, 1, 1 ])
auc = roc_auc_score(y, scores)
auc_flipped = roc_auc_score(y, 1 - scores)
print(round(auc, 3), round(auc_flipped, 3), round(1 - auc, 3))
Why: auc_flipped is what everything below it consumes, so the wrong expression here fails later and somewhere else. Flipping (1 − scores) reverses every ranking, turning each discordant pair into a concordant one.
Worked example
Same scores, but now the labels are inverted so the model ranks positives too low. Its AUC lands below 0.5 — until we flip the scores:
import numpy as np
from sklearn.metrics import roc_auc_score
scores = np.array([0.95,0.90,0.80,0.70,0.55,0.45,0.40,0.30,0.20,0.10])
y = np.array([0, 0, 1, 0, 1, 0, 1, 1, 1, 1 ])
auc = roc_auc_score(y, scores)
auc_flipped = roc_auc_score(y, 1 - scores)
print(round(auc, 3), round(auc_flipped, 3), round(1 - auc, 3))AUC = 0.125 → flip the scores → 0.875
Why: Flipping (1 − scores) reverses every ranking, turning each discordant pair into a concordant one. The flipped AUC is exactly 1 − 0.125 = 0.875.
| model | AUC (verified) |
|---|---|
| as-is (backwards) | 0.125 |
| scores flipped (1 − s) | 0.875 |
| 1 − AUC (matches flip) | 0.875 |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
AUC = 0.125 → flip the scores → 0.875
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Same scores, but now the labels are inverted so the model ranks positives too low. Its AUC lands below 0.5 — until we flip the scores:
Anomaly
Predict first
A student writes this, and it looks reasonable:
ROC plots the two rates; put TPR on the x-axis and FPR on the y-axis, then integrate.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Integrating FPR against TPR measures the area on the OTHER side of the curve.
ROC is FPR on x, TPR on y. Integrate TPR over FPR.
Why: Integrating FPR against TPR measures the area on the OTHER side of the curve. You'd report 0.125 for a genuinely good model and 'conclude' it is worse than random.
Trap
ROC plots the two rates; put TPR on the x-axis and FPR on the y-axis, then integrate.
\[ \text{AUC}_{\text{wrong}} = \int \text{FPR}\; d(\text{TPR}) \]
Axes swapped → area = 1 − 0.875 = 0.125
Why: Integrating FPR against TPR measures the area on the OTHER side of the curve. You'd report 0.125 for a genuinely good model and 'conclude' it is worse than random.
ROC is FPR on x, TPR on y. Integrate TPR over FPR.
\[ \text{AUC} = \int \text{TPR}\; d(\text{FPR}) \]
Correct axes → area = 0.875
Why: np.trapz(tpr, fpr) has y first, x second — tpr integrated over fpr. Mnemonic: the good corner is TOP-LEFT (high TPR, low FPR), so TPR must be the vertical axis.
Notation
Annotate
From Trap: swapping the ROC axes — read this one piece at a time. What is each part doing?
On: \( \text{AUC} = \int \text{TPR}\; d(\text{FPR}) \)
Section
Part 3 of 6 — where ROC lies
Concept
The precision-recall curve plots precision (TP/(TP+FP)) on the y-axis against recall (= TPR = TP/(TP+FN)) on the x-axis, swept over the same thresholds.
Its summary number is the average precision (AP), the PR-curve analogue of AUC — the area under the precision-recall curve.
The key difference from ROC: neither precision nor recall uses TN. PR ignores the true negatives entirely and looks only at the positive class — exactly the class you care about when positives are rare.
Estimation
Predict first
Same descending sweep, but now track precision and recall, and accumulate AP as Σ precision·Δrecall at each positive crossing:
Commit before you compute: what does Build the PR curve and AP from scratch come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: AP steps only on positives, weighting precision by the recall gained
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Each positive crossing lifts recall by 1/P and contributes precision·(that gain).
Worked example
Same descending sweep, but now track precision and recall, and accumulate AP as Σ precision·Δrecall at each positive crossing:
import numpy as np
from sklearn.metrics import average_precision_score
scores = np.array([0.95,0.90,0.80,0.70,0.55,0.45,0.40,0.30,0.20,0.10])
y = np.array([1, 1, 0, 1, 0, 1, 0, 0, 0, 0 ])
P = y.sum(); tp = fp = 0; ap = 0.0; prev_rec = 0.0
for i in np.argsort(-scores):
tp += y[i]; fp += 1 - y[i]
prec = tp / (tp + fp); rec = tp / P
if y[i] == 1:
ap += prec * (rec - prev_rec); prev_rec = rec
print(round(ap, 4), round(average_precision_score(y, scores), 4))AP steps only on positives, weighting precision by the recall gained
Why: Each positive crossing lifts recall by 1/P and contributes precision·(that gain). sklearn defines AP as exactly this sum, Σ (Rₙ − Rₙ₋₁)·Pₙ.
| crossed | label | precision | recall | AP running |
|---|---|---|---|---|
| 0.95 | 1 | 1.000 | 0.250 | 0.250 |
| 0.90 | 1 | 1.000 | 0.500 | 0.500 |
| 0.80 | 0 | 0.667 | 0.500 | 0.500 |
| 0.70 | 1 | 0.750 | 0.750 | 0.688 |
| 0.45 | 1 | 0.667 | 1.000 | 0.854 |
| final AP | — | — | — | 0.8542 |
Discrimination
Sort into buckets
Sort these by label, from memory, without looking back at Build the PR curve and AP from scratch. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Concept
Our hand-accumulated AP is 0.8542, and average_precision_score returns the identical 0.8542. The manual PR curve is faithful.
| method | AP (verified) |
|---|---|
| manual Σ precision·Δrecall | 0.8542 |
| sklearn average_precision_score | 0.8542 |
On this balanced-ish data (4/10 positive) AP 0.854 and AUC 0.875 are close. The gap explodes only when positives become rare — which is the next experiment.
Concept
PR-AUC summarizes the whole curve. If instead you must pick one threshold, the F1 score balances precision and recall at that operating point via their harmonic mean:
\[ F_1 = \frac{2 \cdot \text{precision} \cdot \text{recall}}{\text{precision} + \text{recall}} \]
The harmonic mean punishes imbalance between the two: F1 is high only when precision and recall are both high. A model with precision 1.0 but recall 0.1 scores F1 ≈ 0.18, not 0.55 — one weak side drags it down.
Intuition
Averaging precision and recall the ordinary way would let a lopsided model cheat: precision 1.0, recall 0.1 averages to 0.55 — sounds decent, but the model finds only one positive in ten.
The harmonic mean is dominated by the smaller number. It sits near the weaker of the two, so you can't paper over a terrible recall with a great precision. That is exactly the behavior you want from a single 'is this model actually usable' score.
Same reason a car that goes 100 mph one way and 1 mph back doesn't average 50 mph over the trip — time is dominated by the slow leg. F1 is dominated by the weak metric.
Faded example
Fill in the blanks
Pick the operating threshold, with the scaffolding fading: two lines are gone now — fill both.
import numpy as np
scores = np.array([0.95,0.90,0.80,0.70,0.55,0.45,0.40,0.30,0.20,0.10])
y = np.array([1, 1, 0, 1, 0, 1, 0, 0, 0, 0 ])
P = y.sum(); tp = fp = 0; rows = []
for i in np.argsort(-scores):
tp += y[i]; fp += 1 - y[i]
prec = tp/(tp+fp); rec = tp/P
f1 = 0.0 if prec+rec == 0 else 2precrec/(prec+rec)
rows.append((scores[i], prec, rec, f1))
best = max(rows, key=lambda r: r[3])
print("best F1", round(best[3],3), "at thr", best[0])
Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. At t = 0.45 precision = 0.667 and recall = 1.000, giving the harmonic mean 0.800 — the highest of any threshold.
Worked example
PR-AUC and F1 don't tell you which threshold to deploy. Sweep them, then choose by your requirement: best F1, or the most recall you can get while holding precision ≥ 0.75. Runnable and self-contained:
import numpy as np
scores = np.array([0.95,0.90,0.80,0.70,0.55,0.45,0.40,0.30,0.20,0.10])
y = np.array([1, 1, 0, 1, 0, 1, 0, 0, 0, 0 ])
P = y.sum(); tp = fp = 0; rows = []
for i in np.argsort(-scores):
tp += y[i]; fp += 1 - y[i]
prec = tp/(tp+fp); rec = tp/P
f1 = 0.0 if prec+rec == 0 else 2*prec*rec/(prec+rec)
rows.append((scores[i], prec, rec, f1))
best = max(rows, key=lambda r: r[3])
print("best F1", round(best[3],3), "at thr", best[0])Best F1 = 0.800 at threshold 0.45
Why: At t = 0.45 precision = 0.667 and recall = 1.000, giving the harmonic mean 0.800 — the highest of any threshold. Lowering t further only adds false positives and drops F1.
Need precision ≥ 0.75? Best recall is 0.75 at t = 0.70
Why: Scanning the table for rows with precision ≥ 0.75, the one with the most recall is t = 0.70 (precision 0.75, recall 0.75). The right threshold depends on which error costs you more.
| threshold | precision | recall | F1 |
|---|---|---|---|
| 0.90 | 1.000 | 0.500 | 0.667 |
| 0.70 | 0.750 | 0.750 | 0.750 |
| 0.45 (best F1) | 0.667 | 1.000 | 0.800 |
| 0.40 | 0.571 | 1.000 | 0.727 |
| 0.20 | 0.444 | 1.000 | 0.615 |
Pattern
Step through it
Step through Pick the operating threshold one row at a time. What is driving the change, and what would the row after the last one be?
Intuition
FPR's denominator is FP + TN — the count of all negatives. When 98% of the data is negative, that denominator is enormous, so even hundreds of false positives barely move FPR off zero.
So the ROC curve stays pinned to the left edge and the AUC looks great — while in reality the model floods you with false alarms for every true positive it finds.
Precision's denominator is TP + FP — only the flagged examples. It has no giant TN to hide behind, so it feels every false positive. That is why PR is the honest metric under imbalance.
Fill the middle
Fill in the blanks
From ROC-AUC vs PR-AUC under 2% imbalance — one line has had its right-hand side removed. Put it back.
import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.metrics import roc_auc_score, average_precision_score
X, y = make_classification(n_samples=4000, weights=[0.98, 0.02],
n_informative=5, random_state=1)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.5, random_state=0)
p = LogisticRegression(max_iter=2000).fit(Xtr, ytr).predict_proba(Xte)[:, 1]
print(round(yte.mean(), 4))
print(round(roc_auc_score(yte, p), 3), round(average_precision_score(yte, p), 3))
Why: n_informative is what everything below it consumes, so the wrong expression here fails later and somewhere else. A middling-but-usable-sounding number.
Worked example
Train logistic regression on a synthetic dataset that is only 2% positive, then compare the two summary numbers on held-out data. Fully runnable and seeded:
import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.metrics import roc_auc_score, average_precision_score
X, y = make_classification(n_samples=4000, weights=[0.98, 0.02],
n_informative=5, random_state=1)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.5, random_state=0)
p = LogisticRegression(max_iter=2000).fit(Xtr, ytr).predict_proba(Xte)[:, 1]
print(round(yte.mean(), 4))
print(round(roc_auc_score(yte, p), 3), round(average_precision_score(yte, p), 3))ROC-AUC = 0.736 looks respectable...
Why: A middling-but-usable-sounding number. If this were your only metric you might ship the model.
...but PR-AUC = 0.111 exposes a near-useless detector
Why: The random-guess baseline for AP is the prevalence, 0.022. So 0.111 is only ~5× better than chance — the model barely finds the rare positives. ROC hid that entirely.
| metric | value (verified) | story |
|---|---|---|
| test positive rate | 0.0225 | heavy imbalance |
| ROC-AUC | 0.736 | deceptively rosy |
| PR-AUC (AP) | 0.111 | the honest picture |
| AP baseline (prevalence) | 0.022 | random-guess floor |
Comparison
Comparison matrix
From ROC-AUC vs PR-AUC under 2% imbalance: refill the story column from what you know. The rest of the table is as it appeared.
| metric | value (verified) | story |
|---|---|---|
| test positive rate | 0.0225 | heavy imbalance |
| ROC-AUC | 0.736 | deceptively rosy |
| PR-AUC (AP) | 0.111 | the honest picture |
| AP baseline (prevalence) | 0.022 | random-guess floor |
Anomaly
Predict first
A student writes this, and it looks reasonable:
The rare-event model scores ROC-AUC 0.74, comfortably above 0.5.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: ROC-AUC is dominated by the millions of easy negatives, whose huge count keeps FPR tiny.
On heavy imbalance, report PR-AUC / average precision and compare it to the prevalence baseline.
Why: ROC-AUC is dominated by the millions of easy negatives, whose huge count keeps FPR tiny. The 0.74 is measuring 'can it rank the easy negatives low', not 'can it find the positives'.
Trap
The rare-event model scores ROC-AUC 0.74, comfortably above 0.5.
Ship it — 0.74 AUC means a decent detector
Why: ROC-AUC is dominated by the millions of easy negatives, whose huge count keeps FPR tiny. The 0.74 is measuring 'can it rank the easy negatives low', not 'can it find the positives'.
On heavy imbalance, report PR-AUC / average precision and compare it to the prevalence baseline.
PR-AUC = 0.111 vs 0.022 baseline → weak; do not ship
Why: PR ignores true negatives, so it reflects the real difficulty of the positive class. 0.111 against a 0.022 floor says the detector is only marginally better than guessing.
Break the constraint
Discussion prompt
The rule this trap just fixed:
PR ignores true negatives, so it reflects the real difficulty of the positive class. 0.111 against a 0.022 floor says the detector is only marginally better than guessing.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
ROC-AUC is dominated by the millions of easy negatives, whose huge count keeps FPR tiny. The 0.74 is measuring 'can it rank the easy negatives low', not 'can it find the positives'.
Ranking
Put in order
These are the steps of The curve-selection recipe, scrambled. Put them back in order before the next slide shows you.
P(random positive > random negative) — threshold-free ranking quality1.0, random 0.5, worse < 0.5 (flip it → 1 − AUC)Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
P(random positive > random negative) — threshold-free ranking quality1.0, random 0.5, worse < 0.5 (flip it → 1 − AUC)Edge cases
Discussion prompt
The curve-selection recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
P(random positive > random negative) — threshold-free ranking quality1.0, random 0.5, worse < 0.5 (flip it → 1 − AUC)Elimination
Eliminate the wrong options
A classifier has AUC-ROC = 0.3. The best move is:
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: AUC < 0.5 means the model ranks systematically backwards — it has real signal, just inverted. Flipping every decision turns each discordant pair concordant, giving AUC = 1 − 0.3 = 0.7. We verified exactly this: 0.125 flips to 0.875.
Check
Read the AUC before you answer.
Check your understanding
A classifier has AUC-ROC = 0.3. The best move is:
Answer: A
Why: AUC < 0.5 means the model ranks systematically backwards — it has real signal, just inverted. Flipping every decision turns each discordant pair concordant, giving AUC = 1 − 0.3 = 0.7. We verified exactly this: 0.125 flips to 0.875.
Prediction
Predict first
On a 2%-positive dataset a model shows ROC-AUC = 0.74 but PR-AUC = 0.11. Which better summarizes its ability to find the rare positives?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: PR-AUC = 0.11 — ROC-AUC is inflated by the many easy negatives
Why: FPR's denominator FP+TN is dominated by the 98% negatives, so ROC-AUC stays optimistic no matter how many false alarms occur. PR-AUC uses TP+FP (no TN), so it feels every false positive — 0.11 against a 0.022 prevalence baseline is the honest, weak picture.
Check
Think about which denominator hides the false alarms.
Check your understanding
On a 2%-positive dataset a model shows ROC-AUC = 0.74 but PR-AUC = 0.11. Which better summarizes its ability to find the rare positives?
Answer: A
Why: FPR's denominator FP+TN is dominated by the 98% negatives, so ROC-AUC stays optimistic no matter how many false alarms occur. PR-AUC uses TP+FP (no TN), so it feels every false positive — 0.11 against a 0.022 prevalence baseline is the honest, weak picture.
Section
Part 4 of 6 — trust the probabilities
Intuition
AUC and AP only judge ranking — do positives score above negatives. They say nothing about whether a 0.9 really means a 90% chance.
A model can rank perfectly (AUC = 1) yet output 0.9 on cases that are right only 60% of the time. Its scores are ordered correctly but dishonest as probabilities.
Whenever a downstream decision uses the probability itself — expected cost, a betting threshold, medical risk — you need the scores to be calibrated, not just well-ranked.
Concept
A model is calibrated if its confidence matches reality: among all the cases where it predicts 0.8, about 80% should actually be positive.
reliability diagram — A plot of predicted confidence (x) against observed accuracy (y), one point per confidence bin. Perfect calibration lies on the diagonal. Points BELOW the diagonal mean the model is overconfident — its confidence exceeds its accuracy — which is the typical failure of modern deep nets.
Definition probe
Sort into buckets
Every line below is part of the definition of TP / FP / FN / TN or of reliability diagram — one or the other, never both. Put each where it belongs.
Concept
These are two separate axes of a model, and a model can be good on one and bad on the other:
| well-ranked (high AUC) | poorly-ranked (low AUC) | |
|---|---|---|
| calibrated | the goal — trust the number | honest but weak signal |
| miscalibrated | great order, dishonest values | worst case |
Our ×3-overconfident model lives in the top-right box: its ranking is untouched (temperature scaling proves the argmax never moves), yet its probabilities are dishonest. Fixing calibration leaves AUC exactly where it was — the two axes don't interfere.
Concept
ECE turns the reliability gap into one number. Bin the predictions by confidence, and in each bin take the absolute gap between mean confidence and observed accuracy, weighted by how many points fall in the bin:
\[ \text{ECE} = \sum_{b=1}^{B} \frac{n_b}{n}\,\big|\,\text{acc}_b - \text{conf}_b\,\big| \]
accᵦ is the fraction of bin b that is actually positive; confᵦ is the average predicted probability in that bin; nᵦ/n weights each bin by its share of the data. ECE = 0 is perfect calibration.
Pattern
Predict first
The table runs: 0.0–0.1 | 1285 | 0.023 | 0.181 | 0.158 · 0.4–0.5 | 154 | 0.450 | 0.396 | 0.054 · 0.7–0.8 | 189 | 0.749 | 0.571 | 0.177 · 0.8–0.9 | 268 | 0.855 | 0.649 | 0.205 · 0.9–1.0 | 1198 | 0.976 | 0.841 | 0.136
In ECE bin by bin on an overconfident model, given the rows so far: what is the next one — the row where bin is weighted ECE?
Correct: weighted ECE | — | — | — | 0.1448
| bin | n | conf | acc | gap |
|---|---|---|---|---|
| 0.0–0.1 | 1285 | 0.023 | 0.181 | 0.158 |
| 0.4–0.5 | 154 | 0.450 | 0.396 | 0.054 |
| 0.7–0.8 | 189 | 0.749 | 0.571 | 0.177 |
| 0.8–0.9 | 268 | 0.855 | 0.649 | 0.205 |
| 0.9–1.0 | 1198 | 0.976 | 0.841 | 0.136 |
| weighted ECE | — | — | — | 0.1448 |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. The 0.9–1.0 bin holds 1198 predictions at mean confidence 0.976 but only 84.1% are positive — a 0.136 gap.
Worked example
Make the abstraction concrete. Take the ×3-overconfident model, split its predictions into 10 confidence bins, and print each bin's count, mean confidence, and observed accuracy. Self-contained and seeded:
import numpy as np
np.random.seed(0)
sig = lambda z: 1.0/(1.0+np.exp(-z))
z = np.random.randn(4000)*1.5
labels = (np.random.rand(4000) < sig(z)).astype(float)
probs = sig(z*3.0) # overconfident predictions
edges = np.linspace(0, 1, 11); ece = 0.0
for i in range(10):
m = (probs > edges[i]) & (probs <= edges[i+1])
if m.sum():
conf, acc = probs[m].mean(), labels[m].mean()
ece += m.mean()*abs(acc - conf)
print(f"{edges[i]:.1f}-{edges[i+1]:.1f} n={m.sum():4d} conf={conf:.3f} acc={acc:.3f}")
print("ECE =", round(ece, 4))The extreme bins are where overconfidence bites
Why: The 0.9–1.0 bin holds 1198 predictions at mean confidence 0.976 but only 84.1% are positive — a 0.136 gap. The 0.0–0.1 bin is symmetric: confidence 0.023 but 18.1% positive. The ×3 pushed probabilities to the extremes, past what the data supports.
| bin | n | conf | acc | gap |
|---|---|---|---|---|
| 0.0–0.1 | 1285 | 0.023 | 0.181 | 0.158 |
| 0.4–0.5 | 154 | 0.450 | 0.396 | 0.054 |
| 0.7–0.8 | 189 | 0.749 | 0.571 | 0.177 |
| 0.8–0.9 | 268 | 0.855 | 0.649 | 0.205 |
| 0.9–1.0 | 1198 | 0.976 | 0.841 | 0.136 |
| weighted ECE | — | — | — | 0.1448 |
Pattern
Step through it
Step through ECE bin by bin on an overconfident model one row at a time. What is driving the change, and what would the row after the last one be?
Concept
Calibration is repaired after training, without touching the model's ranking. Three standard tools:
σ(a·s + b) to map scores to calibrated probabilities (two parameters).T > 0 before the sigmoid/softmax. One parameter, cannot change the ranking.Temperature scaling is the workhorse for neural nets: one number, provably preserves accuracy, and directly attacks overconfidence. We derive and sweep it next.
Concept
A model's raw pre-sigmoid output is its logit z. The probability is p = σ(z) = 1/(1 + e^{−z}). Temperature scaling divides the logit by T before the sigmoid:
\[ p_T = \sigma\!\left(\frac{z}{T}\right) = \frac{1}{1 + e^{-z/T}} \]
T > 1 shrinks the logits toward 0, pulling probabilities toward 0.5 — it softens an overconfident model. T < 1 sharpens them. T = 1 leaves the model unchanged.
Explain it
Discussion prompt
Explain The temperature mechanism to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
A model's raw pre-sigmoid output is its logit z. The probability is p = σ(z) = 1/(1 + e^{−z}). Temperature scaling divides the logit by T before the sigmoid:
Missing information
Discussion prompt
Build a well-calibrated latent (label ~ Bernoulli(σ(z))), then make the model overconfident by reporting 3z instead of z. Sweep T and watch ECE. Fully self-contained and seeded:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Multiplying the honest logit z by 3 pushes σ toward 0 and 1 — the model reports 0.99 where reality is ~0.9. That mismatch is exactly what ECE measures.
Worked example
Build a well-calibrated latent (label ~ Bernoulli(σ(z))), then make the model overconfident by reporting 3z instead of z. Sweep T and watch ECE. Fully self-contained and seeded:
import numpy as np
np.random.seed(0)
sig = lambda z: 1.0/(1.0+np.exp(-z))
def ece(probs, labels, bins=10):
e = 0.0; edges = np.linspace(0, 1, bins+1)
for i in range(bins):
m = (probs > edges[i]) & (probs <= edges[i+1])
if m.sum():
e += m.mean()*abs(probs[m].mean() - labels[m].mean())
return e
n = 4000
z = np.random.randn(n)*1.5
labels = (np.random.rand(n) < sig(z)).astype(float)
logits = z*3.0 # overconfident model
for T in [1, 2, 3, 4, 5]:
print(T, round(ece(sig(logits/T), labels), 4))logits = z*3 makes the model overconfident (×3 sharper)
Why: Multiplying the honest logit z by 3 pushes σ toward 0 and 1 — the model reports 0.99 where reality is ~0.9. That mismatch is exactly what ECE measures.
Dividing by T = 3 exactly undoes the ×3 → ECE minimized
Why: σ(logits/T) = σ(3z/3) = σ(z), the honest probability. So T = 3 should give the lowest ECE — and it does, dropping ECE from 0.1448 to 0.0176.
| T | ECE (verified) | note |
|---|---|---|
| 1 (raw) | 0.1448 | badly overconfident |
| 2 | 0.0583 | improving |
| 3 (best) | 0.0176 | undoes the ×3 |
| 4 | 0.0524 | now underconfident |
| 5 | 0.0830 | over-softened |
Error analysis
Annotate
Walk the callouts on Temperature scaling drives ECE down. Each one is a place this is easy to get subtly wrong.
Concept
ECE falls from 0.1448 at T = 1 to its minimum 0.0176 at T = 3, then climbs back to 0.0830 at T = 5. The curve is a valley: too little softening leaves overconfidence, too much creates underconfidence.
The minimum sits exactly at T = 3 because that is the factor by which we inflated the logits. In practice you don't know the true factor — you fit T by minimizing ECE (or NLL) on a held-out set.
Estimation
Predict first
Crucial property: dividing all logits by the same positive T preserves their order, so the argmax — the actual prediction — never moves. Accuracy is untouched:
Commit before you compute: what does Temperature never changes the prediction come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Ranking identical, class predictions identical
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. z/T is a monotone rescaling, so argsort is unchanged and every prediction stays on the same side of 0.5.
Worked example
Crucial property: dividing all logits by the same positive T preserves their order, so the argmax — the actual prediction — never moves. Accuracy is untouched:
import numpy as np
sig = lambda z: 1.0/(1.0+np.exp(-z))
logits = np.array([2.0, -1.0, 0.5, 3.0, -2.5])
p1 = sig(logits/1.0)
p3 = sig(logits/3.0)
print(np.round(p1, 3))
print(np.round(p3, 3))
print((np.argsort(-p1) == np.argsort(-p3)).all())
print((p1 > 0.5).astype(int), (p3 > 0.5).astype(int))Ranking identical, class predictions identical
Why: z/T is a monotone rescaling, so argsort is unchanged and every prediction stays on the same side of 0.5. Softening moves probabilities toward 0.5 but never across the decision boundary here.
| logit z | p at T=1 | p at T=3 | pred (both) |
|---|---|---|---|
| 2.0 | 0.881 | 0.661 | 1 |
| -1.0 | 0.269 | 0.417 | 0 |
| 0.5 | 0.622 | 0.542 | 1 |
| 3.0 | 0.953 | 0.731 | 1 |
| -2.5 | 0.076 | 0.303 | 0 |
Discrimination
Sort into buckets
Sort these by pred (both), from memory, without looking back at Temperature never changes the prediction. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Anomaly
Predict first
A student writes this, and it looks reasonable:
The model outputs 0.99, so this case is almost certainly a positive — trust it.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Confidence and accuracy are different axes.
First measure ECE; if the model is overconfident, recalibrate, then read the probability.
Why: Confidence and accuracy are different axes. An overconfident model (our ECE 0.1448 case) prints 0.99 on plenty of cases it gets wrong. The number is not a probability until the model is calibrated.
Trap
The model outputs 0.99, so this case is almost certainly a positive — trust it.
Treat 0.99 as ~99% reliable
Why: Confidence and accuracy are different axes. An overconfident model (our ECE 0.1448 case) prints 0.99 on plenty of cases it gets wrong. The number is not a probability until the model is calibrated.
First measure ECE; if the model is overconfident, recalibrate, then read the probability.
Fit T, apply σ(z/T), THEN trust 0.99
Why: Only a calibrated 0.99 means ~99% correct. Temperature scaling restores that meaning (ECE 0.1448 → 0.0176) without changing any prediction — same argmax, honest probability.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
s and a true label y (1 = positive). Sorted high score to low:; Fix a threshold. Every example falls into exactly one of four boxes, comparing what the model predicted against the truth.; The two axes of the ROC curve, plus precision, are ratios of those four counts. Learn which denominator each one uses — that choice is the whole story of imbalance.0.74, comfortably above 0.5.Section
Part 5 of 6 — temperature elsewhere
Concept
The exact same T knob controls text generation. Before the softmax over next-token logits, dividing by T reshapes the sampling distribution — same mechanism, different goal.
\[ \Pr(\text{token}_k) = \frac{e^{z_k / T}}{\sum_j e^{z_j / T}} \]
T < 1 sharpens toward the top token (more deterministic, 'greedy'); T > 1 flattens toward uniform (more diverse, more random). T → 0 is pure argmax; T → ∞ is a coin flip over the vocabulary.
Analogy
Discussion prompt
Explain Temperature in LLM sampling by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
The exact same T knob controls text generation. Before the softmax over next-token logits, dividing by T reshapes the sampling distribution — same mechanism, different goal.
Intuition
In calibration, you fit T to a fixed value that makes probabilities honest — you want the number to be trustworthy.
In sampling, you choose T as a creativity dial — low for factual answers, high for brainstorming. Nothing is being calibrated; you're tuning randomness.
Recognizing that σ(z/T) and softmax(z/T) are the same operation is the kind of connection USAAIO loves to test — calibration (this lesson) and softmax temperature (Lesson 17) are one idea.
Counterexample
Discussion prompt
In calibration, you fit T to a fixed value that makes probabilities honest — you want the number to be trustworthy.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
In sampling, you choose T as a creativity dial — low for factual answers, high for brainstorming. Nothing is being calibrated; you're tuning randomness.
Prediction
Predict first
Applying temperature scaling with T > 1 to an overconfident model's logits:
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Softens the probabilities toward 0.5, lowering ECE, with the argmax unchanged
Why: Dividing logits by T > 1 shrinks them toward 0, pulling probabilities toward 0.5 and reducing overconfidence — we saw ECE drop 0.1448 → 0.0176 at T = 3. Because z/T is monotone, the argmax and every 0.5-threshold prediction are unchanged, so accuracy is preserved.
Check
T > 1 applied to logits — think about the argmax.
Check your understanding
Applying temperature scaling with T > 1 to an overconfident model's logits:
Answer: A
Why: Dividing logits by T > 1 shrinks them toward 0, pulling probabilities toward 0.5 and reducing overconfidence — we saw ECE drop 0.1448 → 0.0176 at T = 3. Because z/T is monotone, the argmax and every 0.5-threshold prediction are unchanged, so accuracy is preserved.
Elimination
Eliminate the wrong options
On our data at threshold t = 0.5 the ROC point is (FPR, TPR) = (0.333, 0.75). Which statement is correct?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: TPR = 0.75 = TP/(TP+FN) is the recall — the share of real positives caught. FPR = 0.333 = FP/(FP+TN) is the share of real negatives wrongly flagged. Each ROC point is one threshold's (FPR, TPR) trade-off.
Check
Recall what a single ROC point encodes.
Check your understanding
On our data at threshold t = 0.5 the ROC point is (FPR, TPR) = (0.333, 0.75). Which statement is correct?
Answer: A
Why: TPR = 0.75 = TP/(TP+FN) is the recall — the share of real positives caught. FPR = 0.333 = FP/(FP+TN) is the share of real negatives wrongly flagged. Each ROC point is one threshold's (FPR, TPR) trade-off.
Section
Part 6 of 6 — the project
Concept
Assemble the whole lesson yourself: build an ROC curve and AUC by hand, contrast ROC-AUC with PR-AUC on imbalanced data, then recalibrate an overconfident model with temperature scaling. You derived every piece — now wire them together.
| # | requirement | tool |
|---|---|---|
| 1 | Sweep thresholds → TPR/FPR → AUC | np.argsort, np.trapz |
| 2 | ROC-AUC vs PR-AUC on 2%-positive data | roc_auc_score, average_precision_score |
| 3 | Temperature-scale to the ECE minimum | σ(logits/T), an ece() function |
Build rules: type every line yourself, sort by descending score for the ROC sweep, and confirm your manual AUC matches sklearn before moving on. When a shape error hits, print .shape — read the error, don't delete it.
Worked example
Your turn: sweep thresholds to accumulate TPR/FPR, integrate for the AUC, and check it against sklearn. Predict out loud: will they match exactly?
Hint: walk np.argsort(-scores) (descending). Each example adds to tp if positive or fp if negative; append tp/P and fp/N; finish with np.trapz(tpr, fpr).
import numpy as np
from sklearn.metrics import roc_auc_score
scores = np.array([0.95,0.90,0.80,0.70,0.55,0.45,0.40,0.30,0.20,0.10])
y = np.array([1, 1, 0, 1, 0, 1, 0, 0, 0, 0 ])
P, N = y.sum(), (1-y).sum()
tp = fp = 0; tpr = [0.0]; fpr = [0.0]
for i in np.argsort(-scores):
tp += y[i]; fp += 1 - y[i]
tpr.append(tp/P); fpr.append(fp/N)
print(round(np.trapz(tpr, fpr), 4), round(roc_auc_score(y, scores), 4))| source | AUC (expected) |
|---|---|
| your manual trapezoid | 0.8750 |
| sklearn roc_auc_score | 0.8750 |
Worked example
Your turn: train logistic regression on 2%-positive data and print both AUCs. Predict which one collapses.
Hint: make_classification(weights=[0.98, 0.02], ...), split, predict_proba(...)[:, 1], then roc_auc_score vs average_precision_score. Keep the same random_states to reproduce the numbers.
import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.metrics import roc_auc_score, average_precision_score
X, y = make_classification(n_samples=4000, weights=[0.98, 0.02],
n_informative=5, random_state=1)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.5, random_state=0)
p = LogisticRegression(max_iter=2000).fit(Xtr, ytr).predict_proba(Xte)[:, 1]
print(round(roc_auc_score(yte, p), 3), round(average_precision_score(yte, p), 3))| metric | value (expected) |
|---|---|
| ROC-AUC | 0.736 |
| PR-AUC (AP) | 0.111 |
| prevalence baseline | 0.022 |
Worked example
Your turn: write an ece() function, build an overconfident model (×3 logits), and sweep T. Predict the best T before you run it.
Hint: ece bins predictions into 10 bins and averages |conf − acc| weighted by bin size. The overconfidence factor is 3, so the ECE minimum should land at T ≈ 3.
import numpy as np
np.random.seed(0)
sig = lambda z: 1.0/(1.0+np.exp(-z))
def ece(probs, labels, bins=10):
e = 0.0; edges = np.linspace(0, 1, bins+1)
for i in range(bins):
m = (probs > edges[i]) & (probs <= edges[i+1])
if m.sum():
e += m.mean()*abs(probs[m].mean() - labels[m].mean())
return e
z = np.random.randn(4000)*1.5
labels = (np.random.rand(4000) < sig(z)).astype(float)
logits = z*3.0
for T in [1, 2, 3, 4, 5]:
print(T, round(ece(sig(logits/T), labels), 4))| T | ECE (expected) |
|---|---|
| 1 (raw) | 0.1448 |
| 2 | 0.0583 |
| 3 (best) | 0.0176 |
| 4 | 0.0524 |
| 5 | 0.0830 |
Trade off
Comparison matrix
From Milestone 3 — temperature scaling: every row here is a choice with a cost. Fill the ECE (expected) column, then say which row you would actually pick and what you give up for it.
| T | ECE (expected) |
|---|---|
| 1 (raw) | 0.1448 |
| 2 | 0.0583 |
| 3 (best) | 0.0176 |
| 4 | 0.0524 |
| 5 | 0.0830 |
Concept
import numpy as np
from sklearn.metrics import roc_auc_score, average_precision_score
# 1. ROC-AUC from scratch
scores = np.array([0.95,0.90,0.80,0.70,0.55,0.45,0.40,0.30,0.20,0.10])
y = np.array([1, 1, 0, 1, 0, 1, 0, 0, 0, 0 ])
P, N = y.sum(), (1-y).sum(); tp = fp = 0; T = [0.0]; F = [0.0]
for i in np.argsort(-scores):
tp += y[i]; fp += 1-y[i]; T.append(tp/P); F.append(fp/N)
auc = np.trapz(T, F)
print('ROC-AUC:', round(auc, 4), '== sklearn', bool(np.isclose(auc, roc_auc_score(y, scores))))
print('AP:', round(average_precision_score(y, scores), 4))
# 2. imbalance: ROC-AUC 0.736 rosy vs PR-AUC 0.111 honest
# 3. temperature scaling: ECE 0.1448 (T=1) -> 0.0176 (T=3)| printed line | value (verified) |
|---|---|
| ROC-AUC: ... == sklearn | 0.875, True |
| AP: | 0.8542 |
| imbalance ROC / PR | 0.736 / 0.111 |
| ECE T=1 → T=3 | 0.1448 → 0.0176 |
If your ROC matches sklearn, PR-AUC exposes the imbalance, and temperature scaling minimizes ECE at T = 3 — you can evaluate a model honestly and trust its probabilities.
Comparison
Comparison matrix
From The full program: refill the value (verified) column from what you know. The rest of the table is as it appeared.
| printed line | value (verified) |
|---|---|
| ROC-AUC: ... == sklearn | 0.875, True |
| AP: | 0.8542 |
| imbalance ROC / PR | 0.736 / 0.111 |
| ECE T=1 → T=3 | 0.1448 → 0.0176 |
Concept
Out loud, slides closed: explain (1) what one ROC point represents and why the curve steps up on positives and right on negatives, (2) why ROC-AUC flatters an imbalanced model while PR-AUC doesn't, and (3) what temperature scaling changes — and what it provably leaves alone.
Stretch: reproduce AUC = 0.875 by all three routes (trapezoid, sklearn, concordant-pair count), then plot a reliability diagram before and after temperature scaling. Calibration returns in LLM sampling temperature (Week 37) and uncertainty estimation.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — The confusion matrix · The ROC curve · AUC, three ways · The PR curve & imbalance · Calibration · One knob, two jobs. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
TPR = 0.75, FPR = 0.333, precision = 0.60 at a thresholdAUC = 0.875 three independent waysP(random positive > random negative) and flip a worse-than-random model (AUC → 1 − AUC)0.736 flatters, PR-AUC 0.111 tells the truthECE from 0.1448 to 0.0176 with temperature scaling at T = 3| idea | the one thing to remember |
|---|---|
| ROC axes | FPR on x, TPR on y — the good corner is top-left |
| AUC | P(random positive outranks random negative) |
| imbalance | report PR-AUC vs the prevalence baseline, not ROC-AUC |
| calibration | confidence ≠ accuracy; measure ECE |
| temperature | σ(z/T): T>1 softens, same argmax, one knob for LLM sampling too |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.