Lesson 35: ROC, PR Curves & Calibration

USAAIO Lesson 35, from Week 12, fully worked. It defines the confusion matrix and every rate derived from it, then builds the ROC curve threshold by threshold on one running 10-example dataset. AUC is computed three independent ways - by the trapezoid rule, by sklearn, and by counting concordant pairs - and all three land on 0.875. The precision-recall curve and average precision are derived the same way, and the lesson shows why ROC-AUC misleads under 2% imbalance while PR-AUC, at 0.111, tells the truth. It then defines calibration and the Expected Calibration Error bin by bin, and sweeps temperature scaling to its ECE minimum at T=3. Every snippet runs standalone in a fresh interpreter, and every number was produced by real execution. The lesson runs to 61 slides.

Subject: Machine Learning · 110 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. ROC, PR Curves & Calibration

Title

USAAIO · Lesson 35 · Week 12 (Evaluation)

Threshold-free evaluation, and probabilities you can trust. We build the ROC curve one threshold at a time, compute its AUC three independent ways, expose why it lies under imbalance, then fix an overconfident model with temperature scaling — every number produced by real execution.

2. By the end of this lesson you can

Objectives

  1. Read a confusion matrix at a threshold and compute TPR, FPR, and precision from it by hand
  2. Build the ROC curve point by point by sweeping the threshold, and integrate it for the AUC
  3. Prove AUC = P(random positive scores above random negative) by counting concordant pairs — the same 0.875 three ways
  4. Build the precision-recall curve and average precision, and say exactly why ROC-AUC flatters an imbalanced model while PR-AUC doesn't
  5. Define calibration and the Expected Calibration Error, and drive ECE to its minimum with temperature scaling T

3. What survived from Concentration Inequalities & Generalization?

Warm-up

Discussion prompt

Before we open Lesson 35: ROC, PR Curves & Calibration: without looking back, what was the main idea of Concentration Inequalities & Generalization, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

Markov, Chebyshev, and Hoeffding inequalities, VC dimension as a complexity measure, the VC generalization bound, and why overparameterized networks still generalize (double descent). Verify the inequalities empirically and compute a VC generalization bound.

4. The confusion matrix

Section

Part 1 of 6 — the foundation

5. One threshold turns scores into decisions

Concept

A classifier outputs a score s ∈ [0,1] — how positive it thinks each example is. To make an actual decision you pick a threshold t and call everything with s ≥ t positive.

Change t and the decisions change. Evaluation metrics like accuracy depend on that one arbitrary knob — the whole point of ROC and PR curves is to escape it by sweeping t across all its values at once.

6. Break it if you can: One threshold turns scores into decisions

Counterexample

Discussion prompt

A classifier outputs a score s ∈ [0,1] — how positive it thinks each example is. To make an actual decision you pick a threshold t and call everything with s ≥ t positive.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

7. The data: ten scored examples

Concept

One running example for the whole lesson. Ten examples, each with a model score s and a true label y (1 = positive). Sorted high score to low:

#score strue y
10.951
20.901
30.800
40.701
50.550
60.451
70.400
80.300
90.200
100.100

There are P = 4 positives and N = 6 negatives. Notice the ranking is good but not perfect: example 3 (a negative) scores 0.80, above example 4 (a positive) at 0.70. That one swap is what keeps the AUC below 1.

8. Watch it run: The data: ten scored examples

Pattern

Step through it

Step through The data: ten scored examples one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: # is 1
  2. Step 2: # is 2
  3. Step 3: # is 3
  4. Step 4: # is 4
  5. Step 5: # is 5
  6. Step 6: # is 6
  7. Step 7: # is 7
  8. Step 8: # is 8
  9. Step 9: # is 9
  10. Step 10: # is 10

9. The four confusion-matrix counts

Concept

Fix a threshold. Every example falls into exactly one of four boxes, comparing what the model predicted against the truth.

predicted +predicted −
actually +TP (true positive)FN (false negative)
actually −FP (false positive)TN (true negative)

TP / FP / FN / TN — TP: called positive, was positive (a hit). FP: called positive, was negative (a false alarm). FN: called negative, was positive (a miss). TN: called negative, was negative (a correct rejection). Every metric today is built from just these four numbers.

10. Fill in: predicted − for The four confusion-matrix counts

Comparison

Comparison matrix

From The four confusion-matrix counts: refill the predicted − column from what you know. The rest of the table is as it appeared.

predicted +predicted −
actually +TP (true positive)FN (false negative)
actually −FP (false positive)TN (true negative)

11. Guess the shape of the answer: Count the four boxes at t = 0.5

Estimation

Predict first

Threshold t = 0.5: predict positive when s ≥ 0.5. That flags examples 1–5 (scores 0.95, 0.90, 0.80, 0.70, 0.55). Compare each flag to its true label:

Commit before you compute: what does Count the four boxes at t = 0.5 come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Predicted negative: #6(y=1)✗ #7(y=0)✓ #8(y=0)✓ #9(y=0)✓ #10(y=0)✓

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Example 6 is a positive we missed (FN = 1); the other four are correctly rejected negatives (TN = 4).

12. Count the four boxes at t = 0.5

Worked example

Threshold t = 0.5: predict positive when s ≥ 0.5. That flags examples 1–5 (scores 0.95, 0.90, 0.80, 0.70, 0.55). Compare each flag to its true label:

Predicted positive: #1(y=1)✓ #2(y=1)✓ #3(y=0)✗ #4(y=1)✓ #5(y=0)✗

Why: Three of the five flagged examples are truly positive (TP = 3); two are truly negative false alarms (FP = 2).

Predicted negative: #6(y=1)✗ #7(y=0)✓ #8(y=0)✓ #9(y=0)✓ #10(y=0)✓

Why: Example 6 is a positive we missed (FN = 1); the other four are correctly rejected negatives (TN = 4).

\[ TP = 3, \quad FP = 2, \quad FN = 1, \quad TN = 4 \qquad (TP+FP+FN+TN = 10\;\checkmark) \]

13. Decode the notation: Count the four boxes at t = 0.5

Notation

Annotate

From Count the four boxes at t = 0.5 — read this one piece at a time. What is each part doing?

On: \( TP = 3, \quad FP = 2, \quad FN = 1, \quad TN = 4 \qquad (TP+FP+FN+TN = 10\;\checkmark) \)

  • Three of the five flagged examples are truly positive (TP = 3); two are truly negative false alarms (FP = 2).
  • Example 6 is a positive we missed (FN = 1); the other four are correctly rejected negatives (TN = 4).

14. Three rates read off the boxes

Concept

The two axes of the ROC curve, plus precision, are ratios of those four counts. Learn which denominator each one uses — that choice is the whole story of imbalance.

\[ \text{TPR (recall)} = \frac{TP}{TP+FN}, \qquad \text{FPR} = \frac{FP}{FP+TN}, \qquad \text{precision} = \frac{TP}{TP+FP} \]

TPR (recall) is the fraction of real positives caught. FPR is the fraction of real negatives falsely flagged. Precision is the fraction of flags that were right. TPR and FPR split by the true class; precision splits by the predicted class.

15. By analogy: Three rates read off the boxes

Analogy

Discussion prompt

Explain Three rates read off the boxes by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

The two axes of the ROC curve, plus precision, are ratios of those four counts. Learn which denominator each one uses — that choice is the whole story of imbalance.

16. What has to be given first: The rates at t = 0.5, verified in code

Missing information

Discussion prompt

Plug TP=3, FP=2, FN=1, TN=4 into the three formulas, then confirm with a runnable snippet that rebuilds the whole thing from the raw scores:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

The threshold t = 0.5 gives a single (FPR, TPR) = (0.333, 0.75) point. The ROC curve is what you get by sweeping t and collecting all such points.

17. The rates at t = 0.5, verified in code

Worked example

Plug TP=3, FP=2, FN=1, TN=4 into the three formulas, then confirm with a runnable snippet that rebuilds the whole thing from the raw scores:

import numpy as np
scores = np.array([0.95,0.90,0.80,0.70,0.55,0.45,0.40,0.30,0.20,0.10])
y      = np.array([1,   1,   0,   1,   0,   1,   0,   0,   0,   0  ])
thr = 0.5
pred = (scores >= thr).astype(int)
TP = int(((pred==1)&(y==1)).sum())
FP = int(((pred==1)&(y==0)).sum())
FN = int(((pred==0)&(y==1)).sum())
TN = int(((pred==0)&(y==0)).sum())
print("TP FP FN TN =", TP, FP, FN, TN)
print("TPR =", TP/(TP+FN))
print("FPR =", FP/(FP+TN))
print("precision =", TP/(TP+FP))

TPR = 3/4 = 0.75, FPR = 2/6 = 0.333, precision = 3/5 = 0.60

Why: The threshold t = 0.5 gives a single (FPR, TPR) = (0.333, 0.75) point. The ROC curve is what you get by sweeping t and collecting all such points.

quantityformulavalue (verified)
TP, FP, FN, TNcount boxes3, 2, 1, 4
TPR = TP/(TP+FN)3/(3+1)0.750
FPR = FP/(FP+TN)2/(2+4)0.333
precision = TP/(TP+FP)3/(3+2)0.600

18. What each one costs: The rates at t = 0.5, verified in code

Trade off

Comparison matrix

From The rates at t = 0.5, verified in code: every row here is a choice with a cost. Fill the value (verified) column, then say which row you would actually pick and what you give up for it.

quantityformulavalue (verified)
TP, FP, FN, TNcount boxes3, 2, 1, 4
TPR = TP/(TP+FN)3/(3+1)0.750
FPR = FP/(FP+TN)2/(2+4)0.333
precision = TP/(TP+FP)3/(3+2)0.600

19. The ROC curve

Section

Part 2 of 6 — sweep every threshold

20. Why one point is not enough

Intuition

The (0.333, 0.75) point above depended entirely on our arbitrary t = 0.5. A different threshold gives a different trade-off — lower t catches more positives (higher TPR) but raises false alarms (higher FPR).

Rather than defend one threshold, plot all of them. Slide t from +∞ (nothing flagged) down to −∞ (everything flagged) and trace the (FPR, TPR) point as it moves.

That trace is the ROC curve — Receiver Operating Characteristic. It shows the model's entire achievable frontier of trade-offs in one picture, independent of any single threshold choice.

21. Teach it back: Why one point is not enough

Explain it

Discussion prompt

Explain Why one point is not enough to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Rather than defend one threshold, plot all of them. Slide t from +∞ (nothing flagged) down to −∞ (everything flagged) and trace the (FPR, TPR) point as it moves.

22. How the curve moves as t drops

Concept

Start with t above every score: nothing is flagged, so TP = FP = 0 and the point sits at the origin (0, 0). Now lower t past one example's score at a time.

Each example you cross adds one flag. If that example is a real positive, TP grows → the point steps up. If it is a real negative, FP grows → the point steps right.

So the ROC curve is a staircase from (0,0) to (1,1): up on a positive, right on a negative. A good model front-loads all its ups (climbs the left edge early) before the rights.

23. Predict the next row: Build the ROC curve from scratch

Pattern

Predict first

The table runs: (start) | — | 0 | 0 | 0.000 | 0.000 · 0.95 | 1 | 1 | 0 | 0.250 | 0.000 · 0.90 | 1 | 2 | 0 | 0.500 | 0.000 · 0.80 | 0 | 2 | 1 | 0.500 | 0.167 · 0.70 | 1 | 3 | 1 | 0.750 | 0.167 · 0.55 | 0 | 3 | 2 | 0.750 | 0.333 · 0.45 | 1 | 4 | 2 | 1.000 | 0.333 · 0.40 | 0 | 4 | 3 | 1.000 | 0.500 · 0.30 | 0 | 4 | 4 | 1.000 | 0.667 · 0.20 | 0 | 4 | 5 | 1.000 | 0.833

In Build the ROC curve from scratch, given the rows so far: what is the next one — the row where crossed score is 0.10?

Correct: 0.10 | 0 | 4 | 6 | 1.000 | 1.000

crossed scorelabeltpfpTPRFPR
(start)—000.0000.000
0.951100.2500.000
0.901200.5000.000
0.800210.5000.167
0.701310.7500.167
0.550320.7500.333
0.451421.0000.333
0.400431.0000.500
0.300441.0000.667
0.200451.0000.833
0.100461.0001.000

Why: The relationship between the columns, not the individual numbers, is what generates the next row. Descending score is exactly the order thresholds cross the examples as t drops.

24. Build the ROC curve from scratch

Worked example

Sort by descending score, walk the examples, and after each one record the running (FPR, TPR). This snippet is complete and self-contained — it prints every point on the curve:

import numpy as np
from sklearn.metrics import roc_auc_score
scores = np.array([0.95,0.90,0.80,0.70,0.55,0.45,0.40,0.30,0.20,0.10])
y      = np.array([1,   1,   0,   1,   0,   1,   0,   0,   0,   0  ])
P, N = y.sum(), (1-y).sum()
tp = fp = 0
tpr = [0.0]; fpr = [0.0]
for i in np.argsort(-scores):
    tp += y[i]; fp += 1 - y[i]
    tpr.append(tp / P); fpr.append(fp / N)
print(round(np.trapz(tpr, fpr), 4), round(roc_auc_score(y, scores), 4))

np.argsort(-scores) walks examples high score → low

Why: Descending score is exactly the order thresholds cross the examples as t drops. Each loop pass lowers the effective threshold past one more example.

tp += y[i] steps up; fp += 1 - y[i] steps right

Why: If y[i]=1 the example is positive so TP increments (up); if y[i]=0 it is negative so FP increments (right). Exactly the staircase rule.

crossed scorelabeltpfpTPRFPR
(start)—000.0000.000
0.951100.2500.000
0.901200.5000.000
0.800210.5000.167
0.701310.7500.167
0.550320.7500.333
0.451421.0000.333
0.400431.0000.500
0.300441.0000.667
0.200451.0000.833
0.100461.0001.000

25. Reading the staircase

Concept

Trace the table: the first two crossings are positives, so the point climbs straight up the left edge to TPR = 0.5 with FPR still 0. Then example 3 (score 0.80, a negative) forces the first step right.

After example 6 (the last positive, score 0.45), TPR hits 1.0 — every positive is caught. The remaining four crossings are all negatives, so the curve runs flat across the top from FPR = 0.333 to 1.0.

Figure (svg): An ROC staircase rising from the bottom-left origin up the left edge to TPR 0.5, stepping up and right through the middle, reaching TPR 1.0 and then running flat across the top to the top-right corner, well above the dashed diagonal.

The curve for our data: two steps up, then a mix, reaching the top by FPR 0.333. Area beneath = AUC.

26. AUC, three ways

Section

Part 2 continued — 0.875 every time

27. AUC: the area under the curve

Concept

One number summarizes the whole staircase: the area under the ROC curve, AUC ∈ [0,1]. A curve that hugs the top-left corner encloses nearly the full unit square (AUC → 1); the random diagonal encloses half (AUC = 0.5).

For a staircase, the area is a sum of rectangles: at each step right (width ΔFPR), the height is the current TPR. That is exactly what the trapezoid rule np.trapz(tpr, fpr) computes.

28. What has to happen first: Way 1 — trapezoid area by hand

Ranking

Put in order

Put the moves of Way 1 — trapezoid area by hand into the order they have to happen.

  1. Right-step at 0.80: TPR = 0.5, width 1/6 → 0.5·(1/6)
  2. Right-steps at 0.55, 0.40, 0.30, 0.20, 0.10
  3. Sum of heights 5.25, ÷ 6 = 0.875

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. First negative crossed; the height of the rectangle is the TPR standing at that moment, 0.5.

29. Way 1 — trapezoid area by hand

Worked example

The curve only advances rightward at the four negative crossings. Sum TPR × ΔFPR over just those steps (each ΔFPR = 1/6). Read the TPR at each right-step from the table:

Right-step at 0.80: TPR = 0.5, width 1/6 → 0.5·(1/6)

Why: First negative crossed; the height of the rectangle is the TPR standing at that moment, 0.5.

Right-steps at 0.55, 0.40, 0.30, 0.20, 0.10

Why: Heights 0.75 (at 0.55) then 1.0, 1.0, 1.0, 1.0 for the last four negatives. With the 0.80 step above, that is all 6 negatives — 6 right-steps, each width 1/6.

\[ \text{AUC} = \tfrac{1}{6}\big(0.5 + 0.75 + 1.0 + 1.0 + 1.0 + 1.0\big) = \tfrac{1}{6}(5.25) = 0.875 \]

Sum of heights 5.25, ÷ 6 = 0.875

Why: Heights are the TPR at each of the 6 negative crossings: 0.5 (at 0.80), 0.75 (at 0.55), then 1.0 four times (at 0.40, 0.30, 0.20, 0.10). This is the exact area under the staircase.

30. Work backwards from the answer: Way 1 — trapezoid area by hand

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Sum of heights 5.25, ÷ 6 = 0.875

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

The curve only advances rightward at the four negative crossings. Sum TPR × ΔFPR over just those steps (each ΔFPR = 1/6). Read the TPR at each right-step from the table:

31. Restore the missing line: Way 2 — np.trapz matches sklearn

Fill the middle

Fill in the blanks

From Way 2 — np.trapz matches sklearn — one line has had its right-hand side removed. Put it back.

import numpy as np
from sklearn.metrics import roc_auc_score
scores = np.array([0.95,0.90,0.80,0.70,0.55,0.45,0.40,0.30,0.20,0.10])
y = np.array([1, 1, 0, 1, 0, 1, 0, 0, 0, 0 ])
P, N = y.sum(), (1-y).sum()
tp = fp = 0; tpr = [0.0]; fpr = [0.0]
for i in np.argsort(-scores):
tp += y[i]; fp += 1 - y[i]
tpr.append(tp / P); fpr.append(fp / N)
auc_manual = np.trapz(tpr, fpr)
print(round(auc_manual, 4), round(roc_auc_score(y, scores), 4))

Why: auc_manual is what everything below it consumes, so the wrong expression here fails later and somewhere else. trapz sums trapezoids under the (fpr, tpr) points.

32. Way 2 — np.trapz matches sklearn

Worked example

The same block from before ends by integrating the curve and printing it beside sklearn's roc_auc_score. Both must agree with our hand value 0.875:

import numpy as np
from sklearn.metrics import roc_auc_score
scores = np.array([0.95,0.90,0.80,0.70,0.55,0.45,0.40,0.30,0.20,0.10])
y      = np.array([1,   1,   0,   1,   0,   1,   0,   0,   0,   0  ])
P, N = y.sum(), (1-y).sum()
tp = fp = 0; tpr = [0.0]; fpr = [0.0]
for i in np.argsort(-scores):
    tp += y[i]; fp += 1 - y[i]
    tpr.append(tp / P); fpr.append(fp / N)
auc_manual = np.trapz(tpr, fpr)
print(round(auc_manual, 4), round(roc_auc_score(y, scores), 4))

np.trapz(tpr, fpr) integrates y=tpr over x=fpr

Why: trapz sums trapezoids under the (fpr, tpr) points. On a staircase the trapezoids are rectangles, giving the same 0.875 as the hand sum.

methodAUC (verified)
hand trapezoid0.8750
np.trapz(tpr, fpr)0.8750
sklearn roc_auc_score0.8750

33. Where does each piece belong: Lesson 35: ROC, PR Curves & Calibration

Sorting

Sort into buckets

These are the pieces of Lesson 35: ROC, PR Curves & Calibration, out of order. Put each one back under the part of the lesson it belongs to.

The confusion matrix
One threshold turns scores into decisions; The data: ten scored examples; The four confusion-matrix counts
The ROC curve
Why one point is not enough; How the curve moves as t drops; Build the ROC curve from scratch
AUC, three ways
AUC: the area under the curve; Way 1 — trapezoid area by hand; Way 2 — np.trapz matches sklearn
s1
The confusion matrix is where Lesson 35: ROC, PR Curves & Calibration puts One threshold turns scores into decisions, The data: ten scored examples, The four confusion-matrix counts. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
The ROC curve is where Lesson 35: ROC, PR Curves & Calibration puts Why one point is not enough, How the curve moves as t drops, Build the ROC curve from scratch. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
AUC, three ways is where Lesson 35: ROC, PR Curves & Calibration puts AUC: the area under the curve, Way 1 — trapezoid area by hand, Way 2 — np.trapz matches sklearn. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

34. AUC is a ranking probability

Intuition

Here is the interpretation worth memorizing: AUC is the probability that a randomly chosen positive scores higher than a randomly chosen negative.

That is why AUC is threshold-free — it never fixes a t, it only asks whether the model ranks positives above negatives. A perfect ranker always does (AUC = 1); a coin flip does it half the time (0.5).

\[ \text{AUC} = \Pr\big(\,s^{+} > s^{-}\,\big) \quad\text{over random }(s^{+},s^{-})\text{ pairs} \]

35. Way 3 — count concordant pairs

Worked example

Make that probability literal: form every (positive, negative) pair, count how many the positive outranks, and divide by the total P·N = 4·6 = 24:

import numpy as np
scores = np.array([0.95,0.90,0.80,0.70,0.55,0.45,0.40,0.30,0.20,0.10])
y      = np.array([1,   1,   0,   1,   0,   1,   0,   0,   0,   0  ])
pos = scores[y==1]; neg = scores[y==0]
concordant = 0.0
for sp in pos:
    for sn in neg:
        if sp > sn:    concordant += 1.0
        elif sp == sn: concordant += 0.5
print(concordant, len(pos)*len(neg))
print(round(concordant/(len(pos)*len(neg)), 4))

Only 3 of the 24 pairs are discordant

Why: A pair is discordant when a negative outranks a positive. Here that happens exactly three times: (pos 0.70 < neg 0.80), (pos 0.45 < neg 0.80), and (pos 0.45 < neg 0.55). Every other pair has the positive on top.

concordant = 24 − 3 = 21 → 21/24 = 0.875

Why: There are no ties, so concordant + discordant = 24. The same 0.875 as the trapezoid and sklearn — the ranking view and the area view are one number.

methodvalue (verified)
positive scores[0.95, 0.90, 0.70, 0.45]
negative scores[0.80, 0.55, 0.40, 0.30, 0.20, 0.10]
concordant pairs21 of 24
AUC = 21/240.8750

36. Inspect it line by line: Way 3 — count concordant pairs

Error analysis

Annotate

Walk the callouts on Way 3 — count concordant pairs. Each one is a place this is easy to get subtly wrong.

  • A pair is discordant when a negative outranks a positive. Here that happens exactly three times: (pos 0.70 < neg 0.80), (pos 0.45 < neg 0.80), and (pos 0.45 < neg 0.55). Every other pair has the positive on top.
  • There are no ties, so concordant + discordant = 24. The same 0.875 as the trapezoid and sklearn — the ranking view and the area view are one number.

37. Three reference curves

Concept

Memorize the three landmarks so you can read any AUC instantly:

classifierROC shapeAUC
perfectup the left edge, across the top1.0
randomthe diagonal0.5
worse than randombelow the diagonal< 0.5

A worse-than-random model still has signal — it just points backwards. Flip its predictions and the AUC becomes 1 − AUC > 0.5. Never discard an AUC of 0.3.

38. Restore the missing line: Flipping a backwards classifier

Fill the middle

Fill in the blanks

From Flipping a backwards classifier — one line has had its right-hand side removed. Put it back.

import numpy as np
from sklearn.metrics import roc_auc_score
scores = np.array([0.95,0.90,0.80,0.70,0.55,0.45,0.40,0.30,0.20,0.10])
y = np.array([0, 0, 1, 0, 1, 0, 1, 1, 1, 1 ])
auc = roc_auc_score(y, scores)
auc_flipped = roc_auc_score(y, 1 - scores)
print(round(auc, 3), round(auc_flipped, 3), round(1 - auc, 3))

Why: auc_flipped is what everything below it consumes, so the wrong expression here fails later and somewhere else. Flipping (1 − scores) reverses every ranking, turning each discordant pair into a concordant one.

39. Flipping a backwards classifier

Worked example

Same scores, but now the labels are inverted so the model ranks positives too low. Its AUC lands below 0.5 — until we flip the scores:

import numpy as np
from sklearn.metrics import roc_auc_score
scores = np.array([0.95,0.90,0.80,0.70,0.55,0.45,0.40,0.30,0.20,0.10])
y      = np.array([0,   0,   1,   0,   1,   0,   1,   1,   1,   1  ])
auc         = roc_auc_score(y, scores)
auc_flipped = roc_auc_score(y, 1 - scores)
print(round(auc, 3), round(auc_flipped, 3), round(1 - auc, 3))

AUC = 0.125 → flip the scores → 0.875

Why: Flipping (1 − scores) reverses every ranking, turning each discordant pair into a concordant one. The flipped AUC is exactly 1 − 0.125 = 0.875.

modelAUC (verified)
as-is (backwards)0.125
scores flipped (1 − s)0.875
1 − AUC (matches flip)0.875

40. Work backwards from the answer: Flipping a backwards classifier

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

AUC = 0.125 → flip the scores → 0.875

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Same scores, but now the labels are inverted so the model ranks positives too low. Its AUC lands below 0.5 — until we flip the scores:

41. Something is wrong here: swapping the ROC axes

Anomaly

Predict first

A student writes this, and it looks reasonable:

ROC plots the two rates; put TPR on the x-axis and FPR on the y-axis, then integrate.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Integrating FPR against TPR measures the area on the OTHER side of the curve.

ROC is FPR on x, TPR on y. Integrate TPR over FPR.

Why: Integrating FPR against TPR measures the area on the OTHER side of the curve. You'd report 0.125 for a genuinely good model and 'conclude' it is worse than random.

42. Trap: swapping the ROC axes

Trap

The trap

ROC plots the two rates; put TPR on the x-axis and FPR on the y-axis, then integrate.

\[ \text{AUC}_{\text{wrong}} = \int \text{FPR}\; d(\text{TPR}) \]

Axes swapped → area = 1 − 0.875 = 0.125

Why: Integrating FPR against TPR measures the area on the OTHER side of the curve. You'd report 0.125 for a genuinely good model and 'conclude' it is worse than random.

The fix

ROC is FPR on x, TPR on y. Integrate TPR over FPR.

\[ \text{AUC} = \int \text{TPR}\; d(\text{FPR}) \]

Correct axes → area = 0.875

Why: np.trapz(tpr, fpr) has y first, x second — tpr integrated over fpr. Mnemonic: the good corner is TOP-LEFT (high TPR, low FPR), so TPR must be the vertical axis.

43. Decode the notation: Trap: swapping the ROC axes

Notation

Annotate

From Trap: swapping the ROC axes — read this one piece at a time. What is each part doing?

On: \( \text{AUC} = \int \text{TPR}\; d(\text{FPR}) \)

  • Integrating FPR against TPR measures the area on the OTHER side of the curve. You'd report 0.125 for a genuinely good model and 'conclude' it is worse than random.
  • np.trapz(tpr, fpr) has y first, x second — tpr integrated over fpr. Mnemonic: the good corner is TOP-LEFT (high TPR, low FPR), so TPR must be the vertical axis.

44. The PR curve & imbalance

Section

Part 3 of 6 — where ROC lies

45. Precision-recall: focus on the positives

Concept

The precision-recall curve plots precision (TP/(TP+FP)) on the y-axis against recall (= TPR = TP/(TP+FN)) on the x-axis, swept over the same thresholds.

Its summary number is the average precision (AP), the PR-curve analogue of AUC — the area under the precision-recall curve.

The key difference from ROC: neither precision nor recall uses TN. PR ignores the true negatives entirely and looks only at the positive class — exactly the class you care about when positives are rare.

46. Guess the shape of the answer: Build the PR curve and AP from scratch

Estimation

Predict first

Same descending sweep, but now track precision and recall, and accumulate AP as Σ precision·Δrecall at each positive crossing:

Commit before you compute: what does Build the PR curve and AP from scratch come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: AP steps only on positives, weighting precision by the recall gained

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Each positive crossing lifts recall by 1/P and contributes precision·(that gain).

47. Build the PR curve and AP from scratch

Worked example

Same descending sweep, but now track precision and recall, and accumulate AP as Σ precision·Δrecall at each positive crossing:

import numpy as np
from sklearn.metrics import average_precision_score
scores = np.array([0.95,0.90,0.80,0.70,0.55,0.45,0.40,0.30,0.20,0.10])
y      = np.array([1,   1,   0,   1,   0,   1,   0,   0,   0,   0  ])
P = y.sum(); tp = fp = 0; ap = 0.0; prev_rec = 0.0
for i in np.argsort(-scores):
    tp += y[i]; fp += 1 - y[i]
    prec = tp / (tp + fp); rec = tp / P
    if y[i] == 1:
        ap += prec * (rec - prev_rec); prev_rec = rec
print(round(ap, 4), round(average_precision_score(y, scores), 4))

AP steps only on positives, weighting precision by the recall gained

Why: Each positive crossing lifts recall by 1/P and contributes precision·(that gain). sklearn defines AP as exactly this sum, Σ (Rₙ − Rₙ₋₁)·Pₙ.

crossedlabelprecisionrecallAP running
0.9511.0000.2500.250
0.9011.0000.5000.500
0.8000.6670.5000.500
0.7010.7500.7500.688
0.4510.6671.0000.854
final AP———0.8542

48. Which is which, by label

Discrimination

Sort into buckets

Sort these by label, from memory, without looking back at Build the PR curve and AP from scratch. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

1
0.95; 0.90; 0.70; 0.45
0
0.80
—
final AP
g1
label is "1" for 0.95, 0.90, 0.70, 0.45 — that is what the table on "Build the PR curve and AP from scratch" records, and it is the single property separating this group from the rest.
g2
label is "0" for 0.80 — that is what the table on "Build the PR curve and AP from scratch" records, and it is the single property separating this group from the rest.
g3
label is "—" for final AP — that is what the table on "Build the PR curve and AP from scratch" records, and it is the single property separating this group from the rest.

49. AP matches sklearn exactly

Concept

Our hand-accumulated AP is 0.8542, and average_precision_score returns the identical 0.8542. The manual PR curve is faithful.

methodAP (verified)
manual Σ precision·Δrecall0.8542
sklearn average_precision_score0.8542

On this balanced-ish data (4/10 positive) AP 0.854 and AUC 0.875 are close. The gap explodes only when positives become rare — which is the next experiment.

50. F1: one number at one threshold

Concept

PR-AUC summarizes the whole curve. If instead you must pick one threshold, the F1 score balances precision and recall at that operating point via their harmonic mean:

\[ F_1 = \frac{2 \cdot \text{precision} \cdot \text{recall}}{\text{precision} + \text{recall}} \]

The harmonic mean punishes imbalance between the two: F1 is high only when precision and recall are both high. A model with precision 1.0 but recall 0.1 scores F1 ≈ 0.18, not 0.55 — one weak side drags it down.

51. Why harmonic, not arithmetic, mean

Intuition

Averaging precision and recall the ordinary way would let a lopsided model cheat: precision 1.0, recall 0.1 averages to 0.55 — sounds decent, but the model finds only one positive in ten.

The harmonic mean is dominated by the smaller number. It sits near the weaker of the two, so you can't paper over a terrible recall with a great precision. That is exactly the behavior you want from a single 'is this model actually usable' score.

Same reason a car that goes 100 mph one way and 1 mph back doesn't average 50 mph over the trip — time is dominated by the slow leg. F1 is dominated by the weak metric.

52. Finish it with less help: Pick the operating threshold

Faded example

Fill in the blanks

Pick the operating threshold, with the scaffolding fading: two lines are gone now — fill both.

import numpy as np
scores = np.array([0.95,0.90,0.80,0.70,0.55,0.45,0.40,0.30,0.20,0.10])
y = np.array([1, 1, 0, 1, 0, 1, 0, 0, 0, 0 ])
P = y.sum(); tp = fp = 0; rows = []
for i in np.argsort(-scores):
tp += y[i]; fp += 1 - y[i]
prec = tp/(tp+fp); rec = tp/P
f1 = 0.0 if prec+rec == 0 else 2precrec/(prec+rec)
rows.append((scores[i], prec, rec, f1))
best = max(rows, key=lambda r: r[3])
print("best F1", round(best[3],3), "at thr", best[0])

Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. At t = 0.45 precision = 0.667 and recall = 1.000, giving the harmonic mean 0.800 — the highest of any threshold.

53. Pick the operating threshold

Worked example

PR-AUC and F1 don't tell you which threshold to deploy. Sweep them, then choose by your requirement: best F1, or the most recall you can get while holding precision ≥ 0.75. Runnable and self-contained:

import numpy as np
scores = np.array([0.95,0.90,0.80,0.70,0.55,0.45,0.40,0.30,0.20,0.10])
y      = np.array([1,   1,   0,   1,   0,   1,   0,   0,   0,   0  ])
P = y.sum(); tp = fp = 0; rows = []
for i in np.argsort(-scores):
    tp += y[i]; fp += 1 - y[i]
    prec = tp/(tp+fp); rec = tp/P
    f1 = 0.0 if prec+rec == 0 else 2*prec*rec/(prec+rec)
    rows.append((scores[i], prec, rec, f1))
best = max(rows, key=lambda r: r[3])
print("best F1", round(best[3],3), "at thr", best[0])

Best F1 = 0.800 at threshold 0.45

Why: At t = 0.45 precision = 0.667 and recall = 1.000, giving the harmonic mean 0.800 — the highest of any threshold. Lowering t further only adds false positives and drops F1.

Need precision ≥ 0.75? Best recall is 0.75 at t = 0.70

Why: Scanning the table for rows with precision ≥ 0.75, the one with the most recall is t = 0.70 (precision 0.75, recall 0.75). The right threshold depends on which error costs you more.

thresholdprecisionrecallF1
0.901.0000.5000.667
0.700.7500.7500.750
0.45 (best F1)0.6671.0000.800
0.400.5711.0000.727
0.200.4441.0000.615

54. Watch it run: Pick the operating threshold

Pattern

Step through it

Step through Pick the operating threshold one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: threshold is 0.90
  2. Step 2: threshold is 0.70
  3. Step 3: threshold is 0.45 (best F1)
  4. Step 4: threshold is 0.40
  5. Step 5: threshold is 0.20

55. Why ROC-AUC flatters rare positives

Intuition

FPR's denominator is FP + TN — the count of all negatives. When 98% of the data is negative, that denominator is enormous, so even hundreds of false positives barely move FPR off zero.

So the ROC curve stays pinned to the left edge and the AUC looks great — while in reality the model floods you with false alarms for every true positive it finds.

Precision's denominator is TP + FP — only the flagged examples. It has no giant TN to hide behind, so it feels every false positive. That is why PR is the honest metric under imbalance.

56. Restore the missing line: ROC-AUC vs PR-AUC under 2% imbalance

Fill the middle

Fill in the blanks

From ROC-AUC vs PR-AUC under 2% imbalance — one line has had its right-hand side removed. Put it back.

import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.metrics import roc_auc_score, average_precision_score
X, y = make_classification(n_samples=4000, weights=[0.98, 0.02],
n_informative=5, random_state=1)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.5, random_state=0)
p = LogisticRegression(max_iter=2000).fit(Xtr, ytr).predict_proba(Xte)[:, 1]
print(round(yte.mean(), 4))
print(round(roc_auc_score(yte, p), 3), round(average_precision_score(yte, p), 3))

Why: n_informative is what everything below it consumes, so the wrong expression here fails later and somewhere else. A middling-but-usable-sounding number.

57. ROC-AUC vs PR-AUC under 2% imbalance

Worked example

Train logistic regression on a synthetic dataset that is only 2% positive, then compare the two summary numbers on held-out data. Fully runnable and seeded:

import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.metrics import roc_auc_score, average_precision_score
X, y = make_classification(n_samples=4000, weights=[0.98, 0.02],
                           n_informative=5, random_state=1)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.5, random_state=0)
p = LogisticRegression(max_iter=2000).fit(Xtr, ytr).predict_proba(Xte)[:, 1]
print(round(yte.mean(), 4))
print(round(roc_auc_score(yte, p), 3), round(average_precision_score(yte, p), 3))

ROC-AUC = 0.736 looks respectable...

Why: A middling-but-usable-sounding number. If this were your only metric you might ship the model.

...but PR-AUC = 0.111 exposes a near-useless detector

Why: The random-guess baseline for AP is the prevalence, 0.022. So 0.111 is only ~5× better than chance — the model barely finds the rare positives. ROC hid that entirely.

metricvalue (verified)story
test positive rate0.0225heavy imbalance
ROC-AUC0.736deceptively rosy
PR-AUC (AP)0.111the honest picture
AP baseline (prevalence)0.022random-guess floor

58. Fill in: story for ROC-AUC vs PR-AUC under 2% imbalance

Comparison

Comparison matrix

From ROC-AUC vs PR-AUC under 2% imbalance: refill the story column from what you know. The rest of the table is as it appeared.

metricvalue (verified)story
test positive rate0.0225heavy imbalance
ROC-AUC0.736deceptively rosy
PR-AUC (AP)0.111the honest picture
AP baseline (prevalence)0.022random-guess floor

59. Something is wrong here: judging imbalance by ROC-AUC

Anomaly

Predict first

A student writes this, and it looks reasonable:

The rare-event model scores ROC-AUC 0.74, comfortably above 0.5.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: ROC-AUC is dominated by the millions of easy negatives, whose huge count keeps FPR tiny.

On heavy imbalance, report PR-AUC / average precision and compare it to the prevalence baseline.

Why: ROC-AUC is dominated by the millions of easy negatives, whose huge count keeps FPR tiny. The 0.74 is measuring 'can it rank the easy negatives low', not 'can it find the positives'.

60. Trap: judging imbalance by ROC-AUC

Trap

The trap

The rare-event model scores ROC-AUC 0.74, comfortably above 0.5.

Ship it — 0.74 AUC means a decent detector

Why: ROC-AUC is dominated by the millions of easy negatives, whose huge count keeps FPR tiny. The 0.74 is measuring 'can it rank the easy negatives low', not 'can it find the positives'.

The fix

On heavy imbalance, report PR-AUC / average precision and compare it to the prevalence baseline.

PR-AUC = 0.111 vs 0.022 baseline → weak; do not ship

Why: PR ignores true negatives, so it reflects the real difficulty of the positive class. 0.111 against a 0.022 floor says the detector is only marginally better than guessing.

61. Break it on purpose: judging imbalance by ROC-AUC

Break the constraint

Discussion prompt

The rule this trap just fixed:

PR ignores true negatives, so it reflects the real difficulty of the positive class. 0.111 against a 0.022 floor says the detector is only marginally better than guessing.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

ROC-AUC is dominated by the millions of easy negatives, whose huge count keeps FPR tiny. The 0.74 is measuring 'can it rank the easy negatives low', not 'can it find the positives'.

62. Rebuild the recipe: The curve-selection recipe

Ranking

Put in order

These are the steps of The curve-selection recipe, scrambled. Put them back in order before the next slide shows you.

  1. Build ROC: sweep threshold high→low; step up on a positive, right on a negative; area = AUC
  2. Read AUC as P(random positive > random negative) — threshold-free ranking quality
  3. Reference shapes: perfect 1.0, random 0.5, worse < 0.5 (flip it → 1 − AUC)
  4. Balanced-ish classes? ROC-AUC is fine
  5. Rare positives? Switch to PR-AUC / average precision, and compare it to the prevalence baseline, never ROC-AUC

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

63. The curve-selection recipe

Pattern

  1. Build ROC: sweep threshold high→low; step up on a positive, right on a negative; area = AUC
  2. Read AUC as P(random positive > random negative) — threshold-free ranking quality
  3. Reference shapes: perfect 1.0, random 0.5, worse < 0.5 (flip it → 1 − AUC)
  4. Balanced-ish classes? ROC-AUC is fine
  5. Rare positives? Switch to PR-AUC / average precision, and compare it to the prevalence baseline, never ROC-AUC

64. Where does it stop working: The curve-selection recipe

Edge cases

Discussion prompt

The curve-selection recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Build ROC: sweep threshold high→low; step up on a positive, right on a negative; area = AUC
  2. Read AUC as P(random positive > random negative) — threshold-free ranking quality
  3. Reference shapes: perfect 1.0, random 0.5, worse < 0.5 (flip it → 1 − AUC)
  4. Balanced-ish classes? ROC-AUC is fine
  5. Rare positives? Switch to PR-AUC / average precision, and compare it to the prevalence baseline, never ROC-AUC

65. Rule out three: Check yourself — worse than random

Elimination

Eliminate the wrong options

A classifier has AUC-ROC = 0.3. The best move is:

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Flip its predictions — the flipped classifier has AUC = 1 − 0.3 = 0.7
  • B. Discard it; any AUC below 0.5 is useless noise
  • C. Retrain from scratch — 0.3 means the model learned nothing
  • D. Report 0.3 as-is; it is a valid score

Survives elimination: A

Why: AUC < 0.5 means the model ranks systematically backwards — it has real signal, just inverted. Flipping every decision turns each discordant pair concordant, giving AUC = 1 − 0.3 = 0.7. We verified exactly this: 0.125 flips to 0.875.

66. Check yourself — worse than random

Check

Read the AUC before you answer.

Check your understanding

A classifier has AUC-ROC = 0.3. The best move is:

  • A. Flip its predictions — the flipped classifier has AUC = 1 − 0.3 = 0.7 (correct)
  • B. Discard it; any AUC below 0.5 is useless noise
  • C. Retrain from scratch — 0.3 means the model learned nothing
  • D. Report 0.3 as-is; it is a valid score

Answer: A

Why: AUC < 0.5 means the model ranks systematically backwards — it has real signal, just inverted. Flipping every decision turns each discordant pair concordant, giving AUC = 1 − 0.3 = 0.7. We verified exactly this: 0.125 flips to 0.875.

Why B tempts people
Consistent wrongness is information, not noise. A truly noisy model sits at 0.5; 0.3 is confidently backwards, and inverting recovers 0.7.
Why C tempts people
AUC = 0.5 is the 'learned nothing' point. 0.3 is below it, meaning the model DID learn a real pattern but assigns it the wrong sign — no retraining needed, just a flip.
Why D tempts people
Reporting 0.3 wastes usable signal. The honest and stronger report is the flipped 0.7.

67. Answer it before you see the options: Check yourself — which curve under…

Prediction

Predict first

On a 2%-positive dataset a model shows ROC-AUC = 0.74 but PR-AUC = 0.11. Which better summarizes its ability to find the rare positives?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: PR-AUC = 0.11 — ROC-AUC is inflated by the many easy negatives

Why: FPR's denominator FP+TN is dominated by the 98% negatives, so ROC-AUC stays optimistic no matter how many false alarms occur. PR-AUC uses TP+FP (no TN), so it feels every false positive — 0.11 against a 0.022 prevalence baseline is the honest, weak picture.

68. Check yourself — which curve under imbalance

Check

Think about which denominator hides the false alarms.

Check your understanding

On a 2%-positive dataset a model shows ROC-AUC = 0.74 but PR-AUC = 0.11. Which better summarizes its ability to find the rare positives?

  • A. PR-AUC = 0.11 — ROC-AUC is inflated by the many easy negatives (correct)
  • B. ROC-AUC = 0.74 — it is the higher, more optimistic number
  • C. The average of the two, ≈ 0.43
  • D. Plain accuracy, which is more interpretable

Answer: A

Why: FPR's denominator FP+TN is dominated by the 98% negatives, so ROC-AUC stays optimistic no matter how many false alarms occur. PR-AUC uses TP+FP (no TN), so it feels every false positive — 0.11 against a 0.022 prevalence baseline is the honest, weak picture.

Why B tempts people
Higher is not more honest — ROC-AUC is high precisely BECAUSE the easy, numerous negatives keep FPR near zero. That is the illusion, not the signal.
Why C tempts people
Averaging an inflated metric with an honest one just launders the inflation into the result. Report PR-AUC directly.
Why D tempts people
Accuracy is the worst choice here: predicting 'always negative' scores 98% while finding zero positives (Lesson 26).

69. Calibration

Section

Part 4 of 6 — trust the probabilities

70. Ranking is not the same as honesty

Intuition

AUC and AP only judge ranking — do positives score above negatives. They say nothing about whether a 0.9 really means a 90% chance.

A model can rank perfectly (AUC = 1) yet output 0.9 on cases that are right only 60% of the time. Its scores are ordered correctly but dishonest as probabilities.

Whenever a downstream decision uses the probability itself — expected cost, a betting threshold, medical risk — you need the scores to be calibrated, not just well-ranked.

71. What calibration means

Concept

A model is calibrated if its confidence matches reality: among all the cases where it predicts 0.8, about 80% should actually be positive.

reliability diagram — A plot of predicted confidence (x) against observed accuracy (y), one point per confidence bin. Perfect calibration lies on the diagonal. Points BELOW the diagonal mean the model is overconfident — its confidence exceeds its accuracy — which is the typical failure of modern deep nets.

72. Take the definitions apart: TP / FP / FN / TN vs reliability diagram

Definition probe

Sort into buckets

Every line below is part of the definition of TP / FP / FN / TN or of reliability diagram — one or the other, never both. Put each where it belongs.

TP / FP / FN / TN
called positive, was positive (a hit).; called positive, was negative (a false alarm).; called negative, was positive (a miss).
reliability diagram
A plot of predicted confidence (x) against observed accuracy (y), one point per confidence bin.; Perfect calibration lies on the diagonal.; Points BELOW the diagonal mean the model is overconfident
b1
TP: called positive, was positive (a hit). FP: called positive, was negative (a false alarm). FN: called negative, was positive (a miss). TN: called negative, was negative (a correct rejection). Every metric today is built from just these four numbers.
b2
A plot of predicted confidence (x) against observed accuracy (y), one point per confidence bin. Perfect calibration lies on the diagonal. Points BELOW the diagonal mean the model is overconfident — its confidence exceeds its accuracy — which is the typical failure of modern deep nets.

73. Discrimination and calibration are independent

Concept

These are two separate axes of a model, and a model can be good on one and bad on the other:

well-ranked (high AUC)poorly-ranked (low AUC)
calibratedthe goal — trust the numberhonest but weak signal
miscalibratedgreat order, dishonest valuesworst case

Our ×3-overconfident model lives in the top-right box: its ranking is untouched (temperature scaling proves the argmax never moves), yet its probabilities are dishonest. Fixing calibration leaves AUC exactly where it was — the two axes don't interfere.

74. Expected Calibration Error (ECE)

Concept

ECE turns the reliability gap into one number. Bin the predictions by confidence, and in each bin take the absolute gap between mean confidence and observed accuracy, weighted by how many points fall in the bin:

\[ \text{ECE} = \sum_{b=1}^{B} \frac{n_b}{n}\,\big|\,\text{acc}_b - \text{conf}_b\,\big| \]

accᵦ is the fraction of bin b that is actually positive; confᵦ is the average predicted probability in that bin; nᵦ/n weights each bin by its share of the data. ECE = 0 is perfect calibration.

75. Predict the next row: ECE bin by bin on an overconfident model

Pattern

Predict first

The table runs: 0.0–0.1 | 1285 | 0.023 | 0.181 | 0.158 · 0.4–0.5 | 154 | 0.450 | 0.396 | 0.054 · 0.7–0.8 | 189 | 0.749 | 0.571 | 0.177 · 0.8–0.9 | 268 | 0.855 | 0.649 | 0.205 · 0.9–1.0 | 1198 | 0.976 | 0.841 | 0.136

In ECE bin by bin on an overconfident model, given the rows so far: what is the next one — the row where bin is weighted ECE?

Correct: weighted ECE | — | — | — | 0.1448

binnconfaccgap
0.0–0.112850.0230.1810.158
0.4–0.51540.4500.3960.054
0.7–0.81890.7490.5710.177
0.8–0.92680.8550.6490.205
0.9–1.011980.9760.8410.136
weighted ECE———0.1448

Why: The relationship between the columns, not the individual numbers, is what generates the next row. The 0.9–1.0 bin holds 1198 predictions at mean confidence 0.976 but only 84.1% are positive — a 0.136 gap.

76. ECE bin by bin on an overconfident model

Worked example

Make the abstraction concrete. Take the ×3-overconfident model, split its predictions into 10 confidence bins, and print each bin's count, mean confidence, and observed accuracy. Self-contained and seeded:

import numpy as np
np.random.seed(0)
sig = lambda z: 1.0/(1.0+np.exp(-z))
z = np.random.randn(4000)*1.5
labels = (np.random.rand(4000) < sig(z)).astype(float)
probs = sig(z*3.0)                 # overconfident predictions
edges = np.linspace(0, 1, 11); ece = 0.0
for i in range(10):
    m = (probs > edges[i]) & (probs <= edges[i+1])
    if m.sum():
        conf, acc = probs[m].mean(), labels[m].mean()
        ece += m.mean()*abs(acc - conf)
        print(f"{edges[i]:.1f}-{edges[i+1]:.1f} n={m.sum():4d} conf={conf:.3f} acc={acc:.3f}")
print("ECE =", round(ece, 4))

The extreme bins are where overconfidence bites

Why: The 0.9–1.0 bin holds 1198 predictions at mean confidence 0.976 but only 84.1% are positive — a 0.136 gap. The 0.0–0.1 bin is symmetric: confidence 0.023 but 18.1% positive. The ×3 pushed probabilities to the extremes, past what the data supports.

binnconfaccgap
0.0–0.112850.0230.1810.158
0.4–0.51540.4500.3960.054
0.7–0.81890.7490.5710.177
0.8–0.92680.8550.6490.205
0.9–1.011980.9760.8410.136
weighted ECE———0.1448

77. Watch it run: ECE bin by bin on an overconfident model

Pattern

Step through it

Step through ECE bin by bin on an overconfident model one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: bin is 0.0–0.1
  2. Step 2: bin is 0.4–0.5
  3. Step 3: bin is 0.7–0.8
  4. Step 4: bin is 0.8–0.9
  5. Step 5: bin is 0.9–1.0
  6. Step 6: bin is weighted ECE

78. Three post-hoc fixes

Concept

Calibration is repaired after training, without touching the model's ranking. Three standard tools:

Temperature scaling is the workhorse for neural nets: one number, provably preserves accuracy, and directly attacks overconfidence. We derive and sweep it next.

79. The temperature mechanism

Concept

A model's raw pre-sigmoid output is its logit z. The probability is p = σ(z) = 1/(1 + e^{−z}). Temperature scaling divides the logit by T before the sigmoid:

\[ p_T = \sigma\!\left(\frac{z}{T}\right) = \frac{1}{1 + e^{-z/T}} \]

T > 1 shrinks the logits toward 0, pulling probabilities toward 0.5 — it softens an overconfident model. T < 1 sharpens them. T = 1 leaves the model unchanged.

80. Teach it back: The temperature mechanism

Explain it

Discussion prompt

Explain The temperature mechanism to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

A model's raw pre-sigmoid output is its logit z. The probability is p = σ(z) = 1/(1 + e^{−z}). Temperature scaling divides the logit by T before the sigmoid:

81. What has to be given first: Temperature scaling drives ECE down

Missing information

Discussion prompt

Build a well-calibrated latent (label ~ Bernoulli(σ(z))), then make the model overconfident by reporting 3z instead of z. Sweep T and watch ECE. Fully self-contained and seeded:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

Multiplying the honest logit z by 3 pushes σ toward 0 and 1 — the model reports 0.99 where reality is ~0.9. That mismatch is exactly what ECE measures.

82. Temperature scaling drives ECE down

Worked example

Build a well-calibrated latent (label ~ Bernoulli(σ(z))), then make the model overconfident by reporting 3z instead of z. Sweep T and watch ECE. Fully self-contained and seeded:

import numpy as np
np.random.seed(0)
sig = lambda z: 1.0/(1.0+np.exp(-z))
def ece(probs, labels, bins=10):
    e = 0.0; edges = np.linspace(0, 1, bins+1)
    for i in range(bins):
        m = (probs > edges[i]) & (probs <= edges[i+1])
        if m.sum():
            e += m.mean()*abs(probs[m].mean() - labels[m].mean())
    return e
n = 4000
z = np.random.randn(n)*1.5
labels = (np.random.rand(n) < sig(z)).astype(float)
logits = z*3.0                      # overconfident model
for T in [1, 2, 3, 4, 5]:
    print(T, round(ece(sig(logits/T), labels), 4))

logits = z*3 makes the model overconfident (×3 sharper)

Why: Multiplying the honest logit z by 3 pushes σ toward 0 and 1 — the model reports 0.99 where reality is ~0.9. That mismatch is exactly what ECE measures.

Dividing by T = 3 exactly undoes the ×3 → ECE minimized

Why: σ(logits/T) = σ(3z/3) = σ(z), the honest probability. So T = 3 should give the lowest ECE — and it does, dropping ECE from 0.1448 to 0.0176.

TECE (verified)note
1 (raw)0.1448badly overconfident
20.0583improving
3 (best)0.0176undoes the ×3
40.0524now underconfident
50.0830over-softened

83. Inspect it line by line: Temperature scaling drives ECE down

Error analysis

Annotate

Walk the callouts on Temperature scaling drives ECE down. Each one is a place this is easy to get subtly wrong.

  • Multiplying the honest logit z by 3 pushes σ toward 0 and 1 — the model reports 0.99 where reality is ~0.9. That mismatch is exactly what ECE measures.
  • σ(logits/T) = σ(3z/3) = σ(z), the honest probability. So T = 3 should give the lowest ECE — and it does, dropping ECE from 0.1448 to 0.0176.

84. Reading the ECE sweep

Concept

ECE falls from 0.1448 at T = 1 to its minimum 0.0176 at T = 3, then climbs back to 0.0830 at T = 5. The curve is a valley: too little softening leaves overconfidence, too much creates underconfidence.

The minimum sits exactly at T = 3 because that is the factor by which we inflated the logits. In practice you don't know the true factor — you fit T by minimizing ECE (or NLL) on a held-out set.

85. Guess the shape of the answer: Temperature never changes the prediction

Estimation

Predict first

Crucial property: dividing all logits by the same positive T preserves their order, so the argmax — the actual prediction — never moves. Accuracy is untouched:

Commit before you compute: what does Temperature never changes the prediction come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Ranking identical, class predictions identical

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. z/T is a monotone rescaling, so argsort is unchanged and every prediction stays on the same side of 0.5.

86. Temperature never changes the prediction

Worked example

Crucial property: dividing all logits by the same positive T preserves their order, so the argmax — the actual prediction — never moves. Accuracy is untouched:

import numpy as np
sig = lambda z: 1.0/(1.0+np.exp(-z))
logits = np.array([2.0, -1.0, 0.5, 3.0, -2.5])
p1 = sig(logits/1.0)
p3 = sig(logits/3.0)
print(np.round(p1, 3))
print(np.round(p3, 3))
print((np.argsort(-p1) == np.argsort(-p3)).all())
print((p1 > 0.5).astype(int), (p3 > 0.5).astype(int))

Ranking identical, class predictions identical

Why: z/T is a monotone rescaling, so argsort is unchanged and every prediction stays on the same side of 0.5. Softening moves probabilities toward 0.5 but never across the decision boundary here.

logit zp at T=1p at T=3pred (both)
2.00.8810.6611
-1.00.2690.4170
0.50.6220.5421
3.00.9530.7311
-2.50.0760.3030

87. Which is which, by pred (both)

Discrimination

Sort into buckets

Sort these by pred (both), from memory, without looking back at Temperature never changes the prediction. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

1
2.0; 0.5; 3.0
0
-1.0; -2.5
g1
pred (both) is "1" for 2.0, 0.5, 3.0 — that is what the table on "Temperature never changes the prediction" records, and it is the single property separating this group from the rest.
g2
pred (both) is "0" for -1.0, -2.5 — that is what the table on "Temperature never changes the prediction" records, and it is the single property separating this group from the rest.

88. Something is wrong here: confident means correct

Anomaly

Predict first

A student writes this, and it looks reasonable:

The model outputs 0.99, so this case is almost certainly a positive — trust it.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Confidence and accuracy are different axes.

First measure ECE; if the model is overconfident, recalibrate, then read the probability.

Why: Confidence and accuracy are different axes. An overconfident model (our ECE 0.1448 case) prints 0.99 on plenty of cases it gets wrong. The number is not a probability until the model is calibrated.

89. Trap: confident means correct

Trap

The trap

The model outputs 0.99, so this case is almost certainly a positive — trust it.

Treat 0.99 as ~99% reliable

Why: Confidence and accuracy are different axes. An overconfident model (our ECE 0.1448 case) prints 0.99 on plenty of cases it gets wrong. The number is not a probability until the model is calibrated.

The fix

First measure ECE; if the model is overconfident, recalibrate, then read the probability.

Fit T, apply σ(z/T), THEN trust 0.99

Why: Only a calibrated 0.99 means ~99% correct. Temperature scaling restores that meaning (ECE 0.1448 → 0.0176) without changing any prediction — same argmax, honest probability.

90. Which of these survive contact with Lesson 35: ROC, PR Curves & Calibration?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
One running example for the whole lesson. Ten examples, each with a model score s and a true label y (1 = positive). Sorted high score to low:; Fix a threshold. Every example falls into exactly one of four boxes, comparing what the model predicted against the truth.; The two axes of the ROC curve, plus precision, are ratios of those four counts. Learn which denominator each one uses — that choice is the whole story of imbalance.
Breaks
ROC plots the two rates; put TPR on the x-axis and FPR on the y-axis, then integrate.; The rare-event model scores ROC-AUC 0.74, comfortably above 0.5.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 35: ROC, PR Curves & Calibration puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

91. One knob, two jobs

Section

Part 5 of 6 — temperature elsewhere

92. Temperature in LLM sampling

Concept

The exact same T knob controls text generation. Before the softmax over next-token logits, dividing by T reshapes the sampling distribution — same mechanism, different goal.

\[ \Pr(\text{token}_k) = \frac{e^{z_k / T}}{\sum_j e^{z_j / T}} \]

T < 1 sharpens toward the top token (more deterministic, 'greedy'); T > 1 flattens toward uniform (more diverse, more random). T → 0 is pure argmax; T → ∞ is a coin flip over the vocabulary.

93. By analogy: Temperature in LLM sampling

Analogy

Discussion prompt

Explain Temperature in LLM sampling by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

The exact same T knob controls text generation. Before the softmax over next-token logits, dividing by T reshapes the sampling distribution — same mechanism, different goal.

94. Same math, two purposes

Intuition

In calibration, you fit T to a fixed value that makes probabilities honest — you want the number to be trustworthy.

In sampling, you choose T as a creativity dial — low for factual answers, high for brainstorming. Nothing is being calibrated; you're tuning randomness.

Recognizing that σ(z/T) and softmax(z/T) are the same operation is the kind of connection USAAIO loves to test — calibration (this lesson) and softmax temperature (Lesson 17) are one idea.

95. Break it if you can: Same math, two purposes

Counterexample

Discussion prompt

In calibration, you fit T to a fixed value that makes probabilities honest — you want the number to be trustworthy.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

In sampling, you choose T as a creativity dial — low for factual answers, high for brainstorming. Nothing is being calibrated; you're tuning randomness.

96. Answer it before you see the options: Check yourself — what temperature does

Prediction

Predict first

Applying temperature scaling with T > 1 to an overconfident model's logits:

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: Softens the probabilities toward 0.5, lowering ECE, with the argmax unchanged

Why: Dividing logits by T > 1 shrinks them toward 0, pulling probabilities toward 0.5 and reducing overconfidence — we saw ECE drop 0.1448 → 0.0176 at T = 3. Because z/T is monotone, the argmax and every 0.5-threshold prediction are unchanged, so accuracy is preserved.

97. Check yourself — what temperature does

Check

T > 1 applied to logits — think about the argmax.

Check your understanding

Applying temperature scaling with T > 1 to an overconfident model's logits:

  • A. Softens the probabilities toward 0.5, lowering ECE, with the argmax unchanged (correct)
  • B. Changes which class is predicted for many examples
  • C. Always raises accuracy along with calibration
  • D. Sharpens the probabilities toward 0 and 1

Answer: A

Why: Dividing logits by T > 1 shrinks them toward 0, pulling probabilities toward 0.5 and reducing overconfidence — we saw ECE drop 0.1448 → 0.0176 at T = 3. Because z/T is monotone, the argmax and every 0.5-threshold prediction are unchanged, so accuracy is preserved.

Why B tempts people
Dividing all logits by the same positive T preserves their order, so the argmax never moves — we verified identical predictions at T=1 and T=3.
Why C tempts people
Accuracy is unchanged (same argmax), not raised. Temperature scaling fixes CALIBRATION, not accuracy — that is precisely its guarantee.
Why D tempts people
T > 1 FLATTENS (softens) toward 0.5; it is T < 1 that sharpens toward 0 and 1.

98. Rule out three: Check yourself — reading an AUC point

Elimination

Eliminate the wrong options

On our data at threshold t = 0.5 the ROC point is (FPR, TPR) = (0.333, 0.75). Which statement is correct?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. 75% of true positives are caught and 33% of true negatives are falsely flagged at this threshold
  • B. The model's AUC is 0.75
  • C. 75% of the flagged examples are actually positive
  • D. 33% of all examples are misclassified

Survives elimination: A

Why: TPR = 0.75 = TP/(TP+FN) is the recall — the share of real positives caught. FPR = 0.333 = FP/(FP+TN) is the share of real negatives wrongly flagged. Each ROC point is one threshold's (FPR, TPR) trade-off.

99. Check yourself — reading an AUC point

Check

Recall what a single ROC point encodes.

Check your understanding

On our data at threshold t = 0.5 the ROC point is (FPR, TPR) = (0.333, 0.75). Which statement is correct?

  • A. 75% of true positives are caught and 33% of true negatives are falsely flagged at this threshold (correct)
  • B. The model's AUC is 0.75
  • C. 75% of the flagged examples are actually positive
  • D. 33% of all examples are misclassified

Answer: A

Why: TPR = 0.75 = TP/(TP+FN) is the recall — the share of real positives caught. FPR = 0.333 = FP/(FP+TN) is the share of real negatives wrongly flagged. Each ROC point is one threshold's (FPR, TPR) trade-off.

Why B tempts people
AUC is the area under the WHOLE curve (0.875 here), not the TPR of any single point. One point is one threshold; AUC integrates over all of them.
Why C tempts people
That describes precision = TP/(TP+FP) = 3/5 = 0.60, not TPR. Precision splits by predicted class; TPR splits by true class.
Why D tempts people
The misclassified count is FP + FN = 2 + 1 = 3 of 10 = 30%, and it is not what either ROC axis reports; 0.333 is FPR specifically.

100. Your turn: build it

Section

Part 6 of 6 — the project

101. Project: ROC from scratch → PR → calibration

Concept

Assemble the whole lesson yourself: build an ROC curve and AUC by hand, contrast ROC-AUC with PR-AUC on imbalanced data, then recalibrate an overconfident model with temperature scaling. You derived every piece — now wire them together.

#requirementtool
1Sweep thresholds → TPR/FPR → AUCnp.argsort, np.trapz
2ROC-AUC vs PR-AUC on 2%-positive dataroc_auc_score, average_precision_score
3Temperature-scale to the ECE minimumσ(logits/T), an ece() function

Build rules: type every line yourself, sort by descending score for the ROC sweep, and confirm your manual AUC matches sklearn before moving on. When a shape error hits, print .shape — read the error, don't delete it.

102. Milestone 1 — ROC & AUC by hand

Worked example

Your turn: sweep thresholds to accumulate TPR/FPR, integrate for the AUC, and check it against sklearn. Predict out loud: will they match exactly?

Hint: walk np.argsort(-scores) (descending). Each example adds to tp if positive or fp if negative; append tp/P and fp/N; finish with np.trapz(tpr, fpr).

import numpy as np
from sklearn.metrics import roc_auc_score
scores = np.array([0.95,0.90,0.80,0.70,0.55,0.45,0.40,0.30,0.20,0.10])
y      = np.array([1,   1,   0,   1,   0,   1,   0,   0,   0,   0  ])
P, N = y.sum(), (1-y).sum()
tp = fp = 0; tpr = [0.0]; fpr = [0.0]
for i in np.argsort(-scores):
    tp += y[i]; fp += 1 - y[i]
    tpr.append(tp/P); fpr.append(fp/N)
print(round(np.trapz(tpr, fpr), 4), round(roc_auc_score(y, scores), 4))
sourceAUC (expected)
your manual trapezoid0.8750
sklearn roc_auc_score0.8750

103. Milestone 2 — ROC-AUC vs PR-AUC

Worked example

Your turn: train logistic regression on 2%-positive data and print both AUCs. Predict which one collapses.

Hint: make_classification(weights=[0.98, 0.02], ...), split, predict_proba(...)[:, 1], then roc_auc_score vs average_precision_score. Keep the same random_states to reproduce the numbers.

import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.metrics import roc_auc_score, average_precision_score
X, y = make_classification(n_samples=4000, weights=[0.98, 0.02],
                           n_informative=5, random_state=1)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.5, random_state=0)
p = LogisticRegression(max_iter=2000).fit(Xtr, ytr).predict_proba(Xte)[:, 1]
print(round(roc_auc_score(yte, p), 3), round(average_precision_score(yte, p), 3))
metricvalue (expected)
ROC-AUC0.736
PR-AUC (AP)0.111
prevalence baseline0.022

104. Milestone 3 — temperature scaling

Worked example

Your turn: write an ece() function, build an overconfident model (×3 logits), and sweep T. Predict the best T before you run it.

Hint: ece bins predictions into 10 bins and averages |conf − acc| weighted by bin size. The overconfidence factor is 3, so the ECE minimum should land at T ≈ 3.

import numpy as np
np.random.seed(0)
sig = lambda z: 1.0/(1.0+np.exp(-z))
def ece(probs, labels, bins=10):
    e = 0.0; edges = np.linspace(0, 1, bins+1)
    for i in range(bins):
        m = (probs > edges[i]) & (probs <= edges[i+1])
        if m.sum():
            e += m.mean()*abs(probs[m].mean() - labels[m].mean())
    return e
z = np.random.randn(4000)*1.5
labels = (np.random.rand(4000) < sig(z)).astype(float)
logits = z*3.0
for T in [1, 2, 3, 4, 5]:
    print(T, round(ece(sig(logits/T), labels), 4))
TECE (expected)
1 (raw)0.1448
20.0583
3 (best)0.0176
40.0524
50.0830

105. What each one costs: Milestone 3 — temperature scaling

Trade off

Comparison matrix

From Milestone 3 — temperature scaling: every row here is a choice with a cost. Fill the ECE (expected) column, then say which row you would actually pick and what you give up for it.

TECE (expected)
1 (raw)0.1448
20.0583
3 (best)0.0176
40.0524
50.0830

106. The full program

Concept

import numpy as np
from sklearn.metrics import roc_auc_score, average_precision_score
# 1. ROC-AUC from scratch
scores = np.array([0.95,0.90,0.80,0.70,0.55,0.45,0.40,0.30,0.20,0.10])
y      = np.array([1,   1,   0,   1,   0,   1,   0,   0,   0,   0  ])
P, N = y.sum(), (1-y).sum(); tp = fp = 0; T = [0.0]; F = [0.0]
for i in np.argsort(-scores):
    tp += y[i]; fp += 1-y[i]; T.append(tp/P); F.append(fp/N)
auc = np.trapz(T, F)
print('ROC-AUC:', round(auc, 4), '== sklearn', bool(np.isclose(auc, roc_auc_score(y, scores))))
print('AP:', round(average_precision_score(y, scores), 4))
# 2. imbalance: ROC-AUC 0.736 rosy vs PR-AUC 0.111 honest
# 3. temperature scaling: ECE 0.1448 (T=1) -> 0.0176 (T=3)
printed linevalue (verified)
ROC-AUC: ... == sklearn0.875, True
AP:0.8542
imbalance ROC / PR0.736 / 0.111
ECE T=1 → T=30.1448 → 0.0176

If your ROC matches sklearn, PR-AUC exposes the imbalance, and temperature scaling minimizes ECE at T = 3 — you can evaluate a model honestly and trust its probabilities.

107. Fill in: value (verified) for The full program

Comparison

Comparison matrix

From The full program: refill the value (verified) column from what you know. The rest of the table is as it appeared.

printed linevalue (verified)
ROC-AUC: ... == sklearn0.875, True
AP:0.8542
imbalance ROC / PR0.736 / 0.111
ECE T=1 → T=30.1448 → 0.0176

108. Show it off

Concept

Out loud, slides closed: explain (1) what one ROC point represents and why the curve steps up on positives and right on negatives, (2) why ROC-AUC flatters an imbalanced model while PR-AUC doesn't, and (3) what temperature scaling changes — and what it provably leaves alone.

Stretch: reproduce AUC = 0.875 by all three routes (trapezoid, sklearn, concordant-pair count), then plot a reliability diagram before and after temperature scaling. Calibration returns in LLM sampling temperature (Week 37) and uncertainty estimation.

109. Connect it up: Lesson 35: ROC, PR Curves & Calibration

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — The confusion matrix · The ROC curve · AUC, three ways · The PR curve & imbalance · Calibration · One knob, two jobs. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

110. What you can do now

Recap

ideathe one thing to remember
ROC axesFPR on x, TPR on y — the good corner is top-left
AUCP(random positive outranks random negative)
imbalancereport PR-AUC vs the prevalence baseline, not ROC-AUC
calibrationconfidence ≠ accuracy; measure ECE
temperatureσ(z/T): T>1 softens, same argmax, one knob for LLM sampling too

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 35 (Week 12 — ROC/AUC & Calibration) — Barron · USAAIO Round 2 Preparation, 2026
  2. scikit-learn: roc_auc_score, average_precision_score, precision_recall_curve
  3. Guo et al., On Calibration of Modern Neural Networks (temperature scaling, ECE)
  4. Saito & Rehmsmeier, The Precision-Recall Plot Is More Informative than the ROC Plot on Imbalanced Datasets — PLOS ONE, 2015
  5. Every rate, curve point, AUC, AP, ECE, and temperature produced by real execution — numpy 2.2 + scikit-learn 1.9, verification run July 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108