USAAIO Lesson 32, from Week 11, fully worked. It derives the bias-variance decomposition MSE = Bias² + Variance + Noise line by line, adding and subtracting the mean and showing every cross term vanish, then verifies it numerically on a toy averaging estimator and measures it empirically across polynomial degree on a fixed sin(1.5x) running example, so that the U-shaped test error emerges from real numbers. It then builds cross-validation from the ground up: k-fold written from scratch with a full 5-fold trace and matched to sklearn, along with the stratified, leave-one-out, and time-series variants, the leakage trap, and nested cross-validation. Every snippet runs standalone, and every number came from real execution. The lesson runs to 62 slides.
Subject: Machine Learning · 114 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 32 · Week 11
Why a more powerful model can score worse. We derive MSE = Bias² + Variance + Noise line by line, watch the U-shaped test error emerge from real numbers, then build the cross-validation that estimates it honestly — from scratch, matched to sklearn.
Objectives
MSE = Bias² + Variance + Noise from E[(y − f̂)²] with every cross term shown to vanish — no hand-wavingk folds, and prove it matches sklearn.cross_val_scoreWarm-up
Discussion prompt
Before we open Lesson 32: Bias-Variance & Cross-Validation: without looking back, what was the main idea of t-SNE & UMAP, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
PCA's linear limitation and the manifold hypothesis, t-SNE (KL between pairwise-similarity distributions, t-distributed tails), UMAP's graph-based approach, perplexity/n_neighbors tuning, and why t-SNE embeddings can't be reused downstream. Compare PCA vs t-SNE on digits by 2D cluster separability.
Section
Part 1 of 9 — the setup
Concept
One dataset carries the whole lesson. The true signal is f(x) = sin(1.5x) on [0, 4]. We never see f; we only see noisy samples y = f(x) + ε, where ε is mean-zero noise with σ = 0.3.
\[ y = f(x) + \varepsilon, \qquad f(x) = \sin(1.5x), \qquad \varepsilon \sim \mathcal{N}(0,\, \sigma^2),\ \sigma = 0.3 \]
Our job: fit a polynomial f̂ to the samples that predicts well on new points. The only knob is the polynomial degree — that is our model complexity.
Counterexample
Discussion prompt
One dataset carries the whole lesson. The true signal is f(x) = sin(1.5x) on [0, 4]. We never see f; we only see noisy samples y = f(x) + ε, where ε is mean-zero noise with σ = 0.3.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Our job: fit a polynomial f̂ to the samples that predicts well on new points. The only knob is the polynomial degree — that is our model complexity.
Estimation
Predict first
Take five fixed x values, read off the true signal f(x), then add one draw of noise to get the y we actually observe. Runnable as written:
Commit before you compute: what does One sample of the running data come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: y sits near f but each point is nudged by ε
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The noise vector is [0.529, 0.120, 0.294, 0.672, 0.560], so every observed y is its true f plus that nudge.
Worked example
Take five fixed x values, read off the true signal f(x), then add one draw of noise to get the y we actually observe. Runnable as written:
import numpy as np
np.random.seed(0)
x = np.array([0.5, 1.0, 2.0, 3.0, 3.5])
f = np.sin(1.5 * x) # true signal (unseen in practice)
noise = 0.3 * np.random.randn(5)
y = f + noise # what we actually observe
print(np.round(f, 4))
print(np.round(y, 4))y sits near f but each point is nudged by ε
Why: The noise vector is [0.529, 0.120, 0.294, 0.672, 0.560], so every observed y is its true f plus that nudge. No model can undo ε — that is the irreducible part.
| x | f(x) = sin(1.5x) | ε (noise) | y = f + ε |
|---|---|---|---|
| 0.5 | 0.6816 | +0.5292 | 1.2109 |
| 1.0 | 0.9975 | +0.1200 | 1.1175 |
| 2.0 | 0.1411 | +0.2936 | 0.4347 |
| 3.0 | -0.9775 | +0.6723 | -0.3053 |
| 3.5 | -0.8589 | +0.5603 | -0.2987 |
Comparison
Comparison matrix
From One sample of the running data: refill the f(x) = sin(1.5x) column from what you know. The rest of the table is as it appeared.
| x | f(x) = sin(1.5x) | ε (noise) | y = f + ε |
|---|---|---|---|
| 0.5 | 0.6816 | +0.5292 | 1.2109 |
| 1.0 | 0.9975 | +0.1200 | 1.1175 |
| 2.0 | 0.1411 | +0.2936 | 0.4347 |
| 3.0 | -0.9775 | +0.6723 | -0.3053 |
| 3.5 | -0.8589 | +0.5603 | -0.2987 |
Intuition
Here is the mental shift the whole lesson rests on. Because the training y carry random noise, a different training sample would give a different fitted model f̂. So f̂ is not one fixed function — it is a random object that changes every time you resample the data.
That randomness is why we cannot judge a model by one fit. We must ask two separate questions: on average, is f̂ aimed at the truth? And how much does it jump around from one dataset to the next?
Those two questions are exactly bias and variance. Everything else today is making them precise.
Concept
bias — How far the AVERAGE model f̄ = E[f̂] sits from the truth f. A too-simple model (a straight line for a curve) is aimed wrong no matter how much data you give it — that is high bias, or underfitting.
variance — How much f̂ wobbles around its own average f̄ as the training set changes. A too-flexible model chases the noise, so each dataset yields a wildly different fit — that is high variance, or overfitting.
noise (irreducible error) — The variance σ² of ε itself. Even the perfect model f can't predict ε, so σ² is a floor no model can beat.
Definition probe
Sort into buckets
Every line below is part of the definition of bias or of variance — one or the other, never both. Put each where it belongs.
Picture it
Figure (svg): Two dartboards. Left: darts tightly clustered but far from center (low variance, high bias). Right: darts spread widely but centered on the bullseye on average (high variance, low bias).
Discussion prompt
Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.
Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.
Answer:
Think of each trained model as one throw at a dartboard whose bullseye is the truth f. Retraining on a fresh dataset is another throw.
Intuition
Think of each trained model as one throw at a dartboard whose bullseye is the truth f. Retraining on a fresh dataset is another throw.
Figure (svg): Two dartboards. Left: darts tightly clustered but far from center (low variance, high bias). Right: darts spread widely but centered on the bullseye on average (high variance, low bias).
Analogy
Discussion prompt
Explain The dartboard picture by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Think of each trained model as one throw at a dartboard whose bullseye is the truth f. Retraining on a fresh dataset is another throw.
Section
Part 2 of 9 — every step
Concept
Fix one test input x. Its true label is the random y = f + ε. Our model, trained on a random dataset, predicts f̂. The expected squared error at x averages over BOTH sources of randomness: the noise ε AND the training set (which is what makes f̂ random).
\[ \operatorname{Err}(x) = \mathbb{E}\!\left[(y - \hat f)^2\right], \qquad y = f + \varepsilon \]
We will rewrite this single expectation as three clean pieces. The trick is one line: add and subtract the average model f̄ = E[f̂].
Concept
E[ε] = 0 — the noise is mean-zero (no systematic offset in the labels).Var(ε) = σ² — the noise has a fixed variance, and E[ε²] = σ².ε is independent of the training data, hence independent of f̂.Also write f̄ = E[f̂] for the average prediction over training sets. That is all we need — the rest is algebra.
Explain it
Discussion prompt
Explain The three assumptions we use to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Also write f̄ = E[f̂] for the average prediction over training sets. That is all we need — the rest is algebra.
Concept
Before the algebra, pin down the two quantities precisely so the derivation lands on named objects, not vibes. Bias is the gap between the average model and the truth; variance is the mean-square wobble of the model about its own average:
\[ \operatorname{Bias}(\hat f) = \bar f - f, \qquad \operatorname{Var}(\hat f) = \mathbb{E}\!\left[(\hat f - \bar f)^2\right] \]
Both are averages OVER training sets, at a fixed x
Why: Neither can be computed from one fit — they describe the distribution of f̂ as the data is resampled. That is exactly why Part 3 estimates them by simulating many retrainings.
Translation
\( \operatorname{Bias}(\hat f) = \bar f - f, \qquad \operatorname{Var}(\hat f) = \mathbb{E}\!\left[(\hat f - \bar f)^2\right] \)
Draw it
Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.
Ranking
Put in order
Put the moves of Step 1 — split off the noise into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Replace y by f + ε inside the square, then group the deterministic (f − f̂) apart from the random ε.
Worked example
Substitute y = f + ε and regroup
Why: Replace y by f + ε inside the square, then group the deterministic (f − f̂) apart from the random ε.
\[ \mathbb{E}\!\left[(y - \hat f)^2\right] = \mathbb{E}\!\left[\big((f - \hat f) + \varepsilon\big)^2\right] \]
Expand the square
Why: (a + b)² = a² + 2ab + b² with a = f − f̂ and b = ε.
\[ = \mathbb{E}\!\left[(f - \hat f)^2\right] + 2\,\mathbb{E}\!\left[(f - \hat f)\,\varepsilon\right] + \mathbb{E}\!\left[\varepsilon^2\right] \]
The middle term is 0
Why: ε is independent of f̂, so E[(f − f̂)ε] = E[f − f̂]·E[ε] = E[f − f̂]·0 = 0. The last term is E[ε²] = σ².
\[ = \underbrace{\mathbb{E}\!\left[(f - \hat f)^2\right]}_{\text{model error}} + \underbrace{\sigma^2}_{\text{noise}} \]
Notation
Annotate
From Step 1 — split off the noise — read this one piece at a time. What is each part doing?
On: \( \mathbb{E}\!\left[(y - \hat f)^2\right] = \mathbb{E}\!\left[\big((f - \hat f) + \varepsilon\big)^2\right] \)
Missing information
Discussion prompt
Now crack open the model-error term E[(f − f̂)²]. Insert f̄ = E[f̂] by adding and subtracting it — the move that separates bias from variance:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
f − f̂ = (f − f̄) + (f̄ − f̂). The first piece is a fixed number (bias), the second is the random wobble of f̂ about its mean.
Worked example
Now crack open the model-error term E[(f − f̂)²]. Insert f̄ = E[f̂] by adding and subtracting it — the move that separates bias from variance:
Insert ± f̄
Why: f − f̂ = (f − f̄) + (f̄ − f̂). The first piece is a fixed number (bias), the second is the random wobble of f̂ about its mean.
\[ \mathbb{E}\!\left[(f - \hat f)^2\right] = \mathbb{E}\!\left[\big((f - \bar f) + (\bar f - \hat f)\big)^2\right] \]
Expand the square again
Why: Same (a + b)² rule, now with a = f − f̄ (constant) and b = f̄ − f̂ (random, mean 0).
\[ = \mathbb{E}\!\left[(f - \bar f)^2\right] + 2\,\mathbb{E}\!\left[(f - \bar f)(\bar f - \hat f)\right] + \mathbb{E}\!\left[(\bar f - \hat f)^2\right] \]
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Expand the square again
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Now crack open the model-error term E[(f − f̂)²]. Insert f̄ = E[f̂] by adding and subtracting it — the move that separates bias from variance:
Ranking
Put in order
Put the moves of Step 3 — the last cross term also vanishes into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. (f − f̄) is not random — it has no dependence on the training set — so it factors out of the expectation.
Worked example
Pull the constant out of the cross term
Why: (f − f̄) is not random — it has no dependence on the training set — so it factors out of the expectation.
\[ 2\,\mathbb{E}\!\left[(f - \bar f)(\bar f - \hat f)\right] = 2\,(f - \bar f)\,\mathbb{E}\!\left[\bar f - \hat f\right] \]
The leftover expectation is 0
Why: E[f̄ − f̂] = f̄ − E[f̂] = f̄ − f̄ = 0, because f̄ is by definition the mean of f̂. So the whole cross term dies.
\[ = 2\,(f - \bar f)\cdot 0 = 0 \]
Name the two survivors
Why: E[(f − f̄)²] = (f − f̄)² is the squared bias (a constant); E[(f̄ − f̂)²] is the variance of f̂ by definition.
\[ \mathbb{E}\!\left[(f - \hat f)^2\right] = \underbrace{(f - \bar f)^2}_{\text{Bias}^2} + \underbrace{\mathbb{E}\!\left[(\bar f - \hat f)^2\right]}_{\text{Variance}} \]
Notation
Annotate
From Step 3 — the last cross term also vanishes — read this one piece at a time. What is each part doing?
On: \( \mathbb{E}\!\left[(f - \hat f)^2\right] = \underbrace{(f - \bar f)^2}_{\text{Bias}^2} + \underbrace{\mathbb{E}\!\left[(\bar f - \hat f)^2\right]}_{\text{Variance}} \)
Worked example
Put Step 1 (noise split off) together with Step 3 (bias/variance split). Every cross term was shown to be exactly zero, so nothing is hidden:
\[ \boxed{\;\mathbb{E}\!\left[(y - \hat f)^2\right] = \underbrace{(f - \bar f)^2}_{\text{Bias}^2} + \underbrace{\mathbb{E}\!\left[(\hat f - \bar f)^2\right]}_{\text{Variance}} + \underbrace{\sigma^2}_{\text{Noise}}\;} \]
Read the three terms
Why: Bias² you shrink with a MORE flexible model; Variance you shrink with a LESS flexible model (or more data). σ² you cannot shrink at all. The tension between the first two is the whole game.
Anomaly
Predict first
A student writes this, and it looks reasonable:
After expanding ((f − f̄) + (f̄ − f̂))², keep the cross term 2(f − f̄)(f̄ − f̂) — it looks like it should contribute a Bias×√Variance piece.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: This forgets that we take the EXPECTATION.
Take the expectation of the cross term. (f − f̄) is constant and factors out; E[f̄ − f̂] = 0.
Why: This forgets that we take the EXPECTATION. E[f̄ − f̂] = 0, so the cross term is 0 in expectation — you can't leave a random f̂ sitting in a final answer that is supposed to be an expectation.
Trap
After expanding ((f − f̄) + (f̄ − f̂))², keep the cross term 2(f − f̄)(f̄ − f̂) — it looks like it should contribute a Bias×√Variance piece.
MSE = Bias² + Variance + 2·Bias·(f̄ − f̂) + σ² ✗
Why: This forgets that we take the EXPECTATION. E[f̄ − f̂] = 0, so the cross term is 0 in expectation — you can't leave a random f̂ sitting in a final answer that is supposed to be an expectation.
Take the expectation of the cross term. (f − f̄) is constant and factors out; E[f̄ − f̂] = 0.
MSE = Bias² + Variance + σ² ✓
Why: Exactly three terms. The cross term vanishes precisely because f̄ is DEFINED as E[f̂] — that definition is what makes the decomposition clean. Always finish the expectation before reading off terms.
Section
Part 3 of 9 — the identity is real
Intuition
Before the messy polynomials, pin the identity down on a case simple enough to check by hand. The truth is a constant f = μ = 2, and the noise has σ = 1.
Our estimator: draw n = 3 noisy points and predict their average. By symmetry this estimator is unbiased — on average it hits μ — so Bias² = 0. And the variance of an average of 3 independent draws is σ²/3 = 1/3. So the identity predicts MSE = 0 + 1/3 + 1 = 1.333.
Pattern
Predict first
The table runs: Bias² | 0 | 0.0000 · Variance | σ²/3 = 0.3333 | 0.3342 · Noise σ² | 1 | 1.0000 · sum | 1.3333 | 1.3342
In The unbiased averaging estimator, given the rows so far: what is the next one — the row where term is raw MSE?
Correct: raw MSE | — | 1.3358
| term | predicted | measured |
|---|---|---|
| Bias² | 0 | 0.0000 |
| Variance | σ²/3 = 0.3333 | 0.3342 |
| Noise σ² | 1 | 1.0000 |
| sum | 1.3333 | 1.3342 |
| raw MSE | — | 1.3358 |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. The measured variance 0.3342 matches the theoretical σ²/3 = 0.3333.
Worked example
Simulate 200,000 retrainings. Each retraining draws a fresh 3-point dataset and predicts its mean; we also draw a fresh test label each trial. Measure bias², variance, and the raw MSE, then compare to 0 + 1/3 + 1:
import numpy as np
np.random.seed(0)
mu, sigma, n = 2.0, 1.0, 3
def fhat():
return (mu + sigma * np.random.randn(n)).mean() # predict the mean
T = 200000
preds = np.array([fhat() for _ in range(T)])
ynew = mu + sigma * np.random.randn(T) # fresh test label
mse = np.mean((ynew - preds) ** 2)
bias2 = (preds.mean() - mu) ** 2
var = preds.var()
print(round(bias2, 4), round(var, 4), sigma**2)
print('sum =', round(bias2 + var + sigma**2, 4))
print('mse =', round(mse, 4))bias² = 0.0, variance = 0.334 ≈ σ²/3, σ² = 1
Why: The measured variance 0.3342 matches the theoretical σ²/3 = 0.3333. The estimator is unbiased so bias² rounds to 0.
| term | predicted | measured |
|---|---|---|
| Bias² | 0 | 0.0000 |
| Variance | σ²/3 = 0.3333 | 0.3342 |
| Noise σ² | 1 | 1.0000 |
| sum | 1.3333 | 1.3342 |
| raw MSE | — | 1.3358 |
Trade off
Comparison matrix
From The unbiased averaging estimator: every row here is a choice with a cost. Fill the predicted column, then say which row you would actually pick and what you give up for it.
| term | predicted | measured |
|---|---|---|
| Bias² | 0 | 0.0000 |
| Variance | σ²/3 = 0.3333 | 0.3342 |
| Noise σ² | 1 | 1.0000 |
| sum | 1.3333 | 1.3342 |
| raw MSE | — | 1.3358 |
Fill the middle
Fill in the blanks
From A biased estimator, same truth — one line has had its right-hand side removed. Put it back.
import numpy as np
np.random.seed(1)
mu, sigma = 2.0, 1.0
c = 1.5 # constant predictor, ignores the data
T = 200000
ynew = **mu + sigma * np.random.randn(T)
mse = np.mean((ynew - c) 2)
bias2 = (c - mu) 2
print('bias2 =', bias2, ' var = 0 sigma2 =', sigma2)
print('sum =', bias2 + 0 + sigma**2)
print('mse =', round(mse, 4))
Why: ynew is what everything below it consumes, so the wrong expression here fails later and somewhere else. Trading data for a fixed guess removed all variance but injected bias.
Worked example
Now break the aim on purpose: ignore the data and always predict the constant c = 1.5. The truth is still μ = 2, so Bias = 1.5 − 2 = −0.5 and Bias² = 0.25. Variance is 0 (the prediction never changes). Identity predicts 0.25 + 0 + 1 = 1.25:
import numpy as np
np.random.seed(1)
mu, sigma = 2.0, 1.0
c = 1.5 # constant predictor, ignores the data
T = 200000
ynew = mu + sigma * np.random.randn(T)
mse = np.mean((ynew - c) ** 2)
bias2 = (c - mu) ** 2
print('bias2 =', bias2, ' var = 0 sigma2 =', sigma**2)
print('sum =', bias2 + 0 + sigma**2)
print('mse =', round(mse, 4))bias² = 0.25, variance = 0, sum = 1.25 = measured MSE 1.2498
Why: Trading data for a fixed guess removed all variance but injected bias. The identity holds term for term — this is bias and variance as a genuine tug-of-war.
| estimator | Bias² | Variance | σ² | predicted MSE | measured MSE |
|---|---|---|---|---|---|
| average of 3 | 0.00 | 0.33 | 1.00 | 1.33 | 1.336 |
| constant 1.5 | 0.25 | 0.00 | 1.00 | 1.25 | 1.250 |
Sorting
Sort into buckets
These are the pieces of Lesson 32: Bias-Variance & Cross-Validation, out of order. Put each one back under the part of the lesson it belongs to.
Section
Part 4 of 9 — the U-curve appears
Concept
Back to f(x) = sin(1.5x). The degree of the fitted polynomial is our complexity knob. A low degree can't bend enough to follow the sine — high bias. A high degree bends through every noisy point — high variance.
To measure this we resample: draw many training sets, fit each, and look at how the predictions behave at a grid of test points. The average prediction reveals bias; the spread reveals variance.
Faded example
Fill in the blanks
The measurement machine, with the scaffolding fading: two lines are gone now — fill both.
import numpy as np
np.random.seed(0)
f = **lambda t: np.sin(1.5 * t)**
sigma = 0.3
xt = np.linspace(0.2, 3.8, 50); ft = f(xt) # test grid + truth
def preds_grid(deg, n=12):
x = np.random.uniform(0, 4, n)
y = f(x) + sigma * np.random.randn(n)
return np.polyval(np.polyfit(x, y, deg), xt)
for deg in [1, 2, 3, 5, 9]:
P = np.array([preds_grid(deg) for _ in range(500)])
bias2 = ((P.mean(0) - ft) 2).mean()
var = P.var(0).mean()
print(deg, round(bias2, 4), round(var, 4), round(bias2 + var + sigma2, 4))
Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. As degree rises the average fit hugs the sine (bias vanishes), but the wobble across datasets grows without bound once the polynomial can thread the noise.
Worked example
For a fixed degree, draw 500 training sets of 12 points each, fit a polynomial to each, and predict on a fixed test grid. P is a (500 × grid) array of predictions. Bias² is (mean over datasets − truth)²; variance is the spread across datasets — both averaged over the grid:
import numpy as np
np.random.seed(0)
f = lambda t: np.sin(1.5 * t)
sigma = 0.3
xt = np.linspace(0.2, 3.8, 50); ft = f(xt) # test grid + truth
def preds_grid(deg, n=12):
x = np.random.uniform(0, 4, n)
y = f(x) + sigma * np.random.randn(n)
return np.polyval(np.polyfit(x, y, deg), xt)
for deg in [1, 2, 3, 5, 9]:
P = np.array([preds_grid(deg) for _ in range(500)])
bias2 = ((P.mean(0) - ft) ** 2).mean()
var = P.var(0).mean()
print(deg, round(bias2, 4), round(var, 4), round(bias2 + var + sigma**2, 4))bias² collapses 0.132 → 0.002; variance explodes 0.049 → 79 million
Why: As degree rises the average fit hugs the sine (bias vanishes), but the wobble across datasets grows without bound once the polynomial can thread the noise. Total = bias² + variance + 0.09.
| degree | Bias² | Variance | total = B²+V+σ² |
|---|---|---|---|
| 1 (underfit) | 0.1320 | 0.0494 | 0.2714 |
| 2 | 0.1144 | 0.1781 | 0.3825 |
| 3 (balanced) | 0.0021 | 0.0920 | 0.1841 |
| 5 | 0.0140 | 14.31 | 14.41 |
| 9 (overfit) | 30359 | 79.1 million | 79.1 million |
Concept
Follow the total column: 0.271 → 0.383 → 0.184 → 14.4 → huge. It dips to a minimum at degree 3, then blows up. That dip is the famous U-shaped total error — you overshoot the sweet spot in either direction.
Figure (svg): Error on the vertical axis vs model complexity on the horizontal axis. Bias-squared curve falls from left to right; variance curve rises. Their sum is a U-shaped total-error curve with a minimum in the middle.
Worked example
Which term does more data attack? Hold the degree fixed at 5 (the high-variance regime) and grow the training-set size n = 12 → 30 → 100, remeasuring bias² and variance each time:
import numpy as np
np.random.seed(0)
f = lambda t: np.sin(1.5 * t)
sigma = 0.3
xt = np.linspace(0.2, 3.8, 50); ft = f(xt)
def preds(deg, n):
x = np.random.uniform(0, 4, n)
y = f(x) + sigma * np.random.randn(n)
return np.polyval(np.polyfit(x, y, deg), xt)
for n in [12, 30, 100]:
P = np.array([preds(5, n) for _ in range(400)])
bias2 = ((P.mean(0) - ft) ** 2).mean()
var = P.var(0).mean()
print(n, round(bias2, 4), round(var, 4))variance collapses 55.9 → 0.024 → 0.0045; bias² stays ~0
Why: More data stabilizes the fit (variance falls fast) but cannot fix an aim that was never wrong here — bias is already ~0 at degree 5. So high variance is the one you buy your way out of with data.
| n (train size) | Bias² (deg 5) | Variance (deg 5) |
|---|---|---|
| 12 | 0.1128 | 55.9255 |
| 30 | 0.0001 | 0.0238 |
| 100 | 0.0000 | 0.0045 |
Pattern
Step through it
Step through More data cures variance, not bias one row at a time. What is driving the change, and what would the row after the last one be?
Fill the middle
Fill in the blanks
From Underfit and overfit, one fit each — one line has had its right-hand side removed. Put it back.
import numpy as np
np.random.seed(0)
x = np.array([0.5, 1.0, 2.0, 3.0, 3.5])
f = np.sin(1.5 * x)
y = f + 0.3 * np.random.randn(5)
x0 = 2.0
p1 = np.polyfit(x, y, 1)
p4 = np.polyfit(x, y, 4)
print('deg1 pred @2 =', round(np.polyval(p1, x0), 4))
print('deg4 pred @2 =', round(np.polyval(p4, x0), 4))
print('truth f(2) =', round(np.sin(1.5 * 2.0), 4))
Why: x0 is what everything below it consumes, so the wrong expression here fails later and somewhere else. On THIS dataset both overshoot, but the point of bias/variance is the behavior ACROSS datasets, not one fit — which is exactly why the resampling machine on the previous slide is the honest measurement.
Worked example
Zoom to a single training realization and predict at one test point x0 = 2 (truth f(2) = 0.1411). Degree 1 can't reach; degree 4 is closer — but its wobble across datasets is what the table already flagged:
import numpy as np
np.random.seed(0)
x = np.array([0.5, 1.0, 2.0, 3.0, 3.5])
f = np.sin(1.5 * x)
y = f + 0.3 * np.random.randn(5)
x0 = 2.0
p1 = np.polyfit(x, y, 1)
p4 = np.polyfit(x, y, 4)
print('deg1 pred @2 =', round(np.polyval(p1, x0), 4))
print('deg4 pred @2 =', round(np.polyval(p4, x0), 4))
print('truth f(2) =', round(np.sin(1.5 * 2.0), 4))deg1 = 0.432, deg4 = 0.435, truth = 0.141
Why: On THIS dataset both overshoot, but the point of bias/variance is the behavior ACROSS datasets, not one fit — which is exactly why the resampling machine on the previous slide is the honest measurement.
| model | prediction @ x0=2 | truth f(2) |
|---|---|---|
| degree 1 | 0.4318 | 0.1411 |
| degree 4 | 0.4347 | 0.1411 |
| — across datasets — | deg1: steady & off | deg9: wild |
Section
Part 5 of 9 — train vs test
Estimation
Predict first
The catch: in practice you can't compute bias and variance — they need the unknown f. What you CAN compute is training error and held-out (test) error. Fit degrees 1..12 on one 15-point training set and score both:
Commit before you compute: what does Train error vs test error come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: train error falls monotonically; test error is U-shaped, min at degree 3
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Training error 0.379 → 0.014 always drops with complexity — it is useless for model choice.
Worked example
The catch: in practice you can't compute bias and variance — they need the unknown f. What you CAN compute is training error and held-out (test) error. Fit degrees 1..12 on one 15-point training set and score both:
import numpy as np
np.random.seed(1)
f = lambda t: np.sin(1.5 * t)
sigma = 0.3
xtr = np.sort(np.random.uniform(0, 4, 15))
ytr = f(xtr) + sigma * np.random.randn(15)
xte = np.sort(np.random.uniform(0, 4, 200))
yte = f(xte) + sigma * np.random.randn(200)
for deg in [1, 2, 3, 5, 9, 12]:
c = np.polyfit(xtr, ytr, deg)
tr = np.mean((np.polyval(c, xtr) - ytr) ** 2)
te = np.mean((np.polyval(c, xte) - yte) ** 2)
print(deg, round(tr, 4), round(te, 4))train error falls monotonically; test error is U-shaped, min at degree 3
Why: Training error 0.379 → 0.014 always drops with complexity — it is useless for model choice. Test error 0.237 → 84 million bottoms out at degree 3 (0.1455), then explodes. Only held-out error tells the truth.
| degree | train MSE | test MSE |
|---|---|---|
| 1 | 0.3787 | 0.2367 |
| 2 | 0.1376 | 0.5537 |
| 3 | 0.0769 | 0.1455 ← min |
| 5 | 0.0615 | 0.1183 |
| 9 | 0.0242 | 3074.3 |
| 12 | 0.0136 | 84,366,885 |
Pattern
Step through it
Step through Train error vs test error one row at a time. What is driving the change, and what would the row after the last one be?
Concept
Plot error against training-set size for a fixed model. High bias: train and validation error both plateau high and close together — more data won't help, you need a richer model.
High variance: a large gap — low train error, much higher validation error — that shrinks as you add data. More data (or regularization) is the cure. The gap between the two curves is your single best overfitting gauge.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Degree 12 gets training MSE down to 0.0136, far below degree 3's 0.077. Lower training error is a better fit, so ship degree 12.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Training error ALWAYS falls with complexity — it can be driven to zero by memorizing noise.
Choose complexity by held-out error, at the bottom of the U.
Why: Training error ALWAYS falls with complexity — it can be driven to zero by memorizing noise. Degree 12's TEST MSE is 84 million: it fits the 15 training points and detonates everywhere else. Training error can never justify a model choice.
Trap
Degree 12 gets training MSE down to 0.0136, far below degree 3's 0.077. Lower training error is a better fit, so ship degree 12.
Pick degree 12 on training error ✗
Why: Training error ALWAYS falls with complexity — it can be driven to zero by memorizing noise. Degree 12's TEST MSE is 84 million: it fits the 15 training points and detonates everywhere else. Training error can never justify a model choice.
Choose complexity by held-out error, at the bottom of the U.
Pick degree 3 on test error ✓
Why: Degree 3 minimizes TEST MSE (0.1455) — it sits at the balance of bias² and variance. When you have little data you don't waste a fixed test split; you use cross-validation (next) to estimate that held-out error.
Break the constraint
Discussion prompt
The rule this trap just fixed:
Choose complexity by held-out error, at the bottom of the U.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
Training error ALWAYS falls with complexity — it can be driven to zero by memorizing noise. Degree 12's TEST MSE is 84 million: it fits the 15 training points and detonates everywhere else. Training error can never justify a model choice.
Concept
The U-curve is the classical picture, and it governs the models on this exam: polynomials, ridge, trees, k-NN. But olympiad problems sometimes probe the modern twist — double descent.
In heavily over-parameterized models (huge neural nets), test error rises to a peak near the interpolation threshold — where the model exactly fits the training data — then, surprisingly, falls again as capacity grows further. So 'more complex is worse' holds until interpolation and can reverse beyond it.
The safe takeaway is unchanged: never choose complexity by training error — let held-out error, wherever its true minimum lies, decide.
Section
Part 6 of 9 — build k-fold
Concept
A single train/test split wastes data (the test points never train the model) and gives a noisy score — change which points land in the test set and the number jumps around, exactly the variance problem we've been studying.
k-fold cross-validation fixes both: chop the data into k equal folds, hold out each fold once as validation while training on the other k−1, and average the k scores. Every point is validated exactly once, so no data is wasted and the estimate has lower variance.
Picture it
Figure (svg): Five rows of five boxes. In row i, the i-th box is shaded as the validation fold and the other four are training folds; the shaded box moves one position right each row.
Discussion prompt
Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.
Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.
Answer:
Picture 5 folds. Each row is one CV iteration: one fold (shaded) is held out to score, the rest train. Rotate the held-out fold through all 5 positions, then average the 5 scores.
Intuition
Picture 5 folds. Each row is one CV iteration: one fold (shaded) is held out to score, the rest train. Rotate the held-out fold through all 5 positions, then average the 5 scores.
Figure (svg): Five rows of five boxes. In row i, the i-th box is shaded as the validation fold and the other four are training folds; the shaded box moves one position right each row.
Worked example
Start on the diabetes dataset (442 rows) so the folds and numbers are concrete. np.array_split cuts the index range into 5 nearly equal folds. Predict the sizes before printing:
import numpy as np
from sklearn.datasets import load_diabetes
X, y = load_diabetes(return_X_y=True)
folds = np.array_split(np.arange(len(X)), 5)
print('n =', len(X))
print('fold sizes =', [len(fld) for fld in folds])442 = 89 + 89 + 88 + 88 + 88
Why: np.array_split spreads the remainder: 442 = 5×88 + 2, so the first two folds get one extra row. No point is dropped or duplicated — the folds partition the data.
| quantity | value |
|---|---|
| n (rows) | 442 |
| number of folds | 5 |
| fold sizes | [89, 89, 88, 88, 88] |
Worked example
For each fold i: hold it out as test, concatenate the rest as train, fit Ridge(alpha=1.0), and score. Print every fold's train size, test size, and R² — the full trace:
import numpy as np
from sklearn.linear_model import Ridge
from sklearn.datasets import load_diabetes
X, y = load_diabetes(return_X_y=True)
folds = np.array_split(np.arange(len(X)), 5)
for i in range(5):
te = folds[i]
tr = np.concatenate([folds[j] for j in range(5) if j != i])
m = Ridge(alpha=1.0).fit(X[tr], y[tr])
print(i, len(tr), len(te), round(m.score(X[te], y[te]), 4))five R² scores: 0.322, 0.441, 0.422, 0.425, 0.442
Why: Fold 0 scores lowest (0.3217) — a single split landing there would badly misjudge the model. Averaging the five is what tames that luck-of-the-split variance.
| fold i | train size | test size | R² on held-out |
|---|---|---|---|
| 0 | 353 | 89 | 0.3217 |
| 1 | 353 | 89 | 0.4405 |
| 2 | 354 | 88 | 0.4221 |
| 3 | 354 | 88 | 0.4247 |
| 4 | 354 | 88 | 0.4420 |
Scale up
Step through it
Step through k-fold from scratch — the loop, per fold and watch the numbers move. Now imagine the input ten times bigger: which column is the one that stops this being practical?
Faded example
Fill in the blanks
k-fold from scratch — average & verify, with the scaffolding fading: two lines are gone now — fill both.
import numpy as np
from sklearn.linear_model import Ridge
from sklearn.datasets import load_diabetes
from sklearn.model_selection import cross_val_score, KFold
X, y = load_diabetes(return_X_y=True)
def kfold(model, X, y, k=5):
folds = np.array_split(np.arange(len(X)), k); s = []
for i in range(k):
te = folds[i]
tr = np.concatenate([folds[j] for j in range(k) if j != i])
model.fit(X[tr], y[tr]); s.append(model.score(X[te], y[te]))
return np.mean(s)
mine = kfold(Ridge(alpha=1.0), X, y)
skl = cross_val_score(Ridge(alpha=1.0), X, y, cv=KFold(5, shuffle=False)).mean()
print(round(mine, 4), round(skl, 4), np.isclose(mine, skl))
Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. Our from-scratch average equals sklearn's cross_val_score to the digit, and np.isclose returns True — because we reproduced KFold's contiguous, no-shuffle split exactly.
Worked example
Wrap the loop in a function that returns the mean of the 5 scores, then compare to sklearn.cross_val_score with KFold(5, shuffle=False) — the same no-shuffle split we built:
import numpy as np
from sklearn.linear_model import Ridge
from sklearn.datasets import load_diabetes
from sklearn.model_selection import cross_val_score, KFold
X, y = load_diabetes(return_X_y=True)
def kfold(model, X, y, k=5):
folds = np.array_split(np.arange(len(X)), k); s = []
for i in range(k):
te = folds[i]
tr = np.concatenate([folds[j] for j in range(k) if j != i])
model.fit(X[tr], y[tr]); s.append(model.score(X[te], y[te]))
return np.mean(s)
mine = kfold(Ridge(alpha=1.0), X, y)
skl = cross_val_score(Ridge(alpha=1.0), X, y, cv=KFold(5, shuffle=False)).mean()
print(round(mine, 4), round(skl, 4), np.isclose(mine, skl))mean(0.322, 0.441, 0.422, 0.425, 0.442) = 0.4102 = sklearn
Why: Our from-scratch average equals sklearn's cross_val_score to the digit, and np.isclose returns True — because we reproduced KFold's contiguous, no-shuffle split exactly.
| source | 5-fold mean R² | match |
|---|---|---|
| from scratch | 0.4102 | — |
| sklearn cross_val_score | 0.4102 | True |
| single split (fold 0 only) | 0.3217 | misleading |
Anomaly
Predict first
A student writes this, and it looks reasonable:
Compare your contiguous-split k-fold to cross_val_score with the default cv=5 (or shuffle=True) and expect the same number.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: A shuffled KFold assigns DIFFERENT rows to each fold, so its per-fold scores and average differ from your contiguous split.
Match the splitter exactly: contiguous np.array_split mirrors KFold(k, shuffle=False).
Why: A shuffled KFold assigns DIFFERENT rows to each fold, so its per-fold scores and average differ from your contiguous split. The numbers won't match and you'll think your code is wrong when it isn't.
Trap
Compare your contiguous-split k-fold to cross_val_score with the default cv=5 (or shuffle=True) and expect the same number.
mine (no shuffle) vs KFold(shuffle=True) ✗
Why: A shuffled KFold assigns DIFFERENT rows to each fold, so its per-fold scores and average differ from your contiguous split. The numbers won't match and you'll think your code is wrong when it isn't.
Match the splitter exactly: contiguous np.array_split mirrors KFold(k, shuffle=False).
mine vs KFold(5, shuffle=False) ✓
Why: Same rows in each fold → identical scores → identical average 0.4102, np.isclose True. When verifying a from-scratch algorithm, replicate the reference's data ordering before comparing outputs.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
f̂ to the samples that predicts well on new points. The only knob is the polynomial degree — that is our model complexity.; Think of each trained model as one throw at a dartboard whose bullseye is the truth f. Retraining on a fresh dataset is another throw.; Also write f̄ = E[f̂] for the average prediction over training sets. That is all we need — the rest is algebra.((f − f̄) + (f̄ − f̂))², keep the cross term 2(f − f̄)(f̄ − f̂) — it looks like it should contribute a Bias×√Variance piece.; Degree 12 gets training MSE down to 0.0136, far below degree 3's 0.077. Lower training error is a better fit, so ship degree 12.Section
Part 7 of 9 — pick the right split
Concept
| strategy | use when | key property |
|---|---|---|
| k-fold | the default | each point held out once |
| stratified k-fold | classification | preserves class ratios per fold |
| leave-one-out (LOO) | tiny datasets (k = n) | n fits; low bias, high variance, costly |
| time-series split | temporal data | train only on the PAST |
Plain k-fold assumes rows are exchangeable. When they are not — imbalanced classes, or time order — you must switch splitter or the estimate lies.
Comparison
Comparison matrix
From Four splitters, four situations: refill the use when column from what you know. The rest of the table is as it appeared.
| strategy | use when | key property |
|---|---|---|
| k-fold | the default | each point held out once |
| stratified k-fold | classification | preserves class ratios per fold |
| leave-one-out (LOO) | tiny datasets (k = n) | n fits; low bias, high variance, costly |
| time-series split | temporal data | train only on the PAST |
Fill the middle
Fill in the blanks
From Stratified preserves class balance — one line has had its right-hand side removed. Put it back.
import numpy as np
from sklearn.model_selection import KFold, StratifiedKFold
y = np.array([0]90 + [1]10) # 90/10 imbalance
np.random.seed(0)
y = y[np.random.permutation(len(y))]
def pos_rates(splitter):
return [round(y[te].mean(), 3) for _, te in splitter.split(np.zeros_like(y), y)]
print('plain ', pos_rates(KFold(5, shuffle=True, random_state=0)))
print('strat ', pos_rates(StratifiedKFold(5, shuffle=True, random_state=0)))
Why: y is what everything below it consumes, so the wrong expression here fails later and somewhere else. Plain k-fold's validation folds range from 5% to 15% positives — a fold with too few positives gives a garbage score.
Worked example
Take a 90/10 imbalanced label. Plain k-fold splits randomly, so some folds get too few positives; stratified k-fold forces every fold to the global 10% rate. Watch the per-fold positive rate:
import numpy as np
from sklearn.model_selection import KFold, StratifiedKFold
y = np.array([0]*90 + [1]*10) # 90/10 imbalance
np.random.seed(0)
y = y[np.random.permutation(len(y))]
def pos_rates(splitter):
return [round(y[te].mean(), 3) for _, te in splitter.split(np.zeros_like(y), y)]
print('plain ', pos_rates(KFold(5, shuffle=True, random_state=0)))
print('strat ', pos_rates(StratifiedKFold(5, shuffle=True, random_state=0)))plain wobbles 0.05–0.15; stratified is a flat 0.10 everywhere
Why: Plain k-fold's validation folds range from 5% to 15% positives — a fold with too few positives gives a garbage score. Stratified pins every fold to 10%, matching the whole dataset.
| splitter | per-fold positive rate | spread |
|---|---|---|
| KFold | [0.10, 0.10, 0.15, 0.10, 0.05] | 0.05 – 0.15 |
| StratifiedKFold | [0.10, 0.10, 0.10, 0.10, 0.10] | flat 0.10 |
| global rate | 0.10 | target |
Pattern
Predict first
The table runs: 1 | [0, 1] | [2, 3] · 2 | [0, 1, 2, 3] | [4, 5] · 3 | [0..5] | [6, 7]
In Time-series: never train on the future, given the rows so far: what is the next one — the row where split is 4?
Correct: 4 | [0..7] | [8, 9]
| split | train indices | test indices |
|---|---|---|
| 1 | [0, 1] | [2, 3] |
| 2 | [0, 1, 2, 3] | [4, 5] |
| 3 | [0..5] | [6, 7] |
| 4 | [0..7] | [8, 9] |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. Split 1 trains on [0,1] tests [2,3]; the window grows to train on [0..7] tests [8,9].
Worked example
For ordered data, a random fold would let the model peek at future points to predict the past — leakage. TimeSeriesSplit only ever trains on earlier indices and tests on later ones, growing the training window:
import numpy as np
from sklearn.model_selection import TimeSeriesSplit
idx = np.arange(10)
for tr, te in TimeSeriesSplit(n_splits=4).split(idx):
print('train', list(tr), 'test', list(te))every test block sits strictly AFTER its training block
Why: Split 1 trains on [0,1] tests [2,3]; the window grows to train on [0..7] tests [8,9]. Training indices are always a prefix — the model never sees the future it must predict.
| split | train indices | test indices |
|---|---|---|
| 1 | [0, 1] | [2, 3] |
| 2 | [0, 1, 2, 3] | [4, 5] |
| 3 | [0..5] | [6, 7] |
| 4 | [0..7] | [8, 9] |
Pattern
Step through it
Step through Time-series: never train on the future one row at a time. What is driving the change, and what would the row after the last one be?
Missing information
Discussion prompt
The extreme of k-fold: set k = n, so each fold is a single point. That is leave-one-out (LOO) — n fits, each training on all but one point. It wrings the most training data out of a tiny dataset, at the cost of n model fits. Compare its MSE to 5-fold on diabetes:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
LOO trains on n−1 = 441 points each time, so it is nearly unbiased for the full-data model — a slightly lower MSE estimate. But it runs 442 fits (vs 5) and its estimate has higher variance because the n test sets overlap heavily.
Worked example
The extreme of k-fold: set k = n, so each fold is a single point. That is leave-one-out (LOO) — n fits, each training on all but one point. It wrings the most training data out of a tiny dataset, at the cost of n model fits. Compare its MSE to 5-fold on diabetes:
import numpy as np
from sklearn.linear_model import Ridge
from sklearn.datasets import load_diabetes
from sklearn.model_selection import cross_val_score, LeaveOneOut, KFold
X, y = load_diabetes(return_X_y=True)
loo = cross_val_score(Ridge(1.0), X, y, cv=LeaveOneOut(),
scoring='neg_mean_squared_error')
kf = cross_val_score(Ridge(1.0), X, y, cv=KFold(5, shuffle=False),
scoring='neg_mean_squared_error')
print('LOO folds =', len(loo))
print('LOO MSE =', round(-loo.mean(), 2))
print('5-fold MSE=', round(-kf.mean(), 2))442 folds; LOO MSE 3327.66 vs 5-fold 3420.32
Why: LOO trains on n−1 = 441 points each time, so it is nearly unbiased for the full-data model — a slightly lower MSE estimate. But it runs 442 fits (vs 5) and its estimate has higher variance because the n test sets overlap heavily.
| strategy | folds (fits) | mean MSE | trade |
|---|---|---|---|
| 5-fold | 5 | 3420.32 | cheap, a little biased |
| leave-one-out | 442 | 3327.66 | costly, low bias / high var |
| rule of thumb | 5 or 10 | — | LOO only for tiny n |
Section
Part 8 of 9 — nested CV
Concept
If you try many hyperparameters and report the best score on the same data you used to pick them, that score is optimistically biased — you fit the evaluation set through model selection. This is a form of data leakage.
The cure is separation: tune on inner folds, evaluate on outer folds the tuning never touched. That is nested CV — CV inside CV.
Estimation
Predict first
Grid over Ridge alpha. The cheating protocol picks the alpha that maximizes a single CV and reports that same CV. Nested CV tunes alpha in an inner loop and reports the untouched outer loop. Compare the two estimates:
Commit before you compute: what does Cheating vs nested CV, measured come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: cheat 0.4894 > nested 0.4858
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The cheating score is higher because picking alpha on the same folds it's scored on cherry-picks favorable noise.
Worked example
Grid over Ridge alpha. The cheating protocol picks the alpha that maximizes a single CV and reports that same CV. Nested CV tunes alpha in an inner loop and reports the untouched outer loop. Compare the two estimates:
import numpy as np
from sklearn.linear_model import Ridge
from sklearn.datasets import load_diabetes
from sklearn.model_selection import GridSearchCV, cross_val_score, KFold
X, y = load_diabetes(return_X_y=True)
grid = {'alpha': [0.01, 0.1, 1.0, 10.0, 100.0]}
inner = KFold(5, shuffle=True, random_state=0)
outer = KFold(5, shuffle=True, random_state=0)
gs = GridSearchCV(Ridge(), grid, cv=inner)
nested = cross_val_score(gs, X, y, cv=outer).mean()
best = max(grid['alpha'], key=lambda a: cross_val_score(Ridge(alpha=a), X, y, cv=outer).mean())
cheat = cross_val_score(Ridge(alpha=best), X, y, cv=outer).mean()
print('cheat:', best, round(cheat, 4))
print('nested:', round(nested, 4))cheat 0.4894 > nested 0.4858
Why: The cheating score is higher because picking alpha on the same folds it's scored on cherry-picks favorable noise. Nested CV's outer folds never see the tuning, so 0.4858 is the honest generalization estimate.
| protocol | chosen alpha | reported score | honest? |
|---|---|---|---|
| tune & report on same CV | 0.01 | 0.4894 | no — optimistic |
| nested CV | inner-selected | 0.4858 | yes |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
cheat 0.4894 > nested 0.4858
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Grid over Ridge alpha. The cheating protocol picks the alpha that maximizes a single CV and reports that same CV. Nested CV tunes alpha in an inner loop and reports the untouched outer loop. Compare the two estimates:
Concept
Whatever the search, it belongs in the inner loop of nested CV. The outer loop's only job is to report a number you can trust.
Explain it
Discussion prompt
Explain Searching the hyperparameter space to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Whatever the search, it belongs in the inner loop of nested CV. The outer loop's only job is to report a number you can trust.
Concept
The reason tuning matters is that a hyperparameter like ridge's λ slides you along the same bias-variance curve. Larger λ shrinks the weights toward zero — less flexibility, less variance, more bias. Smaller λ does the reverse.
\[ \lambda \uparrow \;\Longrightarrow\; \text{variance} \downarrow,\ \text{bias} \uparrow \qquad\qquad \lambda \downarrow \;\Longrightarrow\; \text{variance} \uparrow,\ \text{bias} \downarrow \]
So choosing λ by CV is just finding the bottom of the U in a different coordinate. Polynomial degree, tree depth, k in k-NN, λ in ridge — every complexity knob is a bias-variance dial you set with held-out error.
Section
Part 9 of 9
Ranking
Put in order
These are the steps of The bias-variance / CV toolkit, scrambled. Put them back in order before the next slide shows you.
Bias² + Variance + Noise — the cross terms vanish in expectationnWhy: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
Bias² + Variance + Noise — the cross terms vanish in expectationnEdge cases
Discussion prompt
The bias-variance / CV toolkit works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
Bias² + Variance + Noise — the cross terms vanish in expectationnElimination
Eliminate the wrong options
A model has very low training error but high test error. Which term dominates?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Low train / high test is the signature of high variance: the model fit the training noise and fails to generalize. Degree 9–12 in the demo did exactly this — tiny train MSE, test MSE in the millions.
Check
Match the symptom to the term. Think about train vs test.
Check your understanding
A model has very low training error but high test error. Which term dominates?
Answer: A
Why: Low train / high test is the signature of high variance: the model fit the training noise and fails to generalize. Degree 9–12 in the demo did exactly this — tiny train MSE, test MSE in the millions.
Prediction
Predict first
In the derivation, why does the cross term 2·E[(f − f̄)(f̄ − f̂)] equal zero?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: (f − f̄) is constant and E[f̄ − f̂] = 0 since f̄ = E[f̂]
Why: (f − f̄) has no dependence on the training set, so it factors out of the expectation, leaving 2(f − f̄)·E[f̄ − f̂]. Since f̄ is DEFINED as E[f̂], that expectation is f̄ − f̄ = 0.
Check
Recall why the decomposition has exactly three terms.
Check your understanding
In the derivation, why does the cross term 2·E[(f − f̄)(f̄ − f̂)] equal zero?
Answer: A
Why: (f − f̄) has no dependence on the training set, so it factors out of the expectation, leaving 2(f − f̄)·E[f̄ − f̂]. Since f̄ is DEFINED as E[f̂], that expectation is f̄ − f̄ = 0.
Prediction
Predict first
In 5-fold cross-validation, each data point is used for validation:
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Exactly once (and for training in the other 4 folds)
Why: The data partitions into 5 folds; each fold is the validation set once and part of training the other four times. Averaging the 5 held-out scores (0.322, 0.441, 0.422, 0.425, 0.442 → 0.4102) gives the low-variance estimate.
Check
How does 5-fold use each data point?
Check your understanding
In 5-fold cross-validation, each data point is used for validation:
Answer: A
Why: The data partitions into 5 folds; each fold is the validation set once and part of training the other four times. Averaging the 5 held-out scores (0.322, 0.441, 0.422, 0.425, 0.442 → 0.4102) gives the low-variance estimate.
Elimination
Eliminate the wrong options
Nested cross-validation uses an inner and an outer loop in order to:
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: The inner loop selects hyperparameters on each outer training split; the outer test folds never participate in tuning, so the final score is unbiased. In the demo, the cheating estimate 0.4894 was inflated above the honest nested 0.4858.
Check
Why two loops instead of one?
Check your understanding
Nested cross-validation uses an inner and an outer loop in order to:
Answer: A
Why: The inner loop selects hyperparameters on each outer training split; the outer test folds never participate in tuning, so the final score is unbiased. In the demo, the cheating estimate 0.4894 was inflated above the honest nested 0.4858.
Concept
Bring both halves together. Using only NumPy, build k-fold CV on the sin(1.5x) data and use it to pick the polynomial degree — recovering the sweet spot the bias-variance table predicted, without ever peeking at the true f.
| # | requirement | tool |
|---|---|---|
| 1 | generate the noisy sin data | np.random + np.sin |
| 2 | k-fold CV MSE for one degree | np.array_split, np.polyfit |
| 3 | sweep degrees, pick the min-CV degree | np.polyval, min |
Build rules: type every line yourself, seed the RNG so your numbers match, and read errors instead of deleting them — a shape mismatch usually means a fold index went wrong.
Analogy
Discussion prompt
Explain Project: choose the polynomial degree by CV by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Build rules: type every line yourself, seed the RNG so your numbers match, and read errors instead of deleting them — a shape mismatch usually means a fold index went wrong.
Worked example
Your turn: generate 30 noisy points from f(x) = sin(1.5x) on [0, 4] with σ = 0.3 and seed = 2. Predict x's shape and y's rough range before printing.
Hint: np.sort(np.random.uniform(0, 4, 30)) for x; add 0.3 * np.random.randn(30) to sin(1.5x) for y.
import numpy as np
np.random.seed(2)
f = lambda t: np.sin(1.5 * t)
x = np.sort(np.random.uniform(0, 4, 30))
y = f(x) + 0.3 * np.random.randn(30)
print(x.shape, round(y.min(), 3), round(y.max(), 3))| check | value |
|---|---|
| x.shape | (30,) |
| y.min() | -1.124 |
| y.max() | 1.302 |
Pattern
Step through it
Step through Milestone 1 — the data one row at a time. What is driving the change, and what would the row after the last one be?
Worked example
Your turn: write cv_mse(deg, k=5) — split indices into k folds, fit np.polyfit on k−1, score MSE on the held-out fold, average. Predict roughly the degree-3 CV MSE (recall test error bottomed near degree 3).
Hint: folds = np.array_split(np.arange(len(x)), k); tr = np.concatenate([folds[j] for j in range(k) if j != i]); error is np.mean((np.polyval(c, x[te]) - y[te])**2).
import numpy as np
np.random.seed(2)
f = lambda t: np.sin(1.5 * t)
x = np.sort(np.random.uniform(0, 4, 30))
y = f(x) + 0.3 * np.random.randn(30)
def cv_mse(deg, k=5):
folds = np.array_split(np.arange(len(x)), k); errs = []
for i in range(k):
te = folds[i]
tr = np.concatenate([folds[j] for j in range(k) if j != i])
c = np.polyfit(x[tr], y[tr], deg)
errs.append(np.mean((np.polyval(c, x[te]) - y[te]) ** 2))
return np.mean(errs)
print(round(cv_mse(3), 4))| quantity | value |
|---|---|
| cv_mse(3) | 0.3597 |
| k | 5 |
| fits per call | 5 |
Trade off
Comparison matrix
From Milestone 2 — CV MSE for one degree: every row here is a choice with a cost. Fill the value column, then say which row you would actually pick and what you give up for it.
| quantity | value |
|---|---|
| cv_mse(3) | 0.3597 |
| k | 5 |
| fits per call | 5 |
Worked example
Your turn: run cv_mse over degrees 1, 2, 3, 5, 9 and pick the degree with the smallest CV MSE. Predict which one wins before you print.
Hint: build a dict {deg: cv_mse(deg)}, then min(scores, key=scores.get).
import numpy as np
np.random.seed(2)
f = lambda t: np.sin(1.5 * t)
x = np.sort(np.random.uniform(0, 4, 30))
y = f(x) + 0.3 * np.random.randn(30)
def cv_mse(deg, k=5):
folds = np.array_split(np.arange(len(x)), k); errs = []
for i in range(k):
te = folds[i]
tr = np.concatenate([folds[j] for j in range(k) if j != i])
c = np.polyfit(x[tr], y[tr], deg)
errs.append(np.mean((np.polyval(c, x[te]) - y[te]) ** 2))
return np.mean(errs)
scores = {d: round(cv_mse(d), 4) for d in [1, 2, 3, 5, 9]}
print(scores)
print('best degree =', min(scores, key=scores.get))| degree | CV MSE |
|---|---|
| 1 | 0.6838 |
| 2 | 1.5353 |
| 3 | 0.3597 ← min |
| 5 | 8.5814 |
| 9 | 20,759,668 |
Comparison
Comparison matrix
From Milestone 3 — sweep and pick: refill the CV MSE column from what you know. The rest of the table is as it appeared.
| degree | CV MSE |
|---|---|
| 1 | 0.6838 |
| 2 | 1.5353 |
| 3 | 0.3597 ← min |
| 5 | 8.5814 |
| 9 | 20,759,668 |
Concept
import numpy as np
np.random.seed(2)
f = lambda t: np.sin(1.5 * t)
x = np.sort(np.random.uniform(0, 4, 30)) # 1. noisy data
y = f(x) + 0.3 * np.random.randn(30)
def cv_mse(deg, k=5): # 2. k-fold CV MSE
folds = np.array_split(np.arange(len(x)), k); errs = []
for i in range(k):
te = folds[i]
tr = np.concatenate([folds[j] for j in range(k) if j != i])
c = np.polyfit(x[tr], y[tr], deg)
errs.append(np.mean((np.polyval(c, x[te]) - y[te]) ** 2))
return np.mean(errs)
scores = {d: round(cv_mse(d), 4) for d in [1, 2, 3, 5, 9]} # 3. sweep
print(scores)
print('best degree =', min(scores, key=scores.get))| printed line | value |
|---|---|
| scores | {1: 0.6838, 2: 1.5353, 3: 0.3597, 5: 8.5814, 9: 20759668.6189} |
| best degree = | 3 |
CV chose degree 3 — the same sweet spot the bias-variance table found — using only the noisy data. You picked model complexity honestly, without ever seeing the truth.
Counterexample
Discussion prompt
CV chose degree 3 — the same sweet spot the bias-variance table found — using only the noisy data. You picked model complexity honestly, without ever seeing the truth.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Concept
Slides closed, out loud: explain (1) the three terms of MSE = Bias² + Variance + Noise and why the two cross terms vanish, (2) why training error falls monotonically but test error is U-shaped, and (3) why nested CV's outer loop must never see the tuning.
Stretch: swap cv_mse to use StratifiedKFold on a classification target, then wrap the degree sweep in an outer loop to make it a true nested CV. Bias-variance underlies regularization (Week 12) and ensembles (Week 23) — bagging attacks variance, boosting attacks bias.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Two ways to be wrong · Derive the decomposition · Verify it numerically · Bias & variance vs complexity · Diagnose from curves · Cross-validation. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
MSE = Bias² + Variance + Noise from E[(y − f̂)²], showing both cross terms vanish — and verify it numerically (1.33 and 1.25)sin data)sklearn (0.4102) — with stratified / LOO / time-series variants0.4858) isn't inflated by leakage (0.4894)| idea | the one thing to remember |
|---|---|
| decomposition | MSE = Bias² + Variance + Noise; cross terms vanish |
| complexity | test error is U-shaped, not monotone — train error is useless for choice |
| k-fold | each point held out exactly once; average the folds |
| strategy | stratified for classes, time-series for time, LOO for tiny n |
| nested CV | tune inside, evaluate outside — never report the number you tuned on |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.