Lesson 32: Bias-Variance & Cross-Validation

USAAIO Lesson 32, from Week 11, fully worked. It derives the bias-variance decomposition MSE = Bias² + Variance + Noise line by line, adding and subtracting the mean and showing every cross term vanish, then verifies it numerically on a toy averaging estimator and measures it empirically across polynomial degree on a fixed sin(1.5x) running example, so that the U-shaped test error emerges from real numbers. It then builds cross-validation from the ground up: k-fold written from scratch with a full 5-fold trace and matched to sklearn, along with the stratified, leave-one-out, and time-series variants, the leakage trap, and nested cross-validation. Every snippet runs standalone, and every number came from real execution. The lesson runs to 62 slides.

Subject: Machine Learning · 114 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Bias-Variance & Cross-Validation

Title

USAAIO · Lesson 32 · Week 11

Why a more powerful model can score worse. We derive MSE = Bias² + Variance + Noise line by line, watch the U-shaped test error emerge from real numbers, then build the cross-validation that estimates it honestly — from scratch, matched to sklearn.

2. By the end of this lesson you can

Objectives

  1. Derive MSE = Bias² + Variance + Noise from E[(y − f̂)²] with every cross term shown to vanish — no hand-waving
  2. Verify the decomposition numerically on a toy estimator, then measure bias² and variance empirically across polynomial degree
  3. Explain the U-shaped test error: bias falls, variance rises, total dips in the middle — and diagnose underfit vs overfit from train/val curves
  4. Build k-fold cross-validation from scratch, trace all k folds, and prove it matches sklearn.cross_val_score
  5. Pick a CV strategy (stratified, LOO, time-series) and avoid leakage with nested CV

3. What survived from t-SNE & UMAP?

Warm-up

Discussion prompt

Before we open Lesson 32: Bias-Variance & Cross-Validation: without looking back, what was the main idea of t-SNE & UMAP, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

PCA's linear limitation and the manifold hypothesis, t-SNE (KL between pairwise-similarity distributions, t-distributed tails), UMAP's graph-based approach, perplexity/n_neighbors tuning, and why t-SNE embeddings can't be reused downstream. Compare PCA vs t-SNE on digits by 2D cluster separability.

4. Two ways to be wrong

Section

Part 1 of 9 — the setup

5. The running example: a noisy curve

Concept

One dataset carries the whole lesson. The true signal is f(x) = sin(1.5x) on [0, 4]. We never see f; we only see noisy samples y = f(x) + ε, where ε is mean-zero noise with σ = 0.3.

\[ y = f(x) + \varepsilon, \qquad f(x) = \sin(1.5x), \qquad \varepsilon \sim \mathcal{N}(0,\, \sigma^2),\ \sigma = 0.3 \]

Our job: fit a polynomial f̂ to the samples that predicts well on new points. The only knob is the polynomial degree — that is our model complexity.

6. Break it if you can: The running example: a noisy curve

Counterexample

Discussion prompt

One dataset carries the whole lesson. The true signal is f(x) = sin(1.5x) on [0, 4]. We never see f; we only see noisy samples y = f(x) + ε, where ε is mean-zero noise with σ = 0.3.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Our job: fit a polynomial f̂ to the samples that predicts well on new points. The only knob is the polynomial degree — that is our model complexity.

7. Guess the shape of the answer: One sample of the running data

Estimation

Predict first

Take five fixed x values, read off the true signal f(x), then add one draw of noise to get the y we actually observe. Runnable as written:

Commit before you compute: what does One sample of the running data come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: y sits near f but each point is nudged by ε

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The noise vector is [0.529, 0.120, 0.294, 0.672, 0.560], so every observed y is its true f plus that nudge.

8. One sample of the running data

Worked example

Take five fixed x values, read off the true signal f(x), then add one draw of noise to get the y we actually observe. Runnable as written:

import numpy as np
np.random.seed(0)
x = np.array([0.5, 1.0, 2.0, 3.0, 3.5])
f = np.sin(1.5 * x)             # true signal (unseen in practice)
noise = 0.3 * np.random.randn(5)
y = f + noise                   # what we actually observe
print(np.round(f, 4))
print(np.round(y, 4))

y sits near f but each point is nudged by ε

Why: The noise vector is [0.529, 0.120, 0.294, 0.672, 0.560], so every observed y is its true f plus that nudge. No model can undo ε — that is the irreducible part.

xf(x) = sin(1.5x)ε (noise)y = f + ε
0.50.6816+0.52921.2109
1.00.9975+0.12001.1175
2.00.1411+0.29360.4347
3.0-0.9775+0.6723-0.3053
3.5-0.8589+0.5603-0.2987

9. Fill in: f(x) = sin(1.5x) for One sample of the running data

Comparison

Comparison matrix

From One sample of the running data: refill the f(x) = sin(1.5x) column from what you know. The rest of the table is as it appeared.

xf(x) = sin(1.5x)ε (noise)y = f + ε
0.50.6816+0.52921.2109
1.00.9975+0.12001.1175
2.00.1411+0.29360.4347
3.0-0.9775+0.6723-0.3053
3.5-0.8589+0.5603-0.2987

10. The model is a random variable

Intuition

Here is the mental shift the whole lesson rests on. Because the training y carry random noise, a different training sample would give a different fitted model f̂. So f̂ is not one fixed function — it is a random object that changes every time you resample the data.

That randomness is why we cannot judge a model by one fit. We must ask two separate questions: on average, is f̂ aimed at the truth? And how much does it jump around from one dataset to the next?

Those two questions are exactly bias and variance. Everything else today is making them precise.

11. Bias and variance, in words

Concept

bias — How far the AVERAGE model f̄ = E[f̂] sits from the truth f. A too-simple model (a straight line for a curve) is aimed wrong no matter how much data you give it — that is high bias, or underfitting.

variance — How much f̂ wobbles around its own average f̄ as the training set changes. A too-flexible model chases the noise, so each dataset yields a wildly different fit — that is high variance, or overfitting.

noise (irreducible error) — The variance σ² of ε itself. Even the perfect model f can't predict ε, so σ² is a floor no model can beat.

12. Take the definitions apart: bias vs variance

Definition probe

Sort into buckets

Every line below is part of the definition of bias or of variance — one or the other, never both. Put each where it belongs.

bias
How far the AVERAGE model f̄ = E[f̂] sits from the truth f.; A too-simple model (a straight line for a curve) is aimed wrong no matter how much data you give it; that is high bias, or underfitting.
variance
How much f̂ wobbles around its own average f̄ as the training set changes.; A too-flexible model chases the noise, so each dataset yields a wildly different fit; that is high variance, or overfitting.
b1
How far the AVERAGE model f̄ = E[f̂] sits from the truth f. A too-simple model (a straight line for a curve) is aimed wrong no matter how much data you give it — that is high bias, or underfitting.
b2
How much f̂ wobbles around its own average f̄ as the training set changes. A too-flexible model chases the noise, so each dataset yields a wildly different fit — that is high variance, or overfitting.

13. Picture it first: The dartboard picture

Picture it

Figure (svg): Two dartboards. Left: darts tightly clustered but far from center (low variance, high bias). Right: darts spread widely but centered on the bullseye on average (high variance, low bias).

Bias = how far the cluster's center is from the bullseye. Variance = how spread the cluster is. You want both small — but complexity trades one for the other.

Discussion prompt

Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.

Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.

Answer:

Think of each trained model as one throw at a dartboard whose bullseye is the truth f. Retraining on a fresh dataset is another throw.

14. The dartboard picture

Intuition

Think of each trained model as one throw at a dartboard whose bullseye is the truth f. Retraining on a fresh dataset is another throw.

Figure (svg): Two dartboards. Left: darts tightly clustered but far from center (low variance, high bias). Right: darts spread widely but centered on the bullseye on average (high variance, low bias).

Bias = how far the cluster's center is from the bullseye. Variance = how spread the cluster is. You want both small — but complexity trades one for the other.

15. By analogy: The dartboard picture

Analogy

Discussion prompt

Explain The dartboard picture by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Think of each trained model as one throw at a dartboard whose bullseye is the truth f. Retraining on a fresh dataset is another throw.

16. Derive the decomposition

Section

Part 2 of 9 — every step

17. The quantity we decompose

Concept

Fix one test input x. Its true label is the random y = f + ε. Our model, trained on a random dataset, predicts f̂. The expected squared error at x averages over BOTH sources of randomness: the noise ε AND the training set (which is what makes f̂ random).

\[ \operatorname{Err}(x) = \mathbb{E}\!\left[(y - \hat f)^2\right], \qquad y = f + \varepsilon \]

We will rewrite this single expectation as three clean pieces. The trick is one line: add and subtract the average model f̄ = E[f̂].

18. The three assumptions we use

Concept

  1. E[ε] = 0 — the noise is mean-zero (no systematic offset in the labels).
  2. Var(ε) = σ² — the noise has a fixed variance, and E[ε²] = σ².
  3. ε is independent of the training data, hence independent of f̂.

Also write f̄ = E[f̂] for the average prediction over training sets. That is all we need — the rest is algebra.

19. Teach it back: The three assumptions we use

Explain it

Discussion prompt

Explain The three assumptions we use to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Also write f̄ = E[f̂] for the average prediction over training sets. That is all we need — the rest is algebra.

20. Bias and variance, as formulas

Concept

Before the algebra, pin down the two quantities precisely so the derivation lands on named objects, not vibes. Bias is the gap between the average model and the truth; variance is the mean-square wobble of the model about its own average:

\[ \operatorname{Bias}(\hat f) = \bar f - f, \qquad \operatorname{Var}(\hat f) = \mathbb{E}\!\left[(\hat f - \bar f)^2\right] \]

Both are averages OVER training sets, at a fixed x

Why: Neither can be computed from one fit — they describe the distribution of f̂ as the data is resampled. That is exactly why Part 3 estimates them by simulating many retrainings.

21. Say it in words: Bias and variance, as formulas

Translation

\( \operatorname{Bias}(\hat f) = \bar f - f, \qquad \operatorname{Var}(\hat f) = \mathbb{E}\!\left[(\hat f - \bar f)^2\right] \)

Draw it

Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.

22. What has to happen first: Step 1 — split off the noise

Ranking

Put in order

Put the moves of Step 1 — split off the noise into the order they have to happen.

  1. Substitute y = f + ε and regroup
  2. Expand the square
  3. The middle term is 0

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Replace y by f + ε inside the square, then group the deterministic (f − f̂) apart from the random ε.

23. Step 1 — split off the noise

Worked example

Substitute y = f + ε and regroup

Why: Replace y by f + ε inside the square, then group the deterministic (f − f̂) apart from the random ε.

\[ \mathbb{E}\!\left[(y - \hat f)^2\right] = \mathbb{E}\!\left[\big((f - \hat f) + \varepsilon\big)^2\right] \]

Expand the square

Why: (a + b)² = a² + 2ab + b² with a = f − f̂ and b = ε.

\[ = \mathbb{E}\!\left[(f - \hat f)^2\right] + 2\,\mathbb{E}\!\left[(f - \hat f)\,\varepsilon\right] + \mathbb{E}\!\left[\varepsilon^2\right] \]

The middle term is 0

Why: ε is independent of f̂, so E[(f − f̂)ε] = E[f − f̂]·E[ε] = E[f − f̂]·0 = 0. The last term is E[ε²] = σ².

\[ = \underbrace{\mathbb{E}\!\left[(f - \hat f)^2\right]}_{\text{model error}} + \underbrace{\sigma^2}_{\text{noise}} \]

24. Decode the notation: Step 1 — split off the noise

Notation

Annotate

From Step 1 — split off the noise — read this one piece at a time. What is each part doing?

On: \( \mathbb{E}\!\left[(y - \hat f)^2\right] = \mathbb{E}\!\left[\big((f - \hat f) + \varepsilon\big)^2\right] \)

  • Replace y by f + ε inside the square, then group the deterministic (f − f̂) apart from the random ε.
  • (a + b)² = a² + 2ab + b² with a = f − f̂ and b = ε.
  • ε is independent of f̂, so E[(f − f̂)ε] = E[f − f̂]·E[ε] = E[f − f̂]·0 = 0. The last term is E[ε²] = σ².

25. What has to be given first: Step 2 — add and subtract the mean model

Missing information

Discussion prompt

Now crack open the model-error term E[(f − f̂)²]. Insert f̄ = E[f̂] by adding and subtracting it — the move that separates bias from variance:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

f − f̂ = (f − f̄) + (f̄ − f̂). The first piece is a fixed number (bias), the second is the random wobble of f̂ about its mean.

26. Step 2 — add and subtract the mean model

Worked example

Now crack open the model-error term E[(f − f̂)²]. Insert f̄ = E[f̂] by adding and subtracting it — the move that separates bias from variance:

Insert ± f̄

Why: f − f̂ = (f − f̄) + (f̄ − f̂). The first piece is a fixed number (bias), the second is the random wobble of f̂ about its mean.

\[ \mathbb{E}\!\left[(f - \hat f)^2\right] = \mathbb{E}\!\left[\big((f - \bar f) + (\bar f - \hat f)\big)^2\right] \]

Expand the square again

Why: Same (a + b)² rule, now with a = f − f̄ (constant) and b = f̄ − f̂ (random, mean 0).

\[ = \mathbb{E}\!\left[(f - \bar f)^2\right] + 2\,\mathbb{E}\!\left[(f - \bar f)(\bar f - \hat f)\right] + \mathbb{E}\!\left[(\bar f - \hat f)^2\right] \]

27. Work backwards from the answer: Step 2 — add and subtract the mean model

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Expand the square again

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Now crack open the model-error term E[(f − f̂)²]. Insert f̄ = E[f̂] by adding and subtracting it — the move that separates bias from variance:

28. What has to happen first: Step 3 — the last cross term also vanishes

Ranking

Put in order

Put the moves of Step 3 — the last cross term also vanishes into the order they have to happen.

  1. Pull the constant out of the cross term
  2. The leftover expectation is 0
  3. Name the two survivors

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. (f − f̄) is not random — it has no dependence on the training set — so it factors out of the expectation.

29. Step 3 — the last cross term also vanishes

Worked example

Pull the constant out of the cross term

Why: (f − f̄) is not random — it has no dependence on the training set — so it factors out of the expectation.

\[ 2\,\mathbb{E}\!\left[(f - \bar f)(\bar f - \hat f)\right] = 2\,(f - \bar f)\,\mathbb{E}\!\left[\bar f - \hat f\right] \]

The leftover expectation is 0

Why: E[f̄ − f̂] = f̄ − E[f̂] = f̄ − f̄ = 0, because f̄ is by definition the mean of f̂. So the whole cross term dies.

\[ = 2\,(f - \bar f)\cdot 0 = 0 \]

Name the two survivors

Why: E[(f − f̄)²] = (f − f̄)² is the squared bias (a constant); E[(f̄ − f̂)²] is the variance of f̂ by definition.

\[ \mathbb{E}\!\left[(f - \hat f)^2\right] = \underbrace{(f - \bar f)^2}_{\text{Bias}^2} + \underbrace{\mathbb{E}\!\left[(\bar f - \hat f)^2\right]}_{\text{Variance}} \]

30. Decode the notation: Step 3 — the last cross term also vanishes

Notation

Annotate

From Step 3 — the last cross term also vanishes — read this one piece at a time. What is each part doing?

On: \( \mathbb{E}\!\left[(f - \hat f)^2\right] = \underbrace{(f - \bar f)^2}_{\text{Bias}^2} + \underbrace{\mathbb{E}\!\left[(\bar f - \hat f)^2\right]}_{\text{Variance}} \)

  • (f − f̄) is not random — it has no dependence on the training set — so it factors out of the expectation.
  • E[f̄ − f̂] = f̄ − E[f̂] = f̄ − f̄ = 0, because f̄ is by definition the mean of f̂. So the whole cross term dies.
  • E[(f − f̄)²] = (f − f̄)² is the squared bias (a constant); E[(f̄ − f̂)²] is the variance of f̂ by definition.

31. Assemble: the decomposition

Worked example

Put Step 1 (noise split off) together with Step 3 (bias/variance split). Every cross term was shown to be exactly zero, so nothing is hidden:

\[ \boxed{\;\mathbb{E}\!\left[(y - \hat f)^2\right] = \underbrace{(f - \bar f)^2}_{\text{Bias}^2} + \underbrace{\mathbb{E}\!\left[(\hat f - \bar f)^2\right]}_{\text{Variance}} + \underbrace{\sigma^2}_{\text{Noise}}\;} \]

Read the three terms

Why: Bias² you shrink with a MORE flexible model; Variance you shrink with a LESS flexible model (or more data). σ² you cannot shrink at all. The tension between the first two is the whole game.

32. Something is wrong here: which cross term survives?

Anomaly

Predict first

A student writes this, and it looks reasonable:

After expanding ((f − f̄) + (f̄ − f̂))², keep the cross term 2(f − f̄)(f̄ − f̂) — it looks like it should contribute a Bias×√Variance piece.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: This forgets that we take the EXPECTATION.

Take the expectation of the cross term. (f − f̄) is constant and factors out; E[f̄ − f̂] = 0.

Why: This forgets that we take the EXPECTATION. E[f̄ − f̂] = 0, so the cross term is 0 in expectation — you can't leave a random f̂ sitting in a final answer that is supposed to be an expectation.

33. Trap: which cross term survives?

Trap

The trap

After expanding ((f − f̄) + (f̄ − f̂))², keep the cross term 2(f − f̄)(f̄ − f̂) — it looks like it should contribute a Bias×√Variance piece.

MSE = Bias² + Variance + 2·Bias·(f̄ − f̂) + σ² ✗

Why: This forgets that we take the EXPECTATION. E[f̄ − f̂] = 0, so the cross term is 0 in expectation — you can't leave a random f̂ sitting in a final answer that is supposed to be an expectation.

The fix

Take the expectation of the cross term. (f − f̄) is constant and factors out; E[f̄ − f̂] = 0.

MSE = Bias² + Variance + σ² ✓

Why: Exactly three terms. The cross term vanishes precisely because f̄ is DEFINED as E[f̂] — that definition is what makes the decomposition clean. Always finish the expectation before reading off terms.

34. Verify it numerically

Section

Part 3 of 9 — the identity is real

35. A toy where we know the answer

Intuition

Before the messy polynomials, pin the identity down on a case simple enough to check by hand. The truth is a constant f = μ = 2, and the noise has σ = 1.

Our estimator: draw n = 3 noisy points and predict their average. By symmetry this estimator is unbiased — on average it hits μ — so Bias² = 0. And the variance of an average of 3 independent draws is σ²/3 = 1/3. So the identity predicts MSE = 0 + 1/3 + 1 = 1.333.

36. Predict the next row: The unbiased averaging estimator

Pattern

Predict first

The table runs: Bias² | 0 | 0.0000 · Variance | σ²/3 = 0.3333 | 0.3342 · Noise σ² | 1 | 1.0000 · sum | 1.3333 | 1.3342

In The unbiased averaging estimator, given the rows so far: what is the next one — the row where term is raw MSE?

Correct: raw MSE | — | 1.3358

termpredictedmeasured
Bias²00.0000
Varianceσ²/3 = 0.33330.3342
Noise σ²11.0000
sum1.33331.3342
raw MSE—1.3358

Why: The relationship between the columns, not the individual numbers, is what generates the next row. The measured variance 0.3342 matches the theoretical σ²/3 = 0.3333.

37. The unbiased averaging estimator

Worked example

Simulate 200,000 retrainings. Each retraining draws a fresh 3-point dataset and predicts its mean; we also draw a fresh test label each trial. Measure bias², variance, and the raw MSE, then compare to 0 + 1/3 + 1:

import numpy as np
np.random.seed(0)
mu, sigma, n = 2.0, 1.0, 3
def fhat():
    return (mu + sigma * np.random.randn(n)).mean()   # predict the mean
T = 200000
preds = np.array([fhat() for _ in range(T)])
ynew  = mu + sigma * np.random.randn(T)                # fresh test label
mse   = np.mean((ynew - preds) ** 2)
bias2 = (preds.mean() - mu) ** 2
var   = preds.var()
print(round(bias2, 4), round(var, 4), sigma**2)
print('sum  =', round(bias2 + var + sigma**2, 4))
print('mse  =', round(mse, 4))

bias² = 0.0, variance = 0.334 ≈ σ²/3, σ² = 1

Why: The measured variance 0.3342 matches the theoretical σ²/3 = 0.3333. The estimator is unbiased so bias² rounds to 0.

termpredictedmeasured
Bias²00.0000
Varianceσ²/3 = 0.33330.3342
Noise σ²11.0000
sum1.33331.3342
raw MSE—1.3358

38. What each one costs: The unbiased averaging estimator

Trade off

Comparison matrix

From The unbiased averaging estimator: every row here is a choice with a cost. Fill the predicted column, then say which row you would actually pick and what you give up for it.

termpredictedmeasured
Bias²00.0000
Varianceσ²/3 = 0.33330.3342
Noise σ²11.0000
sum1.33331.3342
raw MSE—1.3358

39. Restore the missing line: A biased estimator, same truth

Fill the middle

Fill in the blanks

From A biased estimator, same truth — one line has had its right-hand side removed. Put it back.

import numpy as np
np.random.seed(1)
mu, sigma = 2.0, 1.0
c = 1.5 # constant predictor, ignores the data
T = 200000
ynew = **mu + sigma * np.random.randn(T)
mse = np.mean((ynew - c)
2)
bias2 = (c - mu) 2
print('bias2 =', bias2, ' var = 0 sigma2 =', sigma
2)
print('sum =', bias2 + 0 + sigma**2)
print('mse =', round(mse, 4))

Why: ynew is what everything below it consumes, so the wrong expression here fails later and somewhere else. Trading data for a fixed guess removed all variance but injected bias.

40. A biased estimator, same truth

Worked example

Now break the aim on purpose: ignore the data and always predict the constant c = 1.5. The truth is still μ = 2, so Bias = 1.5 − 2 = −0.5 and Bias² = 0.25. Variance is 0 (the prediction never changes). Identity predicts 0.25 + 0 + 1 = 1.25:

import numpy as np
np.random.seed(1)
mu, sigma = 2.0, 1.0
c = 1.5                     # constant predictor, ignores the data
T = 200000
ynew = mu + sigma * np.random.randn(T)
mse  = np.mean((ynew - c) ** 2)
bias2 = (c - mu) ** 2
print('bias2 =', bias2, ' var = 0  sigma2 =', sigma**2)
print('sum   =', bias2 + 0 + sigma**2)
print('mse   =', round(mse, 4))

bias² = 0.25, variance = 0, sum = 1.25 = measured MSE 1.2498

Why: Trading data for a fixed guess removed all variance but injected bias. The identity holds term for term — this is bias and variance as a genuine tug-of-war.

estimatorBias²Varianceσ²predicted MSEmeasured MSE
average of 30.000.331.001.331.336
constant 1.50.250.001.001.251.250

41. Where does each piece belong: Lesson 32: Bias-Variance & Cross-Validation

Sorting

Sort into buckets

These are the pieces of Lesson 32: Bias-Variance & Cross-Validation, out of order. Put each one back under the part of the lesson it belongs to.

Two ways to be wrong
The running example: a noisy curve; One sample of the running data; The model is a random variable
Derive the decomposition
The quantity we decompose; The three assumptions we use; Bias and variance, as formulas
Verify it numerically
A toy where we know the answer; The unbiased averaging estimator; A biased estimator, same truth
s1
Two ways to be wrong is where Lesson 32: Bias-Variance & Cross-Validation puts The running example: a noisy curve, One sample of the running data, The model is a random variable. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
Derive the decomposition is where Lesson 32: Bias-Variance & Cross-Validation puts The quantity we decompose, The three assumptions we use, Bias and variance, as formulas. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
Verify it numerically is where Lesson 32: Bias-Variance & Cross-Validation puts A toy where we know the answer, The unbiased averaging estimator, A biased estimator, same truth. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

42. Bias & variance vs complexity

Section

Part 4 of 9 — the U-curve appears

43. How complexity moves the two terms

Concept

Back to f(x) = sin(1.5x). The degree of the fitted polynomial is our complexity knob. A low degree can't bend enough to follow the sine — high bias. A high degree bends through every noisy point — high variance.

To measure this we resample: draw many training sets, fit each, and look at how the predictions behave at a grid of test points. The average prediction reveals bias; the spread reveals variance.

44. Finish it with less help: The measurement machine

Faded example

Fill in the blanks

The measurement machine, with the scaffolding fading: two lines are gone now — fill both.

import numpy as np
np.random.seed(0)
f = **lambda t: np.sin(1.5 * t)**
sigma = 0.3
xt = np.linspace(0.2, 3.8, 50); ft = f(xt) # test grid + truth
def preds_grid(deg, n=12):
x = np.random.uniform(0, 4, n)
y = f(x) + sigma * np.random.randn(n)
return np.polyval(np.polyfit(x, y, deg), xt)
for deg in [1, 2, 3, 5, 9]:
P = np.array([preds_grid(deg) for _ in range(500)])
bias2 = ((P.mean(0) - ft) 2).mean()
var =
P.var(0).mean()
print(deg, round(bias2, 4), round(var, 4), round(bias2 + var + sigma
2, 4))

Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. As degree rises the average fit hugs the sine (bias vanishes), but the wobble across datasets grows without bound once the polynomial can thread the noise.

45. The measurement machine

Worked example

For a fixed degree, draw 500 training sets of 12 points each, fit a polynomial to each, and predict on a fixed test grid. P is a (500 × grid) array of predictions. Bias² is (mean over datasets − truth)²; variance is the spread across datasets — both averaged over the grid:

import numpy as np
np.random.seed(0)
f = lambda t: np.sin(1.5 * t)
sigma = 0.3
xt = np.linspace(0.2, 3.8, 50); ft = f(xt)      # test grid + truth
def preds_grid(deg, n=12):
    x = np.random.uniform(0, 4, n)
    y = f(x) + sigma * np.random.randn(n)
    return np.polyval(np.polyfit(x, y, deg), xt)
for deg in [1, 2, 3, 5, 9]:
    P = np.array([preds_grid(deg) for _ in range(500)])
    bias2 = ((P.mean(0) - ft) ** 2).mean()
    var   = P.var(0).mean()
    print(deg, round(bias2, 4), round(var, 4), round(bias2 + var + sigma**2, 4))

bias² collapses 0.132 → 0.002; variance explodes 0.049 → 79 million

Why: As degree rises the average fit hugs the sine (bias vanishes), but the wobble across datasets grows without bound once the polynomial can thread the noise. Total = bias² + variance + 0.09.

degreeBias²Variancetotal = B²+V+σ²
1 (underfit)0.13200.04940.2714
20.11440.17810.3825
3 (balanced)0.00210.09200.1841
50.014014.3114.41
9 (overfit)3035979.1 million79.1 million

46. Reading the table: the U-curve

Concept

Follow the total column: 0.271 → 0.383 → 0.184 → 14.4 → huge. It dips to a minimum at degree 3, then blows up. That dip is the famous U-shaped total error — you overshoot the sweet spot in either direction.

Figure (svg): Error on the vertical axis vs model complexity on the horizontal axis. Bias-squared curve falls from left to right; variance curve rises. Their sum is a U-shaped total-error curve with a minimum in the middle.

Bias² falls, variance rises, their sum is a U. The bottom of the U — here degree 3 — is the model you want.

47. More data cures variance, not bias

Worked example

Which term does more data attack? Hold the degree fixed at 5 (the high-variance regime) and grow the training-set size n = 12 → 30 → 100, remeasuring bias² and variance each time:

import numpy as np
np.random.seed(0)
f = lambda t: np.sin(1.5 * t)
sigma = 0.3
xt = np.linspace(0.2, 3.8, 50); ft = f(xt)
def preds(deg, n):
    x = np.random.uniform(0, 4, n)
    y = f(x) + sigma * np.random.randn(n)
    return np.polyval(np.polyfit(x, y, deg), xt)
for n in [12, 30, 100]:
    P = np.array([preds(5, n) for _ in range(400)])
    bias2 = ((P.mean(0) - ft) ** 2).mean()
    var   = P.var(0).mean()
    print(n, round(bias2, 4), round(var, 4))

variance collapses 55.9 → 0.024 → 0.0045; bias² stays ~0

Why: More data stabilizes the fit (variance falls fast) but cannot fix an aim that was never wrong here — bias is already ~0 at degree 5. So high variance is the one you buy your way out of with data.

n (train size)Bias² (deg 5)Variance (deg 5)
120.112855.9255
300.00010.0238
1000.00000.0045

48. Watch it run: More data cures variance, not bias

Pattern

Step through it

Step through More data cures variance, not bias one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: n (train size) is 12
  2. Step 2: n (train size) is 30
  3. Step 3: n (train size) is 100

49. Restore the missing line: Underfit and overfit, one fit each

Fill the middle

Fill in the blanks

From Underfit and overfit, one fit each — one line has had its right-hand side removed. Put it back.

import numpy as np
np.random.seed(0)
x = np.array([0.5, 1.0, 2.0, 3.0, 3.5])
f = np.sin(1.5 * x)
y = f + 0.3 * np.random.randn(5)
x0 = 2.0
p1 = np.polyfit(x, y, 1)
p4 = np.polyfit(x, y, 4)
print('deg1 pred @2 =', round(np.polyval(p1, x0), 4))
print('deg4 pred @2 =', round(np.polyval(p4, x0), 4))
print('truth f(2) =', round(np.sin(1.5 * 2.0), 4))

Why: x0 is what everything below it consumes, so the wrong expression here fails later and somewhere else. On THIS dataset both overshoot, but the point of bias/variance is the behavior ACROSS datasets, not one fit — which is exactly why the resampling machine on the previous slide is the honest measurement.

50. Underfit and overfit, one fit each

Worked example

Zoom to a single training realization and predict at one test point x0 = 2 (truth f(2) = 0.1411). Degree 1 can't reach; degree 4 is closer — but its wobble across datasets is what the table already flagged:

import numpy as np
np.random.seed(0)
x = np.array([0.5, 1.0, 2.0, 3.0, 3.5])
f = np.sin(1.5 * x)
y = f + 0.3 * np.random.randn(5)
x0 = 2.0
p1 = np.polyfit(x, y, 1)
p4 = np.polyfit(x, y, 4)
print('deg1 pred @2 =', round(np.polyval(p1, x0), 4))
print('deg4 pred @2 =', round(np.polyval(p4, x0), 4))
print('truth  f(2)  =', round(np.sin(1.5 * 2.0), 4))

deg1 = 0.432, deg4 = 0.435, truth = 0.141

Why: On THIS dataset both overshoot, but the point of bias/variance is the behavior ACROSS datasets, not one fit — which is exactly why the resampling machine on the previous slide is the honest measurement.

modelprediction @ x0=2truth f(2)
degree 10.43180.1411
degree 40.43470.1411
— across datasets —deg1: steady & offdeg9: wild

51. Diagnose from curves

Section

Part 5 of 9 — train vs test

52. Guess the shape of the answer: Train error vs test error

Estimation

Predict first

The catch: in practice you can't compute bias and variance — they need the unknown f. What you CAN compute is training error and held-out (test) error. Fit degrees 1..12 on one 15-point training set and score both:

Commit before you compute: what does Train error vs test error come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: train error falls monotonically; test error is U-shaped, min at degree 3

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Training error 0.379 → 0.014 always drops with complexity — it is useless for model choice.

53. Train error vs test error

Worked example

The catch: in practice you can't compute bias and variance — they need the unknown f. What you CAN compute is training error and held-out (test) error. Fit degrees 1..12 on one 15-point training set and score both:

import numpy as np
np.random.seed(1)
f = lambda t: np.sin(1.5 * t)
sigma = 0.3
xtr = np.sort(np.random.uniform(0, 4, 15))
ytr = f(xtr) + sigma * np.random.randn(15)
xte = np.sort(np.random.uniform(0, 4, 200))
yte = f(xte) + sigma * np.random.randn(200)
for deg in [1, 2, 3, 5, 9, 12]:
    c = np.polyfit(xtr, ytr, deg)
    tr = np.mean((np.polyval(c, xtr) - ytr) ** 2)
    te = np.mean((np.polyval(c, xte) - yte) ** 2)
    print(deg, round(tr, 4), round(te, 4))

train error falls monotonically; test error is U-shaped, min at degree 3

Why: Training error 0.379 → 0.014 always drops with complexity — it is useless for model choice. Test error 0.237 → 84 million bottoms out at degree 3 (0.1455), then explodes. Only held-out error tells the truth.

degreetrain MSEtest MSE
10.37870.2367
20.13760.5537
30.07690.1455 ← min
50.06150.1183
90.02423074.3
120.013684,366,885

54. Watch it run: Train error vs test error

Pattern

Step through it

Step through Train error vs test error one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: degree is 1
  2. Step 2: degree is 2
  3. Step 3: degree is 3
  4. Step 4: degree is 5
  5. Step 5: degree is 9
  6. Step 6: degree is 12

55. Learning curves tell you which regime

Concept

Plot error against training-set size for a fixed model. High bias: train and validation error both plateau high and close together — more data won't help, you need a richer model.

High variance: a large gap — low train error, much higher validation error — that shrinks as you add data. More data (or regularization) is the cure. The gap between the two curves is your single best overfitting gauge.

56. Something is wrong here: a more complex model is always better

Anomaly

Predict first

A student writes this, and it looks reasonable:

Degree 12 gets training MSE down to 0.0136, far below degree 3's 0.077. Lower training error is a better fit, so ship degree 12.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Training error ALWAYS falls with complexity — it can be driven to zero by memorizing noise.

Choose complexity by held-out error, at the bottom of the U.

Why: Training error ALWAYS falls with complexity — it can be driven to zero by memorizing noise. Degree 12's TEST MSE is 84 million: it fits the 15 training points and detonates everywhere else. Training error can never justify a model choice.

57. Trap: a more complex model is always better

Trap

The trap

Degree 12 gets training MSE down to 0.0136, far below degree 3's 0.077. Lower training error is a better fit, so ship degree 12.

Pick degree 12 on training error ✗

Why: Training error ALWAYS falls with complexity — it can be driven to zero by memorizing noise. Degree 12's TEST MSE is 84 million: it fits the 15 training points and detonates everywhere else. Training error can never justify a model choice.

The fix

Choose complexity by held-out error, at the bottom of the U.

Pick degree 3 on test error ✓

Why: Degree 3 minimizes TEST MSE (0.1455) — it sits at the balance of bias² and variance. When you have little data you don't waste a fixed test split; you use cross-validation (next) to estimate that held-out error.

58. Break it on purpose: a more complex model is always better

Break the constraint

Discussion prompt

The rule this trap just fixed:

Choose complexity by held-out error, at the bottom of the U.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

Training error ALWAYS falls with complexity — it can be driven to zero by memorizing noise. Degree 12's TEST MSE is 84 million: it fits the 15 training points and detonates everywhere else. Training error can never justify a model choice.

59. Nuance: the classical U isn't the whole story

Concept

The U-curve is the classical picture, and it governs the models on this exam: polynomials, ridge, trees, k-NN. But olympiad problems sometimes probe the modern twist — double descent.

In heavily over-parameterized models (huge neural nets), test error rises to a peak near the interpolation threshold — where the model exactly fits the training data — then, surprisingly, falls again as capacity grows further. So 'more complex is worse' holds until interpolation and can reverse beyond it.

The safe takeaway is unchanged: never choose complexity by training error — let held-out error, wherever its true minimum lies, decide.

60. Cross-validation

Section

Part 6 of 9 — build k-fold

61. Why not just one split?

Concept

A single train/test split wastes data (the test points never train the model) and gives a noisy score — change which points land in the test set and the number jumps around, exactly the variance problem we've been studying.

k-fold cross-validation fixes both: chop the data into k equal folds, hold out each fold once as validation while training on the other k−1, and average the k scores. Every point is validated exactly once, so no data is wasted and the estimate has lower variance.

62. Picture it first: The rotation, drawn

Picture it

Figure (svg): Five rows of five boxes. In row i, the i-th box is shaded as the validation fold and the other four are training folds; the shaded box moves one position right each row.

5-fold CV: the validation fold (shaded) rotates through every position; each point is held out exactly once.

Discussion prompt

Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.

Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.

Answer:

Picture 5 folds. Each row is one CV iteration: one fold (shaded) is held out to score, the rest train. Rotate the held-out fold through all 5 positions, then average the 5 scores.

63. The rotation, drawn

Intuition

Picture 5 folds. Each row is one CV iteration: one fold (shaded) is held out to score, the rest train. Rotate the held-out fold through all 5 positions, then average the 5 scores.

Figure (svg): Five rows of five boxes. In row i, the i-th box is shaded as the validation fold and the other four are training folds; the shaded box moves one position right each row.

5-fold CV: the validation fold (shaded) rotates through every position; each point is held out exactly once.

64. k-fold from scratch — the split

Worked example

Start on the diabetes dataset (442 rows) so the folds and numbers are concrete. np.array_split cuts the index range into 5 nearly equal folds. Predict the sizes before printing:

import numpy as np
from sklearn.datasets import load_diabetes
X, y = load_diabetes(return_X_y=True)
folds = np.array_split(np.arange(len(X)), 5)
print('n =', len(X))
print('fold sizes =', [len(fld) for fld in folds])

442 = 89 + 89 + 88 + 88 + 88

Why: np.array_split spreads the remainder: 442 = 5×88 + 2, so the first two folds get one extra row. No point is dropped or duplicated — the folds partition the data.

quantityvalue
n (rows)442
number of folds5
fold sizes[89, 89, 88, 88, 88]

65. k-fold from scratch — the loop, per fold

Worked example

For each fold i: hold it out as test, concatenate the rest as train, fit Ridge(alpha=1.0), and score. Print every fold's train size, test size, and R² — the full trace:

import numpy as np
from sklearn.linear_model import Ridge
from sklearn.datasets import load_diabetes
X, y = load_diabetes(return_X_y=True)
folds = np.array_split(np.arange(len(X)), 5)
for i in range(5):
    te = folds[i]
    tr = np.concatenate([folds[j] for j in range(5) if j != i])
    m = Ridge(alpha=1.0).fit(X[tr], y[tr])
    print(i, len(tr), len(te), round(m.score(X[te], y[te]), 4))

five R² scores: 0.322, 0.441, 0.422, 0.425, 0.442

Why: Fold 0 scores lowest (0.3217) — a single split landing there would badly misjudge the model. Averaging the five is what tames that luck-of-the-split variance.

fold itrain sizetest sizeR² on held-out
0353890.3217
1353890.4405
2354880.4221
3354880.4247
4354880.4420

66. What happens as it grows: k-fold from scratch — the loop, per fold

Scale up

Step through it

Step through k-fold from scratch — the loop, per fold and watch the numbers move. Now imagine the input ten times bigger: which column is the one that stops this being practical?

  1. Step 1: fold i is 0
  2. Step 2: fold i is 1
  3. Step 3: fold i is 2
  4. Step 4: fold i is 3
  5. Step 5: fold i is 4

67. Finish it with less help: k-fold from scratch — average & verify

Faded example

Fill in the blanks

k-fold from scratch — average & verify, with the scaffolding fading: two lines are gone now — fill both.

import numpy as np
from sklearn.linear_model import Ridge
from sklearn.datasets import load_diabetes
from sklearn.model_selection import cross_val_score, KFold
X, y = load_diabetes(return_X_y=True)
def kfold(model, X, y, k=5):
folds = np.array_split(np.arange(len(X)), k); s = []
for i in range(k):
te = folds[i]
tr = np.concatenate([folds[j] for j in range(k) if j != i])
model.fit(X[tr], y[tr]); s.append(model.score(X[te], y[te]))
return np.mean(s)
mine = kfold(Ridge(alpha=1.0), X, y)
skl = cross_val_score(Ridge(alpha=1.0), X, y, cv=KFold(5, shuffle=False)).mean()
print(round(mine, 4), round(skl, 4), np.isclose(mine, skl))

Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. Our from-scratch average equals sklearn's cross_val_score to the digit, and np.isclose returns True — because we reproduced KFold's contiguous, no-shuffle split exactly.

68. k-fold from scratch — average & verify

Worked example

Wrap the loop in a function that returns the mean of the 5 scores, then compare to sklearn.cross_val_score with KFold(5, shuffle=False) — the same no-shuffle split we built:

import numpy as np
from sklearn.linear_model import Ridge
from sklearn.datasets import load_diabetes
from sklearn.model_selection import cross_val_score, KFold
X, y = load_diabetes(return_X_y=True)
def kfold(model, X, y, k=5):
    folds = np.array_split(np.arange(len(X)), k); s = []
    for i in range(k):
        te = folds[i]
        tr = np.concatenate([folds[j] for j in range(k) if j != i])
        model.fit(X[tr], y[tr]); s.append(model.score(X[te], y[te]))
    return np.mean(s)
mine = kfold(Ridge(alpha=1.0), X, y)
skl  = cross_val_score(Ridge(alpha=1.0), X, y, cv=KFold(5, shuffle=False)).mean()
print(round(mine, 4), round(skl, 4), np.isclose(mine, skl))

mean(0.322, 0.441, 0.422, 0.425, 0.442) = 0.4102 = sklearn

Why: Our from-scratch average equals sklearn's cross_val_score to the digit, and np.isclose returns True — because we reproduced KFold's contiguous, no-shuffle split exactly.

source5-fold mean R²match
from scratch0.4102—
sklearn cross_val_score0.4102True
single split (fold 0 only)0.3217misleading

69. Something is wrong here: shuffle mismatch breaks the comparison

Anomaly

Predict first

A student writes this, and it looks reasonable:

Compare your contiguous-split k-fold to cross_val_score with the default cv=5 (or shuffle=True) and expect the same number.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: A shuffled KFold assigns DIFFERENT rows to each fold, so its per-fold scores and average differ from your contiguous split.

Match the splitter exactly: contiguous np.array_split mirrors KFold(k, shuffle=False).

Why: A shuffled KFold assigns DIFFERENT rows to each fold, so its per-fold scores and average differ from your contiguous split. The numbers won't match and you'll think your code is wrong when it isn't.

70. Trap: shuffle mismatch breaks the comparison

Trap

The trap

Compare your contiguous-split k-fold to cross_val_score with the default cv=5 (or shuffle=True) and expect the same number.

mine (no shuffle) vs KFold(shuffle=True) ✗

Why: A shuffled KFold assigns DIFFERENT rows to each fold, so its per-fold scores and average differ from your contiguous split. The numbers won't match and you'll think your code is wrong when it isn't.

The fix

Match the splitter exactly: contiguous np.array_split mirrors KFold(k, shuffle=False).

mine vs KFold(5, shuffle=False) ✓

Why: Same rows in each fold → identical scores → identical average 0.4102, np.isclose True. When verifying a from-scratch algorithm, replicate the reference's data ordering before comparing outputs.

71. Which of these survive contact with Lesson 32: Bias-Variance & Cross-Validation?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Our job: fit a polynomial f̂ to the samples that predicts well on new points. The only knob is the polynomial degree — that is our model complexity.; Think of each trained model as one throw at a dartboard whose bullseye is the truth f. Retraining on a fresh dataset is another throw.; Also write f̄ = E[f̂] for the average prediction over training sets. That is all we need — the rest is algebra.
Breaks
After expanding ((f − f̄) + (f̄ − f̂))², keep the cross term 2(f − f̄)(f̄ − f̂) — it looks like it should contribute a Bias×√Variance piece.; Degree 12 gets training MSE down to 0.0136, far below degree 3's 0.077. Lower training error is a better fit, so ship degree 12.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 32: Bias-Variance & Cross-Validation puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

72. The CV strategy zoo

Section

Part 7 of 9 — pick the right split

73. Four splitters, four situations

Concept

strategyuse whenkey property
k-foldthe defaulteach point held out once
stratified k-foldclassificationpreserves class ratios per fold
leave-one-out (LOO)tiny datasets (k = n)n fits; low bias, high variance, costly
time-series splittemporal datatrain only on the PAST

Plain k-fold assumes rows are exchangeable. When they are not — imbalanced classes, or time order — you must switch splitter or the estimate lies.

74. Fill in: use when for Four splitters, four situations

Comparison

Comparison matrix

From Four splitters, four situations: refill the use when column from what you know. The rest of the table is as it appeared.

strategyuse whenkey property
k-foldthe defaulteach point held out once
stratified k-foldclassificationpreserves class ratios per fold
leave-one-out (LOO)tiny datasets (k = n)n fits; low bias, high variance, costly
time-series splittemporal datatrain only on the PAST

75. Restore the missing line: Stratified preserves class balance

Fill the middle

Fill in the blanks

From Stratified preserves class balance — one line has had its right-hand side removed. Put it back.

import numpy as np
from sklearn.model_selection import KFold, StratifiedKFold
y = np.array([0]90 + [1]10) # 90/10 imbalance
np.random.seed(0)
y = y[np.random.permutation(len(y))]
def pos_rates(splitter):
return [round(y[te].mean(), 3) for _, te in splitter.split(np.zeros_like(y), y)]
print('plain ', pos_rates(KFold(5, shuffle=True, random_state=0)))
print('strat ', pos_rates(StratifiedKFold(5, shuffle=True, random_state=0)))

Why: y is what everything below it consumes, so the wrong expression here fails later and somewhere else. Plain k-fold's validation folds range from 5% to 15% positives — a fold with too few positives gives a garbage score.

76. Stratified preserves class balance

Worked example

Take a 90/10 imbalanced label. Plain k-fold splits randomly, so some folds get too few positives; stratified k-fold forces every fold to the global 10% rate. Watch the per-fold positive rate:

import numpy as np
from sklearn.model_selection import KFold, StratifiedKFold
y = np.array([0]*90 + [1]*10)            # 90/10 imbalance
np.random.seed(0)
y = y[np.random.permutation(len(y))]
def pos_rates(splitter):
    return [round(y[te].mean(), 3) for _, te in splitter.split(np.zeros_like(y), y)]
print('plain ', pos_rates(KFold(5, shuffle=True, random_state=0)))
print('strat ', pos_rates(StratifiedKFold(5, shuffle=True, random_state=0)))

plain wobbles 0.05–0.15; stratified is a flat 0.10 everywhere

Why: Plain k-fold's validation folds range from 5% to 15% positives — a fold with too few positives gives a garbage score. Stratified pins every fold to 10%, matching the whole dataset.

splitterper-fold positive ratespread
KFold[0.10, 0.10, 0.15, 0.10, 0.05]0.05 – 0.15
StratifiedKFold[0.10, 0.10, 0.10, 0.10, 0.10]flat 0.10
global rate0.10target

77. Predict the next row: Time-series: never train on the future

Pattern

Predict first

The table runs: 1 | [0, 1] | [2, 3] · 2 | [0, 1, 2, 3] | [4, 5] · 3 | [0..5] | [6, 7]

In Time-series: never train on the future, given the rows so far: what is the next one — the row where split is 4?

Correct: 4 | [0..7] | [8, 9]

splittrain indicestest indices
1[0, 1][2, 3]
2[0, 1, 2, 3][4, 5]
3[0..5][6, 7]
4[0..7][8, 9]

Why: The relationship between the columns, not the individual numbers, is what generates the next row. Split 1 trains on [0,1] tests [2,3]; the window grows to train on [0..7] tests [8,9].

78. Time-series: never train on the future

Worked example

For ordered data, a random fold would let the model peek at future points to predict the past — leakage. TimeSeriesSplit only ever trains on earlier indices and tests on later ones, growing the training window:

import numpy as np
from sklearn.model_selection import TimeSeriesSplit
idx = np.arange(10)
for tr, te in TimeSeriesSplit(n_splits=4).split(idx):
    print('train', list(tr), 'test', list(te))

every test block sits strictly AFTER its training block

Why: Split 1 trains on [0,1] tests [2,3]; the window grows to train on [0..7] tests [8,9]. Training indices are always a prefix — the model never sees the future it must predict.

splittrain indicestest indices
1[0, 1][2, 3]
2[0, 1, 2, 3][4, 5]
3[0..5][6, 7]
4[0..7][8, 9]

79. Watch it run: Time-series: never train on the future

Pattern

Step through it

Step through Time-series: never train on the future one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: split is 1
  2. Step 2: split is 2
  3. Step 3: split is 3
  4. Step 4: split is 4

80. What has to be given first: Leave-one-out: k = n

Missing information

Discussion prompt

The extreme of k-fold: set k = n, so each fold is a single point. That is leave-one-out (LOO) — n fits, each training on all but one point. It wrings the most training data out of a tiny dataset, at the cost of n model fits. Compare its MSE to 5-fold on diabetes:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

LOO trains on n−1 = 441 points each time, so it is nearly unbiased for the full-data model — a slightly lower MSE estimate. But it runs 442 fits (vs 5) and its estimate has higher variance because the n test sets overlap heavily.

81. Leave-one-out: k = n

Worked example

The extreme of k-fold: set k = n, so each fold is a single point. That is leave-one-out (LOO) — n fits, each training on all but one point. It wrings the most training data out of a tiny dataset, at the cost of n model fits. Compare its MSE to 5-fold on diabetes:

import numpy as np
from sklearn.linear_model import Ridge
from sklearn.datasets import load_diabetes
from sklearn.model_selection import cross_val_score, LeaveOneOut, KFold
X, y = load_diabetes(return_X_y=True)
loo = cross_val_score(Ridge(1.0), X, y, cv=LeaveOneOut(),
                      scoring='neg_mean_squared_error')
kf  = cross_val_score(Ridge(1.0), X, y, cv=KFold(5, shuffle=False),
                      scoring='neg_mean_squared_error')
print('LOO folds =', len(loo))
print('LOO  MSE  =', round(-loo.mean(), 2))
print('5-fold MSE=', round(-kf.mean(), 2))

442 folds; LOO MSE 3327.66 vs 5-fold 3420.32

Why: LOO trains on n−1 = 441 points each time, so it is nearly unbiased for the full-data model — a slightly lower MSE estimate. But it runs 442 fits (vs 5) and its estimate has higher variance because the n test sets overlap heavily.

strategyfolds (fits)mean MSEtrade
5-fold53420.32cheap, a little biased
leave-one-out4423327.66costly, low bias / high var
rule of thumb5 or 10—LOO only for tiny n

82. Tuning without cheating

Section

Part 8 of 9 — nested CV

83. The tuning trap in one sentence

Concept

If you try many hyperparameters and report the best score on the same data you used to pick them, that score is optimistically biased — you fit the evaluation set through model selection. This is a form of data leakage.

The cure is separation: tune on inner folds, evaluate on outer folds the tuning never touched. That is nested CV — CV inside CV.

84. Guess the shape of the answer: Cheating vs nested CV, measured

Estimation

Predict first

Grid over Ridge alpha. The cheating protocol picks the alpha that maximizes a single CV and reports that same CV. Nested CV tunes alpha in an inner loop and reports the untouched outer loop. Compare the two estimates:

Commit before you compute: what does Cheating vs nested CV, measured come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: cheat 0.4894 > nested 0.4858

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The cheating score is higher because picking alpha on the same folds it's scored on cherry-picks favorable noise.

85. Cheating vs nested CV, measured

Worked example

Grid over Ridge alpha. The cheating protocol picks the alpha that maximizes a single CV and reports that same CV. Nested CV tunes alpha in an inner loop and reports the untouched outer loop. Compare the two estimates:

import numpy as np
from sklearn.linear_model import Ridge
from sklearn.datasets import load_diabetes
from sklearn.model_selection import GridSearchCV, cross_val_score, KFold
X, y = load_diabetes(return_X_y=True)
grid = {'alpha': [0.01, 0.1, 1.0, 10.0, 100.0]}
inner = KFold(5, shuffle=True, random_state=0)
outer = KFold(5, shuffle=True, random_state=0)
gs = GridSearchCV(Ridge(), grid, cv=inner)
nested = cross_val_score(gs, X, y, cv=outer).mean()
best = max(grid['alpha'], key=lambda a: cross_val_score(Ridge(alpha=a), X, y, cv=outer).mean())
cheat = cross_val_score(Ridge(alpha=best), X, y, cv=outer).mean()
print('cheat:', best, round(cheat, 4))
print('nested:', round(nested, 4))

cheat 0.4894 > nested 0.4858

Why: The cheating score is higher because picking alpha on the same folds it's scored on cherry-picks favorable noise. Nested CV's outer folds never see the tuning, so 0.4858 is the honest generalization estimate.

protocolchosen alphareported scorehonest?
tune & report on same CV0.010.4894no — optimistic
nested CVinner-selected0.4858yes

86. Work backwards from the answer: Cheating vs nested CV, measured

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

cheat 0.4894 > nested 0.4858

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Grid over Ridge alpha. The cheating protocol picks the alpha that maximizes a single CV and reports that same CV. Nested CV tunes alpha in an inner loop and reports the untouched outer loop. Compare the two estimates:

87. Searching the hyperparameter space

Concept

Whatever the search, it belongs in the inner loop of nested CV. The outer loop's only job is to report a number you can trust.

88. Teach it back: Searching the hyperparameter space

Explain it

Discussion prompt

Explain Searching the hyperparameter space to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Whatever the search, it belongs in the inner loop of nested CV. The outer loop's only job is to report a number you can trust.

89. A hyperparameter IS a bias-variance dial

Concept

The reason tuning matters is that a hyperparameter like ridge's λ slides you along the same bias-variance curve. Larger λ shrinks the weights toward zero — less flexibility, less variance, more bias. Smaller λ does the reverse.

\[ \lambda \uparrow \;\Longrightarrow\; \text{variance} \downarrow,\ \text{bias} \uparrow \qquad\qquad \lambda \downarrow \;\Longrightarrow\; \text{variance} \uparrow,\ \text{bias} \downarrow \]

So choosing λ by CV is just finding the bottom of the U in a different coordinate. Polynomial degree, tree depth, k in k-NN, λ in ridge — every complexity knob is a bias-variance dial you set with held-out error.

90. Pattern, checks & project

Section

Part 9 of 9

91. Rebuild the recipe: The bias-variance / CV toolkit

Ranking

Put in order

These are the steps of The bias-variance / CV toolkit, scrambled. Put them back in order before the next slide shows you.

  1. Decompose: every test error is Bias² + Variance + Noise — the cross terms vanish in expectation
  2. Diagnose: high train & high test error = bias (underfit); low train, high test = variance (overfit)
  3. Pick complexity at the minimum of held-out error (bottom of the U), never on training error
  4. Estimate held-out error with k-fold CV — stratified for imbalance, time-series for temporal data, LOO for tiny n
  5. Tune hyperparameters in the inner loop of nested CV so the outer estimate stays unbiased

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

92. The bias-variance / CV toolkit

Pattern

  1. Decompose: every test error is Bias² + Variance + Noise — the cross terms vanish in expectation
  2. Diagnose: high train & high test error = bias (underfit); low train, high test = variance (overfit)
  3. Pick complexity at the minimum of held-out error (bottom of the U), never on training error
  4. Estimate held-out error with k-fold CV — stratified for imbalance, time-series for temporal data, LOO for tiny n
  5. Tune hyperparameters in the inner loop of nested CV so the outer estimate stays unbiased

93. Where does it stop working: The bias-variance / CV toolkit

Edge cases

Discussion prompt

The bias-variance / CV toolkit works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Decompose: every test error is Bias² + Variance + Noise — the cross terms vanish in expectation
  2. Diagnose: high train & high test error = bias (underfit); low train, high test = variance (overfit)
  3. Pick complexity at the minimum of held-out error (bottom of the U), never on training error
  4. Estimate held-out error with k-fold CV — stratified for imbalance, time-series for temporal data, LOO for tiny n
  5. Tune hyperparameters in the inner loop of nested CV so the outer estimate stays unbiased

94. Rule out three: Check yourself — the decomposition

Elimination

Eliminate the wrong options

A model has very low training error but high test error. Which term dominates?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. High variance (overfitting)
  • B. High bias (underfitting)
  • C. High irreducible noise σ²
  • D. The model is perfectly tuned

Survives elimination: A

Why: Low train / high test is the signature of high variance: the model fit the training noise and fails to generalize. Degree 9–12 in the demo did exactly this — tiny train MSE, test MSE in the millions.

95. Check yourself — the decomposition

Check

Match the symptom to the term. Think about train vs test.

Check your understanding

A model has very low training error but high test error. Which term dominates?

  • A. High variance (overfitting) (correct)
  • B. High bias (underfitting)
  • C. High irreducible noise σ²
  • D. The model is perfectly tuned

Answer: A

Why: Low train / high test is the signature of high variance: the model fit the training noise and fails to generalize. Degree 9–12 in the demo did exactly this — tiny train MSE, test MSE in the millions.

Why B tempts people
High bias shows as high error on BOTH train and test — the model can't even fit the training data (degree 1's train MSE 0.379 was already high). A large train–test GAP is the opposite.
Why C tempts people
Irreducible noise σ² raises the floor on train and test EQUALLY; it can't create a train–test gap. σ² is the same whichever model you pick.
Why D tempts people
A large train–test gap is the hallmark of overfitting, not good tuning. A well-tuned model has train and test error close together and both low-ish.

96. Answer it before you see the options: Check yourself — the cross term

Prediction

Predict first

In the derivation, why does the cross term 2·E[(f − f̄)(f̄ − f̂)] equal zero?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: (f − f̄) is constant and E[f̄ − f̂] = 0 since f̄ = E[f̂]

Why: (f − f̄) has no dependence on the training set, so it factors out of the expectation, leaving 2(f − f̄)·E[f̄ − f̂]. Since f̄ is DEFINED as E[f̂], that expectation is f̄ − f̄ = 0.

97. Check yourself — the cross term

Check

Recall why the decomposition has exactly three terms.

Check your understanding

In the derivation, why does the cross term 2·E[(f − f̄)(f̄ − f̂)] equal zero?

  • A. (f − f̄) is constant and E[f̄ − f̂] = 0 since f̄ = E[f̂] (correct)
  • B. Because f̂ and the noise ε are independent
  • C. Because bias and variance are always equal in magnitude
  • D. Because squared terms are always non-negative

Answer: A

Why: (f − f̄) has no dependence on the training set, so it factors out of the expectation, leaving 2(f − f̄)·E[f̄ − f̂]. Since f̄ is DEFINED as E[f̂], that expectation is f̄ − f̄ = 0.

Why B tempts people
Independence of f̂ and ε is what kills the OTHER cross term — the one involving ε in Step 1. This cross term has no ε in it; it dies because E[f̄ − f̂] = 0.
Why C tempts people
Bias and variance are generally NOT equal — the demo had bias² 0.132 with variance 0.049 at degree 1. The term vanishes by the mean-zero argument, not by any equality.
Why D tempts people
Non-negativity is irrelevant here — the cross term isn't a square, it's a product that can be positive or negative. It's zero specifically because its expectation is zero.

98. Answer it before you see the options: Check yourself — k-fold

Prediction

Predict first

In 5-fold cross-validation, each data point is used for validation:

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: Exactly once (and for training in the other 4 folds)

Why: The data partitions into 5 folds; each fold is the validation set once and part of training the other four times. Averaging the 5 held-out scores (0.322, 0.441, 0.422, 0.425, 0.442 → 0.4102) gives the low-variance estimate.

99. Check yourself — k-fold

Check

How does 5-fold use each data point?

Check your understanding

In 5-fold cross-validation, each data point is used for validation:

  • A. Exactly once (and for training in the other 4 folds) (correct)
  • B. 5 times
  • C. Never — validation is a separate held-out set
  • D. Only if it lands in the last fold

Answer: A

Why: The data partitions into 5 folds; each fold is the validation set once and part of training the other four times. Averaging the 5 held-out scores (0.322, 0.441, 0.422, 0.425, 0.442 → 0.4102) gives the low-variance estimate.

Why B tempts people
Each point is VALIDATED once; it is used for TRAINING in the other 4 folds. Saying 5 times confuses training use with validation use.
Why C tempts people
That describes a single train/validation split, not k-fold. k-fold's whole point is that there is no separate wasted held-out set — every point validates once.
Why D tempts people
Every fold serves as validation in turn, not just the last one — that is exactly the rotation in the diagram.

100. Rule out three: Check yourself — nested CV

Elimination

Eliminate the wrong options

Nested cross-validation uses an inner and an outer loop in order to:

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Tune hyperparameters (inner) without biasing the generalization estimate (outer)
  • B. Train two models and average their predictions
  • C. Run CV twice to make it faster
  • D. Increase the effective training-set size

Survives elimination: A

Why: The inner loop selects hyperparameters on each outer training split; the outer test folds never participate in tuning, so the final score is unbiased. In the demo, the cheating estimate 0.4894 was inflated above the honest nested 0.4858.

101. Check yourself — nested CV

Check

Why two loops instead of one?

Check your understanding

Nested cross-validation uses an inner and an outer loop in order to:

  • A. Tune hyperparameters (inner) without biasing the generalization estimate (outer) (correct)
  • B. Train two models and average their predictions
  • C. Run CV twice to make it faster
  • D. Increase the effective training-set size

Answer: A

Why: The inner loop selects hyperparameters on each outer training split; the outer test folds never participate in tuning, so the final score is unbiased. In the demo, the cheating estimate 0.4894 was inflated above the honest nested 0.4858.

Why B tempts people
Nested CV is an evaluation protocol, not an ensemble — it doesn't average model predictions, it separates tuning from testing.
Why C tempts people
Nested CV is MORE expensive (an inner CV per outer fold), not faster. Its payoff is an unbiased estimate, not speed.
Why D tempts people
It reuses the same data — it adds no training examples. It only changes WHICH data is allowed to influence tuning vs evaluation.

102. Project: choose the polynomial degree by CV

Concept

Bring both halves together. Using only NumPy, build k-fold CV on the sin(1.5x) data and use it to pick the polynomial degree — recovering the sweet spot the bias-variance table predicted, without ever peeking at the true f.

#requirementtool
1generate the noisy sin datanp.random + np.sin
2k-fold CV MSE for one degreenp.array_split, np.polyfit
3sweep degrees, pick the min-CV degreenp.polyval, min

Build rules: type every line yourself, seed the RNG so your numbers match, and read errors instead of deleting them — a shape mismatch usually means a fold index went wrong.

103. By analogy: Project: choose the polynomial degree by CV

Analogy

Discussion prompt

Explain Project: choose the polynomial degree by CV by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Build rules: type every line yourself, seed the RNG so your numbers match, and read errors instead of deleting them — a shape mismatch usually means a fold index went wrong.

104. Milestone 1 — the data

Worked example

Your turn: generate 30 noisy points from f(x) = sin(1.5x) on [0, 4] with σ = 0.3 and seed = 2. Predict x's shape and y's rough range before printing.

Hint: np.sort(np.random.uniform(0, 4, 30)) for x; add 0.3 * np.random.randn(30) to sin(1.5x) for y.

import numpy as np
np.random.seed(2)
f = lambda t: np.sin(1.5 * t)
x = np.sort(np.random.uniform(0, 4, 30))
y = f(x) + 0.3 * np.random.randn(30)
print(x.shape, round(y.min(), 3), round(y.max(), 3))
checkvalue
x.shape(30,)
y.min()-1.124
y.max()1.302

105. Watch it run: Milestone 1 — the data

Pattern

Step through it

Step through Milestone 1 — the data one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: check is x.shape
  2. Step 2: check is y.min()
  3. Step 3: check is y.max()

106. Milestone 2 — CV MSE for one degree

Worked example

Your turn: write cv_mse(deg, k=5) — split indices into k folds, fit np.polyfit on k−1, score MSE on the held-out fold, average. Predict roughly the degree-3 CV MSE (recall test error bottomed near degree 3).

Hint: folds = np.array_split(np.arange(len(x)), k); tr = np.concatenate([folds[j] for j in range(k) if j != i]); error is np.mean((np.polyval(c, x[te]) - y[te])**2).

import numpy as np
np.random.seed(2)
f = lambda t: np.sin(1.5 * t)
x = np.sort(np.random.uniform(0, 4, 30))
y = f(x) + 0.3 * np.random.randn(30)
def cv_mse(deg, k=5):
    folds = np.array_split(np.arange(len(x)), k); errs = []
    for i in range(k):
        te = folds[i]
        tr = np.concatenate([folds[j] for j in range(k) if j != i])
        c = np.polyfit(x[tr], y[tr], deg)
        errs.append(np.mean((np.polyval(c, x[te]) - y[te]) ** 2))
    return np.mean(errs)
print(round(cv_mse(3), 4))
quantityvalue
cv_mse(3)0.3597
k5
fits per call5

107. What each one costs: Milestone 2 — CV MSE for one degree

Trade off

Comparison matrix

From Milestone 2 — CV MSE for one degree: every row here is a choice with a cost. Fill the value column, then say which row you would actually pick and what you give up for it.

quantityvalue
cv_mse(3)0.3597
k5
fits per call5

108. Milestone 3 — sweep and pick

Worked example

Your turn: run cv_mse over degrees 1, 2, 3, 5, 9 and pick the degree with the smallest CV MSE. Predict which one wins before you print.

Hint: build a dict {deg: cv_mse(deg)}, then min(scores, key=scores.get).

import numpy as np
np.random.seed(2)
f = lambda t: np.sin(1.5 * t)
x = np.sort(np.random.uniform(0, 4, 30))
y = f(x) + 0.3 * np.random.randn(30)
def cv_mse(deg, k=5):
    folds = np.array_split(np.arange(len(x)), k); errs = []
    for i in range(k):
        te = folds[i]
        tr = np.concatenate([folds[j] for j in range(k) if j != i])
        c = np.polyfit(x[tr], y[tr], deg)
        errs.append(np.mean((np.polyval(c, x[te]) - y[te]) ** 2))
    return np.mean(errs)
scores = {d: round(cv_mse(d), 4) for d in [1, 2, 3, 5, 9]}
print(scores)
print('best degree =', min(scores, key=scores.get))
degreeCV MSE
10.6838
21.5353
30.3597 ← min
58.5814
920,759,668

109. Fill in: CV MSE for Milestone 3 — sweep and pick

Comparison

Comparison matrix

From Milestone 3 — sweep and pick: refill the CV MSE column from what you know. The rest of the table is as it appeared.

degreeCV MSE
10.6838
21.5353
30.3597 ← min
58.5814
920,759,668

110. The full program

Concept

import numpy as np
np.random.seed(2)
f = lambda t: np.sin(1.5 * t)
x = np.sort(np.random.uniform(0, 4, 30))     # 1. noisy data
y = f(x) + 0.3 * np.random.randn(30)

def cv_mse(deg, k=5):                          # 2. k-fold CV MSE
    folds = np.array_split(np.arange(len(x)), k); errs = []
    for i in range(k):
        te = folds[i]
        tr = np.concatenate([folds[j] for j in range(k) if j != i])
        c = np.polyfit(x[tr], y[tr], deg)
        errs.append(np.mean((np.polyval(c, x[te]) - y[te]) ** 2))
    return np.mean(errs)

scores = {d: round(cv_mse(d), 4) for d in [1, 2, 3, 5, 9]}  # 3. sweep
print(scores)
print('best degree =', min(scores, key=scores.get))
printed linevalue
scores{1: 0.6838, 2: 1.5353, 3: 0.3597, 5: 8.5814, 9: 20759668.6189}
best degree =3

CV chose degree 3 — the same sweet spot the bias-variance table found — using only the noisy data. You picked model complexity honestly, without ever seeing the truth.

111. Break it if you can: The full program

Counterexample

Discussion prompt

CV chose degree 3 — the same sweet spot the bias-variance table found — using only the noisy data. You picked model complexity honestly, without ever seeing the truth.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

112. Show it off

Concept

Slides closed, out loud: explain (1) the three terms of MSE = Bias² + Variance + Noise and why the two cross terms vanish, (2) why training error falls monotonically but test error is U-shaped, and (3) why nested CV's outer loop must never see the tuning.

Stretch: swap cv_mse to use StratifiedKFold on a classification target, then wrap the degree sweep in an outer loop to make it a true nested CV. Bias-variance underlies regularization (Week 12) and ensembles (Week 23) — bagging attacks variance, boosting attacks bias.

113. Connect it up: Lesson 32: Bias-Variance & Cross-Validation

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Two ways to be wrong · Derive the decomposition · Verify it numerically · Bias & variance vs complexity · Diagnose from curves · Cross-validation. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

114. What you can do now

Recap

ideathe one thing to remember
decompositionMSE = Bias² + Variance + Noise; cross terms vanish
complexitytest error is U-shaped, not monotone — train error is useless for choice
k-foldeach point held out exactly once; average the folds
strategystratified for classes, time-series for time, LOO for tiny n
nested CVtune inside, evaluate outside — never report the number you tuned on

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 32 (Week 11 — Bias-Variance & Cross-Validation) — Barron · USAAIO Round 2 Preparation, 2026
  2. Hastie, Tibshirani & Friedman, The Elements of Statistical Learning, §7.2–7.3 (Bias, Variance and Model Complexity; the Bias-Variance Decomposition) — Springer, 2nd ed.
  3. scikit-learn User Guide — Cross-validation: KFold, StratifiedKFold, LeaveOneOut, TimeSeriesSplit, cross_val_score, GridSearchCV
  4. Every bias², variance, CV score, and per-fold number produced by real execution — numpy 2.2.6 + scikit-learn 1.9.0, verification run July 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108