Lesson 45: Learning Curves, Bias-Variance Diagnosis, and Double Descent

USAAIO Lesson 45, from Week 16, fully worked on ONE running dataset, y = x^2 - x + noise. It derives the bias-variance decomposition of the expected loss term by term, then builds loss-against-epoch and loss-against-training-size curves from scratch and reads them zone by zone, diagnosing high bias against high variance from the gap between them. It derives and measures the targeted cures - dropout, early stopping with patience, and changes in capacity - and produces double descent with a real minimum-norm interpolator that spikes exactly at P = n and then descends again. Every snippet is self-contained and runs in a fresh interpreter, and every number in every trace table was produced by real execution with fixed seeds. The lesson runs to 63 slides.

Subject: Machine Learning · 85 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Learning Curves & Bias–Variance

Title

USAAIO · Lesson 45 · Week 16

Read the shape of a training run and know exactly what to do next. We derive the bias–variance split, build both learning curves from scratch on one dataset, diagnose the disease from the gap, apply each cure and measure it, and end at double descent — where more parameters help.

2. By the end of this lesson you can

Objectives

  1. Derive the bias–variance decomposition E[(y − ŷ)²] = bias² + variance + noise term by term
  2. Build a loss-vs-epoch curve and a loss-vs-training-size curve in code and read each zone
  3. Diagnose high bias (both errors high, no gap) vs high variance (large gap) from the numbers alone
  4. Apply and measure the right cure — capacity for bias; data, dropout, early stopping for variance
  5. Implement early stopping with patience and track the best checkpoint
  6. Reproduce double descent — a real interpolator that spikes at P = n then generalizes again

3. What survived from MLP Architectures — Depth, Width & Activations?

Warm-up

Discussion prompt

Before we open Lesson 45: Learning Curves, Bias-Variance Diagnosis, and Double Descent: without looking back, what was the main idea of MLP Architectures — Depth, Width & Activations, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

universal approximation theorem, depth vs width expressivity, exact parameter counting for any Linear stack, the four core activation functions (sigmoid, tanh, ReLU, GELU) with computed values and gradient-flow consequences, and USAAIO architecture-design guidelines.

4. The decomposition

Section

Part 1 of 6 — where the two errors come from

5. One dataset for the whole lesson

Concept

Every slide fits the same synthetic dataset so the numbers compound. The true function is f(x) = x² − x, and each observed y is f(x) plus Gaussian noise of standard deviation 2.

\[ y_i = \underbrace{x_i^2 - x_i}_{f(x_i)} + \varepsilon_i, \qquad \varepsilon_i \sim \mathcal{N}(0,\, 2^2) \]

quantityvalue
x rangeuniform on [−3, 3]
true f(x)x² − x (a parabola)
noise sd2.0 (irreducible)
split160 train / 40 val (seed 0)

6. Fill in: value for One dataset for the whole lesson

Comparison

Comparison matrix

From One dataset for the whole lesson: refill the value column from what you know. The rest of the table is as it appeared.

quantityvalue
x rangeuniform on [−3, 3]
true f(x)x² − x (a parabola)
noise sd2.0 (irreducible)
split160 train / 40 val (seed 0)

7. Why you can never hit zero error

Intuition

Even a perfect model that knew f(x) = x² − x exactly would still miss each y, because each y carries its own random noise ε that nothing can predict.

So the total error a model makes splits into two very different kinds: error it could remove with a better model, and error that is baked into the data. Naming those pieces precisely is the bias–variance decomposition.

8. Break it if you can: Why you can never hit zero error

Counterexample

Discussion prompt

Even a perfect model that knew f(x) = x² − x exactly would still miss each y, because each y carries its own random noise ε that nothing can predict.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

9. Three sources of error

Concept

Imagine training your model on many different random training sets, each giving a slightly different fitted predictor. At a fixed test point x, three things drive the expected squared error.

10. By analogy: Three sources of error

Analogy

Discussion prompt

Explain Three sources of error by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Imagine training your model on many different random training sets, each giving a slightly different fitted predictor. At a fixed test point x, three things drive the expected squared error.

11. Guess the shape of the answer: Set up the expectation

Estimation

Predict first

Fix a test point x. Let ŷ = ĥ(x) be the (random, training-set-dependent) prediction and y = f + ε. We want the expected squared error over both the noise and the training randomness:

Commit before you compute: what does Set up the expectation come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Insert ±E[ŷ] — add and subtract the mean prediction

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Writing (f − ŷ) as (f − E[ŷ]) + (E[ŷ] − ŷ) is the standard trick that separates a constant offset (bias) from a zero-mean wobble (variance).

12. Set up the expectation

Worked example

Fix a test point x. Let ŷ = ĥ(x) be the (random, training-set-dependent) prediction and y = f + ε. We want the expected squared error over both the noise and the training randomness:

\[ \mathbb{E}\big[(y - \hat y)^2\big] = \mathbb{E}\big[(f + \varepsilon - \hat y)^2\big] \]

Insert ±E[ŷ] — add and subtract the mean prediction

Why: Writing (f − ŷ) as (f − E[ŷ]) + (E[ŷ] − ŷ) is the standard trick that separates a constant offset (bias) from a zero-mean wobble (variance).

\[ = \mathbb{E}\big[(\,\underbrace{f - \mathbb{E}[\hat y]}_{\text{bias}} + \underbrace{\mathbb{E}[\hat y] - \hat y}_{\text{wobble}} + \varepsilon\,)^2\big] \]

13. Work backwards from the answer: Set up the expectation

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Insert ±E[ŷ] — add and subtract the mean prediction

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Fix a test point x. Let ŷ = ĥ(x) be the (random, training-set-dependent) prediction and y = f + ε. We want the expected squared error over both the noise and the training randomness:

14. What has to happen first: Expand the square, kill the cross terms

Ranking

Put in order

Put the moves of Expand the square, kill the cross terms into the order they have to happen.

  1. Square the three-term sum
  2. Cross term A·B vanishes: A is a constant, E[B] = 0
  3. Cross terms with ε vanish: noise is independent, E[ε] = 0

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Every square-of-a-sum produces three squared terms and three cross terms.

15. Expand the square, kill the cross terms

Worked example

Square the three-term sum

Why: Every square-of-a-sum produces three squared terms and three cross terms. We keep the squares and show each cross term averages to zero.

\[ \mathbb{E}\big[(A + B + \varepsilon)^2\big],\quad A = f - \mathbb{E}[\hat y],\ B = \mathbb{E}[\hat y] - \hat y \]

Cross term A·B vanishes: A is a constant, E[B] = 0

Why: A = f − E[ŷ] does not depend on the training draw, and E[B] = E[E[ŷ] − ŷ] = E[ŷ] − E[ŷ] = 0. A constant times a mean-zero quantity has expectation 0.

\[ \mathbb{E}[A\,B] = A\,\mathbb{E}[B] = A\cdot 0 = 0 \]

Cross terms with ε vanish: noise is independent, E[ε] = 0

Why: ε is independent of both f (fixed) and ŷ (built from a different training set), and it is mean-zero, so E[Aε] = E[Bε] = 0.

16. The decomposition, assembled

Worked example

Only the three squares survive

Why: With every cross term gone, the expectation is the sum of the three squared pieces — each with a name.

\[ \mathbb{E}[(y-\hat y)^2] = \underbrace{(f - \mathbb{E}[\hat y])^2}_{\text{bias}^2} + \underbrace{\mathbb{E}[(\hat y - \mathbb{E}[\hat y])^2]}_{\text{variance}} + \underbrace{\sigma^2}_{\text{noise}} \]

Read the levers

Why: You can shrink bias (use a richer model) or shrink variance (use more data / regularization), but σ² = 4 here is a hard floor no method can cross.

\[ \boxed{\;\text{err} = \text{bias}^2 + \text{variance} + \sigma^2\;} \]

termcured byour value
bias²more model capacitylarge for a line, small for a parabola
variancemore data, regularizationlarge for a wiggly high-degree fit
noise σ²nothing — irreducible2² = 4

17. The trade-off in one picture

Concept

Push model complexity up and bias falls but variance rises; push it down and the reverse. Total error is their sum plus the noise floor — a U-shaped curve with a sweet spot in the middle.

Figure (svg): Three curves over increasing model complexity: bias-squared falling, variance rising, and their sum (total error) forming a U with a minimum, sitting above a flat dashed noise floor.

Bias² falls, variance rises; total error is a U above the irreducible noise floor.

18. Teach it back: The trade-off in one picture

Explain it

Discussion prompt

Explain The trade-off in one picture to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Push model complexity up and bias falls but variance rises; push it down and the reverse. Total error is their sum plus the noise floor — a U-shaped curve with a sweet spot in the middle.

19. What has to be given first: See both terms in code — one dataset, three…

Missing information

Discussion prompt

Fit polynomials of degree 1, 2, and 15 to the running dataset and report train and val MSE. This block is fully self-contained — it rebuilds the dataset itself.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

A line cannot bend to a parabola, so both its errors stay high (bias). Degree 2 IS the true shape — lowest val. Degree 15 fits the noise (train drops, but val barely improves and it is fragile).

20. See both terms in code — one dataset, three degrees

Worked example

Fit polynomials of degree 1, 2, and 15 to the running dataset and report train and val MSE. This block is fully self-contained — it rebuilds the dataset itself.

import numpy as np
rng = np.random.default_rng(0)
x = rng.uniform(-3, 3, size=200)
y = x**2 - x + rng.normal(0, 2.0, size=200)   # true f + noise sd=2
xtr, ytr = x[:160], y[:160]
xva, yva = x[160:], y[160:]

def mse(p, xs, ys):
    return np.mean((np.polyval(p, xs) - ys)**2)

for deg in [1, 2, 15]:
    p = np.polyfit(xtr, ytr, deg)
    print(deg, round(mse(p, xtr, ytr), 4), round(mse(p, xva, yva), 4))

deg 1 is biased; deg 2 matches the truth; deg 15 adds variance

Why: A line cannot bend to a parabola, so both its errors stay high (bias). Degree 2 IS the true shape — lowest val. Degree 15 fits the noise (train drops, but val barely improves and it is fragile).

degreetrain MSEval MSEreading
112.71468.1469high bias — too stiff
24.45883.0343sweet spot — true shape
154.29872.7401wiggly; low bias, watch variance

21. What each one costs: See both terms in code — one dataset, three…

Trade off

Comparison matrix

From See both terms in code — one dataset, three degrees: every row here is a choice with a cost. Fill the train MSE column, then say which row you would actually pick and what you give up for it.

degreetrain MSEval MSEreading
112.71468.1469high bias — too stiff
24.45883.0343sweet spot — true shape
154.29872.7401wiggly; low bias, watch variance

22. Reading learning curves

Section

Part 2 of 6 — two axes, two questions

23. Two curves, two different questions

Concept

A learning curve plots error against one of two axes, and each answers a different question about your run.

x-axisquestion it answerswhen to use
training epochsIs the run converging? When does overfitting start?monitor one training run
training-set size nIs this bias or variance? Will more data help?decide whether to collect data

Both plot train and val error together. The gap between them is the signal — never read one curve in isolation.

24. Loss vs epochs — three zones

Concept

In a healthy run both curves fall, then val flattens or turns up while train keeps dropping. That turn is the onset of memorization.

zonetrain lossval lossdiagnosis
earlyfalling fastfalling fastlearning real signal
middlestill fallingflattensgeneralization saturated
latekeeps fallingflat or risingmemorizing noise

Early stopping lives in the middle zone: halt at the val minimum, before the late zone eats your generalization.

25. Predict the next row: Build the loss-vs-epoch curve

Pattern

Predict first

The table runs: 1 | 24.0913 | 14.9322 | early · 20 | 11.2770 | 6.4902 | early · 50 | 5.9076 | 3.3293 | early/mid · 100 | 4.5213 | 2.8274 | val minimum · 200 | 4.4184 | 2.8624 | flat

In Build the loss-vs-epoch curve, given the rows so far: what is the next one — the row where epoch is 400?

Correct: 400 | 4.3722 | 2.8931 | flat — stop earlier

epochtrain MSEval MSEzone
124.091314.9322early
2011.27706.4902early
505.90763.3293early/mid
1004.52132.8274val minimum
2004.41842.8624flat
4004.37222.8931flat — stop earlier

Why: The relationship between the columns, not the individual numbers, is what generates the next row. By epoch 100 val has flattened at ~2.83 (below train because the val set here is easier); training longer to epoch 400 does NOT lower val — the middle/late boundary is around epoch 100.

26. Build the loss-vs-epoch curve

Worked example

Train a small Tanh MLP with Adam on the running dataset and log train and val MSE at chosen epochs. Self-contained; seeds fixed so your numbers match these.

import numpy as np, torch, torch.nn as nn
rng = np.random.default_rng(0)
x = rng.uniform(-3, 3, size=200).astype('float32').reshape(-1, 1)
y = (x**2 - x + rng.normal(0, 2.0, size=(200, 1))).astype('float32')
Xtr, ytr = torch.tensor(x[:160]), torch.tensor(y[:160])
Xva, yva = torch.tensor(x[160:]), torch.tensor(y[160:])
torch.manual_seed(0)
model = nn.Sequential(nn.Linear(1, 32), nn.Tanh(), nn.Linear(32, 1))
opt = torch.optim.Adam(model.parameters(), lr=0.02)
lossf = nn.MSELoss()
for ep in range(1, 401):
    model.train(); opt.zero_grad()
    tl = lossf(model(Xtr), ytr); tl.backward(); opt.step()
    if ep in (1, 20, 50, 100, 200, 400):
        model.eval()
        with torch.no_grad():
            vl = lossf(model(Xva), yva).item()
        print(ep, round(tl.item(), 4), round(vl, 4))

Both losses fall fast, then flatten near the σ² floor

Why: By epoch 100 val has flattened at ~2.83 (below train because the val set here is easier); training longer to epoch 400 does NOT lower val — the middle/late boundary is around epoch 100.

epochtrain MSEval MSEzone
124.091314.9322early
2011.27706.4902early
505.90763.3293early/mid
1004.52132.8274val minimum
2004.41842.8624flat
4004.37222.8931flat — stop earlier

27. The gap is the diagnosis

Intuition

Forget the absolute numbers for a second. Ask only two things: is the floor high? and is there a gap between train and val? Those two yes/no answers name the disease.

High floor + no gap = the model can't fit the signal (bias). Low train + big gap = the model fits the training noise but not the world (variance). That is the entire diagnostic.

28. Bias vs Variance — the diagnosis

Section

Part 3 of 6 — read it off the gap

29. High bias (underfitting): both high, no gap

Concept

A high-bias model is too simple. Train and val error are both high and roughly equal — no gap, but a wrong floor.

\[ \varepsilon_{\text{train}} \approx \varepsilon_{\text{val}} \gg \sigma^2 \]

On the set-size curve, both curves plateau high and refuse to fall as n grows. More data cannot help — the model lacks the capacity to represent f.

30. High variance (overfitting): big gap

Concept

A high-variance model memorizes the training set. Train error is low, val error is far higher. The gap is the signature.

\[ \varepsilon_{\text{train}} \ll \varepsilon_{\text{val}} \]

On the set-size curve, as n grows the val error falls toward the train error and the gap closes. More data is the direct cure for pure variance.

31. Guess the shape of the answer: Build the set-size curve — degree 1 vs…

Estimation

Predict first

Fit a degree-1 model (stiff) and a degree-8 model (flexible) on growing training sets n ∈ {10, 30, 100, 300}, scored on a fixed 100-point val set. Self-contained.

Commit before you compute: what does Build the set-size curve — degree 1 vs degree 8 come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: deg 1 val stays ~12→10.7 (bias floor); deg 8 val collapses 276→4.26 (variance cured by data)

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The degree-1 val barely moves as n grows 30× — a bias floor.

32. Build the set-size curve — degree 1 vs degree 8

Worked example

Fit a degree-1 model (stiff) and a degree-8 model (flexible) on growing training sets n ∈ {10, 30, 100, 300}, scored on a fixed 100-point val set. Self-contained.

import numpy as np
rng = np.random.default_rng(0)
x = rng.uniform(-3, 3, size=400)
y = x**2 - x + rng.normal(0, 2.0, size=400)
xva, yva = x[300:], y[300:]                # fixed 100-point val set
pool_x, pool_y = x[:300], y[:300]

def mse(p, xs, ys):
    return np.mean((np.polyval(p, xs) - ys)**2)

for n in [10, 30, 100, 300]:
    for deg in [1, 8]:
        p = np.polyfit(pool_x[:n], pool_y[:n], deg)
        tr = mse(p, pool_x[:n], pool_y[:n])
        vl = mse(p, xva, yva)
        print(n, deg, round(tr, 4), round(vl, 4))

deg 1 val stays ~12→10.7 (bias floor); deg 8 val collapses 276→4.26 (variance cured by data)

Why: The degree-1 val barely moves as n grows 30× — a bias floor. The degree-8 val plummets from 276.8 (wild overfit at n=10) to 4.26 as data fills in the wiggles.

ndeg-1 traindeg-1 valdeg-8 traindeg-8 val
1011.641012.64840.0306276.7645
3014.988012.13853.18254.3133
10013.961310.88004.41774.3799
30012.081310.67673.88734.2604

33. Watch it run: Build the set-size curve — degree 1 vs degree 8

Pattern

Step through it

Step through Build the set-size curve — degree 1 vs degree 8 one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: n is 10
  2. Step 2: n is 30
  3. Step 3: n is 100
  4. Step 4: n is 300

34. Same numbers, two stories

Intuition

Look down the two val columns. Degree-1's val is a nearly flat line near 11 — pour in 30× the data and nothing improves. That flatness is bias.

Degree-8's val falls off a cliff from 276 to 4.26. The collapse toward the train curve as n grows is variance being cured by data. Two identical experiments, opposite prescriptions.

35. Rebuild the recipe: The diagnose-then-cure recipe

Ranking

Put in order

These are the steps of The diagnose-then-cure recipe, scrambled. Put them back in order before the next slide shows you.

  1. Plot both curves — loss vs epoch (convergence) and loss vs n (bias/variance)
  2. Measure the gap val − train and the floor (level both curves sit at)
  3. High floor, small gap → bias: add capacity, richer features, or train longer; do NOT collect data
  4. Low train, large gap → variance: add data, dropout, weight_decay, or early stopping
  5. Confirm with the set-size curve: gap shrinks with n ⇒ variance; floor stays ⇒ bias

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

36. The diagnose-then-cure recipe

Pattern

  1. Plot both curves — loss vs epoch (convergence) and loss vs n (bias/variance)
  2. Measure the gap val − train and the floor (level both curves sit at)
  3. High floor, small gap → bias: add capacity, richer features, or train longer; do NOT collect data
  4. Low train, large gap → variance: add data, dropout, weight_decay, or early stopping
  5. Confirm with the set-size curve: gap shrinks with n ⇒ variance; floor stays ⇒ bias

37. Where does it stop working: The diagnose-then-cure recipe

Edge cases

Discussion prompt

The diagnose-then-cure recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Plot both curves — loss vs epoch (convergence) and loss vs n (bias/variance)
  2. Measure the gap val − train and the floor (level both curves sit at)
  3. High floor, small gap → bias: add capacity, richer features, or train longer; do NOT collect data
  4. Low train, large gap → variance: add data, dropout, weight_decay, or early stopping
  5. Confirm with the set-size curve: gap shrinks with n ⇒ variance; floor stays ⇒ bias

38. Something is wrong here: "more data always helps"

Anomaly

Predict first

A student writes this, and it looks reasonable:

Train MSE is high (~12) and val MSE is equally high (~12). The model is doing badly, so collect 10× more data to fix it.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Both errors are high AND equal — there is no gap.

Diagnose with the gap FIRST. No gap at a high floor means bias, and bias is a capacity problem.

Why: Both errors are high AND equal — there is no gap. The measured degree-1 val goes 12.65 → 12.14 → 10.88 → 10.68 as n runs 10 → 300: a flat bias floor. Thirty times the data bought almost nothing. Data-collection effort wasted.

39. Trap: "more data always helps"

Trap

The trap

Train MSE is high (~12) and val MSE is equally high (~12). The model is doing badly, so collect 10× more data to fix it.

Add data to a high-bias model

Why: Both errors are high AND equal — there is no gap. The measured degree-1 val goes 12.65 → 12.14 → 10.88 → 10.68 as n runs 10 → 300: a flat bias floor. Thirty times the data bought almost nothing. Data-collection effort wasted.

The fix

Diagnose with the gap FIRST. No gap at a high floor means bias, and bias is a capacity problem.

Gap rule: big gap → data/regularization; no gap + high floor → add capacity

Why: The degree-8 val fell 276 → 4.26 with more data (that's variance, curable by n). The degree-1 floor did not. Only richer models or better features cure bias. Match the cure to the disease.

40. Rule out three: Check 1 — read the curve

Elimination

Eliminate the wrong options

A model shows train MSE = 11.6 and val MSE = 12.6 at n = 10. You grow the training set to n = 300 and both settle near 11–12. What is the primary problem?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. High bias — the model is too simple; add capacity, not data
  • B. High variance — the model is overfitting; add more data
  • C. Both high bias and high variance at once
  • D. No problem — the model has converged healthily

Survives elimination: A

Why: Train and val are close (gap ≈ 1) so there is no overfitting, yet 30× more data leaves the error near 11–12 — the floor will not fall. That flat, high floor with no gap is the textbook high-bias signature (this is exactly the measured degree-1 set-size curve). Cure: more capacity or better features.

41. Check 1 — read the curve

Check

Use the gap-and-floor rule. Don't just look at whether the numbers are 'big'.

Check your understanding

A model shows train MSE = 11.6 and val MSE = 12.6 at n = 10. You grow the training set to n = 300 and both settle near 11–12. What is the primary problem?

  • A. High bias — the model is too simple; add capacity, not data (correct)
  • B. High variance — the model is overfitting; add more data
  • C. Both high bias and high variance at once
  • D. No problem — the model has converged healthily

Answer: A

Why: Train and val are close (gap ≈ 1) so there is no overfitting, yet 30× more data leaves the error near 11–12 — the floor will not fall. That flat, high floor with no gap is the textbook high-bias signature (this is exactly the measured degree-1 set-size curve). Cure: more capacity or better features.

Why B tempts people
High variance requires a LARGE gap (e.g. train 0.03, val 276 — the degree-8 case at n=10). Here the gap is ~1 and adding data did not help, so overfitting is not the issue.
Why C tempts people
Simultaneous high bias and high variance needs BOTH a high floor and a large gap. Here the gap is tiny (~1), so variance is not present.
Why D tempts people
It did converge, but to a bad floor. 'Healthy' means low error; a stable error of ~12 with an irreducible σ² of only 4 signals a structural capacity problem, not health.

42. Cures for overfitting

Section

Part 4 of 6 — dropout & early stopping

43. The cure menu (match to cost)

Concept

Pick by what is cheap: if data is free, collect it. If the model is already large and data is scarce, dropout + early stopping are the fastest levers — and we measure both next.

44. Dropout: the mechanism

Concept

During training, each activation is zeroed with probability p on every forward pass and the survivors are scaled by 1/(1−p). During eval, dropout is off — the scaling already kept the expected magnitude right.

\[ \tilde h_i = \frac{m_i}{1-p}\, h_i, \qquad m_i \sim \text{Bernoulli}(1-p) \]

No neuron can rely on any single partner, so the network learns redundant, distributed features. It is noise injection that lowers the effective capacity at train time.

45. Restore the missing line: Dropout closes the gap — measured

Fill the middle

Fill in the blanks

From Dropout closes the gap — measured — one line has had its right-hand side removed. Put it back.

import numpy as np, torch, torch.nn as nn
rng = np.random.default_rng(0)
x = rng.uniform(-3, 3, size=40).astype('float32').reshape(-1, 1)
y = (x2 - x + rng.normal(0, 2.0, size=(40, 1))).astype('float32')
xv = rng.uniform(-3, 3, size=200).astype('float32').reshape(-1, 1)
yv = (xv
2 - xv + rng.normal(0, 2.0, size=(200, 1))).astype('float32')
Xt, yt, Xvt, yvt = map(torch.tensor, (x, y, xv, yv))
lossf = nn.MSELoss()

def make(drop):
L = [nn.Linear(1, 128), nn.ReLU()]
if drop: L.append(nn.Dropout(drop))
L += [nn.Linear(128, 128), nn.ReLU()]
if drop: L.append(nn.Dropout(drop))
L.append(nn.Linear(128, 1))
return nn.Sequential(*L)

for p in (0.0, 0.5):
torch.manual_seed(0)
m = make(p); opt = torch.optim.Adam(m.parameters(), lr=0.01)
for _ in range(800):
m.train(); opt.zero_grad(); lossf(m(Xt), yt).backward(); opt.step()
m.eval()
with torch.no_grad():
tr = lossf(m(Xt), yt).item(); vl = lossf(m(Xvt), yvt).item()
print(p, round(tr, 4), round(vl, 4), round(vl - tr, 4))

Why: lossf is what everything below it consumes, so the wrong expression here fails later and somewhere else. Without dropout the net drives train MSE to 2.45 but val sits at 7.35 — a 4.9 gap (memorization).

46. Dropout closes the gap — measured

Worked example

Train an overparameterized 128→128→1 ReLU net on just 40 points, once without dropout and once with Dropout(0.5). Same seed, same everything else. Self-contained.

import numpy as np, torch, torch.nn as nn
rng = np.random.default_rng(0)
x  = rng.uniform(-3, 3, size=40).astype('float32').reshape(-1, 1)
y  = (x**2 - x + rng.normal(0, 2.0, size=(40, 1))).astype('float32')
xv = rng.uniform(-3, 3, size=200).astype('float32').reshape(-1, 1)
yv = (xv**2 - xv + rng.normal(0, 2.0, size=(200, 1))).astype('float32')
Xt, yt, Xvt, yvt = map(torch.tensor, (x, y, xv, yv))
lossf = nn.MSELoss()

def make(drop):
    L = [nn.Linear(1, 128), nn.ReLU()]
    if drop: L.append(nn.Dropout(drop))
    L += [nn.Linear(128, 128), nn.ReLU()]
    if drop: L.append(nn.Dropout(drop))
    L.append(nn.Linear(128, 1))
    return nn.Sequential(*L)

for p in (0.0, 0.5):
    torch.manual_seed(0)
    m = make(p); opt = torch.optim.Adam(m.parameters(), lr=0.01)
    for _ in range(800):
        m.train(); opt.zero_grad(); lossf(m(Xt), yt).backward(); opt.step()
    m.eval()
    with torch.no_grad():
        tr = lossf(m(Xt), yt).item(); vl = lossf(m(Xvt), yvt).item()
    print(p, round(tr, 4), round(vl, 4), round(vl - tr, 4))

No dropout: gap = 4.90. Dropout 0.5: gap = 0.36

Why: Without dropout the net drives train MSE to 2.45 but val sits at 7.35 — a 4.9 gap (memorization). Dropout raises train to 4.83 yet drops val to 5.19: it trades a little train fit for a 13× smaller gap.

dropout ptrain MSEval MSEgap (val − train)
0.0 (none)2.44897.34674.8978
0.54.82505.18700.3621

47. Something is wrong here: leaving dropout on at eval

Anomaly

Predict first

A student writes this, and it looks reasonable:

Dropout is part of the model, so evaluate the same way you trained — just call model(Xval) right after the training loop, model still in train mode.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: In train mode dropout keeps randomly zeroing activations, so every val score is noisy and randomly too high — and it changes run to run.

Switch to eval mode and disable grad before validating; switch back to train before the next step.

Why: In train mode dropout keeps randomly zeroing activations, so every val score is noisy and randomly too high — and it changes run to run. You are measuring a crippled network, not the one you will deploy.

48. Trap: leaving dropout on at eval

Trap

The trap

Dropout is part of the model, so evaluate the same way you trained — just call model(Xval) right after the training loop, model still in train mode.

Score val with the module in train() mode

Why: In train mode dropout keeps randomly zeroing activations, so every val score is noisy and randomly too high — and it changes run to run. You are measuring a crippled network, not the one you will deploy.

The fix

Switch to eval mode and disable grad before validating; switch back to train before the next step.

model.eval(); with torch.no_grad(): score val

Why: eval() turns dropout OFF and uses the full, scaled network — the deterministic predictor you actually ship. torch.no_grad() just saves memory. This is why the measured val above is stable and reproducible.

49. Break it on purpose: leaving dropout on at eval

Break the constraint

Discussion prompt

The rule this trap just fixed:

Switch to eval mode and disable grad before validating; switch back to train before the next step.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

In train mode dropout keeps randomly zeroing activations, so every val score is noisy and randomly too high — and it changes run to run. You are measuring a crippled network, not the one you will deploy.

50. Early stopping with patience

Concept

Watch val loss. Remember the best value seen and its epoch. Each epoch that fails to beat it (by a small threshold) increments a wait counter; reset wait to 0 whenever you improve. Stop once wait reaches patience.

Patience gives the run room to escape a temporary plateau — it does not quit on the first bad epoch, only after patience consecutive non-improvements.

51. What has to be given first: Early stopping — measured on a run that…

Missing information

Discussion prompt

Use a deliberately high Adam lr = 0.05 so val bottoms out early then drifts up. patience = 20, improvement threshold 1e-3. Self-contained.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

After epoch 14 no epoch beats 2.8609 by 1e-3, so wait climbs 1..20 while val drifts up. At epoch 34 wait hits patience and training halts — you keep the epoch-14 checkpoint, not the worse epoch-34 weights.

52. Early stopping — measured on a run that overshoots

Worked example

Use a deliberately high Adam lr = 0.05 so val bottoms out early then drifts up. patience = 20, improvement threshold 1e-3. Self-contained.

import numpy as np, torch, torch.nn as nn
rng = np.random.default_rng(0)
x = rng.uniform(-3, 3, size=200).astype('float32').reshape(-1, 1)
y = (x**2 - x + rng.normal(0, 2.0, size=(200, 1))).astype('float32')
Xtr, ytr = torch.tensor(x[:160]), torch.tensor(y[:160])
Xva, yva = torch.tensor(x[160:]), torch.tensor(y[160:])
torch.manual_seed(0)
model = nn.Sequential(nn.Linear(1, 128), nn.ReLU(), nn.Linear(128, 1))
opt = torch.optim.Adam(model.parameters(), lr=0.05)   # high on purpose
lossf = nn.MSELoss()
best_val, best_ep, wait, patience = float('inf'), 0, 0, 20
for ep in range(1, 1001):
    model.train(); opt.zero_grad()
    lossf(model(Xtr), ytr).backward(); opt.step()
    model.eval()
    with torch.no_grad():
        vl = lossf(model(Xva), yva).item()
    if vl < best_val - 1e-3:
        best_val, best_ep, wait = vl, ep, 0
    else:
        wait += 1
    if wait >= patience:
        print('STOP', ep, 'best', round(best_val, 4), 'at', best_ep)
        break

Val bottoms at 2.8609 (epoch 14); 20 quiet epochs later, STOP at 34

Why: After epoch 14 no epoch beats 2.8609 by 1e-3, so wait climbs 1..20 while val drifts up. At epoch 34 wait hits patience and training halts — you keep the epoch-14 checkpoint, not the worse epoch-34 weights.

epochval lossimproved?wait
14.6386yes (best)0
142.8609yes (best)0
153.0364no1
24≈ higherno10
34≈ higherno — STOP20

53. Fill in: wait for Early stopping — measured on a run that…

Comparison

Comparison matrix

From Early stopping — measured on a run that overshoots: refill the wait column from what you know. The rest of the table is as it appeared.

epochval lossimproved?wait
14.6386yes (best)0
142.8609yes (best)0
153.0364no1
24≈ higherno10
34≈ higherno — STOP20

54. Something is wrong here: watching train loss for early stopping

Anomaly

Predict first

A student writes this, and it looks reasonable:

Track train_loss and stop when it stops decreasing — it's the loss you're optimizing, so it's the natural thing to watch.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: With enough capacity train loss keeps sliding down toward 0 and never meaningfully plateaus until the model has fully memorized.

Track val_loss. Stop when the held-out signal stops improving.

Why: With enough capacity train loss keeps sliding down toward 0 and never meaningfully plateaus until the model has fully memorized. You either stop far too late (already overfit) or crank the threshold so tight you stop at epoch 2.

55. Trap: watching train loss for early stopping

Trap

The trap

Track train_loss and stop when it stops decreasing — it's the loss you're optimizing, so it's the natural thing to watch.

Set patience on the train-loss plateau

Why: With enough capacity train loss keeps sliding down toward 0 and never meaningfully plateaus until the model has fully memorized. You either stop far too late (already overfit) or crank the threshold so tight you stop at epoch 2.

The fix

Track val_loss. Stop when the held-out signal stops improving.

Track best val_loss; reset patience on improvement; stop when wait ≥ patience

Why: Val loss is the only quantity that knows about the held-out distribution, so its turn-up is the real onset of memorization (here epoch 14). Train loss has no memory of unseen data.

56. Which of these survive contact with Lesson 45: Learning Curves, Bias-Variance…?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Even a perfect model that knew f(x) = x² − x exactly would still miss each y, because each y carries its own random noise ε that nothing can predict.; A learning curve plots error against one of two axes, and each answers a different question about your run.; In a healthy run both curves fall, then val flattens or turns up while train keeps dropping. That turn is the onset of memorization.
Breaks
Train MSE is high (~12) and val MSE is equally high (~12). The model is doing badly, so collect 10× more data to fix it.; Dropout is part of the model, so evaluate the same way you trained — just call model(Xval) right after the training loop, model still in train mode.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 45: Learning Curves, Bias-Variance Diagnosis, and Double Descent puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

57. Answer it before you see the options: Check 2 — trace the patience counter

Prediction

Predict first

patience = 3, and the per-epoch val losses are [2.1, 1.8, 1.9, 1.7, 1.8, 1.9, 2.0]. At which epoch does early stopping trigger?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: Epoch 7

Why: Trace: ep1 2.1 best (wait 0), ep2 1.8 best (0), ep3 1.9 no (1), ep4 1.7 best (0), ep5 1.8 no (1), ep6 1.9 no (2), ep7 2.0 no (3 ≥ patience → STOP). It fires at epoch 7.

58. Check 2 — trace the patience counter

Check

Walk the counter epoch by epoch. A new best resets wait to 0.

Check your understanding

patience = 3, and the per-epoch val losses are [2.1, 1.8, 1.9, 1.7, 1.8, 1.9, 2.0]. At which epoch does early stopping trigger?

  • A. Epoch 7 (correct)
  • B. Epoch 6
  • C. Epoch 5
  • D. Epoch 4

Answer: A

Why: Trace: ep1 2.1 best (wait 0), ep2 1.8 best (0), ep3 1.9 no (1), ep4 1.7 best (0), ep5 1.8 no (1), ep6 1.9 no (2), ep7 2.0 no (3 ≥ patience → STOP). It fires at epoch 7.

Why B tempts people
At epoch 6 the counter is only 2 (from ep5 and ep6), one short of the threshold of 3.
Why C tempts people
At epoch 5 the counter is 1 — epoch 4 was a new best (1.7) and reset wait to 0.
Why D tempts people
Epoch 4 is itself a new best (1.7), which resets the counter to 0 rather than triggering a stop.

59. Double descent

Section

Part 5 of 6 — past the interpolation threshold

60. The classical U — and where it breaks

Concept

Classical theory says test error is U-shaped in complexity: underfit → sweet spot → overfit, blowing up as you approach enough parameters to fit the data exactly.

That blow-up point — where the number of parameters P equals the number of training points n — is the interpolation threshold. Classical theory stops there. Modern overparameterized models keep going past it, and something surprising happens.

61. Picture it first: Double descent: error falls again past P = n

Picture it

Figure (svg): Test error curve versus number of parameters P: a first descent, a sharp spike at P equals n (the interpolation threshold, marked with a dashed vertical line), then a second descent to a low plateau.

Test error: first descent, a spike at P = n, then a second descent as min-norm smooths the fit.

Discussion prompt

Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.

Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.

Answer:

Push P well beyond n and test error, after spiking at P = n, descends a second time. When infinitely many zero-training-error fits exist, least squares picks the minimum-norm one — the smoothest interpolator, which generalizes.

62. Double descent: error falls again past P = n

Concept

Push P well beyond n and test error, after spiking at P = n, descends a second time. When infinitely many zero-training-error fits exist, least squares picks the minimum-norm one — the smoothest interpolator, which generalizes.

Figure (svg): Test error curve versus number of parameters P: a first descent, a sharp spike at P equals n (the interpolation threshold, marked with a dashed vertical line), then a second descent to a low plateau.

Test error: first descent, a spike at P = n, then a second descent as min-norm smooths the fit.

63. Guess the shape of the answer: Reproduce double descent — a real…

Estimation

Predict first

Fit a step target with min-norm least squares over P Chebyshev features on n = 20 points. Sweep P through the threshold. lstsq returns the minimum-norm solution when the system is under-determined. Self-contained.

Commit before you compute: what does Reproduce double descent — a real interpolator come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Test peaks at P = n = 20 (te ≈ 101, ‖w‖ ≈ 34) then descends to ≈ 0.31 at P = 25

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Below P=20 train error is nonzero (can't interpolate).

64. Reproduce double descent — a real interpolator

Worked example

Fit a step target with min-norm least squares over P Chebyshev features on n = 20 points. Sweep P through the threshold. lstsq returns the minimum-norm solution when the system is under-determined. Self-contained.

import numpy as np
rng = np.random.default_rng(0)
n = 20
xtr = np.sort(rng.uniform(-1, 1, size=n))
ytr = np.sign(xtr)                       # a step: hard for low-degree polys
xte = np.linspace(-1, 1, 1000)
yte = np.sign(xte)

def design(x, P):                        # P Chebyshev features T_0..T_{P-1}
    return np.polynomial.chebyshev.chebvander(x, P - 1)

for P in [3, 8, 15, 19, 20, 21, 25, 40, 100, 300]:
    Ftr, Fte = design(xtr, P), design(xte, P)
    w, *_ = np.linalg.lstsq(Ftr, ytr, rcond=None)   # min-norm solution
    tr = np.mean((Ftr @ w - ytr)**2)
    te = np.mean((Fte @ w - yte)**2)
    print(P, round(tr, 5), round(te, 3), round(np.linalg.norm(w), 2))

Test peaks at P = n = 20 (te ≈ 101, ‖w‖ ≈ 34) then descends to ≈ 0.31 at P = 25

Why: Below P=20 train error is nonzero (can't interpolate). At P≈20 it interpolates but ‖w‖ explodes and test blows up. Past P=20, extra features let lstsq pick a SMALLER-norm interpolator — ‖w‖ falls 34 → 0.34 and test drops to ~0.3–0.9.

P (features)train MSEtest MSE‖w‖regime
80.026970.1841.33underparam (1st descent)
150.000305.3957.45approaching threshold
200.00000101.06234.31P = n → SPIKE
250.000000.3121.21overparam (2nd descent)
3000.000000.9360.34deep overparam, min-norm

65. Work backwards from the answer: Reproduce double descent — a real…

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Test peaks at P = n = 20 (te ≈ 101, ‖w‖ ≈ 34) then descends to ≈ 0.31 at P = 25

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Fit a step target with min-norm least squares over P Chebyshev features on n = 20 points. Sweep P through the threshold. lstsq returns the minimum-norm solution when the system is under-determined. Self-contained.

66. Why the norm tells the story

Intuition

Watch the ‖w‖ column, not just the error. It climbs to 34 right at P = n — the lone interpolator there is forced to use enormous, oscillating coefficients to thread all 20 points exactly.

Add more features and there are now many exact fits; lstsq takes the smallest-norm one, so ‖w‖ collapses to 0.34. Small norm means a smooth function, and smooth functions generalize. That is the whole mechanism behind the second descent.

67. Teach it back: Why the norm tells the story

Explain it

Discussion prompt

Explain Why the norm tells the story to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Watch the ‖w‖ column, not just the error. It climbs to 34 right at P = n — the lone interpolator there is forced to use enormous, oscillating coefficients to thread all 20 points exactly.

68. Why giant models don't overfit

Concept

A transformer with billions of parameters trained on billions of tokens sits deep in the second descent (P ≫ n, with n huge). Gradient descent has an implicit min-norm bias, so it lands on a smooth interpolator that generalizes.

The catch is the spike: you must have n large enough to be past the threshold. Overparameterizing a model on a tiny dataset lands you on the peak, not the plateau — exactly the P = n = 20 disaster above.

69. Rule out three: Check 3 — apply double descent

Elimination

Eliminate the wrong options

A 200M-parameter transformer trained on 10B tokens reaches low test loss. Classical bias–variance theory predicts severe overfitting. Why doesn't it overfit?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. It is deep in the double-descent overparameterized regime; with P ≫ n and n large, gradient descent finds a min-norm interpolator that generalizes
  • B. Weight decay alone removes all overfitting regardless of P and n
  • C. More parameters always lower test error, whatever the dataset size
  • D. Classical theory is right; it is actually overfitting and the test set is mismeasured

Survives elimination: A

Why: With P ≫ n and n very large (10B tokens), the model is far past the interpolation threshold on the second descent. Among the many zero-loss fits, gradient descent's implicit bias selects a minimum-norm (smooth) solution — exactly the ‖w‖ 34 → 0.34 collapse measured above. The large n is what keeps it off the spike.

70. Check 3 — apply double descent

Check

Reason from the mechanism you just measured, not from a slogan.

Check your understanding

A 200M-parameter transformer trained on 10B tokens reaches low test loss. Classical bias–variance theory predicts severe overfitting. Why doesn't it overfit?

  • A. It is deep in the double-descent overparameterized regime; with P ≫ n and n large, gradient descent finds a min-norm interpolator that generalizes (correct)
  • B. Weight decay alone removes all overfitting regardless of P and n
  • C. More parameters always lower test error, whatever the dataset size
  • D. Classical theory is right; it is actually overfitting and the test set is mismeasured

Answer: A

Why: With P ≫ n and n very large (10B tokens), the model is far past the interpolation threshold on the second descent. Among the many zero-loss fits, gradient descent's implicit bias selects a minimum-norm (smooth) solution — exactly the ‖w‖ 34 → 0.34 collapse measured above. The large n is what keeps it off the spike.

Why B tempts people
Weight decay helps, but it is not the reason a 200M-parameter net generalizes. The overparameterized regime plus the min-norm inductive bias is the deeper cause; a net with no weight decay still shows double descent.
Why C tempts people
False — with small n, adding parameters drives you toward the interpolation spike (the P = n = 20 case hit test MSE ≈ 101). More parameters help only when n is large enough to be past the threshold.
Why D tempts people
Double descent is a reproducible, measured phenomenon (Belkin et al., 2019, and the snippet above). Dismissing it contradicts the direct evidence.

71. Your turn: diagnose and cure

Section

Part 6 of 6 — the project

72. Project: full learning-curve analysis

Concept

On the running dataset, build the set-size curve for two model complexities, then implement early stopping on an MLP — reproducing the exact numbers from Parts 3 and 4 yourself.

#milestonetool
1Set-size curve: degree-1 vs degree-8numpy.polyfit, vary n
2Early stopping with patience = 20torch, best-val counter
3Cross-check the curve against sklearnPolynomialFeatures + LinearRegression

Build rules: type every line, run after each milestone, and always use model.train() for the step and model.eval() + torch.no_grad() for validation — never mix them.

73. By analogy: Project: full learning-curve analysis

Analogy

Discussion prompt

Explain Project: full learning-curve analysis by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

On the running dataset, build the set-size curve for two model complexities, then implement early stopping on an MLP — reproducing the exact numbers from Parts 3 and 4 yourself.

74. Milestone 1 — the set-size curve

Worked example

Your turn: fit degree-1 and degree-8 polynomials for n ∈ {10, 30, 100, 300} on a fixed 100-point val set. Predict aloud: which degree's val error collapses as n grows?

Hint: rebuild the dataset with default_rng(0), hold out x[300:] as val, and loop n over the pool. Use np.polyfit and np.polyval.

import numpy as np
rng = np.random.default_rng(0)
x = rng.uniform(-3, 3, size=400)
y = x**2 - x + rng.normal(0, 2.0, size=400)
xva, yva = x[300:], y[300:]
pool_x, pool_y = x[:300], y[:300]

def mse(p, xs, ys):
    return np.mean((np.polyval(p, xs) - ys)**2)

for n in [10, 30, 100, 300]:
    p1 = np.polyfit(pool_x[:n], pool_y[:n], 1)
    p8 = np.polyfit(pool_x[:n], pool_y[:n], 8)
    print(n, round(mse(p1, xva, yva), 3), round(mse(p8, xva, yva), 3))
ndeg-1 valdeg-8 val
1012.648276.765
3012.1394.313
10010.8804.380
30010.6774.260

75. Watch it run: Milestone 1 — the set-size curve

Pattern

Step through it

Step through Milestone 1 — the set-size curve one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: n is 10
  2. Step 2: n is 30
  3. Step 3: n is 100
  4. Step 4: n is 300

76. Milestone 2 — early stopping with patience

Worked example

Your turn: implement the patience loop on the 200-point MLP with lr = 0.05, patience = 20. Predict: which epoch holds the best val, and roughly when does it stop?

Hint: best_val = float('inf'), wait = 0; reset wait on improvement (threshold 1e-3); break when wait >= patience. Remember eval() + no_grad() for the val score.

import numpy as np, torch, torch.nn as nn
rng = np.random.default_rng(0)
x = rng.uniform(-3, 3, size=200).astype('float32').reshape(-1, 1)
y = (x**2 - x + rng.normal(0, 2.0, size=(200, 1))).astype('float32')
Xtr, ytr = torch.tensor(x[:160]), torch.tensor(y[:160])
Xva, yva = torch.tensor(x[160:]), torch.tensor(y[160:])
torch.manual_seed(0)
m = nn.Sequential(nn.Linear(1, 128), nn.ReLU(), nn.Linear(128, 1))
opt = torch.optim.Adam(m.parameters(), lr=0.05); lf = nn.MSELoss()
best, bep, wait = float('inf'), 0, 0
for ep in range(1, 1001):
    m.train(); opt.zero_grad(); lf(m(Xtr), ytr).backward(); opt.step()
    m.eval()
    with torch.no_grad(): vl = lf(m(Xva), yva).item()
    if vl < best - 1e-3: best, bep, wait = vl, ep, 0
    else: wait += 1
    if wait >= 20:
        print('STOP', ep, 'best', round(best, 4), 'at', bep); break
outputvalue
stop epoch34
best val2.8609
best epoch14

77. Watch it run: Milestone 2 — early stopping with patience

Pattern

Step through it

Step through Milestone 2 — early stopping with patience one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: output is stop epoch
  2. Step 2: output is best val
  3. Step 3: output is best epoch

78. Milestone 3 — cross-check against sklearn

Worked example

Your turn: rebuild the set-size fits with an sklearn Pipeline and confirm the val numbers match your polyfit results to the decimal. Predict: exact match or just close?

Hint: make_pipeline(PolynomialFeatures(deg), LinearRegression()); reshape x to (-1, 1); score with mean squared residual on the same val set.

import numpy as np
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import PolynomialFeatures
from sklearn.linear_model import LinearRegression
rng = np.random.default_rng(0)
x = rng.uniform(-3, 3, size=400).reshape(-1, 1)
y = (x**2 - x).ravel() + rng.normal(0, 2.0, size=400)
xva, yva = x[300:], y[300:]
px, py = x[:300], y[:300]
for n in [30, 300]:
    for deg in [1, 8]:
        m = make_pipeline(PolynomialFeatures(deg), LinearRegression())
        m.fit(px[:n], py[:n])
        vl = np.mean((m.predict(xva) - yva)**2)
        print(n, deg, round(vl, 4))
ndegsklearn valpolyfit val (M1)
30112.138512.139
3084.31334.313
300110.676710.677
30084.26044.260

79. What each one costs: Milestone 3 — cross-check against sklearn

Trade off

Comparison matrix

From Milestone 3 — cross-check against sklearn: every row here is a choice with a cost. Fill the polyfit val (M1) column, then say which row you would actually pick and what you give up for it.

ndegsklearn valpolyfit val (M1)
30112.138512.139
3084.31334.313
300110.676710.677
30084.26044.260

80. The full program

Concept

import numpy as np, torch, torch.nn as nn
rng = np.random.default_rng(0)
x = rng.uniform(-3, 3, size=400).astype('float32').reshape(-1, 1)
y = (x**2 - x + rng.normal(0, 2.0, size=(400, 1))).astype('float32')

# 1. set-size curve: degree 1 (bias) vs degree 8 (variance)
xva, yva = x[300:].ravel(), y[300:].ravel()
px, py = x[:300].ravel(), y[:300].ravel()
def mse(p, xs, ys): return np.mean((np.polyval(p, xs) - ys)**2)
for n in (30, 300):
    for deg in (1, 8):
        p = np.polyfit(px[:n], py[:n], deg)
        print('n', n, 'deg', deg, 'val', round(mse(p, xva, yva), 4))

# 2. early stopping on a fresh 200-point split
rng2 = np.random.default_rng(0)
xe = rng2.uniform(-3, 3, size=200).astype('float32').reshape(-1, 1)
ye = (xe**2 - xe + rng2.normal(0, 2.0, size=(200, 1))).astype('float32')
Xtr, ytr = torch.tensor(xe[:160]), torch.tensor(ye[:160])
Xva, yv = torch.tensor(xe[160:]), torch.tensor(ye[160:])
torch.manual_seed(0)
m = nn.Sequential(nn.Linear(1, 128), nn.ReLU(), nn.Linear(128, 1))
opt = torch.optim.Adam(m.parameters(), lr=0.05); lf = nn.MSELoss()
best, bep, wait = float('inf'), 0, 0
for ep in range(1, 1001):
    m.train(); opt.zero_grad(); lf(m(Xtr), ytr).backward(); opt.step()
    m.eval()
    with torch.no_grad(): vl = lf(m(Xva), yv).item()
    if vl < best - 1e-3: best, bep, wait = vl, ep, 0
    else: wait += 1
    if wait >= 20:
        print('early stop', ep, 'best', round(best, 4), 'at', bep); break
printed linevalue
n 30 deg 1 val12.1385 (bias floor)
n 30 deg 8 val4.3133 (variance, will improve)
n 300 deg 1 val10.6767 (still bias floor)
n 300 deg 8 val4.2604 (variance cured by data)
early stopep 34, best 2.8609 at ep 14

If your degree-1 val barely moves while degree-8's collapses, and early stopping halts at epoch 34 keeping the epoch-14 weights — you can read a learning curve and prescribe the fix.

81. Fill in: value for The full program

Comparison

Comparison matrix

From The full program: refill the value column from what you know. The rest of the table is as it appeared.

printed linevalue
n 30 deg 1 val12.1385 (bias floor)
n 30 deg 8 val4.3133 (variance, will improve)
n 300 deg 1 val10.6767 (still bias floor)
n 300 deg 8 val4.2604 (variance cured by data)
early stopep 34, best 2.8609 at ep 14

82. Show it off

Concept

Slides closed, out loud: given only train MSE and val MSE, explain (1) which disease you have, (2) the one experiment you'd run to confirm it, and (3) which cure you'd reach for first.

Stretch: sweep polynomial degree 1–20 on a 30-point set and plot bias² and variance separately; then push the double-descent sweep to P = 1000 and confirm ‖w‖ keeps shrinking. Finally, say in one sentence why a huge transformer sits on the second descent.

83. Break it if you can: Show it off

Counterexample

Discussion prompt

Slides closed, out loud: given only train MSE and val MSE, explain (1) which disease you have, (2) the one experiment you'd run to confirm it, and (3) which cure you'd reach for first.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

84. Connect it up: Lesson 45: Learning Curves, Bias-Variance Diagnosis, and Double Descent

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — The decomposition · Reading learning curves · Bias vs Variance — the diagnosis · Cures for overfitting · Double descent · Your turn: diagnose and cure. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

85. What you can do now

Recap

signalthe one thing to remember
loss vs epochgap opening = overfitting; both flat & high = underfitting
loss vs ngap shrinks with data ⇒ variance; floor stays ⇒ bias
dropouttrain mode zeroes activations; eval mode is off — always eval() to score
early stoppingwatch val, reset patience on improvement, keep the best checkpoint
double descentpast P = n the min-norm fit is smooth ⇒ error falls again if n is large

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 45 (Week 16 — Learning Curves & Bias-Variance) — Barron · USAAIO Round 2 Preparation, 2026
  2. Belkin, Hsu, Ma, Mandal — Reconciling modern machine-learning practice and the classical bias–variance trade-off (double descent) — PNAS 116(32), 2019
  3. Srivastava et al. — Dropout: A Simple Way to Prevent Neural Networks from Overfitting — JMLR 15, 2014
  4. numpy.polyfit / numpy.linalg.lstsq
  5. Every learning-curve number, gap, stop-epoch, and double-descent value produced by real execution — numpy 2.2.6 + torch 2.7.1+cpu + scikit-learn 1.9, fixed seeds, verification run July 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108