USAAIO Lesson 45, from Week 16, fully worked on ONE running dataset, y = x^2 - x + noise. It derives the bias-variance decomposition of the expected loss term by term, then builds loss-against-epoch and loss-against-training-size curves from scratch and reads them zone by zone, diagnosing high bias against high variance from the gap between them. It derives and measures the targeted cures - dropout, early stopping with patience, and changes in capacity - and produces double descent with a real minimum-norm interpolator that spikes exactly at P = n and then descends again. Every snippet is self-contained and runs in a fresh interpreter, and every number in every trace table was produced by real execution with fixed seeds. The lesson runs to 63 slides.
Subject: Machine Learning · 85 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 45 · Week 16
Read the shape of a training run and know exactly what to do next. We derive the bias–variance split, build both learning curves from scratch on one dataset, diagnose the disease from the gap, apply each cure and measure it, and end at double descent — where more parameters help.
Objectives
E[(y − ŷ)²] = bias² + variance + noise term by termP = n then generalizes againWarm-up
Discussion prompt
Before we open Lesson 45: Learning Curves, Bias-Variance Diagnosis, and Double Descent: without looking back, what was the main idea of MLP Architectures — Depth, Width & Activations, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
universal approximation theorem, depth vs width expressivity, exact parameter counting for any Linear stack, the four core activation functions (sigmoid, tanh, ReLU, GELU) with computed values and gradient-flow consequences, and USAAIO architecture-design guidelines.
Section
Part 1 of 6 — where the two errors come from
Concept
Every slide fits the same synthetic dataset so the numbers compound. The true function is f(x) = x² − x, and each observed y is f(x) plus Gaussian noise of standard deviation 2.
\[ y_i = \underbrace{x_i^2 - x_i}_{f(x_i)} + \varepsilon_i, \qquad \varepsilon_i \sim \mathcal{N}(0,\, 2^2) \]
| quantity | value |
|---|---|
| x range | uniform on [−3, 3] |
| true f(x) | x² − x (a parabola) |
| noise sd | 2.0 (irreducible) |
| split | 160 train / 40 val (seed 0) |
Comparison
Comparison matrix
From One dataset for the whole lesson: refill the value column from what you know. The rest of the table is as it appeared.
| quantity | value |
|---|---|
| x range | uniform on [−3, 3] |
| true f(x) | x² − x (a parabola) |
| noise sd | 2.0 (irreducible) |
| split | 160 train / 40 val (seed 0) |
Intuition
Even a perfect model that knew f(x) = x² − x exactly would still miss each y, because each y carries its own random noise ε that nothing can predict.
So the total error a model makes splits into two very different kinds: error it could remove with a better model, and error that is baked into the data. Naming those pieces precisely is the bias–variance decomposition.
Counterexample
Discussion prompt
Even a perfect model that knew f(x) = x² − x exactly would still miss each y, because each y carries its own random noise ε that nothing can predict.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Concept
Imagine training your model on many different random training sets, each giving a slightly different fitted predictor. At a fixed test point x, three things drive the expected squared error.
f(x). A too-simple model is biased everywhere.σ² from ε; no model can touch it.Analogy
Discussion prompt
Explain Three sources of error by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Imagine training your model on many different random training sets, each giving a slightly different fitted predictor. At a fixed test point x, three things drive the expected squared error.
Estimation
Predict first
Fix a test point x. Let ŷ = ĥ(x) be the (random, training-set-dependent) prediction and y = f + ε. We want the expected squared error over both the noise and the training randomness:
Commit before you compute: what does Set up the expectation come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Insert ±E[ŷ] — add and subtract the mean prediction
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Writing (f − ŷ) as (f − E[ŷ]) + (E[ŷ] − ŷ) is the standard trick that separates a constant offset (bias) from a zero-mean wobble (variance).
Worked example
Fix a test point x. Let ŷ = ĥ(x) be the (random, training-set-dependent) prediction and y = f + ε. We want the expected squared error over both the noise and the training randomness:
\[ \mathbb{E}\big[(y - \hat y)^2\big] = \mathbb{E}\big[(f + \varepsilon - \hat y)^2\big] \]
Insert ±E[ŷ] — add and subtract the mean prediction
Why: Writing (f − ŷ) as (f − E[ŷ]) + (E[ŷ] − ŷ) is the standard trick that separates a constant offset (bias) from a zero-mean wobble (variance).
\[ = \mathbb{E}\big[(\,\underbrace{f - \mathbb{E}[\hat y]}_{\text{bias}} + \underbrace{\mathbb{E}[\hat y] - \hat y}_{\text{wobble}} + \varepsilon\,)^2\big] \]
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Insert ±E[ŷ] — add and subtract the mean prediction
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Fix a test point x. Let ŷ = ĥ(x) be the (random, training-set-dependent) prediction and y = f + ε. We want the expected squared error over both the noise and the training randomness:
Ranking
Put in order
Put the moves of Expand the square, kill the cross terms into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Every square-of-a-sum produces three squared terms and three cross terms.
Worked example
Square the three-term sum
Why: Every square-of-a-sum produces three squared terms and three cross terms. We keep the squares and show each cross term averages to zero.
\[ \mathbb{E}\big[(A + B + \varepsilon)^2\big],\quad A = f - \mathbb{E}[\hat y],\ B = \mathbb{E}[\hat y] - \hat y \]
Cross term A·B vanishes: A is a constant, E[B] = 0
Why: A = f − E[ŷ] does not depend on the training draw, and E[B] = E[E[ŷ] − ŷ] = E[ŷ] − E[ŷ] = 0. A constant times a mean-zero quantity has expectation 0.
\[ \mathbb{E}[A\,B] = A\,\mathbb{E}[B] = A\cdot 0 = 0 \]
Cross terms with ε vanish: noise is independent, E[ε] = 0
Why: ε is independent of both f (fixed) and ŷ (built from a different training set), and it is mean-zero, so E[Aε] = E[Bε] = 0.
Worked example
Only the three squares survive
Why: With every cross term gone, the expectation is the sum of the three squared pieces — each with a name.
\[ \mathbb{E}[(y-\hat y)^2] = \underbrace{(f - \mathbb{E}[\hat y])^2}_{\text{bias}^2} + \underbrace{\mathbb{E}[(\hat y - \mathbb{E}[\hat y])^2]}_{\text{variance}} + \underbrace{\sigma^2}_{\text{noise}} \]
Read the levers
Why: You can shrink bias (use a richer model) or shrink variance (use more data / regularization), but σ² = 4 here is a hard floor no method can cross.
\[ \boxed{\;\text{err} = \text{bias}^2 + \text{variance} + \sigma^2\;} \]
| term | cured by | our value |
|---|---|---|
| bias² | more model capacity | large for a line, small for a parabola |
| variance | more data, regularization | large for a wiggly high-degree fit |
| noise σ² | nothing — irreducible | 2² = 4 |
Concept
Push model complexity up and bias falls but variance rises; push it down and the reverse. Total error is their sum plus the noise floor — a U-shaped curve with a sweet spot in the middle.
Figure (svg): Three curves over increasing model complexity: bias-squared falling, variance rising, and their sum (total error) forming a U with a minimum, sitting above a flat dashed noise floor.
Explain it
Discussion prompt
Explain The trade-off in one picture to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Push model complexity up and bias falls but variance rises; push it down and the reverse. Total error is their sum plus the noise floor — a U-shaped curve with a sweet spot in the middle.
Missing information
Discussion prompt
Fit polynomials of degree 1, 2, and 15 to the running dataset and report train and val MSE. This block is fully self-contained — it rebuilds the dataset itself.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
A line cannot bend to a parabola, so both its errors stay high (bias). Degree 2 IS the true shape — lowest val. Degree 15 fits the noise (train drops, but val barely improves and it is fragile).
Worked example
Fit polynomials of degree 1, 2, and 15 to the running dataset and report train and val MSE. This block is fully self-contained — it rebuilds the dataset itself.
import numpy as np
rng = np.random.default_rng(0)
x = rng.uniform(-3, 3, size=200)
y = x**2 - x + rng.normal(0, 2.0, size=200) # true f + noise sd=2
xtr, ytr = x[:160], y[:160]
xva, yva = x[160:], y[160:]
def mse(p, xs, ys):
return np.mean((np.polyval(p, xs) - ys)**2)
for deg in [1, 2, 15]:
p = np.polyfit(xtr, ytr, deg)
print(deg, round(mse(p, xtr, ytr), 4), round(mse(p, xva, yva), 4))deg 1 is biased; deg 2 matches the truth; deg 15 adds variance
Why: A line cannot bend to a parabola, so both its errors stay high (bias). Degree 2 IS the true shape — lowest val. Degree 15 fits the noise (train drops, but val barely improves and it is fragile).
| degree | train MSE | val MSE | reading |
|---|---|---|---|
| 1 | 12.7146 | 8.1469 | high bias — too stiff |
| 2 | 4.4588 | 3.0343 | sweet spot — true shape |
| 15 | 4.2987 | 2.7401 | wiggly; low bias, watch variance |
Trade off
Comparison matrix
From See both terms in code — one dataset, three degrees: every row here is a choice with a cost. Fill the train MSE column, then say which row you would actually pick and what you give up for it.
| degree | train MSE | val MSE | reading |
|---|---|---|---|
| 1 | 12.7146 | 8.1469 | high bias — too stiff |
| 2 | 4.4588 | 3.0343 | sweet spot — true shape |
| 15 | 4.2987 | 2.7401 | wiggly; low bias, watch variance |
Section
Part 2 of 6 — two axes, two questions
Concept
A learning curve plots error against one of two axes, and each answers a different question about your run.
| x-axis | question it answers | when to use |
|---|---|---|
| training epochs | Is the run converging? When does overfitting start? | monitor one training run |
| training-set size n | Is this bias or variance? Will more data help? | decide whether to collect data |
Both plot train and val error together. The gap between them is the signal — never read one curve in isolation.
Concept
In a healthy run both curves fall, then val flattens or turns up while train keeps dropping. That turn is the onset of memorization.
| zone | train loss | val loss | diagnosis |
|---|---|---|---|
| early | falling fast | falling fast | learning real signal |
| middle | still falling | flattens | generalization saturated |
| late | keeps falling | flat or rising | memorizing noise |
Early stopping lives in the middle zone: halt at the val minimum, before the late zone eats your generalization.
Pattern
Predict first
The table runs: 1 | 24.0913 | 14.9322 | early · 20 | 11.2770 | 6.4902 | early · 50 | 5.9076 | 3.3293 | early/mid · 100 | 4.5213 | 2.8274 | val minimum · 200 | 4.4184 | 2.8624 | flat
In Build the loss-vs-epoch curve, given the rows so far: what is the next one — the row where epoch is 400?
Correct: 400 | 4.3722 | 2.8931 | flat — stop earlier
| epoch | train MSE | val MSE | zone |
|---|---|---|---|
| 1 | 24.0913 | 14.9322 | early |
| 20 | 11.2770 | 6.4902 | early |
| 50 | 5.9076 | 3.3293 | early/mid |
| 100 | 4.5213 | 2.8274 | val minimum |
| 200 | 4.4184 | 2.8624 | flat |
| 400 | 4.3722 | 2.8931 | flat — stop earlier |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. By epoch 100 val has flattened at ~2.83 (below train because the val set here is easier); training longer to epoch 400 does NOT lower val — the middle/late boundary is around epoch 100.
Worked example
Train a small Tanh MLP with Adam on the running dataset and log train and val MSE at chosen epochs. Self-contained; seeds fixed so your numbers match these.
import numpy as np, torch, torch.nn as nn
rng = np.random.default_rng(0)
x = rng.uniform(-3, 3, size=200).astype('float32').reshape(-1, 1)
y = (x**2 - x + rng.normal(0, 2.0, size=(200, 1))).astype('float32')
Xtr, ytr = torch.tensor(x[:160]), torch.tensor(y[:160])
Xva, yva = torch.tensor(x[160:]), torch.tensor(y[160:])
torch.manual_seed(0)
model = nn.Sequential(nn.Linear(1, 32), nn.Tanh(), nn.Linear(32, 1))
opt = torch.optim.Adam(model.parameters(), lr=0.02)
lossf = nn.MSELoss()
for ep in range(1, 401):
model.train(); opt.zero_grad()
tl = lossf(model(Xtr), ytr); tl.backward(); opt.step()
if ep in (1, 20, 50, 100, 200, 400):
model.eval()
with torch.no_grad():
vl = lossf(model(Xva), yva).item()
print(ep, round(tl.item(), 4), round(vl, 4))Both losses fall fast, then flatten near the σ² floor
Why: By epoch 100 val has flattened at ~2.83 (below train because the val set here is easier); training longer to epoch 400 does NOT lower val — the middle/late boundary is around epoch 100.
| epoch | train MSE | val MSE | zone |
|---|---|---|---|
| 1 | 24.0913 | 14.9322 | early |
| 20 | 11.2770 | 6.4902 | early |
| 50 | 5.9076 | 3.3293 | early/mid |
| 100 | 4.5213 | 2.8274 | val minimum |
| 200 | 4.4184 | 2.8624 | flat |
| 400 | 4.3722 | 2.8931 | flat — stop earlier |
Intuition
Forget the absolute numbers for a second. Ask only two things: is the floor high? and is there a gap between train and val? Those two yes/no answers name the disease.
High floor + no gap = the model can't fit the signal (bias). Low train + big gap = the model fits the training noise but not the world (variance). That is the entire diagnostic.
Section
Part 3 of 6 — read it off the gap
Concept
A high-bias model is too simple. Train and val error are both high and roughly equal — no gap, but a wrong floor.
\[ \varepsilon_{\text{train}} \approx \varepsilon_{\text{val}} \gg \sigma^2 \]
On the set-size curve, both curves plateau high and refuse to fall as n grows. More data cannot help — the model lacks the capacity to represent f.
Concept
A high-variance model memorizes the training set. Train error is low, val error is far higher. The gap is the signature.
\[ \varepsilon_{\text{train}} \ll \varepsilon_{\text{val}} \]
On the set-size curve, as n grows the val error falls toward the train error and the gap closes. More data is the direct cure for pure variance.
Estimation
Predict first
Fit a degree-1 model (stiff) and a degree-8 model (flexible) on growing training sets n ∈ {10, 30, 100, 300}, scored on a fixed 100-point val set. Self-contained.
Commit before you compute: what does Build the set-size curve — degree 1 vs degree 8 come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: deg 1 val stays ~12→10.7 (bias floor); deg 8 val collapses 276→4.26 (variance cured by data)
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The degree-1 val barely moves as n grows 30× — a bias floor.
Worked example
Fit a degree-1 model (stiff) and a degree-8 model (flexible) on growing training sets n ∈ {10, 30, 100, 300}, scored on a fixed 100-point val set. Self-contained.
import numpy as np
rng = np.random.default_rng(0)
x = rng.uniform(-3, 3, size=400)
y = x**2 - x + rng.normal(0, 2.0, size=400)
xva, yva = x[300:], y[300:] # fixed 100-point val set
pool_x, pool_y = x[:300], y[:300]
def mse(p, xs, ys):
return np.mean((np.polyval(p, xs) - ys)**2)
for n in [10, 30, 100, 300]:
for deg in [1, 8]:
p = np.polyfit(pool_x[:n], pool_y[:n], deg)
tr = mse(p, pool_x[:n], pool_y[:n])
vl = mse(p, xva, yva)
print(n, deg, round(tr, 4), round(vl, 4))deg 1 val stays ~12→10.7 (bias floor); deg 8 val collapses 276→4.26 (variance cured by data)
Why: The degree-1 val barely moves as n grows 30× — a bias floor. The degree-8 val plummets from 276.8 (wild overfit at n=10) to 4.26 as data fills in the wiggles.
| n | deg-1 train | deg-1 val | deg-8 train | deg-8 val |
|---|---|---|---|---|
| 10 | 11.6410 | 12.6484 | 0.0306 | 276.7645 |
| 30 | 14.9880 | 12.1385 | 3.1825 | 4.3133 |
| 100 | 13.9613 | 10.8800 | 4.4177 | 4.3799 |
| 300 | 12.0813 | 10.6767 | 3.8873 | 4.2604 |
Pattern
Step through it
Step through Build the set-size curve — degree 1 vs degree 8 one row at a time. What is driving the change, and what would the row after the last one be?
Intuition
Look down the two val columns. Degree-1's val is a nearly flat line near 11 — pour in 30× the data and nothing improves. That flatness is bias.
Degree-8's val falls off a cliff from 276 to 4.26. The collapse toward the train curve as n grows is variance being cured by data. Two identical experiments, opposite prescriptions.
Ranking
Put in order
These are the steps of The diagnose-then-cure recipe, scrambled. Put them back in order before the next slide shows you.
val − train and the floor (level both curves sit at)weight_decay, or early stoppingWhy: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
val − train and the floor (level both curves sit at)weight_decay, or early stoppingEdge cases
Discussion prompt
The diagnose-then-cure recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
val − train and the floor (level both curves sit at)weight_decay, or early stoppingAnomaly
Predict first
A student writes this, and it looks reasonable:
Train MSE is high (~12) and val MSE is equally high (~12). The model is doing badly, so collect 10× more data to fix it.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Both errors are high AND equal — there is no gap.
Diagnose with the gap FIRST. No gap at a high floor means bias, and bias is a capacity problem.
Why: Both errors are high AND equal — there is no gap. The measured degree-1 val goes 12.65 → 12.14 → 10.88 → 10.68 as n runs 10 → 300: a flat bias floor. Thirty times the data bought almost nothing. Data-collection effort wasted.
Trap
Train MSE is high (~12) and val MSE is equally high (~12). The model is doing badly, so collect 10× more data to fix it.
Add data to a high-bias model
Why: Both errors are high AND equal — there is no gap. The measured degree-1 val goes 12.65 → 12.14 → 10.88 → 10.68 as n runs 10 → 300: a flat bias floor. Thirty times the data bought almost nothing. Data-collection effort wasted.
Diagnose with the gap FIRST. No gap at a high floor means bias, and bias is a capacity problem.
Gap rule: big gap → data/regularization; no gap + high floor → add capacity
Why: The degree-8 val fell 276 → 4.26 with more data (that's variance, curable by n). The degree-1 floor did not. Only richer models or better features cure bias. Match the cure to the disease.
Elimination
Eliminate the wrong options
A model shows train MSE = 11.6 and val MSE = 12.6 at n = 10. You grow the training set to n = 300 and both settle near 11–12. What is the primary problem?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Train and val are close (gap ≈ 1) so there is no overfitting, yet 30× more data leaves the error near 11–12 — the floor will not fall. That flat, high floor with no gap is the textbook high-bias signature (this is exactly the measured degree-1 set-size curve). Cure: more capacity or better features.
Check
Use the gap-and-floor rule. Don't just look at whether the numbers are 'big'.
Check your understanding
A model shows train MSE = 11.6 and val MSE = 12.6 at n = 10. You grow the training set to n = 300 and both settle near 11–12. What is the primary problem?
Answer: A
Why: Train and val are close (gap ≈ 1) so there is no overfitting, yet 30× more data leaves the error near 11–12 — the floor will not fall. That flat, high floor with no gap is the textbook high-bias signature (this is exactly the measured degree-1 set-size curve). Cure: more capacity or better features.
Section
Part 4 of 6 — dropout & early stopping
Concept
weight_decay)Pick by what is cheap: if data is free, collect it. If the model is already large and data is scarce, dropout + early stopping are the fastest levers — and we measure both next.
Concept
During training, each activation is zeroed with probability p on every forward pass and the survivors are scaled by 1/(1−p). During eval, dropout is off — the scaling already kept the expected magnitude right.
\[ \tilde h_i = \frac{m_i}{1-p}\, h_i, \qquad m_i \sim \text{Bernoulli}(1-p) \]
No neuron can rely on any single partner, so the network learns redundant, distributed features. It is noise injection that lowers the effective capacity at train time.
Fill the middle
Fill in the blanks
From Dropout closes the gap — measured — one line has had its right-hand side removed. Put it back.
import numpy as np, torch, torch.nn as nn
rng = np.random.default_rng(0)
x = rng.uniform(-3, 3, size=40).astype('float32').reshape(-1, 1)
y = (x2 - x + rng.normal(0, 2.0, size=(40, 1))).astype('float32')
xv = rng.uniform(-3, 3, size=200).astype('float32').reshape(-1, 1)
yv = (xv2 - xv + rng.normal(0, 2.0, size=(200, 1))).astype('float32')
Xt, yt, Xvt, yvt = map(torch.tensor, (x, y, xv, yv))
lossf = nn.MSELoss()
def make(drop):
L = [nn.Linear(1, 128), nn.ReLU()]
if drop: L.append(nn.Dropout(drop))
L += [nn.Linear(128, 128), nn.ReLU()]
if drop: L.append(nn.Dropout(drop))
L.append(nn.Linear(128, 1))
return nn.Sequential(*L)
for p in (0.0, 0.5):
torch.manual_seed(0)
m = make(p); opt = torch.optim.Adam(m.parameters(), lr=0.01)
for _ in range(800):
m.train(); opt.zero_grad(); lossf(m(Xt), yt).backward(); opt.step()
m.eval()
with torch.no_grad():
tr = lossf(m(Xt), yt).item(); vl = lossf(m(Xvt), yvt).item()
print(p, round(tr, 4), round(vl, 4), round(vl - tr, 4))
Why: lossf is what everything below it consumes, so the wrong expression here fails later and somewhere else. Without dropout the net drives train MSE to 2.45 but val sits at 7.35 — a 4.9 gap (memorization).
Worked example
Train an overparameterized 128→128→1 ReLU net on just 40 points, once without dropout and once with Dropout(0.5). Same seed, same everything else. Self-contained.
import numpy as np, torch, torch.nn as nn
rng = np.random.default_rng(0)
x = rng.uniform(-3, 3, size=40).astype('float32').reshape(-1, 1)
y = (x**2 - x + rng.normal(0, 2.0, size=(40, 1))).astype('float32')
xv = rng.uniform(-3, 3, size=200).astype('float32').reshape(-1, 1)
yv = (xv**2 - xv + rng.normal(0, 2.0, size=(200, 1))).astype('float32')
Xt, yt, Xvt, yvt = map(torch.tensor, (x, y, xv, yv))
lossf = nn.MSELoss()
def make(drop):
L = [nn.Linear(1, 128), nn.ReLU()]
if drop: L.append(nn.Dropout(drop))
L += [nn.Linear(128, 128), nn.ReLU()]
if drop: L.append(nn.Dropout(drop))
L.append(nn.Linear(128, 1))
return nn.Sequential(*L)
for p in (0.0, 0.5):
torch.manual_seed(0)
m = make(p); opt = torch.optim.Adam(m.parameters(), lr=0.01)
for _ in range(800):
m.train(); opt.zero_grad(); lossf(m(Xt), yt).backward(); opt.step()
m.eval()
with torch.no_grad():
tr = lossf(m(Xt), yt).item(); vl = lossf(m(Xvt), yvt).item()
print(p, round(tr, 4), round(vl, 4), round(vl - tr, 4))No dropout: gap = 4.90. Dropout 0.5: gap = 0.36
Why: Without dropout the net drives train MSE to 2.45 but val sits at 7.35 — a 4.9 gap (memorization). Dropout raises train to 4.83 yet drops val to 5.19: it trades a little train fit for a 13× smaller gap.
| dropout p | train MSE | val MSE | gap (val − train) |
|---|---|---|---|
| 0.0 (none) | 2.4489 | 7.3467 | 4.8978 |
| 0.5 | 4.8250 | 5.1870 | 0.3621 |
Anomaly
Predict first
A student writes this, and it looks reasonable:
Dropout is part of the model, so evaluate the same way you trained — just call model(Xval) right after the training loop, model still in train mode.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: In train mode dropout keeps randomly zeroing activations, so every val score is noisy and randomly too high — and it changes run to run.
Switch to eval mode and disable grad before validating; switch back to train before the next step.
Why: In train mode dropout keeps randomly zeroing activations, so every val score is noisy and randomly too high — and it changes run to run. You are measuring a crippled network, not the one you will deploy.
Trap
Dropout is part of the model, so evaluate the same way you trained — just call model(Xval) right after the training loop, model still in train mode.
Score val with the module in train() mode
Why: In train mode dropout keeps randomly zeroing activations, so every val score is noisy and randomly too high — and it changes run to run. You are measuring a crippled network, not the one you will deploy.
Switch to eval mode and disable grad before validating; switch back to train before the next step.
model.eval(); with torch.no_grad(): score val
Why: eval() turns dropout OFF and uses the full, scaled network — the deterministic predictor you actually ship. torch.no_grad() just saves memory. This is why the measured val above is stable and reproducible.
Break the constraint
Discussion prompt
The rule this trap just fixed:
Switch to eval mode and disable grad before validating; switch back to train before the next step.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
In train mode dropout keeps randomly zeroing activations, so every val score is noisy and randomly too high — and it changes run to run. You are measuring a crippled network, not the one you will deploy.
Concept
Watch val loss. Remember the best value seen and its epoch. Each epoch that fails to beat it (by a small threshold) increments a wait counter; reset wait to 0 whenever you improve. Stop once wait reaches patience.
Patience gives the run room to escape a temporary plateau — it does not quit on the first bad epoch, only after patience consecutive non-improvements.
Missing information
Discussion prompt
Use a deliberately high Adam lr = 0.05 so val bottoms out early then drifts up. patience = 20, improvement threshold 1e-3. Self-contained.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
After epoch 14 no epoch beats 2.8609 by 1e-3, so wait climbs 1..20 while val drifts up. At epoch 34 wait hits patience and training halts — you keep the epoch-14 checkpoint, not the worse epoch-34 weights.
Worked example
Use a deliberately high Adam lr = 0.05 so val bottoms out early then drifts up. patience = 20, improvement threshold 1e-3. Self-contained.
import numpy as np, torch, torch.nn as nn
rng = np.random.default_rng(0)
x = rng.uniform(-3, 3, size=200).astype('float32').reshape(-1, 1)
y = (x**2 - x + rng.normal(0, 2.0, size=(200, 1))).astype('float32')
Xtr, ytr = torch.tensor(x[:160]), torch.tensor(y[:160])
Xva, yva = torch.tensor(x[160:]), torch.tensor(y[160:])
torch.manual_seed(0)
model = nn.Sequential(nn.Linear(1, 128), nn.ReLU(), nn.Linear(128, 1))
opt = torch.optim.Adam(model.parameters(), lr=0.05) # high on purpose
lossf = nn.MSELoss()
best_val, best_ep, wait, patience = float('inf'), 0, 0, 20
for ep in range(1, 1001):
model.train(); opt.zero_grad()
lossf(model(Xtr), ytr).backward(); opt.step()
model.eval()
with torch.no_grad():
vl = lossf(model(Xva), yva).item()
if vl < best_val - 1e-3:
best_val, best_ep, wait = vl, ep, 0
else:
wait += 1
if wait >= patience:
print('STOP', ep, 'best', round(best_val, 4), 'at', best_ep)
breakVal bottoms at 2.8609 (epoch 14); 20 quiet epochs later, STOP at 34
Why: After epoch 14 no epoch beats 2.8609 by 1e-3, so wait climbs 1..20 while val drifts up. At epoch 34 wait hits patience and training halts — you keep the epoch-14 checkpoint, not the worse epoch-34 weights.
| epoch | val loss | improved? | wait |
|---|---|---|---|
| 1 | 4.6386 | yes (best) | 0 |
| 14 | 2.8609 | yes (best) | 0 |
| 15 | 3.0364 | no | 1 |
| 24 | ≈ higher | no | 10 |
| 34 | ≈ higher | no — STOP | 20 |
Comparison
Comparison matrix
From Early stopping — measured on a run that overshoots: refill the wait column from what you know. The rest of the table is as it appeared.
| epoch | val loss | improved? | wait |
|---|---|---|---|
| 1 | 4.6386 | yes (best) | 0 |
| 14 | 2.8609 | yes (best) | 0 |
| 15 | 3.0364 | no | 1 |
| 24 | ≈ higher | no | 10 |
| 34 | ≈ higher | no — STOP | 20 |
Anomaly
Predict first
A student writes this, and it looks reasonable:
Track train_loss and stop when it stops decreasing — it's the loss you're optimizing, so it's the natural thing to watch.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: With enough capacity train loss keeps sliding down toward 0 and never meaningfully plateaus until the model has fully memorized.
Track val_loss. Stop when the held-out signal stops improving.
Why: With enough capacity train loss keeps sliding down toward 0 and never meaningfully plateaus until the model has fully memorized. You either stop far too late (already overfit) or crank the threshold so tight you stop at epoch 2.
Trap
Track train_loss and stop when it stops decreasing — it's the loss you're optimizing, so it's the natural thing to watch.
Set patience on the train-loss plateau
Why: With enough capacity train loss keeps sliding down toward 0 and never meaningfully plateaus until the model has fully memorized. You either stop far too late (already overfit) or crank the threshold so tight you stop at epoch 2.
Track val_loss. Stop when the held-out signal stops improving.
Track best val_loss; reset patience on improvement; stop when wait ≥ patience
Why: Val loss is the only quantity that knows about the held-out distribution, so its turn-up is the real onset of memorization (here epoch 14). Train loss has no memory of unseen data.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
f(x) = x² − x exactly would still miss each y, because each y carries its own random noise ε that nothing can predict.; A learning curve plots error against one of two axes, and each answers a different question about your run.; In a healthy run both curves fall, then val flattens or turns up while train keeps dropping. That turn is the onset of memorization.~12) and val MSE is equally high (~12). The model is doing badly, so collect 10× more data to fix it.; Dropout is part of the model, so evaluate the same way you trained — just call model(Xval) right after the training loop, model still in train mode.Prediction
Predict first
patience = 3, and the per-epoch val losses are [2.1, 1.8, 1.9, 1.7, 1.8, 1.9, 2.0]. At which epoch does early stopping trigger?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Epoch 7
Why: Trace: ep1 2.1 best (wait 0), ep2 1.8 best (0), ep3 1.9 no (1), ep4 1.7 best (0), ep5 1.8 no (1), ep6 1.9 no (2), ep7 2.0 no (3 ≥ patience → STOP). It fires at epoch 7.
Check
Walk the counter epoch by epoch. A new best resets wait to 0.
Check your understanding
patience = 3, and the per-epoch val losses are [2.1, 1.8, 1.9, 1.7, 1.8, 1.9, 2.0]. At which epoch does early stopping trigger?
Answer: A
Why: Trace: ep1 2.1 best (wait 0), ep2 1.8 best (0), ep3 1.9 no (1), ep4 1.7 best (0), ep5 1.8 no (1), ep6 1.9 no (2), ep7 2.0 no (3 ≥ patience → STOP). It fires at epoch 7.
Section
Part 5 of 6 — past the interpolation threshold
Concept
Classical theory says test error is U-shaped in complexity: underfit → sweet spot → overfit, blowing up as you approach enough parameters to fit the data exactly.
That blow-up point — where the number of parameters P equals the number of training points n — is the interpolation threshold. Classical theory stops there. Modern overparameterized models keep going past it, and something surprising happens.
Picture it
Figure (svg): Test error curve versus number of parameters P: a first descent, a sharp spike at P equals n (the interpolation threshold, marked with a dashed vertical line), then a second descent to a low plateau.
Discussion prompt
Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.
Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.
Answer:
Push P well beyond n and test error, after spiking at P = n, descends a second time. When infinitely many zero-training-error fits exist, least squares picks the minimum-norm one — the smoothest interpolator, which generalizes.
Concept
Push P well beyond n and test error, after spiking at P = n, descends a second time. When infinitely many zero-training-error fits exist, least squares picks the minimum-norm one — the smoothest interpolator, which generalizes.
Figure (svg): Test error curve versus number of parameters P: a first descent, a sharp spike at P equals n (the interpolation threshold, marked with a dashed vertical line), then a second descent to a low plateau.
Estimation
Predict first
Fit a step target with min-norm least squares over P Chebyshev features on n = 20 points. Sweep P through the threshold. lstsq returns the minimum-norm solution when the system is under-determined. Self-contained.
Commit before you compute: what does Reproduce double descent — a real interpolator come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Test peaks at P = n = 20 (te ≈ 101, ‖w‖ ≈ 34) then descends to ≈ 0.31 at P = 25
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Below P=20 train error is nonzero (can't interpolate).
Worked example
Fit a step target with min-norm least squares over P Chebyshev features on n = 20 points. Sweep P through the threshold. lstsq returns the minimum-norm solution when the system is under-determined. Self-contained.
import numpy as np
rng = np.random.default_rng(0)
n = 20
xtr = np.sort(rng.uniform(-1, 1, size=n))
ytr = np.sign(xtr) # a step: hard for low-degree polys
xte = np.linspace(-1, 1, 1000)
yte = np.sign(xte)
def design(x, P): # P Chebyshev features T_0..T_{P-1}
return np.polynomial.chebyshev.chebvander(x, P - 1)
for P in [3, 8, 15, 19, 20, 21, 25, 40, 100, 300]:
Ftr, Fte = design(xtr, P), design(xte, P)
w, *_ = np.linalg.lstsq(Ftr, ytr, rcond=None) # min-norm solution
tr = np.mean((Ftr @ w - ytr)**2)
te = np.mean((Fte @ w - yte)**2)
print(P, round(tr, 5), round(te, 3), round(np.linalg.norm(w), 2))Test peaks at P = n = 20 (te ≈ 101, ‖w‖ ≈ 34) then descends to ≈ 0.31 at P = 25
Why: Below P=20 train error is nonzero (can't interpolate). At P≈20 it interpolates but ‖w‖ explodes and test blows up. Past P=20, extra features let lstsq pick a SMALLER-norm interpolator — ‖w‖ falls 34 → 0.34 and test drops to ~0.3–0.9.
| P (features) | train MSE | test MSE | ‖w‖ | regime |
|---|---|---|---|---|
| 8 | 0.02697 | 0.184 | 1.33 | underparam (1st descent) |
| 15 | 0.00030 | 5.395 | 7.45 | approaching threshold |
| 20 | 0.00000 | 101.062 | 34.31 | P = n → SPIKE |
| 25 | 0.00000 | 0.312 | 1.21 | overparam (2nd descent) |
| 300 | 0.00000 | 0.936 | 0.34 | deep overparam, min-norm |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Test peaks at P = n = 20 (te ≈ 101, ‖w‖ ≈ 34) then descends to ≈ 0.31 at P = 25
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Fit a step target with min-norm least squares over P Chebyshev features on n = 20 points. Sweep P through the threshold. lstsq returns the minimum-norm solution when the system is under-determined. Self-contained.
Intuition
Watch the ‖w‖ column, not just the error. It climbs to 34 right at P = n — the lone interpolator there is forced to use enormous, oscillating coefficients to thread all 20 points exactly.
Add more features and there are now many exact fits; lstsq takes the smallest-norm one, so ‖w‖ collapses to 0.34. Small norm means a smooth function, and smooth functions generalize. That is the whole mechanism behind the second descent.
Explain it
Discussion prompt
Explain Why the norm tells the story to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Watch the ‖w‖ column, not just the error. It climbs to 34 right at P = n — the lone interpolator there is forced to use enormous, oscillating coefficients to thread all 20 points exactly.
Concept
A transformer with billions of parameters trained on billions of tokens sits deep in the second descent (P ≫ n, with n huge). Gradient descent has an implicit min-norm bias, so it lands on a smooth interpolator that generalizes.
The catch is the spike: you must have n large enough to be past the threshold. Overparameterizing a model on a tiny dataset lands you on the peak, not the plateau — exactly the P = n = 20 disaster above.
Elimination
Eliminate the wrong options
A 200M-parameter transformer trained on 10B tokens reaches low test loss. Classical bias–variance theory predicts severe overfitting. Why doesn't it overfit?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: With P ≫ n and n very large (10B tokens), the model is far past the interpolation threshold on the second descent. Among the many zero-loss fits, gradient descent's implicit bias selects a minimum-norm (smooth) solution — exactly the ‖w‖ 34 → 0.34 collapse measured above. The large n is what keeps it off the spike.
Check
Reason from the mechanism you just measured, not from a slogan.
Check your understanding
A 200M-parameter transformer trained on 10B tokens reaches low test loss. Classical bias–variance theory predicts severe overfitting. Why doesn't it overfit?
Answer: A
Why: With P ≫ n and n very large (10B tokens), the model is far past the interpolation threshold on the second descent. Among the many zero-loss fits, gradient descent's implicit bias selects a minimum-norm (smooth) solution — exactly the ‖w‖ 34 → 0.34 collapse measured above. The large n is what keeps it off the spike.
Section
Part 6 of 6 — the project
Concept
On the running dataset, build the set-size curve for two model complexities, then implement early stopping on an MLP — reproducing the exact numbers from Parts 3 and 4 yourself.
| # | milestone | tool |
|---|---|---|
| 1 | Set-size curve: degree-1 vs degree-8 | numpy.polyfit, vary n |
| 2 | Early stopping with patience = 20 | torch, best-val counter |
| 3 | Cross-check the curve against sklearn | PolynomialFeatures + LinearRegression |
Build rules: type every line, run after each milestone, and always use model.train() for the step and model.eval() + torch.no_grad() for validation — never mix them.
Analogy
Discussion prompt
Explain Project: full learning-curve analysis by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
On the running dataset, build the set-size curve for two model complexities, then implement early stopping on an MLP — reproducing the exact numbers from Parts 3 and 4 yourself.
Worked example
Your turn: fit degree-1 and degree-8 polynomials for n ∈ {10, 30, 100, 300} on a fixed 100-point val set. Predict aloud: which degree's val error collapses as n grows?
Hint: rebuild the dataset with default_rng(0), hold out x[300:] as val, and loop n over the pool. Use np.polyfit and np.polyval.
import numpy as np
rng = np.random.default_rng(0)
x = rng.uniform(-3, 3, size=400)
y = x**2 - x + rng.normal(0, 2.0, size=400)
xva, yva = x[300:], y[300:]
pool_x, pool_y = x[:300], y[:300]
def mse(p, xs, ys):
return np.mean((np.polyval(p, xs) - ys)**2)
for n in [10, 30, 100, 300]:
p1 = np.polyfit(pool_x[:n], pool_y[:n], 1)
p8 = np.polyfit(pool_x[:n], pool_y[:n], 8)
print(n, round(mse(p1, xva, yva), 3), round(mse(p8, xva, yva), 3))| n | deg-1 val | deg-8 val |
|---|---|---|
| 10 | 12.648 | 276.765 |
| 30 | 12.139 | 4.313 |
| 100 | 10.880 | 4.380 |
| 300 | 10.677 | 4.260 |
Pattern
Step through it
Step through Milestone 1 — the set-size curve one row at a time. What is driving the change, and what would the row after the last one be?
Worked example
Your turn: implement the patience loop on the 200-point MLP with lr = 0.05, patience = 20. Predict: which epoch holds the best val, and roughly when does it stop?
Hint: best_val = float('inf'), wait = 0; reset wait on improvement (threshold 1e-3); break when wait >= patience. Remember eval() + no_grad() for the val score.
import numpy as np, torch, torch.nn as nn
rng = np.random.default_rng(0)
x = rng.uniform(-3, 3, size=200).astype('float32').reshape(-1, 1)
y = (x**2 - x + rng.normal(0, 2.0, size=(200, 1))).astype('float32')
Xtr, ytr = torch.tensor(x[:160]), torch.tensor(y[:160])
Xva, yva = torch.tensor(x[160:]), torch.tensor(y[160:])
torch.manual_seed(0)
m = nn.Sequential(nn.Linear(1, 128), nn.ReLU(), nn.Linear(128, 1))
opt = torch.optim.Adam(m.parameters(), lr=0.05); lf = nn.MSELoss()
best, bep, wait = float('inf'), 0, 0
for ep in range(1, 1001):
m.train(); opt.zero_grad(); lf(m(Xtr), ytr).backward(); opt.step()
m.eval()
with torch.no_grad(): vl = lf(m(Xva), yva).item()
if vl < best - 1e-3: best, bep, wait = vl, ep, 0
else: wait += 1
if wait >= 20:
print('STOP', ep, 'best', round(best, 4), 'at', bep); break| output | value |
|---|---|
| stop epoch | 34 |
| best val | 2.8609 |
| best epoch | 14 |
Pattern
Step through it
Step through Milestone 2 — early stopping with patience one row at a time. What is driving the change, and what would the row after the last one be?
Worked example
Your turn: rebuild the set-size fits with an sklearn Pipeline and confirm the val numbers match your polyfit results to the decimal. Predict: exact match or just close?
Hint: make_pipeline(PolynomialFeatures(deg), LinearRegression()); reshape x to (-1, 1); score with mean squared residual on the same val set.
import numpy as np
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import PolynomialFeatures
from sklearn.linear_model import LinearRegression
rng = np.random.default_rng(0)
x = rng.uniform(-3, 3, size=400).reshape(-1, 1)
y = (x**2 - x).ravel() + rng.normal(0, 2.0, size=400)
xva, yva = x[300:], y[300:]
px, py = x[:300], y[:300]
for n in [30, 300]:
for deg in [1, 8]:
m = make_pipeline(PolynomialFeatures(deg), LinearRegression())
m.fit(px[:n], py[:n])
vl = np.mean((m.predict(xva) - yva)**2)
print(n, deg, round(vl, 4))| n | deg | sklearn val | polyfit val (M1) |
|---|---|---|---|
| 30 | 1 | 12.1385 | 12.139 |
| 30 | 8 | 4.3133 | 4.313 |
| 300 | 1 | 10.6767 | 10.677 |
| 300 | 8 | 4.2604 | 4.260 |
Trade off
Comparison matrix
From Milestone 3 — cross-check against sklearn: every row here is a choice with a cost. Fill the polyfit val (M1) column, then say which row you would actually pick and what you give up for it.
| n | deg | sklearn val | polyfit val (M1) |
|---|---|---|---|
| 30 | 1 | 12.1385 | 12.139 |
| 30 | 8 | 4.3133 | 4.313 |
| 300 | 1 | 10.6767 | 10.677 |
| 300 | 8 | 4.2604 | 4.260 |
Concept
import numpy as np, torch, torch.nn as nn
rng = np.random.default_rng(0)
x = rng.uniform(-3, 3, size=400).astype('float32').reshape(-1, 1)
y = (x**2 - x + rng.normal(0, 2.0, size=(400, 1))).astype('float32')
# 1. set-size curve: degree 1 (bias) vs degree 8 (variance)
xva, yva = x[300:].ravel(), y[300:].ravel()
px, py = x[:300].ravel(), y[:300].ravel()
def mse(p, xs, ys): return np.mean((np.polyval(p, xs) - ys)**2)
for n in (30, 300):
for deg in (1, 8):
p = np.polyfit(px[:n], py[:n], deg)
print('n', n, 'deg', deg, 'val', round(mse(p, xva, yva), 4))
# 2. early stopping on a fresh 200-point split
rng2 = np.random.default_rng(0)
xe = rng2.uniform(-3, 3, size=200).astype('float32').reshape(-1, 1)
ye = (xe**2 - xe + rng2.normal(0, 2.0, size=(200, 1))).astype('float32')
Xtr, ytr = torch.tensor(xe[:160]), torch.tensor(ye[:160])
Xva, yv = torch.tensor(xe[160:]), torch.tensor(ye[160:])
torch.manual_seed(0)
m = nn.Sequential(nn.Linear(1, 128), nn.ReLU(), nn.Linear(128, 1))
opt = torch.optim.Adam(m.parameters(), lr=0.05); lf = nn.MSELoss()
best, bep, wait = float('inf'), 0, 0
for ep in range(1, 1001):
m.train(); opt.zero_grad(); lf(m(Xtr), ytr).backward(); opt.step()
m.eval()
with torch.no_grad(): vl = lf(m(Xva), yv).item()
if vl < best - 1e-3: best, bep, wait = vl, ep, 0
else: wait += 1
if wait >= 20:
print('early stop', ep, 'best', round(best, 4), 'at', bep); break| printed line | value |
|---|---|
| n 30 deg 1 val | 12.1385 (bias floor) |
| n 30 deg 8 val | 4.3133 (variance, will improve) |
| n 300 deg 1 val | 10.6767 (still bias floor) |
| n 300 deg 8 val | 4.2604 (variance cured by data) |
| early stop | ep 34, best 2.8609 at ep 14 |
If your degree-1 val barely moves while degree-8's collapses, and early stopping halts at epoch 34 keeping the epoch-14 weights — you can read a learning curve and prescribe the fix.
Comparison
Comparison matrix
From The full program: refill the value column from what you know. The rest of the table is as it appeared.
| printed line | value |
|---|---|
| n 30 deg 1 val | 12.1385 (bias floor) |
| n 30 deg 8 val | 4.3133 (variance, will improve) |
| n 300 deg 1 val | 10.6767 (still bias floor) |
| n 300 deg 8 val | 4.2604 (variance cured by data) |
| early stop | ep 34, best 2.8609 at ep 14 |
Concept
Slides closed, out loud: given only train MSE and val MSE, explain (1) which disease you have, (2) the one experiment you'd run to confirm it, and (3) which cure you'd reach for first.
Stretch: sweep polynomial degree 1–20 on a 30-point set and plot bias² and variance separately; then push the double-descent sweep to P = 1000 and confirm ‖w‖ keeps shrinking. Finally, say in one sentence why a huge transformer sits on the second descent.
Counterexample
Discussion prompt
Slides closed, out loud: given only train MSE and val MSE, explain (1) which disease you have, (2) the one experiment you'd run to confirm it, and (3) which cure you'd reach for first.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — The decomposition · Reading learning curves · Bias vs Variance — the diagnosis · Cures for overfitting · Double descent · Your turn: diagnose and cure. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
error = bias² + variance + σ² and name which lever moves each termP = n = 20, then ‖w‖ fell 34 → 0.34 on the second descent| signal | the one thing to remember |
|---|---|
| loss vs epoch | gap opening = overfitting; both flat & high = underfitting |
| loss vs n | gap shrinks with data ⇒ variance; floor stays ⇒ bias |
| dropout | train mode zeroes activations; eval mode is off — always eval() to score |
| early stopping | watch val, reset patience on improvement, keep the best checkpoint |
| double descent | past P = n the min-norm fit is smooth ⇒ error falls again if n is large |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.