USAAIO Lesson 34, from Week 12 on generalization theory, fully worked. It derives Markov's inequality from the layer-cake, or indicator, trick with no steps skipped, obtains Chebyshev's by applying Markov to the squared deviation, and states Hoeffding's, proving its exponential decay in n empirically on a biased coin. It then solves the sample-complexity inversion n >= ln(2/delta)/(2 eps^2) by hand. From there it builds up shattering and VC dimension, from a one-dimensional threshold to a two-dimensional line, which shatters 3 points but fails XOR, managing at best 3 of 4. It covers Sauer's lemma and the VC generalization bound term by term, and demonstrates double descent with a real least-norm fit whose test error peaks at the interpolation threshold. One running example, a coin with p=0.3, threads the whole deck, and every inequality and every printed number was produced by real execution. The lesson runs to 62 slides.
Subject: Machine Learning · 115 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 34 · Week 12 (Generalization Theory)
Why a model that fits the sample should work on new data. We derive Markov from one indicator, get Chebyshev for free, prove Hoeffding decays exponentially, build VC dimension from shattering, read the generalization bound, and watch double descent break the classical story — every step shown, every number run.
Objectives
P(X ≥ a) ≤ E[X]/a from the indicator inequality a·1[X≥a] ≤ X, and know exactly which assumption it usesP(|X−μ| ≥ kσ) ≤ 1/k² by applying Markov to the squared deviation (X−μ)²P(|X̄−E[X]| ≥ t) ≤ 2e^{−2nt²} and invert it for the sample size n ≥ ln(2/δ)/(2ε²)3 points but fails XOR (VC = d+1)gen ≤ train + √(VC·ln n / n) and explain why double descent beats it in practiceWarm-up
Discussion prompt
Before we open Lesson 34: Concentration Inequalities & Generalization: without looking back, what was the main idea of Gradient Checking & Debugging, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
central finite differences and their O(ε²) accuracy, gradient checking by relative error, the common backprop bugs (broadcasting, wrong axis, missing zero_grad, sign), vanishing/exploding-gradient diagnosis, and NaN debugging. Build gradient_check and use it to catch a deliberately buggy gradient.
Section
Part 1 of 6
Concept
One example threads the whole lesson. We flip a coin whose true heads-probability is p = 0.3 (we call heads 1, tails 0). We never see p; we only see flips and form the sample mean X̄.
\[ \bar X = \frac{1}{n}\sum_{i=1}^{n} X_i, \qquad X_i \in \{0,1\}, \qquad \mathbb{E}[X_i] = p = 0.3 \]
X̄ is our estimate of p, and in ML it plays the role of training error estimating true error. The whole lesson asks: how far can X̄ stray from p?
Counterexample
Discussion prompt
One example threads the whole lesson. We flip a coin whose true heads-probability is p = 0.3 (we call heads 1, tails 0). We never see p; we only see flips and form the sample mean X̄.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
X̄ is our estimate of p, and in ML it plays the role of training error estimating true error. The whole lesson asks: how far can X̄ stray from p?
Estimation
Predict first
Flip the p = 0.3 coin 20 times with a fixed seed and read off X̄. Watch it miss 0.3. Runnable as written:
Commit before you compute: what does One sample already strays come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: 8 heads in 20 flips → X̄ = 0.40, not 0.30
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. A finite sample is noisy: the estimate landed 0.10 above the truth.
Worked example
Flip the p = 0.3 coin 20 times with a fixed seed and read off X̄. Watch it miss 0.3. Runnable as written:
import numpy as np
rng = np.random.default_rng(0)
p = 0.3
flips = rng.binomial(1, p, size=20) # 20 coin flips, heads = 1
print('flips:', flips)
print('sum :', flips.sum())
print('Xbar :', flips.mean())8 heads in 20 flips → X̄ = 0.40, not 0.30
Why: A finite sample is noisy: the estimate landed 0.10 above the truth. Concentration inequalities put a ceiling on how often — and how far — that happens.
| quantity | value (verified) |
|---|---|
| flips (seed 0) | [0 0 0 0 1 1 0 1 0 1 1 0 1 0 1 0 1 0 0 0] |
| sum of heads | 8 |
| X̄ = sum / 20 | 0.40 |
| true p | 0.30 |
Comparison
Comparison matrix
From One sample already strays: refill the value (verified) column from what you know. The rest of the table is as it appeared.
| quantity | value (verified) |
|---|---|
| flips (seed 0) | [0 0 0 0 1 1 0 1 0 1 1 0 1 0 1 0 1 0 0 0] |
| sum of heads | 8 |
| X̄ = sum / 20 | 0.40 |
| true p | 0.30 |
Intuition
A learner picks the model that looks best on the training sample. If sample statistics could wander arbitrarily far from the truth, a low training error would tell you nothing about future data.
Generalization theory earns its keep by proving the opposite: with enough data, X̄ is pinned near p — so training error is pinned near true error. 'Pinned' is exactly what a concentration inequality guarantees.
We build three of them, weakest to strongest: Markov (mean only), Chebyshev (adds variance), Hoeffding (adds boundedness). Each new assumption buys a tighter bound.
Analogy
Discussion prompt
Explain Why this is the whole game by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
A learner picks the model that looks best on the training sample. If sample statistics could wander arbitrarily far from the truth, a low training error would tell you nothing about future data.
Section
Part 2 of 6 — the weakest bound
Concept
Let X ≥ 0 and pick a level a > 0. Define the indicator 1[X ≥ a], which is 1 when the event happens and 0 otherwise. The entire proof rests on one pointwise inequality:
\[ a \cdot \mathbf{1}[X \ge a] \;\le\; X \qquad \text{(true for every outcome)} \]
Check both cases: if X ≥ a the left side is a and a ≤ X; if X < a the left side is 0 ≤ X (since X ≥ 0). Either way it holds. Now just take expectations.
Explain it
Discussion prompt
Explain The one inequality behind Markov to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Let X ≥ 0 and pick a level a > 0. Define the indicator 1[X ≥ a], which is 1 when the event happens and 0 otherwise. The entire proof rests on one pointwise inequality:
Ranking
Put in order
Put the moves of Derive Markov, step by step into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Holds for every outcome, as we checked case by case.
Worked example
Start from the pointwise inequality
Why: Holds for every outcome, as we checked case by case.
\[ a \cdot \mathbf{1}[X \ge a] \le X \]
Take expectations of both sides
Why: Expectation is monotone: if U ≤ V pointwise then E[U] ≤ E[V]. The inequality survives.
\[ \mathbb{E}\big[a \cdot \mathbf{1}[X \ge a]\big] \le \mathbb{E}[X] \]
Pull the constant a out; E of an indicator is a probability
Why: E[1[A]] = P(A) by definition, so E[a·1[X≥a]] = a·P(X≥a).
\[ a \, P(X \ge a) \le \mathbb{E}[X] \]
Divide both sides by a > 0
Why: Dividing by a positive number keeps the direction. This IS Markov's inequality.
\[ \boxed{\,P(X \ge a) \le \dfrac{\mathbb{E}[X]}{a}\,} \]
Notation
Annotate
From Derive Markov, step by step — read this one piece at a time. What is each part doing?
On: \( a \cdot \mathbf{1}[X \ge a] \le X \)
Concept
The only ingredients were X ≥ 0 and that E[X] exists. No variance, no distribution shape, no boundedness. That universality is Markov's strength and its weakness: it must hold for every non-negative variable with that mean, so it can't be tight for any particular one.
one-sided vs two-sided — Markov bounds only the RIGHT tail P(X ≥ a) of a non-negative variable. To bound how far X falls on BOTH sides of its mean, we need the deviation itself — that is the job of Chebyshev.
Missing information
Discussion prompt
Take X ~ Exponential(1), so E[X] = 1. Draw two million samples and compare the empirical right-tail P(X ≥ a) to the Markov ceiling 1/a. Complete snippet:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
At a=4 the truth is 0.0183 but Markov only promises ≤ 0.25 — a 14× gap. Valid, but loose, exactly because it ignores the exponential's shape.
Worked example
Take X ~ Exponential(1), so E[X] = 1. Draw two million samples and compare the empirical right-tail P(X ≥ a) to the Markov ceiling 1/a. Complete snippet:
import numpy as np
rng = np.random.default_rng(0)
X = rng.exponential(1.0, 2_000_000) # E[X] = 1
for a in [2, 4, 5]:
emp = (X >= a).mean()
print(f'a={a}: P(X>={a})={emp:.4f} Markov 1/a={1/a:.4f}')Every empirical tail sits far under 1/a
Why: At a=4 the truth is 0.0183 but Markov only promises ≤ 0.25 — a 14× gap. Valid, but loose, exactly because it ignores the exponential's shape.
| a | empirical P(X ≥ a) | Markov bound 1/a |
|---|---|---|
| 2 | 0.1352 | 0.5000 |
| 4 | 0.0183 | 0.2500 |
| 5 | 0.0067 | 0.2000 |
Trade off
Comparison matrix
From Markov holds — and how loosely: every row here is a choice with a cost. Fill the empirical P(X ≥ a) column, then say which row you would actually pick and what you give up for it.
| a | empirical P(X ≥ a) | Markov bound 1/a |
|---|---|---|
| 2 | 0.1352 | 0.5000 |
| 4 | 0.0183 | 0.2500 |
| 5 | 0.0067 | 0.2000 |
Anomaly
Predict first
A student writes this, and it looks reasonable:
Markov gives P(X ≥ 4) ≤ 0.25, so roughly a quarter of the mass sits above 4.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: An upper bound is a CEILING, not an estimate.
Read 0.25 as a worst-case ceiling. To get near the true 0.018, feed the bound more structure.
Why: An upper bound is a CEILING, not an estimate. The true value is 0.0183 — about 14× smaller. Markov used only the mean, so it is deliberately pessimistic.
Trap
Markov gives P(X ≥ 4) ≤ 0.25, so roughly a quarter of the mass sits above 4.
Treat 0.25 as ≈ the probability
Why: An upper bound is a CEILING, not an estimate. The true value is 0.0183 — about 14× smaller. Markov used only the mean, so it is deliberately pessimistic.
Read 0.25 as a worst-case ceiling. To get near the true 0.018, feed the bound more structure.
Add variance → Chebyshev; add boundedness → Hoeffding
Why: Each extra assumption tightens the bound. Markov's job is not to be tight — it is to hold when you know almost nothing, and to be the lemma the other bounds are built from.
Section
Part 3 of 6 — variance buys two sides
Concept
We want a two-sided bound on |X − μ|. The variable |X − μ| can be positive from either side, so we square it: (X − μ)² is non-negative — precisely what Markov needs.
And its mean is famous: E[(X − μ)²] = Var(X) = σ². So apply Markov to Y = (X − μ)² at level a = (kσ)². The next slide does it move by move.
Fill the middle
Fill in the blanks
From Derive Chebyshev from Markov — finish the line. Write what belongs on the right of the equals sign before you look.
P(|X-\mu| \ge k\sigma) = P\big((X-\mu)^2 \ge (k\sigma)^2\big)
Why: Producing the right-hand side unprompted is the difference between recognising this line and being able to use it. |X − μ| ≥ kσ happens for exactly the same outcomes as (X − μ)² ≥ (kσ)² — squaring both non-negative sides preserves the event.
Worked example
The events are identical
Why: |X − μ| ≥ kσ happens for exactly the same outcomes as (X − μ)² ≥ (kσ)² — squaring both non-negative sides preserves the event.
\[ P(|X-\mu| \ge k\sigma) = P\big((X-\mu)^2 \ge (k\sigma)^2\big) \]
Apply Markov to Y = (X − μ)² at level (kσ)²
Why: Y ≥ 0, so Markov gives P(Y ≥ a) ≤ E[Y]/a with E[Y] = σ² and a = k²σ².
\[ P\big((X-\mu)^2 \ge k^2\sigma^2\big) \le \frac{\mathbb{E}[(X-\mu)^2]}{k^2\sigma^2} = \frac{\sigma^2}{k^2\sigma^2} \]
Cancel σ²
Why: The variance cancels top and bottom, leaving a bound that depends only on k. This is Chebyshev.
\[ \boxed{\,P(|X-\mu| \ge k\sigma) \le \dfrac{1}{k^2}\,} \]
Translation
\( P\big((X-\mu)^2 \ge k^2\sigma^2\big) \le \frac{\mathbb{E}[(X-\mu)^2]}{k^2\sigma^2} = \frac{\sigma^2}{k^2\sigma^2} \)
Draw it
Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.
Pattern
Predict first
The table runs: P(|X − 1| ≥ 2) | 0.0498 · P(Y ≥ 4), Y = (X−1)² | 0.0498 · E[Y] (= Var = σ²) | 1.0
In Verify Chebyshev really is Markov-on-Y, given the rows so far: what is the next one — the row where quantity is Markov E[Y]/4 = 1/k²?
Correct: Markov E[Y]/4 = 1/k² | 0.25
| quantity | value (verified) |
|---|---|
| P(|X − 1| ≥ 2) | 0.0498 |
| P(Y ≥ 4), Y = (X−1)² | 0.0498 |
| E[Y] (= Var = σ²) | 1.0 |
| Markov E[Y]/4 = 1/k² | 0.25 |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. Squaring the deviation maps the two-sided event onto a one-sided event of Y, exactly as claimed, and E[Y] recovers the variance 1.
Worked example
Let's confirm the derivation numerically on Exponential(1) (μ = σ = 1) at k = 2. We check that P(|X−1| ≥ 2) equals P(Y ≥ 4) and that E[Y]/4 = 1/4:
import numpy as np
rng = np.random.default_rng(0)
X = rng.exponential(1.0, 2_000_000) # mu = sigma = 1
Y = (X - 1.0)**2
print('P(|X-1|>=2) :', round((np.abs(X-1)>=2).mean(), 4))
print('P(Y>=4) :', round((Y >= 4).mean(), 4))
print('E[Y] :', round(Y.mean(), 4))
print('Markov E[Y]/4:', round(Y.mean()/4, 4))The two events give the same 0.0498; E[Y] ≈ 1 = σ²
Why: Squaring the deviation maps the two-sided event onto a one-sided event of Y, exactly as claimed, and E[Y] recovers the variance 1.
| quantity | value (verified) |
|---|---|
| P(|X − 1| ≥ 2) | 0.0498 |
| P(Y ≥ 4), Y = (X−1)² | 0.0498 |
| E[Y] (= Var = σ²) | 1.0 |
| Markov E[Y]/4 = 1/k² | 0.25 |
Discrimination
Sort into buckets
Sort these by value (verified), from memory, without looking back at Verify Chebyshev really is Markov-on-Y. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Fill the middle
Fill in the blanks
From Chebyshev across several k — one line has had its right-hand side removed. Put it back.
import numpy as np
rng = np.random.default_rng(0)
X = rng.exponential(1.0, 2_000_000) # mu = 1, sigma = 1
for k in [2, 3, 4]:
emp = (np.abs(X - 1.0) >= k).mean()
print(f'k=___: P(|X-1|>=___)=___ 1/k^2=___')
Why: X is what everything below it consumes, so the wrong expression here fails later and somewhere else. Chebyshev is tighter than Markov but still worst-case over all distributions with σ = 1, so a specific one beats it comfortably.
Worked example
Sweep k = 2, 3, 4 on the same Exponential(1). Each empirical two-sided tail must sit under 1/k². Runnable:
import numpy as np
rng = np.random.default_rng(0)
X = rng.exponential(1.0, 2_000_000) # mu = 1, sigma = 1
for k in [2, 3, 4]:
emp = (np.abs(X - 1.0) >= k).mean()
print(f'k={k}: P(|X-1|>={k})={emp:.4f} 1/k^2={1/k**2:.4f}')All three hold; the exponential's fat right tail keeps them loose
Why: Chebyshev is tighter than Markov but still worst-case over all distributions with σ = 1, so a specific one beats it comfortably.
| k | empirical P(|X−1| ≥ k) | Chebyshev 1/k² |
|---|---|---|
| 2 | 0.0498 | 0.2500 |
| 3 | 0.0183 | 0.1111 |
| 4 | 0.0067 | 0.0625 |
Pattern
Step through it
Step through Chebyshev across several k one row at a time. What is driving the change, and what would the row after the last one be?
Anomaly
Predict first
A student writes this, and it looks reasonable:
To bound P(|X−μ| ≥ kσ), apply Markov directly to |X − μ|: P(|X−μ| ≥ kσ) ≤ E|X−μ| / (kσ).
It is wrong. Say what breaks — and say it before you turn the page.
Correct: This is a valid bound, but it uses the mean absolute deviation E|X−μ|, not the variance — it does NOT give the clean 1/k².
Apply Markov to the squared deviation (X − μ)², whose mean is exactly σ².
Why: This is a valid bound, but it uses the mean absolute deviation E|X−μ|, not the variance — it does NOT give the clean 1/k². You've thrown away the σ² cancellation that makes Chebyshev memorable and distribution-free in k.
Trap
To bound P(|X−μ| ≥ kσ), apply Markov directly to |X − μ|: P(|X−μ| ≥ kσ) ≤ E|X−μ| / (kσ).
Stop at E|X − μ| / (kσ)
Why: This is a valid bound, but it uses the mean absolute deviation E|X−μ|, not the variance — it does NOT give the clean 1/k². You've thrown away the σ² cancellation that makes Chebyshev memorable and distribution-free in k.
Apply Markov to the squared deviation (X − μ)², whose mean is exactly σ².
P((X−μ)² ≥ k²σ²) ≤ σ²/(k²σ²) = 1/k²
Why: Squaring makes the mean equal Var, and the σ² cancels — that cancellation is the whole point. The '²' in Chebyshev is not optional.
Section
Part 4 of 6 — averages of bounded variables
Concept
For a single variable, Chebyshev gives P(|X−μ| ≥ t) ≤ σ²/t² — decay only like 1/t². But we average n flips, and averaging shrinks variance: Var(X̄) = σ²/n. Feeding that into Chebyshev already gives σ²/(n t²), decaying like 1/n.
Hoeffding does far better. For an average of n bounded i.i.d. variables the tail decays exponentially in n — the difference between 'need thousands of samples' and 'need a few hundred'.
Concept
For i.i.d. X₁,…,Xₙ each bounded in [0, 1] with mean E[X], the sample mean X̄ satisfies, for any t > 0:
\[ P\big(|\bar X - \mathbb{E}[X]| \ge t\big) \;\le\; 2\,\exp\!\big(-2 n t^2\big) \]
The 2 is for the two tails; the exponent carries n (more data) and t² (bigger deviations are far rarer). Our coin is {0,1}-valued, so it lands squarely in Hoeffding's [0,1] hypothesis — this is the bound built for it.
Faded example
Fill in the blanks
Hoeffding vs the coin's real tail, with the scaffolding fading: two lines are gone now — fill both.
import numpy as np
rng = np.random.default_rng(0)
p, n, t = 0.3, 100, 0.1
means = rng.binomial(n, p, size=200_000) / n # each is Xbar of n flips
emp = (np.abs(means - p) >= t).mean()
bound = *2np.exp(-2nt2)
print('empirical P(|Xbar-p|>=t):', round(emp, 4))
print('Hoeffding 2exp(-2 n t^2):', round(bound, 4))
Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. The bound holds with margin. It is looser than the truth (Hoeffding is distribution-free over all [0,1] variables) but far tighter than what Markov or Chebyshev would give here.
Worked example
Simulate 200000 experiments, each n = 100 flips of the p = 0.3 coin, and measure how often X̄ misses p by ≥ 0.1. Compare to Hoeffding. Complete snippet:
import numpy as np
rng = np.random.default_rng(0)
p, n, t = 0.3, 100, 0.1
means = rng.binomial(n, p, size=200_000) / n # each is Xbar of n flips
emp = (np.abs(means - p) >= t).mean()
bound = 2*np.exp(-2*n*t**2)
print('empirical P(|Xbar-p|>=t):', round(emp, 4))
print('Hoeffding 2exp(-2 n t^2):', round(bound, 4))Empirical 0.0302 sits under the Hoeffding ceiling 0.2707
Why: The bound holds with margin. It is looser than the truth (Hoeffding is distribution-free over all [0,1] variables) but far tighter than what Markov or Chebyshev would give here.
| quantity | value (verified) |
|---|---|
| n, t | 100, 0.10 |
| empirical P(|X̄ − p| ≥ 0.1) | 0.0302 |
| Hoeffding 2e^{−2nt²} | 0.2707 |
| exact binomial tail | 0.0299 |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Empirical 0.0302 sits under the Hoeffding ceiling 0.2707
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Simulate 200000 experiments, each n = 100 flips of the p = 0.3 coin, and measure how often X̄ misses p by ≥ 0.1. Compare to Hoeffding. Complete snippet:
Faded example
Fill in the blanks
The exponential collapse in n, with the scaffolding fading: two lines are gone now — fill both.
import numpy as np
t = 0.1
for n in [50, 100, 200, 500, 1000]:
bound = *2np.exp(-2nt2)
print(f'n=___: 2exp(-2 n t^2) = ___')
Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. Each extra sample multiplies the bound by exp(−2t²) < 1, so the tail decays geometrically in n.
Worked example
Fix t = 0.1 and grow n. The whole reason Hoeffding matters is how violently the ceiling drops. Runnable:
import numpy as np
t = 0.1
for n in [50, 100, 200, 500, 1000]:
bound = 2*np.exp(-2*n*t**2)
print(f'n={n:4d}: 2exp(-2 n t^2) = {bound:.6f}')From n=50 to n=500 the ceiling falls from 0.74 to 0.00009
Why: Each extra sample multiplies the bound by exp(−2t²) < 1, so the tail decays geometrically in n. That exponential is what makes learning from finitely many samples possible.
| n | 2·exp(−2·n·0.1²) (verified) |
|---|---|
| 50 | 0.735759 |
| 100 | 0.270671 |
| 200 | 0.036631 |
| 500 | 0.000091 |
| 1000 | 0.000000 |
Pattern
Step through it
Step through The exponential collapse in n one row at a time. What is driving the change, and what would the row after the last one be?
Intuition
Replace 'coin' with 'a fixed classifier' and 'Xᵢ' with 'is example i misclassified?' (a {0,1} variable). Then X̄ is the training error and E[X] is the true error.
Hoeffding now reads: training error is within t of true error with probability ≥ 1 − 2e^{−2nt²}. That is the Probably Approximately Correct guarantee — 'approximately' is t, 'probably' is the 1 − δ.
So the practical question becomes: how many samples n do I need to promise accuracy ε with confidence 1 − δ? Invert the bound.
Ranking
Put in order
Put the moves of Invert Hoeffding for sample complexity into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. We want P(|X̄ − E[X]| ≥ ε) ≤ δ.
Worked example
Demand the failure probability be ≤ δ
Why: We want P(|X̄ − E[X]| ≥ ε) ≤ δ. Hoeffding's ceiling is 2e^{−2nε²}, so it suffices to force that ceiling below δ.
\[ 2\,e^{-2 n \varepsilon^2} \le \delta \]
Divide by 2 and take logs
Why: ln is increasing, so it preserves ≤; ln of the exponential returns its exponent.
\[ -2 n \varepsilon^2 \le \ln\!\frac{\delta}{2} \;\Longleftrightarrow\; -2 n \varepsilon^2 \le -\ln\!\frac{2}{\delta} \]
Multiply by −1 (flip the inequality), divide by 2ε²
Why: Multiplying by a negative flips ≥ to the sample-size condition. Dividing by the positive 2ε² keeps direction.
\[ \boxed{\,n \ge \dfrac{\ln(2/\delta)}{2\varepsilon^2}\,} \]
Notation
Annotate
From Invert Hoeffding for sample complexity — read this one piece at a time. What is each part doing?
On: \( \boxed{\,n \ge \dfrac{\ln(2/\delta)}{2\varepsilon^2}\,} \)
Fill the middle
Fill in the blanks
From Sample complexity, computed and checked — one line has had its right-hand side removed. Put it back.
import numpy as np
eps, delta = 0.05, 0.05
n_req = np.log(2/delta) / (2*eps2)
n = int(np.ceil(n_req))**
print('ln(2/delta) :', round(np.log(2/delta), 4))
print('n >= ... :', round(n_req, 2), '-> ceil', n)
print('check bound at n :', round(2np.exp(-2neps*2), 4), '<= 0.05')
Why: n is what everything below it consumes, so the wrong expression here fails later and somewhere else. The ceiling formula demands 738 flips to promise ±0.05 accuracy 95% of the time.
Worked example
Ask for ε = 0.05 accuracy at δ = 0.05 confidence. Compute n, round up, and plug back to confirm the ceiling really drops under δ. Runnable:
import numpy as np
eps, delta = 0.05, 0.05
n_req = np.log(2/delta) / (2*eps**2)
n = int(np.ceil(n_req))
print('ln(2/delta) :', round(np.log(2/delta), 4))
print('n >= ... :', round(n_req, 2), '-> ceil', n)
print('check bound at n :', round(2*np.exp(-2*n*eps**2), 4), '<= 0.05')n ≥ 737.78 → 738 samples; the bound at 738 is 0.0499 ≤ 0.05
Why: The ceiling formula demands 738 flips to promise ±0.05 accuracy 95% of the time. Plugging 738 back gives 0.0499, just under δ — the inversion is exact.
| quantity | value (verified) |
|---|---|
| ln(2/δ) | 3.6889 |
| 2ε² | 0.005 |
| n ≥ ln(2/δ)/(2ε²) | 737.78 → 738 |
| 2e^{−2·738·ε²} | 0.0499 (≤ 0.05 ✓) |
Concept
Each bound adds one assumption and tightens. Read the ladder top to bottom — it is the single most exam-tested idea in this lesson.
| bound | needs | tail on X̄ | decay in n |
|---|---|---|---|
| Markov | X ≥ 0, mean | E[X]/a | — (single var) |
| Chebyshev | + variance σ² | σ²/(n t²) | 1/n |
| Hoeffding | + bounded [0,1], i.i.d. | 2e^{−2nt²} | exponential |
More structure ⇒ tighter bound. When your variable is bounded — as {0,1} losses always are — reach for Hoeffding.
Anomaly
Predict first
A student writes this, and it looks reasonable:
The coin's X̄ obeys 2e^{−2nt²}, so use the same bound for the average of n draws from Exponential(1).
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Hoeffding REQUIRES each variable lie in a bounded range [a,b] — the exponent is really 2nt²/(b−a)².
Match the tool to the support. For an unbounded variable with known variance, use Chebyshev on the average.
Why: Hoeffding REQUIRES each variable lie in a bounded range [a,b] — the exponent is really 2nt²/(b−a)². The exponential is unbounded (b = ∞), so (b−a)² = ∞ makes the exponent 0 and the bound the useless '≤ 2'. The clean {0,1} version silently assumed b−a = 1.
Trap
The coin's X̄ obeys 2e^{−2nt²}, so use the same bound for the average of n draws from Exponential(1).
Apply 2e^{−2nt²} to the exponential average
Why: Hoeffding REQUIRES each variable lie in a bounded range [a,b] — the exponent is really 2nt²/(b−a)². The exponential is unbounded (b = ∞), so (b−a)² = ∞ makes the exponent 0 and the bound the useless '≤ 2'. The clean {0,1} version silently assumed b−a = 1.
Match the tool to the support. For an unbounded variable with known variance, use Chebyshev on the average.
P(|X̄ − μ| ≥ t) ≤ σ²/(n t²) via Chebyshev
Why: Averaging shrinks variance to σ²/n, so Chebyshev still gives a valid 1/n bound with no boundedness needed. Hoeffding's exponential speed is a reward for boundedness — you cannot claim it for free.
Break the constraint
Discussion prompt
The rule this trap just fixed:
Averaging shrinks variance to σ²/n, so Chebyshev still gives a valid 1/n bound with no boundedness needed. Hoeffding's exponential speed is a reward for boundedness — you cannot claim it for free.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
Hoeffding REQUIRES each variable lie in a bounded range [a,b] — the exponent is really 2nt²/(b−a)². The exponential is unbounded (b = ∞), so (b−a)² = ∞ makes the exponent 0 and the bound the useless '≤ 2'. The clean {0,1} version silently assumed b−a = 1.
Section
Part 5 of 6 — shattering
Intuition
Hoeffding pins the training error of one fixed classifier near its true error. But a learner searches over a whole class H and keeps whichever looks best — so it can get lucky on some hypothesis in H by chance.
The more hypotheses H contains, the more chances to get lucky, and the wider the possible train–true gap. We need a number for 'how rich is H'. Counting hypotheses fails (often infinite), so we count something smarter: how many points H can label every possible way.
Concept
shatter — A hypothesis class H shatters a set of m points if, for EVERY one of the 2^m ways to label them +/−, some hypothesis in H realizes that exact labeling. If even one labeling is impossible, H does not shatter the set.
Shattering is an all-or-nothing test on a specific point set. m points have 2ᵐ labelings; H shatters them only if it can hit all 2ᵐ.
Definition probe
Sort into buckets
Every line below is part of the definition of one-sided vs two-sided or of shatter — one or the other, never both. Put each where it belongs.
Estimation
Predict first
Simplest class: a 1-D threshold h(x) = sign(x − θ) (with an optional flip). Two points x = 1, 2 have 2² = 4 labelings. We brute-force over thresholds and check all four are realizable:
Commit before you compute: what does A threshold shatters 2 points on a line come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: All 4 labelings realizable → the threshold class shatters 2 points
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Sliding θ between, below, or above the points, plus the sign flip, reaches every labeling.
Worked example
Simplest class: a 1-D threshold h(x) = sign(x − θ) (with an optional flip). Two points x = 1, 2 have 2² = 4 labelings. We brute-force over thresholds and check all four are realizable:
import numpy as np
pts = np.array([1.0, 2.0])
def h(x, theta, sign):
return np.where(x > theta, sign, -sign)
targets = [(-1,-1), (-1,1), (1,-1), (1,1)]
for tgt in targets:
ok = any(tuple(h(pts, th, s)) == tgt
for th in np.linspace(0, 3, 31) for s in (1, -1))
print('labeling', tgt, 'realizable:', ok)All 4 labelings realizable → the threshold class shatters 2 points
Why: Sliding θ between, below, or above the points, plus the sign flip, reaches every labeling. So VC of 1-D thresholds is at least 2... but is it more?
| labeling of (x=1, x=2) | realizable? (verified) |
|---|---|
| (−1, −1) | True |
| (−1, +1) | True |
| (+1, −1) | True |
| (+1, +1) | True |
Invariant
Step through it
Step through A threshold shatters 2 points on a line one row at a time. One of these columns never changes — find it, and say why it cannot.
Worked example
Now three points 1, 2, 3. The alternating labeling (+, −, +) needs the classifier to flip sign twice — a single threshold can't. Count how many of the 8 labelings are actually reachable:
import numpy as np
from itertools import product
pts = np.array([1.0, 2.0, 3.0])
def h(x, theta, sign):
return np.where(x > theta, sign, -sign)
good = 0
for tgt in product((1, -1), repeat=3):
ok = any(tuple(h(pts, th, s)) == tgt
for th in np.linspace(0, 4, 401) for s in (1, -1))
good += ok
print('realizable labelings of 3 points:', good, 'of 8')Only 6 of 8 labelings work — (+,−,+) and (−,+,−) are impossible
Why: A threshold splits the line into one 'low' side and one 'high' side; it cannot carve out a middle group. So 1-D thresholds shatter 2 points but not 3 — their VC dimension is exactly 2.
| labeling of (1, 2, 3) | realizable? |
|---|---|
| (+, +, +) / (−, −, −) / … | 6 of these work |
| (+, −, +) | impossible |
| (−, +, −) | impossible |
| total realizable | 6 of 8 |
Discrimination
Sort into buckets
Sort these by realizable?, from memory, without looking back at A threshold FAILS on 3 points. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Concept
VC dimension — The VC dimension of H is the size of the LARGEST point set H can shatter. Formally: the largest m such that SOME set of m points is shattered by H. It is a single integer summarizing the capacity of the whole class.
Two subtleties that trip people up: you need only find one set of m points that is shattered (not all sets); and to prove VC < m+1, you must show no set of m+1 points can be shattered.
Worked example
Move to 2-D linear classifiers (a separating line, d = 2). Take a triangle — three non-collinear points — and check that a line can realize all 2³ = 8 labelings. We fit a hard-margin LinearSVC per labeling:
import numpy as np
from itertools import product
from sklearn.svm import LinearSVC
P = np.array([[0., 0.], [2., 0.], [0., 2.]]) # non-collinear triangle
n_sep = 0
for labels in product([0, 1], repeat=3):
if len(set(labels)) == 1:
n_sep += 1 # single class: trivial
continue
clf = LinearSVC(C=1e6, max_iter=100000).fit(P, np.array(labels))
n_sep += (clf.predict(P) == np.array(labels)).mean() == 1.0
print('labelings a line can realize:', n_sep, 'of', 2**3)8 of 8 labelings realized → a line shatters 3 points
Why: For every +/− split of the triangle, some line separates the + from the − points (single-class labelings are trivially 'separated' by any line). So VC of a 2-D linear classifier is at least 3.
| labelings tried | realized by a line (verified) |
|---|---|
| 2 single-class (all +, all −) | 2 (trivial) |
| 6 mixed labelings | 6 (SVC acc = 1.0) |
| total | 8 of 8 → shattered |
Fill the middle
Fill in the blanks
From A line CANNOT shatter 4 points — XOR — one line has had its right-hand side removed. Put it back.
import numpy as np
X = np.array([[0.,0.],[1.,1.],[1.,0.],[0.,1.]])
y = np.array([0, 0, 1, 1]) # XOR: diagonal pairs share a class
g = np.linspace(-3, 3, 61)
best = 0.0
for a in g:
for b in g:
for c in g:
pred = ((aX[:,0] + bX[:,1] + c) > 0).astype(int)
best = max(best, (pred == y).mean(), (1-pred == y).mean())
print('best line accuracy on XOR:', best)
Why: g is what everything below it consumes, so the wrong expression here fails later and somewhere else. No line separates the two diagonal classes, so the (+,−,+,−)-style XOR labeling is unrealizable.
Worked example
Four points in the XOR arrangement defeat every line. Grid-search all lines sign(a·x + b·y + c) and record the best accuracy any line achieves on the XOR labeling:
import numpy as np
X = np.array([[0.,0.],[1.,1.],[1.,0.],[0.,1.]])
y = np.array([0, 0, 1, 1]) # XOR: diagonal pairs share a class
g = np.linspace(-3, 3, 61)
best = 0.0
for a in g:
for b in g:
for c in g:
pred = ((a*X[:,0] + b*X[:,1] + c) > 0).astype(int)
best = max(best, (pred == y).mean(), (1-pred == y).mean())
print('best line accuracy on XOR:', best)Best a line can do on XOR is 0.75 — it always misses ≥ 1 of 4
Why: No line separates the two diagonal classes, so the (+,−,+,−)-style XOR labeling is unrealizable. A line fails to shatter 4 points; combined with shattering 3, VC(line in ℝ²) = 3.
| point set | best line accuracy (verified) |
|---|---|
| 3-point triangle (all labelings) | 1.00 → shattered |
| 4-point XOR | 0.75 → NOT shattered |
| ⇒ VC dimension of a 2-D line | 3 = d + 1 |
Comparison
Comparison matrix
From A line CANNOT shatter 4 points — XOR: refill the best line accuracy (verified) column from what you know. The rest of the table is as it appeared.
| point set | best line accuracy (verified) |
|---|---|
| 3-point triangle (all labelings) | 1.00 → shattered |
| 4-point XOR | 0.75 → NOT shattered |
| ⇒ VC dimension of a 2-D line | 3 = d + 1 |
Concept
The two experiments (shatter 3, fail 4) are the d = 2 case of a general theorem for linear classifiers in ℝᵈ (with a bias term):
\[ \text{VC}\big(\text{linear classifiers in } \mathbb{R}^d\big) = d + 1 \]
The +1 is the bias/offset degree of freedom — the same ones-column that let a regression line leave the origin. On a plane (d = 2) that gives 3, matching what we just brute-forced.
Anomaly
Predict first
A student writes this, and it looks reasonable:
A linear classifier in ℝ² has 2 parameters of direction, so its VC dimension is 2.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: This drops the bias term. A line through the ORIGIN (no offset) shatters only 2 points, but a general line has a third degree of freedom — its offset c — and shatters 3.
Count the offset. A general line a·x + b·y + c has d + 1 = 3 free parameters.
Why: This drops the bias term. A line through the ORIGIN (no offset) shatters only 2 points, but a general line has a third degree of freedom — its offset c — and shatters 3. Our brute force realized all 8 labelings of 3 points, so 2 is too small.
Trap
A linear classifier in ℝ² has 2 parameters of direction, so its VC dimension is 2.
Answer VC = d = 2
Why: This drops the bias term. A line through the ORIGIN (no offset) shatters only 2 points, but a general line has a third degree of freedom — its offset c — and shatters 3. Our brute force realized all 8 labelings of 3 points, so 2 is too small.
Count the offset. A general line a·x + b·y + c has d + 1 = 3 free parameters.
VC = d + 1 = 3
Why: The '+1' is the bias c. It is the same extra degree of freedom as the ones-column in regression, and it is exactly what lets the line shatter one more point than d alone would.
Section
Part 6 of 6 — and where it breaks
Concept
A class with VC dimension d can label at most Πₕ(m) distinct ways on any m points. Sauer's lemma bounds that growth function by a polynomial once m > d:
\[ \Pi_H(m) \;\le\; \sum_{i=0}^{d} \binom{m}{i} \;\le\; \Big(\tfrac{e\,m}{d}\Big)^{d} \quad (m > d) \]
This is the key that turns capacity into a bound: below m = d the class labels all 2ᵐ ways (shattering), but past d the count grows only polynomially, not exponentially — so infinitely many hypotheses still behave like a small finite set.
Pattern
Predict first
The table runs: 3 | 8 | 8 · 4 | 15 | 16 · 5 | 26 | 32
In Watch the growth go sub-exponential, given the rows so far: what is the next one — the row where m is 10?
Correct: 10 | 176 | 1024
| m | Sauer Σ_{i≤3} C(m,i) | 2ᵐ |
|---|---|---|
| 3 | 8 | 8 |
| 4 | 15 | 16 |
| 5 | 26 | 32 |
| 10 | 176 | 1024 |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. Up to the VC dimension the class realizes every labeling (8 = 2³).
Worked example
For a line in ℝ² (d = 3), tabulate Sauer's sum Σ_{i≤3} C(m, i) against the naive 2ᵐ. Watch them agree up to m = 3, then diverge:
from math import comb
d = 3 # VC dimension of a line in R^2
for m in [3, 4, 5, 10]:
sauer = sum(comb(m, i) for i in range(d+1))
print(f'm={m:2d}: growth <= {sauer:4d} vs 2^m = {2**m}')At m=3 both are 8 (shattering); by m=10, 176 vs 1024
Why: Up to the VC dimension the class realizes every labeling (8 = 2³). Beyond it, Sauer's polynomial (176) falls far below the exponential (1024) — capacity is effectively finite.
| m | Sauer Σ_{i≤3} C(m,i) | 2ᵐ |
|---|---|---|
| 3 | 8 | 8 |
| 4 | 15 | 16 |
| 5 | 26 | 32 |
| 10 | 176 | 1024 |
Pattern
Step through it
Step through Watch the growth go sub-exponential one row at a time. What is driving the change, and what would the row after the last one be?
Intuition
Hoeffding controls one hypothesis. A learner tries many, so we pay a union bound: the chance that some hypothesis in H is fooled is at most the sum of the individual failure chances.
If H behaved like M distinct hypotheses, the total failure budget is M · 2e^{−2nt²}. Forcing that below δ and solving for t puts a √(ln M / n) term in the gap — the price of choosing.
The trouble: M is usually infinite. Sauer's lemma rescues it — on n points, H acts like only Πₕ(n) ≤ (en/VC)^{VC} effective hypotheses, so ln M becomes ≈ VC·ln n. That substitution is where the VC term is born.
Concept
Plugging Sauer's polynomial growth into a Hoeffding-style union argument yields the headline result. With probability ≥ 1 − δ, for every hypothesis in H at once:
\[ \underbrace{\text{err}_{\text{true}}}_{\text{gen}} \;\le\; \underbrace{\text{err}_{\text{train}}}_{\text{fit}} \;+\; O\!\left(\sqrt{\frac{\text{VC}\cdot \ln n}{n}}\right) \]
The gap term shrinks with data n and grows with capacity VC. It is the rigorous form of 'enough data beats a complex model' — and the reason Hoeffding was worth proving: it is the engine inside the square root.
Explain it
Discussion prompt
Explain The VC generalization bound to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Plugging Sauer's polynomial growth into a Hoeffding-style union argument yields the headline result. With probability ≥ 1 − δ, for every hypothesis in H at once:
Missing information
Discussion prompt
Exam-style: n = 1000 examples, a model of VC dimension 50. Evaluate the gap term √(VC·ln n / n), and sweep to see how data and capacity move it. Runnable:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
The bound is pessimistic (0.59 is a huge gap), yet increasing n 100× shrinks it to 0.076 — data closes the gap. Raising VC to 1000 blows it past 1 (vacuous), showing capacity's cost.
Worked example
Exam-style: n = 1000 examples, a model of VC dimension 50. Evaluate the gap term √(VC·ln n / n), and sweep to see how data and capacity move it. Runnable:
import numpy as np
for vc, n in [(3, 100), (50, 1000), (50, 100_000), (1000, 1000)]:
term = np.sqrt(vc*np.log(n)/n)
print(f'VC={vc:5d}, n={n:6d}: sqrt(VC ln n / n) = {term:.4f}')VC=50, n=1000 → 0.59; hold VC, grow n to 100k → 0.076
Why: The bound is pessimistic (0.59 is a huge gap), yet increasing n 100× shrinks it to 0.076 — data closes the gap. Raising VC to 1000 blows it past 1 (vacuous), showing capacity's cost.
| VC | n | √(VC·ln n / n) (verified) |
|---|---|---|
| 3 | 100 | 0.3717 |
| 50 | 1000 | 0.5877 |
| 50 | 100000 | 0.0759 |
| 1000 | 1000 | 2.6283 (vacuous) |
Ranking
Put in order
These are the steps of The concentration → generalization pipeline, scrambled. Put them back in order before the next slide shows you.
X̄ (train err) concentrates near E[X] (true err) by Hoeffding 2e^{−2nt²}ℝᵈ is d+12ᵐ labelings into a polynomial (em/d)ᵈ once m > VC≤ √(VC·ln n / n) — shrinks with n, grows with VCn: invert to n ≥ ln(2/δ)/(2ε²) per hypothesis; capacity multiplies itWhy: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
X̄ (train err) concentrates near E[X] (true err) by Hoeffding 2e^{−2nt²}ℝᵈ is d+12ᵐ labelings into a polynomial (em/d)ᵈ once m > VC≤ √(VC·ln n / n) — shrinks with n, grows with VCn: invert to n ≥ ln(2/δ)/(2ε²) per hypothesis; capacity multiplies itEdge cases
Discussion prompt
The concentration → generalization pipeline works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
X̄ (train err) concentrates near E[X] (true err) by Hoeffding 2e^{−2nt²}ℝᵈ is d+12ᵐ labelings into a polynomial (em/d)ᵈ once m > VC≤ √(VC·ln n / n) — shrinks with n, grows with VCn: invert to n ≥ ln(2/δ)/(2ε²) per hypothesis; capacity multiplies itElimination
Eliminate the wrong options
Which inequality needs ONLY that X ≥ 0 and E[X] exists — no variance, no boundedness?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Markov's derivation used only a·1[X≥a] ≤ X, which needs X ≥ 0 and a finite mean — nothing else. That minimal assumption set is exactly why it is the loosest of the three and serves as the lemma for the others.
Check
Which bound survives on the least information?
Check your understanding
Which inequality needs ONLY that X ≥ 0 and E[X] exists — no variance, no boundedness?
Answer: A
Why: Markov's derivation used only a·1[X≥a] ≤ X, which needs X ≥ 0 and a finite mean — nothing else. That minimal assumption set is exactly why it is the loosest of the three and serves as the lemma for the others.
Prediction
Predict first
Chebyshev's P(|X−μ| ≥ kσ) ≤ 1/k² is obtained by applying Markov to which non-negative variable?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: (X − μ)², at level (kσ)²
Why: Squaring the deviation makes it non-negative (so Markov applies) and gives it mean E[(X−μ)²] = σ². Markov at level (kσ)² then yields σ²/(k²σ²) = 1/k² after the variance cancels.
Check
Recall how we got a two-sided bound from a one-sided lemma.
Check your understanding
Chebyshev's P(|X−μ| ≥ kσ) ≤ 1/k² is obtained by applying Markov to which non-negative variable?
Answer: A
Why: Squaring the deviation makes it non-negative (so Markov applies) and gives it mean E[(X−μ)²] = σ². Markov at level (kσ)² then yields σ²/(k²σ²) = 1/k² after the variance cancels.
Prediction
Predict first
The VC dimension of a linear classifier in ℝ² (a general line with offset) is:
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: 3 — it shatters some 3 points but no 4 (XOR fails)
Why: VC = d + 1 = 3. Our brute force realized all 8 labelings of a triangle (shatters 3) but capped at 0.75 accuracy on the 4-point XOR (cannot shatter 4). The largest shattered set has size 3.
Check
Count degrees of freedom, and remember the offset.
Check your understanding
The VC dimension of a linear classifier in ℝ² (a general line with offset) is:
Answer: A
Why: VC = d + 1 = 3. Our brute force realized all 8 labelings of a triangle (shatters 3) but capped at 0.75 accuracy on the 4-point XOR (cannot shatter 4). The largest shattered set has size 3.
Elimination
Eliminate the wrong options
By the VC bound gap ≈ √(VC·ln n / n), the train–true gap shrinks when:
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: VC sits in the numerator and n dominates the denominator, so more data shrinks the gap and more capacity widens it. Our sweep confirmed it: VC=50 gap fell from 0.59 at n=1000 to 0.076 at n=100000.
Check
Read the square root: what is in the numerator, what is in the denominator?
Check your understanding
By the VC bound gap ≈ √(VC·ln n / n), the train–true gap shrinks when:
Answer: A
Why: VC sits in the numerator and n dominates the denominator, so more data shrinks the gap and more capacity widens it. Our sweep confirmed it: VC=50 gap fell from 0.59 at n=1000 to 0.076 at n=100000.
Concept
Modern neural nets have VC dimension far larger than n — plug that in and the bound is > 1, i.e. vacuous. Classical theory predicts disaster. Yet these overparameterized nets generalize beautifully. Something is missing from the worst-case picture.
The empirical resolution is double descent: as capacity grows past the point where the model exactly fits the training data (the interpolation threshold), test error rises to a spike — then falls again. We can watch it happen in code.
Estimation
Predict first
Fit a least-norm linear model on random tanh features, sweeping the feature count p past the sample size n = 40. Watch test MSE spike at p ≈ n then descend. Runnable as written:
Commit before you compute: what does Double descent, from a real fit come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Test MSE dives, spikes near p ≈ n = 40, then descends again
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. At p = 40 the model exactly interpolates the 40 training points and test error explodes (~38); push p to 400 and the least-norm solution becomes smooth again, MSE back near 0.75.
Worked example
Fit a least-norm linear model on random tanh features, sweeping the feature count p past the sample size n = 40. Watch test MSE spike at p ≈ n then descend. Runnable as written:
import numpy as np
rng = np.random.default_rng(0)
n, d, sigma, ntest = 40, 5, 0.3, 2000
beta = rng.standard_normal(d)
Xtr = rng.standard_normal((n, d)); ytr = Xtr @ beta + sigma*rng.standard_normal(n)
Xte = rng.standard_normal((ntest, d)); yte = Xte @ beta + sigma*rng.standard_normal(ntest)
for p in [5, 20, 40, 60, 400]:
W = rng.standard_normal((d, p)) / np.sqrt(d)
Ptr, Pte = np.tanh(Xtr @ W), np.tanh(Xte @ W)
w = np.linalg.pinv(Ptr) @ ytr # least-norm fit
print(f'p={p:3d}: test MSE = {np.mean((Pte @ w - yte)**2):.3f}')Test MSE dives, spikes near p ≈ n = 40, then descends again
Why: At p = 40 the model exactly interpolates the 40 training points and test error explodes (~38); push p to 400 and the least-norm solution becomes smooth again, MSE back near 0.75. Classical VC sees only the rising middle.
| p (features) | test MSE (verified) |
|---|---|
| 5 (under-param) | 0.197 |
| 20 | 0.601 |
| 40 (≈ interpolation) | 37.930 (spike) |
| 60 | 4.843 |
| 400 (over-param) | 0.746 |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Test MSE dives, spikes near p ≈ n = 40, then descends again
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Fit a least-norm linear model on random tanh features, sweeping the feature count p past the sample size n = 40. Watch test MSE spike at p ≈ n then descend. Runnable as written:
Anomaly
Predict first
A student writes this, and it looks reasonable:
The p = 40 fit had test MSE 37.9, so adding still more features must push test error even higher.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: That reads only the FIRST descent and the spike.
Recognize the second descent: past the interpolation threshold, more capacity plus an implicit smoothness bias lowers test error again.
Why: That reads only the FIRST descent and the spike. Our own sweep refutes it: at p = 60 the MSE is already back down to 4.84, and at p = 400 it is 0.75 — an order of magnitude BELOW the p = 40 spike. Test error is non-monotonic in capacity.
Trap
The p = 40 fit had test MSE 37.9, so adding still more features must push test error even higher.
Extrapolate the rising curve past p = n
Why: That reads only the FIRST descent and the spike. Our own sweep refutes it: at p = 60 the MSE is already back down to 4.84, and at p = 400 it is 0.75 — an order of magnitude BELOW the p = 40 spike. Test error is non-monotonic in capacity.
Recognize the second descent: past the interpolation threshold, more capacity plus an implicit smoothness bias lowers test error again.
Read the whole double-descent curve
Why: Error falls (under-param), spikes at p ≈ n, then falls again (over-param). The least-norm solution among many interpolators is the smoothest one, so huge p is safe — the regime modern nets actually live in.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
X̄ is our estimate of p, and in ML it plays the role of training error estimating true error. The whole lesson asks: how far can X̄ stray from p?; Check both cases: if X ≥ a the left side is a and a ≤ X; if X < a the left side is 0 ≤ X (since X ≥ 0). Either way it holds. Now just take expectations.; We want a two-sided bound on |X − μ|. The variable |X − μ| can be positive from either side, so we square it: (X − μ)² is non-negative — precisely what Markov needs.P(X ≥ 4) ≤ 0.25, so roughly a quarter of the mass sits above 4.; To bound P(|X−μ| ≥ kσ), apply Markov directly to |X − μ|: P(|X−μ| ≥ kσ) ≤ E|X−μ| / (kσ).Intuition
Right at p ≈ n the model is forced to thread every noisy point with essentially one solution — a wild, high-curvature fit. That is the spike.
Past that, many solutions interpolate the data, and the least-norm / gradient-descent one picks the smoothest among them. That implicit bias toward simple functions — not raw parameter count — is what VC's worst-case bound cannot see.
So VC bounds are still true (they are worst-case guarantees), just not tight for the specific, well-regularized functions modern training actually finds.
Analogy
Discussion prompt
Explain Why over-parameterization can help by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Right at p ≈ n the model is forced to thread every noisy point with essentially one solution — a wild, high-curvature fit. That is the spike.
Section
The project
Concept
Rebuild the lesson's four pillars yourself: two empirical tail checks, one sample-complexity inversion, and one capacity computation. You've derived every piece — now assemble it.
| # | milestone | tool |
|---|---|---|
| 1 | Markov: empirical P(X≥4) ≤ E[X]/4 | rng.exponential |
| 2 | Chebyshev: empirical P(|X−1|≥2) ≤ 1/4 | np.abs, .mean() |
| 3 | Hoeffding: bound + invert for n | np.exp, np.log |
| 4 | VC gap term √(VC·ln n / n) | np.sqrt |
Build rules: type every line, draw ≥ 10⁶ samples for stable tail estimates, and confirm each empirical probability sits under its bound before moving on.
Counterexample
Discussion prompt
Rebuild the lesson's four pillars yourself: two empirical tail checks, one sample-complexity inversion, and one capacity computation. You've derived every piece — now assemble it.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Build rules: type every line, draw ≥ 10⁶ samples for stable tail estimates, and confirm each empirical probability sits under its bound before moving on.
Worked example
Your turn: sample Exponential(1) and check P(X ≥ 4) against the Markov ceiling E[X]/4 = 1/4. Predict which is larger before you run it.
Hint: (X >= 4).mean() is the empirical tail; the bound is 1/4. They should straddle — empirical well below the bound.
import numpy as np
rng = np.random.default_rng(0)
X = rng.exponential(1.0, 2_000_000)
print('empirical P(X>=4):', round((X >= 4).mean(), 4))
print('Markov bound 1/4 :', 1/4)| quantity | expected value |
|---|---|
| empirical P(X ≥ 4) | 0.0183 |
| Markov bound 1/4 | 0.25 |
| holds (empirical ≤ bound)? | yes |
Worked example
Your turn: on the same Exponential(1) (μ = σ = 1), check the two-sided tail P(|X − 1| ≥ 2) against 1/k² = 1/4 at k = 2. Predict: under the bound?
Hint: (np.abs(X - 1) >= 2).mean() versus 1 / 2**2. Reuse the X you already drew.
import numpy as np
rng = np.random.default_rng(0)
X = rng.exponential(1.0, 2_000_000)
print('empirical P(|X-1|>=2):', round((np.abs(X - 1) >= 2).mean(), 4))
print('Chebyshev 1/k^2 :', 1/2**2)| quantity | expected value |
|---|---|
| empirical P(|X − 1| ≥ 2) | 0.0498 |
| Chebyshev bound 1/4 | 0.25 |
| holds? | yes |
Worked example
Your turn: compute the Hoeffding bound at n = 1000, t = 0.05, then invert it to find the n needed for ε = 0.05, δ = 0.05. Predict whether the inverted n is in the hundreds or thousands.
Hint: bound = 2*np.exp(-2*n*t**2); inverse n = ln(2/δ)/(2ε²), rounded up with np.ceil.
import numpy as np
print('Hoeffding n=1000,t=0.05:', round(2*np.exp(-2*1000*0.05**2), 4))
eps, delta = 0.05, 0.05
n = int(np.ceil(np.log(2/delta) / (2*eps**2)))
print('samples needed n :', n)
print('check bound at n :', round(2*np.exp(-2*n*eps**2), 4))| quantity | expected value |
|---|---|
| Hoeffding (n=1000, t=0.05) | 0.0135 |
| n = ⌈ln(2/δ)/(2ε²)⌉ | 738 |
| bound at n = 738 | 0.0499 (≤ 0.05 ✓) |
Worked example
Your turn: compute the gap term √(VC·ln n / n) for VC = 50, n = 1000, then again at n = 100000 to watch data close the gap. Predict the direction of change.
Hint: np.sqrt(vc*np.log(n)/n). Growing n should shrink the term.
import numpy as np
vc = 50
for n in [1000, 100_000]:
print(f'n={n:6d}: gap term = {np.sqrt(vc*np.log(n)/n):.4f}')| VC | n | gap term (expected) |
|---|---|---|
| 50 | 1000 | 0.5877 |
| 50 | 100000 | 0.0759 |
| direction | n↑ | gap ↓ |
Trade off
Comparison matrix
From Milestone 4 — the VC gap term: every row here is a choice with a cost. Fill the n column, then say which row you would actually pick and what you give up for it.
| VC | n | gap term (expected) |
|---|---|---|
| 50 | 1000 | 0.5877 |
| 50 | 100000 | 0.0759 |
| direction | n↑ | gap ↓ |
Concept
import numpy as np
rng = np.random.default_rng(0)
X = rng.exponential(1.0, 2_000_000) # E[X]=1, sigma=1
print('Markov P(X>=4) :', round((X>=4).mean(),4), '<=', 1/4)
print('Chebyshev P(|X-1|>=2):', round((np.abs(X-1)>=2).mean(),4), '<=', 1/4)
print('Hoeffding (n=1000,t=0.05):', round(2*np.exp(-2*1000*0.05**2),4))
print('VC gap (VC=50,n=1000) :', round(np.sqrt(50*np.log(1000)/1000),4))| printed line | value (verified) |
|---|---|
| Markov P(X≥4) | 0.0183 ≤ 0.25 |
| Chebyshev P(|X−1|≥2) | 0.0498 ≤ 0.25 |
| Hoeffding (n=1000, t=0.05) | 0.0135 |
| VC gap (VC=50, n=1000) | 0.5877 |
If every empirical tail sits under its bound, Hoeffding reads 0.0135, and the VC gap reads 0.5877 — you can reason rigorously about when a model generalizes.
Comparison
Comparison matrix
From The full program: refill the value (verified) column from what you know. The rest of the table is as it appeared.
| printed line | value (verified) |
|---|---|
| Markov P(X≥4) | 0.0183 ≤ 0.25 |
| Chebyshev P(|X−1|≥2) | 0.0498 ≤ 0.25 |
| Hoeffding (n=1000, t=0.05) | 0.0135 |
| VC gap (VC=50, n=1000) | 0.5877 |
Concept
Slides closed, out loud: explain (1) the one indicator inequality that proves Markov, (2) why squaring the deviation turns Markov into Chebyshev, and (3) what 'shatter' means and why a line stops at 3 points.
Stretch: use Hoeffding to bound how far a fixed classifier's training error can drift from its true error, then re-run the double-descent sweep at p = 41 and p = 120 and confirm the spike is real — the exact worst-case-vs-practice gap the olympiad likes to probe.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — The question generalization answers · Markov — from one indicator · Chebyshev — Markov on (X−μ)² · Hoeffding — the exponential engine · VC dimension — measuring capacity · The generalization bound. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
a·1[X≥a] ≤ X and know it needs only X ≥ 0 and a mean1/k² by applying Markov to (X−μ)², and Hoeffding 2e^{−2nt²} for bounded i.i.d. averagesn ≥ ln(2/δ)/(2ε²) — here 738 flips for ε = δ = 0.053 points, fails XOR, so VC = d+1√(VC·ln n / n), know it is loose, and explain double descent past the interpolation threshold| idea | the one thing to remember |
|---|---|
| Markov | P(X≥a) ≤ E[X]/a — weakest, from one indicator |
| Chebyshev | Markov on (X−μ)² ⇒ 1/k² |
| Hoeffding | 2e^{−2nt²} — exponential in n; invert for n |
| VC (line in ℝᵈ) | d + 1; largest shattered set |
| gen bound | √(VC·ln n / n) — loose; double descent beats it |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.