Lesson 34: Concentration Inequalities & Generalization

USAAIO Lesson 34, from Week 12 on generalization theory, fully worked. It derives Markov's inequality from the layer-cake, or indicator, trick with no steps skipped, obtains Chebyshev's by applying Markov to the squared deviation, and states Hoeffding's, proving its exponential decay in n empirically on a biased coin. It then solves the sample-complexity inversion n >= ln(2/delta)/(2 eps^2) by hand. From there it builds up shattering and VC dimension, from a one-dimensional threshold to a two-dimensional line, which shatters 3 points but fails XOR, managing at best 3 of 4. It covers Sauer's lemma and the VC generalization bound term by term, and demonstrates double descent with a real least-norm fit whose test error peaks at the interpolation threshold. One running example, a coin with p=0.3, threads the whole deck, and every inequality and every printed number was produced by real execution. The lesson runs to 62 slides.

Subject: Machine Learning · 115 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Concentration & Generalization

Title

USAAIO · Lesson 34 · Week 12 (Generalization Theory)

Why a model that fits the sample should work on new data. We derive Markov from one indicator, get Chebyshev for free, prove Hoeffding decays exponentially, build VC dimension from shattering, read the generalization bound, and watch double descent break the classical story — every step shown, every number run.

2. By the end of this lesson you can

Objectives

  1. Derive Markov P(X ≥ a) ≤ E[X]/a from the indicator inequality a·1[X≥a] ≤ X, and know exactly which assumption it uses
  2. Get Chebyshev P(|X−μ| ≥ kσ) ≤ 1/k² by applying Markov to the squared deviation (X−μ)²
  3. State Hoeffding P(|X̄−E[X]| ≥ t) ≤ 2e^{−2nt²} and invert it for the sample size n ≥ ln(2/δ)/(2ε²)
  4. Define shattering and VC dimension, and show a line shatters 3 points but fails XOR (VC = d+1)
  5. Read the VC generalization bound gen ≤ train + √(VC·ln n / n) and explain why double descent beats it in practice

3. What survived from Gradient Checking & Debugging?

Warm-up

Discussion prompt

Before we open Lesson 34: Concentration Inequalities & Generalization: without looking back, what was the main idea of Gradient Checking & Debugging, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

central finite differences and their O(ε²) accuracy, gradient checking by relative error, the common backprop bugs (broadcasting, wrong axis, missing zero_grad, sign), vanishing/exploding-gradient diagnosis, and NaN debugging. Build gradient_check and use it to catch a deliberately buggy gradient.

4. The question generalization answers

Section

Part 1 of 6

5. The running example: a biased coin

Concept

One example threads the whole lesson. We flip a coin whose true heads-probability is p = 0.3 (we call heads 1, tails 0). We never see p; we only see flips and form the sample mean X̄.

\[ \bar X = \frac{1}{n}\sum_{i=1}^{n} X_i, \qquad X_i \in \{0,1\}, \qquad \mathbb{E}[X_i] = p = 0.3 \]

X̄ is our estimate of p, and in ML it plays the role of training error estimating true error. The whole lesson asks: how far can X̄ stray from p?

6. Break it if you can: The running example: a biased coin

Counterexample

Discussion prompt

One example threads the whole lesson. We flip a coin whose true heads-probability is p = 0.3 (we call heads 1, tails 0). We never see p; we only see flips and form the sample mean X̄.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

X̄ is our estimate of p, and in ML it plays the role of training error estimating true error. The whole lesson asks: how far can X̄ stray from p?

7. Guess the shape of the answer: One sample already strays

Estimation

Predict first

Flip the p = 0.3 coin 20 times with a fixed seed and read off X̄. Watch it miss 0.3. Runnable as written:

Commit before you compute: what does One sample already strays come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: 8 heads in 20 flips → X̄ = 0.40, not 0.30

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. A finite sample is noisy: the estimate landed 0.10 above the truth.

8. One sample already strays

Worked example

Flip the p = 0.3 coin 20 times with a fixed seed and read off X̄. Watch it miss 0.3. Runnable as written:

import numpy as np
rng = np.random.default_rng(0)
p = 0.3
flips = rng.binomial(1, p, size=20)   # 20 coin flips, heads = 1
print('flips:', flips)
print('sum  :', flips.sum())
print('Xbar :', flips.mean())

8 heads in 20 flips → X̄ = 0.40, not 0.30

Why: A finite sample is noisy: the estimate landed 0.10 above the truth. Concentration inequalities put a ceiling on how often — and how far — that happens.

quantityvalue (verified)
flips (seed 0)[0 0 0 0 1 1 0 1 0 1 1 0 1 0 1 0 1 0 0 0]
sum of heads8
X̄ = sum / 200.40
true p0.30

9. Fill in: value (verified) for One sample already strays

Comparison

Comparison matrix

From One sample already strays: refill the value (verified) column from what you know. The rest of the table is as it appeared.

quantityvalue (verified)
flips (seed 0)[0 0 0 0 1 1 0 1 0 1 1 0 1 0 1 0 1 0 0 0]
sum of heads8
X̄ = sum / 200.40
true p0.30

10. Why this is the whole game

Intuition

A learner picks the model that looks best on the training sample. If sample statistics could wander arbitrarily far from the truth, a low training error would tell you nothing about future data.

Generalization theory earns its keep by proving the opposite: with enough data, X̄ is pinned near p — so training error is pinned near true error. 'Pinned' is exactly what a concentration inequality guarantees.

We build three of them, weakest to strongest: Markov (mean only), Chebyshev (adds variance), Hoeffding (adds boundedness). Each new assumption buys a tighter bound.

11. By analogy: Why this is the whole game

Analogy

Discussion prompt

Explain Why this is the whole game by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

A learner picks the model that looks best on the training sample. If sample statistics could wander arbitrarily far from the truth, a low training error would tell you nothing about future data.

12. Markov — from one indicator

Section

Part 2 of 6 — the weakest bound

13. The one inequality behind Markov

Concept

Let X ≥ 0 and pick a level a > 0. Define the indicator 1[X ≥ a], which is 1 when the event happens and 0 otherwise. The entire proof rests on one pointwise inequality:

\[ a \cdot \mathbf{1}[X \ge a] \;\le\; X \qquad \text{(true for every outcome)} \]

Check both cases: if X ≥ a the left side is a and a ≤ X; if X < a the left side is 0 ≤ X (since X ≥ 0). Either way it holds. Now just take expectations.

14. Teach it back: The one inequality behind Markov

Explain it

Discussion prompt

Explain The one inequality behind Markov to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Let X ≥ 0 and pick a level a > 0. Define the indicator 1[X ≥ a], which is 1 when the event happens and 0 otherwise. The entire proof rests on one pointwise inequality:

15. What has to happen first: Derive Markov, step by step

Ranking

Put in order

Put the moves of Derive Markov, step by step into the order they have to happen.

  1. Start from the pointwise inequality
  2. Take expectations of both sides
  3. Pull the constant a out; E of an indicator is a probability
  4. Divide both sides by a > 0

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Holds for every outcome, as we checked case by case.

16. Derive Markov, step by step

Worked example

Start from the pointwise inequality

Why: Holds for every outcome, as we checked case by case.

\[ a \cdot \mathbf{1}[X \ge a] \le X \]

Take expectations of both sides

Why: Expectation is monotone: if U ≤ V pointwise then E[U] ≤ E[V]. The inequality survives.

\[ \mathbb{E}\big[a \cdot \mathbf{1}[X \ge a]\big] \le \mathbb{E}[X] \]

Pull the constant a out; E of an indicator is a probability

Why: E[1[A]] = P(A) by definition, so E[a·1[X≥a]] = a·P(X≥a).

\[ a \, P(X \ge a) \le \mathbb{E}[X] \]

Divide both sides by a > 0

Why: Dividing by a positive number keeps the direction. This IS Markov's inequality.

\[ \boxed{\,P(X \ge a) \le \dfrac{\mathbb{E}[X]}{a}\,} \]

17. Decode the notation: Derive Markov, step by step

Notation

Annotate

From Derive Markov, step by step — read this one piece at a time. What is each part doing?

On: \( a \cdot \mathbf{1}[X \ge a] \le X \)

  • Holds for every outcome, as we checked case by case.
  • Expectation is monotone: if U ≤ V pointwise then E[U] ≤ E[V]. The inequality survives.
  • E[1[A]] = P(A) by definition, so E[a·1[X≥a]] = a·P(X≥a).

18. What Markov assumes — and doesn't

Concept

The only ingredients were X ≥ 0 and that E[X] exists. No variance, no distribution shape, no boundedness. That universality is Markov's strength and its weakness: it must hold for every non-negative variable with that mean, so it can't be tight for any particular one.

one-sided vs two-sided — Markov bounds only the RIGHT tail P(X ≥ a) of a non-negative variable. To bound how far X falls on BOTH sides of its mean, we need the deviation itself — that is the job of Chebyshev.

19. What has to be given first: Markov holds — and how loosely

Missing information

Discussion prompt

Take X ~ Exponential(1), so E[X] = 1. Draw two million samples and compare the empirical right-tail P(X ≥ a) to the Markov ceiling 1/a. Complete snippet:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

At a=4 the truth is 0.0183 but Markov only promises ≤ 0.25 — a 14× gap. Valid, but loose, exactly because it ignores the exponential's shape.

20. Markov holds — and how loosely

Worked example

Take X ~ Exponential(1), so E[X] = 1. Draw two million samples and compare the empirical right-tail P(X ≥ a) to the Markov ceiling 1/a. Complete snippet:

import numpy as np
rng = np.random.default_rng(0)
X = rng.exponential(1.0, 2_000_000)   # E[X] = 1
for a in [2, 4, 5]:
    emp = (X >= a).mean()
    print(f'a={a}: P(X>={a})={emp:.4f}  Markov 1/a={1/a:.4f}')

Every empirical tail sits far under 1/a

Why: At a=4 the truth is 0.0183 but Markov only promises ≤ 0.25 — a 14× gap. Valid, but loose, exactly because it ignores the exponential's shape.

aempirical P(X ≥ a)Markov bound 1/a
20.13520.5000
40.01830.2500
50.00670.2000

21. What each one costs: Markov holds — and how loosely

Trade off

Comparison matrix

From Markov holds — and how loosely: every row here is a choice with a cost. Fill the empirical P(X ≥ a) column, then say which row you would actually pick and what you give up for it.

aempirical P(X ≥ a)Markov bound 1/a
20.13520.5000
40.01830.2500
50.00670.2000

22. Something is wrong here: reading a bound as an estimate

Anomaly

Predict first

A student writes this, and it looks reasonable:

Markov gives P(X ≥ 4) ≤ 0.25, so roughly a quarter of the mass sits above 4.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: An upper bound is a CEILING, not an estimate.

Read 0.25 as a worst-case ceiling. To get near the true 0.018, feed the bound more structure.

Why: An upper bound is a CEILING, not an estimate. The true value is 0.0183 — about 14× smaller. Markov used only the mean, so it is deliberately pessimistic.

23. Trap: reading a bound as an estimate

Trap

The trap

Markov gives P(X ≥ 4) ≤ 0.25, so roughly a quarter of the mass sits above 4.

Treat 0.25 as ≈ the probability

Why: An upper bound is a CEILING, not an estimate. The true value is 0.0183 — about 14× smaller. Markov used only the mean, so it is deliberately pessimistic.

The fix

Read 0.25 as a worst-case ceiling. To get near the true 0.018, feed the bound more structure.

Add variance → Chebyshev; add boundedness → Hoeffding

Why: Each extra assumption tightens the bound. Markov's job is not to be tight — it is to hold when you know almost nothing, and to be the lemma the other bounds are built from.

24. Chebyshev — Markov on (X−μ)²

Section

Part 3 of 6 — variance buys two sides

25. The trick: square the deviation

Concept

We want a two-sided bound on |X − μ|. The variable |X − μ| can be positive from either side, so we square it: (X − μ)² is non-negative — precisely what Markov needs.

And its mean is famous: E[(X − μ)²] = Var(X) = σ². So apply Markov to Y = (X − μ)² at level a = (kσ)². The next slide does it move by move.

26. Complete the line: Derive Chebyshev from Markov

Fill the middle

Fill in the blanks

From Derive Chebyshev from Markov — finish the line. Write what belongs on the right of the equals sign before you look.

P(|X-\mu| \ge k\sigma) = P\big((X-\mu)^2 \ge (k\sigma)^2\big)

Why: Producing the right-hand side unprompted is the difference between recognising this line and being able to use it. |X − μ| ≥ kσ happens for exactly the same outcomes as (X − μ)² ≥ (kσ)² — squaring both non-negative sides preserves the event.

27. Derive Chebyshev from Markov

Worked example

The events are identical

Why: |X − μ| ≥ kσ happens for exactly the same outcomes as (X − μ)² ≥ (kσ)² — squaring both non-negative sides preserves the event.

\[ P(|X-\mu| \ge k\sigma) = P\big((X-\mu)^2 \ge (k\sigma)^2\big) \]

Apply Markov to Y = (X − μ)² at level (kσ)²

Why: Y ≥ 0, so Markov gives P(Y ≥ a) ≤ E[Y]/a with E[Y] = σ² and a = k²σ².

\[ P\big((X-\mu)^2 \ge k^2\sigma^2\big) \le \frac{\mathbb{E}[(X-\mu)^2]}{k^2\sigma^2} = \frac{\sigma^2}{k^2\sigma^2} \]

Cancel σ²

Why: The variance cancels top and bottom, leaving a bound that depends only on k. This is Chebyshev.

\[ \boxed{\,P(|X-\mu| \ge k\sigma) \le \dfrac{1}{k^2}\,} \]

28. Say it in words: Derive Chebyshev from Markov

Translation

\( P\big((X-\mu)^2 \ge k^2\sigma^2\big) \le \frac{\mathbb{E}[(X-\mu)^2]}{k^2\sigma^2} = \frac{\sigma^2}{k^2\sigma^2} \)

Draw it

Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.

29. Predict the next row: Verify Chebyshev really is Markov-on-Y

Pattern

Predict first

The table runs: P(|X − 1| ≥ 2) | 0.0498 · P(Y ≥ 4), Y = (X−1)² | 0.0498 · E[Y] (= Var = σ²) | 1.0

In Verify Chebyshev really is Markov-on-Y, given the rows so far: what is the next one — the row where quantity is Markov E[Y]/4 = 1/k²?

Correct: Markov E[Y]/4 = 1/k² | 0.25

quantityvalue (verified)
P(|X − 1| ≥ 2)0.0498
P(Y ≥ 4), Y = (X−1)²0.0498
E[Y] (= Var = σ²)1.0
Markov E[Y]/4 = 1/k²0.25

Why: The relationship between the columns, not the individual numbers, is what generates the next row. Squaring the deviation maps the two-sided event onto a one-sided event of Y, exactly as claimed, and E[Y] recovers the variance 1.

30. Verify Chebyshev really is Markov-on-Y

Worked example

Let's confirm the derivation numerically on Exponential(1) (μ = σ = 1) at k = 2. We check that P(|X−1| ≥ 2) equals P(Y ≥ 4) and that E[Y]/4 = 1/4:

import numpy as np
rng = np.random.default_rng(0)
X = rng.exponential(1.0, 2_000_000)   # mu = sigma = 1
Y = (X - 1.0)**2
print('P(|X-1|>=2)  :', round((np.abs(X-1)>=2).mean(), 4))
print('P(Y>=4)      :', round((Y >= 4).mean(), 4))
print('E[Y]         :', round(Y.mean(), 4))
print('Markov E[Y]/4:', round(Y.mean()/4, 4))

The two events give the same 0.0498; E[Y] ≈ 1 = σ²

Why: Squaring the deviation maps the two-sided event onto a one-sided event of Y, exactly as claimed, and E[Y] recovers the variance 1.

quantityvalue (verified)
P(|X − 1| ≥ 2)0.0498
P(Y ≥ 4), Y = (X−1)²0.0498
E[Y] (= Var = σ²)1.0
Markov E[Y]/4 = 1/k²0.25

31. Which is which, by value (verified)

Discrimination

Sort into buckets

Sort these by value (verified), from memory, without looking back at Verify Chebyshev really is Markov-on-Y. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

0.0498
P(|X − 1| ≥ 2); P(Y ≥ 4), Y = (X−1)²
1.0
E[Y] (= Var = σ²)
0.25
Markov E[Y]/4 = 1/k²
g1
value (verified) is "0.0498" for P(|X − 1| ≥ 2), P(Y ≥ 4), Y = (X−1)² — that is what the table on "Verify Chebyshev really is Markov-on-Y" records, and it is the single property separating this group from the rest.
g2
value (verified) is "1.0" for E[Y] (= Var = σ²) — that is what the table on "Verify Chebyshev really is Markov-on-Y" records, and it is the single property separating this group from the rest.
g3
value (verified) is "0.25" for Markov E[Y]/4 = 1/k² — that is what the table on "Verify Chebyshev really is Markov-on-Y" records, and it is the single property separating this group from the rest.

32. Restore the missing line: Chebyshev across several k

Fill the middle

Fill in the blanks

From Chebyshev across several k — one line has had its right-hand side removed. Put it back.

import numpy as np
rng = np.random.default_rng(0)
X = rng.exponential(1.0, 2_000_000) # mu = 1, sigma = 1
for k in [2, 3, 4]:
emp = (np.abs(X - 1.0) >= k).mean()
print(f'k=___: P(|X-1|>=___)=___ 1/k^2=___')

Why: X is what everything below it consumes, so the wrong expression here fails later and somewhere else. Chebyshev is tighter than Markov but still worst-case over all distributions with σ = 1, so a specific one beats it comfortably.

33. Chebyshev across several k

Worked example

Sweep k = 2, 3, 4 on the same Exponential(1). Each empirical two-sided tail must sit under 1/k². Runnable:

import numpy as np
rng = np.random.default_rng(0)
X = rng.exponential(1.0, 2_000_000)   # mu = 1, sigma = 1
for k in [2, 3, 4]:
    emp = (np.abs(X - 1.0) >= k).mean()
    print(f'k={k}: P(|X-1|>={k})={emp:.4f}  1/k^2={1/k**2:.4f}')

All three hold; the exponential's fat right tail keeps them loose

Why: Chebyshev is tighter than Markov but still worst-case over all distributions with σ = 1, so a specific one beats it comfortably.

kempirical P(|X−1| ≥ k)Chebyshev 1/k²
20.04980.2500
30.01830.1111
40.00670.0625

34. Watch it run: Chebyshev across several k

Pattern

Step through it

Step through Chebyshev across several k one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: k is 2
  2. Step 2: k is 3
  3. Step 3: k is 4

35. Something is wrong here: forgetting to square the deviation

Anomaly

Predict first

A student writes this, and it looks reasonable:

To bound P(|X−μ| ≥ kσ), apply Markov directly to |X − μ|: P(|X−μ| ≥ kσ) ≤ E|X−μ| / (kσ).

It is wrong. Say what breaks — and say it before you turn the page.

Correct: This is a valid bound, but it uses the mean absolute deviation E|X−μ|, not the variance — it does NOT give the clean 1/k².

Apply Markov to the squared deviation (X − μ)², whose mean is exactly σ².

Why: This is a valid bound, but it uses the mean absolute deviation E|X−μ|, not the variance — it does NOT give the clean 1/k². You've thrown away the σ² cancellation that makes Chebyshev memorable and distribution-free in k.

36. Trap: forgetting to square the deviation

Trap

The trap

To bound P(|X−μ| ≥ kσ), apply Markov directly to |X − μ|: P(|X−μ| ≥ kσ) ≤ E|X−μ| / (kσ).

Stop at E|X − μ| / (kσ)

Why: This is a valid bound, but it uses the mean absolute deviation E|X−μ|, not the variance — it does NOT give the clean 1/k². You've thrown away the σ² cancellation that makes Chebyshev memorable and distribution-free in k.

The fix

Apply Markov to the squared deviation (X − μ)², whose mean is exactly σ².

P((X−μ)² ≥ k²σ²) ≤ σ²/(k²σ²) = 1/k²

Why: Squaring makes the mean equal Var, and the σ² cancels — that cancellation is the whole point. The '²' in Chebyshev is not optional.

37. Hoeffding — the exponential engine

Section

Part 4 of 6 — averages of bounded variables

38. Why we need a third bound

Concept

For a single variable, Chebyshev gives P(|X−μ| ≥ t) ≤ σ²/t² — decay only like 1/t². But we average n flips, and averaging shrinks variance: Var(X̄) = σ²/n. Feeding that into Chebyshev already gives σ²/(n t²), decaying like 1/n.

Hoeffding does far better. For an average of n bounded i.i.d. variables the tail decays exponentially in n — the difference between 'need thousands of samples' and 'need a few hundred'.

39. Hoeffding's inequality

Concept

For i.i.d. X₁,…,Xₙ each bounded in [0, 1] with mean E[X], the sample mean X̄ satisfies, for any t > 0:

\[ P\big(|\bar X - \mathbb{E}[X]| \ge t\big) \;\le\; 2\,\exp\!\big(-2 n t^2\big) \]

The 2 is for the two tails; the exponent carries n (more data) and t² (bigger deviations are far rarer). Our coin is {0,1}-valued, so it lands squarely in Hoeffding's [0,1] hypothesis — this is the bound built for it.

40. Finish it with less help: Hoeffding vs the coin's real tail

Faded example

Fill in the blanks

Hoeffding vs the coin's real tail, with the scaffolding fading: two lines are gone now — fill both.

import numpy as np
rng = np.random.default_rng(0)
p, n, t = 0.3, 100, 0.1
means = rng.binomial(n, p, size=200_000) / n # each is Xbar of n flips
emp = (np.abs(means - p) >= t).mean()
bound = *2np.exp(-2nt2)
print('empirical P(|Xbar-p|>=t):', round(emp, 4))
print('Hoeffding 2exp(-2 n t^2):', round(bound, 4))

Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. The bound holds with margin. It is looser than the truth (Hoeffding is distribution-free over all [0,1] variables) but far tighter than what Markov or Chebyshev would give here.

41. Hoeffding vs the coin's real tail

Worked example

Simulate 200000 experiments, each n = 100 flips of the p = 0.3 coin, and measure how often X̄ misses p by ≥ 0.1. Compare to Hoeffding. Complete snippet:

import numpy as np
rng = np.random.default_rng(0)
p, n, t = 0.3, 100, 0.1
means = rng.binomial(n, p, size=200_000) / n   # each is Xbar of n flips
emp = (np.abs(means - p) >= t).mean()
bound = 2*np.exp(-2*n*t**2)
print('empirical P(|Xbar-p|>=t):', round(emp, 4))
print('Hoeffding 2exp(-2 n t^2):', round(bound, 4))

Empirical 0.0302 sits under the Hoeffding ceiling 0.2707

Why: The bound holds with margin. It is looser than the truth (Hoeffding is distribution-free over all [0,1] variables) but far tighter than what Markov or Chebyshev would give here.

quantityvalue (verified)
n, t100, 0.10
empirical P(|X̄ − p| ≥ 0.1)0.0302
Hoeffding 2e^{−2nt²}0.2707
exact binomial tail0.0299

42. Work backwards from the answer: Hoeffding vs the coin's real tail

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Empirical 0.0302 sits under the Hoeffding ceiling 0.2707

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Simulate 200000 experiments, each n = 100 flips of the p = 0.3 coin, and measure how often X̄ misses p by ≥ 0.1. Compare to Hoeffding. Complete snippet:

43. Finish it with less help: The exponential collapse in n

Faded example

Fill in the blanks

The exponential collapse in n, with the scaffolding fading: two lines are gone now — fill both.

import numpy as np
t = 0.1
for n in [50, 100, 200, 500, 1000]:
bound = *2np.exp(-2nt2)
print(f'n=___: 2exp(-2 n t^2) = ___')

Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. Each extra sample multiplies the bound by exp(−2t²) < 1, so the tail decays geometrically in n.

44. The exponential collapse in n

Worked example

Fix t = 0.1 and grow n. The whole reason Hoeffding matters is how violently the ceiling drops. Runnable:

import numpy as np
t = 0.1
for n in [50, 100, 200, 500, 1000]:
    bound = 2*np.exp(-2*n*t**2)
    print(f'n={n:4d}: 2exp(-2 n t^2) = {bound:.6f}')

From n=50 to n=500 the ceiling falls from 0.74 to 0.00009

Why: Each extra sample multiplies the bound by exp(−2t²) < 1, so the tail decays geometrically in n. That exponential is what makes learning from finitely many samples possible.

n2·exp(−2·n·0.1²) (verified)
500.735759
1000.270671
2000.036631
5000.000091
10000.000000

45. Watch it run: The exponential collapse in n

Pattern

Step through it

Step through The exponential collapse in n one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: n is 50
  2. Step 2: n is 100
  3. Step 3: n is 200
  4. Step 4: n is 500
  5. Step 5: n is 1000

46. This is exactly PAC learning

Intuition

Replace 'coin' with 'a fixed classifier' and 'Xᵢ' with 'is example i misclassified?' (a {0,1} variable). Then X̄ is the training error and E[X] is the true error.

Hoeffding now reads: training error is within t of true error with probability ≥ 1 − 2e^{−2nt²}. That is the Probably Approximately Correct guarantee — 'approximately' is t, 'probably' is the 1 − δ.

So the practical question becomes: how many samples n do I need to promise accuracy ε with confidence 1 − δ? Invert the bound.

47. What has to happen first: Invert Hoeffding for sample complexity

Ranking

Put in order

Put the moves of Invert Hoeffding for sample complexity into the order they have to happen.

  1. Demand the failure probability be ≤ δ
  2. Divide by 2 and take logs
  3. Multiply by −1 (flip the inequality), divide by 2ε²

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. We want P(|X̄ − E[X]| ≥ ε) ≤ δ.

48. Invert Hoeffding for sample complexity

Worked example

Demand the failure probability be ≤ δ

Why: We want P(|X̄ − E[X]| ≥ ε) ≤ δ. Hoeffding's ceiling is 2e^{−2nε²}, so it suffices to force that ceiling below δ.

\[ 2\,e^{-2 n \varepsilon^2} \le \delta \]

Divide by 2 and take logs

Why: ln is increasing, so it preserves ≤; ln of the exponential returns its exponent.

\[ -2 n \varepsilon^2 \le \ln\!\frac{\delta}{2} \;\Longleftrightarrow\; -2 n \varepsilon^2 \le -\ln\!\frac{2}{\delta} \]

Multiply by −1 (flip the inequality), divide by 2ε²

Why: Multiplying by a negative flips ≥ to the sample-size condition. Dividing by the positive 2ε² keeps direction.

\[ \boxed{\,n \ge \dfrac{\ln(2/\delta)}{2\varepsilon^2}\,} \]

49. Decode the notation: Invert Hoeffding for sample complexity

Notation

Annotate

From Invert Hoeffding for sample complexity — read this one piece at a time. What is each part doing?

On: \( \boxed{\,n \ge \dfrac{\ln(2/\delta)}{2\varepsilon^2}\,} \)

  • We want P(|X̄ − E[X]| ≥ ε) ≤ δ. Hoeffding's ceiling is 2e^{−2nε²}, so it suffices to force that ceiling below δ.
  • ln is increasing, so it preserves ≤; ln of the exponential returns its exponent.
  • Multiplying by a negative flips ≥ to the sample-size condition. Dividing by the positive 2ε² keeps direction.

50. Restore the missing line: Sample complexity, computed and checked

Fill the middle

Fill in the blanks

From Sample complexity, computed and checked — one line has had its right-hand side removed. Put it back.

import numpy as np
eps, delta = 0.05, 0.05
n_req = np.log(2/delta) / (2*eps2)
n =
int(np.ceil(n_req))**
print('ln(2/delta) :', round(np.log(2/delta), 4))
print('n >= ... :', round(n_req, 2), '-> ceil', n)
print('check bound at n :', round(2np.exp(-2neps*2), 4), '<= 0.05')

Why: n is what everything below it consumes, so the wrong expression here fails later and somewhere else. The ceiling formula demands 738 flips to promise ±0.05 accuracy 95% of the time.

51. Sample complexity, computed and checked

Worked example

Ask for ε = 0.05 accuracy at δ = 0.05 confidence. Compute n, round up, and plug back to confirm the ceiling really drops under δ. Runnable:

import numpy as np
eps, delta = 0.05, 0.05
n_req = np.log(2/delta) / (2*eps**2)
n = int(np.ceil(n_req))
print('ln(2/delta)      :', round(np.log(2/delta), 4))
print('n >= ...         :', round(n_req, 2), '-> ceil', n)
print('check bound at n :', round(2*np.exp(-2*n*eps**2), 4), '<= 0.05')

n ≥ 737.78 → 738 samples; the bound at 738 is 0.0499 ≤ 0.05

Why: The ceiling formula demands 738 flips to promise ±0.05 accuracy 95% of the time. Plugging 738 back gives 0.0499, just under δ — the inversion is exact.

quantityvalue (verified)
ln(2/δ)3.6889
2ε²0.005
n ≥ ln(2/δ)/(2ε²)737.78 → 738
2e^{−2·738·ε²}0.0499 (≤ 0.05 ✓)

52. The three inequalities, ranked

Concept

Each bound adds one assumption and tightens. Read the ladder top to bottom — it is the single most exam-tested idea in this lesson.

boundneedstail on X̄decay in n
MarkovX ≥ 0, meanE[X]/a— (single var)
Chebyshev+ variance σ²σ²/(n t²)1/n
Hoeffding+ bounded [0,1], i.i.d.2e^{−2nt²}exponential

More structure ⇒ tighter bound. When your variable is bounded — as {0,1} losses always are — reach for Hoeffding.

53. Something is wrong here: Hoeffding on an unbounded variable

Anomaly

Predict first

A student writes this, and it looks reasonable:

The coin's X̄ obeys 2e^{−2nt²}, so use the same bound for the average of n draws from Exponential(1).

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Hoeffding REQUIRES each variable lie in a bounded range [a,b] — the exponent is really 2nt²/(b−a)².

Match the tool to the support. For an unbounded variable with known variance, use Chebyshev on the average.

Why: Hoeffding REQUIRES each variable lie in a bounded range [a,b] — the exponent is really 2nt²/(b−a)². The exponential is unbounded (b = ∞), so (b−a)² = ∞ makes the exponent 0 and the bound the useless '≤ 2'. The clean {0,1} version silently assumed b−a = 1.

54. Trap: Hoeffding on an unbounded variable

Trap

The trap

The coin's X̄ obeys 2e^{−2nt²}, so use the same bound for the average of n draws from Exponential(1).

Apply 2e^{−2nt²} to the exponential average

Why: Hoeffding REQUIRES each variable lie in a bounded range [a,b] — the exponent is really 2nt²/(b−a)². The exponential is unbounded (b = ∞), so (b−a)² = ∞ makes the exponent 0 and the bound the useless '≤ 2'. The clean {0,1} version silently assumed b−a = 1.

The fix

Match the tool to the support. For an unbounded variable with known variance, use Chebyshev on the average.

P(|X̄ − μ| ≥ t) ≤ σ²/(n t²) via Chebyshev

Why: Averaging shrinks variance to σ²/n, so Chebyshev still gives a valid 1/n bound with no boundedness needed. Hoeffding's exponential speed is a reward for boundedness — you cannot claim it for free.

55. Break it on purpose: Hoeffding on an unbounded variable

Break the constraint

Discussion prompt

The rule this trap just fixed:

Averaging shrinks variance to σ²/n, so Chebyshev still gives a valid 1/n bound with no boundedness needed. Hoeffding's exponential speed is a reward for boundedness — you cannot claim it for free.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

Hoeffding REQUIRES each variable lie in a bounded range [a,b] — the exponent is really 2nt²/(b−a)². The exponential is unbounded (b = ∞), so (b−a)² = ∞ makes the exponent 0 and the bound the useless '≤ 2'. The clean {0,1} version silently assumed b−a = 1.

56. VC dimension — measuring capacity

Section

Part 5 of 6 — shattering

57. From one classifier to a whole class

Intuition

Hoeffding pins the training error of one fixed classifier near its true error. But a learner searches over a whole class H and keeps whichever looks best — so it can get lucky on some hypothesis in H by chance.

The more hypotheses H contains, the more chances to get lucky, and the wider the possible train–true gap. We need a number for 'how rich is H'. Counting hypotheses fails (often infinite), so we count something smarter: how many points H can label every possible way.

58. Shattering

Concept

shatter — A hypothesis class H shatters a set of m points if, for EVERY one of the 2^m ways to label them +/−, some hypothesis in H realizes that exact labeling. If even one labeling is impossible, H does not shatter the set.

Shattering is an all-or-nothing test on a specific point set. m points have 2ᵐ labelings; H shatters them only if it can hit all 2ᵐ.

59. Take the definitions apart: one-sided vs two-sided vs shatter

Definition probe

Sort into buckets

Every line below is part of the definition of one-sided vs two-sided or of shatter — one or the other, never both. Put each where it belongs.

one-sided vs two-sided
Markov bounds only the RIGHT tail P(X ≥ a) of a non-negative variable.; To bound how far X falls on BOTH sides of its mean, we need the deviation itself; that is the job of Chebyshev.
shatter
A hypothesis class H shatters a set of m points if, for EVERY one of the 2^m ways to label them +/−, some hypothesis in H realizes that exact labeling.; If even one labeling is impossible, H does not shatter the set.
b1
Markov bounds only the RIGHT tail P(X ≥ a) of a non-negative variable. To bound how far X falls on BOTH sides of its mean, we need the deviation itself — that is the job of Chebyshev.
b2
A hypothesis class H shatters a set of m points if, for EVERY one of the 2^m ways to label them +/−, some hypothesis in H realizes that exact labeling. If even one labeling is impossible, H does not shatter the set.

60. Guess the shape of the answer: A threshold shatters 2 points on a line

Estimation

Predict first

Simplest class: a 1-D threshold h(x) = sign(x − θ) (with an optional flip). Two points x = 1, 2 have 2² = 4 labelings. We brute-force over thresholds and check all four are realizable:

Commit before you compute: what does A threshold shatters 2 points on a line come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: All 4 labelings realizable → the threshold class shatters 2 points

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Sliding θ between, below, or above the points, plus the sign flip, reaches every labeling.

61. A threshold shatters 2 points on a line

Worked example

Simplest class: a 1-D threshold h(x) = sign(x − θ) (with an optional flip). Two points x = 1, 2 have 2² = 4 labelings. We brute-force over thresholds and check all four are realizable:

import numpy as np
pts = np.array([1.0, 2.0])
def h(x, theta, sign):
    return np.where(x > theta, sign, -sign)
targets = [(-1,-1), (-1,1), (1,-1), (1,1)]
for tgt in targets:
    ok = any(tuple(h(pts, th, s)) == tgt
             for th in np.linspace(0, 3, 31) for s in (1, -1))
    print('labeling', tgt, 'realizable:', ok)

All 4 labelings realizable → the threshold class shatters 2 points

Why: Sliding θ between, below, or above the points, plus the sign flip, reaches every labeling. So VC of 1-D thresholds is at least 2... but is it more?

labeling of (x=1, x=2)realizable? (verified)
(−1, −1)True
(−1, +1)True
(+1, −1)True
(+1, +1)True

62. What stays fixed: A threshold shatters 2 points on a line

Invariant

Step through it

Step through A threshold shatters 2 points on a line one row at a time. One of these columns never changes — find it, and say why it cannot.

  1. Step 1: labeling of (x=1, x=2) is (−1, −1)
  2. Step 2: labeling of (x=1, x=2) is (−1, +1)
  3. Step 3: labeling of (x=1, x=2) is (+1, −1)
  4. Step 4: labeling of (x=1, x=2) is (+1, +1)

63. A threshold FAILS on 3 points

Worked example

Now three points 1, 2, 3. The alternating labeling (+, −, +) needs the classifier to flip sign twice — a single threshold can't. Count how many of the 8 labelings are actually reachable:

import numpy as np
from itertools import product
pts = np.array([1.0, 2.0, 3.0])
def h(x, theta, sign):
    return np.where(x > theta, sign, -sign)
good = 0
for tgt in product((1, -1), repeat=3):
    ok = any(tuple(h(pts, th, s)) == tgt
             for th in np.linspace(0, 4, 401) for s in (1, -1))
    good += ok
print('realizable labelings of 3 points:', good, 'of 8')

Only 6 of 8 labelings work — (+,−,+) and (−,+,−) are impossible

Why: A threshold splits the line into one 'low' side and one 'high' side; it cannot carve out a middle group. So 1-D thresholds shatter 2 points but not 3 — their VC dimension is exactly 2.

labeling of (1, 2, 3)realizable?
(+, +, +) / (−, −, −) / …6 of these work
(+, −, +)impossible
(−, +, −)impossible
total realizable6 of 8

64. Which is which, by realizable?

Discrimination

Sort into buckets

Sort these by realizable?, from memory, without looking back at A threshold FAILS on 3 points. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

6 of these work
(+, +, +) / (−, −, −) / …
impossible
(+, −, +); (−, +, −)
6 of 8
total realizable
g1
realizable? is "6 of these work" for (+, +, +) / (−, −, −) / … — that is what the table on "A threshold FAILS on 3 points" records, and it is the single property separating this group from the rest.
g2
realizable? is "impossible" for (+, −, +), (−, +, −) — that is what the table on "A threshold FAILS on 3 points" records, and it is the single property separating this group from the rest.
g3
realizable? is "6 of 8" for total realizable — that is what the table on "A threshold FAILS on 3 points" records, and it is the single property separating this group from the rest.

65. VC dimension, defined

Concept

VC dimension — The VC dimension of H is the size of the LARGEST point set H can shatter. Formally: the largest m such that SOME set of m points is shattered by H. It is a single integer summarizing the capacity of the whole class.

Two subtleties that trip people up: you need only find one set of m points that is shattered (not all sets); and to prove VC < m+1, you must show no set of m+1 points can be shattered.

66. A line in ℝ² shatters 3 points

Worked example

Move to 2-D linear classifiers (a separating line, d = 2). Take a triangle — three non-collinear points — and check that a line can realize all 2³ = 8 labelings. We fit a hard-margin LinearSVC per labeling:

import numpy as np
from itertools import product
from sklearn.svm import LinearSVC
P = np.array([[0., 0.], [2., 0.], [0., 2.]])   # non-collinear triangle
n_sep = 0
for labels in product([0, 1], repeat=3):
    if len(set(labels)) == 1:
        n_sep += 1                              # single class: trivial
        continue
    clf = LinearSVC(C=1e6, max_iter=100000).fit(P, np.array(labels))
    n_sep += (clf.predict(P) == np.array(labels)).mean() == 1.0
print('labelings a line can realize:', n_sep, 'of', 2**3)

8 of 8 labelings realized → a line shatters 3 points

Why: For every +/− split of the triangle, some line separates the + from the − points (single-class labelings are trivially 'separated' by any line). So VC of a 2-D linear classifier is at least 3.

labelings triedrealized by a line (verified)
2 single-class (all +, all −)2 (trivial)
6 mixed labelings6 (SVC acc = 1.0)
total8 of 8 → shattered

67. Restore the missing line: A line CANNOT shatter 4 points — XOR

Fill the middle

Fill in the blanks

From A line CANNOT shatter 4 points — XOR — one line has had its right-hand side removed. Put it back.

import numpy as np
X = np.array([[0.,0.],[1.,1.],[1.,0.],[0.,1.]])
y = np.array([0, 0, 1, 1]) # XOR: diagonal pairs share a class
g = np.linspace(-3, 3, 61)
best = 0.0
for a in g:
for b in g:
for c in g:
pred = ((aX[:,0] + bX[:,1] + c) > 0).astype(int)
best = max(best, (pred == y).mean(), (1-pred == y).mean())
print('best line accuracy on XOR:', best)

Why: g is what everything below it consumes, so the wrong expression here fails later and somewhere else. No line separates the two diagonal classes, so the (+,−,+,−)-style XOR labeling is unrealizable.

68. A line CANNOT shatter 4 points — XOR

Worked example

Four points in the XOR arrangement defeat every line. Grid-search all lines sign(a·x + b·y + c) and record the best accuracy any line achieves on the XOR labeling:

import numpy as np
X = np.array([[0.,0.],[1.,1.],[1.,0.],[0.,1.]])
y = np.array([0, 0, 1, 1])                  # XOR: diagonal pairs share a class
g = np.linspace(-3, 3, 61)
best = 0.0
for a in g:
    for b in g:
        for c in g:
            pred = ((a*X[:,0] + b*X[:,1] + c) > 0).astype(int)
            best = max(best, (pred == y).mean(), (1-pred == y).mean())
print('best line accuracy on XOR:', best)

Best a line can do on XOR is 0.75 — it always misses ≥ 1 of 4

Why: No line separates the two diagonal classes, so the (+,−,+,−)-style XOR labeling is unrealizable. A line fails to shatter 4 points; combined with shattering 3, VC(line in ℝ²) = 3.

point setbest line accuracy (verified)
3-point triangle (all labelings)1.00 → shattered
4-point XOR0.75 → NOT shattered
⇒ VC dimension of a 2-D line3 = d + 1

69. Fill in: best line accuracy (verified) for A line CANNOT shatter 4 points — XOR

Comparison

Comparison matrix

From A line CANNOT shatter 4 points — XOR: refill the best line accuracy (verified) column from what you know. The rest of the table is as it appeared.

point setbest line accuracy (verified)
3-point triangle (all labelings)1.00 → shattered
4-point XOR0.75 → NOT shattered
⇒ VC dimension of a 2-D line3 = d + 1

70. The general rule: VC = d + 1

Concept

The two experiments (shatter 3, fail 4) are the d = 2 case of a general theorem for linear classifiers in ℝᵈ (with a bias term):

\[ \text{VC}\big(\text{linear classifiers in } \mathbb{R}^d\big) = d + 1 \]

The +1 is the bias/offset degree of freedom — the same ones-column that let a regression line leave the origin. On a plane (d = 2) that gives 3, matching what we just brute-forced.

71. Something is wrong here: VC = d, forgetting the +1

Anomaly

Predict first

A student writes this, and it looks reasonable:

A linear classifier in ℝ² has 2 parameters of direction, so its VC dimension is 2.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: This drops the bias term. A line through the ORIGIN (no offset) shatters only 2 points, but a general line has a third degree of freedom — its offset c — and shatters 3.

Count the offset. A general line a·x + b·y + c has d + 1 = 3 free parameters.

Why: This drops the bias term. A line through the ORIGIN (no offset) shatters only 2 points, but a general line has a third degree of freedom — its offset c — and shatters 3. Our brute force realized all 8 labelings of 3 points, so 2 is too small.

72. Trap: VC = d, forgetting the +1

Trap

The trap

A linear classifier in ℝ² has 2 parameters of direction, so its VC dimension is 2.

Answer VC = d = 2

Why: This drops the bias term. A line through the ORIGIN (no offset) shatters only 2 points, but a general line has a third degree of freedom — its offset c — and shatters 3. Our brute force realized all 8 labelings of 3 points, so 2 is too small.

The fix

Count the offset. A general line a·x + b·y + c has d + 1 = 3 free parameters.

VC = d + 1 = 3

Why: The '+1' is the bias c. It is the same extra degree of freedom as the ones-column in regression, and it is exactly what lets the line shatter one more point than d alone would.

73. The generalization bound

Section

Part 6 of 6 — and where it breaks

74. Sauer's lemma: capacity is polynomial

Concept

A class with VC dimension d can label at most Πₕ(m) distinct ways on any m points. Sauer's lemma bounds that growth function by a polynomial once m > d:

\[ \Pi_H(m) \;\le\; \sum_{i=0}^{d} \binom{m}{i} \;\le\; \Big(\tfrac{e\,m}{d}\Big)^{d} \quad (m > d) \]

This is the key that turns capacity into a bound: below m = d the class labels all 2ᵐ ways (shattering), but past d the count grows only polynomially, not exponentially — so infinitely many hypotheses still behave like a small finite set.

75. Predict the next row: Watch the growth go sub-exponential

Pattern

Predict first

The table runs: 3 | 8 | 8 · 4 | 15 | 16 · 5 | 26 | 32

In Watch the growth go sub-exponential, given the rows so far: what is the next one — the row where m is 10?

Correct: 10 | 176 | 1024

mSauer Σ_{i≤3} C(m,i)2ᵐ
388
41516
52632
101761024

Why: The relationship between the columns, not the individual numbers, is what generates the next row. Up to the VC dimension the class realizes every labeling (8 = 2³).

76. Watch the growth go sub-exponential

Worked example

For a line in ℝ² (d = 3), tabulate Sauer's sum Σ_{i≤3} C(m, i) against the naive 2ᵐ. Watch them agree up to m = 3, then diverge:

from math import comb
d = 3   # VC dimension of a line in R^2
for m in [3, 4, 5, 10]:
    sauer = sum(comb(m, i) for i in range(d+1))
    print(f'm={m:2d}: growth <= {sauer:4d}   vs 2^m = {2**m}')

At m=3 both are 8 (shattering); by m=10, 176 vs 1024

Why: Up to the VC dimension the class realizes every labeling (8 = 2³). Beyond it, Sauer's polynomial (176) falls far below the exponential (1024) — capacity is effectively finite.

mSauer Σ_{i≤3} C(m,i)2ᵐ
388
41516
52632
101761024

77. Watch it run: Watch the growth go sub-exponential

Pattern

Step through it

Step through Watch the growth go sub-exponential one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: m is 3
  2. Step 2: m is 4
  3. Step 3: m is 5
  4. Step 4: m is 10

78. The union bound: paying for the search

Intuition

Hoeffding controls one hypothesis. A learner tries many, so we pay a union bound: the chance that some hypothesis in H is fooled is at most the sum of the individual failure chances.

If H behaved like M distinct hypotheses, the total failure budget is M · 2e^{−2nt²}. Forcing that below δ and solving for t puts a √(ln M / n) term in the gap — the price of choosing.

The trouble: M is usually infinite. Sauer's lemma rescues it — on n points, H acts like only Πₕ(n) ≤ (en/VC)^{VC} effective hypotheses, so ln M becomes ≈ VC·ln n. That substitution is where the VC term is born.

79. The VC generalization bound

Concept

Plugging Sauer's polynomial growth into a Hoeffding-style union argument yields the headline result. With probability ≥ 1 − δ, for every hypothesis in H at once:

\[ \underbrace{\text{err}_{\text{true}}}_{\text{gen}} \;\le\; \underbrace{\text{err}_{\text{train}}}_{\text{fit}} \;+\; O\!\left(\sqrt{\frac{\text{VC}\cdot \ln n}{n}}\right) \]

The gap term shrinks with data n and grows with capacity VC. It is the rigorous form of 'enough data beats a complex model' — and the reason Hoeffding was worth proving: it is the engine inside the square root.

80. Teach it back: The VC generalization bound

Explain it

Discussion prompt

Explain The VC generalization bound to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Plugging Sauer's polynomial growth into a Hoeffding-style union argument yields the headline result. With probability ≥ 1 − δ, for every hypothesis in H at once:

81. What has to be given first: Compute a generalization gap

Missing information

Discussion prompt

Exam-style: n = 1000 examples, a model of VC dimension 50. Evaluate the gap term √(VC·ln n / n), and sweep to see how data and capacity move it. Runnable:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

The bound is pessimistic (0.59 is a huge gap), yet increasing n 100× shrinks it to 0.076 — data closes the gap. Raising VC to 1000 blows it past 1 (vacuous), showing capacity's cost.

82. Compute a generalization gap

Worked example

Exam-style: n = 1000 examples, a model of VC dimension 50. Evaluate the gap term √(VC·ln n / n), and sweep to see how data and capacity move it. Runnable:

import numpy as np
for vc, n in [(3, 100), (50, 1000), (50, 100_000), (1000, 1000)]:
    term = np.sqrt(vc*np.log(n)/n)
    print(f'VC={vc:5d}, n={n:6d}: sqrt(VC ln n / n) = {term:.4f}')

VC=50, n=1000 → 0.59; hold VC, grow n to 100k → 0.076

Why: The bound is pessimistic (0.59 is a huge gap), yet increasing n 100× shrinks it to 0.076 — data closes the gap. Raising VC to 1000 blows it past 1 (vacuous), showing capacity's cost.

VCn√(VC·ln n / n) (verified)
31000.3717
5010000.5877
501000000.0759
100010002.6283 (vacuous)

83. Rebuild the recipe: The concentration → generalization pipeline

Ranking

Put in order

These are the steps of The concentration → generalization pipeline, scrambled. Put them back in order before the next slide shows you.

  1. Bound one hypothesis: X̄ (train err) concentrates near E[X] (true err) by Hoeffding 2e^{−2nt²}
  2. Measure the class: capacity = VC dimension = largest shattered set; linear in ℝᵈ is d+1
  3. Tame the search: Sauer turns the 2ᵐ labelings into a polynomial (em/d)ᵈ once m > VC
  4. Union-bound the whole class: gap ≤ √(VC·ln n / n) — shrinks with n, grows with VC
  5. Get n: invert to n ≥ ln(2/δ)/(2ε²) per hypothesis; capacity multiplies it

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

84. The concentration → generalization pipeline

Pattern

  1. Bound one hypothesis: X̄ (train err) concentrates near E[X] (true err) by Hoeffding 2e^{−2nt²}
  2. Measure the class: capacity = VC dimension = largest shattered set; linear in ℝᵈ is d+1
  3. Tame the search: Sauer turns the 2ᵐ labelings into a polynomial (em/d)ᵈ once m > VC
  4. Union-bound the whole class: gap ≤ √(VC·ln n / n) — shrinks with n, grows with VC
  5. Get n: invert to n ≥ ln(2/δ)/(2ε²) per hypothesis; capacity multiplies it

85. Where does it stop working: The concentration → generalization pipeline

Edge cases

Discussion prompt

The concentration → generalization pipeline works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Bound one hypothesis: X̄ (train err) concentrates near E[X] (true err) by Hoeffding 2e^{−2nt²}
  2. Measure the class: capacity = VC dimension = largest shattered set; linear in ℝᵈ is d+1
  3. Tame the search: Sauer turns the 2ᵐ labelings into a polynomial (em/d)ᵈ once m > VC
  4. Union-bound the whole class: gap ≤ √(VC·ln n / n) — shrinks with n, grows with VC
  5. Get n: invert to n ≥ ln(2/δ)/(2ε²) per hypothesis; capacity multiplies it

86. Rule out three: Check yourself — least assumptions

Elimination

Eliminate the wrong options

Which inequality needs ONLY that X ≥ 0 and E[X] exists — no variance, no boundedness?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Markov
  • B. Chebyshev
  • C. Hoeffding
  • D. The VC generalization bound

Survives elimination: A

Why: Markov's derivation used only a·1[X≥a] ≤ X, which needs X ≥ 0 and a finite mean — nothing else. That minimal assumption set is exactly why it is the loosest of the three and serves as the lemma for the others.

87. Check yourself — least assumptions

Check

Which bound survives on the least information?

Check your understanding

Which inequality needs ONLY that X ≥ 0 and E[X] exists — no variance, no boundedness?

  • A. Markov (correct)
  • B. Chebyshev
  • C. Hoeffding
  • D. The VC generalization bound

Answer: A

Why: Markov's derivation used only a·1[X≥a] ≤ X, which needs X ≥ 0 and a finite mean — nothing else. That minimal assumption set is exactly why it is the loosest of the three and serves as the lemma for the others.

Why B tempts people
Chebyshev additionally needs the variance σ²: it is literally Markov applied to (X−μ)², whose mean is Var(X). No variance, no Chebyshev.
Why C tempts people
Hoeffding requires the variables to be BOUNDED (here in [0,1]) and i.i.d. — much more structure than Markov, which is what buys its exponential tail.
Why D tempts people
The VC bound is about the capacity of a whole hypothesis class across n samples, not the tail of a single non-negative random variable.

88. Answer it before you see the options: Check yourself — the Chebyshev trick

Prediction

Predict first

Chebyshev's P(|X−μ| ≥ kσ) ≤ 1/k² is obtained by applying Markov to which non-negative variable?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: (X − μ)², at level (kσ)²

Why: Squaring the deviation makes it non-negative (so Markov applies) and gives it mean E[(X−μ)²] = σ². Markov at level (kσ)² then yields σ²/(k²σ²) = 1/k² after the variance cancels.

89. Check yourself — the Chebyshev trick

Check

Recall how we got a two-sided bound from a one-sided lemma.

Check your understanding

Chebyshev's P(|X−μ| ≥ kσ) ≤ 1/k² is obtained by applying Markov to which non-negative variable?

  • A. (X − μ)², at level (kσ)² (correct)
  • B. X itself, at level kσ
  • C. |X|, at level kσ
  • D. e^{X}, at level e^{kσ}

Answer: A

Why: Squaring the deviation makes it non-negative (so Markov applies) and gives it mean E[(X−μ)²] = σ². Markov at level (kσ)² then yields σ²/(k²σ²) = 1/k² after the variance cancels.

Why B tempts people
Markov on X gives a one-sided right-tail bound E[X]/(kσ), not the two-sided deviation bound — and X need not even be centered or non-negative.
Why C tempts people
Markov on |X| bounds P(|X| ≥ kσ), a statement about magnitude around 0, not deviation |X−μ| around the mean; and its mean is E|X|, not σ².
Why D tempts people
Applying Markov to e^{X} is the Chernoff/Hoeffding route (the moment-generating function), not the Chebyshev derivation, which uses the second moment (X−μ)².

90. Answer it before you see the options: Check yourself — VC of a line

Prediction

Predict first

The VC dimension of a linear classifier in ℝ² (a general line with offset) is:

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: 3 — it shatters some 3 points but no 4 (XOR fails)

Why: VC = d + 1 = 3. Our brute force realized all 8 labelings of a triangle (shatters 3) but capped at 0.75 accuracy on the 4-point XOR (cannot shatter 4). The largest shattered set has size 3.

91. Check yourself — VC of a line

Check

Count degrees of freedom, and remember the offset.

Check your understanding

The VC dimension of a linear classifier in ℝ² (a general line with offset) is:

  • A. 3 — it shatters some 3 points but no 4 (XOR fails) (correct)
  • B. 2 — one per input dimension
  • C. 4 — it can label any 4 points
  • D. infinite — lines are very flexible

Answer: A

Why: VC = d + 1 = 3. Our brute force realized all 8 labelings of a triangle (shatters 3) but capped at 0.75 accuracy on the 4-point XOR (cannot shatter 4). The largest shattered set has size 3.

Why B tempts people
That drops the '+1' bias term. A line through the origin shatters only 2, but a general line's offset c adds a degree of freedom, letting it shatter 3.
Why C tempts people
A line cannot shatter 4 points — XOR is the counterexample, where the best any line achieves is 3 of 4 correct (0.75), never all 8 labelings.
Why D tempts people
Linear classifiers have FINITE VC dimension (d+1). Infinite VC belongs to far richer classes, e.g. 1-D thresholds on a cleverly ordered set, or sine classifiers.

92. Rule out three: Check yourself — closing the gap

Elimination

Eliminate the wrong options

By the VC bound gap ≈ √(VC·ln n / n), the train–true gap shrinks when:

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. n grows, or VC dimension falls
  • B. VC dimension grows
  • C. the learning rate increases
  • D. the training error decreases

Survives elimination: A

Why: VC sits in the numerator and n dominates the denominator, so more data shrinks the gap and more capacity widens it. Our sweep confirmed it: VC=50 gap fell from 0.59 at n=1000 to 0.076 at n=100000.

93. Check yourself — closing the gap

Check

Read the square root: what is in the numerator, what is in the denominator?

Check your understanding

By the VC bound gap ≈ √(VC·ln n / n), the train–true gap shrinks when:

  • A. n grows, or VC dimension falls (correct)
  • B. VC dimension grows
  • C. the learning rate increases
  • D. the training error decreases

Answer: A

Why: VC sits in the numerator and n dominates the denominator, so more data shrinks the gap and more capacity widens it. Our sweep confirmed it: VC=50 gap fell from 0.59 at n=1000 to 0.076 at n=100000.

Why B tempts people
Higher VC is MORE capacity to overfit, so it WIDENS the bound — at VC=1000, n=1000 the term exceeds 2.6 and becomes vacuous.
Why C tempts people
The learning rate is an optimization hyperparameter; it does not appear anywhere in a capacity-vs-data generalization bound.
Why D tempts people
Lower training error can come WITH a larger gap (classic overfitting). The bound's gap term depends on VC and n, not on the training error itself.

94. Where the classical story breaks

Concept

Modern neural nets have VC dimension far larger than n — plug that in and the bound is > 1, i.e. vacuous. Classical theory predicts disaster. Yet these overparameterized nets generalize beautifully. Something is missing from the worst-case picture.

The empirical resolution is double descent: as capacity grows past the point where the model exactly fits the training data (the interpolation threshold), test error rises to a spike — then falls again. We can watch it happen in code.

95. Guess the shape of the answer: Double descent, from a real fit

Estimation

Predict first

Fit a least-norm linear model on random tanh features, sweeping the feature count p past the sample size n = 40. Watch test MSE spike at p ≈ n then descend. Runnable as written:

Commit before you compute: what does Double descent, from a real fit come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Test MSE dives, spikes near p ≈ n = 40, then descends again

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. At p = 40 the model exactly interpolates the 40 training points and test error explodes (~38); push p to 400 and the least-norm solution becomes smooth again, MSE back near 0.75.

96. Double descent, from a real fit

Worked example

Fit a least-norm linear model on random tanh features, sweeping the feature count p past the sample size n = 40. Watch test MSE spike at p ≈ n then descend. Runnable as written:

import numpy as np
rng = np.random.default_rng(0)
n, d, sigma, ntest = 40, 5, 0.3, 2000
beta = rng.standard_normal(d)
Xtr = rng.standard_normal((n, d)); ytr = Xtr @ beta + sigma*rng.standard_normal(n)
Xte = rng.standard_normal((ntest, d)); yte = Xte @ beta + sigma*rng.standard_normal(ntest)
for p in [5, 20, 40, 60, 400]:
    W = rng.standard_normal((d, p)) / np.sqrt(d)
    Ptr, Pte = np.tanh(Xtr @ W), np.tanh(Xte @ W)
    w = np.linalg.pinv(Ptr) @ ytr            # least-norm fit
    print(f'p={p:3d}: test MSE = {np.mean((Pte @ w - yte)**2):.3f}')

Test MSE dives, spikes near p ≈ n = 40, then descends again

Why: At p = 40 the model exactly interpolates the 40 training points and test error explodes (~38); push p to 400 and the least-norm solution becomes smooth again, MSE back near 0.75. Classical VC sees only the rising middle.

p (features)test MSE (verified)
5 (under-param)0.197
200.601
40 (≈ interpolation)37.930 (spike)
604.843
400 (over-param)0.746

97. Work backwards from the answer: Double descent, from a real fit

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Test MSE dives, spikes near p ≈ n = 40, then descends again

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Fit a least-norm linear model on random tanh features, sweeping the feature count p past the sample size n = 40. Watch test MSE spike at p ≈ n then descend. Runnable as written:

98. Something is wrong here: 'more parameters ⇒ always worse test error'

Anomaly

Predict first

A student writes this, and it looks reasonable:

The p = 40 fit had test MSE 37.9, so adding still more features must push test error even higher.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: That reads only the FIRST descent and the spike.

Recognize the second descent: past the interpolation threshold, more capacity plus an implicit smoothness bias lowers test error again.

Why: That reads only the FIRST descent and the spike. Our own sweep refutes it: at p = 60 the MSE is already back down to 4.84, and at p = 400 it is 0.75 — an order of magnitude BELOW the p = 40 spike. Test error is non-monotonic in capacity.

99. Trap: 'more parameters ⇒ always worse test error'

Trap

The trap

The p = 40 fit had test MSE 37.9, so adding still more features must push test error even higher.

Extrapolate the rising curve past p = n

Why: That reads only the FIRST descent and the spike. Our own sweep refutes it: at p = 60 the MSE is already back down to 4.84, and at p = 400 it is 0.75 — an order of magnitude BELOW the p = 40 spike. Test error is non-monotonic in capacity.

The fix

Recognize the second descent: past the interpolation threshold, more capacity plus an implicit smoothness bias lowers test error again.

Read the whole double-descent curve

Why: Error falls (under-param), spikes at p ≈ n, then falls again (over-param). The least-norm solution among many interpolators is the smoothest one, so huge p is safe — the regime modern nets actually live in.

100. Which of these survive contact with Lesson 34: Concentration Inequalities &…?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
X̄ is our estimate of p, and in ML it plays the role of training error estimating true error. The whole lesson asks: how far can X̄ stray from p?; Check both cases: if X ≥ a the left side is a and a ≤ X; if X < a the left side is 0 ≤ X (since X ≥ 0). Either way it holds. Now just take expectations.; We want a two-sided bound on |X − μ|. The variable |X − μ| can be positive from either side, so we square it: (X − μ)² is non-negative — precisely what Markov needs.
Breaks
Markov gives P(X ≥ 4) ≤ 0.25, so roughly a quarter of the mass sits above 4.; To bound P(|X−μ| ≥ kσ), apply Markov directly to |X − μ|: P(|X−μ| ≥ kσ) ≤ E|X−μ| / (kσ).
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 34: Concentration Inequalities & Generalization puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

101. Why over-parameterization can help

Intuition

Right at p ≈ n the model is forced to thread every noisy point with essentially one solution — a wild, high-curvature fit. That is the spike.

Past that, many solutions interpolate the data, and the least-norm / gradient-descent one picks the smoothest among them. That implicit bias toward simple functions — not raw parameter count — is what VC's worst-case bound cannot see.

So VC bounds are still true (they are worst-case guarantees), just not tight for the specific, well-regularized functions modern training actually finds.

102. By analogy: Why over-parameterization can help

Analogy

Discussion prompt

Explain Why over-parameterization can help by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Right at p ≈ n the model is forced to thread every noisy point with essentially one solution — a wild, high-curvature fit. That is the spike.

103. Your turn: verify the bounds

Section

The project

104. Project: concentration & capacity in code

Concept

Rebuild the lesson's four pillars yourself: two empirical tail checks, one sample-complexity inversion, and one capacity computation. You've derived every piece — now assemble it.

#milestonetool
1Markov: empirical P(X≥4) ≤ E[X]/4rng.exponential
2Chebyshev: empirical P(|X−1|≥2) ≤ 1/4np.abs, .mean()
3Hoeffding: bound + invert for nnp.exp, np.log
4VC gap term √(VC·ln n / n)np.sqrt

Build rules: type every line, draw ≥ 10⁶ samples for stable tail estimates, and confirm each empirical probability sits under its bound before moving on.

105. Break it if you can: Project: concentration & capacity in code

Counterexample

Discussion prompt

Rebuild the lesson's four pillars yourself: two empirical tail checks, one sample-complexity inversion, and one capacity computation. You've derived every piece — now assemble it.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Build rules: type every line, draw ≥ 10⁶ samples for stable tail estimates, and confirm each empirical probability sits under its bound before moving on.

106. Milestone 1 — Markov empirically

Worked example

Your turn: sample Exponential(1) and check P(X ≥ 4) against the Markov ceiling E[X]/4 = 1/4. Predict which is larger before you run it.

Hint: (X >= 4).mean() is the empirical tail; the bound is 1/4. They should straddle — empirical well below the bound.

import numpy as np
rng = np.random.default_rng(0)
X = rng.exponential(1.0, 2_000_000)
print('empirical P(X>=4):', round((X >= 4).mean(), 4))
print('Markov bound 1/4 :', 1/4)
quantityexpected value
empirical P(X ≥ 4)0.0183
Markov bound 1/40.25
holds (empirical ≤ bound)?yes

107. Milestone 2 — Chebyshev empirically

Worked example

Your turn: on the same Exponential(1) (μ = σ = 1), check the two-sided tail P(|X − 1| ≥ 2) against 1/k² = 1/4 at k = 2. Predict: under the bound?

Hint: (np.abs(X - 1) >= 2).mean() versus 1 / 2**2. Reuse the X you already drew.

import numpy as np
rng = np.random.default_rng(0)
X = rng.exponential(1.0, 2_000_000)
print('empirical P(|X-1|>=2):', round((np.abs(X - 1) >= 2).mean(), 4))
print('Chebyshev 1/k^2      :', 1/2**2)
quantityexpected value
empirical P(|X − 1| ≥ 2)0.0498
Chebyshev bound 1/40.25
holds?yes

108. Milestone 3 — Hoeffding and its inverse

Worked example

Your turn: compute the Hoeffding bound at n = 1000, t = 0.05, then invert it to find the n needed for ε = 0.05, δ = 0.05. Predict whether the inverted n is in the hundreds or thousands.

Hint: bound = 2*np.exp(-2*n*t**2); inverse n = ln(2/δ)/(2ε²), rounded up with np.ceil.

import numpy as np
print('Hoeffding n=1000,t=0.05:', round(2*np.exp(-2*1000*0.05**2), 4))
eps, delta = 0.05, 0.05
n = int(np.ceil(np.log(2/delta) / (2*eps**2)))
print('samples needed n       :', n)
print('check bound at n       :', round(2*np.exp(-2*n*eps**2), 4))
quantityexpected value
Hoeffding (n=1000, t=0.05)0.0135
n = ⌈ln(2/δ)/(2ε²)⌉738
bound at n = 7380.0499 (≤ 0.05 ✓)

109. Milestone 4 — the VC gap term

Worked example

Your turn: compute the gap term √(VC·ln n / n) for VC = 50, n = 1000, then again at n = 100000 to watch data close the gap. Predict the direction of change.

Hint: np.sqrt(vc*np.log(n)/n). Growing n should shrink the term.

import numpy as np
vc = 50
for n in [1000, 100_000]:
    print(f'n={n:6d}: gap term = {np.sqrt(vc*np.log(n)/n):.4f}')
VCngap term (expected)
5010000.5877
501000000.0759
directionn↑gap ↓

110. What each one costs: Milestone 4 — the VC gap term

Trade off

Comparison matrix

From Milestone 4 — the VC gap term: every row here is a choice with a cost. Fill the n column, then say which row you would actually pick and what you give up for it.

VCngap term (expected)
5010000.5877
501000000.0759
directionn↑gap ↓

111. The full program

Concept

import numpy as np
rng = np.random.default_rng(0)
X = rng.exponential(1.0, 2_000_000)                 # E[X]=1, sigma=1

print('Markov    P(X>=4)   :', round((X>=4).mean(),4),        '<=', 1/4)
print('Chebyshev P(|X-1|>=2):', round((np.abs(X-1)>=2).mean(),4), '<=', 1/4)
print('Hoeffding (n=1000,t=0.05):', round(2*np.exp(-2*1000*0.05**2),4))
print('VC gap    (VC=50,n=1000) :', round(np.sqrt(50*np.log(1000)/1000),4))
printed linevalue (verified)
Markov P(X≥4)0.0183 ≤ 0.25
Chebyshev P(|X−1|≥2)0.0498 ≤ 0.25
Hoeffding (n=1000, t=0.05)0.0135
VC gap (VC=50, n=1000)0.5877

If every empirical tail sits under its bound, Hoeffding reads 0.0135, and the VC gap reads 0.5877 — you can reason rigorously about when a model generalizes.

112. Fill in: value (verified) for The full program

Comparison

Comparison matrix

From The full program: refill the value (verified) column from what you know. The rest of the table is as it appeared.

printed linevalue (verified)
Markov P(X≥4)0.0183 ≤ 0.25
Chebyshev P(|X−1|≥2)0.0498 ≤ 0.25
Hoeffding (n=1000, t=0.05)0.0135
VC gap (VC=50, n=1000)0.5877

113. Show it off

Concept

Slides closed, out loud: explain (1) the one indicator inequality that proves Markov, (2) why squaring the deviation turns Markov into Chebyshev, and (3) what 'shatter' means and why a line stops at 3 points.

Stretch: use Hoeffding to bound how far a fixed classifier's training error can drift from its true error, then re-run the double-descent sweep at p = 41 and p = 120 and confirm the spike is real — the exact worst-case-vs-practice gap the olympiad likes to probe.

114. Connect it up: Lesson 34: Concentration Inequalities & Generalization

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — The question generalization answers · Markov — from one indicator · Chebyshev — Markov on (X−μ)² · Hoeffding — the exponential engine · VC dimension — measuring capacity · The generalization bound. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

115. What you can do now

Recap

ideathe one thing to remember
MarkovP(X≥a) ≤ E[X]/a — weakest, from one indicator
ChebyshevMarkov on (X−μ)² ⇒ 1/k²
Hoeffding2e^{−2nt²} — exponential in n; invert for n
VC (line in ℝᵈ)d + 1; largest shattered set
gen bound√(VC·ln n / n) — loose; double descent beats it

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 34 (Week 12 — Generalization Theory) — Barron · USAAIO Round 2 Preparation, 2026
  2. Shalev-Shwartz & Ben-David, Understanding Machine Learning, Ch. 4 (Concentration) & Ch. 6 (VC dimension) — Cambridge University Press, 2014
  3. Mohri, Rostamizadeh & Talwalkar, Foundations of Machine Learning, Ch. 2–3 — MIT Press, 2nd ed.
  4. Belkin et al., Reconciling modern machine-learning practice and the bias–variance trade-off (double descent)
  5. Every tail probability, VC/shattering check, and double-descent MSE produced by real execution — numpy 2.2.6 + scikit-learn 1.9.0 + scipy 1.16, verification run July 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108