Uniform Convergence, No Free Lunch, and VC Dimension

Session 1b: Hoeffding's inequality and the one-over-root-m window, epsilon-representative samples and the four-inequality ERM lemma, the agnostic finite-class bound derived by re-running the five-step template with only one step changed, no-free-lunch and the bias-complexity decomposition, VC dimension with the two-half proof discipline for thresholds, intervals and rectangles, the fundamental theorem and its rates, and a five-question routine for reading a learning-theory paper.

Subject: Statistical Learning Theory · 61 slides · symbolic lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Uniform Convergence, No Free Lunch, and VC

Title

Statistical Learning Theory · Session 1b

Running the template a second time, and meeting the measure that replaces class size

2. What this half adds

Objectives

Part one built a five-step template on one theorem. A template you have seen once is an anecdote. This half runs it again on a different theorem, which is the only way to find out whether it was really a template.

Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) chapters 4-6 — the reference for this half

3. Hoeffding, and What It Buys

Section

Section 1

4. What was missing

Warm-up

Two minutes, before any new machinery.

Discussion prompt

In the realizable proof, the per-hypothesis bound was the probability that a bad hypothesis survives every sample point. Why does that argument have nothing to say once realizability is dropped?

Hint: The old bound was about a probability being exactly zero. What is the analogous quantity when it is not?

Answer:

Because it was a statement about surviving with zero mistakes, and once nothing is perfect the minimiser makes mistakes too. There is no longer an event called 'looked perfect' to bound.

What replaces it must control the gap between empirical and true risk for a hypothesis, whatever that gap happens to be. That is a two-sided statement about a sample mean, and the standard tool for it is Hoeffding's inequality.

Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) §4.1

5. Hoeffding's inequality

Concept

For independent variables each confined to a bounded range, the sample mean is exponentially unlikely to sit far from its expectation.

\[ \Pr\left[\left|\frac{1}{m}\sum_{i=1}^m Z_i - \mu\right| > \varepsilon\right] \le 2\exp\left(\frac{-2m\varepsilon^2}{(b-a)^2}\right) \]

For the zero-one loss the range has width one, so the bound is twice the exponential of minus two m epsilon squared. That is the only form used in this course.

Hoeffding, Probability Inequalities for Sums of Bounded Random Variables, Journal of the American Statistical Association 58(301), 13-30 (1963) — the original paper; the companion deck derives it from Markov via Chernoff

6. What the inequality is saying

Picture it

Three sample sizes. The centre never moves; only the width does.

Figure (svg): Three concentration curves of decreasing width around the true risk, for sample sizes twenty-five, one hundred and four hundred

The window has width about one over the square root of m.

That square root is the source of every rate in the agnostic half of the subject. If you remember one picture from this deck, this is the one.

7. Estimate the window

Estimation

No computation. Just the square root.

Predict first

A sample of 10000 points estimates the risk of one fixed hypothesis to within roughly what?

  • About 0.1
  • About 0.01
  • About 0.001
  • About 0.0001

Correct: About 0.01.

Why: The window scales as one over the square root of m, and the square root of ten thousand is one hundred, so the accuracy is on the order of one hundredth. Getting to a thousandth would need a million points, which is the practical meaning of the square-root rate.

8. Two things to notice in the exponent

Notation

This bound will appear on nearly every slide of the next month.

Annotate

On: \( 2\exp\left(-2m\varepsilon^2\right) \)

  • The leading factor of two is the price of a two-sided statement. Forgetting it is one of the most common slips in student proofs, and it propagates into the final constant.
  • Epsilon is squared. That is why inverting for m gives one over epsilon squared, whereas the realizable bound gave one over epsilon.
  • There is no dependence on the distribution anywhere, which is what makes the bound usable when the distribution is unknown.

Point three is the quiet miracle of the whole approach: a bound that holds for every distribution, without knowing any of them.

9. Uniform Convergence

Section

Section 2

10. An epsilon-representative sample

Concept

A sample is representative if the empirical risk is close to the true risk for every hypothesis in the class at once, not merely for a hypothesis chosen in advance.

\[ \forall h \in \mathcal{H}: \; |L_S(h) - L_{\mathcal{D}}(h)| \le \varepsilon \]

The words at once are doing all the work. This is a property of the sample, not of any hypothesis, and it is exactly the property that makes selection safe.

Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) Definition 4.1

11. Worked example: the ERM lemma in four inequalities

Worked example

Claim: if the sample is representative to within half epsilon, then every empirical risk minimiser is within epsilon of the best hypothesis in the class.

Start at the true risk of the chosen hypothesis, and move to its empirical risk

Why: Legal because the sample is representative, and representativeness covers every hypothesis including this one.

\[ L_{\mathcal{D}}(h_S) \le L_S(h_S) + \tfrac{\varepsilon}{2} \]

Swap the chosen hypothesis for the best one

Why: Legal because the chosen hypothesis minimises empirical risk, so its empirical risk is no larger than anyone's.

\[ \le L_S(h^{\star}) + \tfrac{\varepsilon}{2} \]

Move back from empirical to true risk for the best hypothesis

Why: Representativeness again, in the other direction.

\[ \le L_{\mathcal{D}}(h^{\star}) + \varepsilon \]

Verify: that each step used exactly one fact

Why: Representative, minimises, representative. If a step in your own proof cannot be justified in one word, it is doing two things and should be split.

12. The chain, with its justifications

Picture it

Three inequalities and three one-word reasons.

Figure (svg): A four rung ladder from the true risk of the chosen hypothesis down to the true risk of the best hypothesis plus epsilon, each rung labelled with its justification

Representative, ERM minimises, representative.

Notice where the halves come from: representativeness is used twice, so the budget is split in two. Anywhere you see a half epsilon in this subject, look for a step that gets used twice.

13. Reconstruct the middle inequality

Fill the middle

The two ends of the chain are given.

Fill in the blanks

L_L_S(h^{\star}) + \tfrac{\varepsilon}{2}}(h_S) \le L_S(h_S) + \tfrac______ \le ___ \le L____}(h^___) + \varepsilon

Why: The middle term swaps the chosen hypothesis for the best one at the same empirical risk level, which is licensed by the fact that empirical risk minimisation returns something with the smallest empirical risk. It is the only step in the chain that is not about representativeness.

14. Now re-run the template

Concept

Step one: the bad event is that the sample fails to be representative. Step two: Hoeffding for one fixed hypothesis. Step three: union bound over the class. Step four: invert.

\[ \Pr\left[\exists h: |L_S(h)-L_{\mathcal{D}}(h)| > \tfrac{\varepsilon}{2}\right] \le 2|\mathcal{H}|\exp\left(-\tfrac{m\varepsilon^2}{2}\right) \]

Only step two changed. The other three are word for word what they were in part one, which is what makes the word template honest.

Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) Corollary 4.6

15. Worked example: the agnostic finite-class bound

Worked example

Set the failure probability below delta and solve, exactly as before.

Write the bound from the union of Hoeffding statements

Why: Half epsilon goes into Hoeffding because the ERM lemma needs a half-epsilon-representative sample.

\[ 2|\mathcal{H}|\exp\left(-\tfrac{m\varepsilon^2}{2}\right) \le \delta \]

Take logarithms and rearrange

Why: The same algebra as move four in part one.

\[ m \ge \frac{2\log(2|\mathcal{H}|/\delta)}{\varepsilon^2} \]

Verify: the dependence on each parameter

Why: Logarithmic in the class size and in one over delta, as before, and now quadratic in one over epsilon. Only the epsilon dependence changed, and it changed because epsilon is squared inside Hoeffding's exponent.

16. What the square actually costs

Picture it

A hundred hypotheses, epsilon a tenth, delta a twentieth, both settings.

Figure (svg): Two bars comparing seventy-seven samples in the realizable case against one thousand six hundred and fifty-nine in the agnostic case

Seventy-seven against one thousand six hundred and fifty-nine.

A factor of about twenty-one, from one modelling assumption. That is the concrete meaning of the slogan that agnostic costs a square.

17. The two runs of the template, side by side

Comparison

Fill in the agnostic column. Three of the four rows are unchanged.

Comparison matrix

template steprealizableagnostic
bad eventsome bad hypothesis has zero empirical riskthe sample fails to be representative
one hypothesissurvives with probability at most e to the minus eps mHoeffding: two e to the minus two m eps squared
extend to the classunion boundunion bound, unchanged
invertlog over epstwo log over eps squared

One row changed. That is what it means for the five steps to be a template rather than a summary of one proof.

18. Explain the square to a classmate

Explain it

Two sentences.

Discussion prompt

Why does the agnostic case cost one over epsilon squared while the realizable case costs only one over epsilon?

Hint: Look at where epsilon sits in each exponent.

Answer:

Because in the realizable case the evidence is a run of zero mistakes, and the probability of that decays geometrically in m, giving a linear equation in m when you take logarithms.

In the agnostic case the evidence is that a sample mean is close to its expectation, and Hoeffding's exponent carries epsilon squared, so inverting gives one over epsilon squared. It is the square in the exponent, nothing deeper.

Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) §4.2

19. Trap: keeping realizability after it has been dropped

Trap

The trap

Writing an agnostic proof and then using a realizable step inside it.

\[ L_S(h_S) = 0 \]

Assert that the minimiser has zero empirical risk

Why: True only when some hypothesis is perfect, which is exactly the assumption the agnostic setting removes.

Watch the rest of the proof produce the realizable rate

Why: Getting one over epsilon in an agnostic theorem is the symptom. If your rate is too good, look for a smuggled assumption.

The fix

Keep the empirical risk of the minimiser as an unknown quantity.

\[ L_S(h_S) \le L_S(h) \quad \forall h \in \mathcal{H} \]

Use only what ERM actually guarantees

Why: That its empirical risk is minimal, not that it is zero. This is the inequality the chain in section two used.

Check the rate at the end

Why: An agnostic theorem should come out with epsilon squared. A mismatch between the setting and the rate is a reliable bug detector.

20. No Free Lunch

Section

Section 3

21. The theorem

Concept

For any learning algorithm at all, on binary classification with the zero-one loss, if the sample is at most half the size of the domain then there is a distribution that defeats it.

Defeats it precisely: the learner's true error is at least one eighth with probability at least one seventh, even though some function in play achieves zero error.

Note the order of quantifiers. The distribution is chosen after the algorithm. That is what makes the theorem possible and it is what the PAC definition deliberately forbids.

Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) Theorem 5.1

22. The adversary labels what you did not see

Picture it

At most half the domain is in the sample; the rest is free for the adversary.

Figure (svg): The domain split into a seen half where the learner has evidence and an unseen half where an adversary is free to choose the labels

With no restriction on the class, the unseen half carries no information.

The proof is an averaging argument: over all labellings of the unseen part, the learner is right half the time on average, so some labelling is at least as bad as average.

23. Why is this not a contradiction?

Socratic

It looks like it contradicts everything proved so far. It does not.

Discussion prompt

Part one proved that any finite class is learnable. No-free-lunch says learning is impossible. Reconcile them precisely.

Hint: What is the hypothesis class in each of the two statements?

Answer:

The finite-class theorem quantifies over distributions for a fixed class. No-free-lunch quantifies over distributions after the algorithm, with the class being all functions.

So the two statements are about different classes. Finite classes are learnable; the class of all functions is not.

The correct reading of no-free-lunch is therefore not 'learning is impossible' but 'learning without a restriction is impossible'. It is a theorem about the necessity of inductive bias.

Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) §5.1

24. The bias and complexity decomposition

Concept

Split the true risk of what the learner returns into two parts: how good the best member of the class is, and how far the learner falls short of it.

\[ L_{\mathcal{D}}(h_S) = \underbrace{\min_{h \in \mathcal{H}} L_{\mathcal{D}}(h)}_{\text{approximation}} + \underbrace{L_{\mathcal{D}}(h_S) - \min_{h \in \mathcal{H}} L_{\mathcal{D}}(h)}_{\text{estimation}} \]

Enlarging the class lowers the first and raises the second. That tension is the conceptual spine of the entire course, and every complexity measure later is a way of quantifying the second term.

Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) §5.2

25. The tradeoff, drawn

Picture it

Approximation falls, estimation rises, and the total has a minimum somewhere in between.

Figure (svg): Three curves against class size: a falling approximation error, a rising estimation error, and their sum with a minimum in the middle

Neither extreme is where you want to be.

Note that the position of the minimum depends on the sample size. More data flattens the estimation curve and moves the best class to the right, which is the theoretical statement behind 'bigger models need more data'.

26. Which term does each change affect?

Sorting

Keeping the two terms apart is the practical use of this decomposition.

Sort into buckets

Approximation error, or estimation error?

approximation
adding more layers to the network; the true labelling function is not in the class
estimation
collecting ten times as much data; the sample happened to be unrepresentative
appr
These are properties of the class alone, and they do not change with the data. A richer class lowers this term; a mismatch between the class and the truth raises it.
est
These are properties of the sample and the size of the class relative to it. They vanish as the sample grows, whereas approximation error does not.

The diagnostic value: if training error is already high, the approximation term is your problem and more data will not help.

27. VC Dimension, in Preview

Section

Section 4

28. Restriction and shattering

Concept

Take a finite set of points and ask which labellings of them the class can produce. That set of labellings is the restriction of the class to those points.

The set is shattered when the class produces every possible labelling of it, so the restriction has two to the power of the number of points.

VC dimension — The size of the largest set that the class shatters, or infinity if arbitrarily large sets are shattered.

Vapnik & Chervonenkis, On the Uniform Convergence of Relative Frequencies of Events to Their Probabilities, Theory of Probability and its Applications 16(2), 264-280 (1971) — the original result; UML chapter 6 for the modern presentation

29. Worked example: thresholds have VC dimension one

Worked example

Take the class of thresholds that predict one to the right of a cut point and zero to the left.

Lower half: exhibit one shattered point

Why: A single point can be labelled zero by putting the cut to its right, and one by putting the cut to its left. Both labellings are achieved, so a set of size one is shattered.

Upper half: show no pair is shattered

Why: Take two points with the first to the left of the second. Ask for the labelling one then zero.

Derive the contradiction

Why: Labelling the left point one puts the cut to its left; labelling the right point zero puts the cut to its right. The cut cannot be on both sides of a point.

\[ \mathrm{VCdim} = 1 \]

Verify: that both halves were actually done

Why: One construction and one impossibility. A VC argument with only the first half proves a lower bound and nothing more, and that omission is the commonest way these exercises lose marks.

30. Predict the VC dimension

Prediction

Axis-aligned rectangles in the plane, labelling a point one when it is inside the rectangle.

Predict first

What is the VC dimension?

  • 2
  • 3
  • 4
  • infinite

Correct: 4.

Why: Four points arranged as a diamond can be shattered: for any subset, the smallest axis-aligned rectangle containing that subset excludes the others. For five points, take the leftmost, rightmost, topmost and bottommost; the fifth lies inside their bounding box, so any rectangle containing those four contains it, and the labelling that excludes only the fifth is impossible.

31. Four yes, five no

Picture it

The two halves of the rectangle argument, side by side.

Figure (svg): On the left four points in a diamond that a rectangle can shatter, on the right five points where the innermost cannot be excluded

Construct a shattered set of size d; then rule out every set of size d plus one.

That two-half shape is the pattern for every VC computation you will be asked for. Write both halves, always.

32. Match the class to its VC dimension

Matching

The four you are most likely to meet in the first homework.

Match the pairs

  • l1. thresholds on the real line
  • l2. intervals on the real line
  • l3. axis-aligned rectangles in the plane
  • l4. halfspaces in d dimensions
  • r1. 1
  • r2. 2
  • r3. 4
  • r4. d plus one

Why: The last one is proved with Radon's theorem, which says any d plus two points in d dimensions can be split into two sets whose convex hulls intersect — and intersecting hulls cannot be separated by a halfspace. Worth knowing the name even before you see the proof.

33. Pattern: proving any VC claim

Pattern

Two halves, always, and they are different kinds of argument.

  1. Lower bound. Construct one specific set of size d and exhibit, or argue for, all two-to-the-d labellings. A construction, so be concrete.
  2. Upper bound. Take an ARBITRARY set of size d plus one and show some labelling is unreachable. A case analysis or a geometric obstruction, and it must cover every arrangement.
  3. Do not swap them. Exhibiting one unshatterable set of size d plus one proves nothing; you must handle all of them.
  4. State the answer as an equality only once both halves are written.

The asymmetry between the two halves — one existential, one universal — is exactly where students lose marks, and it is visible in the definition if you read the quantifiers.

Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) chapter 6

34. Check: what does the upper half require?

Check

Think about the quantifiers.

Check your understanding

To prove the VC dimension of a class is at most 3, what must you show?

  • A. That every set of four points fails to be shattered (correct)
  • B. That some set of four points fails to be shattered
  • C. That some set of three points is shattered
  • D. That every set of three points is shattered

Answer: A

Why: The VC dimension is the size of the largest shattered set, so bounding it above by three means no set of size four is shattered — a universal claim over all four-point sets.

Why B tempts people
One unshatterable four-point set says nothing; there might be a different four-point set that is shattered, which would make the dimension at least four.
Why C tempts people
That is the lower half, and it proves the dimension is at least three rather than at most.
Why D tempts people
Far stronger than needed and usually false. The definition asks only for the existence of one shattered set of the given size.

35. The fundamental theorem

Concept

For binary classification with the zero-one loss, five statements about a class are all equivalent: it has the uniform convergence property, empirical risk minimisation is a successful learner for it, it is PAC learnable, it is agnostic PAC learnable, and its VC dimension is finite.

That is an unusually strong result. It says the single number you computed for rectangles decides everything.

Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) Theorem 6.7

36. Five statements, one condition

Picture it

Any one of them implies all the others.

Figure (svg): Five boxes arranged in a ring, labelled uniform convergence, ERM succeeds, PAC learnable, agnostic PAC learnable and finite VC dimension, all marked equivalent

Finite VC dimension is the computable one, which is why it is the one you check.

37. The quantitative version

Concept

The theorem also gives rates, with the VC dimension standing exactly where the logarithm of the class size stood in part one.

\[ m^{\text{agnostic}} = \Theta\!\left(\frac{d + \log(1/\delta)}{\varepsilon^2}\right) \]

\[ m^{\text{realizable}} = O\!\left(\frac{d\log(1/\varepsilon) + \log(1/\delta)}{\varepsilon}\right) \]

Note the extra logarithmic factor in the realizable rate. It is often suppressed in lecture notes, and it is genuinely there in the standard upper bound.

Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) Theorem 6.8

38. Compare the two theorems term by term

Analogy

The finite-class bound from part one against the VC bound.

Match the pairs

  • l1. log of the class size
  • l2. the union bound over the class
  • l3. finiteness of the class
  • l4. one over eps squared, agnostic
  • r1. the VC dimension d
  • r2. the growth function, via Sauer and Shelah
  • r3. finiteness of the VC dimension
  • r4. one over eps squared, unchanged

Why: The last row is the point: moving to infinite classes changes what you pay for complexity, and changes nothing about the dependence on accuracy. The epsilon rate comes from the concentration inequality, and that did not change.

39. The succession of complexity measures

Picture it

Five names, in the order the course will introduce them.

Figure (svg): A five step progression from class size through the growth function and Sauer Shelah to symmetrization and Rademacher complexity

Each one replaces the previous with something smaller or more data dependent.

You do not need any of them today. You need to recognise them when the professor says them, and to know that each is answering the same question: what should stand in for the size of the class?

40. Which is the honest summary of the fundamental theorem?

Elimination

One sentence, for your own notes.

Eliminate the wrong options

Which statement would you write down?

  • A. For binary classification with zero-one loss, a class is learnable exactly when its VC dimension is finite.
  • B. Every class with finite VC dimension is learnable by any algorithm.
  • C. Finite VC dimension implies learnability, and this is true for every loss function.
  • D. A class is learnable exactly when it is finite.

Survives elimination: A

Why: The two qualifications that matter are the setting, binary classification with zero-one loss, and the fact that the successful learner named is ERM. Dropping either turns a true statement into a false one.

41. Worked example: intervals, both halves

Worked example

Prove that the class of intervals on the real line has VC dimension exactly two.

Lower half: construct a shattered pair

Why: Take the points 0 and 1. The empty interval gives zero zero, a small interval around 0 gives one zero, a small one around 1 gives zero one, and a wide one gives one one.

\[ 2^2 = 4 \text{ labellings, all realised} \]

Upper half: take an arbitrary triple

Why: Not a convenient one. Any three points with the first strictly left of the second and the second of the third.

Exhibit the unreachable labelling

Why: Ask for one, zero, one. An interval containing the outer two contains everything between them, hence the middle point.

\[ \mathrm{VCdim} = 2 \]

Verify: that the upper half really was arbitrary

Why: Nothing in the argument used the positions of the three points, only their order. If your upper half needs specific coordinates, it is not yet a proof.

42. Intervals: the labelling that cannot happen

Picture it

Three points, asked for one, zero, one.

Figure (svg): Three points on a line labelled one, zero, one, with a shaded interval showing that covering both outer points forces covering the middle

Convexity of an interval is the whole obstruction.

The same argument generalises: a class of convex sets can never separate a point that lies between two positives.

43. Find the flaw in this VC argument

Error analysis

A student's attempt to show the VC dimension of intervals is at most two.

Annotate

On: \( \text{The points } 0,1,2 \text{ cannot be shattered, so } \mathrm{VCdim} \le 2. \)

  • One specific triple is exhibited, and it genuinely cannot be shattered.
  • But the upper half is a universal claim: NO set of three points may be shattered. Exhibiting one such set proves nothing about the others.
  • The repair is to take an arbitrary ordered triple and argue from the order alone, which is what the previous slide did.

This is the single most common error in first VC homeworks, and it is invisible unless you read the quantifier in the definition carefully.

44. Which half does each statement establish?

Discrimination

Sorting these correctly is most of the discipline.

Sort into buckets

Lower bound on the VC dimension, or upper bound?

lower bound
here is a set of size 4 and all sixteen of its labellings; these three specific points can be shattered
upper bound
any set of size 5 has a point inside the others' bounding box; for every ordered triple, the labelling one zero one is unreachable
low
An existence claim: some set of that size is shattered. One concrete construction is enough, and it should be concrete.
up
A universal claim: no set of that size is shattered. It must cover every arrangement, so it argues from a structural property rather than from coordinates.

45. What replaces the class size: counting behaviours

Concept

An infinite class can still produce only finitely many labellings of a finite sample. The growth function counts those labellings, at worst, over all samples of a given size.

\[ \tau_{\mathcal{H}}(m) = \max_{|C| = m} \left|\mathcal{H}_C\right| \]

For thresholds it is m plus one; for intervals it is the number of ways to choose two gaps out of m plus one, plus the empty labelling. Both are polynomial, and neither is anything like the size of the class.

Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) §6.5

46. Infinitely many hypotheses, four behaviours

Picture it

Three points, and every threshold that exists.

Figure (svg): Three points on a line shown with four different threshold positions, producing only four distinct labellings

Only the gaps between points matter, so the count is m plus one.

This is the observation that makes infinite classes tractable, and Sauer and Shelah turn it into a general bound: the growth function is at most a polynomial of degree the VC dimension.

47. Why does counting behaviours help at all?

Socratic

It is not obvious that this rescues the union bound.

Discussion prompt

The union bound was over hypotheses. The growth function counts labellings of a sample. How can one replace the other, when the bound has to hold before the sample is drawn?

Hint: What has to happen before the quantity of interest depends only on finitely many points?

Answer:

It cannot directly, and that gap is exactly what symmetrization repairs. You first replace the true risk by the empirical risk on a second, independent ghost sample.

After that swap, the whole event depends only on the labels the class assigns to 2m points, so two hypotheses that agree on those points are interchangeable. Now you may union bound over behaviours instead of hypotheses.

So the order is: symmetrize first, then count. Getting that order backwards is why the argument looks circular the first time you meet it.

Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) §6.5

48. Symmetrization, in four lines

Picture it

The trick that lets a finite count replace an infinite class.

Figure (svg): A four step flow from wanting to compare against the true risk, through drawing an independent ghost sample, to everything depending on two m points

Ghost sample first; counting second.

You will see this proof in full within a couple of weeks. Knowing its shape now means you will be following the argument rather than transcribing it.

49. What happens as the VC dimension grows?

Edge cases

Push the parameter and see what the theory says.

Discussion prompt

The agnostic rate is the VC dimension plus a log term, all over epsilon squared. What does that predict as the dimension grows towards infinity, and does that match no-free-lunch?

Hint: What is the VC dimension of the class of all binary functions on an infinite domain?

Answer:

The required sample grows linearly in the dimension, so an infinite VC dimension means no finite sample suffices, for any epsilon and delta.

That is exactly no-free-lunch, restated quantitatively. The class of all functions on an infinite domain has infinite VC dimension, and the fundamental theorem then says it is not learnable.

So the two results are not separate facts. No-free-lunch is the qualitative shadow of the lower bound in the fundamental theorem.

Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) Theorem 6.7

50. Which claim about no-free-lunch survives?

Two truths and a lie

Three readings of the theorem, one of them defensible.

Eliminate the wrong options

Which is right?

  • A. It shows that some restriction on the hypothesis class is necessary for learning to be possible at all.
  • B. It shows that no learning algorithm is better than any other.
  • C. It shows that machine learning cannot work in practice.
  • D. It applies only to the zero-one loss and so is a technicality.

Survives elimination: A

Why: The theorem is best read as a statement about what the definition of learning must include, rather than as a pessimistic result. It is the formal reason the hypothesis class appears in every theorem in the course.

51. How sure are you?

Commit first

Answer, then rate your confidence.

Predict first

A class has VC dimension 5. Is it PAC learnable?

  • Yes, and the fundamental theorem gives the rate
  • Only if it is also finite
  • Only in the realizable case
  • Not enough information; it depends on the distribution

Correct: Yes, and the fundamental theorem gives the rate.

Why: Finite VC dimension is equivalent to learnability for binary classification with the zero-one loss, in both the realizable and agnostic settings, and the rate is the dimension plus a log term over epsilon squared. The fourth option inverts the definition: PAC learnability already quantifies over all distributions, so no distribution-specific information is needed or allowed.

52. Order the paper-reading routine

Ranking

Five questions, and the order matters more than it looks.

Put in order

  1. Identify the setting: domain, labels, loss, hypothesis class
  2. List the assumptions and mark which are unusual
  3. Write the main theorem in one sentence, with its rate
  4. Decide which template step carries the new idea
  5. Ask what breaks if each assumption is dropped

Why: The setting first, because the theorem statement is unreadable without it. The rate before the new idea, because an unusually good rate is the clue that points at which assumption is doing the work. And the last question is what turns reading into understanding, so it cannot come first.

53. Reading a Paper

Section

Section 5

54. Your final project will be a paper, so build the routine now

Concept

A statistical learning paper is not read front to back. Five questions, answered in order, get you most of the way, and they reuse the template you already have.

Bousquet, Boucheron & Lugosi, Introduction to Statistical Learning Theory — a good second voice, and short enough to practise the routine on

55. The five-question routine

Picture it

Answer these on one page before reading any proof in detail.

Figure (svg): A five step flow through the setting, the assumptions, the main theorem and its rate, the template step the paper innovates on, and what breaks without each assumption

The fourth question is the one that tells you what the paper is actually for.

Most papers innovate on exactly one of the five template steps. Identifying which one turns a forty-page paper into one idea plus a lot of bookkeeping.

56. Practise the routine on a paper you have not read

Step zero

Pick anything from your professor's reading list.

Discussion prompt

Without reading past the abstract and the theorem statements, answer: what is the setting, what are the assumptions, and what is the rate?

Hint: You are allowed to read only the abstract and the theorem statements for this exercise.

Answer:

The setting is almost always in the first two paragraphs of section two, and it is usually four objects: domain, labels, loss, class.

The assumptions are frequently in a displayed list, and the interesting ones are the ones a reader would not have guessed: bounded loss, sub-Gaussian noise, a margin condition.

The rate is in the abstract, and it is worth writing down as a formula rather than as words, so you can compare it against the rates you already know.

If the rate is better than the standard one over the square root of m, some assumption is buying that, and finding which one is the most useful single thing you can do with the paper.

Mohri, Rostamizadeh & Talwalkar, Foundations of Machine Learning, 2nd edition — its bibliographic notes are a good map of which paper did what

57. Pattern: five mistakes to watch for in your own proofs

Pattern

Four of these appeared in part one as traps. Keep the list beside you while writing homework.

mistakehow to catch it
applying concentration to the hypothesis the algorithm chosethe bound has no complexity term in it
dropping the factor of two on a two-sided tailcompare against the one-sided version and see if it halved
confusing probability over the sample with probability over a test pointname the random object at every probability sign
using realizability after the setting dropped itan agnostic theorem coming out with a one over eps rate
union bounding over an infinite classthe bound is infinite, hence vacuous
  1. Name the random object at every probability you write.
  2. Check the rate against the setting before believing a result.
  3. Ask what the proof never used, and you will find the generalisation.

Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) chapters 2-6

58. Check: diagnose the bug

Check

A classmate's agnostic bound.

Check your understanding

A student proves an agnostic bound and gets m at least log of the class size over delta, divided by epsilon. What is almost certainly wrong?

  • A. They used a realizable step; an agnostic rate should be one over epsilon squared (correct)
  • B. They forgot the union bound
  • C. They used the wrong logarithm base
  • D. Nothing; that is the correct agnostic rate

Answer: A

Why: The epsilon dependence is the fingerprint of which concentration argument was used. A linear dependence on one over epsilon comes from the zero-mistakes argument, which is only available under realizability.

Why B tempts people
Forgetting the union bound would remove the class size from the bound entirely, and the class size is present here.
Why C tempts people
The base of the logarithm changes only a constant factor and never the exponent on epsilon.
Why D tempts people
That is the realizable rate. An agnostic theorem cannot achieve it in general, and the fundamental theorem's lower bound says so.

59. Exit ticket

Exit ticket

One honest answer, and it sets the agenda for session two.

Predict first

Which of these would you least want to be asked to reproduce on Tuesday?

  • Hoeffding's inequality and what it gives you
  • The four-inequality ERM chain
  • Re-running the template for the agnostic bound
  • No-free-lunch and why it is not a contradiction
  • A VC dimension computation with both halves
  • The fundamental theorem and its rates

Correct: Whichever you named opens the next session.

Why: If it is the fifth, that is the most likely thing to appear on the first homework, and the two-half discipline is drillable in about twenty minutes across the four standard classes.

60. One page, both halves of the session

Connect it up

Twenty minutes before Tuesday.

Draw it

On one page: the five template steps down the left, and two columns to their right — how the realizable proof performs each step, and how the agnostic proof does. Then, underneath, the four VC classes with their dimensions and the one-line obstruction for each upper half.

That page is the whole of session one. Bring it, and bring the questions it exposes.

61. Session one, complete

Recap

The template held up under a second run, which is the evidence that it is worth memorising.

resultrate or value
Hoeffding, zero-one losstwo e to the minus two m eps squared
finite class, realizablelog of class over delta, over eps
finite class, agnostictwice log of twice class over delta, over eps squared
numeric comparison at eps 0.177 against 1659
thresholds, intervals, rectangles, halfspaces1, 2, 4, and d plus one
fundamental theorem, agnosticd plus log one over delta, over eps squared

Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) chapters 4-6 — read these before session two, and Bousquet sections 1 to 3 as a second voice

Sources

  1. Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors)
  2. Bousquet, Boucheron & Lugosi, Introduction to Statistical Learning Theory — Advanced Lectures on Machine Learning, Springer LNCS 3176 (2004), pp. 169-207
  3. Mohri, Rostamizadeh & Talwalkar, Foundations of Machine Learning, 2nd edition — MIT Press, 2018
  4. Hoeffding, Probability Inequalities for Sums of Bounded Random Variables, Journal of the American Statistical Association 58(301), 13-30 (1963)
  5. Vapnik & Chervonenkis, On the Uniform Convergence of Relative Frequencies of Events to Their Probabilities, Theory of Probability and its Applications 16(2), 264-280 (1971)
  6. Companion deck in this library: Concentration Inequalities and Generalization, which derives Markov, Chebyshev and Hoeffding with no steps skipped — slides/decks/usaaio/lesson-34-concentration-generalization.json

Want this taught 1-on-1? Alexander tutors Statistical Learning Theory — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108