Session 1b: Hoeffding's inequality and the one-over-root-m window, epsilon-representative samples and the four-inequality ERM lemma, the agnostic finite-class bound derived by re-running the five-step template with only one step changed, no-free-lunch and the bias-complexity decomposition, VC dimension with the two-half proof discipline for thresholds, intervals and rectangles, the fundamental theorem and its rates, and a five-question routine for reading a learning-theory paper.
Subject: Statistical Learning Theory · 61 slides · symbolic lesson
Open the interactive version of this deck · Homework for this lesson
Title
Statistical Learning Theory · Session 1b
Running the template a second time, and meeting the measure that replaces class size
Objectives
Part one built a five-step template on one theorem. A template you have seen once is an anecdote. This half runs it again on a different theorem, which is the only way to find out whether it was really a template.
Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) chapters 4-6 — the reference for this half
Section
Section 1
Warm-up
Two minutes, before any new machinery.
Discussion prompt
In the realizable proof, the per-hypothesis bound was the probability that a bad hypothesis survives every sample point. Why does that argument have nothing to say once realizability is dropped?
Hint: The old bound was about a probability being exactly zero. What is the analogous quantity when it is not?
Answer:
Because it was a statement about surviving with zero mistakes, and once nothing is perfect the minimiser makes mistakes too. There is no longer an event called 'looked perfect' to bound.
What replaces it must control the gap between empirical and true risk for a hypothesis, whatever that gap happens to be. That is a two-sided statement about a sample mean, and the standard tool for it is Hoeffding's inequality.
Concept
For independent variables each confined to a bounded range, the sample mean is exponentially unlikely to sit far from its expectation.
\[ \Pr\left[\left|\frac{1}{m}\sum_{i=1}^m Z_i - \mu\right| > \varepsilon\right] \le 2\exp\left(\frac{-2m\varepsilon^2}{(b-a)^2}\right) \]
For the zero-one loss the range has width one, so the bound is twice the exponential of minus two m epsilon squared. That is the only form used in this course.
Hoeffding, Probability Inequalities for Sums of Bounded Random Variables, Journal of the American Statistical Association 58(301), 13-30 (1963) — the original paper; the companion deck derives it from Markov via Chernoff
Picture it
Three sample sizes. The centre never moves; only the width does.
Figure (svg): Three concentration curves of decreasing width around the true risk, for sample sizes twenty-five, one hundred and four hundred
That square root is the source of every rate in the agnostic half of the subject. If you remember one picture from this deck, this is the one.
Estimation
No computation. Just the square root.
Predict first
A sample of 10000 points estimates the risk of one fixed hypothesis to within roughly what?
Correct: About 0.01.
Why: The window scales as one over the square root of m, and the square root of ten thousand is one hundred, so the accuracy is on the order of one hundredth. Getting to a thousandth would need a million points, which is the practical meaning of the square-root rate.
Notation
This bound will appear on nearly every slide of the next month.
Annotate
On: \( 2\exp\left(-2m\varepsilon^2\right) \)
Point three is the quiet miracle of the whole approach: a bound that holds for every distribution, without knowing any of them.
Section
Section 2
Concept
A sample is representative if the empirical risk is close to the true risk for every hypothesis in the class at once, not merely for a hypothesis chosen in advance.
\[ \forall h \in \mathcal{H}: \; |L_S(h) - L_{\mathcal{D}}(h)| \le \varepsilon \]
The words at once are doing all the work. This is a property of the sample, not of any hypothesis, and it is exactly the property that makes selection safe.
Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) Definition 4.1
Worked example
Claim: if the sample is representative to within half epsilon, then every empirical risk minimiser is within epsilon of the best hypothesis in the class.
Start at the true risk of the chosen hypothesis, and move to its empirical risk
Why: Legal because the sample is representative, and representativeness covers every hypothesis including this one.
\[ L_{\mathcal{D}}(h_S) \le L_S(h_S) + \tfrac{\varepsilon}{2} \]
Swap the chosen hypothesis for the best one
Why: Legal because the chosen hypothesis minimises empirical risk, so its empirical risk is no larger than anyone's.
\[ \le L_S(h^{\star}) + \tfrac{\varepsilon}{2} \]
Move back from empirical to true risk for the best hypothesis
Why: Representativeness again, in the other direction.
\[ \le L_{\mathcal{D}}(h^{\star}) + \varepsilon \]
Verify: that each step used exactly one fact
Why: Representative, minimises, representative. If a step in your own proof cannot be justified in one word, it is doing two things and should be split.
Picture it
Three inequalities and three one-word reasons.
Figure (svg): A four rung ladder from the true risk of the chosen hypothesis down to the true risk of the best hypothesis plus epsilon, each rung labelled with its justification
Notice where the halves come from: representativeness is used twice, so the budget is split in two. Anywhere you see a half epsilon in this subject, look for a step that gets used twice.
Fill the middle
The two ends of the chain are given.
Fill in the blanks
L_L_S(h^{\star}) + \tfrac{\varepsilon}{2}}(h_S) \le L_S(h_S) + \tfrac______ \le ___ \le L____}(h^___) + \varepsilon
Why: The middle term swaps the chosen hypothesis for the best one at the same empirical risk level, which is licensed by the fact that empirical risk minimisation returns something with the smallest empirical risk. It is the only step in the chain that is not about representativeness.
Concept
Step one: the bad event is that the sample fails to be representative. Step two: Hoeffding for one fixed hypothesis. Step three: union bound over the class. Step four: invert.
\[ \Pr\left[\exists h: |L_S(h)-L_{\mathcal{D}}(h)| > \tfrac{\varepsilon}{2}\right] \le 2|\mathcal{H}|\exp\left(-\tfrac{m\varepsilon^2}{2}\right) \]
Only step two changed. The other three are word for word what they were in part one, which is what makes the word template honest.
Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) Corollary 4.6
Worked example
Set the failure probability below delta and solve, exactly as before.
Write the bound from the union of Hoeffding statements
Why: Half epsilon goes into Hoeffding because the ERM lemma needs a half-epsilon-representative sample.
\[ 2|\mathcal{H}|\exp\left(-\tfrac{m\varepsilon^2}{2}\right) \le \delta \]
Take logarithms and rearrange
Why: The same algebra as move four in part one.
\[ m \ge \frac{2\log(2|\mathcal{H}|/\delta)}{\varepsilon^2} \]
Verify: the dependence on each parameter
Why: Logarithmic in the class size and in one over delta, as before, and now quadratic in one over epsilon. Only the epsilon dependence changed, and it changed because epsilon is squared inside Hoeffding's exponent.
Picture it
A hundred hypotheses, epsilon a tenth, delta a twentieth, both settings.
Figure (svg): Two bars comparing seventy-seven samples in the realizable case against one thousand six hundred and fifty-nine in the agnostic case
A factor of about twenty-one, from one modelling assumption. That is the concrete meaning of the slogan that agnostic costs a square.
Comparison
Fill in the agnostic column. Three of the four rows are unchanged.
Comparison matrix
| template step | realizable | agnostic |
|---|---|---|
| bad event | some bad hypothesis has zero empirical risk | the sample fails to be representative |
| one hypothesis | survives with probability at most e to the minus eps m | Hoeffding: two e to the minus two m eps squared |
| extend to the class | union bound | union bound, unchanged |
| invert | log over eps | two log over eps squared |
One row changed. That is what it means for the five steps to be a template rather than a summary of one proof.
Explain it
Two sentences.
Discussion prompt
Why does the agnostic case cost one over epsilon squared while the realizable case costs only one over epsilon?
Hint: Look at where epsilon sits in each exponent.
Answer:
Because in the realizable case the evidence is a run of zero mistakes, and the probability of that decays geometrically in m, giving a linear equation in m when you take logarithms.
In the agnostic case the evidence is that a sample mean is close to its expectation, and Hoeffding's exponent carries epsilon squared, so inverting gives one over epsilon squared. It is the square in the exponent, nothing deeper.
Trap
Writing an agnostic proof and then using a realizable step inside it.
\[ L_S(h_S) = 0 \]
Assert that the minimiser has zero empirical risk
Why: True only when some hypothesis is perfect, which is exactly the assumption the agnostic setting removes.
Watch the rest of the proof produce the realizable rate
Why: Getting one over epsilon in an agnostic theorem is the symptom. If your rate is too good, look for a smuggled assumption.
Keep the empirical risk of the minimiser as an unknown quantity.
\[ L_S(h_S) \le L_S(h) \quad \forall h \in \mathcal{H} \]
Use only what ERM actually guarantees
Why: That its empirical risk is minimal, not that it is zero. This is the inequality the chain in section two used.
Check the rate at the end
Why: An agnostic theorem should come out with epsilon squared. A mismatch between the setting and the rate is a reliable bug detector.
Section
Section 3
Concept
For any learning algorithm at all, on binary classification with the zero-one loss, if the sample is at most half the size of the domain then there is a distribution that defeats it.
Defeats it precisely: the learner's true error is at least one eighth with probability at least one seventh, even though some function in play achieves zero error.
Note the order of quantifiers. The distribution is chosen after the algorithm. That is what makes the theorem possible and it is what the PAC definition deliberately forbids.
Picture it
At most half the domain is in the sample; the rest is free for the adversary.
Figure (svg): The domain split into a seen half where the learner has evidence and an unseen half where an adversary is free to choose the labels
The proof is an averaging argument: over all labellings of the unseen part, the learner is right half the time on average, so some labelling is at least as bad as average.
Socratic
It looks like it contradicts everything proved so far. It does not.
Discussion prompt
Part one proved that any finite class is learnable. No-free-lunch says learning is impossible. Reconcile them precisely.
Hint: What is the hypothesis class in each of the two statements?
Answer:
The finite-class theorem quantifies over distributions for a fixed class. No-free-lunch quantifies over distributions after the algorithm, with the class being all functions.
So the two statements are about different classes. Finite classes are learnable; the class of all functions is not.
The correct reading of no-free-lunch is therefore not 'learning is impossible' but 'learning without a restriction is impossible'. It is a theorem about the necessity of inductive bias.
Concept
Split the true risk of what the learner returns into two parts: how good the best member of the class is, and how far the learner falls short of it.
\[ L_{\mathcal{D}}(h_S) = \underbrace{\min_{h \in \mathcal{H}} L_{\mathcal{D}}(h)}_{\text{approximation}} + \underbrace{L_{\mathcal{D}}(h_S) - \min_{h \in \mathcal{H}} L_{\mathcal{D}}(h)}_{\text{estimation}} \]
Enlarging the class lowers the first and raises the second. That tension is the conceptual spine of the entire course, and every complexity measure later is a way of quantifying the second term.
Picture it
Approximation falls, estimation rises, and the total has a minimum somewhere in between.
Figure (svg): Three curves against class size: a falling approximation error, a rising estimation error, and their sum with a minimum in the middle
Note that the position of the minimum depends on the sample size. More data flattens the estimation curve and moves the best class to the right, which is the theoretical statement behind 'bigger models need more data'.
Sorting
Keeping the two terms apart is the practical use of this decomposition.
Sort into buckets
Approximation error, or estimation error?
The diagnostic value: if training error is already high, the approximation term is your problem and more data will not help.
Section
Section 4
Concept
Take a finite set of points and ask which labellings of them the class can produce. That set of labellings is the restriction of the class to those points.
The set is shattered when the class produces every possible labelling of it, so the restriction has two to the power of the number of points.
VC dimension — The size of the largest set that the class shatters, or infinity if arbitrarily large sets are shattered.
Vapnik & Chervonenkis, On the Uniform Convergence of Relative Frequencies of Events to Their Probabilities, Theory of Probability and its Applications 16(2), 264-280 (1971) — the original result; UML chapter 6 for the modern presentation
Worked example
Take the class of thresholds that predict one to the right of a cut point and zero to the left.
Lower half: exhibit one shattered point
Why: A single point can be labelled zero by putting the cut to its right, and one by putting the cut to its left. Both labellings are achieved, so a set of size one is shattered.
Upper half: show no pair is shattered
Why: Take two points with the first to the left of the second. Ask for the labelling one then zero.
Derive the contradiction
Why: Labelling the left point one puts the cut to its left; labelling the right point zero puts the cut to its right. The cut cannot be on both sides of a point.
\[ \mathrm{VCdim} = 1 \]
Verify: that both halves were actually done
Why: One construction and one impossibility. A VC argument with only the first half proves a lower bound and nothing more, and that omission is the commonest way these exercises lose marks.
Prediction
Axis-aligned rectangles in the plane, labelling a point one when it is inside the rectangle.
Predict first
What is the VC dimension?
Correct: 4.
Why: Four points arranged as a diamond can be shattered: for any subset, the smallest axis-aligned rectangle containing that subset excludes the others. For five points, take the leftmost, rightmost, topmost and bottommost; the fifth lies inside their bounding box, so any rectangle containing those four contains it, and the labelling that excludes only the fifth is impossible.
Picture it
The two halves of the rectangle argument, side by side.
Figure (svg): On the left four points in a diamond that a rectangle can shatter, on the right five points where the innermost cannot be excluded
That two-half shape is the pattern for every VC computation you will be asked for. Write both halves, always.
Matching
The four you are most likely to meet in the first homework.
Match the pairs
Why: The last one is proved with Radon's theorem, which says any d plus two points in d dimensions can be split into two sets whose convex hulls intersect — and intersecting hulls cannot be separated by a halfspace. Worth knowing the name even before you see the proof.
Pattern
Two halves, always, and they are different kinds of argument.
The asymmetry between the two halves — one existential, one universal — is exactly where students lose marks, and it is visible in the definition if you read the quantifiers.
Check
Think about the quantifiers.
Check your understanding
To prove the VC dimension of a class is at most 3, what must you show?
Answer: A
Why: The VC dimension is the size of the largest shattered set, so bounding it above by three means no set of size four is shattered — a universal claim over all four-point sets.
Concept
For binary classification with the zero-one loss, five statements about a class are all equivalent: it has the uniform convergence property, empirical risk minimisation is a successful learner for it, it is PAC learnable, it is agnostic PAC learnable, and its VC dimension is finite.
That is an unusually strong result. It says the single number you computed for rectangles decides everything.
Picture it
Any one of them implies all the others.
Figure (svg): Five boxes arranged in a ring, labelled uniform convergence, ERM succeeds, PAC learnable, agnostic PAC learnable and finite VC dimension, all marked equivalent
Concept
The theorem also gives rates, with the VC dimension standing exactly where the logarithm of the class size stood in part one.
\[ m^{\text{agnostic}} = \Theta\!\left(\frac{d + \log(1/\delta)}{\varepsilon^2}\right) \]
\[ m^{\text{realizable}} = O\!\left(\frac{d\log(1/\varepsilon) + \log(1/\delta)}{\varepsilon}\right) \]
Note the extra logarithmic factor in the realizable rate. It is often suppressed in lecture notes, and it is genuinely there in the standard upper bound.
Analogy
The finite-class bound from part one against the VC bound.
Match the pairs
Why: The last row is the point: moving to infinite classes changes what you pay for complexity, and changes nothing about the dependence on accuracy. The epsilon rate comes from the concentration inequality, and that did not change.
Picture it
Five names, in the order the course will introduce them.
Figure (svg): A five step progression from class size through the growth function and Sauer Shelah to symmetrization and Rademacher complexity
You do not need any of them today. You need to recognise them when the professor says them, and to know that each is answering the same question: what should stand in for the size of the class?
Elimination
One sentence, for your own notes.
Eliminate the wrong options
Which statement would you write down?
Survives elimination: A
Why: The two qualifications that matter are the setting, binary classification with zero-one loss, and the fact that the successful learner named is ERM. Dropping either turns a true statement into a false one.
Worked example
Prove that the class of intervals on the real line has VC dimension exactly two.
Lower half: construct a shattered pair
Why: Take the points 0 and 1. The empty interval gives zero zero, a small interval around 0 gives one zero, a small one around 1 gives zero one, and a wide one gives one one.
\[ 2^2 = 4 \text{ labellings, all realised} \]
Upper half: take an arbitrary triple
Why: Not a convenient one. Any three points with the first strictly left of the second and the second of the third.
Exhibit the unreachable labelling
Why: Ask for one, zero, one. An interval containing the outer two contains everything between them, hence the middle point.
\[ \mathrm{VCdim} = 2 \]
Verify: that the upper half really was arbitrary
Why: Nothing in the argument used the positions of the three points, only their order. If your upper half needs specific coordinates, it is not yet a proof.
Picture it
Three points, asked for one, zero, one.
Figure (svg): Three points on a line labelled one, zero, one, with a shaded interval showing that covering both outer points forces covering the middle
The same argument generalises: a class of convex sets can never separate a point that lies between two positives.
Error analysis
A student's attempt to show the VC dimension of intervals is at most two.
Annotate
On: \( \text{The points } 0,1,2 \text{ cannot be shattered, so } \mathrm{VCdim} \le 2. \)
This is the single most common error in first VC homeworks, and it is invisible unless you read the quantifier in the definition carefully.
Discrimination
Sorting these correctly is most of the discipline.
Sort into buckets
Lower bound on the VC dimension, or upper bound?
Concept
An infinite class can still produce only finitely many labellings of a finite sample. The growth function counts those labellings, at worst, over all samples of a given size.
\[ \tau_{\mathcal{H}}(m) = \max_{|C| = m} \left|\mathcal{H}_C\right| \]
For thresholds it is m plus one; for intervals it is the number of ways to choose two gaps out of m plus one, plus the empty labelling. Both are polynomial, and neither is anything like the size of the class.
Picture it
Three points, and every threshold that exists.
Figure (svg): Three points on a line shown with four different threshold positions, producing only four distinct labellings
This is the observation that makes infinite classes tractable, and Sauer and Shelah turn it into a general bound: the growth function is at most a polynomial of degree the VC dimension.
Socratic
It is not obvious that this rescues the union bound.
Discussion prompt
The union bound was over hypotheses. The growth function counts labellings of a sample. How can one replace the other, when the bound has to hold before the sample is drawn?
Hint: What has to happen before the quantity of interest depends only on finitely many points?
Answer:
It cannot directly, and that gap is exactly what symmetrization repairs. You first replace the true risk by the empirical risk on a second, independent ghost sample.
After that swap, the whole event depends only on the labels the class assigns to 2m points, so two hypotheses that agree on those points are interchangeable. Now you may union bound over behaviours instead of hypotheses.
So the order is: symmetrize first, then count. Getting that order backwards is why the argument looks circular the first time you meet it.
Picture it
The trick that lets a finite count replace an infinite class.
Figure (svg): A four step flow from wanting to compare against the true risk, through drawing an independent ghost sample, to everything depending on two m points
You will see this proof in full within a couple of weeks. Knowing its shape now means you will be following the argument rather than transcribing it.
Edge cases
Push the parameter and see what the theory says.
Discussion prompt
The agnostic rate is the VC dimension plus a log term, all over epsilon squared. What does that predict as the dimension grows towards infinity, and does that match no-free-lunch?
Hint: What is the VC dimension of the class of all binary functions on an infinite domain?
Answer:
The required sample grows linearly in the dimension, so an infinite VC dimension means no finite sample suffices, for any epsilon and delta.
That is exactly no-free-lunch, restated quantitatively. The class of all functions on an infinite domain has infinite VC dimension, and the fundamental theorem then says it is not learnable.
So the two results are not separate facts. No-free-lunch is the qualitative shadow of the lower bound in the fundamental theorem.
Two truths and a lie
Three readings of the theorem, one of them defensible.
Eliminate the wrong options
Which is right?
Survives elimination: A
Why: The theorem is best read as a statement about what the definition of learning must include, rather than as a pessimistic result. It is the formal reason the hypothesis class appears in every theorem in the course.
Commit first
Answer, then rate your confidence.
Predict first
A class has VC dimension 5. Is it PAC learnable?
Correct: Yes, and the fundamental theorem gives the rate.
Why: Finite VC dimension is equivalent to learnability for binary classification with the zero-one loss, in both the realizable and agnostic settings, and the rate is the dimension plus a log term over epsilon squared. The fourth option inverts the definition: PAC learnability already quantifies over all distributions, so no distribution-specific information is needed or allowed.
Ranking
Five questions, and the order matters more than it looks.
Put in order
Why: The setting first, because the theorem statement is unreadable without it. The rate before the new idea, because an unusually good rate is the clue that points at which assumption is doing the work. And the last question is what turns reading into understanding, so it cannot come first.
Section
Section 5
Concept
A statistical learning paper is not read front to back. Five questions, answered in order, get you most of the way, and they reuse the template you already have.
Bousquet, Boucheron & Lugosi, Introduction to Statistical Learning Theory — a good second voice, and short enough to practise the routine on
Picture it
Answer these on one page before reading any proof in detail.
Figure (svg): A five step flow through the setting, the assumptions, the main theorem and its rate, the template step the paper innovates on, and what breaks without each assumption
Most papers innovate on exactly one of the five template steps. Identifying which one turns a forty-page paper into one idea plus a lot of bookkeeping.
Step zero
Pick anything from your professor's reading list.
Discussion prompt
Without reading past the abstract and the theorem statements, answer: what is the setting, what are the assumptions, and what is the rate?
Hint: You are allowed to read only the abstract and the theorem statements for this exercise.
Answer:
The setting is almost always in the first two paragraphs of section two, and it is usually four objects: domain, labels, loss, class.
The assumptions are frequently in a displayed list, and the interesting ones are the ones a reader would not have guessed: bounded loss, sub-Gaussian noise, a margin condition.
The rate is in the abstract, and it is worth writing down as a formula rather than as words, so you can compare it against the rates you already know.
If the rate is better than the standard one over the square root of m, some assumption is buying that, and finding which one is the most useful single thing you can do with the paper.
Mohri, Rostamizadeh & Talwalkar, Foundations of Machine Learning, 2nd edition — its bibliographic notes are a good map of which paper did what
Pattern
Four of these appeared in part one as traps. Keep the list beside you while writing homework.
| mistake | how to catch it |
|---|---|
| applying concentration to the hypothesis the algorithm chose | the bound has no complexity term in it |
| dropping the factor of two on a two-sided tail | compare against the one-sided version and see if it halved |
| confusing probability over the sample with probability over a test point | name the random object at every probability sign |
| using realizability after the setting dropped it | an agnostic theorem coming out with a one over eps rate |
| union bounding over an infinite class | the bound is infinite, hence vacuous |
Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) chapters 2-6
Check
A classmate's agnostic bound.
Check your understanding
A student proves an agnostic bound and gets m at least log of the class size over delta, divided by epsilon. What is almost certainly wrong?
Answer: A
Why: The epsilon dependence is the fingerprint of which concentration argument was used. A linear dependence on one over epsilon comes from the zero-mistakes argument, which is only available under realizability.
Exit ticket
One honest answer, and it sets the agenda for session two.
Predict first
Which of these would you least want to be asked to reproduce on Tuesday?
Correct: Whichever you named opens the next session.
Why: If it is the fifth, that is the most likely thing to appear on the first homework, and the two-half discipline is drillable in about twenty minutes across the four standard classes.
Connect it up
Twenty minutes before Tuesday.
Draw it
On one page: the five template steps down the left, and two columns to their right — how the realizable proof performs each step, and how the agnostic proof does. Then, underneath, the four VC classes with their dimensions and the one-line obstruction for each upper half.
That page is the whole of session one. Bring it, and bring the questions it exposes.
Recap
The template held up under a second run, which is the evidence that it is worth memorising.
| result | rate or value |
|---|---|
| Hoeffding, zero-one loss | two e to the minus two m eps squared |
| finite class, realizable | log of class over delta, over eps |
| finite class, agnostic | twice log of twice class over delta, over eps squared |
| numeric comparison at eps 0.1 | 77 against 1659 |
| thresholds, intervals, rectangles, halfspaces | 1, 2, 4, and d plus one |
| fundamental theorem, agnostic | d plus log one over delta, over eps squared |
Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) chapters 4-6 — read these before session two, and Bousquet sections 1 to 3 as a second voice
Want this taught 1-on-1? Alexander tutors Statistical Learning Theory — $55/session, free consultation.