Session 1a of a master's statistical learning theory course, built for a student who follows algorithm proofs but not learning-theory ones: the formal setup, why empirical risk minimisation can fail, the finite realizable sample-complexity theorem presented as four named moves each mapped onto its algorithm-proof twin, and the PAC and agnostic PAC definitions with their quantifiers read in order.
Subject: Statistical Learning Theory · 60 slides · symbolic lesson
Open the interactive version of this deck · Homework for this lesson
Title
Statistical Learning Theory · Session 1a
Four moves that recur in nearly every sample-complexity proof you will meet
Objectives
You told me the proofs are the gap, and that algorithm proofs are not. That is a good position to be in, because the two are more alike than the textbooks make them look.
Part two of this session then runs the same template a second time, which is how you find out it really is a template.
Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) chapters 2-4 — the reference this deck follows most closely
Section
Section 1
Warm-up
Two minutes before any notation.
Discussion prompt
In an algorithm proof — say, proving binary search correct — what are the standard parts? Name them.
Hint: What has to be true before, what has to be true after, and what stays true throughout?
Answer:
A precondition, a postcondition, an invariant that holds at every iteration, a per-step argument that the invariant is preserved, and a termination or bound argument.
Hold that list. By the end of this deck each part will have a counterpart in the learning-theory proof, and the counterpart is exact rather than a loose analogy.
Concept
Six objects. Every theorem in the course is a statement about them, so it is worth being pedantic for two minutes.
| object | what it is | known to the learner? |
|---|---|---|
| the domain | the set of possible inputs | yes |
| the label set | today, just two labels | yes |
| the distribution | how inputs are generated | no, never |
| the labelling function | the truth being learned | no |
| the training sample | m pairs drawn independently | yes |
| the hypothesis class | the predictors allowed, fixed in advance | yes |
The third and fourth rows are the whole difficulty. Everything else is bookkeeping.
Picture it
Draw this once on the first page of your notes and refer back to it all semester.
Figure (svg): A pipeline from an unknown distribution to a training sample to a learner to an output hypothesis, with true risk and empirical risk contrasted below
Note where the distribution sits: it generates the sample and it defines the true risk, but the learner never observes it. That is why the subject exists.
Concept
The true risk of a predictor is the probability it is wrong on a fresh draw. The empirical risk is the fraction of the training sample it gets wrong.
\[ L_{\mathcal{D}}(h) = \Pr_{x \sim \mathcal{D}}\left[h(x) \neq f(x)\right] \]
\[ L_S(h) = \frac{1}{m}\left|\{ i : h(x_i) \neq y_i \}\right| \]
You can compute the second one. You want a guarantee about the first.
Socratic
Take a real minute here. This is the sentence that makes everything else make sense.
Discussion prompt
For a hypothesis chosen in advance, its empirical risk is an average of m independent zero-one variables, so a concentration inequality applies directly. Why can you not do that to the hypothesis your algorithm actually output?
Hint: What does the algorithm know that a hypothesis fixed in advance does not?
Answer:
Because the output hypothesis was selected using the sample. It is a function of S, so it is not independent of S, and the terms being averaged are no longer independent draws from a fixed distribution.
Concretely: the algorithm searched for the hypothesis that looks best on this particular sample, so of course it looks good on it. That is selection, not evidence.
Every major tool in the syllabus — the union bound, the growth function, symmetrization, Rademacher complexity — exists to buy back a statement about the chosen hypothesis from statements about fixed ones.
Bousquet, Boucheron & Lugosi, Introduction to Statistical Learning Theory §2
Picture it
Two hypotheses, and only one of them is a legal target for a concentration inequality.
Figure (svg): Two panels contrasting a hypothesis fixed before seeing the sample, to which concentration applies, against the hypothesis chosen using the sample, to which it does not
When you meet a proof that seems to be doing something unnecessarily elaborate, it is almost always paying for this crossing.
Two truths and a lie
Let h be some hypothesis and let h_S be the output of empirical risk minimisation on S.
Eliminate the wrong options
Which one may you assert without further argument?
Survives elimination: A
Why: Only the first involves a hypothesis that is independent of the sample. Recognising which of these you are entitled to at any moment in a proof is most of what reading these proofs consists of.
Section
Section 2
Concept
The obvious algorithm: return any predictor in the class that makes the fewest mistakes on the training sample.
\[ h_S \in \operatorname*{argmin}_{h \in \mathcal{H}} L_S(h) \]
It is the right starting point and it is not automatically safe. The next slide is the counterexample everybody meets first.
Prediction
Let the class be all functions from the domain to the two labels, with no restriction at all, and let the true label be 1 everywhere.
Predict first
Consider the predictor that outputs the training label on training points and 0 everywhere else. What are its two risks?
Correct: Empirical risk 0, true risk 1.
Why: It reproduces every training label exactly, so it makes no training mistakes. On a continuous distribution a fresh point almost surely misses the finite training set, so the predictor outputs 0 while the truth is 1, and it is wrong with probability one. It is a legal minimiser of the empirical risk and it is worthless.
Picture it
Perfect on five points, wrong on the continuum.
Figure (svg): A line with five filled sample points labelled one and many hollow points elsewhere labelled zero, annotated with empirical risk zero and true risk one
Nothing about this is pathological. It is what unrestricted empirical risk minimisation is entitled to do.
Concept
Choosing the hypothesis class before seeing the data is what rules the memoriser out. That choice is inductive bias, and it is not an unfortunate approximation: without it there is nothing to prove.
inductive bias — The restriction on which predictors the learner may return, committed to before any data is seen.
The no-free-lunch theorem in part two turns this observation into a theorem: no bias, no learning, for any algorithm whatsoever.
Explain it to yourself
Worth answering out loud, because the proof depends on it in a place that is easy to miss.
Discussion prompt
The theorems all say the class is chosen before the sample. What exactly would break if you were allowed to choose it afterwards?
Hint: Look ahead to the union bound and ask what it is quantifying over.
Answer:
The union bound in step three is taken over a class that does not depend on S. If the class were chosen after seeing S, the set of hypotheses you are union-bounding over would itself be random, and the bound would not apply.
Operationally, choosing the class after looking at the data is exactly how the memoriser gets back in: you would simply pick the class containing it.
This is the same discipline as fixing a significance level before running a test, and it fails in the same way when violated.
Section
Section 3
Concept
Assume the class is finite, and assume some member of it labels the data perfectly. Then a modest sample suffices.
\[ m \ge \frac{\log(|\mathcal{H}|/\delta)}{\varepsilon} \]
Under that condition, with probability at least one minus delta over the draw of the sample, every empirical risk minimiser has true risk at most epsilon.
Read the statement twice before the proof. Note where the probability lives: over the sample, not over anything else.
Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) Corollary 2.3
Picture it
Learn these as moves with names. They recur all semester.
Figure (svg): A four step flow: define failure over hypotheses, bound one fixed bad hypothesis, union bound over the class, then invert for the sample size
Worked example
The goal is a statement about the algorithm. The first move converts it into a statement about the class, which is what we can actually reason about.
Call a hypothesis bad if its true risk exceeds epsilon
Why: This defines a subset of the class, with no reference to the algorithm at all.
\[ \mathcal{H}_{\text{bad}} = \{ h \in \mathcal{H} : L_{\mathcal{D}}(h) > \varepsilon \} \]
Observe what must have happened if the algorithm failed
Why: Realizability means some hypothesis has zero empirical risk, so the minimiser also has zero empirical risk. If its true risk exceeds epsilon, then some bad hypothesis had zero empirical risk.
\[ \{ L_{\mathcal{D}}(h_S) > \varepsilon \} \subseteq \{ \exists h \in \mathcal{H}_{\text{bad}} : L_S(h) = 0 \} \]
Take probabilities of both sides
Why: A subset has no larger probability, so bounding the right-hand event bounds the failure.
Verify: that the right-hand event no longer mentions the algorithm
Why: It does not. That is the whole purpose of this move, and it is the same manoeuvre as replacing a claim about a loop's output with a claim about its invariant.
Picture it
The same logical shape as an algorithm proof's postcondition step.
Figure (svg): Three boxes showing the algorithm failing implies some bad hypothesis survived the sample, which implies a bound on the probability
In an algorithm proof you replace 'the loop produced the right answer' with 'the invariant held and the loop exited'. Identical move.
Worked example
Now fix a single bad hypothesis. It is fixed, so concentration is legal.
Write the probability that it survives all m points
Why: Each point is drawn independently, and the hypothesis is wrong on any given point with probability more than epsilon.
\[ \Pr\left[L_S(h) = 0\right] = \left(1 - L_{\mathcal{D}}(h)\right)^m < (1-\varepsilon)^m \]
Relax to an exponential
Why: The inequality one minus x is at most the exponential of minus x holds for every real x, and it makes the algebra in move 4 trivial.
\[ (1-\varepsilon)^m \le e^{-\varepsilon m} \]
Verify: that the bound behaves correctly at the extremes
Why: At epsilon equal to zero it gives one, which is right, since a hypothesis with no error certainly survives. As m grows it goes to zero geometrically, which matches the intuition that each extra sample is another independent chance to catch the impostor.
Picture it
With epsilon a tenth, the exact survival probability and its exponential bound.
Figure (svg): Two decaying curves against sample size, the exact one below and the exponential bound above
The relaxation costs a little, and in the homework you compute exactly how much: seventy-three samples become seventy-seven.
Analogy
Match each learning-theory move to the algorithm-proof move it corresponds to.
Match the pairs
Why: The correspondence is structural rather than decorative. In both cases the hard intellectual work is the per-unit bound in the second row, and the third row is a mechanical accumulation over units — iterations in one case, hypotheses in the other.
Worked example
You do not know which bad hypothesis will fool you, so you pay for all of them.
Apply the union bound over the bad set
Why: The probability that at least one of several events occurs is at most the sum of their probabilities. No independence is needed, which is exactly why this tool is used so heavily.
\[ \Pr\left[\exists h \in \mathcal{H}_{\text{bad}} : L_S(h)=0\right] \le \sum_{h \in \mathcal{H}_{\text{bad}}} e^{-\varepsilon m} \]
Bound the number of terms by the size of the whole class
Why: The bad set is a subset, so its size is at most the size of the class. This is where finiteness is used, and it is the only place.
\[ \le |\mathcal{H}| e^{-\varepsilon m} \]
Verify: that no independence was assumed anywhere
Why: It was not. The bad events overlap heavily in general, and the union bound does not care. That looseness is the price of its generality.
Picture it
The bad events overlap, and the union bound charges you for the overlaps twice.
Figure (svg): Overlapping shaded regions inside a rectangle of all samples, each region being the event that one bad hypothesis looks perfect
The looseness here is what later tools attack. The growth function replaces the class size with the number of behaviours the class can exhibit on m points, which is vastly smaller.
Worked example
Demand that the failure probability be at most delta and solve.
Set the bound below delta
Why: This is the definition of what we promised.
\[ |\mathcal{H}| e^{-\varepsilon m} \le \delta \]
Take logarithms and rearrange
Why: Nothing subtle; this is the recurrence-solving step of the proof.
\[ m \ge \frac{\log(|\mathcal{H}|/\delta)}{\varepsilon} \]
Verify: that each dependence points the right way
Why: Larger class needs more data, and only logarithmically, which is the good news. Smaller epsilon needs more data, linearly. Smaller delta needs more data, logarithmically, so high confidence is cheap. If any of those had come out backwards, the algebra would be wrong.
Picture it
Three inputs, three very different costs.
Figure (svg): Three bars comparing how sample size grows with class size, with accuracy and with confidence
Socratic
This is the question that turns a proof into a template.
Discussion prompt
Read back over the four moves. What property of the hypothesis class did the argument actually rely on?
Hint: Go through the four moves and mark every place the class appears.
Answer:
Only that it is finite. Nothing about what the hypotheses are, what the domain is, or how the class is parameterised.
That is why the same proof gives a bound for any finite class at all, and it is why the search for a replacement for the class size is the obvious next question — which is exactly what VC dimension answers.
Getting into the habit of asking this at the end of every proof is the single most useful reading habit in this subject.
Pattern
Written out once, so you can check any proof in the course against it.
| template step | in this proof | the algorithm-proof twin |
|---|---|---|
| fix targets | eps and delta given | state the specification |
| bad event | true risk above eps | negate the postcondition |
| one hypothesis | survives with prob at most e to the minus eps m | one iteration preserves the invariant |
| extend | union bound, cost is the class size | induct over iterations |
| invert | solve for m | solve the recurrence |
Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) chapters 2-4
Check
Think about where the assumption was actually used.
Check your understanding
If the data are not perfectly labellable by any hypothesis in the class, which of the four moves stops working?
Answer: A
Why: Realizability is what licensed the step from the algorithm failing to some bad hypothesis having zero empirical risk. Without it the minimiser has some positive empirical risk and the characterisation of the bad event collapses.
Trap
A proof attempt that looks completely reasonable.
\[ \Pr\left[|L_S(h_S) - L_{\mathcal{D}}(h_S)| > \varepsilon\right] \le 2e^{-2m\varepsilon^2} \]
Apply Hoeffding to the hypothesis the algorithm returned
Why: Hoeffding needs the summands to be independent draws. Here the hypothesis was selected using those very draws.
Conclude a bound with no dependence on the class at all
Why: That should be the alarm: it would prove that learning is possible with no restriction on the class, which the memoriser refutes.
Apply it only to hypotheses fixed in advance, then pay to extend.
\[ \forall h \in \mathcal{H}: \; \Pr\left[|L_S(h) - L_{\mathcal{D}}(h)| > \varepsilon\right] \le 2e^{-2m\varepsilon^2} \]
Quantify over the class before drawing the sample
Why: Each statement is about a fixed hypothesis, which is legal, and the quantifier sits outside the probability.
Then union bound to get one statement covering all of them at once
Why: The class size reappears, and that is correct: it is the price of covering the algorithm's choice.
Sanity-check any bound that has no complexity term in it
Why: If a generalisation bound does not mention the class in any way, you have almost certainly made this mistake.
Error analysis
A student's proof sketch for the realizable finite case.
Annotate
On: \( \Pr[L_{\mathcal{D}}(h_S) > \varepsilon] = \Pr[L_S(h_S)=0 \text{ and } L_{\mathcal{D}}(h_S) > \varepsilon] \le e^{-\varepsilon m} \)
The giveaway is the missing class size. A realizable bound without one would say that any class at all is learnable from a constant number of samples.
Trap
Reading the guarantee as a statement about individual predictions.
Say that the learner is wrong at most a delta fraction of the time
Why: Delta is not an error rate on test points. It is the probability that the training sample was unlucky.
Conflate epsilon and delta
Why: Epsilon is the error rate on fresh points. Delta is how often the whole training run fails to deliver that error rate.
Two nested statements, with different randomness in each.
\[ \Pr_{S}\Big[\; \underbrace{L_{\mathcal{D}}(h_S)}_{\text{over a fresh } x} \le \varepsilon \;\Big] \ge 1-\delta \]
Inner: the error rate on a fresh point, for the hypothesis you ended up with
Why: This is epsilon, and the fresh point has already been averaged over.
Outer: over which training samples you might have drawn
Why: This is delta. On a bad draw the guarantee simply does not hold, and nothing bounds how bad it is.
Picture it
Almost every misreading of a PAC statement is these two being merged.
Figure (svg): Two panels distinguishing randomness over the training sample from randomness over a fresh test point, with delta attached to the first
Whenever a proof writes a probability, name which of the two it is before reading on. It takes a second and it prevents most of the confusion.
Trap
Extending the argument to thresholds on the real line, which is an infinite class.
\[ \Pr\left[\exists h : L_S(h)=0, L_{\mathcal{D}}(h)>\varepsilon\right] \le \sum_{h \in \mathcal{H}} e^{-\varepsilon m} = \infty \]
Write the same sum over an uncountable class
Why: The sum diverges, so the bound is vacuous. It is not wrong, it just says nothing.
Conclude that thresholds are not learnable
Why: They are, very easily. The tool was wrong, not the claim.
Count behaviours on the sample, not hypotheses in the class.
Infinitely many thresholds produce only m plus one distinct labellings of m points, because only which points fall either side matters.
\[ \tau_{\mathcal{H}}(m) = m+1 \]
Replace the class size by the growth function
Why: That is exactly what the next few lectures build, and Sauer and Shelah bound the growth function by a polynomial in m of degree the VC dimension.
Recognise the pattern
Why: Whenever a bound is vacuous, ask what quantity was over-counted rather than abandoning the approach.
Concept
Take thresholds restricted to a grid of one thousand candidate values, which is what any implementation actually searches. The class is finite, so the theorem applies directly.
\[ |\mathcal{H}| = 1000, \quad \varepsilon = 0.05, \quad \delta = 0.01 \]
\[ m \ge \frac{\log(1000/0.01)}{0.05} = \frac{11.51}{0.05} \approx 231 \]
Two hundred and thirty-one labelled examples for a five percent error guarantee at ninety-nine percent confidence. That is a small number, and the reason is the logarithm.
Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) Corollary 2.3
Estimation
Now enlarge the class by a factor of a thousand, to a million hypotheses, keeping epsilon at a tenth.
Predict first
Roughly how many samples does a million-hypothesis class need?
Correct: About 170.
Why: The class size enters through a logarithm, so a thousandfold increase adds only the logarithm of a thousand, about seven, divided by epsilon. The exact figure is 169. This is the single most encouraging fact in the subject, and it is why finiteness alone is enough.
Picture it
A thousand hypotheses, a million hypotheses, and the cost of doubling.
Figure (svg): Three bars comparing sample sizes for a thousand and a million hypotheses and the marginal cost of doubling a class
Which immediately raises the question the course spends its first month on: what is the right notion of size for a class that is infinite?
Fill the middle
Move four, with the middle removed.
Fill in the blanks
|\mathcal-\varepsilon m \le \log(\delta/|\mathcal{H}|)|e^\frac{\log(|\mathcal{H}|/\delta)}{\varepsilon} \le \delta \;\Rightarrow\; ___ \;\Rightarrow\; m \ge ___
Why: Taking logarithms turns the exponential into a linear inequality in m, and dividing by epsilon finishes it. Note that the inequality flips when you divide by the negative sign on the exponent, which is the one place to be careful.
Ranking
The lines of the finite realizable proof, shuffled.
Put in order
Why: The order is forced by dependency rather than by taste: the single-hypothesis bound in step three is only legal because step one defined the bad set without reference to the sample, and step four only makes sense once there is a per-hypothesis quantity to sum.
Picture it
Each step is licensed by the one above it.
Figure (svg): A four step dependency chain showing that defining the bad set without the algorithm licenses the fixed-hypothesis bound, which licenses the union bound, which produces the inequality in m
Explain it
Two sentences, out loud.
Discussion prompt
Why does the proof pay for every hypothesis in the class, rather than just the one the algorithm will end up choosing?
Hint: At the moment you apply the bound, has the algorithm run yet?
Answer:
Because which one the algorithm chooses depends on the sample, and the bound has to hold before the sample is drawn. You cannot condition on a choice that has not been made yet.
So you insure against all of them at once. The premium is the size of the class, and the whole later theory is about negotiating that premium down.
Commit first
Answer, then rate your confidence honestly.
Predict first
The finite realizable theorem is proved. Does it tell you anything about a class of one hundred hypotheses when the data are NOT perfectly labellable by any of them?
Correct: No.
Why: The theorem's conclusion is derived from a chain that begins with the minimiser achieving zero empirical risk. Drop that and the chain has no first link. A statement for the non-realizable case has to be proved separately, which is what the agnostic bound in part two does, at a cost of one over epsilon squared rather than one over epsilon.
Discrimination
Deciding which setting a problem lives in is the first thing to do when reading any theorem.
Sort into buckets
Which setting does each description belong to?
Section
Section 4
Concept
A class is probably approximately correct learnable if there is a function giving, for each accuracy and confidence, a sufficient sample size that works against every distribution.
\[ \forall \varepsilon, \delta \in (0,1), \; \forall \mathcal{D}, \; \forall f \in \mathcal{H}: \; m \ge m_{\mathcal{H}}(\varepsilon,\delta) \Rightarrow \Pr_{S}\left[L_{\mathcal{D}}(h_S) \le \varepsilon\right] \ge 1-\delta \]
Read the quantifiers left to right, out loud, every time. Most confusion about this definition is a quantifier read in the wrong order.
Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) Definition 3.1
Picture it
Five rows, and the order is the content.
Figure (svg): Five nested bars listing the quantifiers of the PAC definition in order, with the universal quantifier over distributions highlighted
That ordering is what makes the definition demanding. Swapping the third and fourth rows would give a much weaker and much less useful notion.
Notation
Every piece of this carries weight.
Annotate
On: \( \Pr_{S \sim \mathcal{D}^m}\left[L_{\mathcal{D}}(h_S) \le \varepsilon\right] \ge 1-\delta \)
Point four is worth pausing on. The guarantee permits the learner to return something terrible on an unlucky sample, at most a delta fraction of the time.
Sorting
Keeping the two sources of randomness separate is most of the battle in this subject.
Sort into buckets
In each expression, what is random?
If you can answer this question at every line of a proof, you will not get lost in one.
Concept
Now the distribution is over input and label pairs together, so the same input may carry different labels on different draws. No hypothesis need be perfect, and no hypothesis need even be good.
\[ L_{\mathcal{D}}(h_S) \le \min_{h \in \mathcal{H}} L_{\mathcal{D}}(h) + \varepsilon \]
The promise is now relative: do almost as well as the best member of the class. Whether that is any good depends on the class, which is where the bias and complexity tradeoff comes in.
Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) Definition 3.3
Comparison
Four rows. Fill in what changes.
Comparison matrix
| realizable | agnostic | |
|---|---|---|
| distribution over | inputs, plus a labelling function | input and label pairs jointly |
| best achievable risk | zero | the minimum over the class, generally positive |
| the guarantee | true risk at most eps | within eps of the best in class |
| sample complexity in eps | one over eps | one over eps squared |
The last row is the practical headline: the same accuracy costs quadratically more data once you give up perfection.
Picture it
Both curves show the sample needed as the accuracy demand tightens.
Figure (svg): Two curves of sample size against accuracy, the agnostic one growing much faster as accuracy tightens
The slogan is worth memorising: agnostic costs a square, because you can no longer use the zero-mistakes trick.
Edge cases
Be precise. Naming the exact line is the skill.
Discussion prompt
In the agnostic setting, which line of the finite-class proof is the first one that becomes false, and what would you need to replace it?
Hint: Which line used the word zero?
Answer:
The line asserting that the minimiser has zero empirical risk. Without a perfect hypothesis in the class, the minimiser has some positive empirical risk and the event 'a bad hypothesis survived every point' is no longer implied by failure.
What you need instead is control of the difference between empirical and true risk, for every hypothesis simultaneously. That is the uniform convergence property, and it is the subject of part two.
Notice the shape of the repair: the second and third moves survive unchanged, and only the first is rewritten. That is typical.
Concept
A habit worth forming now: before reading any proof in this course, answer four questions about the statement.
The fourth question is the one that most reliably tells you what the proof will have to do.
Bousquet, Boucheron & Lugosi, Introduction to Statistical Learning Theory §1
Explain it to yourself
Answer all four questions about the finite realizable bound, out loud.
Discussion prompt
For that theorem: what is quantified, what is the probability over, which assumptions are load-bearing, and what would a counterexample look like?
Hint: Take them one at a time and write the answers down; do not do this in your head.
Answer:
Quantified: over every distribution and every labelling consistent with the class, and over every accuracy and confidence. The class is fixed first.
The probability is over the draw of the training sample, and over nothing else.
Load-bearing: finiteness of the class, realizability, and independence of the draws. Convenience: the exponential relaxation, which only changes constants.
A counterexample would be a finite class, a distribution, and a labelling in the class, on which some empirical risk minimiser has true risk above epsilon with probability more than delta despite a sample of the stated size. The memoriser is not one, because its class is not finite.
Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) Corollary 2.3
Real world
You have trained models. Connect the theory to something you have actually seen.
Discussion prompt
Hyperparameter search evaluates many candidate models on a validation set and keeps the best. Which part of today's material is that, and what does it predict?
Hint: What plays the role of the hypothesis class when you are choosing between fifty learning rates?
Answer:
It is exactly the setting of the theorem, with the validation set as the sample and the set of candidate configurations as the hypothesis class.
So the validation score of the selected model is optimistically biased, and the bias grows with the logarithm of the number of configurations tried. This is why a held-out test set, untouched by the search, is not a formality.
The bound even tells you roughly how much: with a validation set of size m and k configurations, the gap scales like the square root of log k over m, which is the agnostic rate from part two.
Mohri, Rostamizadeh & Talwalkar, Foundations of Machine Learning, 2nd edition chapter 2
Exit ticket
One honest answer before part two.
Predict first
Which of these is still least comfortable?
Correct: Whichever you named is where part two starts.
Why: If the answer is the second option, that is the most important one on the list and worth going back over before anything else, because every technique later in the syllabus is a response to it.
Connect it up
Ten minutes, on the inside cover of your notebook, before the next lecture.
Draw it
Write the five template steps down the left of a page. Beside each, write the corresponding algorithm-proof move. Then, in a third column, write the line of the finite-class proof that performs it.
Take that page to every lecture. When a proof loses you, find the step you are on before rereading anything.
Recap
One setup, one counterexample, one theorem, and one template that the rest of the course reuses.
| quantity | expression | how it scales |
|---|---|---|
| finite realizable sample size | log of class size over delta, all over eps | log in class size, linear in one over eps |
| survival of one bad hypothesis | at most e to the minus eps m | geometric in m |
| union bound cost | the class size | the thing VC dimension later replaces |
| agnostic sample size | see part two | one over eps squared |
Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) chapters 2-3 — read these two before the next session
Want this taught 1-on-1? Alexander tutors Statistical Learning Theory — $55/session, free consultation.