The Learning Problem and the Proof Template

Session 1a of a master's statistical learning theory course, built for a student who follows algorithm proofs but not learning-theory ones: the formal setup, why empirical risk minimisation can fail, the finite realizable sample-complexity theorem presented as four named moves each mapped onto its algorithm-proof twin, and the PAC and agnostic PAC definitions with their quantifiers read in order.

Subject: Statistical Learning Theory · 60 slides · symbolic lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. The Learning Problem and the Proof Template

Title

Statistical Learning Theory · Session 1a

Four moves that recur in nearly every sample-complexity proof you will meet

2. What this session is for

Objectives

You told me the proofs are the gap, and that algorithm proofs are not. That is a good position to be in, because the two are more alike than the textbooks make them look.

Part two of this session then runs the same template a second time, which is how you find out it really is a template.

Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) chapters 2-4 — the reference this deck follows most closely

3. The Setup

Section

Section 1

4. What you already have

Warm-up

Two minutes before any notation.

Discussion prompt

In an algorithm proof — say, proving binary search correct — what are the standard parts? Name them.

Hint: What has to be true before, what has to be true after, and what stays true throughout?

Answer:

A precondition, a postcondition, an invariant that holds at every iteration, a per-step argument that the invariant is preserved, and a termination or bound argument.

Hold that list. By the end of this deck each part will have a counterpart in the learning-theory proof, and the counterpart is exact rather than a loose analogy.

Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) chapter 2

5. The objects, named once

Concept

Six objects. Every theorem in the course is a statement about them, so it is worth being pedantic for two minutes.

objectwhat it isknown to the learner?
the domainthe set of possible inputsyes
the label settoday, just two labelsyes
the distributionhow inputs are generatedno, never
the labelling functionthe truth being learnedno
the training samplem pairs drawn independentlyyes
the hypothesis classthe predictors allowed, fixed in advanceyes

The third and fourth rows are the whole difficulty. Everything else is bookkeeping.

Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) §2.1

6. The setup as one picture

Picture it

Draw this once on the first page of your notes and refer back to it all semester.

Figure (svg): A pipeline from an unknown distribution to a training sample to a learner to an output hypothesis, with true risk and empirical risk contrasted below

The learner sees only the middle of the chain.

Note where the distribution sits: it generates the sample and it defines the true risk, but the learner never observes it. That is why the subject exists.

7. Two risks, one observable

Concept

The true risk of a predictor is the probability it is wrong on a fresh draw. The empirical risk is the fraction of the training sample it gets wrong.

\[ L_{\mathcal{D}}(h) = \Pr_{x \sim \mathcal{D}}\left[h(x) \neq f(x)\right] \]

\[ L_S(h) = \frac{1}{m}\left|\{ i : h(x_i) \neq y_i \}\right| \]

You can compute the second one. You want a guarantee about the first.

Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) §2.2

8. The observation that the whole course is built on

Socratic

Take a real minute here. This is the sentence that makes everything else make sense.

Discussion prompt

For a hypothesis chosen in advance, its empirical risk is an average of m independent zero-one variables, so a concentration inequality applies directly. Why can you not do that to the hypothesis your algorithm actually output?

Hint: What does the algorithm know that a hypothesis fixed in advance does not?

Answer:

Because the output hypothesis was selected using the sample. It is a function of S, so it is not independent of S, and the terms being averaged are no longer independent draws from a fixed distribution.

Concretely: the algorithm searched for the hypothesis that looks best on this particular sample, so of course it looks good on it. That is selection, not evidence.

Every major tool in the syllabus — the union bound, the growth function, symmetrization, Rademacher complexity — exists to buy back a statement about the chosen hypothesis from statements about fixed ones.

Bousquet, Boucheron & Lugosi, Introduction to Statistical Learning Theory §2

9. Fixed, versus chosen

Picture it

Two hypotheses, and only one of them is a legal target for a concentration inequality.

Figure (svg): Two panels contrasting a hypothesis fixed before seeing the sample, to which concentration applies, against the hypothesis chosen using the sample, to which it does not

The right-hand panel is the entire research programme of the field.

When you meet a proof that seems to be doing something unnecessarily elaborate, it is almost always paying for this crossing.

10. Which statement is safe?

Two truths and a lie

Let h be some hypothesis and let h_S be the output of empirical risk minimisation on S.

Eliminate the wrong options

Which one may you assert without further argument?

  • A. For a hypothesis fixed before S is drawn, the empirical risk concentrates around the true risk.
  • B. For the output hypothesis, the empirical risk concentrates around the true risk.
  • C. If the empirical risk is zero then the true risk is small.
  • D. Concentration applies to the output hypothesis provided the sample is large.

Survives elimination: A

Why: Only the first involves a hypothesis that is independent of the sample. Recognising which of these you are entitled to at any moment in a proof is most of what reading these proofs consists of.

11. ERM, and Why It Can Fail

Section

Section 2

12. Empirical risk minimisation

Concept

The obvious algorithm: return any predictor in the class that makes the fewest mistakes on the training sample.

\[ h_S \in \operatorname*{argmin}_{h \in \mathcal{H}} L_S(h) \]

It is the right starting point and it is not automatically safe. The next slide is the counterexample everybody meets first.

Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) §2.2

13. Predict the failure

Prediction

Let the class be all functions from the domain to the two labels, with no restriction at all, and let the true label be 1 everywhere.

Predict first

Consider the predictor that outputs the training label on training points and 0 everywhere else. What are its two risks?

  • Empirical risk 0, true risk 0
  • Empirical risk 0, true risk 1
  • Empirical risk 1, true risk 0
  • Both about one half

Correct: Empirical risk 0, true risk 1.

Why: It reproduces every training label exactly, so it makes no training mistakes. On a continuous distribution a fresh point almost surely misses the finite training set, so the predictor outputs 0 while the truth is 1, and it is wrong with probability one. It is a legal minimiser of the empirical risk and it is worthless.

14. The memoriser

Picture it

Perfect on five points, wrong on the continuum.

Figure (svg): A line with five filled sample points labelled one and many hollow points elsewhere labelled zero, annotated with empirical risk zero and true risk one

Zero training error, achieved by cheating.

Nothing about this is pathological. It is what unrestricted empirical risk minimisation is entitled to do.

15. So the class must be restricted, and that restriction is prior knowledge

Concept

Choosing the hypothesis class before seeing the data is what rules the memoriser out. That choice is inductive bias, and it is not an unfortunate approximation: without it there is nothing to prove.

inductive bias — The restriction on which predictors the learner may return, committed to before any data is seen.

The no-free-lunch theorem in part two turns this observation into a theorem: no bias, no learning, for any algorithm whatsoever.

Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) §2.3

16. Why does the class have to be fixed in advance?

Explain it to yourself

Worth answering out loud, because the proof depends on it in a place that is easy to miss.

Discussion prompt

The theorems all say the class is chosen before the sample. What exactly would break if you were allowed to choose it afterwards?

Hint: Look ahead to the union bound and ask what it is quantifying over.

Answer:

The union bound in step three is taken over a class that does not depend on S. If the class were chosen after seeing S, the set of hypotheses you are union-bounding over would itself be random, and the bound would not apply.

Operationally, choosing the class after looking at the data is exactly how the memoriser gets back in: you would simply pick the class containing it.

This is the same discipline as fixing a significance level before running a test, and it fails in the same way when violated.

Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) §2.3

17. The First Theorem

Section

Section 3

18. Finite class, realizable case

Concept

Assume the class is finite, and assume some member of it labels the data perfectly. Then a modest sample suffices.

\[ m \ge \frac{\log(|\mathcal{H}|/\delta)}{\varepsilon} \]

Under that condition, with probability at least one minus delta over the draw of the sample, every empirical risk minimiser has true risk at most epsilon.

Read the statement twice before the proof. Note where the probability lives: over the sample, not over anything else.

Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) Corollary 2.3

19. The four moves

Picture it

Learn these as moves with names. They recur all semester.

Figure (svg): A four step flow: define failure over hypotheses, bound one fixed bad hypothesis, union bound over the class, then invert for the sample size

Every sample-complexity proof in the syllabus is a variation on these four.

20. Worked example: move 1, define failure over hypotheses

Worked example

The goal is a statement about the algorithm. The first move converts it into a statement about the class, which is what we can actually reason about.

Call a hypothesis bad if its true risk exceeds epsilon

Why: This defines a subset of the class, with no reference to the algorithm at all.

\[ \mathcal{H}_{\text{bad}} = \{ h \in \mathcal{H} : L_{\mathcal{D}}(h) > \varepsilon \} \]

Observe what must have happened if the algorithm failed

Why: Realizability means some hypothesis has zero empirical risk, so the minimiser also has zero empirical risk. If its true risk exceeds epsilon, then some bad hypothesis had zero empirical risk.

\[ \{ L_{\mathcal{D}}(h_S) > \varepsilon \} \subseteq \{ \exists h \in \mathcal{H}_{\text{bad}} : L_S(h) = 0 \} \]

Take probabilities of both sides

Why: A subset has no larger probability, so bounding the right-hand event bounds the failure.

Verify: that the right-hand event no longer mentions the algorithm

Why: It does not. That is the whole purpose of this move, and it is the same manoeuvre as replacing a claim about a loop's output with a claim about its invariant.

21. Move 1 as a reduction

Picture it

The same logical shape as an algorithm proof's postcondition step.

Figure (svg): Three boxes showing the algorithm failing implies some bad hypothesis survived the sample, which implies a bound on the probability

Downwards is implication, so upwards is a bound on probability.

In an algorithm proof you replace 'the loop produced the right answer' with 'the invariant held and the loop exited'. Identical move.

22. Worked example: move 2, bound one fixed bad hypothesis

Worked example

Now fix a single bad hypothesis. It is fixed, so concentration is legal.

Write the probability that it survives all m points

Why: Each point is drawn independently, and the hypothesis is wrong on any given point with probability more than epsilon.

\[ \Pr\left[L_S(h) = 0\right] = \left(1 - L_{\mathcal{D}}(h)\right)^m < (1-\varepsilon)^m \]

Relax to an exponential

Why: The inequality one minus x is at most the exponential of minus x holds for every real x, and it makes the algebra in move 4 trivial.

\[ (1-\varepsilon)^m \le e^{-\varepsilon m} \]

Verify: that the bound behaves correctly at the extremes

Why: At epsilon equal to zero it gives one, which is right, since a hypothesis with no error certainly survives. As m grows it goes to zero geometrically, which matches the intuition that each extra sample is another independent chance to catch the impostor.

23. Exact against relaxed

Picture it

With epsilon a tenth, the exact survival probability and its exponential bound.

Figure (svg): Two decaying curves against sample size, the exact one below and the exponential bound above

The bound sits above the truth, which is the direction you need.

The relaxation costs a little, and in the homework you compute exactly how much: seventy-three samples become seventy-seven.

24. Move 2 has an algorithmic twin

Analogy

Match each learning-theory move to the algorithm-proof move it corresponds to.

Match the pairs

  • l1. define the bad event over hypotheses
  • l2. bound the probability for one fixed hypothesis
  • l3. union bound over the whole class
  • l4. set the total below delta and solve for m
  • r1. state the postcondition, and what must hold if it fails
  • r2. the per-iteration argument that the invariant is preserved
  • r3. induct over the iterations, accumulating the per-step cost
  • r4. solve the recurrence for the running time

Why: The correspondence is structural rather than decorative. In both cases the hard intellectual work is the per-unit bound in the second row, and the third row is a mechanical accumulation over units — iterations in one case, hypotheses in the other.

25. Worked example: move 3, the union bound

Worked example

You do not know which bad hypothesis will fool you, so you pay for all of them.

Apply the union bound over the bad set

Why: The probability that at least one of several events occurs is at most the sum of their probabilities. No independence is needed, which is exactly why this tool is used so heavily.

\[ \Pr\left[\exists h \in \mathcal{H}_{\text{bad}} : L_S(h)=0\right] \le \sum_{h \in \mathcal{H}_{\text{bad}}} e^{-\varepsilon m} \]

Bound the number of terms by the size of the whole class

Why: The bad set is a subset, so its size is at most the size of the class. This is where finiteness is used, and it is the only place.

\[ \le |\mathcal{H}| e^{-\varepsilon m} \]

Verify: that no independence was assumed anywhere

Why: It was not. The bad events overlap heavily in general, and the union bound does not care. That looseness is the price of its generality.

26. Paying for every blob

Picture it

The bad events overlap, and the union bound charges you for the overlaps twice.

Figure (svg): Overlapping shaded regions inside a rectangle of all samples, each region being the event that one bad hypothesis looks perfect

Crude, general, and enough.

The looseness here is what later tools attack. The growth function replaces the class size with the number of behaviours the class can exhibit on m points, which is vastly smaller.

27. Worked example: move 4, invert for the sample size

Worked example

Demand that the failure probability be at most delta and solve.

Set the bound below delta

Why: This is the definition of what we promised.

\[ |\mathcal{H}| e^{-\varepsilon m} \le \delta \]

Take logarithms and rearrange

Why: Nothing subtle; this is the recurrence-solving step of the proof.

\[ m \ge \frac{\log(|\mathcal{H}|/\delta)}{\varepsilon} \]

Verify: that each dependence points the right way

Why: Larger class needs more data, and only logarithmically, which is the good news. Smaller epsilon needs more data, linearly. Smaller delta needs more data, logarithmically, so high confidence is cheap. If any of those had come out backwards, the algebra would be wrong.

28. Sample complexity, term by term

Picture it

Three inputs, three very different costs.

Figure (svg): Three bars comparing how sample size grows with class size, with accuracy and with confidence

Accuracy is the expensive one, and in the agnostic case it gets worse.

29. What did the proof never use?

Socratic

This is the question that turns a proof into a template.

Discussion prompt

Read back over the four moves. What property of the hypothesis class did the argument actually rely on?

Hint: Go through the four moves and mark every place the class appears.

Answer:

Only that it is finite. Nothing about what the hypotheses are, what the domain is, or how the class is parameterised.

That is why the same proof gives a bound for any finite class at all, and it is why the search for a replacement for the class size is the obvious next question — which is exactly what VC dimension answers.

Getting into the habit of asking this at the end of every proof is the single most useful reading habit in this subject.

Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) §2.3

30. Pattern: the five-step template

Pattern

Written out once, so you can check any proof in the course against it.

  1. Fix the targets. An accuracy epsilon and a confidence delta, both arbitrary.
  2. Define the bad event over hypotheses, so the algorithm drops out of the statement.
  3. Bound one fixed hypothesis with a concentration inequality. This is where the probability actually enters.
  4. Extend to the whole class, by union bound now, by a smarter complexity measure later.
  5. Invert the resulting inequality for the sample size.
template stepin this proofthe algorithm-proof twin
fix targetseps and delta givenstate the specification
bad eventtrue risk above epsnegate the postcondition
one hypothesissurvives with prob at most e to the minus eps mone iteration preserves the invariant
extendunion bound, cost is the class sizeinduct over iterations
invertsolve for msolve the recurrence

Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) chapters 2-4

31. Check: which move fails without realizability?

Check

Think about where the assumption was actually used.

Check your understanding

If the data are not perfectly labellable by any hypothesis in the class, which of the four moves stops working?

  • A. Move 1, because the minimiser no longer has zero empirical risk (correct)
  • B. Move 2, because concentration no longer applies to a fixed hypothesis
  • C. Move 3, because the union bound needs independence
  • D. Move 4, because the logarithm is no longer defined

Answer: A

Why: Realizability is what licensed the step from the algorithm failing to some bad hypothesis having zero empirical risk. Without it the minimiser has some positive empirical risk and the characterisation of the bad event collapses.

Why B tempts people
Concentration for a fixed hypothesis needs only independent bounded samples, which realizability has nothing to do with.
Why C tempts people
The union bound never needed independence. That is precisely why it is used here.
Why D tempts people
The algebra in move 4 is unaffected. The problem is upstream, in what the bound is a bound on.

32. Trap: applying a concentration inequality to the output hypothesis

Trap

The trap

A proof attempt that looks completely reasonable.

\[ \Pr\left[|L_S(h_S) - L_{\mathcal{D}}(h_S)| > \varepsilon\right] \le 2e^{-2m\varepsilon^2} \]

Apply Hoeffding to the hypothesis the algorithm returned

Why: Hoeffding needs the summands to be independent draws. Here the hypothesis was selected using those very draws.

Conclude a bound with no dependence on the class at all

Why: That should be the alarm: it would prove that learning is possible with no restriction on the class, which the memoriser refutes.

The fix

Apply it only to hypotheses fixed in advance, then pay to extend.

\[ \forall h \in \mathcal{H}: \; \Pr\left[|L_S(h) - L_{\mathcal{D}}(h)| > \varepsilon\right] \le 2e^{-2m\varepsilon^2} \]

Quantify over the class before drawing the sample

Why: Each statement is about a fixed hypothesis, which is legal, and the quantifier sits outside the probability.

Then union bound to get one statement covering all of them at once

Why: The class size reappears, and that is correct: it is the price of covering the algorithm's choice.

Sanity-check any bound that has no complexity term in it

Why: If a generalisation bound does not mention the class in any way, you have almost certainly made this mistake.

33. Find the flawed step

Error analysis

A student's proof sketch for the realizable finite case.

Annotate

On: \( \Pr[L_{\mathcal{D}}(h_S) > \varepsilon] = \Pr[L_S(h_S)=0 \text{ and } L_{\mathcal{D}}(h_S) > \varepsilon] \le e^{-\varepsilon m} \)

  • The first equality is fine under realizability: the minimiser does achieve zero empirical risk.
  • The inequality is not. It applies the single-hypothesis bound to h_S, which is not fixed.
  • The repair is the union bound, and it costs a factor equal to the size of the class.

The giveaway is the missing class size. A realizable bound without one would say that any class at all is learnable from a constant number of samples.

34. Trap: confusing the two sources of randomness

Trap

The trap

Reading the guarantee as a statement about individual predictions.

Say that the learner is wrong at most a delta fraction of the time

Why: Delta is not an error rate on test points. It is the probability that the training sample was unlucky.

Conflate epsilon and delta

Why: Epsilon is the error rate on fresh points. Delta is how often the whole training run fails to deliver that error rate.

The fix

Two nested statements, with different randomness in each.

\[ \Pr_{S}\Big[\; \underbrace{L_{\mathcal{D}}(h_S)}_{\text{over a fresh } x} \le \varepsilon \;\Big] \ge 1-\delta \]

Inner: the error rate on a fresh point, for the hypothesis you ended up with

Why: This is epsilon, and the fresh point has already been averaged over.

Outer: over which training samples you might have drawn

Why: This is delta. On a bad draw the guarantee simply does not hold, and nothing bounds how bad it is.

35. Two randomnesses, kept apart

Picture it

Almost every misreading of a PAC statement is these two being merged.

Figure (svg): Two panels distinguishing randomness over the training sample from randomness over a fresh test point, with delta attached to the first

Delta belongs to the outer probability; epsilon to the inner one.

Whenever a proof writes a probability, name which of the two it is before reading on. It takes a second and it prevents most of the confusion.

36. Trap: union bounding an infinite class

Trap

The trap

Extending the argument to thresholds on the real line, which is an infinite class.

\[ \Pr\left[\exists h : L_S(h)=0, L_{\mathcal{D}}(h)>\varepsilon\right] \le \sum_{h \in \mathcal{H}} e^{-\varepsilon m} = \infty \]

Write the same sum over an uncountable class

Why: The sum diverges, so the bound is vacuous. It is not wrong, it just says nothing.

Conclude that thresholds are not learnable

Why: They are, very easily. The tool was wrong, not the claim.

The fix

Count behaviours on the sample, not hypotheses in the class.

Infinitely many thresholds produce only m plus one distinct labellings of m points, because only which points fall either side matters.

\[ \tau_{\mathcal{H}}(m) = m+1 \]

Replace the class size by the growth function

Why: That is exactly what the next few lectures build, and Sauer and Shelah bound the growth function by a polynomial in m of degree the VC dimension.

Recognise the pattern

Why: Whenever a bound is vacuous, ask what quantity was over-counted rather than abandoning the approach.

37. A concrete finite class, and a real number

Concept

Take thresholds restricted to a grid of one thousand candidate values, which is what any implementation actually searches. The class is finite, so the theorem applies directly.

\[ |\mathcal{H}| = 1000, \quad \varepsilon = 0.05, \quad \delta = 0.01 \]

\[ m \ge \frac{\log(1000/0.01)}{0.05} = \frac{11.51}{0.05} \approx 231 \]

Two hundred and thirty-one labelled examples for a five percent error guarantee at ninety-nine percent confidence. That is a small number, and the reason is the logarithm.

Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) Corollary 2.3

38. Estimate before computing

Estimation

Now enlarge the class by a factor of a thousand, to a million hypotheses, keeping epsilon at a tenth.

Predict first

Roughly how many samples does a million-hypothesis class need?

  • About 170
  • About 1700
  • About 170 thousand
  • About a million

Correct: About 170.

Why: The class size enters through a logarithm, so a thousandfold increase adds only the logarithm of a thousand, about seven, divided by epsilon. The exact figure is 169. This is the single most encouraging fact in the subject, and it is why finiteness alone is enough.

39. The logarithm is doing all the work

Picture it

A thousand hypotheses, a million hypotheses, and the cost of doubling.

Figure (svg): Three bars comparing sample sizes for a thousand and a million hypotheses and the marginal cost of doubling a class

Doubling the class costs about seven extra samples at a tenth accuracy.

Which immediately raises the question the course spends its first month on: what is the right notion of size for a class that is infinite?

40. Invert the bound yourself

Fill the middle

Move four, with the middle removed.

Fill in the blanks

|\mathcal-\varepsilon m \le \log(\delta/|\mathcal{H}|)|e^\frac{\log(|\mathcal{H}|/\delta)}{\varepsilon} \le \delta \;\Rightarrow\; ___ \;\Rightarrow\; m \ge ___

Why: Taking logarithms turns the exponential into a linear inequality in m, and dividing by epsilon finishes it. Note that the inequality flips when you divide by the negative sign on the exponent, which is the one place to be careful.

41. Put the proof in order

Ranking

The lines of the finite realizable proof, shuffled.

Put in order

  1. Fix epsilon and delta, and define the bad set of hypotheses
  2. Note that failure implies some bad hypothesis had zero empirical risk
  3. Bound the survival probability of one fixed bad hypothesis
  4. Union bound over the bad set, and enlarge it to the whole class
  5. Require the result to be at most delta and solve for m

Why: The order is forced by dependency rather than by taste: the single-hypothesis bound in step three is only legal because step one defined the bad set without reference to the sample, and step four only makes sense once there is a per-hypothesis quantity to sum.

42. Why the order is forced

Picture it

Each step is licensed by the one above it.

Figure (svg): A four step dependency chain showing that defining the bad set without the algorithm licenses the fixed-hypothesis bound, which licenses the union bound, which produces the inequality in m

Not a stylistic order; a dependency order.

43. Explain the union bound step to a classmate

Explain it

Two sentences, out loud.

Discussion prompt

Why does the proof pay for every hypothesis in the class, rather than just the one the algorithm will end up choosing?

Hint: At the moment you apply the bound, has the algorithm run yet?

Answer:

Because which one the algorithm chooses depends on the sample, and the bound has to hold before the sample is drawn. You cannot condition on a choice that has not been made yet.

So you insure against all of them at once. The premium is the size of the class, and the whole later theory is about negotiating that premium down.

Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) §2.3

44. How sure are you?

Commit first

Answer, then rate your confidence honestly.

Predict first

The finite realizable theorem is proved. Does it tell you anything about a class of one hundred hypotheses when the data are NOT perfectly labellable by any of them?

  • No — the theorem assumes realizability and says nothing otherwise
  • Yes — the same bound holds with a slightly worse constant
  • Yes — it holds with eps replaced by eps squared
  • Only if the best hypothesis has error below eps

Correct: No.

Why: The theorem's conclusion is derived from a chain that begins with the minimiser achieving zero empirical risk. Drop that and the chain has no first link. A statement for the non-realizable case has to be proved separately, which is what the agnostic bound in part two does, at a cost of one over epsilon squared rather than one over epsilon.

45. Realizable, or agnostic?

Discrimination

Deciding which setting a problem lives in is the first thing to do when reading any theorem.

Sort into buckets

Which setting does each description belong to?

realizable
labels are generated by an unknown function in the class; empirical risk minimisation achieves zero training error
agnostic
the same input can appear with different labels; the best achievable error within the class is 12 percent
real
Some member of the class is perfect, so zero training error is attainable and the survives-every-sample argument is available.
agn
No member of the class need be perfect, and the labelling may even be genuinely noisy. The goal becomes competitive rather than absolute.

46. PAC Learnability

Section

Section 4

47. The definition, with its quantifiers in order

Concept

A class is probably approximately correct learnable if there is a function giving, for each accuracy and confidence, a sufficient sample size that works against every distribution.

\[ \forall \varepsilon, \delta \in (0,1), \; \forall \mathcal{D}, \; \forall f \in \mathcal{H}: \; m \ge m_{\mathcal{H}}(\varepsilon,\delta) \Rightarrow \Pr_{S}\left[L_{\mathcal{D}}(h_S) \le \varepsilon\right] \ge 1-\delta \]

Read the quantifiers left to right, out loud, every time. Most confusion about this definition is a quantifier read in the wrong order.

Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) Definition 3.1

48. The quantifier ladder

Picture it

Five rows, and the order is the content.

Figure (svg): Five nested bars listing the quantifiers of the PAC definition in order, with the universal quantifier over distributions highlighted

The learner is chosen before the distribution, and must work against all of them.

That ordering is what makes the definition demanding. Swapping the third and fourth rows would give a much weaker and much less useful notion.

49. Decode the definition symbol by symbol

Notation

Every piece of this carries weight.

Annotate

On: \( \Pr_{S \sim \mathcal{D}^m}\left[L_{\mathcal{D}}(h_S) \le \varepsilon\right] \ge 1-\delta \)

  • The probability is over the draw of the training sample, and over nothing else. The distribution and the labelling are fixed but unknown; they are not random.
  • The superscript m says the sample is m independent draws. Independence is what every concentration inequality in the course needs.
  • The quantity inside is the true risk of the OUTPUT hypothesis, so it is itself a random variable — it depends on S through the learner.
  • Delta is the failure probability of the whole procedure, not of any individual prediction. Confusing those two is the commonest misreading of this line.

Point four is worth pausing on. The guarantee permits the learner to return something terrible on an unlucky sample, at most a delta fraction of the time.

50. Probability over what?

Sorting

Keeping the two sources of randomness separate is most of the battle in this subject.

Sort into buckets

In each expression, what is random?

not random at all
the true risk of a fixed hypothesis
random through the sample
the empirical risk of a fixed hypothesis; the true risk of the output hypothesis; the failure event in the PAC definition
none
A fixed hypothesis and a fixed distribution give one number. The definition of true risk already integrates over the test point, so nothing random survives.
sample
These vary as the training sample varies. Every probability statement in a PAC theorem is over exactly this randomness.
point
This bucket is empty here, and deliberately so: the fresh test point has already been averaged out inside the definition of true risk. Confusing that averaging with the outer probability over samples is the single most common error in these proofs.

If you can answer this question at every line of a proof, you will not get lost in one.

51. Agnostic PAC drops realizability

Concept

Now the distribution is over input and label pairs together, so the same input may carry different labels on different draws. No hypothesis need be perfect, and no hypothesis need even be good.

\[ L_{\mathcal{D}}(h_S) \le \min_{h \in \mathcal{H}} L_{\mathcal{D}}(h) + \varepsilon \]

The promise is now relative: do almost as well as the best member of the class. Whether that is any good depends on the class, which is where the bias and complexity tradeoff comes in.

Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) Definition 3.3

52. Realizable against agnostic

Comparison

Four rows. Fill in what changes.

Comparison matrix

realizableagnostic
distribution overinputs, plus a labelling functioninput and label pairs jointly
best achievable riskzerothe minimum over the class, generally positive
the guaranteetrue risk at most epswithin eps of the best in class
sample complexity in epsone over epsone over eps squared

The last row is the practical headline: the same accuracy costs quadratically more data once you give up perfection.

53. The price of dropping realizability

Picture it

Both curves show the sample needed as the accuracy demand tightens.

Figure (svg): Two curves of sample size against accuracy, the agnostic one growing much faster as accuracy tightens

One over eps against one over eps squared.

The slogan is worth memorising: agnostic costs a square, because you can no longer use the zero-mistakes trick.

54. Where exactly does the old proof break?

Edge cases

Be precise. Naming the exact line is the skill.

Discussion prompt

In the agnostic setting, which line of the finite-class proof is the first one that becomes false, and what would you need to replace it?

Hint: Which line used the word zero?

Answer:

The line asserting that the minimiser has zero empirical risk. Without a perfect hypothesis in the class, the minimiser has some positive empirical risk and the event 'a bad hypothesis survived every point' is no longer implied by failure.

What you need instead is control of the difference between empirical and true risk, for every hypothesis simultaneously. That is the uniform convergence property, and it is the subject of part two.

Notice the shape of the repair: the second and third moves survive unchanged, and only the first is rewritten. That is typical.

Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) §4.1

55. Reading a theorem statement before its proof

Concept

A habit worth forming now: before reading any proof in this course, answer four questions about the statement.

  1. What is quantified universally, and in what order?
  2. What is the probability over? Name the random object.
  3. Which assumptions are load-bearing, and which are convenience?
  4. What would a counterexample look like if the theorem were false?

The fourth question is the one that most reliably tells you what the proof will have to do.

Bousquet, Boucheron & Lugosi, Introduction to Statistical Learning Theory §1

56. Apply the reading habit to the theorem you just proved

Explain it to yourself

Answer all four questions about the finite realizable bound, out loud.

Discussion prompt

For that theorem: what is quantified, what is the probability over, which assumptions are load-bearing, and what would a counterexample look like?

Hint: Take them one at a time and write the answers down; do not do this in your head.

Answer:

Quantified: over every distribution and every labelling consistent with the class, and over every accuracy and confidence. The class is fixed first.

The probability is over the draw of the training sample, and over nothing else.

Load-bearing: finiteness of the class, realizability, and independence of the draws. Convenience: the exponential relaxation, which only changes constants.

A counterexample would be a finite class, a distribution, and a labelling in the class, on which some empirical risk minimiser has true risk above epsilon with probability more than delta despite a sample of the stated size. The memoriser is not one, because its class is not finite.

Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) Corollary 2.3

57. Where this bites in practice

Real world

You have trained models. Connect the theory to something you have actually seen.

Discussion prompt

Hyperparameter search evaluates many candidate models on a validation set and keeps the best. Which part of today's material is that, and what does it predict?

Hint: What plays the role of the hypothesis class when you are choosing between fifty learning rates?

Answer:

It is exactly the setting of the theorem, with the validation set as the sample and the set of candidate configurations as the hypothesis class.

So the validation score of the selected model is optimistically biased, and the bias grows with the logarithm of the number of configurations tried. This is why a held-out test set, untouched by the search, is not a formality.

The bound even tells you roughly how much: with a validation set of size m and k configurations, the gap scales like the square root of log k over m, which is the agnostic rate from part two.

Mohri, Rostamizadeh & Talwalkar, Foundations of Machine Learning, 2nd edition chapter 2

58. Exit ticket

Exit ticket

One honest answer before part two.

Predict first

Which of these is still least comfortable?

  • The setup and what each object is
  • Why concentration cannot be applied to the output hypothesis
  • The four moves of the finite-class proof
  • Reading the PAC quantifiers in order
  • What changes in the agnostic setting

Correct: Whichever you named is where part two starts.

Why: If the answer is the second option, that is the most important one on the list and worth going back over before anything else, because every technique later in the syllabus is a response to it.

59. Write the template out

Connect it up

Ten minutes, on the inside cover of your notebook, before the next lecture.

Draw it

Write the five template steps down the left of a page. Beside each, write the corresponding algorithm-proof move. Then, in a third column, write the line of the finite-class proof that performs it.

Take that page to every lecture. When a proof loses you, find the step you are on before rereading anything.

60. Part one, in summary

Recap

One setup, one counterexample, one theorem, and one template that the rest of the course reuses.

quantityexpressionhow it scales
finite realizable sample sizelog of class size over delta, all over epslog in class size, linear in one over eps
survival of one bad hypothesisat most e to the minus eps mgeometric in m
union bound costthe class sizethe thing VC dimension later replaces
agnostic sample sizesee part twoone over eps squared

Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors) chapters 2-3 — read these two before the next session

Sources

  1. Shalev-Shwartz & Ben-David, Understanding Machine Learning: From Theory to Algorithms (free PDF from the authors)
  2. Bousquet, Boucheron & Lugosi, Introduction to Statistical Learning Theory — Advanced Lectures on Machine Learning, Springer LNCS 3176 (2004), pp. 169-207
  3. Mohri, Rostamizadeh & Talwalkar, Foundations of Machine Learning, 2nd edition — MIT Press, 2018
  4. Hoeffding, Probability Inequalities for Sums of Bounded Random Variables, Journal of the American Statistical Association 58(301), 13-30 (1963)
  5. Companion deck in this library: Concentration Inequalities and Generalization, which derives Markov, Chebyshev and Hoeffding in full — slides/decks/usaaio/lesson-34-concentration-generalization.json

Want this taught 1-on-1? Alexander tutors Statistical Learning Theory — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108