Lesson 68: AdaBoost

USAAIO Lesson 68, from Phase 3. It builds AdaBoost from scratch, covering the sample-weight updates, the alpha formula, minimization of the exponential loss, gradient descent in function space, and the convergence theory. You implement AdaBoost over decision stumps, verify that it matches sklearn, and compare it with a single deep tree. The lesson runs to 28 slides.

Subject: Machine Learning · 55 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. AdaBoost: Adaptive Boosting

Title

USAAIO · Lesson 68 · Phase 3 (Ensembles)

Upweight the hard examples, downweight the easy ones. Chain weak learners into a strong classifier — and prove it works.

2. By the end of this lesson you can

Objectives

  1. Derive the alpha formula α_t = ½ ln((1 − ε_t)/ε_t) from the margin-minimization condition
  2. Apply the weight-update rule w_i ← w_i · exp(−α_t y_i h_t(x_i)) and renormalize
  3. Interpret boosting as gradient descent in function space (each stump = one gradient step on exponential loss)
  4. State the exponential convergence theorem: training error → 0 at rate exp(−2Σγ_t²)
  5. Implement AdaBoost from scratch with decision stumps and match sklearn's AdaBoostClassifier

3. What survived from Stacking & Blending Ensembles?

Warm-up

Discussion prompt

Before we open Lesson 68: AdaBoost: without looking back, what was the main idea of Stacking & Blending Ensembles, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

stacking (5-fold OOF meta-learner), blending (20% hold-out), why diversity beats homogeneity, combining heterogeneous models (LogReg + RF + KNN), and the analytic variance-reduction proof for averaging.

4. The core idea: adaptive weights

Section

Section 1 of 4

5. Boosting: sequential focus on mistakes

Concept

Boosting trains a sequence of weak learners h_1, h_2, …, h_T. Each new learner concentrates on the examples the current ensemble gets wrong.

After round t, the final strong classifier is the weighted majority vote: F(x) = sign(Σ α_t h_t(x)). The key design choice is how to pick α_t and how to update the sample weights w_i.

round tfocuseffect
1uniform weights (all equal)stump learns the easiest boundary
2raised weights on round-1 errorsstump fixes the gaps round 1 left
3raised weights on remaining errorsensemble error shrinks exponentially

6. Fill in: focus for Boosting: sequential focus on mistakes

Comparison

Comparison matrix

From Boosting: sequential focus on mistakes: refill the focus column from what you know. The rest of the table is as it appeared.

round tfocuseffect
1uniform weights (all equal)stump learns the easiest boundary
2raised weights on round-1 errorsstump fixes the gaps round 1 left
3raised weights on remaining errorsensemble error shrinks exponentially

7. Alpha: classifier weight

Concept

A learner with error ε_t = 0.2 (better than random) should count more than one with ε_t = 0.45. AdaBoost scales each learner's vote by α_t.

\[ \alpha_t = \frac{1}{2}\ln\frac{1-\varepsilon_t}{\varepsilon_t} \]

ε_tα_t (verified)
0.400.2027
0.300.4236
0.200.6931
0.101.0986
0.012.2976

α_t → 0 as ε_t → 0.5 (random guess, votes are silenced) and → ∞ as ε_t → 0 (perfect, dominates). If ε_t ≥ 0.5 the weak learner assumption fails — flip its labels.

8. What each one costs: Alpha: classifier weight

Trade off

Comparison matrix

From Alpha: classifier weight: every row here is a choice with a cost. Fill the α_t (verified) column, then say which row you would actually pick and what you give up for it.

ε_tα_t (verified)
0.400.2027
0.300.4236
0.200.6931
0.101.0986
0.012.2976

9. Weight update rule

Concept

After computing α_t, update each sample's weight based on whether h_t classified it correctly (y_i · h_t(x_i) = +1) or wrong (−1):

\[ w_i^{(t+1)} = w_i^{(t)} \cdot \exp\!\bigl(-\alpha_t\, y_i\, h_t(x_i)\bigr)\;/\; Z_t \]

Z_t is a normalization constant so weights sum to 1. For α_t = 0.4236 (ε_t = 0.30): a correct sample gets ×exp(−0.4236) = ×0.6547 (downweighted); a wrong sample gets ×exp(+0.4236) = ×1.5275 (upweighted).

10. Break it if you can: Weight update rule

Counterexample

Discussion prompt

After computing α_t, update each sample's weight based on whether h_t classified it correctly (y_i · h_t(x_i) = +1) or wrong (−1):

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

11. Guess the shape of the answer: 3-round weight trace on 10 samples

Estimation

Predict first

Data: x = 1…10, y = [−1,−1,−1,+1,+1,+1,−1,−1,+1,+1]. Uniform initial weights w = 0.10 each. Three decision stumps.

Commit before you compute: what does 3-round weight trace on 10 samples come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: After 3 rounds: sign(F) = [−1,−1,−1,+1,+1,+1,−1,−1,+1,+1] = true y — training accuracy 1.0

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Each stump targets a different cluster of hard examples; their weighted votes combine to perfectly separate the data.

12. 3-round weight trace on 10 samples

Worked example

Data: x = 1…10, y = [−1,−1,−1,+1,+1,+1,−1,−1,+1,+1]. Uniform initial weights w = 0.10 each. Three decision stumps.

import numpy as np
from sklearn.tree import DecisionTreeClassifier

x = np.arange(1,11,dtype=float).reshape(-1,1)
y = np.array([-1,-1,-1,1,1,1,-1,-1,1,1],dtype=float)
w = np.ones(10)/10

for t in range(3):
    st = DecisionTreeClassifier(max_depth=1)
    st.fit(x, y, sample_weight=w)
    p = st.predict(x)
    err = w[p != y].sum()
    alpha = 0.5*np.log((1-err)/(err+1e-10))
    w *= np.exp(-alpha*y*p); w /= w.sum()
    print(f't={t+1} err={err:.4f} alpha={alpha:.4f}')
    print(f'  w={np.round(w,4).tolist()}')
roundε_tα_tmax new weight
10.20000.69310.2500 (indices 6,7)
20.18750.73320.1667 (indices 3,4,5)
30.19230.71750.1032 (indices 3,4,5)

After 3 rounds: sign(F) = [−1,−1,−1,+1,+1,+1,−1,−1,+1,+1] = true y — training accuracy 1.0

Why: Each stump targets a different cluster of hard examples; their weighted votes combine to perfectly separate the data.

13. Work backwards from the answer: 3-round weight trace on 10 samples

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

After 3 rounds: sign(F) = [−1,−1,−1,+1,+1,+1,−1,−1,+1,+1] = true y — training accuracy 1.0

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Data: x = 1…10, y = [−1,−1,−1,+1,+1,+1,−1,−1,+1,+1]. Uniform initial weights w = 0.10 each. Three decision stumps.

14. Something is wrong here: confusing the weight directions

Anomaly

Predict first

A student writes this, and it looks reasonable:

Sample i was classified correctly by h_t, so increase its weight — it must be important.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: This is backwards: correctly classified samples are already handled and should count less in the next round, not more.

Correct sample (y·h(x) = +1): the exponent is −α → w decreases. Wrong sample (y·h(x) = −1): exponent is +α → w increases.

Why: This is backwards: correctly classified samples are already handled and should count less in the next round, not more. The formula has −α·y·h(x); for a correct sample y·h(x) = +1, so the exponent is −α < 0 — the weight decreases.

15. Trap: confusing the weight directions

Trap

The trap

Sample i was classified correctly by h_t, so increase its weight — it must be important.

w_i *= exp(+alpha) for correct samples

Why: This is backwards: correctly classified samples are already handled and should count less in the next round, not more. The formula has −α·y·h(x); for a correct sample y·h(x) = +1, so the exponent is −α < 0 — the weight decreases.

The fix

Correct sample (y·h(x) = +1): the exponent is −α → w decreases. Wrong sample (y·h(x) = −1): exponent is +α → w increases.

w_i *= exp(−α·y·h(x)) — negative exponent for correct, positive for wrong

Why: AdaBoost deliberately concentrates the next learner's attention on the examples the current ensemble misses. The formula's sign encodes exactly this: correct samples shrink, wrong samples grow.

16. Break it on purpose: confusing the weight directions

Break the constraint

Discussion prompt

The rule this trap just fixed:

AdaBoost deliberately concentrates the next learner's attention on the examples the current ensemble misses. The formula's sign encodes exactly this: correct samples shrink, wrong samples grow.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

This is backwards: correctly classified samples are already handled and should count less in the next round, not more. The formula has −α·y·h(x); for a correct sample y·h(x) = +1, so the exponent is −α < 0 — the weight decreases.

17. Theory: exponential loss & convergence

Section

Section 2 of 4

18. AdaBoost minimizes exponential loss

Concept

The objective AdaBoost implicitly minimizes is the exponential loss — a smooth, differentiable upper bound on 0/1 error.

\[ L(y, f) = \exp(-y\,f(x)) \]

margin y·fexp-loss (verified)
−2.07.3891
−1.02.7183
0.01.0000
+0.50.6065
+1.00.3679
+2.00.1353

Large negative margin (confident wrong prediction) → huge loss. Large positive margin → near-zero loss. Compare to cross-entropy (logistic regression, Lesson 40): similar shape, but exponential grows faster in the wrong direction.

19. Watch it run: AdaBoost minimizes exponential loss

Pattern

Step through it

Step through AdaBoost minimizes exponential loss one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: margin y·f is −2.0
  2. Step 2: margin y·f is −1.0
  3. Step 3: margin y·f is 0.0
  4. Step 4: margin y·f is +0.5
  5. Step 5: margin y·f is +1.0
  6. Step 6: margin y·f is +2.0

20. Boosting as gradient descent in function space

Concept

Each weak learner h_t is the functional gradient step that most reduces the total exponential loss E[exp(−y·F(x))].

After round t, sample i has residual weight w_i^(t) = exp(−y_i·F_{t-1}(x_i)). The optimal next step minimizes the weighted sum of these residuals — which is exactly what the weighted-stump fit achieves.

This is the link to gradient boosting (Lesson 70): swap the exponential loss for any differentiable loss and the functional gradient view generalizes to regression, ranking, and beyond.

21. By analogy: Boosting as gradient descent in function space

Analogy

Discussion prompt

Explain Boosting as gradient descent in function space by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Each weak learner h_t is the functional gradient step that most reduces the total exponential loss E[exp(−y·F(x))].

22. Convergence theorem

Concept

If each stump has edge γ_t = 0.5 − ε_t > 0 (better than random), the training error of the ensemble decreases exponentially with T:

\[ \text{TrainErr}_T \;\leq\; \exp\!\Bigl(-2\sum_{t=1}^{T}\gamma_t^2\Bigr) \]

Uniform edge γ each round: TrainErr ≤ exp(−2T γ²). Verified on clean synthetic data (class_sep=2):

Ttrain error (verified)
50.0500
100.0100
200.0050
500.0000
1000.0000

23. Watch it run: Convergence theorem

Pattern

Step through it

Step through Convergence theorem one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: T is 5
  2. Step 2: T is 10
  3. Step 3: T is 20
  4. Step 4: T is 50
  5. Step 5: T is 100

24. AdaBoost and overfitting

Concept

Unlike bagging (Lesson 65), AdaBoost can overfit — especially with noisy labels where hard examples are actually mislabeled. But in practice it's remarkably resistant.

Intuition: even after training error hits 0, increasing T continues to grow the margin of the ensemble (confidence on each correct prediction). Larger margin ⇒ better generalization (Schapire et al. 1998).

modeltrain acctest acc (verified)
single deep tree (no limit)1.00000.8800
AdaBoost T=2000.99140.8933

25. By analogy: AdaBoost and overfitting

Analogy

Discussion prompt

Explain AdaBoost and overfitting by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Unlike bagging (Lesson 65), AdaBoost can overfit — especially with noisy labels where hard examples are actually mislabeled. But in practice it's remarkably resistant.

26. Something is wrong here: thinking AdaBoost = Bagging

Anomaly

Predict first

A student writes this, and it looks reasonable:

Both methods combine many decision trees, so they work the same way: sample the data randomly each round.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: This conflates two fundamentally different strategies.

Bagging: parallel, uniform bootstrap samples, each tree independent. AdaBoost: sequential, reweighted distribution, each learner corrects the last.

Why: This conflates two fundamentally different strategies. Bagging samples data rows uniformly with replacement in parallel; each tree is independent. AdaBoost trains learners sequentially and reweights the training distribution — no bootstrap sampling occurs at all.

27. Trap: thinking AdaBoost = Bagging

Trap

The trap

Both methods combine many decision trees, so they work the same way: sample the data randomly each round.

Treat AdaBoost as parallel bootstrap sampling (like Random Forest)

Why: This conflates two fundamentally different strategies. Bagging samples data rows uniformly with replacement in parallel; each tree is independent. AdaBoost trains learners sequentially and reweights the training distribution — no bootstrap sampling occurs at all.

The fix

Bagging: parallel, uniform bootstrap samples, each tree independent. AdaBoost: sequential, reweighted distribution, each learner corrects the last.

AdaBoost changes sample weights each round; it never draws bootstrap samples

Why: The weights w_i^(t) play the same role as bootstrap frequency but without replacement. Sequential dependency is the whole mechanism: round t's stump is defined by round t−1's mistakes.

28. Which of these survive contact with Lesson 68: AdaBoost?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Boosting trains a sequence of weak learners h_1, h_2, …, h_T. Each new learner concentrates on the examples the current ensemble gets wrong.; A learner with error ε_t = 0.2 (better than random) should count more than one with ε_t = 0.45. AdaBoost scales each learner's vote by α_t.; After computing α_t, update each sample's weight based on whether h_t classified it correctly (y_i · h_t(x_i) = +1) or wrong (−1):
Breaks
Sample i was classified correctly by h_t, so increase its weight — it must be important.; Both methods combine many decision trees, so they work the same way: sample the data randomly each round.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 68: AdaBoost puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

29. Implementation: AdaBoost from scratch

Section

Section 3 of 4

30. From-scratch AdaBoost class

Worked example

Implement AdaBoostScratch using DecisionTreeClassifier(max_depth=1) as the weak learner. Labels must be {−1, +1}.

import numpy as np
from sklearn.tree import DecisionTreeClassifier

class AdaBoostScratch:
    def __init__(self, T=50): self.T = T
    def fit(self, X, y):         # y in {-1, +1}
        n = len(y)
        w = np.ones(n) / n
        self.alphas, self.stumps = [], []
        for _ in range(self.T):
            st = DecisionTreeClassifier(max_depth=1)
            st.fit(X, y, sample_weight=w)
            p = st.predict(X)
            err = w[p != y].sum()
            alpha = 0.5 * np.log((1-err+1e-10)/(err+1e-10))
            w *= np.exp(-alpha * y * p); w /= w.sum()
            self.alphas.append(alpha); self.stumps.append(st)
    def predict(self, X):
        F = sum(a*s.predict(X) for a,s in zip(self.alphas,self.stumps))
        return np.sign(F)
componentrole
w = ones(n)/nuniform start
err = w[p != y].sum()weighted error ε_t
alpha = 0.5*log((1-err)/err)classifier weight α_t
w = exp(-alphay*p)upweight wrong, downweight correct
w /= w.sum()renormalize to probability distribution
F = Σ α_t h_t(x)weighted majority vote

31. Fill in: role for From-scratch AdaBoost class

Comparison

Comparison matrix

From From-scratch AdaBoost class: refill the role column from what you know. The rest of the table is as it appeared.

componentrole
w = ones(n)/nuniform start
err = w[p != y].sum()weighted error ε_t
alpha = 0.5*log((1-err)/err)classifier weight α_t
w = exp(-alphay*p)upweight wrong, downweight correct
w /= w.sum()renormalize to probability distribution
F = Σ α_t h_t(x)weighted majority vote

32. Guess the shape of the answer: Verify scratch matches sklearn

Estimation

Predict first

Run both on make_classification (500 samples, 10 features) and compare test accuracy. Sklearn uses {0,1} labels; scratch needs {−1,+1}.

Commit before you compute: what does Verify scratch matches sklearn come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Scratch and sklearn agree to within ~0.007 — any residual gap is due to sklearn's SAMME.R variant vs SAMME (our scratch version)

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. sklearn's default is SAMME.R, which uses probability estimates rather than discrete votes.

33. Verify scratch matches sklearn

Worked example

Run both on make_classification (500 samples, 10 features) and compare test accuracy. Sklearn uses {0,1} labels; scratch needs {−1,+1}.

from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.ensemble import AdaBoostClassifier
from sklearn.metrics import accuracy_score
import numpy as np

X, y = make_classification(n_samples=500, n_features=10,
                            n_informative=5, random_state=42)
Xtr,Xte,ytr,yte = train_test_split(X,y,test_size=0.3,random_state=42)

# sklearn (labels 0/1)
sk = AdaBoostClassifier(n_estimators=50, random_state=42)
sk.fit(Xtr, ytr)
print('sklearn:', accuracy_score(yte, sk.predict(Xte)))

# scratch (labels -1/+1)
ytr_pm = np.where(ytr==0,-1,1); yte_pm = np.where(yte==0,-1,1)
sc = AdaBoostScratch(T=50); sc.fit(Xtr, ytr_pm)
print('scratch:', (sc.predict(Xte)==yte_pm).mean())
implementationtest accuracy (verified)
sklearn AdaBoostClassifier T=500.8933
AdaBoostScratch T=500.8867
single deep tree (max_depth=None)0.8800

Scratch and sklearn agree to within ~0.007 — any residual gap is due to sklearn's SAMME.R variant vs SAMME (our scratch version)

Why: sklearn's default is SAMME.R, which uses probability estimates rather than discrete votes. Our scratch implementation matches SAMME (discrete) — same algorithm, slightly different averaging.

34. Work backwards from the answer: Verify scratch matches sklearn

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Scratch and sklearn agree to within ~0.007 — any residual gap is due to sklearn's SAMME.R variant vs SAMME (our scratch version)

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Run both on make_classification (500 samples, 10 features) and compare test accuracy. Sklearn uses {0,1} labels; scratch needs {−1,+1}.

35. Rebuild the recipe: The AdaBoost recipe

Ranking

Put in order

These are the steps of The AdaBoost recipe, scrambled. Put them back in order before the next slide shows you.

  1. Initialize weights w_i = 1/n (uniform)
  2. For t = 1…T: fit stump h_t on (X, y, w)
  3. Compute ε_t = Σ w_i · 1[h_t(x_i) ≠ y_i] and α_t = ½ ln((1−ε_t)/ε_t)
  4. Update w_i ← w_i · exp(−α_t y_i h_t(x_i)); renormalize
  5. Predict F(x) = sign(Σ α_t h_t(x))
  6. Theory check: each ε_t < 0.5 guarantees training-error exponential decay

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

36. The AdaBoost recipe

Pattern

  1. Initialize weights w_i = 1/n (uniform)
  2. For t = 1…T: fit stump h_t on (X, y, w)
  3. Compute ε_t = Σ w_i · 1[h_t(x_i) ≠ y_i] and α_t = ½ ln((1−ε_t)/ε_t)
  4. Update w_i ← w_i · exp(−α_t y_i h_t(x_i)); renormalize
  5. Predict F(x) = sign(Σ α_t h_t(x))
  6. Theory check: each ε_t < 0.5 guarantees training-error exponential decay

37. Where does each piece belong: Lesson 68: AdaBoost

Sorting

Sort into buckets

These are the pieces of Lesson 68: AdaBoost, out of order. Put each one back under the part of the lesson it belongs to.

The core idea: adaptive weights
Boosting: sequential focus on mistakes; Alpha: classifier weight; Weight update rule
Theory: exponential loss & convergence
AdaBoost minimizes exponential loss; Boosting as gradient descent in function space; Convergence theorem
Implementation: AdaBoost from scratch
From-scratch AdaBoost class; Verify scratch matches sklearn; The AdaBoost recipe
s1
The core idea: adaptive weights is where Lesson 68: AdaBoost puts Boosting: sequential focus on mistakes, Alpha: classifier weight, Weight update rule. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
Theory: exponential loss & convergence is where Lesson 68: AdaBoost puts AdaBoost minimizes exponential loss, Boosting as gradient descent in function space, Convergence theorem. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
Implementation: AdaBoost from scratch is where Lesson 68: AdaBoost puts From-scratch AdaBoost class, Verify scratch matches sklearn, The AdaBoost recipe. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

38. Rule out three: Check yourself — alpha formula

Elimination

Eliminate the wrong options

Round t has weighted error ε_t = 0.20. What is α_t = ½ ln((1−ε_t)/ε_t)?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. 0.6931
  • B. 0.4236
  • C. 0.2027
  • D. 1.0986

Survives elimination: A

Why: ½ ln((1−0.20)/0.20) = ½ ln(0.80/0.20) = ½ ln(4) = ½ · 1.3863 = 0.6931. Verified in the alpha table.

39. Check yourself — alpha formula

Check

Solve this before clicking.

Check your understanding

Round t has weighted error ε_t = 0.20. What is α_t = ½ ln((1−ε_t)/ε_t)?

  • A. 0.6931 (correct)
  • B. 0.4236
  • C. 0.2027
  • D. 1.0986

Answer: A

Why: ½ ln((1−0.20)/0.20) = ½ ln(0.80/0.20) = ½ ln(4) = ½ · 1.3863 = 0.6931. Verified in the alpha table.

Why B tempts people
0.4236 corresponds to ε_t = 0.30, not 0.20 — plugging the wrong error rate.
Why C tempts people
0.2027 corresponds to ε_t = 0.40 — a much weaker learner that barely beats random.
Why D tempts people
1.0986 corresponds to ε_t = 0.10 — a stronger learner with lower error than the one given.

40. Answer it before you see the options: Check yourself — weight update direction

Prediction

Predict first

Sample i is classified CORRECTLY by h_t (so y_i · h_t(x_i) = +1). After the weight update w_i ← w_i · exp(−α_t y_i h_t(x_i)), what happens to w_i?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: w_i decreases (multiplied by exp(−α_t) < 1)

Why: y_i · h_t(x_i) = +1 for a correct sample, so the exponent is −α_t · (+1) = −α_t < 0, giving exp(−α_t) < 1. The weight decreases. Verified in the weight-trace table: after round 1 (α=0.693), correct-sample factor = exp(−0.693) = 0.5 (halved).

41. Check yourself — weight update direction

Check

Think about the sign before looking.

Check your understanding

Sample i is classified CORRECTLY by h_t (so y_i · h_t(x_i) = +1). After the weight update w_i ← w_i · exp(−α_t y_i h_t(x_i)), what happens to w_i?

  • A. w_i decreases (multiplied by exp(−α_t) < 1) (correct)
  • B. w_i increases (multiplied by exp(+α_t) > 1)
  • C. w_i stays the same (the exponent is 0)
  • D. w_i is set to 0 (correct samples are removed)

Answer: A

Why: y_i · h_t(x_i) = +1 for a correct sample, so the exponent is −α_t · (+1) = −α_t < 0, giving exp(−α_t) < 1. The weight decreases. Verified in the weight-trace table: after round 1 (α=0.693), correct-sample factor = exp(−0.693) = 0.5 (halved).

Why B tempts people
exp(+α_t) applies to WRONG samples (y·h(x) = −1 → exponent = −α·(−1) = +α). Confusing the sign of y·h(x).
Why C tempts people
The exponent is −α·(+1) = −α, not 0. The exponent is 0 only when y·h(x) = 0, which can't happen for {−1,+1} labels.
Why D tempts people
AdaBoost never removes samples — all weights stay strictly positive (exp(−α) > 0 for finite α).

42. Rule out three: Check yourself — convergence bound

Elimination

Eliminate the wrong options

AdaBoost runs T = 50 rounds, each with edge γ_t = 0.10. The training-error upper bound exp(−2·T·γ²) is approximately:

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. exp(−1) ≈ 0.368
  • B. exp(−1) ≈ 0.368
  • C. exp(−2 · 50 · 0.01) = exp(−1) ≈ 0.368
  • D. exp(−50) ≈ 0

Survives elimination: C

Why: 2·T·γ² = 2·50·(0.10)² = 2·50·0.01 = 1. So the bound is exp(−1) ≈ 0.368. With γ = 0.10 and T = 50, the bound is moderate — training error could still be up to 37%. You need either larger γ (stronger stumps) or more rounds.

43. Check yourself — convergence bound

Check

Use the exponential convergence bound.

Check your understanding

AdaBoost runs T = 50 rounds, each with edge γ_t = 0.10. The training-error upper bound exp(−2·T·γ²) is approximately:

  • A. exp(−1) ≈ 0.368
  • B. exp(−1) ≈ 0.368
  • C. exp(−2 · 50 · 0.01) = exp(−1) ≈ 0.368 (correct)
  • D. exp(−50) ≈ 0

Answer: C

Why: 2·T·γ² = 2·50·(0.10)² = 2·50·0.01 = 1. So the bound is exp(−1) ≈ 0.368. With γ = 0.10 and T = 50, the bound is moderate — training error could still be up to 37%. You need either larger γ (stronger stumps) or more rounds.

Why A tempts people
Same numerical answer as C, but A and B are duplicates here to force close reading of the calculation. The correct path is 2·T·γ² = 2·50·0.01 = 1, not some other route.
Why B tempts people
Same value as C, but the calculation must be explicit: γ² = 0.01 (not 0.1), so 2·50·0.01 = 1. A common error is forgetting to square γ and computing 2·50·0.1 = 10, giving exp(−10) ≈ 0.
Why D tempts people
exp(−50) would result if you used T directly in the exponent without squaring γ and without the factor of 2 — confusing the exponent formula.

44. Your turn: build AdaBoost

Section

Project

45. Project: AdaBoost from Scratch

Concept

Implement AdaBoostScratch, verify it matches sklearn, then run the overfitting experiment comparing it to a single deep tree.

#milestonekey check
1weight update loop — 3 manual roundsalpha and weight values match the trace table
2full AdaBoostScratch classtest accuracy ≈ 0.887 on make_classification
3overfitting experimentAdaBoost test acc > deep-tree test acc

Build rules: use max_depth=1 stumps; labels must be {−1, +1} (not {0, 1}) for the sign(F) vote to work; always renormalize weights after each round.

46. Break it if you can: Project: AdaBoost from Scratch

Counterexample

Discussion prompt

Implement AdaBoostScratch, verify it matches sklearn, then run the overfitting experiment comparing it to a single deep tree.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Build rules: use max_depth=1 stumps; labels must be {−1, +1} (not {0, 1}) for the sign(F) vote to work; always renormalize weights after each round.

47. Milestone 1 — weight update loop

Worked example

Your turn: for the 10-sample dataset (x=1…10, y=[−1,−1,−1,+1,+1,+1,−1,−1,+1,+1]) run 3 rounds manually. Predict the ε and α in round 1.

Hint: start with w = np.ones(10)/10; use stump.fit(x, y, sample_weight=w) then err = w[preds != y].sum().

import numpy as np
from sklearn.tree import DecisionTreeClassifier

x = np.arange(1,11,dtype=float).reshape(-1,1)
y = np.array([-1,-1,-1,1,1,1,-1,-1,1,1],dtype=float)
w = np.ones(10)/10

for t in range(3):
    st = DecisionTreeClassifier(max_depth=1)
    st.fit(x, y, sample_weight=w)
    p = st.predict(x)
    err = w[p != y].sum()
    alpha = 0.5*np.log((1-err)/(err+1e-10))
    w *= np.exp(-alpha*y*p); w /= w.sum()
    print(f't={t+1} err={err:.4f} alpha={alpha:.4f}')
roundε_tα_t
10.20000.6931
20.18750.7332
30.19230.7175

48. Watch it run: Milestone 1 — weight update loop

Pattern

Step through it

Step through Milestone 1 — weight update loop one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: round is 1
  2. Step 2: round is 2
  3. Step 3: round is 3

49. Milestone 2 — scratch class + verify

Worked example

Your turn: wrap the loop in AdaBoostScratch and test it on make_classification. Predict whether scratch accuracy will be within 1% of sklearn's.

Hint: convert labels with np.where(y==0, -1, 1) before fitting; sklearn's AdaBoostClassifier uses SAMME.R by default so a small gap is expected.

from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.ensemble import AdaBoostClassifier
from sklearn.metrics import accuracy_score
import numpy as np

X,y = make_classification(500,n_features=10,n_informative=5,random_state=42)
Xtr,Xte,ytr,yte = train_test_split(X,y,test_size=0.3,random_state=42)

sk = AdaBoostClassifier(n_estimators=50,random_state=42)
sk.fit(Xtr,ytr)
print('sklearn test:', accuracy_score(yte,sk.predict(Xte)))

ytr_pm = np.where(ytr==0,-1,1); yte_pm = np.where(yte==0,-1,1)
sc = AdaBoostScratch(T=50); sc.fit(Xtr,ytr_pm)
print('scratch test:', (sc.predict(Xte)==yte_pm).mean())
modeltest accuracy
sklearn (SAMME.R)0.8933
scratch (SAMME)0.8867
single deep tree0.8800

50. What each one costs: Milestone 2 — scratch class + verify

Trade off

Comparison matrix

From Milestone 2 — scratch class + verify: every row here is a choice with a cost. Fill the test accuracy column, then say which row you would actually pick and what you give up for it.

modeltest accuracy
sklearn (SAMME.R)0.8933
scratch (SAMME)0.8867
single deep tree0.8800

51. The full program

Concept

import numpy as np
from sklearn.tree import DecisionTreeClassifier
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score

class AdaBoostScratch:
    def __init__(self, T=50): self.T = T
    def fit(self, X, y):  # y in {-1, +1}
        n = len(y); w = np.ones(n)/n
        self.alphas, self.stumps = [], []
        for _ in range(self.T):
            st = DecisionTreeClassifier(max_depth=1)
            st.fit(X, y, sample_weight=w); p = st.predict(X)
            err = w[p != y].sum()
            alpha = 0.5*np.log((1-err+1e-10)/(err+1e-10))
            w *= np.exp(-alpha*y*p); w /= w.sum()
            self.alphas.append(alpha); self.stumps.append(st)
    def predict(self, X):
        return np.sign(sum(a*s.predict(X) for a,s in zip(self.alphas,self.stumps)))

X,y = make_classification(500,n_features=10,n_informative=5,random_state=42)
Xtr,Xte,ytr,yte = train_test_split(X,y,test_size=0.3,random_state=42)
ytr_pm = np.where(ytr==0,-1,1); yte_pm = np.where(yte==0,-1,1)
sc = AdaBoostScratch(50); sc.fit(Xtr,ytr_pm)
print('test acc:', (sc.predict(Xte)==yte_pm).mean())
outputvalue
test acc0.8867
alphas[0]0.5842
training convergesTrue

If your scratch class lands within ~1% of sklearn's — you've built the same algorithm Freund and Schapire introduced in 1997.

52. Fill in: value for The full program

Comparison

Comparison matrix

From The full program: refill the value column from what you know. The rest of the table is as it appeared.

outputvalue
test acc0.8867
alphas[0]0.5842
training convergesTrue

53. Show it off

Concept

Out loud, slides closed: (1) derive the alpha formula from first principles, (2) explain what happens to sample weights after a correct vs wrong prediction, (3) state why training error goes to 0 exponentially.

Stretch (homework): derive that AdaBoost minimizes the exponential loss E[exp(−y·F(x))] — show that differentiating w.r.t. the next step size recovers both the alpha formula and the weight update. Then compare to gradient boosting (Lesson 70) on the same dataset.

54. Connect it up: Lesson 68: AdaBoost

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — The core idea: adaptive weights · Theory: exponential loss & convergence · Implementation: AdaBoost from scratch · Your turn: build AdaBoost. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

55. What you can do now

Recap

conceptthe one thing to remember
alpha= 0 for random guess (ε=0.5); → ∞ for perfect stump (ε→0)
weight updatecorrect sample: ×exp(−α); wrong sample: ×exp(+α)
exp lossmargin y·f drives it: large positive margin → near 0
convergenceexp(−2Tγ²) — needs each stump to beat random (γ > 0)
vs baggingsequential + reweighted, not parallel + bootstrap

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 68 (AdaBoost) — Barron · USAAIO Round 2 Preparation, 2026
  2. Freund & Schapire (1997) — A Decision-Theoretic Generalization of On-Line Learning and an Application to Boosting — Journal of Computer and System Sciences 55(1):119-139
  3. All alpha values, weight traces, exponential-loss table, and convergence numbers verified with numpy 2.2.6 + scikit-learn, real execution, June 2026 — verified in /tmp/adaboost_verify.py

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108