USAAIO Lesson 68, from Phase 3. It builds AdaBoost from scratch, covering the sample-weight updates, the alpha formula, minimization of the exponential loss, gradient descent in function space, and the convergence theory. You implement AdaBoost over decision stumps, verify that it matches sklearn, and compare it with a single deep tree. The lesson runs to 28 slides.
Subject: Machine Learning · 55 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 68 · Phase 3 (Ensembles)
Upweight the hard examples, downweight the easy ones. Chain weak learners into a strong classifier — and prove it works.
Objectives
AdaBoostClassifierWarm-up
Discussion prompt
Before we open Lesson 68: AdaBoost: without looking back, what was the main idea of Stacking & Blending Ensembles, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
stacking (5-fold OOF meta-learner), blending (20% hold-out), why diversity beats homogeneity, combining heterogeneous models (LogReg + RF + KNN), and the analytic variance-reduction proof for averaging.
Section
Section 1 of 4
Concept
Boosting trains a sequence of weak learners h_1, h_2, …, h_T. Each new learner concentrates on the examples the current ensemble gets wrong.
After round t, the final strong classifier is the weighted majority vote: F(x) = sign(Σ α_t h_t(x)). The key design choice is how to pick α_t and how to update the sample weights w_i.
| round t | focus | effect |
|---|---|---|
| 1 | uniform weights (all equal) | stump learns the easiest boundary |
| 2 | raised weights on round-1 errors | stump fixes the gaps round 1 left |
| 3 | raised weights on remaining errors | ensemble error shrinks exponentially |
Comparison
Comparison matrix
From Boosting: sequential focus on mistakes: refill the focus column from what you know. The rest of the table is as it appeared.
| round t | focus | effect |
|---|---|---|
| 1 | uniform weights (all equal) | stump learns the easiest boundary |
| 2 | raised weights on round-1 errors | stump fixes the gaps round 1 left |
| 3 | raised weights on remaining errors | ensemble error shrinks exponentially |
Concept
A learner with error ε_t = 0.2 (better than random) should count more than one with ε_t = 0.45. AdaBoost scales each learner's vote by α_t.
\[ \alpha_t = \frac{1}{2}\ln\frac{1-\varepsilon_t}{\varepsilon_t} \]
| ε_t | α_t (verified) |
|---|---|
| 0.40 | 0.2027 |
| 0.30 | 0.4236 |
| 0.20 | 0.6931 |
| 0.10 | 1.0986 |
| 0.01 | 2.2976 |
α_t → 0 as ε_t → 0.5 (random guess, votes are silenced) and → ∞ as ε_t → 0 (perfect, dominates). If ε_t ≥ 0.5 the weak learner assumption fails — flip its labels.
Trade off
Comparison matrix
From Alpha: classifier weight: every row here is a choice with a cost. Fill the α_t (verified) column, then say which row you would actually pick and what you give up for it.
| ε_t | α_t (verified) |
|---|---|
| 0.40 | 0.2027 |
| 0.30 | 0.4236 |
| 0.20 | 0.6931 |
| 0.10 | 1.0986 |
| 0.01 | 2.2976 |
Concept
After computing α_t, update each sample's weight based on whether h_t classified it correctly (y_i · h_t(x_i) = +1) or wrong (−1):
\[ w_i^{(t+1)} = w_i^{(t)} \cdot \exp\!\bigl(-\alpha_t\, y_i\, h_t(x_i)\bigr)\;/\; Z_t \]
Z_t is a normalization constant so weights sum to 1. For α_t = 0.4236 (ε_t = 0.30): a correct sample gets ×exp(−0.4236) = ×0.6547 (downweighted); a wrong sample gets ×exp(+0.4236) = ×1.5275 (upweighted).
Counterexample
Discussion prompt
After computing α_t, update each sample's weight based on whether h_t classified it correctly (y_i · h_t(x_i) = +1) or wrong (−1):
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Estimation
Predict first
Data: x = 1…10, y = [−1,−1,−1,+1,+1,+1,−1,−1,+1,+1]. Uniform initial weights w = 0.10 each. Three decision stumps.
Commit before you compute: what does 3-round weight trace on 10 samples come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: After 3 rounds: sign(F) = [−1,−1,−1,+1,+1,+1,−1,−1,+1,+1] = true y — training accuracy 1.0
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Each stump targets a different cluster of hard examples; their weighted votes combine to perfectly separate the data.
Worked example
Data: x = 1…10, y = [−1,−1,−1,+1,+1,+1,−1,−1,+1,+1]. Uniform initial weights w = 0.10 each. Three decision stumps.
import numpy as np
from sklearn.tree import DecisionTreeClassifier
x = np.arange(1,11,dtype=float).reshape(-1,1)
y = np.array([-1,-1,-1,1,1,1,-1,-1,1,1],dtype=float)
w = np.ones(10)/10
for t in range(3):
st = DecisionTreeClassifier(max_depth=1)
st.fit(x, y, sample_weight=w)
p = st.predict(x)
err = w[p != y].sum()
alpha = 0.5*np.log((1-err)/(err+1e-10))
w *= np.exp(-alpha*y*p); w /= w.sum()
print(f't={t+1} err={err:.4f} alpha={alpha:.4f}')
print(f' w={np.round(w,4).tolist()}')| round | ε_t | α_t | max new weight |
|---|---|---|---|
| 1 | 0.2000 | 0.6931 | 0.2500 (indices 6,7) |
| 2 | 0.1875 | 0.7332 | 0.1667 (indices 3,4,5) |
| 3 | 0.1923 | 0.7175 | 0.1032 (indices 3,4,5) |
After 3 rounds: sign(F) = [−1,−1,−1,+1,+1,+1,−1,−1,+1,+1] = true y — training accuracy 1.0
Why: Each stump targets a different cluster of hard examples; their weighted votes combine to perfectly separate the data.
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
After 3 rounds: sign(F) = [−1,−1,−1,+1,+1,+1,−1,−1,+1,+1] = true y — training accuracy 1.0
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Data: x = 1…10, y = [−1,−1,−1,+1,+1,+1,−1,−1,+1,+1]. Uniform initial weights w = 0.10 each. Three decision stumps.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Sample i was classified correctly by h_t, so increase its weight — it must be important.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: This is backwards: correctly classified samples are already handled and should count less in the next round, not more.
Correct sample (y·h(x) = +1): the exponent is −α → w decreases. Wrong sample (y·h(x) = −1): exponent is +α → w increases.
Why: This is backwards: correctly classified samples are already handled and should count less in the next round, not more. The formula has −α·y·h(x); for a correct sample y·h(x) = +1, so the exponent is −α < 0 — the weight decreases.
Trap
Sample i was classified correctly by h_t, so increase its weight — it must be important.
w_i *= exp(+alpha) for correct samples
Why: This is backwards: correctly classified samples are already handled and should count less in the next round, not more. The formula has −α·y·h(x); for a correct sample y·h(x) = +1, so the exponent is −α < 0 — the weight decreases.
Correct sample (y·h(x) = +1): the exponent is −α → w decreases. Wrong sample (y·h(x) = −1): exponent is +α → w increases.
w_i *= exp(−α·y·h(x)) — negative exponent for correct, positive for wrong
Why: AdaBoost deliberately concentrates the next learner's attention on the examples the current ensemble misses. The formula's sign encodes exactly this: correct samples shrink, wrong samples grow.
Break the constraint
Discussion prompt
The rule this trap just fixed:
AdaBoost deliberately concentrates the next learner's attention on the examples the current ensemble misses. The formula's sign encodes exactly this: correct samples shrink, wrong samples grow.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
This is backwards: correctly classified samples are already handled and should count less in the next round, not more. The formula has −α·y·h(x); for a correct sample y·h(x) = +1, so the exponent is −α < 0 — the weight decreases.
Section
Section 2 of 4
Concept
The objective AdaBoost implicitly minimizes is the exponential loss — a smooth, differentiable upper bound on 0/1 error.
\[ L(y, f) = \exp(-y\,f(x)) \]
| margin y·f | exp-loss (verified) |
|---|---|
| −2.0 | 7.3891 |
| −1.0 | 2.7183 |
| 0.0 | 1.0000 |
| +0.5 | 0.6065 |
| +1.0 | 0.3679 |
| +2.0 | 0.1353 |
Large negative margin (confident wrong prediction) → huge loss. Large positive margin → near-zero loss. Compare to cross-entropy (logistic regression, Lesson 40): similar shape, but exponential grows faster in the wrong direction.
Pattern
Step through it
Step through AdaBoost minimizes exponential loss one row at a time. What is driving the change, and what would the row after the last one be?
Concept
Each weak learner h_t is the functional gradient step that most reduces the total exponential loss E[exp(−y·F(x))].
After round t, sample i has residual weight w_i^(t) = exp(−y_i·F_{t-1}(x_i)). The optimal next step minimizes the weighted sum of these residuals — which is exactly what the weighted-stump fit achieves.
This is the link to gradient boosting (Lesson 70): swap the exponential loss for any differentiable loss and the functional gradient view generalizes to regression, ranking, and beyond.
Analogy
Discussion prompt
Explain Boosting as gradient descent in function space by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Each weak learner h_t is the functional gradient step that most reduces the total exponential loss E[exp(−y·F(x))].
Concept
If each stump has edge γ_t = 0.5 − ε_t > 0 (better than random), the training error of the ensemble decreases exponentially with T:
\[ \text{TrainErr}_T \;\leq\; \exp\!\Bigl(-2\sum_{t=1}^{T}\gamma_t^2\Bigr) \]
Uniform edge γ each round: TrainErr ≤ exp(−2T γ²). Verified on clean synthetic data (class_sep=2):
| T | train error (verified) |
|---|---|
| 5 | 0.0500 |
| 10 | 0.0100 |
| 20 | 0.0050 |
| 50 | 0.0000 |
| 100 | 0.0000 |
Pattern
Step through it
Step through Convergence theorem one row at a time. What is driving the change, and what would the row after the last one be?
Concept
Unlike bagging (Lesson 65), AdaBoost can overfit — especially with noisy labels where hard examples are actually mislabeled. But in practice it's remarkably resistant.
Intuition: even after training error hits 0, increasing T continues to grow the margin of the ensemble (confidence on each correct prediction). Larger margin ⇒ better generalization (Schapire et al. 1998).
| model | train acc | test acc (verified) |
|---|---|---|
| single deep tree (no limit) | 1.0000 | 0.8800 |
| AdaBoost T=200 | 0.9914 | 0.8933 |
Analogy
Discussion prompt
Explain AdaBoost and overfitting by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Unlike bagging (Lesson 65), AdaBoost can overfit — especially with noisy labels where hard examples are actually mislabeled. But in practice it's remarkably resistant.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Both methods combine many decision trees, so they work the same way: sample the data randomly each round.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: This conflates two fundamentally different strategies.
Bagging: parallel, uniform bootstrap samples, each tree independent. AdaBoost: sequential, reweighted distribution, each learner corrects the last.
Why: This conflates two fundamentally different strategies. Bagging samples data rows uniformly with replacement in parallel; each tree is independent. AdaBoost trains learners sequentially and reweights the training distribution — no bootstrap sampling occurs at all.
Trap
Both methods combine many decision trees, so they work the same way: sample the data randomly each round.
Treat AdaBoost as parallel bootstrap sampling (like Random Forest)
Why: This conflates two fundamentally different strategies. Bagging samples data rows uniformly with replacement in parallel; each tree is independent. AdaBoost trains learners sequentially and reweights the training distribution — no bootstrap sampling occurs at all.
Bagging: parallel, uniform bootstrap samples, each tree independent. AdaBoost: sequential, reweighted distribution, each learner corrects the last.
AdaBoost changes sample weights each round; it never draws bootstrap samples
Why: The weights w_i^(t) play the same role as bootstrap frequency but without replacement. Sequential dependency is the whole mechanism: round t's stump is defined by round t−1's mistakes.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
Section
Section 3 of 4
Worked example
Implement AdaBoostScratch using DecisionTreeClassifier(max_depth=1) as the weak learner. Labels must be {−1, +1}.
import numpy as np
from sklearn.tree import DecisionTreeClassifier
class AdaBoostScratch:
def __init__(self, T=50): self.T = T
def fit(self, X, y): # y in {-1, +1}
n = len(y)
w = np.ones(n) / n
self.alphas, self.stumps = [], []
for _ in range(self.T):
st = DecisionTreeClassifier(max_depth=1)
st.fit(X, y, sample_weight=w)
p = st.predict(X)
err = w[p != y].sum()
alpha = 0.5 * np.log((1-err+1e-10)/(err+1e-10))
w *= np.exp(-alpha * y * p); w /= w.sum()
self.alphas.append(alpha); self.stumps.append(st)
def predict(self, X):
F = sum(a*s.predict(X) for a,s in zip(self.alphas,self.stumps))
return np.sign(F)| component | role |
|---|---|
| w = ones(n)/n | uniform start |
| err = w[p != y].sum() | weighted error ε_t |
| alpha = 0.5*log((1-err)/err) | classifier weight α_t |
| w = exp(-alphay*p) | upweight wrong, downweight correct |
| w /= w.sum() | renormalize to probability distribution |
| F = Σ α_t h_t(x) | weighted majority vote |
Comparison
Comparison matrix
From From-scratch AdaBoost class: refill the role column from what you know. The rest of the table is as it appeared.
| component | role |
|---|---|
| w = ones(n)/n | uniform start |
| err = w[p != y].sum() | weighted error ε_t |
| alpha = 0.5*log((1-err)/err) | classifier weight α_t |
| w = exp(-alphay*p) | upweight wrong, downweight correct |
| w /= w.sum() | renormalize to probability distribution |
| F = Σ α_t h_t(x) | weighted majority vote |
Estimation
Predict first
Run both on make_classification (500 samples, 10 features) and compare test accuracy. Sklearn uses {0,1} labels; scratch needs {−1,+1}.
Commit before you compute: what does Verify scratch matches sklearn come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Scratch and sklearn agree to within ~0.007 — any residual gap is due to sklearn's SAMME.R variant vs SAMME (our scratch version)
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. sklearn's default is SAMME.R, which uses probability estimates rather than discrete votes.
Worked example
Run both on make_classification (500 samples, 10 features) and compare test accuracy. Sklearn uses {0,1} labels; scratch needs {−1,+1}.
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.ensemble import AdaBoostClassifier
from sklearn.metrics import accuracy_score
import numpy as np
X, y = make_classification(n_samples=500, n_features=10,
n_informative=5, random_state=42)
Xtr,Xte,ytr,yte = train_test_split(X,y,test_size=0.3,random_state=42)
# sklearn (labels 0/1)
sk = AdaBoostClassifier(n_estimators=50, random_state=42)
sk.fit(Xtr, ytr)
print('sklearn:', accuracy_score(yte, sk.predict(Xte)))
# scratch (labels -1/+1)
ytr_pm = np.where(ytr==0,-1,1); yte_pm = np.where(yte==0,-1,1)
sc = AdaBoostScratch(T=50); sc.fit(Xtr, ytr_pm)
print('scratch:', (sc.predict(Xte)==yte_pm).mean())| implementation | test accuracy (verified) |
|---|---|
| sklearn AdaBoostClassifier T=50 | 0.8933 |
| AdaBoostScratch T=50 | 0.8867 |
| single deep tree (max_depth=None) | 0.8800 |
Scratch and sklearn agree to within ~0.007 — any residual gap is due to sklearn's SAMME.R variant vs SAMME (our scratch version)
Why: sklearn's default is SAMME.R, which uses probability estimates rather than discrete votes. Our scratch implementation matches SAMME (discrete) — same algorithm, slightly different averaging.
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Scratch and sklearn agree to within ~0.007 — any residual gap is due to sklearn's SAMME.R variant vs SAMME (our scratch version)
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Run both on make_classification (500 samples, 10 features) and compare test accuracy. Sklearn uses {0,1} labels; scratch needs {−1,+1}.
Ranking
Put in order
These are the steps of The AdaBoost recipe, scrambled. Put them back in order before the next slide shows you.
Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
Sorting
Sort into buckets
These are the pieces of Lesson 68: AdaBoost, out of order. Put each one back under the part of the lesson it belongs to.
Elimination
Eliminate the wrong options
Round t has weighted error ε_t = 0.20. What is α_t = ½ ln((1−ε_t)/ε_t)?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: ½ ln((1−0.20)/0.20) = ½ ln(0.80/0.20) = ½ ln(4) = ½ · 1.3863 = 0.6931. Verified in the alpha table.
Check
Solve this before clicking.
Check your understanding
Round t has weighted error ε_t = 0.20. What is α_t = ½ ln((1−ε_t)/ε_t)?
Answer: A
Why: ½ ln((1−0.20)/0.20) = ½ ln(0.80/0.20) = ½ ln(4) = ½ · 1.3863 = 0.6931. Verified in the alpha table.
Prediction
Predict first
Sample i is classified CORRECTLY by h_t (so y_i · h_t(x_i) = +1). After the weight update w_i ← w_i · exp(−α_t y_i h_t(x_i)), what happens to w_i?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: w_i decreases (multiplied by exp(−α_t) < 1)
Why: y_i · h_t(x_i) = +1 for a correct sample, so the exponent is −α_t · (+1) = −α_t < 0, giving exp(−α_t) < 1. The weight decreases. Verified in the weight-trace table: after round 1 (α=0.693), correct-sample factor = exp(−0.693) = 0.5 (halved).
Check
Think about the sign before looking.
Check your understanding
Sample i is classified CORRECTLY by h_t (so y_i · h_t(x_i) = +1). After the weight update w_i ← w_i · exp(−α_t y_i h_t(x_i)), what happens to w_i?
Answer: A
Why: y_i · h_t(x_i) = +1 for a correct sample, so the exponent is −α_t · (+1) = −α_t < 0, giving exp(−α_t) < 1. The weight decreases. Verified in the weight-trace table: after round 1 (α=0.693), correct-sample factor = exp(−0.693) = 0.5 (halved).
Elimination
Eliminate the wrong options
AdaBoost runs T = 50 rounds, each with edge γ_t = 0.10. The training-error upper bound exp(−2·T·γ²) is approximately:
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: C
Why: 2·T·γ² = 2·50·(0.10)² = 2·50·0.01 = 1. So the bound is exp(−1) ≈ 0.368. With γ = 0.10 and T = 50, the bound is moderate — training error could still be up to 37%. You need either larger γ (stronger stumps) or more rounds.
Check
Use the exponential convergence bound.
Check your understanding
AdaBoost runs T = 50 rounds, each with edge γ_t = 0.10. The training-error upper bound exp(−2·T·γ²) is approximately:
Answer: C
Why: 2·T·γ² = 2·50·(0.10)² = 2·50·0.01 = 1. So the bound is exp(−1) ≈ 0.368. With γ = 0.10 and T = 50, the bound is moderate — training error could still be up to 37%. You need either larger γ (stronger stumps) or more rounds.
Section
Project
Concept
Implement AdaBoostScratch, verify it matches sklearn, then run the overfitting experiment comparing it to a single deep tree.
| # | milestone | key check |
|---|---|---|
| 1 | weight update loop — 3 manual rounds | alpha and weight values match the trace table |
| 2 | full AdaBoostScratch class | test accuracy ≈ 0.887 on make_classification |
| 3 | overfitting experiment | AdaBoost test acc > deep-tree test acc |
Build rules: use max_depth=1 stumps; labels must be {−1, +1} (not {0, 1}) for the sign(F) vote to work; always renormalize weights after each round.
Counterexample
Discussion prompt
Implement AdaBoostScratch, verify it matches sklearn, then run the overfitting experiment comparing it to a single deep tree.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Build rules: use max_depth=1 stumps; labels must be {−1, +1} (not {0, 1}) for the sign(F) vote to work; always renormalize weights after each round.
Worked example
Your turn: for the 10-sample dataset (x=1…10, y=[−1,−1,−1,+1,+1,+1,−1,−1,+1,+1]) run 3 rounds manually. Predict the ε and α in round 1.
Hint: start with w = np.ones(10)/10; use stump.fit(x, y, sample_weight=w) then err = w[preds != y].sum().
import numpy as np
from sklearn.tree import DecisionTreeClassifier
x = np.arange(1,11,dtype=float).reshape(-1,1)
y = np.array([-1,-1,-1,1,1,1,-1,-1,1,1],dtype=float)
w = np.ones(10)/10
for t in range(3):
st = DecisionTreeClassifier(max_depth=1)
st.fit(x, y, sample_weight=w)
p = st.predict(x)
err = w[p != y].sum()
alpha = 0.5*np.log((1-err)/(err+1e-10))
w *= np.exp(-alpha*y*p); w /= w.sum()
print(f't={t+1} err={err:.4f} alpha={alpha:.4f}')| round | ε_t | α_t |
|---|---|---|
| 1 | 0.2000 | 0.6931 |
| 2 | 0.1875 | 0.7332 |
| 3 | 0.1923 | 0.7175 |
Pattern
Step through it
Step through Milestone 1 — weight update loop one row at a time. What is driving the change, and what would the row after the last one be?
Worked example
Your turn: wrap the loop in AdaBoostScratch and test it on make_classification. Predict whether scratch accuracy will be within 1% of sklearn's.
Hint: convert labels with np.where(y==0, -1, 1) before fitting; sklearn's AdaBoostClassifier uses SAMME.R by default so a small gap is expected.
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.ensemble import AdaBoostClassifier
from sklearn.metrics import accuracy_score
import numpy as np
X,y = make_classification(500,n_features=10,n_informative=5,random_state=42)
Xtr,Xte,ytr,yte = train_test_split(X,y,test_size=0.3,random_state=42)
sk = AdaBoostClassifier(n_estimators=50,random_state=42)
sk.fit(Xtr,ytr)
print('sklearn test:', accuracy_score(yte,sk.predict(Xte)))
ytr_pm = np.where(ytr==0,-1,1); yte_pm = np.where(yte==0,-1,1)
sc = AdaBoostScratch(T=50); sc.fit(Xtr,ytr_pm)
print('scratch test:', (sc.predict(Xte)==yte_pm).mean())| model | test accuracy |
|---|---|
| sklearn (SAMME.R) | 0.8933 |
| scratch (SAMME) | 0.8867 |
| single deep tree | 0.8800 |
Trade off
Comparison matrix
From Milestone 2 — scratch class + verify: every row here is a choice with a cost. Fill the test accuracy column, then say which row you would actually pick and what you give up for it.
| model | test accuracy |
|---|---|
| sklearn (SAMME.R) | 0.8933 |
| scratch (SAMME) | 0.8867 |
| single deep tree | 0.8800 |
Concept
import numpy as np
from sklearn.tree import DecisionTreeClassifier
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score
class AdaBoostScratch:
def __init__(self, T=50): self.T = T
def fit(self, X, y): # y in {-1, +1}
n = len(y); w = np.ones(n)/n
self.alphas, self.stumps = [], []
for _ in range(self.T):
st = DecisionTreeClassifier(max_depth=1)
st.fit(X, y, sample_weight=w); p = st.predict(X)
err = w[p != y].sum()
alpha = 0.5*np.log((1-err+1e-10)/(err+1e-10))
w *= np.exp(-alpha*y*p); w /= w.sum()
self.alphas.append(alpha); self.stumps.append(st)
def predict(self, X):
return np.sign(sum(a*s.predict(X) for a,s in zip(self.alphas,self.stumps)))
X,y = make_classification(500,n_features=10,n_informative=5,random_state=42)
Xtr,Xte,ytr,yte = train_test_split(X,y,test_size=0.3,random_state=42)
ytr_pm = np.where(ytr==0,-1,1); yte_pm = np.where(yte==0,-1,1)
sc = AdaBoostScratch(50); sc.fit(Xtr,ytr_pm)
print('test acc:', (sc.predict(Xte)==yte_pm).mean())| output | value |
|---|---|
| test acc | 0.8867 |
| alphas[0] | 0.5842 |
| training converges | True |
If your scratch class lands within ~1% of sklearn's — you've built the same algorithm Freund and Schapire introduced in 1997.
Comparison
Comparison matrix
From The full program: refill the value column from what you know. The rest of the table is as it appeared.
| output | value |
|---|---|
| test acc | 0.8867 |
| alphas[0] | 0.5842 |
| training converges | True |
Concept
Out loud, slides closed: (1) derive the alpha formula from first principles, (2) explain what happens to sample weights after a correct vs wrong prediction, (3) state why training error goes to 0 exponentially.
Stretch (homework): derive that AdaBoost minimizes the exponential loss E[exp(−y·F(x))] — show that differentiating w.r.t. the next step size recovers both the alpha formula and the weight update. Then compare to gradient boosting (Lesson 70) on the same dataset.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — The core idea: adaptive weights · Theory: exponential loss & convergence · Implementation: AdaBoost from scratch · Your turn: build AdaBoost. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
AdaBoostScratch matches sklearn within ~0.7% (SAMME vs SAMME.R)| concept | the one thing to remember |
|---|---|
| alpha | = 0 for random guess (ε=0.5); → ∞ for perfect stump (ε→0) |
| weight update | correct sample: ×exp(−α); wrong sample: ×exp(+α) |
| exp loss | margin y·f drives it: large positive margin → near 0 |
| convergence | exp(−2Tγ²) — needs each stump to beat random (γ > 0) |
| vs bagging | sequential + reweighted, not parallel + bootstrap |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.