USAAIO Lesson 50, from Phase 2. It covers the gradient-boosting framework F_m = F_{m-1} + eta*h_m and pseudo-residuals as the negative gradients of the loss - y − F for MSE, and y − sigma(F) for classification - then builds a three-round GBM regressor from scratch using an sklearn DecisionTree. It goes on to the XGBoost improvements, namely L2 leaf regularization, column subsampling, and approximate splits, and to LightGBM's histogram-based leaf-wise growth, ending with a cross-validated hyperparameter sweep. You build GradientBoostingRegressorFromScratch on make_regression data and compare the GBM with a random forest. The lesson runs to 29 slides.
Subject: Machine Learning · 60 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 50 · Phase 2
From residuals to an ensemble: the boosting framework, pseudo-residuals, and modern implementations — XGBoost, LightGBM.
Objectives
DecisionTreeRegressorWarm-up
Discussion prompt
Before we open Lesson 50: Gradient Boosting: without looking back, what was the main idea of Random Forests, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
bootstrap aggregation, random feature subsets (sqrt(d)), out-of-bag error as free validation, MDI vs permutation vs SHAP feature importance, and Random Forest vs single tree vs boosting tradeoffs. Implement RandomForestClassifier from scratch and compare OOB error to k-fold CV.
Section
Part 1 of 3
Concept
Boosting builds an ensemble by fitting each new learner to the mistakes of the current model — it is an additive procedure, not an average.
\[ F_m(x) = F_{m-1}(x) + \eta \, h_m(x) \]
Counterexample
Discussion prompt
Boosting builds an ensemble by fitting each new learner to the mistakes of the current model — it is an additive procedure, not an average.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Concept
Gradient boosting frames h_m as fitting the negative gradient of the loss L(y, F) with respect to F — this is the steepest-descent direction in function space.
\[ r_{im} = -\frac{\partial L(y_i,\, F_{m-1}(x_i))}{\partial F_{m-1}(x_i)} \]
| loss | pseudo-residual r_im | intuition |
|---|---|---|
| MSE ½(y−F)² | y − F | plain residual |
| Log-loss (binary) | y − σ(F) | prob error |
Comparison
Comparison matrix
From Pseudo-residuals: the negative gradient of the loss: refill the pseudo-residual r_im column from what you know. The rest of the table is as it appeared.
| loss | pseudo-residual r_im | intuition |
|---|---|---|
| MSE ½(y−F)² | y − F | plain residual |
| Log-loss (binary) | y − σ(F) | prob error |
Concept
For mean-squared error the derivation is one line — and explains why the homework says 'just the residuals!'
\[ L(y,F) = \frac{1}{2}(y - F)^2 \;\Longrightarrow\; -\frac{\partial L}{\partial F} = -\bigl(-(y-F)\bigr) = y - F \]
Fitting a tree to y − F (the ordinary residual) is gradient descent in function space — no separate math is needed.
Analogy
Discussion prompt
Explain Deriving the MSE pseudo-residual by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
For mean-squared error the derivation is one line — and explains why the homework says 'just the residuals!'
Concept
For binary classification GBM outputs a log-odds score F. The loss is log-loss, and the pseudo-residual involves the sigmoid.
\[ -\frac{\partial}{\partial F}\bigl[-y\ln\sigma(F)-(1-y)\ln(1-\sigma(F))\bigr] = y - \sigma(F) \]
| F (log-odds) | σ(F) | y | pseudo = y−σ(F) |
|---|---|---|---|
| 0.0 | 0.5000 | 1 | 0.5000 |
| 1.0 | 0.7311 | 1 | 0.2689 |
| -1.0 | 0.2689 | 0 | -0.2689 |
Trade off
Comparison matrix
From Classification pseudo-residuals: y − σ(F): every row here is a choice with a cost. Fill the pseudo = y−σ(F) column, then say which row you would actually pick and what you give up for it.
| F (log-odds) | σ(F) | y | pseudo = y−σ(F) |
|---|---|---|---|
| 0.0 | 0.5000 | 1 | 0.5000 |
| 1.0 | 0.7311 | 1 | 0.2689 |
| -1.0 | 0.2689 | 0 | -0.2689 |
Anomaly
Predict first
A student writes this, and it looks reasonable:
The GBM loop always fits the next tree to the residuals y − F, regardless of whether the task is regression or classification.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: This is only right for MSE regression.
The pseudo-residual formula depends on the loss function; derive it as −∂L/∂F for each task.
Why: This is only right for MSE regression. The formula y − F treats F as a predicted value, but for classification F is a log-odds score — the pseudo-residual must go through the sigmoid: r = y − σ(F).
Trap
The GBM loop always fits the next tree to the residuals y − F, regardless of whether the task is regression or classification.
Use r = y − F in both cases
Why: This is only right for MSE regression. The formula y − F treats F as a predicted value, but for classification F is a log-odds score — the pseudo-residual must go through the sigmoid: r = y − σ(F).
The pseudo-residual formula depends on the loss function; derive it as −∂L/∂F for each task.
Regression (MSE): r = y − F. Classification (log-loss): r = y − σ(F)
Why: Both follow from −∂L/∂F. MSE gives the plain residual; log-loss passes F through the sigmoid first. These are the same for regression only by coincidence.
Break the constraint
Discussion prompt
The rule this trap just fixed:
Both follow from −∂L/∂F. MSE gives the plain residual; log-loss passes F through the sigmoid first. These are the same for regression only by coincidence.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
This is only right for MSE regression. The formula y − F treats F as a predicted value, but for classification F is a log-odds score — the pseudo-residual must go through the sigmoid: r = y − σ(F).
Section
Part 2 of 3
Concept
Using depth-1 stumps and a small η (< 1) keeps each step modest — many weak learners combine into one strong one.
Explain it
Discussion prompt
Explain The from-scratch GBM algorithm to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Using depth-1 stumps and a small η (< 1) keeps each step modest — many weak learners combine into one strong one.
Estimation
Predict first
Trace gradient boosting on a tiny dataset with y = 2x and η = 0.5 so you can follow every update by hand.
Commit before you compute: what does Worked example — 3-round GBM (3 samples) come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: MSE falls 2.67 → 1.17 → 0.32 → 0.11 across three rounds
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Each round reduces the unexplained variance.
Worked example
Trace gradient boosting on a tiny dataset with y = 2x and η = 0.5 so you can follow every update by hand.
import numpy as np
from sklearn.tree import DecisionTreeRegressor
X = np.array([[1.0],[2.0],[3.0]])
y = np.array([2.0, 4.0, 6.0]) # y = 2x, clean
eta = 0.5
F = np.full(3, y.mean()) # F0 = 4.0 for all
for m in range(1, 4):
r = y - F # pseudo-residuals (MSE)
stump = DecisionTreeRegressor(max_depth=1, random_state=0)
stump.fit(X, r)
F = F + eta * stump.predict(X)
print(f'm={m}: r={r.round(3)}, F={F.round(3)}, MSE={np.mean((y-F)**2):.4f}')| m | r (pseudo-resids) | F after update | MSE |
|---|---|---|---|
| 0 (init) | [−2, 0, +2] | [4, 4, 4] | 2.6667 |
| 1 | [−2, 0, +2] | [3, 4.5, 4.5] | 1.1667 |
| 2 | [−1, −0.5, +1.5] | [2.625, 4.125, 5.25] | 0.3229 |
| 3 | [−0.625, −0.125, +0.75] | [2.438, 3.938, 5.625] | 0.1120 |
Round 1: stump splits on x≤1.5; left leaf predicts −2, right leaf predicts +1
Why: The tree fits the residuals, not y directly. η=0.5 moves only half-way to the residual to avoid over-correcting in a single step.
MSE falls 2.67 → 1.17 → 0.32 → 0.11 across three rounds
Why: Each round reduces the unexplained variance. With more rounds (and smaller η) F_M converges toward y.
Error analysis
Annotate
Walk the callouts on Worked example — 3-round GBM (3 samples). Each one is a place this is easy to get subtly wrong.
Missing information
Discussion prompt
Apply the same loop to the load_diabetes dataset and compare training RMSE vs sklearn's GradientBoostingRegressor.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Boosting directly targets residuals, so it tends to slightly outperform bagging-based forests on tabular data with the same depth — but needs careful regularization (n_estimators, learning_rate) to not overfit.
Worked example
Apply the same loop to the load_diabetes dataset and compare training RMSE vs sklearn's GradientBoostingRegressor.
from sklearn.datasets import load_diabetes
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeRegressor
from sklearn.ensemble import GradientBoostingRegressor
from sklearn.metrics import mean_squared_error
import numpy as np
X, y = load_diabetes(return_X_y=True)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.2, random_state=42)
# sklearn GBM
gbm = GradientBoostingRegressor(n_estimators=100, learning_rate=0.1, max_depth=3, random_state=42)
gbm.fit(Xtr, ytr)
print('GBM RMSE:', round(np.sqrt(mean_squared_error(yte, gbm.predict(Xte))), 2))| model | test RMSE |
|---|---|
| GBM (sklearn) | 53.73 |
| Random Forest | 54.33 |
GBM RMSE 53.73 vs Random Forest 54.33 on the diabetes holdout
Why: Boosting directly targets residuals, so it tends to slightly outperform bagging-based forests on tabular data with the same depth — but needs careful regularization (n_estimators, learning_rate) to not overfit.
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
GBM RMSE 53.73 vs Random Forest 54.33 on the diabetes holdout
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Apply the same loop to the load_diabetes dataset and compare training RMSE vs sklearn's GradientBoostingRegressor.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Use a large learning_rate (0.5) and many estimators — more trees always means a better model.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: The train RMSE collapses, which looks like perfect fit.
Use a small learning_rate and pick n_estimators by cross-validation.
Why: The train RMSE collapses, which looks like perfect fit. But test RMSE is 66.34 — worse than 10-estimator GBM (62.39). The model memorised training noise.
Trap
Use a large learning_rate (0.5) and many estimators — more trees always means a better model.
n_estimators=500, learning_rate=0.5, depth=5: train RMSE → 0.01
Why: The train RMSE collapses, which looks like perfect fit. But test RMSE is 66.34 — worse than 10-estimator GBM (62.39). The model memorised training noise.
Use a small learning_rate and pick n_estimators by cross-validation.
n_estimators=10, learning_rate=0.1, depth=5: train RMSE=40.27, test RMSE=62.39
Why: Low learning rate with cross-validated n_estimators trades per-step progress for generalisation. The learning_rate / n_estimators tradeoff: halving η and doubling M usually helps.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
Section
Part 3 of 3
Concept
XGBoost (Chen & Guestrin, 2016) extends vanilla GBM with three algorithmic improvements that dominate ML competitions.
| improvement | what it does | parameter |
|---|---|---|
| L2 leaf reg | penalises large leaf weights: Ω(h) = λΣwⱼ² | reg_lambda |
| Column subsampling | random feature subset per tree (like RF) | colsample_bytree |
| Approx split finding | histogram bins the feature values; O(n) not O(n log n) | built-in |
The regularised objective is Obj = ΣL(yᵢ, Fᵢ) + Σ[γTⱼ + ½λwⱼ²] where T is the number of leaves.
Concept
With L2 regularisation, the optimal leaf weight w* for a leaf j covering a set of samples J is closed-form:
\[ w_j^* = -\frac{\sum_{i \in J} g_i}{\sum_{i \in J} h_i + \lambda} \]
where gᵢ = ∂L/∂Fᵢ (gradient) and hᵢ = ∂²L/∂Fᵢ² (Hessian). For MSE: gᵢ = F−y, hᵢ = 1, so w* = mean residual shrunk by λ/(n_leaf + λ).
Explain it
Discussion prompt
Explain XGBoost leaf weight formula to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
With L2 regularisation, the optimal leaf weight w* for a leaf j covering a set of samples J is closed-form:
Concept
LightGBM accelerates gradient boosting for large datasets with two architectural changes.
| technique | what it means | speed gain |
|---|---|---|
| Histogram features | bin continuous features into 256 buckets; find splits in O(bins) not O(n) | ~10× on large n |
| Leaf-wise growth | grow the single leaf with highest loss reduction next (not level-by-level) | fewer nodes, better fit per leaf |
| GOSS sampling | keep large-gradient samples, subsample small-gradient ones | fewer samples per round |
Leaf-wise can overfit on small datasets — num_leaves and min_child_samples are the main regularizers to tune.
Comparison
Comparison matrix
From LightGBM: histogram boosting + leaf-wise growth: refill the speed gain column from what you know. The rest of the table is as it appeared.
| technique | what it means | speed gain |
|---|---|---|
| Histogram features | bin continuous features into 256 buckets; find splits in O(bins) not O(n) | ~10× on large n |
| Leaf-wise growth | grow the single leaf with highest loss reduction next (not level-by-level) | fewer nodes, better fit per leaf |
| GOSS sampling | keep large-gradient samples, subsample small-gradient ones | fewer samples per round |
Estimation
Predict first
Sweep n_estimators with 5-fold CV to find the best number of boosting rounds — this is the standard model-selection approach (Lesson 32 bias-variance in practice).
Commit before you compute: what does Hyperparameter sweep with cross-validation come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Optimal n_estimators ≈ 50: CV-RMSE bottoms at 57.23 then starts to drift up
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Beyond the optimum, additional trees chase training-set noise.
Worked example
Sweep n_estimators with 5-fold CV to find the best number of boosting rounds — this is the standard model-selection approach (Lesson 32 bias-variance in practice).
from sklearn.model_selection import cross_val_score
from sklearn.ensemble import GradientBoostingRegressor
from sklearn.datasets import load_diabetes
import numpy as np
X, y = load_diabetes(return_X_y=True)
for n_est in [10, 50, 100]:
cv = -cross_val_score(
GradientBoostingRegressor(n_estimators=n_est, learning_rate=0.1,
max_depth=3, random_state=42),
X, y, cv=5, scoring='neg_root_mean_squared_error')
print(f'n_estimators={n_est:3d} CV-RMSE={cv.mean():.2f} ± {cv.std():.2f}')| n_estimators | CV-RMSE (mean) | CV-RMSE (std) |
|---|---|---|
| 10 | 60.50 | 1.93 |
| 50 | 57.23 | 2.01 |
| 100 | 58.57 | 2.10 |
Optimal n_estimators ≈ 50: CV-RMSE bottoms at 57.23 then starts to drift up
Why: Beyond the optimum, additional trees chase training-set noise. Cross-validation reveals this where held-out loss increases. Lower learning_rate shifts the optimum to a higher n_estimators.
Pattern
Step through it
Step through Hyperparameter sweep with cross-validation one row at a time. What is driving the change, and what would the row after the last one be?
Ranking
Put in order
These are the steps of The gradient boosting recipe, scrambled. Put them back in order before the next slide shows you.
Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
Edge cases
Discussion prompt
The gradient boosting recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
Elimination
Eliminate the wrong options
A GBM is trained for binary classification with log-loss. At iteration m, sample i has log-odds score F = 1.0 and true label y = 1. What is the pseudo-residual?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: For log-loss the pseudo-residual is y − σ(F). With F=1.0, σ(1.0)=0.7311. So r = 1 − 0.731 = 0.269. The sign is positive because the model is under-confident (probability 0.73 for a positive label).
Check
Derive before clicking.
Check your understanding
A GBM is trained for binary classification with log-loss. At iteration m, sample i has log-odds score F = 1.0 and true label y = 1. What is the pseudo-residual?
Answer: A
Why: For log-loss the pseudo-residual is y − σ(F). With F=1.0, σ(1.0)=0.7311. So r = 1 − 0.731 = 0.269. The sign is positive because the model is under-confident (probability 0.73 for a positive label).
Prediction
Predict first
A leaf covers 4 samples, all with MSE gradient gᵢ = −1.0 and Hessian hᵢ = 1.0. With λ = 1, what is the optimal leaf weight w*?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: w* = 0.8
Why: w* = −ΣgⱼI / (ΣhⱼI + λ) = −(4 × −1.0) / (4 × 1.0 + 1) = 4 / 5 = 0.8. The negative gradient sums to +4; the denominator adds the λ=1 shrinkage term.
Check
Apply the formula.
Check your understanding
A leaf covers 4 samples, all with MSE gradient gᵢ = −1.0 and Hessian hᵢ = 1.0. With λ = 1, what is the optimal leaf weight w*?
Answer: A
Why: w* = −ΣgⱼI / (ΣhⱼI + λ) = −(4 × −1.0) / (4 × 1.0 + 1) = 4 / 5 = 0.8. The negative gradient sums to +4; the denominator adds the λ=1 shrinkage term.
Elimination
Eliminate the wrong options
Which statement best describes LightGBM's leaf-wise growth strategy?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: LightGBM selects the single leaf with the highest gain and splits it next. This gives deeper, more asymmetric trees than level-wise growth (which is what sklearn's GBM and standard XGBoost use). It achieves better fit per leaf count but can overfit on small data — controlled by num_leaves and min_child_samples.
Check
Which claim is accurate?
Check your understanding
Which statement best describes LightGBM's leaf-wise growth strategy?
Answer: A
Why: LightGBM selects the single leaf with the highest gain and splits it next. This gives deeper, more asymmetric trees than level-wise growth (which is what sklearn's GBM and standard XGBoost use). It achieves better fit per leaf count but can overfit on small data — controlled by num_leaves and min_child_samples.
Section
Project
Concept
Implement a 3-round gradient boosting regressor using DecisionTreeRegressor depth-1 stumps, then compare with sklearn's GBM and a Random Forest on load_diabetes.
| # | requirement | tool |
|---|---|---|
| 1 | 3-round from-scratch GBM loop | np.mean, DecisionTreeRegressor |
| 2 | sklearn GBM baseline | GradientBoostingRegressor |
| 3 | n_estimators CV sweep | cross_val_score |
Predict: will from-scratch RMSE after 3 rounds match sklearn's GBM at 100 rounds? Why or why not?
Analogy
Discussion prompt
Explain Project: GBM regressor from scratch by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Implement a 3-round gradient boosting regressor using DecisionTreeRegressor depth-1 stumps, then compare with sklearn's GBM and a Random Forest on load_diabetes.
Missing information
Discussion prompt
Your turn: implement F₀ = mean, then 3 rounds of pseudo-residual + depth-1 stump + update. Print MSE at each round.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
3 stumps with η=0.5 take large steps and leave substantial residual. Sklearn's GBM at 100 rounds with η=0.1 makes 100 fine adjustments — same budget, better outcome.
Worked example
Your turn: implement F₀ = mean, then 3 rounds of pseudo-residual + depth-1 stump + update. Print MSE at each round.
Hint: F = np.full(n, y.mean()), then loop r = y - F; stump.fit(X, r); F += eta * stump.predict(X).
import numpy as np
from sklearn.datasets import load_diabetes
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeRegressor
from sklearn.metrics import mean_squared_error
X, y = load_diabetes(return_X_y=True)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.2, random_state=42)
eta = 0.5
F_tr = np.full(len(ytr), ytr.mean())
F_te = np.full(len(yte), ytr.mean())
for m in range(1, 4):
r = ytr - F_tr
t = DecisionTreeRegressor(max_depth=1, random_state=0)
t.fit(Xtr, r)
F_tr += eta * t.predict(Xtr)
F_te += eta * t.predict(Xte)
print(f'm={m}: train RMSE={np.sqrt(mean_squared_error(ytr,F_tr)):.2f} test RMSE={np.sqrt(mean_squared_error(yte,F_te)):.2f}')| round m | train RMSE | test RMSE |
|---|---|---|
| 0 (init) | 77.82 | 77.82 |
| 1 | 73.14 | 73.91 |
| 2 | 69.21 | 70.53 |
| 3 | 65.88 | 67.44 |
After 3 rounds with η=0.5: test RMSE ≈ 67.4 — higher than sklearn's 100-round GBM (53.73)
Why: 3 stumps with η=0.5 take large steps and leave substantial residual. Sklearn's GBM at 100 rounds with η=0.1 makes 100 fine adjustments — same budget, better outcome.
Pattern
Step through it
Step through Milestone 1 — from-scratch GBM loop one row at a time. What is driving the change, and what would the row after the last one be?
Estimation
Predict first
Your turn: run sklearn GradientBoostingRegressor and sweep n_estimators with 5-fold CV. Predict which n_estimators minimises CV-RMSE.
Commit before you compute: what does Milestone 2 — sklearn GBM and hyperparameter sweep come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: CV minimum at n_estimators=50: RMSE 57.23; n_estimators=100 is already slightly overfitting
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The U-shaped CV curve is the bias-variance tradeoff (Lesson 32) in action: too few trees = underfitting; too many = overfitting on training folds.
Worked example
Your turn: run sklearn GradientBoostingRegressor and sweep n_estimators with 5-fold CV. Predict which n_estimators minimises CV-RMSE.
Hint: cross_val_score(..., scoring='neg_root_mean_squared_error'); negate the scores to get RMSE.
from sklearn.model_selection import cross_val_score
from sklearn.ensemble import GradientBoostingRegressor
from sklearn.datasets import load_diabetes
import numpy as np
X, y = load_diabetes(return_X_y=True)
for n_est in [10, 50, 100]:
cv = -cross_val_score(
GradientBoostingRegressor(n_estimators=n_est, learning_rate=0.1,
max_depth=3, random_state=42),
X, y, cv=5, scoring='neg_root_mean_squared_error')
print(f'n_estimators={n_est:3d} CV-RMSE={cv.mean():.2f} ± {cv.std():.2f}')| n_estimators | CV-RMSE | std |
|---|---|---|
| 10 | 60.50 | ±1.93 |
| 50 | 57.23 | ±2.01 |
| 100 | 58.57 | ±2.10 |
CV minimum at n_estimators=50: RMSE 57.23; n_estimators=100 is already slightly overfitting
Why: The U-shaped CV curve is the bias-variance tradeoff (Lesson 32) in action: too few trees = underfitting; too many = overfitting on training folds.
Trade off
Comparison matrix
From Milestone 2 — sklearn GBM and hyperparameter sweep: every row here is a choice with a cost. Fill the CV-RMSE column, then say which row you would actually pick and what you give up for it.
| n_estimators | CV-RMSE | std |
|---|---|---|
| 10 | 60.50 | ±1.93 |
| 50 | 57.23 | ±2.01 |
| 100 | 58.57 | ±2.10 |
Concept
import numpy as np
from sklearn.datasets import load_diabetes
from sklearn.model_selection import train_test_split, cross_val_score
from sklearn.tree import DecisionTreeRegressor
from sklearn.ensemble import GradientBoostingRegressor, RandomForestRegressor
from sklearn.metrics import mean_squared_error
X, y = load_diabetes(return_X_y=True)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.2, random_state=42)
# from-scratch GBM (3 rounds, eta=0.5)
F = np.full(len(yte), ytr.mean())
Ftr = np.full(len(ytr), ytr.mean())
for _ in range(3):
r = ytr - Ftr
t = DecisionTreeRegressor(max_depth=1, random_state=0)
t.fit(Xtr, r); Ftr += 0.5*t.predict(Xtr); F += 0.5*t.predict(Xte)
print('Scratch 3-round RMSE:', round(np.sqrt(mean_squared_error(yte, F)), 2))
# sklearn GBM (100 rounds)
gbm = GradientBoostingRegressor(n_estimators=100, learning_rate=0.1, max_depth=3, random_state=42)
gbm.fit(Xtr, ytr)
print('GBM RMSE:', round(np.sqrt(mean_squared_error(yte, gbm.predict(Xte))), 2))
# Random Forest
rf = RandomForestRegressor(n_estimators=100, random_state=42)
rf.fit(Xtr, ytr)
print('RF RMSE:', round(np.sqrt(mean_squared_error(yte, rf.predict(Xte))), 2))| model | test RMSE |
|---|---|
| Scratch GBM (3 rounds, η=0.5) | ~67.4 |
| sklearn GBM (100 rounds, η=0.1) | 53.73 |
| Random Forest (100 trees) | 54.33 |
Boosting beats bagging at the same tree count because it targets residuals directly — but only when properly regularised with small η and cross-validated depth/n_estimators.
Comparison
Comparison matrix
From The full program: refill the test RMSE column from what you know. The rest of the table is as it appeared.
| model | test RMSE |
|---|---|
| Scratch GBM (3 rounds, η=0.5) | ~67.4 |
| sklearn GBM (100 rounds, η=0.1) | 53.73 |
| Random Forest (100 trees) | 54.33 |
Concept
Out loud, slides closed: (1) state the GBM update rule and name every term; (2) derive that MSE pseudo-residuals = y − F; (3) explain why you cannot use the same formula for classification.
Stretch (homework): tune learning_rate and max_depth together with GridSearchCV; implement early stopping by monitoring a validation RMSE curve; read Chen & Guestrin 2016 Section 2 on the regularised objective.
Counterexample
Discussion prompt
Out loud, slides closed: (1) state the GBM update rule and name every term; (2) derive that MSE pseudo-residuals = y − F; (3) explain why you cannot use the same formula for classification.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — The boosting framework · From scratch: 3-round GBM · XGBoost and LightGBM · Your turn: build it. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
| idea | the one thing to remember |
|---|---|
| GBM update | F_m = F_{m-1} + η·h_m; h_m fits pseudo-residuals |
| MSE residual | −∂L/∂F = y − F (the plain residual, by coincidence) |
| Classification | pseudo = y − σ(F); sigmoid is essential |
| Regularise | small η + CV n_estimators; depth ≤ 5 |
| XGBoost | L2 leaf reg: w* = −Σg/(Σh+λ) |
| LightGBM | histogram bins + leaf-wise = fast on large data |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.