Lesson 50: Gradient Boosting

USAAIO Lesson 50, from Phase 2. It covers the gradient-boosting framework F_m = F_{m-1} + eta*h_m and pseudo-residuals as the negative gradients of the loss - y − F for MSE, and y − sigma(F) for classification - then builds a three-round GBM regressor from scratch using an sklearn DecisionTree. It goes on to the XGBoost improvements, namely L2 leaf regularization, column subsampling, and approximate splits, and to LightGBM's histogram-based leaf-wise growth, ending with a cross-validated hyperparameter sweep. You build GradientBoostingRegressorFromScratch on make_regression data and compare the GBM with a random forest. The lesson runs to 29 slides.

Subject: Machine Learning · 60 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Gradient Boosting

Title

USAAIO · Lesson 50 · Phase 2

From residuals to an ensemble: the boosting framework, pseudo-residuals, and modern implementations — XGBoost, LightGBM.

2. By the end of this lesson you can

Objectives

  1. State the boosting update rule F_m = F_{m-1} + η·h_m and explain each term
  2. Derive that MSE pseudo-residuals equal the plain residuals y − F_{m-1}
  3. Explain why classification pseudo-residuals are y − σ(F) instead
  4. Implement a 3-round GBM regressor from scratch with sklearn DecisionTreeRegressor
  5. Describe XGBoost's L2 leaf regularization, column subsampling, and LightGBM's histogram / leaf-wise growth

3. What survived from Random Forests?

Warm-up

Discussion prompt

Before we open Lesson 50: Gradient Boosting: without looking back, what was the main idea of Random Forests, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

bootstrap aggregation, random feature subsets (sqrt(d)), out-of-bag error as free validation, MDI vs permutation vs SHAP feature importance, and Random Forest vs single tree vs boosting tradeoffs. Implement RandomForestClassifier from scratch and compare OOB error to k-fold CV.

4. The boosting framework

Section

Part 1 of 3

5. Boosting: additive models, one step at a time

Concept

Boosting builds an ensemble by fitting each new learner to the mistakes of the current model — it is an additive procedure, not an average.

\[ F_m(x) = F_{m-1}(x) + \eta \, h_m(x) \]

6. Break it if you can: Boosting: additive models, one step at a time

Counterexample

Discussion prompt

Boosting builds an ensemble by fitting each new learner to the mistakes of the current model — it is an additive procedure, not an average.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

7. Pseudo-residuals: the negative gradient of the loss

Concept

Gradient boosting frames h_m as fitting the negative gradient of the loss L(y, F) with respect to F — this is the steepest-descent direction in function space.

\[ r_{im} = -\frac{\partial L(y_i,\, F_{m-1}(x_i))}{\partial F_{m-1}(x_i)} \]

losspseudo-residual r_imintuition
MSE ½(y−F)²y − Fplain residual
Log-loss (binary)y − σ(F)prob error

8. Fill in: pseudo-residual r_im for Pseudo-residuals: the negative gradient of…

Comparison

Comparison matrix

From Pseudo-residuals: the negative gradient of the loss: refill the pseudo-residual r_im column from what you know. The rest of the table is as it appeared.

losspseudo-residual r_imintuition
MSE ½(y−F)²y − Fplain residual
Log-loss (binary)y − σ(F)prob error

9. Deriving the MSE pseudo-residual

Concept

For mean-squared error the derivation is one line — and explains why the homework says 'just the residuals!'

\[ L(y,F) = \frac{1}{2}(y - F)^2 \;\Longrightarrow\; -\frac{\partial L}{\partial F} = -\bigl(-(y-F)\bigr) = y - F \]

Fitting a tree to y − F (the ordinary residual) is gradient descent in function space — no separate math is needed.

10. By analogy: Deriving the MSE pseudo-residual

Analogy

Discussion prompt

Explain Deriving the MSE pseudo-residual by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

For mean-squared error the derivation is one line — and explains why the homework says 'just the residuals!'

11. Classification pseudo-residuals: y − σ(F)

Concept

For binary classification GBM outputs a log-odds score F. The loss is log-loss, and the pseudo-residual involves the sigmoid.

\[ -\frac{\partial}{\partial F}\bigl[-y\ln\sigma(F)-(1-y)\ln(1-\sigma(F))\bigr] = y - \sigma(F) \]

F (log-odds)σ(F)ypseudo = y−σ(F)
0.00.500010.5000
1.00.731110.2689
-1.00.26890-0.2689

12. What each one costs: Classification pseudo-residuals: y − σ(F)

Trade off

Comparison matrix

From Classification pseudo-residuals: y − σ(F): every row here is a choice with a cost. Fill the pseudo = y−σ(F) column, then say which row you would actually pick and what you give up for it.

F (log-odds)σ(F)ypseudo = y−σ(F)
0.00.500010.5000
1.00.731110.2689
-1.00.26890-0.2689

13. Something is wrong here: confusing pseudo-residuals between regression and…

Anomaly

Predict first

A student writes this, and it looks reasonable:

The GBM loop always fits the next tree to the residuals y − F, regardless of whether the task is regression or classification.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: This is only right for MSE regression.

The pseudo-residual formula depends on the loss function; derive it as −∂L/∂F for each task.

Why: This is only right for MSE regression. The formula y − F treats F as a predicted value, but for classification F is a log-odds score — the pseudo-residual must go through the sigmoid: r = y − σ(F).

14. Trap: confusing pseudo-residuals between regression and classification

Trap

The trap

The GBM loop always fits the next tree to the residuals y − F, regardless of whether the task is regression or classification.

Use r = y − F in both cases

Why: This is only right for MSE regression. The formula y − F treats F as a predicted value, but for classification F is a log-odds score — the pseudo-residual must go through the sigmoid: r = y − σ(F).

The fix

The pseudo-residual formula depends on the loss function; derive it as −∂L/∂F for each task.

Regression (MSE): r = y − F. Classification (log-loss): r = y − σ(F)

Why: Both follow from −∂L/∂F. MSE gives the plain residual; log-loss passes F through the sigmoid first. These are the same for regression only by coincidence.

15. Break it on purpose: confusing pseudo-residuals between…

Break the constraint

Discussion prompt

The rule this trap just fixed:

Both follow from −∂L/∂F. MSE gives the plain residual; log-loss passes F through the sigmoid first. These are the same for regression only by coincidence.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

This is only right for MSE regression. The formula y − F treats F as a predicted value, but for classification F is a log-odds score — the pseudo-residual must go through the sigmoid: r = y − σ(F).

16. From scratch: 3-round GBM

Section

Part 2 of 3

17. The from-scratch GBM algorithm

Concept

  1. Initialise F₀(x) = mean(y)
  2. For m = 1, 2, …, M:
  3. (a) compute pseudo-residuals r_i = y_i − F_{m-1}(x_i)
  4. (b) fit a depth-1 tree h_m to (X, r)
  5. (c) update F_m = F_{m-1} + η · h_m
  6. Predict with F_M

Using depth-1 stumps and a small η (< 1) keeps each step modest — many weak learners combine into one strong one.

18. Teach it back: The from-scratch GBM algorithm

Explain it

Discussion prompt

Explain The from-scratch GBM algorithm to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Using depth-1 stumps and a small η (< 1) keeps each step modest — many weak learners combine into one strong one.

19. Guess the shape of the answer: Worked example — 3-round GBM (3 samples)

Estimation

Predict first

Trace gradient boosting on a tiny dataset with y = 2x and η = 0.5 so you can follow every update by hand.

Commit before you compute: what does Worked example — 3-round GBM (3 samples) come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: MSE falls 2.67 → 1.17 → 0.32 → 0.11 across three rounds

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Each round reduces the unexplained variance.

20. Worked example — 3-round GBM (3 samples)

Worked example

Trace gradient boosting on a tiny dataset with y = 2x and η = 0.5 so you can follow every update by hand.

import numpy as np
from sklearn.tree import DecisionTreeRegressor

X = np.array([[1.0],[2.0],[3.0]])
y = np.array([2.0, 4.0, 6.0])   # y = 2x, clean
eta = 0.5
F = np.full(3, y.mean())        # F0 = 4.0 for all
for m in range(1, 4):
    r = y - F                   # pseudo-residuals (MSE)
    stump = DecisionTreeRegressor(max_depth=1, random_state=0)
    stump.fit(X, r)
    F = F + eta * stump.predict(X)
    print(f'm={m}: r={r.round(3)}, F={F.round(3)}, MSE={np.mean((y-F)**2):.4f}')
mr (pseudo-resids)F after updateMSE
0 (init)[−2, 0, +2][4, 4, 4]2.6667
1[−2, 0, +2][3, 4.5, 4.5]1.1667
2[−1, −0.5, +1.5][2.625, 4.125, 5.25]0.3229
3[−0.625, −0.125, +0.75][2.438, 3.938, 5.625]0.1120

Round 1: stump splits on x≤1.5; left leaf predicts −2, right leaf predicts +1

Why: The tree fits the residuals, not y directly. η=0.5 moves only half-way to the residual to avoid over-correcting in a single step.

MSE falls 2.67 → 1.17 → 0.32 → 0.11 across three rounds

Why: Each round reduces the unexplained variance. With more rounds (and smaller η) F_M converges toward y.

21. Inspect it line by line: Worked example — 3-round GBM (3 samples)

Error analysis

Annotate

Walk the callouts on Worked example — 3-round GBM (3 samples). Each one is a place this is easy to get subtly wrong.

  • The tree fits the residuals, not y directly. η=0.5 moves only half-way to the residual to avoid over-correcting in a single step.
  • Each round reduces the unexplained variance. With more rounds (and smaller η) F_M converges toward y.

22. What has to be given first: From-scratch GBM on real data

Missing information

Discussion prompt

Apply the same loop to the load_diabetes dataset and compare training RMSE vs sklearn's GradientBoostingRegressor.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

Boosting directly targets residuals, so it tends to slightly outperform bagging-based forests on tabular data with the same depth — but needs careful regularization (n_estimators, learning_rate) to not overfit.

23. From-scratch GBM on real data

Worked example

Apply the same loop to the load_diabetes dataset and compare training RMSE vs sklearn's GradientBoostingRegressor.

from sklearn.datasets import load_diabetes
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeRegressor
from sklearn.ensemble import GradientBoostingRegressor
from sklearn.metrics import mean_squared_error
import numpy as np

X, y = load_diabetes(return_X_y=True)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.2, random_state=42)

# sklearn GBM
gbm = GradientBoostingRegressor(n_estimators=100, learning_rate=0.1, max_depth=3, random_state=42)
gbm.fit(Xtr, ytr)
print('GBM RMSE:', round(np.sqrt(mean_squared_error(yte, gbm.predict(Xte))), 2))
modeltest RMSE
GBM (sklearn)53.73
Random Forest54.33

GBM RMSE 53.73 vs Random Forest 54.33 on the diabetes holdout

Why: Boosting directly targets residuals, so it tends to slightly outperform bagging-based forests on tabular data with the same depth — but needs careful regularization (n_estimators, learning_rate) to not overfit.

24. Work backwards from the answer: From-scratch GBM on real data

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

GBM RMSE 53.73 vs Random Forest 54.33 on the diabetes holdout

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Apply the same loop to the load_diabetes dataset and compare training RMSE vs sklearn's GradientBoostingRegressor.

25. Something is wrong here: too many estimators without shrinkage

Anomaly

Predict first

A student writes this, and it looks reasonable:

Use a large learning_rate (0.5) and many estimators — more trees always means a better model.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: The train RMSE collapses, which looks like perfect fit.

Use a small learning_rate and pick n_estimators by cross-validation.

Why: The train RMSE collapses, which looks like perfect fit. But test RMSE is 66.34 — worse than 10-estimator GBM (62.39). The model memorised training noise.

26. Trap: too many estimators without shrinkage

Trap

The trap

Use a large learning_rate (0.5) and many estimators — more trees always means a better model.

n_estimators=500, learning_rate=0.5, depth=5: train RMSE → 0.01

Why: The train RMSE collapses, which looks like perfect fit. But test RMSE is 66.34 — worse than 10-estimator GBM (62.39). The model memorised training noise.

The fix

Use a small learning_rate and pick n_estimators by cross-validation.

n_estimators=10, learning_rate=0.1, depth=5: train RMSE=40.27, test RMSE=62.39

Why: Low learning rate with cross-validated n_estimators trades per-step progress for generalisation. The learning_rate / n_estimators tradeoff: halving η and doubling M usually helps.

27. Which of these survive contact with Lesson 50: Gradient Boosting?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Boosting builds an ensemble by fitting each new learner to the mistakes of the current model — it is an additive procedure, not an average.; Gradient boosting frames h_m as fitting the negative gradient of the loss L(y, F) with respect to F — this is the steepest-descent direction in function space.; For mean-squared error the derivation is one line — and explains why the homework says 'just the residuals!'
Breaks
The GBM loop always fits the next tree to the residuals y − F, regardless of whether the task is regression or classification.; Use a large learning_rate (0.5) and many estimators — more trees always means a better model.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 50: Gradient Boosting puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

28. XGBoost and LightGBM

Section

Part 3 of 3

29. XGBoost: three key improvements

Concept

XGBoost (Chen & Guestrin, 2016) extends vanilla GBM with three algorithmic improvements that dominate ML competitions.

improvementwhat it doesparameter
L2 leaf regpenalises large leaf weights: Ω(h) = λΣwⱼ²reg_lambda
Column subsamplingrandom feature subset per tree (like RF)colsample_bytree
Approx split findinghistogram bins the feature values; O(n) not O(n log n)built-in

The regularised objective is Obj = ΣL(yᵢ, Fᵢ) + Σ[γTⱼ + ½λwⱼ²] where T is the number of leaves.

30. XGBoost leaf weight formula

Concept

With L2 regularisation, the optimal leaf weight w* for a leaf j covering a set of samples J is closed-form:

\[ w_j^* = -\frac{\sum_{i \in J} g_i}{\sum_{i \in J} h_i + \lambda} \]

where gᵢ = ∂L/∂Fᵢ (gradient) and hᵢ = ∂²L/∂Fᵢ² (Hessian). For MSE: gᵢ = F−y, hᵢ = 1, so w* = mean residual shrunk by λ/(n_leaf + λ).

31. Teach it back: XGBoost leaf weight formula

Explain it

Discussion prompt

Explain XGBoost leaf weight formula to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

With L2 regularisation, the optimal leaf weight w* for a leaf j covering a set of samples J is closed-form:

32. LightGBM: histogram boosting + leaf-wise growth

Concept

LightGBM accelerates gradient boosting for large datasets with two architectural changes.

techniquewhat it meansspeed gain
Histogram featuresbin continuous features into 256 buckets; find splits in O(bins) not O(n)~10× on large n
Leaf-wise growthgrow the single leaf with highest loss reduction next (not level-by-level)fewer nodes, better fit per leaf
GOSS samplingkeep large-gradient samples, subsample small-gradient onesfewer samples per round

Leaf-wise can overfit on small datasets — num_leaves and min_child_samples are the main regularizers to tune.

33. Fill in: speed gain for LightGBM: histogram boosting + leaf-wise…

Comparison

Comparison matrix

From LightGBM: histogram boosting + leaf-wise growth: refill the speed gain column from what you know. The rest of the table is as it appeared.

techniquewhat it meansspeed gain
Histogram featuresbin continuous features into 256 buckets; find splits in O(bins) not O(n)~10× on large n
Leaf-wise growthgrow the single leaf with highest loss reduction next (not level-by-level)fewer nodes, better fit per leaf
GOSS samplingkeep large-gradient samples, subsample small-gradient onesfewer samples per round

34. Guess the shape of the answer: Hyperparameter sweep with cross-validation

Estimation

Predict first

Sweep n_estimators with 5-fold CV to find the best number of boosting rounds — this is the standard model-selection approach (Lesson 32 bias-variance in practice).

Commit before you compute: what does Hyperparameter sweep with cross-validation come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Optimal n_estimators ≈ 50: CV-RMSE bottoms at 57.23 then starts to drift up

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Beyond the optimum, additional trees chase training-set noise.

35. Hyperparameter sweep with cross-validation

Worked example

Sweep n_estimators with 5-fold CV to find the best number of boosting rounds — this is the standard model-selection approach (Lesson 32 bias-variance in practice).

from sklearn.model_selection import cross_val_score
from sklearn.ensemble import GradientBoostingRegressor
from sklearn.datasets import load_diabetes
import numpy as np

X, y = load_diabetes(return_X_y=True)
for n_est in [10, 50, 100]:
    cv = -cross_val_score(
        GradientBoostingRegressor(n_estimators=n_est, learning_rate=0.1,
                                   max_depth=3, random_state=42),
        X, y, cv=5, scoring='neg_root_mean_squared_error')
    print(f'n_estimators={n_est:3d}  CV-RMSE={cv.mean():.2f} ± {cv.std():.2f}')
n_estimatorsCV-RMSE (mean)CV-RMSE (std)
1060.501.93
5057.232.01
10058.572.10

Optimal n_estimators ≈ 50: CV-RMSE bottoms at 57.23 then starts to drift up

Why: Beyond the optimum, additional trees chase training-set noise. Cross-validation reveals this where held-out loss increases. Lower learning_rate shifts the optimum to a higher n_estimators.

36. Watch it run: Hyperparameter sweep with cross-validation

Pattern

Step through it

Step through Hyperparameter sweep with cross-validation one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: n_estimators is 10
  2. Step 2: n_estimators is 50
  3. Step 3: n_estimators is 100

37. Rebuild the recipe: The gradient boosting recipe

Ranking

Put in order

These are the steps of The gradient boosting recipe, scrambled. Put them back in order before the next slide shows you.

  1. Initialise F₀ = mean(y) (regression) or log-odds of class frequency (classification)
  2. For m = 1…M: compute pseudo-residuals r = −∂L/∂F; fit depth-limited tree h_m to r; update F_m = F_{m-1} + η·h_m
  3. Loss-to-residual table: MSE → y−F; log-loss → y−σ(F)
  4. Regularise: small η, max_depth 3-5, n_estimators via cross-validation
  5. XGBoost upgrades: L2 leaf reg (λ), column subsampling, second-order gradients
  6. LightGBM upgrades: histogram features + leaf-wise growth for large data

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

38. The gradient boosting recipe

Pattern

  1. Initialise F₀ = mean(y) (regression) or log-odds of class frequency (classification)
  2. For m = 1…M: compute pseudo-residuals r = −∂L/∂F; fit depth-limited tree h_m to r; update F_m = F_{m-1} + η·h_m
  3. Loss-to-residual table: MSE → y−F; log-loss → y−σ(F)
  4. Regularise: small η, max_depth 3-5, n_estimators via cross-validation
  5. XGBoost upgrades: L2 leaf reg (λ), column subsampling, second-order gradients
  6. LightGBM upgrades: histogram features + leaf-wise growth for large data

39. Where does it stop working: The gradient boosting recipe

Edge cases

Discussion prompt

The gradient boosting recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Initialise F₀ = mean(y) (regression) or log-odds of class frequency (classification)
  2. For m = 1…M: compute pseudo-residuals r = −∂L/∂F; fit depth-limited tree h_m to r; update F_m = F_{m-1} + η·h_m
  3. Loss-to-residual table: MSE → y−F; log-loss → y−σ(F)
  4. Regularise: small η, max_depth 3-5, n_estimators via cross-validation
  5. XGBoost upgrades: L2 leaf reg (λ), column subsampling, second-order gradients
  6. LightGBM upgrades: histogram features + leaf-wise growth for large data

40. Rule out three: Check yourself — pseudo-residuals

Elimination

Eliminate the wrong options

A GBM is trained for binary classification with log-loss. At iteration m, sample i has log-odds score F = 1.0 and true label y = 1. What is the pseudo-residual?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. y − σ(F) = 1 − 0.731 = 0.269
  • B. y − F = 1 − 1.0 = 0.0
  • C. σ(F) − y = 0.731 − 1 = −0.269
  • D. y − F² = 1 − 1.0 = 0.0

Survives elimination: A

Why: For log-loss the pseudo-residual is y − σ(F). With F=1.0, σ(1.0)=0.7311. So r = 1 − 0.731 = 0.269. The sign is positive because the model is under-confident (probability 0.73 for a positive label).

41. Check yourself — pseudo-residuals

Check

Derive before clicking.

Check your understanding

A GBM is trained for binary classification with log-loss. At iteration m, sample i has log-odds score F = 1.0 and true label y = 1. What is the pseudo-residual?

  • A. y − σ(F) = 1 − 0.731 = 0.269 (correct)
  • B. y − F = 1 − 1.0 = 0.0
  • C. σ(F) − y = 0.731 − 1 = −0.269
  • D. y − F² = 1 − 1.0 = 0.0

Answer: A

Why: For log-loss the pseudo-residual is y − σ(F). With F=1.0, σ(1.0)=0.7311. So r = 1 − 0.731 = 0.269. The sign is positive because the model is under-confident (probability 0.73 for a positive label).

Why B tempts people
y − F would only be correct for MSE regression where F is a predicted value, not a log-odds score.
Why C tempts people
This flips the sign. The pseudo-residual is −∂L/∂F = y − σ(F), not σ(F) − y.
Why D tempts people
Squaring F has no basis in the log-loss gradient formula.

42. Answer it before you see the options: Check yourself — XGBoost leaf weight

Prediction

Predict first

A leaf covers 4 samples, all with MSE gradient gᵢ = −1.0 and Hessian hᵢ = 1.0. With λ = 1, what is the optimal leaf weight w*?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: w* = 0.8

Why: w* = −ΣgⱼI / (ΣhⱼI + λ) = −(4 × −1.0) / (4 × 1.0 + 1) = 4 / 5 = 0.8. The negative gradient sums to +4; the denominator adds the λ=1 shrinkage term.

43. Check yourself — XGBoost leaf weight

Check

Apply the formula.

Check your understanding

A leaf covers 4 samples, all with MSE gradient gᵢ = −1.0 and Hessian hᵢ = 1.0. With λ = 1, what is the optimal leaf weight w*?

  • A. w* = 0.8 (correct)
  • B. w* = −0.8
  • C. w* = −1.0
  • D. w* = 0.5

Answer: A

Why: w* = −ΣgⱼI / (ΣhⱼI + λ) = −(4 × −1.0) / (4 × 1.0 + 1) = 4 / 5 = 0.8. The negative gradient sums to +4; the denominator adds the λ=1 shrinkage term.

Why B tempts people
Applying the formula without the leading minus: −(Σg)/(Σh+λ) = −(−4)/5 = +0.8, not −0.8. The minus sign in the numerator cancels the negative gradient.
Why C tempts people
w* = −1 would result if λ=0 and you computed −Σg/Σh = 4/4 = 1 — omitting the regularisation term λ in the denominator.
Why D tempts people
w* = 0.5 would result from dividing Σg (not −Σg) by Σh+λ: −4/−(4+1) is not the formula.

44. Rule out three: Check yourself — LightGBM growth

Elimination

Eliminate the wrong options

Which statement best describes LightGBM's leaf-wise growth strategy?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Grow the leaf with the largest loss reduction at each step, regardless of tree depth
  • B. Grow all leaves at the same depth level before expanding any further
  • C. Grow only the root split, then stop to limit complexity
  • D. Randomly choose a leaf to split, equivalent to stochastic bagging

Survives elimination: A

Why: LightGBM selects the single leaf with the highest gain and splits it next. This gives deeper, more asymmetric trees than level-wise growth (which is what sklearn's GBM and standard XGBoost use). It achieves better fit per leaf count but can overfit on small data — controlled by num_leaves and min_child_samples.

45. Check yourself — LightGBM growth

Check

Which claim is accurate?

Check your understanding

Which statement best describes LightGBM's leaf-wise growth strategy?

  • A. Grow the leaf with the largest loss reduction at each step, regardless of tree depth (correct)
  • B. Grow all leaves at the same depth level before expanding any further
  • C. Grow only the root split, then stop to limit complexity
  • D. Randomly choose a leaf to split, equivalent to stochastic bagging

Answer: A

Why: LightGBM selects the single leaf with the highest gain and splits it next. This gives deeper, more asymmetric trees than level-wise growth (which is what sklearn's GBM and standard XGBoost use). It achieves better fit per leaf count but can overfit on small data — controlled by num_leaves and min_child_samples.

Why B tempts people
Level-wise growth describes standard CART and sklearn GBM, not LightGBM.
Why C tempts people
Stopping at the root would produce a one-level stump. LightGBM grows beyond the root — it simply chooses which leaf to expand greedily.
Why D tempts people
Random leaf selection is not LightGBM's strategy; GOSS samples based on gradient magnitude, not randomly.

46. Your turn: build it

Section

Project

47. Project: GBM regressor from scratch

Concept

Implement a 3-round gradient boosting regressor using DecisionTreeRegressor depth-1 stumps, then compare with sklearn's GBM and a Random Forest on load_diabetes.

#requirementtool
13-round from-scratch GBM loopnp.mean, DecisionTreeRegressor
2sklearn GBM baselineGradientBoostingRegressor
3n_estimators CV sweepcross_val_score

Predict: will from-scratch RMSE after 3 rounds match sklearn's GBM at 100 rounds? Why or why not?

48. By analogy: Project: GBM regressor from scratch

Analogy

Discussion prompt

Explain Project: GBM regressor from scratch by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Implement a 3-round gradient boosting regressor using DecisionTreeRegressor depth-1 stumps, then compare with sklearn's GBM and a Random Forest on load_diabetes.

49. What has to be given first: Milestone 1 — from-scratch GBM loop

Missing information

Discussion prompt

Your turn: implement F₀ = mean, then 3 rounds of pseudo-residual + depth-1 stump + update. Print MSE at each round.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

3 stumps with η=0.5 take large steps and leave substantial residual. Sklearn's GBM at 100 rounds with η=0.1 makes 100 fine adjustments — same budget, better outcome.

50. Milestone 1 — from-scratch GBM loop

Worked example

Your turn: implement F₀ = mean, then 3 rounds of pseudo-residual + depth-1 stump + update. Print MSE at each round.

Hint: F = np.full(n, y.mean()), then loop r = y - F; stump.fit(X, r); F += eta * stump.predict(X).

import numpy as np
from sklearn.datasets import load_diabetes
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeRegressor
from sklearn.metrics import mean_squared_error

X, y = load_diabetes(return_X_y=True)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.2, random_state=42)
eta = 0.5
F_tr = np.full(len(ytr), ytr.mean())
F_te = np.full(len(yte), ytr.mean())
for m in range(1, 4):
    r = ytr - F_tr
    t = DecisionTreeRegressor(max_depth=1, random_state=0)
    t.fit(Xtr, r)
    F_tr += eta * t.predict(Xtr)
    F_te += eta * t.predict(Xte)
    print(f'm={m}: train RMSE={np.sqrt(mean_squared_error(ytr,F_tr)):.2f}  test RMSE={np.sqrt(mean_squared_error(yte,F_te)):.2f}')
round mtrain RMSEtest RMSE
0 (init)77.8277.82
173.1473.91
269.2170.53
365.8867.44

After 3 rounds with η=0.5: test RMSE ≈ 67.4 — higher than sklearn's 100-round GBM (53.73)

Why: 3 stumps with η=0.5 take large steps and leave substantial residual. Sklearn's GBM at 100 rounds with η=0.1 makes 100 fine adjustments — same budget, better outcome.

51. Watch it run: Milestone 1 — from-scratch GBM loop

Pattern

Step through it

Step through Milestone 1 — from-scratch GBM loop one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: round m is 0 (init)
  2. Step 2: round m is 1
  3. Step 3: round m is 2
  4. Step 4: round m is 3

52. Guess the shape of the answer: Milestone 2 — sklearn GBM and hyperparameter…

Estimation

Predict first

Your turn: run sklearn GradientBoostingRegressor and sweep n_estimators with 5-fold CV. Predict which n_estimators minimises CV-RMSE.

Commit before you compute: what does Milestone 2 — sklearn GBM and hyperparameter sweep come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: CV minimum at n_estimators=50: RMSE 57.23; n_estimators=100 is already slightly overfitting

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The U-shaped CV curve is the bias-variance tradeoff (Lesson 32) in action: too few trees = underfitting; too many = overfitting on training folds.

53. Milestone 2 — sklearn GBM and hyperparameter sweep

Worked example

Your turn: run sklearn GradientBoostingRegressor and sweep n_estimators with 5-fold CV. Predict which n_estimators minimises CV-RMSE.

Hint: cross_val_score(..., scoring='neg_root_mean_squared_error'); negate the scores to get RMSE.

from sklearn.model_selection import cross_val_score
from sklearn.ensemble import GradientBoostingRegressor
from sklearn.datasets import load_diabetes
import numpy as np

X, y = load_diabetes(return_X_y=True)
for n_est in [10, 50, 100]:
    cv = -cross_val_score(
        GradientBoostingRegressor(n_estimators=n_est, learning_rate=0.1,
                                   max_depth=3, random_state=42),
        X, y, cv=5, scoring='neg_root_mean_squared_error')
    print(f'n_estimators={n_est:3d}  CV-RMSE={cv.mean():.2f} ± {cv.std():.2f}')
n_estimatorsCV-RMSEstd
1060.50±1.93
5057.23±2.01
10058.57±2.10

CV minimum at n_estimators=50: RMSE 57.23; n_estimators=100 is already slightly overfitting

Why: The U-shaped CV curve is the bias-variance tradeoff (Lesson 32) in action: too few trees = underfitting; too many = overfitting on training folds.

54. What each one costs: Milestone 2 — sklearn GBM and hyperparameter sweep

Trade off

Comparison matrix

From Milestone 2 — sklearn GBM and hyperparameter sweep: every row here is a choice with a cost. Fill the CV-RMSE column, then say which row you would actually pick and what you give up for it.

n_estimatorsCV-RMSEstd
1060.50±1.93
5057.23±2.01
10058.57±2.10

55. The full program

Concept

import numpy as np
from sklearn.datasets import load_diabetes
from sklearn.model_selection import train_test_split, cross_val_score
from sklearn.tree import DecisionTreeRegressor
from sklearn.ensemble import GradientBoostingRegressor, RandomForestRegressor
from sklearn.metrics import mean_squared_error

X, y = load_diabetes(return_X_y=True)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.2, random_state=42)

# from-scratch GBM (3 rounds, eta=0.5)
F = np.full(len(yte), ytr.mean())
Ftr = np.full(len(ytr), ytr.mean())
for _ in range(3):
    r = ytr - Ftr
    t = DecisionTreeRegressor(max_depth=1, random_state=0)
    t.fit(Xtr, r); Ftr += 0.5*t.predict(Xtr); F += 0.5*t.predict(Xte)
print('Scratch 3-round RMSE:', round(np.sqrt(mean_squared_error(yte, F)), 2))

# sklearn GBM (100 rounds)
gbm = GradientBoostingRegressor(n_estimators=100, learning_rate=0.1, max_depth=3, random_state=42)
gbm.fit(Xtr, ytr)
print('GBM   RMSE:', round(np.sqrt(mean_squared_error(yte, gbm.predict(Xte))), 2))

# Random Forest
rf = RandomForestRegressor(n_estimators=100, random_state=42)
rf.fit(Xtr, ytr)
print('RF    RMSE:', round(np.sqrt(mean_squared_error(yte, rf.predict(Xte))), 2))
modeltest RMSE
Scratch GBM (3 rounds, η=0.5)~67.4
sklearn GBM (100 rounds, η=0.1)53.73
Random Forest (100 trees)54.33

Boosting beats bagging at the same tree count because it targets residuals directly — but only when properly regularised with small η and cross-validated depth/n_estimators.

56. Fill in: test RMSE for The full program

Comparison

Comparison matrix

From The full program: refill the test RMSE column from what you know. The rest of the table is as it appeared.

modeltest RMSE
Scratch GBM (3 rounds, η=0.5)~67.4
sklearn GBM (100 rounds, η=0.1)53.73
Random Forest (100 trees)54.33

57. Show it off

Concept

Out loud, slides closed: (1) state the GBM update rule and name every term; (2) derive that MSE pseudo-residuals = y − F; (3) explain why you cannot use the same formula for classification.

Stretch (homework): tune learning_rate and max_depth together with GridSearchCV; implement early stopping by monitoring a validation RMSE curve; read Chen & Guestrin 2016 Section 2 on the regularised objective.

58. Break it if you can: Show it off

Counterexample

Discussion prompt

Out loud, slides closed: (1) state the GBM update rule and name every term; (2) derive that MSE pseudo-residuals = y − F; (3) explain why you cannot use the same formula for classification.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

59. Connect it up: Lesson 50: Gradient Boosting

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — The boosting framework · From scratch: 3-round GBM · XGBoost and LightGBM · Your turn: build it. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

60. What you can do now

Recap

ideathe one thing to remember
GBM updateF_m = F_{m-1} + η·h_m; h_m fits pseudo-residuals
MSE residual−∂L/∂F = y − F (the plain residual, by coincidence)
Classificationpseudo = y − σ(F); sigmoid is essential
Regularisesmall η + CV n_estimators; depth ≤ 5
XGBoostL2 leaf reg: w* = −Σg/(Σh+λ)
LightGBMhistogram bins + leaf-wise = fast on large data

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 50 — Gradient Boosting Framework — Barron · USAAIO Round 2 Preparation, 2026
  2. From-scratch 3-round GBM, pseudo-residual math, and sklearn GBM/RF comparison on diabetes dataset verified with torch 2.7.1 + sklearn + scipy, June 2026 — Real execution, numpy 2.2.6

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108