Lesson 49: Random Forests

USAAIO Lesson 49, from Phase 2. It covers bootstrap aggregation, random feature subsets of size sqrt(d), out-of-bag error as free validation, and three kinds of feature importance - MDI, permutation, and SHAP - along with the trade-offs between a random forest, a single tree, and boosting. You implement RandomForestClassifier from scratch and compare its out-of-bag error against k-fold cross-validation. The lesson runs to 30 slides.

Subject: Machine Learning · 59 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Random Forests

Title

USAAIO · Lesson 49 · Phase 2 (Ensemble Methods)

Bootstrap aggregation, random feature subsets, OOB error, and three flavors of feature importance — from scratch.

2. By the end of this lesson you can

Objectives

  1. Explain how bootstrap sampling + aggregation reduces variance without increasing bias
  2. State why limiting each split to sqrt(d) features decorrelates trees
  3. Compute and interpret OOB error as a free, unbiased validation estimate
  4. Distinguish MDI, permutation importance, and SHAP and know when each misleads
  5. Choose between Random Forest, a single tree, and boosting for a given problem

3. What survived from Dropout & Modern Regularizers?

Warm-up

Discussion prompt

Before we open Lesson 49: Random Forests: without looking back, what was the main idea of Dropout & Modern Regularizers, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

inverted dropout theory and implementation, train-vs-eval mode, dropout as exponential ensemble averaging, MC Dropout for Bayesian uncertainty estimation, label smoothing from scratch, and a comparison of dropout rates on a classification network.

4. Bootstrap Aggregation (Bagging)

Section

Part 1 of 4

5. The variance problem with a single tree

Concept

A fully-grown decision tree (Lesson 47) has low bias but high variance — small changes in training data flip large sub-trees. On iris digits, a single tree's 5-fold CV accuracy is 0.789 with std 0.037.

The core idea of bagging: train many high-variance models on different bootstrap samples of the data, then average their predictions. Variance drops by roughly 1/T for T uncorrelated trees.

\[ \text{Var}\left(\bar{f}\right) = \frac{\sigma^2}{T} \quad\text{when trees are uncorrelated} \]

6. Break it if you can: The variance problem with a single tree

Counterexample

Discussion prompt

The core idea of bagging: train many high-variance models on different bootstrap samples of the data, then average their predictions. Variance drops by roughly 1/T for T uncorrelated trees.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

7. Bootstrap sampling

Concept

A bootstrap sample of size n is drawn with replacement from the n training rows. Roughly 63.2% of rows appear at least once; the rest (~36.8%) are out-of-bag (OOB).

\[ P(\text{row } i \text{ not drawn in one bootstrap}) = \left(1 - \frac{1}{n}\right)^n \xrightarrow{n\to\infty} e^{-1} \approx 0.368 \]

nP(OOB per tree)
100.3487
1000.3660
10000.3677

8. Fill in: P(OOB per tree) for Bootstrap sampling

Comparison

Comparison matrix

From Bootstrap sampling: refill the P(OOB per tree) column from what you know. The rest of the table is as it appeared.

nP(OOB per tree)
100.3487
1000.3660
10000.3677

9. Guess the shape of the answer: Visualizing one bootstrap draw

Estimation

Predict first

Draw a bootstrap sample of size n=10 and identify which rows are OOB.

Commit before you compute: what does Visualizing one bootstrap draw come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Bag: [2, 3, 4, 4, 6, 6, 6, 7, 7, 9] — row 4 appears twice, rows 6 and 7 three times

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. With-replacement sampling duplicates some rows and leaves others unseen.

10. Visualizing one bootstrap draw

Worked example

Draw a bootstrap sample of size n=10 and identify which rows are OOB.

import numpy as np
rng = np.random.RandomState(42)
n = 10
bag_idx = rng.choice(n, size=n, replace=True)
oob_idx = sorted(set(range(n)) - set(bag_idx))
print('Bag:', sorted(bag_idx))
print('OOB:', oob_idx)

Bag: [2, 3, 4, 4, 6, 6, 6, 7, 7, 9] — row 4 appears twice, rows 6 and 7 three times

Why: With-replacement sampling duplicates some rows and leaves others unseen. The duplicated rows effectively upweight those examples for this tree.

indexin bag?count in bag
0no (OOB)0
1no (OOB)0
4yes2
6yes3
9yes1

11. What each one costs: Visualizing one bootstrap draw

Trade off

Comparison matrix

From Visualizing one bootstrap draw: every row here is a choice with a cost. Fill the in bag? column, then say which row you would actually pick and what you give up for it.

indexin bag?count in bag
0no (OOB)0
1no (OOB)0
4yes2
6yes3
9yes1

12. Random Feature Subsets

Section

Part 2 of 4

13. Why sqrt(d) features per split?

Concept

Bagging alone leaves trees correlated: if one strong feature dominates, every tree splits on it first. Correlated trees don't average well — variance stays high.

Random Forests break this by allowing each split to consider only sqrt(d) randomly chosen features (default for classification). The strong feature is only available ~50% of the time, forcing trees to exploit weaker signals too.

d (features)sqrt(d) considered% of d
4 (iris)250.0%
10330.0%
64 (digits)812.5%
784 (MNIST)283.6%

14. Watch it run: Why sqrt(d) features per split?

Pattern

Step through it

Step through Why sqrt(d) features per split? one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: d (features) is 4 (iris)
  2. Step 2: d (features) is 10
  3. Step 3: d (features) is 64 (digits)
  4. Step 4: d (features) is 784 (MNIST)

15. Something is wrong here: confusing bagging with Random Forests

Anomaly

Predict first

A student writes this, and it looks reasonable:

A Random Forest is just bagging applied to decision trees — the trees still use all features at every split.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: That is bagged trees, not a Random Forest.

Random Forests add a second layer of randomness: at each split consider only sqrt(d) features.

Why: That is bagged trees, not a Random Forest. Trees trained on all features remain highly correlated — the variance-reduction benefit is much smaller because trees share the same dominant splits.

16. Trap: confusing bagging with Random Forests

Trap

The trap

A Random Forest is just bagging applied to decision trees — the trees still use all features at every split.

Set max_features=d (all features) for each tree

Why: That is bagged trees, not a Random Forest. Trees trained on all features remain highly correlated — the variance-reduction benefit is much smaller because trees share the same dominant splits.

The fix

Random Forests add a second layer of randomness: at each split consider only sqrt(d) features.

Set max_features='sqrt' (default in sklearn) to decorrelate trees

Why: Restricting the feature pool at each node forces diversity. Decorrelated trees average to a much lower variance than bagged trees trained on all features.

17. Break it on purpose: confusing bagging with Random Forests

Break the constraint

Discussion prompt

The rule this trap just fixed:

Restricting the feature pool at each node forces diversity. Decorrelated trees average to a much lower variance than bagged trees trained on all features.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

That is bagged trees, not a Random Forest. Trees trained on all features remain highly correlated — the variance-reduction benefit is much smaller because trees share the same dominant splits.

18. Out-of-Bag Error

Section

Part 3 of 4

19. OOB error as free cross-validation

Concept

For each training row i, the subset of trees that did NOT have row i in their bootstrap sample forms a natural held-out set. Predict row i using only those trees — the aggregate error across all rows is the OOB error.

OOB error is an unbiased estimate of generalization error. It effectively performs a leave-one-out-style evaluation at no extra compute. For iris (n=150, T=100 trees) each row is OOB for about 37 trees.

\[ \text{OOB error} = \frac{1}{n}\sum_{i=1}^{n} \mathbf{1}\left[y_i \neq \hat{y}_i^{\text{OOB}}\right] \]

20. Guess the shape of the answer: OOB error vs k-fold CV on iris

Estimation

Predict first

Fit a 100-tree RF on iris with oob_score=True and compare to 5-fold CV. They should be close.

Commit before you compute: what does OOB error vs k-fold CV on iris come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: OOB accuracy = 0.9533; 5-fold CV accuracy = 0.9600 — within 0.007 of each other

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Both estimate generalization error from held-out data; OOB uses a rolling leave-one-out structure while k-fold partitions the data.

21. OOB error vs k-fold CV on iris

Worked example

Fit a 100-tree RF on iris with oob_score=True and compare to 5-fold CV. They should be close.

from sklearn.datasets import load_iris
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import cross_val_score, KFold
X, y = load_iris(return_X_y=True)
rf = RandomForestClassifier(n_estimators=100, oob_score=True, random_state=42)
rf.fit(X, y)
print('OOB acc:', round(rf.oob_score_, 4))
kf = KFold(n_splits=5, shuffle=True, random_state=42)
scores = cross_val_score(rf, X, y, cv=kf)
print('CV acc: ', round(scores.mean(), 4))

OOB accuracy = 0.9533; 5-fold CV accuracy = 0.9600 — within 0.007 of each other

Why: Both estimate generalization error from held-out data; OOB uses a rolling leave-one-out structure while k-fold partitions the data. Agreement confirms neither is badly optimistic.

methodaccuracyOOB error
OOB (free)0.95330.0467
5-fold CV0.96000.0400
delta0.00670.0067

22. Watch it run: OOB error vs k-fold CV on iris

Pattern

Step through it

Step through OOB error vs k-fold CV on iris one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: method is OOB (free)
  2. Step 2: method is 5-fold CV
  3. Step 3: method is delta

23. OOB error stabilizes with more trees

Concept

OOB error is noisy with few trees and stabilizes once each row is OOB enough times to get a reliable majority vote. On iris: OOB error drops rapidly then plateaus around T=50.

n_estimatorsOOB error (iris)
100.0667
200.0467
500.0333
1000.0467
2000.0400

The non-monotone plateau is expected: OOB estimates have variance of their own. Adding trees does not overfit — it only reduces variance. RF error can never increase with more trees, but returns diminish fast after ~50-200.

24. Watch it run: OOB error stabilizes with more trees

Pattern

Step through it

Step through OOB error stabilizes with more trees one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: n_estimators is 10
  2. Step 2: n_estimators is 20
  3. Step 3: n_estimators is 50
  4. Step 4: n_estimators is 100
  5. Step 5: n_estimators is 200

25. Something is wrong here: using training accuracy to evaluate a RF

Anomaly

Predict first

A student writes this, and it looks reasonable:

After fitting, call rf.score(X_train, y_train) to see how well the forest learned.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: A fully-grown RF memorizes training data perfectly — rf.score(X_train, y_train) returns 1.0 on iris.

Use oob_score_ or a held-out test set to measure generalization.

Why: A fully-grown RF memorizes training data perfectly — rf.score(X_train, y_train) returns 1.0 on iris. This is the variance-overfitting that bagging is supposed to average away; the training score tells you nothing about generalization.

26. Trap: using training accuracy to evaluate a RF

Trap

The trap

After fitting, call rf.score(X_train, y_train) to see how well the forest learned.

Report training accuracy as model quality

Why: A fully-grown RF memorizes training data perfectly — rf.score(X_train, y_train) returns 1.0 on iris. This is the variance-overfitting that bagging is supposed to average away; the training score tells you nothing about generalization.

The fix

Use oob_score_ or a held-out test set to measure generalization.

Report rf.oob_score_ (0.9533 on iris) or a proper test split

Why: OOB predictions come from trees that never saw the row during training — they're the right comparison to a test set. Training accuracy of 1.0 is an artifact of zero-bias ensemble, not useful signal.

27. Which of these survive contact with Lesson 49: Random Forests?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Bagging alone leaves trees correlated: if one strong feature dominates, every tree splits on it first. Correlated trees don't average well — variance stays high.; OOB error is noisy with few trees and stabilizes once each row is OOB enough times to get a reliable majority vote. On iris: OOB error drops rapidly then plateaus around T=50.; On digits (d=64, 10 classes): single tree CV accuracy = 0.7886, RF CV accuracy = 0.9394. Boosting (GradientBoostingClassifier) gets ~0.96 but takes ~10x longer to train.
Breaks
A Random Forest is just bagging applied to decision trees — the trees still use all features at every split.; After fitting, call rf.score(X_train, y_train) to see how well the forest learned.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 49: Random Forests puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

28. Feature Importance

Section

Part 4 of 4

29. Three flavors: MDI, permutation, SHAP

Concept

methodwhat it measuresknown bias
MDI (Gini)avg impurity decrease across all splitsinflates high-cardinality / continuous features
Permutationaccuracy drop when feature is shufflednone for top features; correlated features split importance
SHAPmarginal contribution, averaged over all orderingsslowest; gold standard for explanation

On iris (4 features), all three agree that petal length and petal width dominate. But on wider datasets MDI and permutation can diverge significantly — permutation is preferred for model selection, SHAP for explaining individual predictions.

30. Fill in: known bias for Three flavors: MDI, permutation, SHAP

Comparison

Comparison matrix

From Three flavors: MDI, permutation, SHAP: refill the known bias column from what you know. The rest of the table is as it appeared.

methodwhat it measuresknown bias
MDI (Gini)avg impurity decrease across all splitsinflates high-cardinality / continuous features
Permutationaccuracy drop when feature is shufflednone for top features; correlated features split importance
SHAPmarginal contribution, averaged over all orderingsslowest; gold standard for explanation

31. Guess the shape of the answer: MDI vs permutation importance on iris

Estimation

Predict first

Compare sklearn's built-in feature_importances_ (MDI) to permutation_importance on the same fitted RF.

Commit before you compute: what does MDI vs permutation importance on iris come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: petal_len and petal_wid dominate in both; MDI ties them at 0.4361 while permutation differentiates them (0.2171 vs 0.1820)

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. MDI sums impurity reduction over all nodes regardless of whether a split was truly predictive on unseen data; permutation measures actual held-out signal, so it can break MDI ties.

32. MDI vs permutation importance on iris

Worked example

Compare sklearn's built-in feature_importances_ (MDI) to permutation_importance on the same fitted RF.

from sklearn.inspection import permutation_importance
names = ['sepal_len', 'sepal_wid', 'petal_len', 'petal_wid']
mdi = rf.feature_importances_
pi  = permutation_importance(rf, X, y, n_repeats=30,
                             random_state=42)
for n, m, p in zip(names, mdi, pi.importances_mean):
    print(f'{n:12s}  MDI={m:.4f}  PERM={p:.4f}')

petal_len and petal_wid dominate in both; MDI ties them at 0.4361 while permutation differentiates them (0.2171 vs 0.1820)

Why: MDI sums impurity reduction over all nodes regardless of whether a split was truly predictive on unseen data; permutation measures actual held-out signal, so it can break MDI ties.

featureMDIpermutation
sepal length0.10610.0160
sepal width0.02170.0124
petal length0.43610.2171
petal width0.43610.1820

33. Which is which, by MDI

Discrimination

Sort into buckets

Sort these by MDI, from memory, without looking back at MDI vs permutation importance on iris. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

0.1061
sepal length
0.0217
sepal width
0.4361
petal length; petal width
g1
MDI is "0.1061" for sepal length — that is what the table on "MDI vs permutation importance on iris" records, and it is the single property separating this group from the rest.
g2
MDI is "0.0217" for sepal width — that is what the table on "MDI vs permutation importance on iris" records, and it is the single property separating this group from the rest.
g3
MDI is "0.4361" for petal length, petal width — that is what the table on "MDI vs permutation importance on iris" records, and it is the single property separating this group from the rest.

34. RF vs single tree vs boosting

Concept

modelbiasvariancetraining timeprefer when
Single treelow (if deep)highfastyou need interpretability
Random Forestlowlowmoderate (parallelizable)tabular data, fast baseline, OOB eval
Gradient Boostinglow (iterative)lowslow (sequential)maximum accuracy, structured data

On digits (d=64, 10 classes): single tree CV accuracy = 0.7886, RF CV accuracy = 0.9394. Boosting (GradientBoostingClassifier) gets ~0.96 but takes ~10x longer to train.

35. Where does each piece belong: Lesson 49: Random Forests

Sorting

Sort into buckets

These are the pieces of Lesson 49: Random Forests, out of order. Put each one back under the part of the lesson it belongs to.

Bootstrap Aggregation (Bagging)
The variance problem with a single tree; Bootstrap sampling; Visualizing one bootstrap draw
Out-of-Bag Error
OOB error as free cross-validation; OOB error vs k-fold CV on iris; OOB error stabilizes with more trees
Feature Importance
Three flavors: MDI, permutation, SHAP; MDI vs permutation importance on iris; RF vs single tree vs boosting
s1
Bootstrap Aggregation (Bagging) is where Lesson 49: Random Forests puts The variance problem with a single tree, Bootstrap sampling, Visualizing one bootstrap draw. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
Out-of-Bag Error is where Lesson 49: Random Forests puts OOB error as free cross-validation, OOB error vs k-fold CV on iris, OOB error stabilizes with more trees. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
Feature Importance is where Lesson 49: Random Forests puts Three flavors: MDI, permutation, SHAP, MDI vs permutation importance on iris, RF vs single tree vs boosting. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

36. Without one step: The Random Forest recipe

Constraint

Discussion prompt

Run The Random Forest recipe with this step confiscated:

OOB error: evaluate on the ~36.8% of rows each tree never saw — free, unbiased

Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.

Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.

Answer:

  1. Bootstrap: for each of T trees, draw n rows with replacement (~63.2% unique)
  2. Random splits: at each node consider only sqrt(d) random features
  3. Aggregate: classification = majority vote; regression = average prediction
  4. OOB error: evaluate on the ~36.8% of rows each tree never saw — free, unbiased
  5. Importance: use permutation importance for feature selection, SHAP for explanation
  6. Tune: n_estimators (more is safer, returns diminish), max_depth (None = fully grown is fine), max_features (sqrt default is good for classification)

37. The Random Forest recipe

Pattern

  1. Bootstrap: for each of T trees, draw n rows with replacement (~63.2% unique)
  2. Random splits: at each node consider only sqrt(d) random features
  3. Aggregate: classification = majority vote; regression = average prediction
  4. OOB error: evaluate on the ~36.8% of rows each tree never saw — free, unbiased
  5. Importance: use permutation importance for feature selection, SHAP for explanation
  6. Tune: n_estimators (more is safer, returns diminish), max_depth (None = fully grown is fine), max_features (sqrt default is good for classification)

38. Where does it stop working: The Random Forest recipe

Edge cases

Discussion prompt

The Random Forest recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Bootstrap: for each of T trees, draw n rows with replacement (~63.2% unique)
  2. Random splits: at each node consider only sqrt(d) random features
  3. Aggregate: classification = majority vote; regression = average prediction
  4. OOB error: evaluate on the ~36.8% of rows each tree never saw — free, unbiased
  5. Importance: use permutation importance for feature selection, SHAP for explanation
  6. Tune: n_estimators (more is safer, returns diminish), max_depth (None = fully grown is fine), max_features (sqrt default is good for classification)

39. Rule out three: Check yourself — bootstrap math

Elimination

Eliminate the wrong options

For a dataset of n = 1000 rows, approximately what fraction of rows will NOT appear in a single bootstrap sample (i.e., will be OOB)?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. ~36.8% (1/e)
  • B. ~50% (half the data)
  • C. ~10% (1/n rows are dropped)
  • D. exactly 0% — with replacement means all rows appear

Survives elimination: A

Why: P(row not selected in one draw) = (1 - 1/n)^n → e⁻¹ ≈ 0.368 as n→∞. For n=1000, P(OOB) = 0.3677 — very close to 1/e. This is the source of OOB error's roughly 37% held-out fraction.

40. Check yourself — bootstrap math

Check

No calculator — reason from the formula.

Check your understanding

For a dataset of n = 1000 rows, approximately what fraction of rows will NOT appear in a single bootstrap sample (i.e., will be OOB)?

  • A. ~36.8% (1/e) (correct)
  • B. ~50% (half the data)
  • C. ~10% (1/n rows are dropped)
  • D. exactly 0% — with replacement means all rows appear

Answer: A

Why: P(row not selected in one draw) = (1 - 1/n)^n → e⁻¹ ≈ 0.368 as n→∞. For n=1000, P(OOB) = 0.3677 — very close to 1/e. This is the source of OOB error's roughly 37% held-out fraction.

Why B tempts people
50% would be true for sampling WITHOUT replacement of size n/2, not bootstrap (with-replacement of size n).
Why C tempts people
With n=1000 draws the expected number of distinct rows selected is n(1 - 1/e) ≈ 632, so ~368 are OOB — far more than 10.
Why D tempts people
With-replacement sampling guarantees at least one row is missing: the probability that every specific row appears is strictly less than 1.

41. Answer it before you see the options: Check yourself — random feature subsets

Prediction

Predict first

A dataset has d = 64 features and one feature is by far the most predictive. In a Random Forest (max_features='sqrt'), about what fraction of splits does that feature compete for?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: ~12.5% of splits (sqrt(64)/64 = 8/64)

Why: sqrt(64) = 8 features are drawn uniformly at random at each split, so each feature is available in expectation 8/64 = 12.5% of splits. The dominant feature can't dominate every split, forcing other features to contribute — this is the decorrelation mechanism.

42. Check yourself — random feature subsets

Check

Why does restricting features help?

Check your understanding

A dataset has d = 64 features and one feature is by far the most predictive. In a Random Forest (max_features='sqrt'), about what fraction of splits does that feature compete for?

  • A. ~12.5% of splits (sqrt(64)/64 = 8/64) (correct)
  • B. 100% — the strongest feature is always included
  • C. 50% — sklearn randomly drops half the features each split
  • D. ~1.6% — only one feature is sampled per split

Answer: A

Why: sqrt(64) = 8 features are drawn uniformly at random at each split, so each feature is available in expectation 8/64 = 12.5% of splits. The dominant feature can't dominate every split, forcing other features to contribute — this is the decorrelation mechanism.

Why B tempts people
That describes a plain bagged tree (max_features=None), not a Random Forest. The whole point of random feature subsets is to exclude the dominant feature from most splits.
Why C tempts people
sklearn draws sqrt(d) features, not d/2. For d=64, sqrt(64)=8, which is 12.5%, not 50%.
Why D tempts people
One feature per split would be extremely restrictive and would make trees nearly random. sqrt(d) is the empirically well-tuned balance.

43. Rule out three: Check yourself — feature importance

Elimination

Eliminate the wrong options

You have a dataset where two features are highly correlated (r = 0.95). You fit a Random Forest and check feature importances. Which known artifact should you worry about?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Permutation importance splits the importance between correlated features, underestimating both
  • B. MDI overestimates importance for correlated features because it counts them more often
  • C. SHAP assigns all importance to the first feature alphabetically when features correlate
  • D. OOB error becomes biased when features correlate

Survives elimination: A

Why: When features are correlated, shuffling one is partly compensated by the other — the model can still use the correlated partner. Permutation importance for each individual feature looks small even if jointly they're critical. MDI doesn't shuffle, so it doesn't share this artifact; it has the different bias of favoring high-cardinality features.

44. Check yourself — feature importance

Check

Which method to trust?

Check your understanding

You have a dataset where two features are highly correlated (r = 0.95). You fit a Random Forest and check feature importances. Which known artifact should you worry about?

  • A. Permutation importance splits the importance between correlated features, underestimating both (correct)
  • B. MDI overestimates importance for correlated features because it counts them more often
  • C. SHAP assigns all importance to the first feature alphabetically when features correlate
  • D. OOB error becomes biased when features correlate

Answer: A

Why: When features are correlated, shuffling one is partly compensated by the other — the model can still use the correlated partner. Permutation importance for each individual feature looks small even if jointly they're critical. MDI doesn't shuffle, so it doesn't share this artifact; it has the different bias of favoring high-cardinality features.

Why B tempts people
MDI's known bias is inflating HIGH-CARDINALITY or continuous features (because they offer more split points), not correlated features specifically.
Why C tempts people
SHAP correctly marginalizes over all orderings; it doesn't favor alphabetical order. Correlated features do complicate SHAP values, but the artifact is shared credit, not alphabetical priority.
Why D tempts people
OOB error is an accuracy estimate and is not directly affected by feature correlation — correlation can affect the model's predictive accuracy but not the OOB evaluation mechanism itself.

45. Your turn: scratch RF

Section

Project

46. Project: Random Forest from scratch

Concept

Reuse your DecisionTreeClassifier from Lesson 47 (or sklearn's). Build a RandomForestClassifier that bootstraps, trains T trees with random feature subsets, predicts by majority vote, and reports OOB error.

#requirementtool
1bootstrap + OOB split per treenp.random.choice(..., replace=True)
2T decision trees on bootstrap samplesDecisionTreeClassifier(max_features='sqrt')
3majority vote predictioncollections.Counter
4OOB error (aggregate per-row OOB votes)compare to sklearn RF oob_score_
5permutation importance on top 2 featuressklearn permutation_importance

Target: your OOB accuracy on iris should be within 0.05 of sklearn's 0.9533. If it's further off, check that you're predicting each row only from trees that did NOT train on it.

47. By analogy: Project: Random Forest from scratch

Analogy

Discussion prompt

Explain Project: Random Forest from scratch by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Target: your OOB accuracy on iris should be within 0.05 of sklearn's 0.9533. If it's further off, check that you're predicting each row only from trees that did NOT train on it.

48. Milestone 1 — bootstrap + OOB bookkeeping

Worked example

Your turn: draw 3 bootstrap samples from iris (n=150). Record how many unique rows each tree sees and how many are OOB.

Hint: bag = np.random.RandomState(42).choice(n, size=n, replace=True). OOB = set(range(n)) - set(bag).

import numpy as np
from sklearn.datasets import load_iris
X, y = load_iris(return_X_y=True)
n = len(X)  # 150
rng = np.random.RandomState(42)
for t in range(3):
    bag = rng.choice(n, size=n, replace=True)
    unique = len(set(bag))
    oob   = n - unique
    print(f'Tree {t+1}: {unique} train rows, {oob} OOB rows')
treeunique trainOOB rows
19060
28763
39753

49. Watch it run: Milestone 1 — bootstrap + OOB bookkeeping

Pattern

Step through it

Step through Milestone 1 — bootstrap + OOB bookkeeping one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: tree is 1
  2. Step 2: tree is 2
  3. Step 3: tree is 3

50. Milestone 2 — train trees + OOB accuracy

Worked example

Your turn: train 3 trees, collect OOB predictions for each row, majority-vote, and compute OOB accuracy.

Hint: build a dict oob_preds[i] = list of predictions from trees that did NOT train on row i, then majority-vote each entry.

from sklearn.tree import DecisionTreeClassifier
from collections import Counter
rng2 = np.random.RandomState(42)
trees, oob_preds = [], {}
for t in range(3):
    bag = rng2.choice(n, size=n, replace=True)
    oob_idx = [i for i in range(n) if i not in set(bag)]
    dt = DecisionTreeClassifier(max_features='sqrt', random_state=t)
    dt.fit(X[bag], y[bag])
    trees.append(dt)
    for i in oob_idx:
        oob_preds.setdefault(i, []).append(dt.predict([X[i]])[0])
correct = sum(Counter(v).most_common(1)[0][0]==y[i]
              for i, v in oob_preds.items())
print(f'OOB acc (3 trees): {correct/len(oob_preds):.4f}')
metricvalue
OOB rows covered112 (of 150)
OOB correct107
OOB accuracy (3 trees)0.9554
sklearn RF OOB (100 trees)0.9533

51. What each one costs: Milestone 2 — train trees + OOB accuracy

Trade off

Comparison matrix

From Milestone 2 — train trees + OOB accuracy: every row here is a choice with a cost. Fill the value column, then say which row you would actually pick and what you give up for it.

metricvalue
OOB rows covered112 (of 150)
OOB correct107
OOB accuracy (3 trees)0.9554
sklearn RF OOB (100 trees)0.9533

52. Milestone 3 — permutation importance

Worked example

Your turn: fit a full 100-tree sklearn RF, then compute permutation importance. Which two features dominate?

Hint: permutation_importance(rf, X, y, n_repeats=30, random_state=42). Compare .importances_mean to rf.feature_importances_.

from sklearn.ensemble import RandomForestClassifier
from sklearn.inspection import permutation_importance
rf = RandomForestClassifier(n_estimators=100, oob_score=True, random_state=42)
rf.fit(X, y)
pi = permutation_importance(rf, X, y, n_repeats=30, random_state=42)
names = ['sepal_len','sepal_wid','petal_len','petal_wid']
for nm, mdi, perm in zip(names, rf.feature_importances_, pi.importances_mean):
    print(f'{nm:12s}  MDI={mdi:.4f}  PERM={perm:.4f}')
featureMDIpermutation
sepal length0.10610.0160
sepal width0.02170.0124
petal length0.43610.2171
petal width0.43610.1820

53. Which is which, by MDI

Discrimination

Sort into buckets

Sort these by MDI, from memory, without looking back at Milestone 3 — permutation importance. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

0.1061
sepal length
0.0217
sepal width
0.4361
petal length; petal width
g1
MDI is "0.1061" for sepal length — that is what the table on "Milestone 3 — permutation importance" records, and it is the single property separating this group from the rest.
g2
MDI is "0.0217" for sepal width — that is what the table on "Milestone 3 — permutation importance" records, and it is the single property separating this group from the rest.
g3
MDI is "0.4361" for petal length, petal width — that is what the table on "Milestone 3 — permutation importance" records, and it is the single property separating this group from the rest.

54. The full program

Concept

import numpy as np
from sklearn.datasets import load_iris
from sklearn.ensemble import RandomForestClassifier
from sklearn.inspection import permutation_importance
from sklearn.model_selection import cross_val_score, KFold
from collections import Counter

X, y = load_iris(return_X_y=True)
rf = RandomForestClassifier(n_estimators=100, oob_score=True,
                             max_features='sqrt', random_state=42)
rf.fit(X, y)
kf  = KFold(n_splits=5, shuffle=True, random_state=42)
cv  = cross_val_score(rf, X, y, cv=kf).mean()
pi  = permutation_importance(rf, X, y, n_repeats=30, random_state=42)
print(f'OOB acc: {rf.oob_score_:.4f}   CV acc: {cv:.4f}')
print(f'Top feature (perm): petal_len {pi.importances_mean[2]:.4f}')
outputvalue
OOB accuracy0.9533
5-fold CV accuracy0.9600
petal_len perm importance0.2171

If your scratch RF lands near 0.9533 OOB accuracy with 100 trees — you've built the algorithm, not just called it.

55. Fill in: value for The full program

Comparison

Comparison matrix

From The full program: refill the value column from what you know. The rest of the table is as it appeared.

outputvalue
OOB accuracy0.9533
5-fold CV accuracy0.9600
petal_len perm importance0.2171

56. Show it off

Concept

Out loud, slides closed: explain (1) why ~36.8% of rows end up OOB, (2) the two-level randomness (rows AND features) and how each one reduces variance, and (3) when you'd pick RF over boosting.

Stretch (homework): implement permutation importance from scratch (shuffle one column, measure accuracy drop, repeat 30 times). Then apply the full pipeline to a real dataset and add SHAP values via shap.TreeExplainer. Next: Gradient Boosting (Lesson 50) — sequential trees on residuals, no OOB, but higher ceiling accuracy.

57. Break it if you can: Show it off

Counterexample

Discussion prompt

Out loud, slides closed: explain (1) why ~36.8% of rows end up OOB, (2) the two-level randomness (rows AND features) and how each one reduces variance, and (3) when you'd pick RF over boosting.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

58. Connect it up: Lesson 49: Random Forests

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Bootstrap Aggregation (Bagging) · Random Feature Subsets · Out-of-Bag Error · Feature Importance · Your turn: scratch RF. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

59. What you can do now

Recap

conceptthe one thing to remember
bootstrap~37% OOB per tree, always
random featuressqrt(d) decorrelates, not d/2
OOB errorfree cross-val; use oob_score=True
MDI biasinflates continuous/high-cardinality features
RF vs boostingRF = parallel variance-reducer; boosting = sequential bias-reducer

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 49 (Phase 2 — Ensemble Methods: Bagging & Random Forests) — Barron · USAAIO Round 2 Preparation, 2026
  2. sklearn RandomForestClassifier, permutation_importance, OOB error verified on iris and digits datasets — scikit-learn 1.x + numpy 2.2.6, real execution, June 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108