USAAIO Lesson 49, from Phase 2. It covers bootstrap aggregation, random feature subsets of size sqrt(d), out-of-bag error as free validation, and three kinds of feature importance - MDI, permutation, and SHAP - along with the trade-offs between a random forest, a single tree, and boosting. You implement RandomForestClassifier from scratch and compare its out-of-bag error against k-fold cross-validation. The lesson runs to 30 slides.
Subject: Machine Learning · 59 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 49 · Phase 2 (Ensemble Methods)
Bootstrap aggregation, random feature subsets, OOB error, and three flavors of feature importance — from scratch.
Objectives
Warm-up
Discussion prompt
Before we open Lesson 49: Random Forests: without looking back, what was the main idea of Dropout & Modern Regularizers, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
inverted dropout theory and implementation, train-vs-eval mode, dropout as exponential ensemble averaging, MC Dropout for Bayesian uncertainty estimation, label smoothing from scratch, and a comparison of dropout rates on a classification network.
Section
Part 1 of 4
Concept
A fully-grown decision tree (Lesson 47) has low bias but high variance — small changes in training data flip large sub-trees. On iris digits, a single tree's 5-fold CV accuracy is 0.789 with std 0.037.
The core idea of bagging: train many high-variance models on different bootstrap samples of the data, then average their predictions. Variance drops by roughly 1/T for T uncorrelated trees.
\[ \text{Var}\left(\bar{f}\right) = \frac{\sigma^2}{T} \quad\text{when trees are uncorrelated} \]
Counterexample
Discussion prompt
The core idea of bagging: train many high-variance models on different bootstrap samples of the data, then average their predictions. Variance drops by roughly 1/T for T uncorrelated trees.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Concept
A bootstrap sample of size n is drawn with replacement from the n training rows. Roughly 63.2% of rows appear at least once; the rest (~36.8%) are out-of-bag (OOB).
\[ P(\text{row } i \text{ not drawn in one bootstrap}) = \left(1 - \frac{1}{n}\right)^n \xrightarrow{n\to\infty} e^{-1} \approx 0.368 \]
| n | P(OOB per tree) |
|---|---|
| 10 | 0.3487 |
| 100 | 0.3660 |
| 1000 | 0.3677 |
Comparison
Comparison matrix
From Bootstrap sampling: refill the P(OOB per tree) column from what you know. The rest of the table is as it appeared.
| n | P(OOB per tree) |
|---|---|
| 10 | 0.3487 |
| 100 | 0.3660 |
| 1000 | 0.3677 |
Estimation
Predict first
Draw a bootstrap sample of size n=10 and identify which rows are OOB.
Commit before you compute: what does Visualizing one bootstrap draw come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Bag: [2, 3, 4, 4, 6, 6, 6, 7, 7, 9] — row 4 appears twice, rows 6 and 7 three times
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. With-replacement sampling duplicates some rows and leaves others unseen.
Worked example
Draw a bootstrap sample of size n=10 and identify which rows are OOB.
import numpy as np
rng = np.random.RandomState(42)
n = 10
bag_idx = rng.choice(n, size=n, replace=True)
oob_idx = sorted(set(range(n)) - set(bag_idx))
print('Bag:', sorted(bag_idx))
print('OOB:', oob_idx)Bag: [2, 3, 4, 4, 6, 6, 6, 7, 7, 9] — row 4 appears twice, rows 6 and 7 three times
Why: With-replacement sampling duplicates some rows and leaves others unseen. The duplicated rows effectively upweight those examples for this tree.
| index | in bag? | count in bag |
|---|---|---|
| 0 | no (OOB) | 0 |
| 1 | no (OOB) | 0 |
| 4 | yes | 2 |
| 6 | yes | 3 |
| 9 | yes | 1 |
Trade off
Comparison matrix
From Visualizing one bootstrap draw: every row here is a choice with a cost. Fill the in bag? column, then say which row you would actually pick and what you give up for it.
| index | in bag? | count in bag |
|---|---|---|
| 0 | no (OOB) | 0 |
| 1 | no (OOB) | 0 |
| 4 | yes | 2 |
| 6 | yes | 3 |
| 9 | yes | 1 |
Section
Part 2 of 4
Concept
Bagging alone leaves trees correlated: if one strong feature dominates, every tree splits on it first. Correlated trees don't average well — variance stays high.
Random Forests break this by allowing each split to consider only sqrt(d) randomly chosen features (default for classification). The strong feature is only available ~50% of the time, forcing trees to exploit weaker signals too.
| d (features) | sqrt(d) considered | % of d |
|---|---|---|
| 4 (iris) | 2 | 50.0% |
| 10 | 3 | 30.0% |
| 64 (digits) | 8 | 12.5% |
| 784 (MNIST) | 28 | 3.6% |
Pattern
Step through it
Step through Why sqrt(d) features per split? one row at a time. What is driving the change, and what would the row after the last one be?
Anomaly
Predict first
A student writes this, and it looks reasonable:
A Random Forest is just bagging applied to decision trees — the trees still use all features at every split.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: That is bagged trees, not a Random Forest.
Random Forests add a second layer of randomness: at each split consider only sqrt(d) features.
Why: That is bagged trees, not a Random Forest. Trees trained on all features remain highly correlated — the variance-reduction benefit is much smaller because trees share the same dominant splits.
Trap
A Random Forest is just bagging applied to decision trees — the trees still use all features at every split.
Set max_features=d (all features) for each tree
Why: That is bagged trees, not a Random Forest. Trees trained on all features remain highly correlated — the variance-reduction benefit is much smaller because trees share the same dominant splits.
Random Forests add a second layer of randomness: at each split consider only sqrt(d) features.
Set max_features='sqrt' (default in sklearn) to decorrelate trees
Why: Restricting the feature pool at each node forces diversity. Decorrelated trees average to a much lower variance than bagged trees trained on all features.
Break the constraint
Discussion prompt
The rule this trap just fixed:
Restricting the feature pool at each node forces diversity. Decorrelated trees average to a much lower variance than bagged trees trained on all features.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
That is bagged trees, not a Random Forest. Trees trained on all features remain highly correlated — the variance-reduction benefit is much smaller because trees share the same dominant splits.
Section
Part 3 of 4
Concept
For each training row i, the subset of trees that did NOT have row i in their bootstrap sample forms a natural held-out set. Predict row i using only those trees — the aggregate error across all rows is the OOB error.
OOB error is an unbiased estimate of generalization error. It effectively performs a leave-one-out-style evaluation at no extra compute. For iris (n=150, T=100 trees) each row is OOB for about 37 trees.
\[ \text{OOB error} = \frac{1}{n}\sum_{i=1}^{n} \mathbf{1}\left[y_i \neq \hat{y}_i^{\text{OOB}}\right] \]
Estimation
Predict first
Fit a 100-tree RF on iris with oob_score=True and compare to 5-fold CV. They should be close.
Commit before you compute: what does OOB error vs k-fold CV on iris come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: OOB accuracy = 0.9533; 5-fold CV accuracy = 0.9600 — within 0.007 of each other
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Both estimate generalization error from held-out data; OOB uses a rolling leave-one-out structure while k-fold partitions the data.
Worked example
Fit a 100-tree RF on iris with oob_score=True and compare to 5-fold CV. They should be close.
from sklearn.datasets import load_iris
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import cross_val_score, KFold
X, y = load_iris(return_X_y=True)
rf = RandomForestClassifier(n_estimators=100, oob_score=True, random_state=42)
rf.fit(X, y)
print('OOB acc:', round(rf.oob_score_, 4))
kf = KFold(n_splits=5, shuffle=True, random_state=42)
scores = cross_val_score(rf, X, y, cv=kf)
print('CV acc: ', round(scores.mean(), 4))OOB accuracy = 0.9533; 5-fold CV accuracy = 0.9600 — within 0.007 of each other
Why: Both estimate generalization error from held-out data; OOB uses a rolling leave-one-out structure while k-fold partitions the data. Agreement confirms neither is badly optimistic.
| method | accuracy | OOB error |
|---|---|---|
| OOB (free) | 0.9533 | 0.0467 |
| 5-fold CV | 0.9600 | 0.0400 |
| delta | 0.0067 | 0.0067 |
Pattern
Step through it
Step through OOB error vs k-fold CV on iris one row at a time. What is driving the change, and what would the row after the last one be?
Concept
OOB error is noisy with few trees and stabilizes once each row is OOB enough times to get a reliable majority vote. On iris: OOB error drops rapidly then plateaus around T=50.
| n_estimators | OOB error (iris) |
|---|---|
| 10 | 0.0667 |
| 20 | 0.0467 |
| 50 | 0.0333 |
| 100 | 0.0467 |
| 200 | 0.0400 |
The non-monotone plateau is expected: OOB estimates have variance of their own. Adding trees does not overfit — it only reduces variance. RF error can never increase with more trees, but returns diminish fast after ~50-200.
Pattern
Step through it
Step through OOB error stabilizes with more trees one row at a time. What is driving the change, and what would the row after the last one be?
Anomaly
Predict first
A student writes this, and it looks reasonable:
After fitting, call rf.score(X_train, y_train) to see how well the forest learned.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: A fully-grown RF memorizes training data perfectly — rf.score(X_train, y_train) returns 1.0 on iris.
Use oob_score_ or a held-out test set to measure generalization.
Why: A fully-grown RF memorizes training data perfectly — rf.score(X_train, y_train) returns 1.0 on iris. This is the variance-overfitting that bagging is supposed to average away; the training score tells you nothing about generalization.
Trap
After fitting, call rf.score(X_train, y_train) to see how well the forest learned.
Report training accuracy as model quality
Why: A fully-grown RF memorizes training data perfectly — rf.score(X_train, y_train) returns 1.0 on iris. This is the variance-overfitting that bagging is supposed to average away; the training score tells you nothing about generalization.
Use oob_score_ or a held-out test set to measure generalization.
Report rf.oob_score_ (0.9533 on iris) or a proper test split
Why: OOB predictions come from trees that never saw the row during training — they're the right comparison to a test set. Training accuracy of 1.0 is an artifact of zero-bias ensemble, not useful signal.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
rf.score(X_train, y_train) to see how well the forest learned.Section
Part 4 of 4
Concept
| method | what it measures | known bias |
|---|---|---|
| MDI (Gini) | avg impurity decrease across all splits | inflates high-cardinality / continuous features |
| Permutation | accuracy drop when feature is shuffled | none for top features; correlated features split importance |
| SHAP | marginal contribution, averaged over all orderings | slowest; gold standard for explanation |
On iris (4 features), all three agree that petal length and petal width dominate. But on wider datasets MDI and permutation can diverge significantly — permutation is preferred for model selection, SHAP for explaining individual predictions.
Comparison
Comparison matrix
From Three flavors: MDI, permutation, SHAP: refill the known bias column from what you know. The rest of the table is as it appeared.
| method | what it measures | known bias |
|---|---|---|
| MDI (Gini) | avg impurity decrease across all splits | inflates high-cardinality / continuous features |
| Permutation | accuracy drop when feature is shuffled | none for top features; correlated features split importance |
| SHAP | marginal contribution, averaged over all orderings | slowest; gold standard for explanation |
Estimation
Predict first
Compare sklearn's built-in feature_importances_ (MDI) to permutation_importance on the same fitted RF.
Commit before you compute: what does MDI vs permutation importance on iris come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: petal_len and petal_wid dominate in both; MDI ties them at 0.4361 while permutation differentiates them (0.2171 vs 0.1820)
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. MDI sums impurity reduction over all nodes regardless of whether a split was truly predictive on unseen data; permutation measures actual held-out signal, so it can break MDI ties.
Worked example
Compare sklearn's built-in feature_importances_ (MDI) to permutation_importance on the same fitted RF.
from sklearn.inspection import permutation_importance
names = ['sepal_len', 'sepal_wid', 'petal_len', 'petal_wid']
mdi = rf.feature_importances_
pi = permutation_importance(rf, X, y, n_repeats=30,
random_state=42)
for n, m, p in zip(names, mdi, pi.importances_mean):
print(f'{n:12s} MDI={m:.4f} PERM={p:.4f}')petal_len and petal_wid dominate in both; MDI ties them at 0.4361 while permutation differentiates them (0.2171 vs 0.1820)
Why: MDI sums impurity reduction over all nodes regardless of whether a split was truly predictive on unseen data; permutation measures actual held-out signal, so it can break MDI ties.
| feature | MDI | permutation |
|---|---|---|
| sepal length | 0.1061 | 0.0160 |
| sepal width | 0.0217 | 0.0124 |
| petal length | 0.4361 | 0.2171 |
| petal width | 0.4361 | 0.1820 |
Discrimination
Sort into buckets
Sort these by MDI, from memory, without looking back at MDI vs permutation importance on iris. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Concept
| model | bias | variance | training time | prefer when |
|---|---|---|---|---|
| Single tree | low (if deep) | high | fast | you need interpretability |
| Random Forest | low | low | moderate (parallelizable) | tabular data, fast baseline, OOB eval |
| Gradient Boosting | low (iterative) | low | slow (sequential) | maximum accuracy, structured data |
On digits (d=64, 10 classes): single tree CV accuracy = 0.7886, RF CV accuracy = 0.9394. Boosting (GradientBoostingClassifier) gets ~0.96 but takes ~10x longer to train.
Sorting
Sort into buckets
These are the pieces of Lesson 49: Random Forests, out of order. Put each one back under the part of the lesson it belongs to.
Constraint
Discussion prompt
Run The Random Forest recipe with this step confiscated:
OOB error: evaluate on the ~36.8% of rows each tree never saw — free, unbiased
Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.
Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.
Answer:
sqrt(d) random featuresn_estimators (more is safer, returns diminish), max_depth (None = fully grown is fine), max_features (sqrt default is good for classification)Pattern
sqrt(d) random featuresn_estimators (more is safer, returns diminish), max_depth (None = fully grown is fine), max_features (sqrt default is good for classification)Edge cases
Discussion prompt
The Random Forest recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
sqrt(d) random featuresn_estimators (more is safer, returns diminish), max_depth (None = fully grown is fine), max_features (sqrt default is good for classification)Elimination
Eliminate the wrong options
For a dataset of n = 1000 rows, approximately what fraction of rows will NOT appear in a single bootstrap sample (i.e., will be OOB)?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: P(row not selected in one draw) = (1 - 1/n)^n → e⁻¹ ≈ 0.368 as n→∞. For n=1000, P(OOB) = 0.3677 — very close to 1/e. This is the source of OOB error's roughly 37% held-out fraction.
Check
No calculator — reason from the formula.
Check your understanding
For a dataset of n = 1000 rows, approximately what fraction of rows will NOT appear in a single bootstrap sample (i.e., will be OOB)?
Answer: A
Why: P(row not selected in one draw) = (1 - 1/n)^n → e⁻¹ ≈ 0.368 as n→∞. For n=1000, P(OOB) = 0.3677 — very close to 1/e. This is the source of OOB error's roughly 37% held-out fraction.
Prediction
Predict first
A dataset has d = 64 features and one feature is by far the most predictive. In a Random Forest (max_features='sqrt'), about what fraction of splits does that feature compete for?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: ~12.5% of splits (sqrt(64)/64 = 8/64)
Why: sqrt(64) = 8 features are drawn uniformly at random at each split, so each feature is available in expectation 8/64 = 12.5% of splits. The dominant feature can't dominate every split, forcing other features to contribute — this is the decorrelation mechanism.
Check
Why does restricting features help?
Check your understanding
A dataset has d = 64 features and one feature is by far the most predictive. In a Random Forest (max_features='sqrt'), about what fraction of splits does that feature compete for?
Answer: A
Why: sqrt(64) = 8 features are drawn uniformly at random at each split, so each feature is available in expectation 8/64 = 12.5% of splits. The dominant feature can't dominate every split, forcing other features to contribute — this is the decorrelation mechanism.
Elimination
Eliminate the wrong options
You have a dataset where two features are highly correlated (r = 0.95). You fit a Random Forest and check feature importances. Which known artifact should you worry about?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: When features are correlated, shuffling one is partly compensated by the other — the model can still use the correlated partner. Permutation importance for each individual feature looks small even if jointly they're critical. MDI doesn't shuffle, so it doesn't share this artifact; it has the different bias of favoring high-cardinality features.
Check
Which method to trust?
Check your understanding
You have a dataset where two features are highly correlated (r = 0.95). You fit a Random Forest and check feature importances. Which known artifact should you worry about?
Answer: A
Why: When features are correlated, shuffling one is partly compensated by the other — the model can still use the correlated partner. Permutation importance for each individual feature looks small even if jointly they're critical. MDI doesn't shuffle, so it doesn't share this artifact; it has the different bias of favoring high-cardinality features.
Section
Project
Concept
Reuse your DecisionTreeClassifier from Lesson 47 (or sklearn's). Build a RandomForestClassifier that bootstraps, trains T trees with random feature subsets, predicts by majority vote, and reports OOB error.
| # | requirement | tool |
|---|---|---|
| 1 | bootstrap + OOB split per tree | np.random.choice(..., replace=True) |
| 2 | T decision trees on bootstrap samples | DecisionTreeClassifier(max_features='sqrt') |
| 3 | majority vote prediction | collections.Counter |
| 4 | OOB error (aggregate per-row OOB votes) | compare to sklearn RF oob_score_ |
| 5 | permutation importance on top 2 features | sklearn permutation_importance |
Target: your OOB accuracy on iris should be within 0.05 of sklearn's 0.9533. If it's further off, check that you're predicting each row only from trees that did NOT train on it.
Analogy
Discussion prompt
Explain Project: Random Forest from scratch by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Target: your OOB accuracy on iris should be within 0.05 of sklearn's 0.9533. If it's further off, check that you're predicting each row only from trees that did NOT train on it.
Worked example
Your turn: draw 3 bootstrap samples from iris (n=150). Record how many unique rows each tree sees and how many are OOB.
Hint: bag = np.random.RandomState(42).choice(n, size=n, replace=True). OOB = set(range(n)) - set(bag).
import numpy as np
from sklearn.datasets import load_iris
X, y = load_iris(return_X_y=True)
n = len(X) # 150
rng = np.random.RandomState(42)
for t in range(3):
bag = rng.choice(n, size=n, replace=True)
unique = len(set(bag))
oob = n - unique
print(f'Tree {t+1}: {unique} train rows, {oob} OOB rows')| tree | unique train | OOB rows |
|---|---|---|
| 1 | 90 | 60 |
| 2 | 87 | 63 |
| 3 | 97 | 53 |
Pattern
Step through it
Step through Milestone 1 — bootstrap + OOB bookkeeping one row at a time. What is driving the change, and what would the row after the last one be?
Worked example
Your turn: train 3 trees, collect OOB predictions for each row, majority-vote, and compute OOB accuracy.
Hint: build a dict oob_preds[i] = list of predictions from trees that did NOT train on row i, then majority-vote each entry.
from sklearn.tree import DecisionTreeClassifier
from collections import Counter
rng2 = np.random.RandomState(42)
trees, oob_preds = [], {}
for t in range(3):
bag = rng2.choice(n, size=n, replace=True)
oob_idx = [i for i in range(n) if i not in set(bag)]
dt = DecisionTreeClassifier(max_features='sqrt', random_state=t)
dt.fit(X[bag], y[bag])
trees.append(dt)
for i in oob_idx:
oob_preds.setdefault(i, []).append(dt.predict([X[i]])[0])
correct = sum(Counter(v).most_common(1)[0][0]==y[i]
for i, v in oob_preds.items())
print(f'OOB acc (3 trees): {correct/len(oob_preds):.4f}')| metric | value |
|---|---|
| OOB rows covered | 112 (of 150) |
| OOB correct | 107 |
| OOB accuracy (3 trees) | 0.9554 |
| sklearn RF OOB (100 trees) | 0.9533 |
Trade off
Comparison matrix
From Milestone 2 — train trees + OOB accuracy: every row here is a choice with a cost. Fill the value column, then say which row you would actually pick and what you give up for it.
| metric | value |
|---|---|
| OOB rows covered | 112 (of 150) |
| OOB correct | 107 |
| OOB accuracy (3 trees) | 0.9554 |
| sklearn RF OOB (100 trees) | 0.9533 |
Worked example
Your turn: fit a full 100-tree sklearn RF, then compute permutation importance. Which two features dominate?
Hint: permutation_importance(rf, X, y, n_repeats=30, random_state=42). Compare .importances_mean to rf.feature_importances_.
from sklearn.ensemble import RandomForestClassifier
from sklearn.inspection import permutation_importance
rf = RandomForestClassifier(n_estimators=100, oob_score=True, random_state=42)
rf.fit(X, y)
pi = permutation_importance(rf, X, y, n_repeats=30, random_state=42)
names = ['sepal_len','sepal_wid','petal_len','petal_wid']
for nm, mdi, perm in zip(names, rf.feature_importances_, pi.importances_mean):
print(f'{nm:12s} MDI={mdi:.4f} PERM={perm:.4f}')| feature | MDI | permutation |
|---|---|---|
| sepal length | 0.1061 | 0.0160 |
| sepal width | 0.0217 | 0.0124 |
| petal length | 0.4361 | 0.2171 |
| petal width | 0.4361 | 0.1820 |
Discrimination
Sort into buckets
Sort these by MDI, from memory, without looking back at Milestone 3 — permutation importance. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Concept
import numpy as np
from sklearn.datasets import load_iris
from sklearn.ensemble import RandomForestClassifier
from sklearn.inspection import permutation_importance
from sklearn.model_selection import cross_val_score, KFold
from collections import Counter
X, y = load_iris(return_X_y=True)
rf = RandomForestClassifier(n_estimators=100, oob_score=True,
max_features='sqrt', random_state=42)
rf.fit(X, y)
kf = KFold(n_splits=5, shuffle=True, random_state=42)
cv = cross_val_score(rf, X, y, cv=kf).mean()
pi = permutation_importance(rf, X, y, n_repeats=30, random_state=42)
print(f'OOB acc: {rf.oob_score_:.4f} CV acc: {cv:.4f}')
print(f'Top feature (perm): petal_len {pi.importances_mean[2]:.4f}')| output | value |
|---|---|
| OOB accuracy | 0.9533 |
| 5-fold CV accuracy | 0.9600 |
| petal_len perm importance | 0.2171 |
If your scratch RF lands near 0.9533 OOB accuracy with 100 trees — you've built the algorithm, not just called it.
Comparison
Comparison matrix
From The full program: refill the value column from what you know. The rest of the table is as it appeared.
| output | value |
|---|---|
| OOB accuracy | 0.9533 |
| 5-fold CV accuracy | 0.9600 |
| petal_len perm importance | 0.2171 |
Concept
Out loud, slides closed: explain (1) why ~36.8% of rows end up OOB, (2) the two-level randomness (rows AND features) and how each one reduces variance, and (3) when you'd pick RF over boosting.
Stretch (homework): implement permutation importance from scratch (shuffle one column, measure accuracy drop, repeat 30 times). Then apply the full pipeline to a real dataset and add SHAP values via shap.TreeExplainer. Next: Gradient Boosting (Lesson 50) — sequential trees on residuals, no OOB, but higher ceiling accuracy.
Counterexample
Discussion prompt
Out loud, slides closed: explain (1) why ~36.8% of rows end up OOB, (2) the two-level randomness (rows AND features) and how each one reduces variance, and (3) when you'd pick RF over boosting.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Bootstrap Aggregation (Bagging) · Random Feature Subsets · Out-of-Bag Error · Feature Importance · Your turn: scratch RF. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
| concept | the one thing to remember |
|---|---|
| bootstrap | ~37% OOB per tree, always |
| random features | sqrt(d) decorrelates, not d/2 |
| OOB error | free cross-val; use oob_score=True |
| MDI bias | inflates continuous/high-cardinality features |
| RF vs boosting | RF = parallel variance-reducer; boosting = sequential bias-reducer |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.