USAAIO Lesson 67, from Phase 3. It covers stacking, using a 5-fold out-of-fold meta-learner, and blending, using a 20% hold-out, then explains why diversity beats homogeneity and how to combine heterogeneous models such as logistic regression, a random forest, and kNN. It closes with the analytic variance-reduction proof for averaging. All the numbers were verified with sklearn 1.x and numpy 2.2.6. The lesson runs to 28 slides.
Subject: Machine Learning · 52 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 67 · Phase 3
Move past single-model limits: stack diverse learners with out-of-fold predictions, blend on a hold-out set, and prove analytically why diversity reduces variance.
Objectives
Warm-up
Discussion prompt
Before we open Lesson 67: Stacking & Blending Ensembles: without looking back, what was the main idea of Mixed Precision, Loss Scaling & Gradient Accumulation, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
Build an AMP training loop verified against full-precision.
Section
Part 1 of 4
Concept
A single model with variance sigma^2 can be replaced by the average of M independent models with the same variance — but the average's variance shrinks to sigma^2 / M (Lesson 32 callback).
\[ \mathrm{Var}\!\left(\bar{f}\right) = \rho\,\sigma^2 + \frac{1-\rho}{M}\,\sigma^2 \]
rho is the pairwise correlation of model errors. At rho = 0 (perfectly uncorrelated), Var shrinks to sigma^2/M. At rho = 1 (identical models), nothing helps.
Counterexample
Discussion prompt
A single model with variance sigma^2 can be replaced by the average of M independent models with the same variance — but the average's variance shrinks to sigma^2 / M (Lesson 32 callback).
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Concept
| rho | M=5, sigma2=0.10 | ratio vs single |
|---|---|---|
| 0.0 | 0.0200 | 0.20x |
| 0.3 | 0.0440 | 0.44x |
| 0.5 | 0.0600 | 0.60x |
| 0.9 | 0.0920 | 0.92x |
| 1.0 | 0.1000 | 1.00x (no gain) |
Diverse models (low rho) give the biggest variance reduction. This is why stacking heterogeneous base learners outperforms stacking five copies of the same model.
Comparison
Comparison matrix
From Variance vs correlation trade-off: refill the M=5, sigma2=0.10 column from what you know. The rest of the table is as it appeared.
| rho | M=5, sigma2=0.10 | ratio vs single |
|---|---|---|
| 0.0 | 0.0200 | 0.20x |
| 0.3 | 0.0440 | 0.44x |
| 0.5 | 0.0600 | 0.60x |
| 0.9 | 0.0920 | 0.92x |
| 1.0 | 0.1000 | 1.00x (no gain) |
Section
Part 2 of 4
Concept
Layer 1 — base models (diverse). Layer 2 — meta-learner trained on the base models' predictions. The trick: base models must predict on data they have not seen, or the meta-learner overfits.
Estimation
Predict first
Load load_digits (1797 samples, 10 classes). Base models: LogReg, RF, KNN. Meta-learner: LogReg on concatenated predict_proba OOF features.
Commit before you compute: what does Stacking from scratch — 5-fold OOF come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Stacking accuracy: 0.9861 vs best single (LogReg) 0.9722
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The meta-learner sees calibrated probability vectors from three diverse models.
Worked example
Load load_digits (1797 samples, 10 classes). Base models: LogReg, RF, KNN. Meta-learner: LogReg on concatenated predict_proba OOF features.
import numpy as np
from sklearn.datasets import load_digits
from sklearn.model_selection import StratifiedKFold, train_test_split
from sklearn.linear_model import LogisticRegression
from sklearn.ensemble import RandomForestClassifier
from sklearn.neighbors import KNeighborsClassifier
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.base import clone
from sklearn.metrics import accuracy_score
X, y = load_digits(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42, stratify=y)
base_specs = [
('LR', Pipeline([('sc', StandardScaler()),
('clf', LogisticRegression(max_iter=2000, random_state=42))])),
('RF', RandomForestClassifier(n_estimators=100, random_state=42)),
('KNN', Pipeline([('sc', StandardScaler()),
('clf', KNeighborsClassifier(n_neighbors=3))])),
]
n_cls, n_base = 10, 3
skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
oof_proba = np.zeros((len(X_train), n_base * n_cls))
test_proba = np.zeros((len(X_test), n_base * n_cls))
for j, (name, spec) in enumerate(base_specs):
fold_tp = np.zeros((len(X_test), n_cls, 5))
for fi, (tr, val) in enumerate(skf.split(X_train, y_train)):
m = clone(spec)
m.fit(X_train[tr], y_train[tr])
oof_proba[val, j*n_cls:(j+1)*n_cls] = m.predict_proba(X_train[val])
fold_tp[:, :, fi] = m.predict_proba(X_test)
test_proba[:, j*n_cls:(j+1)*n_cls] = fold_tp.mean(axis=2)
meta = LogisticRegression(max_iter=3000, C=0.1, random_state=42)
meta.fit(oof_proba, y_train)
print(accuracy_score(y_test, meta.predict(test_proba)))Stacking accuracy: 0.9861 vs best single (LogReg) 0.9722
Why: The meta-learner sees calibrated probability vectors from three diverse models. It learns that when KNN and RF agree but LogReg is uncertain, it can safely boost confidence — an information source a single model cannot exploit.
| model | accuracy |
|---|---|
| LogReg (alone) | 0.9722 |
| RF (alone) | 0.9611 |
| KNN (alone) | 0.9667 |
| Majority vote | 0.9806 |
| Stacking (proba OOF) | 0.9861 |
Trade off
Comparison matrix
From Stacking from scratch — 5-fold OOF: every row here is a choice with a cost. Fill the accuracy column, then say which row you would actually pick and what you give up for it.
| model | accuracy |
|---|---|
| LogReg (alone) | 0.9722 |
| RF (alone) | 0.9611 |
| KNN (alone) | 0.9667 |
| Majority vote | 0.9806 |
| Stacking (proba OOF) | 0.9861 |
Concept
Each OOF fold predicts exactly once for every training sample. With n_train = 1437 and K = 5, the folds are:
| fold | train rows | val rows (OOF) |
|---|---|---|
| 1 | 1149 | 288 |
| 2 | 1149 | 288 |
| 3 | 1150 | 287 |
| 4 | 1150 | 287 |
| 5 | 1150 | 287 |
The 1437 OOF predictions are then used as training data for the meta-learner — no sample was predicted by a model it was trained on.
Pattern
Step through it
Step through 5-fold split sizes one row at a time. What is driving the change, and what would the row after the last one be?
Anomaly
Predict first
A student writes this, and it looks reasonable:
Fit each base model on the full training set. Use those in-sample predictions as meta-features.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Each base model has seen every training sample.
Use out-of-fold (OOF) predictions: each training sample is predicted by a base model that was never trained on that sample.
Why: Each base model has seen every training sample. Its in-sample predictions are nearly perfect for well-fit models — the meta-learner trains on overfit signals, then sees test data where those signals are noisy. The meta-learner overfits to the base models' memorized outputs.
Trap
Fit each base model on the full training set. Use those in-sample predictions as meta-features.
base.fit(X_train, y_train) --> meta.fit(base.predict_proba(X_train), y_train)
Why: Each base model has seen every training sample. Its in-sample predictions are nearly perfect for well-fit models — the meta-learner trains on overfit signals, then sees test data where those signals are noisy. The meta-learner overfits to the base models' memorized outputs.
Use out-of-fold (OOF) predictions: each training sample is predicted by a base model that was never trained on that sample.
5-fold CV --> oof_proba[val] = base_fold.predict_proba(X_train[val]) --> meta.fit(oof_proba, y_train)
Why: OOF predictions are honest — they simulate test-time uncertainty. The meta-learner trains on signals that generalize, not memorized perfect probabilities.
Break the constraint
Discussion prompt
The rule this trap just fixed:
OOF predictions are honest — they simulate test-time uncertainty. The meta-learner trains on signals that generalize, not memorized perfect probabilities.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
Each base model has seen every training sample. Its in-sample predictions are nearly perfect for well-fit models — the meta-learner trains on overfit signals, then sees test data where those signals are noisy. The meta-learner overfits to the base models' memorized outputs.
Section
Part 3 of 4
Concept
| property | stacking (K-fold OOF) | blending (hold-out) |
|---|---|---|
| meta-train size | all n_train samples | ~20% of n_train |
| data efficiency | high | lower (wastes 80% for base) |
| leakage risk | none (OOF) | none (separate split) |
| computation | K x n_base model fits | 1 x n_base model fits |
| typical use | competition accuracy | quick baseline |
Blending is faster (train each base model once) but the meta-learner sees only ~20% of the data. Stacking uses all training data for the meta-learner, at the cost of K times more base-model fits.
Analogy
Discussion prompt
Explain Blending vs stacking by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Blending is faster (train each base model once) but the meta-learner sees only ~20% of the data. Stacking uses all training data for the meta-learner, at the cost of K times more base-model fits.
Estimation
Predict first
Reserve 20% of training data for meta-learner training. Fit base models on the remaining 80%, then predict the 20% hold-out and the test set.
Commit before you compute: what does Blending — hold-out 20% for meta come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Blending accuracy: 0.9861 (equal to stacking on this dataset)
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. On load_digits the meta-learner converges with fewer samples.
Worked example
Reserve 20% of training data for meta-learner training. Fit base models on the remaining 80%, then predict the 20% hold-out and the test set.
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
from sklearn.base import clone
from sklearn.metrics import accuracy_score
import numpy as np
X_bl, X_bval, y_bl, y_bval = train_test_split(
X_train, y_train, test_size=0.2, random_state=42, stratify=y_train)
print(f'Base train: {len(y_bl)}, Meta train: {len(y_bval)}')
bl_val_p, bl_test_p = [], []
for name, spec in base_specs:
m = clone(spec); m.fit(X_bl, y_bl)
bl_val_p.append(m.predict_proba(X_bval))
bl_test_p.append(m.predict_proba(X_test))
meta_b = LogisticRegression(max_iter=2000, C=0.1, random_state=42)
meta_b.fit(np.hstack(bl_val_p), y_bval)
print(accuracy_score(y_test, meta_b.predict(np.hstack(bl_test_p))))Blending accuracy: 0.9861 (equal to stacking on this dataset)
Why: On load_digits the meta-learner converges with fewer samples. On noisier or higher-dimensional problems, stacking's larger meta-training set typically wins.
| method | meta train n | accuracy |
|---|---|---|
| blending (20% hold-out) | 288 | 0.9861 |
| stacking (5-fold OOF) | 1437 | 0.9861 |
| best single model | — | 0.9722 |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Blending accuracy: 0.9861 (equal to stacking on this dataset)
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Reserve 20% of training data for meta-learner training. Fit base models on the remaining 80%, then predict the 20% hold-out and the test set.
Anomaly
Predict first
A student writes this, and it looks reasonable:
After fitting the meta-learner on the 20% hold-out, report accuracy on that same 20% as the ensemble's performance.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: The meta-learner was trained on that hold-out set — reporting accuracy there is optimistic in-sample error.
Evaluate the full stack (base models + meta-learner) on the held-out test set that neither the base models nor the meta-learner touched during training.
Why: The meta-learner was trained on that hold-out set — reporting accuracy there is optimistic in-sample error. You've measured how well the meta-learner memorized its own training set, not how it generalizes.
Trap
After fitting the meta-learner on the 20% hold-out, report accuracy on that same 20% as the ensemble's performance.
meta.fit(bl_val_feat, y_bval) --> score on bl_val
Why: The meta-learner was trained on that hold-out set — reporting accuracy there is optimistic in-sample error. You've measured how well the meta-learner memorized its own training set, not how it generalizes.
Evaluate the full stack (base models + meta-learner) on the held-out test set that neither the base models nor the meta-learner touched during training.
test_feat = [m.predict_proba(X_test) for m in base_models] --> meta.predict(test_feat)
Why: The test set is neutral to both layers. Reporting accuracy there gives an unbiased estimate of ensemble generalization.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
Section
Part 4 of 4
Concept
Two models are diverse when they make different errors on the same samples. The meta-learner can then correct one model's mistakes using the other's confidence.
Stacking five identical LogReg models gives rho close to 1 — the variance formula predicts almost zero gain (Lesson 67's analytic table).
Analogy
Discussion prompt
Explain What makes base models diverse? by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Two models are diverse when they make different errors on the same samples. The meta-learner can then correct one model's mistakes using the other's confidence.
Concept
AutoML systems (H2O AutoML, AutoSklearn) automate the stacking process: search over model families, tune hyperparameters, and assemble a stacking ensemble — often producing near-Pareto-optimal ensembles with no manual tuning.
| system | base models | meta strategy |
|---|---|---|
| H2O AutoML | GLM, GBM, RF, DNN, XGBoost | Stacked Ensemble (OOF) |
| AutoSklearn | sklearn models + preprocessing | Ensemble selection from evaluated configs |
| manual stacking | your choice | your meta-learner |
Understanding OOF stacking from scratch (this lesson) is prerequisite knowledge for interpreting what AutoML does and diagnosing when it fails.
Comparison
Comparison matrix
From AutoML: automated stacking at scale: refill the meta strategy column from what you know. The rest of the table is as it appeared.
| system | base models | meta strategy |
|---|---|---|
| H2O AutoML | GLM, GBM, RF, DNN, XGBoost | Stacked Ensemble (OOF) |
| AutoSklearn | sklearn models + preprocessing | Ensemble selection from evaluated configs |
| manual stacking | your choice | your meta-learner |
Ranking
Put in order
These are the steps of The stacking/blending recipe, scrambled. Put them back in order before the next slide shows you.
predict_proba → concatenate → meta-learnerC=0.1) on probability vectors; never train meta on in-fold predsWhy: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
predict_proba → concatenate → meta-learnerC=0.1) on probability vectors; never train meta on in-fold predsEdge cases
Discussion prompt
The stacking/blending recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
predict_proba → concatenate → meta-learnerC=0.1) on probability vectors; never train meta on in-fold predsElimination
Eliminate the wrong options
Why must meta-learner training features come from out-of-fold (OOF) predictions rather than in-sample predictions?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: A well-fit base model achieves near-perfect accuracy on its own training samples — these in-sample probabilities are overfit signals. The meta-learner would learn to trust overconfident probabilities that don't appear at test time. OOF predictions come from a held-out fold, accurately reflecting how the model behaves on unseen data.
Check
Identify the key step.
Check your understanding
Why must meta-learner training features come from out-of-fold (OOF) predictions rather than in-sample predictions?
Answer: A
Why: A well-fit base model achieves near-perfect accuracy on its own training samples — these in-sample probabilities are overfit signals. The meta-learner would learn to trust overconfident probabilities that don't appear at test time. OOF predictions come from a held-out fold, accurately reflecting how the model behaves on unseen data.
Prediction
Predict first
M=5 diverse models each have variance sigma^2=0.10 and pairwise error correlation rho=0.3. What is Var(average)?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: 0.0440
Why: Var(avg) = rhosigma^2 + (1-rho)sigma^2/M = 0.30.10 + 0.70.10/5 = 0.030 + 0.014 = 0.0440. This is 2.27x lower than a single model's 0.10.
Check
Apply the formula.
Check your understanding
M=5 diverse models each have variance sigma^2=0.10 and pairwise error correlation rho=0.3. What is Var(average)?
Answer: A
Why: Var(avg) = rhosigma^2 + (1-rho)sigma^2/M = 0.30.10 + 0.70.10/5 = 0.030 + 0.014 = 0.0440. This is 2.27x lower than a single model's 0.10.
Elimination
Eliminate the wrong options
Compared with 5-fold stacking, blending with a 20% hold-out is:
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Blending trains each base model only once (on 80% of the data), so it requires 1/5 the compute of 5-fold stacking. But the meta-learner trains on only 20% of the training labels (~288 samples here vs 1437 for stacking). On small datasets this is a meaningful accuracy gap.
Check
Compare the two approaches.
Check your understanding
Compared with 5-fold stacking, blending with a 20% hold-out is:
Answer: A
Why: Blending trains each base model only once (on 80% of the data), so it requires 1/5 the compute of 5-fold stacking. But the meta-learner trains on only 20% of the training labels (~288 samples here vs 1437 for stacking). On small datasets this is a meaningful accuracy gap.
Section
Project
Concept
Implement a stacking ensemble on load_digits with LogReg, RF, and KNN as base models and LogReg as the meta-learner. Prove it beats the best individual model.
| # | milestone | key call |
|---|---|---|
| 1 | 5-fold OOF proba features | StratifiedKFold + clone + predict_proba |
| 2 | meta-learner training | LogReg(C=0.1).fit(oof_proba, y_train) |
| 3 | test-time prediction | test_proba = avg of fold predictions |
| 4 | accuracy table | accuracy_score vs singles + vote |
Build rules: never let a base model predict its own training fold; use predict_proba not predict for meta features; evaluate only on the untouched test set.
Counterexample
Discussion prompt
Implement a stacking ensemble on load_digits with LogReg, RF, and KNN as base models and LogReg as the meta-learner. Prove it beats the best individual model.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Build rules: never let a base model predict its own training fold; use predict_proba not predict for meta features; evaluate only on the untouched test set.
Worked example
Your turn: set up 5-fold CV and collect OOF probability vectors for all three base models.
Hint: oof_proba = np.zeros((n_train, n_base * n_cls)) — each model contributes 10 probability columns. Use clone(spec) inside the fold loop so each fold gets a fresh model.
import numpy as np
from sklearn.datasets import load_digits
from sklearn.model_selection import StratifiedKFold, train_test_split
from sklearn.linear_model import LogisticRegression
from sklearn.ensemble import RandomForestClassifier
from sklearn.neighbors import KNeighborsClassifier
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.base import clone
X, y = load_digits(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42, stratify=y)
base_specs = [
('LR', Pipeline([('sc', StandardScaler()),
('clf', LogisticRegression(max_iter=2000, random_state=42))])),
('RF', RandomForestClassifier(n_estimators=100, random_state=42)),
('KNN', Pipeline([('sc', StandardScaler()),
('clf', KNeighborsClassifier(n_neighbors=3))])),
]
n_cls, n_base = 10, 3
skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
oof_proba = np.zeros((len(X_train), n_base * n_cls))
print(oof_proba.shape) # (1437, 30)| array | shape | meaning |
|---|---|---|
| oof_proba | (1437, 30) | 10 proba cols per base model x 3 models |
| block j cols | j10 : (j+1)10 | proba for base model j |
| fold val rows | 288 or 287 | OOF rows filled each fold |
Worked example
Your turn: fill oof_proba across folds, then train the meta-learner.
Hint: inner loop for fi, (tr, val) in enumerate(skf.split(...)) — fit clone(spec) on X_train[tr], predict X_train[val], store in oof_proba[val, j*10:(j+1)*10].
test_proba = np.zeros((len(X_test), n_base * n_cls))
for j, (name, spec) in enumerate(base_specs):
fold_tp = np.zeros((len(X_test), n_cls, 5))
for fi, (tr, val) in enumerate(skf.split(X_train, y_train)):
m = clone(spec)
m.fit(X_train[tr], y_train[tr])
oof_proba[val, j*n_cls:(j+1)*n_cls] = m.predict_proba(X_train[val])
fold_tp[:, :, fi] = m.predict_proba(X_test)
test_proba[:, j*n_cls:(j+1)*n_cls] = fold_tp.mean(axis=2)
meta = LogisticRegression(max_iter=3000, C=0.1, random_state=42)
meta.fit(oof_proba, y_train)
print(meta.coef_.shape) # (10, 30) - 10-class meta-learner| step | output shape |
|---|---|
| oof_proba after loop | (1437, 30) |
| test_proba (avg over folds) | (360, 30) |
| meta.coef_ | (10, 30) |
Trade off
Comparison matrix
From Milestone 2 — fill OOF and train meta: every row here is a choice with a cost. Fill the output shape column, then say which row you would actually pick and what you give up for it.
| step | output shape |
|---|---|
| oof_proba after loop | (1437, 30) |
| test_proba (avg over folds) | (360, 30) |
| meta.coef_ | (10, 30) |
Worked example
Your turn: predict the test set with the stacked meta-learner and compare against individual models and majority vote.
Hint: meta.predict(test_proba) gives the ensemble prediction. Compare with each base model trained on full X_train.
from sklearn.metrics import accuracy_score
from scipy import stats
stack_acc = accuracy_score(y_test, meta.predict(test_proba))
print(f'Stack: {stack_acc:.4f}')
indiv, mv_preds = {}, []
for name, spec in base_specs:
m = clone(spec); m.fit(X_train, y_train)
preds = m.predict(X_test)
indiv[name] = accuracy_score(y_test, preds)
mv_preds.append(preds)
mv = stats.mode(np.column_stack(mv_preds), axis=1).mode.flatten()
print('MV:', accuracy_score(y_test, mv))
print(indiv)| model | accuracy |
|---|---|
| LogReg (alone) | 0.9722 |
| RF (alone) | 0.9611 |
| KNN (alone) | 0.9667 |
| Majority vote | 0.9806 |
| Stack (proba OOF) | 0.9861 |
Comparison
Comparison matrix
From Milestone 3 — evaluate and compare: refill the accuracy column from what you know. The rest of the table is as it appeared.
| model | accuracy |
|---|---|
| LogReg (alone) | 0.9722 |
| RF (alone) | 0.9611 |
| KNN (alone) | 0.9667 |
| Majority vote | 0.9806 |
| Stack (proba OOF) | 0.9861 |
Concept
Out loud, slides closed: explain (1) why OOF predictions prevent meta-learner overfitting, (2) the variance-reduction formula and what rho represents, and (3) one trade-off between stacking and blending.
Stretch (homework): implement blending as an alternative and compare test accuracy; prove the variance formula for averaging M models with equal sigma^2 and zero correlation; try adding a 4th base model (GradientBoosting) and observe the accuracy change.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Why ensemble at all? · Stacking: OOF meta-learner · Blending: hold-out meta-train · Diversity & heterogeneous models · Your turn: build it. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
predict_proba meta-features (0.9861 vs 0.9722 single)| idea | the one thing to remember |
|---|---|
| OOF stacking | predict on held-out fold — never on your own training rows |
| blending | faster (1 fit per base), but meta sees only 20% of labels |
| diversity | rho near 0 = full 1/M variance reduction; rho near 1 = no gain |
| meta features | use predict_proba vectors, not hard labels |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.