Lesson 67: Stacking & Blending Ensembles

USAAIO Lesson 67, from Phase 3. It covers stacking, using a 5-fold out-of-fold meta-learner, and blending, using a 20% hold-out, then explains why diversity beats homogeneity and how to combine heterogeneous models such as logistic regression, a random forest, and kNN. It closes with the analytic variance-reduction proof for averaging. All the numbers were verified with sklearn 1.x and numpy 2.2.6. The lesson runs to 28 slides.

Subject: Machine Learning · 52 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Stacking & Blending Ensembles

Title

USAAIO · Lesson 67 · Phase 3

Move past single-model limits: stack diverse learners with out-of-fold predictions, blend on a hold-out set, and prove analytically why diversity reduces variance.

2. By the end of this lesson you can

Objectives

  1. Implement stacking from scratch with 5-fold OOF and a logistic-regression meta-learner
  2. Implement blending with a 20% hold-out set and explain its data efficiency trade-off
  3. Explain why diversity — not sheer quantity — drives ensemble gains
  4. Combine heterogeneous models (LogReg + RF + KNN) into a single stacked classifier
  5. Prove the variance-reduction formula for averaging M models with pairwise correlation rho

3. What survived from Mixed Precision, Loss Scaling & Gradient Accumulation?

Warm-up

Discussion prompt

Before we open Lesson 67: Stacking & Blending Ensembles: without looking back, what was the main idea of Mixed Precision, Loss Scaling & Gradient Accumulation, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

Build an AMP training loop verified against full-precision.

4. Why ensemble at all?

Section

Part 1 of 4

5. The bias-variance argument for ensembles

Concept

A single model with variance sigma^2 can be replaced by the average of M independent models with the same variance — but the average's variance shrinks to sigma^2 / M (Lesson 32 callback).

\[ \mathrm{Var}\!\left(\bar{f}\right) = \rho\,\sigma^2 + \frac{1-\rho}{M}\,\sigma^2 \]

rho is the pairwise correlation of model errors. At rho = 0 (perfectly uncorrelated), Var shrinks to sigma^2/M. At rho = 1 (identical models), nothing helps.

6. Break it if you can: The bias-variance argument for ensembles

Counterexample

Discussion prompt

A single model with variance sigma^2 can be replaced by the average of M independent models with the same variance — but the average's variance shrinks to sigma^2 / M (Lesson 32 callback).

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

7. Variance vs correlation trade-off

Concept

rhoM=5, sigma2=0.10ratio vs single
0.00.02000.20x
0.30.04400.44x
0.50.06000.60x
0.90.09200.92x
1.00.10001.00x (no gain)

Diverse models (low rho) give the biggest variance reduction. This is why stacking heterogeneous base learners outperforms stacking five copies of the same model.

8. Fill in: M=5, sigma2=0.10 for Variance vs correlation trade-off

Comparison

Comparison matrix

From Variance vs correlation trade-off: refill the M=5, sigma2=0.10 column from what you know. The rest of the table is as it appeared.

rhoM=5, sigma2=0.10ratio vs single
0.00.02000.20x
0.30.04400.44x
0.50.06000.60x
0.90.09200.92x
1.00.10001.00x (no gain)

9. Stacking: OOF meta-learner

Section

Part 2 of 4

10. Stacking architecture

Concept

Layer 1 — base models (diverse). Layer 2 — meta-learner trained on the base models' predictions. The trick: base models must predict on data they have not seen, or the meta-learner overfits.

  1. Split train into K folds
  2. For each fold: fit base models on the other K-1 folds, predict the held-out fold (out-of-fold, OOF)
  3. OOF predictions cover all training samples without leakage
  4. Train meta-learner on OOF predictions
  5. At test time: refit base models on all training data, then combine with the meta-learner

11. Guess the shape of the answer: Stacking from scratch — 5-fold OOF

Estimation

Predict first

Load load_digits (1797 samples, 10 classes). Base models: LogReg, RF, KNN. Meta-learner: LogReg on concatenated predict_proba OOF features.

Commit before you compute: what does Stacking from scratch — 5-fold OOF come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Stacking accuracy: 0.9861 vs best single (LogReg) 0.9722

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The meta-learner sees calibrated probability vectors from three diverse models.

12. Stacking from scratch — 5-fold OOF

Worked example

Load load_digits (1797 samples, 10 classes). Base models: LogReg, RF, KNN. Meta-learner: LogReg on concatenated predict_proba OOF features.

import numpy as np
from sklearn.datasets import load_digits
from sklearn.model_selection import StratifiedKFold, train_test_split
from sklearn.linear_model import LogisticRegression
from sklearn.ensemble import RandomForestClassifier
from sklearn.neighbors import KNeighborsClassifier
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.base import clone
from sklearn.metrics import accuracy_score

X, y = load_digits(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42, stratify=y)

base_specs = [
    ('LR', Pipeline([('sc', StandardScaler()),
                     ('clf', LogisticRegression(max_iter=2000, random_state=42))])),
    ('RF', RandomForestClassifier(n_estimators=100, random_state=42)),
    ('KNN', Pipeline([('sc', StandardScaler()),
                      ('clf', KNeighborsClassifier(n_neighbors=3))])),
]
n_cls, n_base = 10, 3
skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)

oof_proba = np.zeros((len(X_train), n_base * n_cls))
test_proba = np.zeros((len(X_test),  n_base * n_cls))

for j, (name, spec) in enumerate(base_specs):
    fold_tp = np.zeros((len(X_test), n_cls, 5))
    for fi, (tr, val) in enumerate(skf.split(X_train, y_train)):
        m = clone(spec)
        m.fit(X_train[tr], y_train[tr])
        oof_proba[val, j*n_cls:(j+1)*n_cls] = m.predict_proba(X_train[val])
        fold_tp[:, :, fi] = m.predict_proba(X_test)
    test_proba[:, j*n_cls:(j+1)*n_cls] = fold_tp.mean(axis=2)

meta = LogisticRegression(max_iter=3000, C=0.1, random_state=42)
meta.fit(oof_proba, y_train)
print(accuracy_score(y_test, meta.predict(test_proba)))

Stacking accuracy: 0.9861 vs best single (LogReg) 0.9722

Why: The meta-learner sees calibrated probability vectors from three diverse models. It learns that when KNN and RF agree but LogReg is uncertain, it can safely boost confidence — an information source a single model cannot exploit.

modelaccuracy
LogReg (alone)0.9722
RF (alone)0.9611
KNN (alone)0.9667
Majority vote0.9806
Stacking (proba OOF)0.9861

13. What each one costs: Stacking from scratch — 5-fold OOF

Trade off

Comparison matrix

From Stacking from scratch — 5-fold OOF: every row here is a choice with a cost. Fill the accuracy column, then say which row you would actually pick and what you give up for it.

modelaccuracy
LogReg (alone)0.9722
RF (alone)0.9611
KNN (alone)0.9667
Majority vote0.9806
Stacking (proba OOF)0.9861

14. 5-fold split sizes

Concept

Each OOF fold predicts exactly once for every training sample. With n_train = 1437 and K = 5, the folds are:

foldtrain rowsval rows (OOF)
11149288
21149288
31150287
41150287
51150287

The 1437 OOF predictions are then used as training data for the meta-learner — no sample was predicted by a model it was trained on.

15. Watch it run: 5-fold split sizes

Pattern

Step through it

Step through 5-fold split sizes one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: fold is 1
  2. Step 2: fold is 2
  3. Step 3: fold is 3
  4. Step 4: fold is 4
  5. Step 5: fold is 5

16. Something is wrong here: training meta-learner on in-fold predictions

Anomaly

Predict first

A student writes this, and it looks reasonable:

Fit each base model on the full training set. Use those in-sample predictions as meta-features.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Each base model has seen every training sample.

Use out-of-fold (OOF) predictions: each training sample is predicted by a base model that was never trained on that sample.

Why: Each base model has seen every training sample. Its in-sample predictions are nearly perfect for well-fit models — the meta-learner trains on overfit signals, then sees test data where those signals are noisy. The meta-learner overfits to the base models' memorized outputs.

17. Trap: training meta-learner on in-fold predictions

Trap

The trap

Fit each base model on the full training set. Use those in-sample predictions as meta-features.

base.fit(X_train, y_train) --> meta.fit(base.predict_proba(X_train), y_train)

Why: Each base model has seen every training sample. Its in-sample predictions are nearly perfect for well-fit models — the meta-learner trains on overfit signals, then sees test data where those signals are noisy. The meta-learner overfits to the base models' memorized outputs.

The fix

Use out-of-fold (OOF) predictions: each training sample is predicted by a base model that was never trained on that sample.

5-fold CV --> oof_proba[val] = base_fold.predict_proba(X_train[val]) --> meta.fit(oof_proba, y_train)

Why: OOF predictions are honest — they simulate test-time uncertainty. The meta-learner trains on signals that generalize, not memorized perfect probabilities.

18. Break it on purpose: training meta-learner on in-fold predictions

Break the constraint

Discussion prompt

The rule this trap just fixed:

OOF predictions are honest — they simulate test-time uncertainty. The meta-learner trains on signals that generalize, not memorized perfect probabilities.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

Each base model has seen every training sample. Its in-sample predictions are nearly perfect for well-fit models — the meta-learner trains on overfit signals, then sees test data where those signals are noisy. The meta-learner overfits to the base models' memorized outputs.

19. Blending: hold-out meta-train

Section

Part 3 of 4

20. Blending vs stacking

Concept

propertystacking (K-fold OOF)blending (hold-out)
meta-train sizeall n_train samples~20% of n_train
data efficiencyhighlower (wastes 80% for base)
leakage risknone (OOF)none (separate split)
computationK x n_base model fits1 x n_base model fits
typical usecompetition accuracyquick baseline

Blending is faster (train each base model once) but the meta-learner sees only ~20% of the data. Stacking uses all training data for the meta-learner, at the cost of K times more base-model fits.

21. By analogy: Blending vs stacking

Analogy

Discussion prompt

Explain Blending vs stacking by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Blending is faster (train each base model once) but the meta-learner sees only ~20% of the data. Stacking uses all training data for the meta-learner, at the cost of K times more base-model fits.

22. Guess the shape of the answer: Blending — hold-out 20% for meta

Estimation

Predict first

Reserve 20% of training data for meta-learner training. Fit base models on the remaining 80%, then predict the 20% hold-out and the test set.

Commit before you compute: what does Blending — hold-out 20% for meta come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Blending accuracy: 0.9861 (equal to stacking on this dataset)

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. On load_digits the meta-learner converges with fewer samples.

23. Blending — hold-out 20% for meta

Worked example

Reserve 20% of training data for meta-learner training. Fit base models on the remaining 80%, then predict the 20% hold-out and the test set.

from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
from sklearn.base import clone
from sklearn.metrics import accuracy_score
import numpy as np

X_bl, X_bval, y_bl, y_bval = train_test_split(
    X_train, y_train, test_size=0.2, random_state=42, stratify=y_train)
print(f'Base train: {len(y_bl)}, Meta train: {len(y_bval)}')

bl_val_p, bl_test_p = [], []
for name, spec in base_specs:
    m = clone(spec); m.fit(X_bl, y_bl)
    bl_val_p.append(m.predict_proba(X_bval))
    bl_test_p.append(m.predict_proba(X_test))

meta_b = LogisticRegression(max_iter=2000, C=0.1, random_state=42)
meta_b.fit(np.hstack(bl_val_p), y_bval)
print(accuracy_score(y_test, meta_b.predict(np.hstack(bl_test_p))))

Blending accuracy: 0.9861 (equal to stacking on this dataset)

Why: On load_digits the meta-learner converges with fewer samples. On noisier or higher-dimensional problems, stacking's larger meta-training set typically wins.

methodmeta train naccuracy
blending (20% hold-out)2880.9861
stacking (5-fold OOF)14370.9861
best single model—0.9722

24. Work backwards from the answer: Blending — hold-out 20% for meta

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Blending accuracy: 0.9861 (equal to stacking on this dataset)

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Reserve 20% of training data for meta-learner training. Fit base models on the remaining 80%, then predict the 20% hold-out and the test set.

25. Something is wrong here: evaluating blending on the blend-val set

Anomaly

Predict first

A student writes this, and it looks reasonable:

After fitting the meta-learner on the 20% hold-out, report accuracy on that same 20% as the ensemble's performance.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: The meta-learner was trained on that hold-out set — reporting accuracy there is optimistic in-sample error.

Evaluate the full stack (base models + meta-learner) on the held-out test set that neither the base models nor the meta-learner touched during training.

Why: The meta-learner was trained on that hold-out set — reporting accuracy there is optimistic in-sample error. You've measured how well the meta-learner memorized its own training set, not how it generalizes.

26. Trap: evaluating blending on the blend-val set

Trap

The trap

After fitting the meta-learner on the 20% hold-out, report accuracy on that same 20% as the ensemble's performance.

meta.fit(bl_val_feat, y_bval) --> score on bl_val

Why: The meta-learner was trained on that hold-out set — reporting accuracy there is optimistic in-sample error. You've measured how well the meta-learner memorized its own training set, not how it generalizes.

The fix

Evaluate the full stack (base models + meta-learner) on the held-out test set that neither the base models nor the meta-learner touched during training.

test_feat = [m.predict_proba(X_test) for m in base_models] --> meta.predict(test_feat)

Why: The test set is neutral to both layers. Reporting accuracy there gives an unbiased estimate of ensemble generalization.

27. Which of these survive contact with Lesson 67: Stacking & Blending Ensembles?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Diverse models (low rho) give the biggest variance reduction. This is why stacking heterogeneous base learners outperforms stacking five copies of the same model.; Each OOF fold predicts exactly once for every training sample. With n_train = 1437 and K = 5, the folds are:; Two models are diverse when they make different errors on the same samples. The meta-learner can then correct one model's mistakes using the other's confidence.
Breaks
Fit each base model on the full training set. Use those in-sample predictions as meta-features.; After fitting the meta-learner on the 20% hold-out, report accuracy on that same 20% as the ensemble's performance.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 67: Stacking & Blending Ensembles puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

28. Diversity & heterogeneous models

Section

Part 4 of 4

29. What makes base models diverse?

Concept

Two models are diverse when they make different errors on the same samples. The meta-learner can then correct one model's mistakes using the other's confidence.

Stacking five identical LogReg models gives rho close to 1 — the variance formula predicts almost zero gain (Lesson 67's analytic table).

30. By analogy: What makes base models diverse?

Analogy

Discussion prompt

Explain What makes base models diverse? by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Two models are diverse when they make different errors on the same samples. The meta-learner can then correct one model's mistakes using the other's confidence.

31. AutoML: automated stacking at scale

Concept

AutoML systems (H2O AutoML, AutoSklearn) automate the stacking process: search over model families, tune hyperparameters, and assemble a stacking ensemble — often producing near-Pareto-optimal ensembles with no manual tuning.

systembase modelsmeta strategy
H2O AutoMLGLM, GBM, RF, DNN, XGBoostStacked Ensemble (OOF)
AutoSklearnsklearn models + preprocessingEnsemble selection from evaluated configs
manual stackingyour choiceyour meta-learner

Understanding OOF stacking from scratch (this lesson) is prerequisite knowledge for interpreting what AutoML does and diagnosing when it fails.

32. Fill in: meta strategy for AutoML: automated stacking at scale

Comparison

Comparison matrix

From AutoML: automated stacking at scale: refill the meta strategy column from what you know. The rest of the table is as it appeared.

systembase modelsmeta strategy
H2O AutoMLGLM, GBM, RF, DNN, XGBoostStacked Ensemble (OOF)
AutoSklearnsklearn models + preprocessingEnsemble selection from evaluated configs
manual stackingyour choiceyour meta-learner

33. Rebuild the recipe: The stacking/blending recipe

Ranking

Put in order

These are the steps of The stacking/blending recipe, scrambled. Put them back in order before the next slide shows you.

  1. Choose diverse base models — different biases, not five copies of one
  2. Stacking: 5-fold stratified CV → OOF predict_proba → concatenate → meta-learner
  3. Blending: split train 80/20 → base on 80%, meta on 20% predictions
  4. Meta-learner: LogReg (C=0.1) on probability vectors; never train meta on in-fold preds
  5. Test-time: refit each base on all train (stacking) or just the 80% (blending), then meta
  6. Evaluate on a test set untouched by both layers

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

34. The stacking/blending recipe

Pattern

  1. Choose diverse base models — different biases, not five copies of one
  2. Stacking: 5-fold stratified CV → OOF predict_proba → concatenate → meta-learner
  3. Blending: split train 80/20 → base on 80%, meta on 20% predictions
  4. Meta-learner: LogReg (C=0.1) on probability vectors; never train meta on in-fold preds
  5. Test-time: refit each base on all train (stacking) or just the 80% (blending), then meta
  6. Evaluate on a test set untouched by both layers

35. Where does it stop working: The stacking/blending recipe

Edge cases

Discussion prompt

The stacking/blending recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Choose diverse base models — different biases, not five copies of one
  2. Stacking: 5-fold stratified CV → OOF predict_proba → concatenate → meta-learner
  3. Blending: split train 80/20 → base on 80%, meta on 20% predictions
  4. Meta-learner: LogReg (C=0.1) on probability vectors; never train meta on in-fold preds
  5. Test-time: refit each base on all train (stacking) or just the 80% (blending), then meta
  6. Evaluate on a test set untouched by both layers

36. Rule out three: Check yourself — OOF purpose

Elimination

Eliminate the wrong options

Why must meta-learner training features come from out-of-fold (OOF) predictions rather than in-sample predictions?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. In-sample predictions are overfit; OOF predictions simulate honest test-time uncertainty
  • B. OOF predictions are more accurate than in-sample predictions
  • C. In-sample predictions require more memory than OOF predictions
  • D. The meta-learner cannot accept probability vectors, only labels

Survives elimination: A

Why: A well-fit base model achieves near-perfect accuracy on its own training samples — these in-sample probabilities are overfit signals. The meta-learner would learn to trust overconfident probabilities that don't appear at test time. OOF predictions come from a held-out fold, accurately reflecting how the model behaves on unseen data.

37. Check yourself — OOF purpose

Check

Identify the key step.

Check your understanding

Why must meta-learner training features come from out-of-fold (OOF) predictions rather than in-sample predictions?

  • A. In-sample predictions are overfit; OOF predictions simulate honest test-time uncertainty (correct)
  • B. OOF predictions are more accurate than in-sample predictions
  • C. In-sample predictions require more memory than OOF predictions
  • D. The meta-learner cannot accept probability vectors, only labels

Answer: A

Why: A well-fit base model achieves near-perfect accuracy on its own training samples — these in-sample probabilities are overfit signals. The meta-learner would learn to trust overconfident probabilities that don't appear at test time. OOF predictions come from a held-out fold, accurately reflecting how the model behaves on unseen data.

Why B tempts people
OOF predictions are typically less accurate (the base model hasn't seen those samples) — the point is that they're honest, not more accurate.
Why C tempts people
Memory usage is not the reason; both formats have the same shape. The reason is statistical, not computational.
Why D tempts people
Probability vectors (predict_proba output) are the preferred meta-features — richer signal than hard labels.

38. Answer it before you see the options: Check yourself — variance reduction

Prediction

Predict first

M=5 diverse models each have variance sigma^2=0.10 and pairwise error correlation rho=0.3. What is Var(average)?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: 0.0440

Why: Var(avg) = rhosigma^2 + (1-rho)sigma^2/M = 0.30.10 + 0.70.10/5 = 0.030 + 0.014 = 0.0440. This is 2.27x lower than a single model's 0.10.

39. Check yourself — variance reduction

Check

Apply the formula.

Check your understanding

M=5 diverse models each have variance sigma^2=0.10 and pairwise error correlation rho=0.3. What is Var(average)?

  • A. 0.0440 (correct)
  • B. 0.0200
  • C. 0.0300
  • D. 0.0500

Answer: A

Why: Var(avg) = rhosigma^2 + (1-rho)sigma^2/M = 0.30.10 + 0.70.10/5 = 0.030 + 0.014 = 0.0440. This is 2.27x lower than a single model's 0.10.

Why B tempts people
0.0200 is the rho=0 (uncorrelated) result. With rho=0.3, the correlated term adds 0.3*0.10=0.030.
Why C tempts people
0.0300 is just the correlated term rhosigma^2 alone — the formula adds (1-rho)sigma^2/M on top.
Why D tempts people
0.0500 = sigma^2/M (the fully uncorrelated formula) applied at rho=0.5, not rho=0.3.

40. Rule out three: Check yourself — stacking vs blending

Elimination

Eliminate the wrong options

Compared with 5-fold stacking, blending with a 20% hold-out is:

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Faster to train (1 pass per base model) but gives the meta-learner fewer training samples
  • B. More data-efficient because all training samples are used exactly once
  • C. Identical to stacking when K=5 and hold-out fraction is 0.2
  • D. Always less accurate because it introduces label leakage

Survives elimination: A

Why: Blending trains each base model only once (on 80% of the data), so it requires 1/5 the compute of 5-fold stacking. But the meta-learner trains on only 20% of the training labels (~288 samples here vs 1437 for stacking). On small datasets this is a meaningful accuracy gap.

41. Check yourself — stacking vs blending

Check

Compare the two approaches.

Check your understanding

Compared with 5-fold stacking, blending with a 20% hold-out is:

  • A. Faster to train (1 pass per base model) but gives the meta-learner fewer training samples (correct)
  • B. More data-efficient because all training samples are used exactly once
  • C. Identical to stacking when K=5 and hold-out fraction is 0.2
  • D. Always less accurate because it introduces label leakage

Answer: A

Why: Blending trains each base model only once (on 80% of the data), so it requires 1/5 the compute of 5-fold stacking. But the meta-learner trains on only 20% of the training labels (~288 samples here vs 1437 for stacking). On small datasets this is a meaningful accuracy gap.

Why B tempts people
Blending is less data-efficient: 80% goes only to base models and is never used for meta-learning. Stacking uses every training sample for both tasks (through K folds).
Why C tempts people
K=5 and hold-out=0.2 happen to give the same fold size, but the procedures differ: stacking fits each base model 5 times and aggregates; blending fits once.
Why D tempts people
Blending introduces no label leakage as long as the hold-out set is strictly separate from the base-model training set.

42. Your turn: build it

Section

Project

43. Project: stacking ensemble from scratch

Concept

Implement a stacking ensemble on load_digits with LogReg, RF, and KNN as base models and LogReg as the meta-learner. Prove it beats the best individual model.

#milestonekey call
15-fold OOF proba featuresStratifiedKFold + clone + predict_proba
2meta-learner trainingLogReg(C=0.1).fit(oof_proba, y_train)
3test-time predictiontest_proba = avg of fold predictions
4accuracy tableaccuracy_score vs singles + vote

Build rules: never let a base model predict its own training fold; use predict_proba not predict for meta features; evaluate only on the untouched test set.

44. Break it if you can: Project: stacking ensemble from scratch

Counterexample

Discussion prompt

Implement a stacking ensemble on load_digits with LogReg, RF, and KNN as base models and LogReg as the meta-learner. Prove it beats the best individual model.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Build rules: never let a base model predict its own training fold; use predict_proba not predict for meta features; evaluate only on the untouched test set.

45. Milestone 1 — OOF proba features

Worked example

Your turn: set up 5-fold CV and collect OOF probability vectors for all three base models.

Hint: oof_proba = np.zeros((n_train, n_base * n_cls)) — each model contributes 10 probability columns. Use clone(spec) inside the fold loop so each fold gets a fresh model.

import numpy as np
from sklearn.datasets import load_digits
from sklearn.model_selection import StratifiedKFold, train_test_split
from sklearn.linear_model import LogisticRegression
from sklearn.ensemble import RandomForestClassifier
from sklearn.neighbors import KNeighborsClassifier
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.base import clone

X, y = load_digits(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42, stratify=y)

base_specs = [
    ('LR', Pipeline([('sc', StandardScaler()),
                     ('clf', LogisticRegression(max_iter=2000, random_state=42))])),
    ('RF', RandomForestClassifier(n_estimators=100, random_state=42)),
    ('KNN', Pipeline([('sc', StandardScaler()),
                      ('clf', KNeighborsClassifier(n_neighbors=3))])),
]
n_cls, n_base = 10, 3
skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
oof_proba = np.zeros((len(X_train), n_base * n_cls))
print(oof_proba.shape)  # (1437, 30)
arrayshapemeaning
oof_proba(1437, 30)10 proba cols per base model x 3 models
block j colsj10 : (j+1)10proba for base model j
fold val rows288 or 287OOF rows filled each fold

46. Milestone 2 — fill OOF and train meta

Worked example

Your turn: fill oof_proba across folds, then train the meta-learner.

Hint: inner loop for fi, (tr, val) in enumerate(skf.split(...)) — fit clone(spec) on X_train[tr], predict X_train[val], store in oof_proba[val, j*10:(j+1)*10].

test_proba = np.zeros((len(X_test), n_base * n_cls))

for j, (name, spec) in enumerate(base_specs):
    fold_tp = np.zeros((len(X_test), n_cls, 5))
    for fi, (tr, val) in enumerate(skf.split(X_train, y_train)):
        m = clone(spec)
        m.fit(X_train[tr], y_train[tr])
        oof_proba[val, j*n_cls:(j+1)*n_cls] = m.predict_proba(X_train[val])
        fold_tp[:, :, fi] = m.predict_proba(X_test)
    test_proba[:, j*n_cls:(j+1)*n_cls] = fold_tp.mean(axis=2)

meta = LogisticRegression(max_iter=3000, C=0.1, random_state=42)
meta.fit(oof_proba, y_train)
print(meta.coef_.shape)  # (10, 30) - 10-class meta-learner
stepoutput shape
oof_proba after loop(1437, 30)
test_proba (avg over folds)(360, 30)
meta.coef_(10, 30)

47. What each one costs: Milestone 2 — fill OOF and train meta

Trade off

Comparison matrix

From Milestone 2 — fill OOF and train meta: every row here is a choice with a cost. Fill the output shape column, then say which row you would actually pick and what you give up for it.

stepoutput shape
oof_proba after loop(1437, 30)
test_proba (avg over folds)(360, 30)
meta.coef_(10, 30)

48. Milestone 3 — evaluate and compare

Worked example

Your turn: predict the test set with the stacked meta-learner and compare against individual models and majority vote.

Hint: meta.predict(test_proba) gives the ensemble prediction. Compare with each base model trained on full X_train.

from sklearn.metrics import accuracy_score
from scipy import stats

stack_acc = accuracy_score(y_test, meta.predict(test_proba))
print(f'Stack: {stack_acc:.4f}')

indiv, mv_preds = {}, []
for name, spec in base_specs:
    m = clone(spec); m.fit(X_train, y_train)
    preds = m.predict(X_test)
    indiv[name] = accuracy_score(y_test, preds)
    mv_preds.append(preds)
mv = stats.mode(np.column_stack(mv_preds), axis=1).mode.flatten()
print('MV:', accuracy_score(y_test, mv))
print(indiv)
modelaccuracy
LogReg (alone)0.9722
RF (alone)0.9611
KNN (alone)0.9667
Majority vote0.9806
Stack (proba OOF)0.9861

49. Fill in: accuracy for Milestone 3 — evaluate and compare

Comparison

Comparison matrix

From Milestone 3 — evaluate and compare: refill the accuracy column from what you know. The rest of the table is as it appeared.

modelaccuracy
LogReg (alone)0.9722
RF (alone)0.9611
KNN (alone)0.9667
Majority vote0.9806
Stack (proba OOF)0.9861

50. Show it off

Concept

Out loud, slides closed: explain (1) why OOF predictions prevent meta-learner overfitting, (2) the variance-reduction formula and what rho represents, and (3) one trade-off between stacking and blending.

Stretch (homework): implement blending as an alternative and compare test accuracy; prove the variance formula for averaging M models with equal sigma^2 and zero correlation; try adding a 4th base model (GradientBoosting) and observe the accuracy change.

51. Connect it up: Lesson 67: Stacking & Blending Ensembles

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Why ensemble at all? · Stacking: OOF meta-learner · Blending: hold-out meta-train · Diversity & heterogeneous models · Your turn: build it. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

52. What you can do now

Recap

ideathe one thing to remember
OOF stackingpredict on held-out fold — never on your own training rows
blendingfaster (1 fit per base), but meta sees only 20% of labels
diversityrho near 0 = full 1/M variance reduction; rho near 1 = no gain
meta featuresuse predict_proba vectors, not hard labels

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 67 — Stacking, Blending, Ensembles, Diversity — Barron · USAAIO Round 2 Preparation, 2026
  2. Stacking (5-fold OOF proba meta) 0.9861, blending 0.9861, best single 0.9722 on load_digits — sklearn + numpy 2.2.6, real execution, June 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108