Lesson 71: Kaggle Competition Pipeline

USAAIO Lesson 71, from Phase 4, on the end-to-end competition workflow. It covers exploratory data analysis - distributions, correlations, missing values, and outliers - then interaction and polynomial feature engineering, model selection across the linear, tree, and neural families with 5-fold cross-validation, stacking ensembles, and using cross-validation scores to predict where you will land on the leaderboard. It was verified on the sklearn diabetes dataset with numpy 2.2.6 and scikit-learn. The lesson runs to 30 slides.

Subject: Machine Learning · 58 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Kaggle Competition Pipeline

Title

USAAIO · Lesson 71 · Phase 4

End-to-end: EDA → feature engineering → model selection → ensemble → CV-based leaderboard estimate. The full workflow in one session.

2. By the end of this lesson you can

Objectives

  1. Diagnose distributions, correlations, missing values, and outliers in a competition dataset
  2. Engineer interaction features and polynomial features and measure whether they help
  3. Compare linear, tree-based, and neural models using 5-fold CV R²
  4. Build a stacking ensemble of diverse base learners with a meta-learner
  5. Use cross-validation to estimate your holdout (leaderboard) score before submission

3. EDA: what the data is telling you

Section

Part 1 of 4

4. Why EDA before anything else

Concept

Competitions are lost to silent data problems — a skewed target, a near-duplicate feature, or outliers that drag every model toward one region of the space.

A 30-minute EDA on the diabetes dataset surfaces everything you need: 442 samples, 10 features, no missing values, and 3 BMI outliers by IQR.

5. Break it if you can: Why EDA before anything else

Counterexample

Discussion prompt

Competitions are lost to silent data problems — a skewed target, a near-duplicate feature, or outliers that drag every model toward one region of the space.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

A 30-minute EDA on the diabetes dataset surfaces everything you need: 442 samples, 10 features, no missing values, and 3 BMI outliers by IQR.

6. Guess the shape of the answer: EDA on the diabetes dataset

Estimation

Predict first

Load load_diabetes, print shape and target summary, check for missing values, compute feature–target correlations, and flag BMI outliers.

Commit before you compute: what does EDA on the diabetes dataset come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Shape (442, 10); target 25–346, mean 152.1; 0 missing; top features bmi r=0.586, s5 r=0.566, bp r=0.441; 3 BMI outliers

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. No imputation needed here; bmi and s5 dominate the correlation ranking — they are the first candidates for feature engineering.

7. EDA on the diabetes dataset

Worked example

Load load_diabetes, print shape and target summary, check for missing values, compute feature–target correlations, and flag BMI outliers.

from sklearn.datasets import load_diabetes
import numpy as np
data = load_diabetes()
X, y, names = data.data, data.target, list(data.feature_names)
print(X.shape, y.min(), y.max(), y.mean().round(1))
print('missing:', np.isnan(X).sum())
corrs = [np.corrcoef(X[:,i], y)[0,1] for i in range(10)]
top3 = sorted(range(10), key=lambda i: abs(corrs[i]), reverse=True)[:3]
for i in top3: print(f'{names[i]}: r={corrs[i]:.3f}')
bmi = X[:, names.index('bmi')]
q1,q3 = np.percentile(bmi,[25,75]); iqr=q3-q1
print('bmi outliers:', ((bmi<q1-1.5*iqr)|(bmi>q3+1.5*iqr)).sum())

Shape (442, 10); target 25–346, mean 152.1; 0 missing; top features bmi r=0.586, s5 r=0.566, bp r=0.441; 3 BMI outliers

Why: No imputation needed here; bmi and s5 dominate the correlation ranking — they are the first candidates for feature engineering.

featurer with targetnote
bmi0.586strongest predictor
s50.566log serum triglycerides
bp0.441blood pressure
s3−0.395negative correlation
other 6<0.22weaker signal

8. Fill in: r with target for EDA on the diabetes dataset

Comparison

Comparison matrix

From EDA on the diabetes dataset: refill the r with target column from what you know. The rest of the table is as it appeared.

featurer with targetnote
bmi0.586strongest predictor
s50.566log serum triglycerides
bp0.441blood pressure
s3−0.395negative correlation
other 6<0.22weaker signal

9. Outlier strategies: drop vs. clip vs. encode

Concept

strategywhen to userisk
drop rowoutlier is a data-entry errorloses real examples
clip to fence (IQR±1.5)genuine but extreme; tree models don't need itdistorts real extremes
log / power transformright-skewed target or featurechanges scale, affects interpretability
encode as binary flagmissingness is informative (MNAR)adds a feature; safe default

For diabetes the 3 BMI outliers are genuine patients — leave them in and let tree-based models handle them naturally.

10. What each one costs: Outlier strategies: drop vs. clip vs. encode

Trade off

Comparison matrix

From Outlier strategies: drop vs. clip vs. encode: every row here is a choice with a cost. Fill the when to use column, then say which row you would actually pick and what you give up for it.

strategywhen to userisk
drop rowoutlier is a data-entry errorloses real examples
clip to fence (IQR±1.5)genuine but extreme; tree models don't need itdistorts real extremes
log / power transformright-skewed target or featurechanges scale, affects interpretability
encode as binary flagmissingness is informative (MNAR)adds a feature; safe default

11. Feature engineering: building signal

Section

Part 2 of 4

12. Three feature engineering primitives

Concept

Rule of thumb: add a feature only if CV score improves. A raw feature matrix with 10 columns can explode to 65 degree-2 polynomial terms — every one adds variance and needs regularization.

13. By analogy: Three feature engineering primitives

Analogy

Discussion prompt

Explain Three feature engineering primitives by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Rule of thumb: add a feature only if CV score improves. A raw feature matrix with 10 columns can explode to 65 degree-2 polynomial terms — every one adds variance and needs regularization.

14. What has to be given first: Interaction feature: does bmi × s5 help?

Missing information

Discussion prompt

Hypothesis: high BMI combined with high triglycerides (s5) predicts worse outcomes synergistically. Add bmi × s5 and re-run Ridge CV.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

bmi and s5 are already in the model; their product adds colinearity without new signal for a linear model. The feature is useless here — and CV told us that without touching the test set.

15. Interaction feature: does bmi × s5 help?

Worked example

Hypothesis: high BMI combined with high triglycerides (s5) predicts worse outcomes synergistically. Add bmi × s5 and re-run Ridge CV.

from sklearn.datasets import load_diabetes
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import Ridge
from sklearn.model_selection import cross_val_score, KFold
import numpy as np
data = load_diabetes(); X, y = data.data, data.target
names = list(data.feature_names)
bmi_s5 = X[:, names.index('bmi')] * X[:, names.index('s5')]
X_aug = np.column_stack([X, bmi_s5])
sc = StandardScaler(); cv = KFold(5, shuffle=True, random_state=42)
base = cross_val_score(Ridge(1.0), sc.fit_transform(X), y, cv=cv, scoring='r2')
aug  = cross_val_score(Ridge(1.0), sc.fit_transform(X_aug), y, cv=cv, scoring='r2')
print('base:', base.mean().round(3), 'aug:', aug.mean().round(3))

base R²=0.479; aug R²=0.479 — the interaction adds nothing for Ridge on this dataset

Why: bmi and s5 are already in the model; their product adds colinearity without new signal for a linear model. The feature is useless here — and CV told us that without touching the test set.

modelfeaturesCV R²
Ridgeraw 100.479
Ridgeraw 10 + bmi×s50.479
verdict—interaction ≈ no gain

16. Work backwards from the answer: Interaction feature: does bmi × s5 help?

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

base R²=0.479; aug R²=0.479 — the interaction adds nothing for Ridge on this dataset

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Hypothesis: high BMI combined with high triglycerides (s5) predicts worse outcomes synergistically. Add bmi × s5 and re-run Ridge CV.

17. Something is wrong here: evaluating features on the training set

Anomaly

Predict first

A student writes this, and it looks reasonable:

Add bmi × s5, fit a Ridge on all 442 samples, get training R²=0.521 — the feature helped.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Training R² always weakly increases when you add features — the model just interpolates the noise.

Always measure feature value with held-out or CV score, not training score.

Why: Training R² always weakly increases when you add features — the model just interpolates the noise. This is pure overfitting detection failure.

18. Trap: evaluating features on the training set

Trap

The trap

Add bmi × s5, fit a Ridge on all 442 samples, get training R²=0.521 — the feature helped.

Report training R² as evidence the interaction improves the model

Why: Training R² always weakly increases when you add features — the model just interpolates the noise. This is pure overfitting detection failure.

The fix

Always measure feature value with held-out or CV score, not training score.

Run cross_val_score before and after adding the feature; CV R²=0.479 vs 0.479 — feature rejected

Why: CV measures generalization. The feature that looks useful on training data but flat on CV is adding variance without signal — drop it.

19. Break it on purpose: evaluating features on the training set

Break the constraint

Discussion prompt

The rule this trap just fixed:

CV measures generalization. The feature that looks useful on training data but flat on CV is adding variance without signal — drop it.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

Training R² always weakly increases when you add features — the model just interpolates the noise. This is pure overfitting detection failure.

20. Model selection: linear, tree, neural

Section

Part 3 of 4

21. Three model families — their trade-offs

Concept

familyexamplestrengthweakness
linearRidge, Lassofast, interpretable, regularizablecan't model interactions without manual features
tree-basedRF, GradBoostnon-linear, handles outliers/interactionscan overfit small datasets
neuralMLPflexible universal approximatorneeds more data, hyperparameter-sensitive

For 442 samples the tree-based and neural models face a high-variance regime — you need heavy regularization or ensemble averaging to beat Ridge (Lesson 32).

22. Teach it back: Three model families — their trade-offs

Explain it

Discussion prompt

Explain Three model families — their trade-offs to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

For 442 samples the tree-based and neural models face a high-variance regime — you need heavy regularization or ensemble averaging to beat Ridge (Lesson 32).

23. Guess the shape of the answer: 5-fold CV comparison: all three families

Estimation

Predict first

Scale the data, then run 5-fold CV R² for Ridge, DecisionTree (depth 3), RandomForest (50 trees), GradientBoosting, and MLP.

Commit before you compute: what does 5-fold CV comparison: all three families come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Ridge leads at 0.479; GBM second at 0.458; RF at 0.444; DTree at 0.328; MLP at 0.023 — the small dataset punishes MLP hard

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. MLP's 0.023 isn't a bug — it simply lacks enough samples to converge reliably with a 64-unit hidden layer.

24. 5-fold CV comparison: all three families

Worked example

Scale the data, then run 5-fold CV R² for Ridge, DecisionTree (depth 3), RandomForest (50 trees), GradientBoosting, and MLP.

from sklearn.datasets import load_diabetes
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import Ridge
from sklearn.tree import DecisionTreeRegressor
from sklearn.ensemble import RandomForestRegressor, GradientBoostingRegressor
from sklearn.neural_network import MLPRegressor
from sklearn.model_selection import cross_val_score, KFold
data = load_diabetes(); X, y = data.data, data.target
X = StandardScaler().fit_transform(X)
cv = KFold(5, shuffle=True, random_state=42)
for name, m in [
    ('Ridge',  Ridge(1.0)),
    ('DTree',  DecisionTreeRegressor(max_depth=3, random_state=42)),
    ('RF',     RandomForestRegressor(50, max_depth=3, random_state=42)),
    ('GBM',    GradientBoostingRegressor(50, max_depth=2, random_state=42)),
    ('MLP',    MLPRegressor((64,), max_iter=500, random_state=42)),
]:
    s = cross_val_score(m, X, y, cv=cv, scoring='r2')
    print(f'{name}: {s.mean():.3f} ± {s.std():.3f}')

Ridge leads at 0.479; GBM second at 0.458; RF at 0.444; DTree at 0.328; MLP at 0.023 — the small dataset punishes MLP hard

Why: MLP's 0.023 isn't a bug — it simply lacks enough samples to converge reliably with a 64-unit hidden layer. GBM's depth-2 trees generalize better than depth-3 DTree because shallow trees regularize by construction.

modelCV R²CV std
Ridge0.479±0.083
GradBoost (d=2)0.458±0.078
RandomForest (d=3)0.444±0.084
DecisionTree (d=3)0.328±0.077
MLP (64,)0.023±0.101

25. Work backwards from the answer: 5-fold CV comparison: all three families

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Ridge leads at 0.479; GBM second at 0.458; RF at 0.444; DTree at 0.328; MLP at 0.023 — the small dataset punishes MLP hard

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Scale the data, then run 5-fold CV R² for Ridge, DecisionTree (depth 3), RandomForest (50 trees), GradientBoosting, and MLP.

26. Something is wrong here: picking the model with the highest training R²

Anomaly

Predict first

A student writes this, and it looks reasonable:

Fit all five models on the full dataset, pick the one with the highest training R² — that's the best model.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: MLP memorizes the small dataset completely (Lesson 32 bias-variance).

Select by CV score on held-out folds, never training score.

Why: MLP memorizes the small dataset completely (Lesson 32 bias-variance). Its CV R²=0.023 means it predicts near-baseline on new data.

27. Trap: picking the model with the highest training R²

Trap

The trap

Fit all five models on the full dataset, pick the one with the highest training R² — that's the best model.

Select MLP because it achieves training R²≈0.95 on 442 samples

Why: MLP memorizes the small dataset completely (Lesson 32 bias-variance). Its CV R²=0.023 means it predicts near-baseline on new data.

The fix

Select by CV score on held-out folds, never training score.

Select Ridge (CV R²=0.479) — lower training R² but far better generalization than MLP's CV R²=0.023

Why: CV measures expected leaderboard performance. Training R² measures memorization. They diverge most for flexible models on small data.

28. Which of these survive contact with Lesson 71: Kaggle Competition Pipeline?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Competitions are lost to silent data problems — a skewed target, a near-duplicate feature, or outliers that drag every model toward one region of the space.; For diabetes the 3 BMI outliers are genuine patients — leave them in and let tree-based models handle them naturally.; For 442 samples the tree-based and neural models face a high-variance regime — you need heavy regularization or ensemble averaging to beat Ridge (Lesson 32).
Breaks
Add bmi × s5, fit a Ridge on all 442 samples, get training R²=0.521 — the feature helped.; Fit all five models on the full dataset, pick the one with the highest training R² — that's the best model.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 71: Kaggle Competition Pipeline puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

29. Ensemble: stacking diverse models

Section

Part 4 of 4

30. Why stacking works

Concept

Stacking trains a meta-learner to combine out-of-fold predictions from diverse base learners. The meta-learner learns when to trust each base model.

  1. Split training data into K folds
  2. For each fold: train base learners on K-1 folds, predict on the held-out fold
  3. Collect K held-out predictions → level-1 feature matrix (n × B)
  4. Train a meta-learner (usually Ridge) on the level-1 matrix to predict y
  5. At test time: all base learners predict, meta-learner combines

Diversity is essential — stacking Ridge + RF + GBM helps because each model captures different structure. Stacking three copies of Ridge gains nothing.

31. Teach it back: Why stacking works

Explain it

Discussion prompt

Explain Why stacking works to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Diversity is essential — stacking Ridge + RF + GBM helps because each model captures different structure. Stacking three copies of Ridge gains nothing.

32. Guess the shape of the answer: StackingRegressor: Ridge + RF + GBM

Estimation

Predict first

Build a stacking ensemble with Ridge, RandomForest, and GradientBoosting as base learners; use Ridge as the meta-learner. Measure CV and holdout R².

Commit before you compute: what does StackingRegressor: Ridge + RF + GBM come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: CV R²=0.480; holdout R²=0.472 — stacking matches best single model (Ridge at 0.479) and slightly beats GB+RF alone

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. On 442 samples, the base models are correlated enough that stacking gain is modest.

33. StackingRegressor: Ridge + RF + GBM

Worked example

Build a stacking ensemble with Ridge, RandomForest, and GradientBoosting as base learners; use Ridge as the meta-learner. Measure CV and holdout R².

from sklearn.datasets import load_diabetes
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import Ridge
from sklearn.ensemble import RandomForestRegressor, GradientBoostingRegressor, StackingRegressor
from sklearn.model_selection import cross_val_score, KFold, train_test_split
from sklearn.metrics import r2_score
import numpy as np
data = load_diabetes(); X, y = data.data, data.target
X = StandardScaler().fit_transform(X)
base = [('ridge', Ridge(1.0)),
        ('rf',    RandomForestRegressor(50, max_depth=3, random_state=42)),
        ('gb',    GradientBoostingRegressor(50, max_depth=2, random_state=42))]
stack = StackingRegressor(base, final_estimator=Ridge(0.5), cv=5)
cv5 = KFold(5, shuffle=True, random_state=42)
print('CV:', cross_val_score(stack, X, y, cv=cv5, scoring='r2').mean().round(3))
Xtr,Xte,ytr,yte = train_test_split(X, y, test_size=0.2, random_state=42)
stack.fit(Xtr, ytr)
print('holdout R²:', r2_score(yte, stack.predict(Xte)).round(3))

CV R²=0.480; holdout R²=0.472 — stacking matches best single model (Ridge at 0.479) and slightly beats GB+RF alone

Why: On 442 samples, the base models are correlated enough that stacking gain is modest. The holdout score (0.472) slightly below CV (0.480) is the typical small optimism gap.

modelCV R²holdout R²
Ridge (best single)0.4790.454
GradBoost0.4580.485
RandomForest0.4440.459
Stacking (all 3)0.4800.472

34. Fill in: holdout R² for StackingRegressor: Ridge + RF + GBM

Comparison

Comparison matrix

From StackingRegressor: Ridge + RF + GBM: refill the holdout R² column from what you know. The rest of the table is as it appeared.

modelCV R²holdout R²
Ridge (best single)0.4790.454
GradBoost0.4580.485
RandomForest0.4440.459
Stacking (all 3)0.4800.472

35. CV score as leaderboard estimate

Concept

Before submitting, estimate your public leaderboard score from CV: the K-fold mean is your unbiased estimate; the std tells you variance across folds.

\[ \widehat{\text{score}}_{\text{leaderboard}} \approx \overline{\text{CV}} \pm 1\text{-}2\,\sigma_{\text{fold}} \]

Our stacking model: CV 0.480 ± 0.090 → predict leaderboard near 0.39–0.57. Holdout truth was 0.472 — within one fold std. CV is calibrated.

36. By analogy: CV score as leaderboard estimate

Analogy

Discussion prompt

Explain CV score as leaderboard estimate by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Before submitting, estimate your public leaderboard score from CV: the K-fold mean is your unbiased estimate; the std tells you variance across folds.

37. Permutation feature importance

Concept

After selecting a model, identify which features drive it: permutation importance shuffles each feature in turn and measures R² drop (Lesson 55). Works on any model — unlike tree split-count importance.

featuremean R² dropstd
bmi0.3176±0.040
s50.2778±0.026
bp0.0313±0.005
others<0.03—

The permutation ranking matches the raw correlation ranking (bmi, s5, bp) — consistent signal, no spurious interactions are being exploited.

38. Watch it run: Permutation feature importance

Pattern

Step through it

Step through Permutation feature importance one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: feature is bmi
  2. Step 2: feature is s5
  3. Step 3: feature is bp
  4. Step 4: feature is others

39. Without one step: The competition pipeline recipe

Constraint

Discussion prompt

Run The competition pipeline recipe with this step confiscated:

Model selection: compare linear/tree/neural families by CV, never training score

Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.

Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.

Answer:

  1. EDA: distributions, correlation table, missing value map, IQR outlier check
  2. Feature engineering: add interactions/polynomials only if CV score improves
  3. Baseline: Ridge or GBM; scale inputs; run 5-fold CV R²
  4. Model selection: compare linear/tree/neural families by CV, never training score
  5. Ensemble: stack diverse models (Ridge + RF + GBM); meta-learner = Ridge
  6. Leaderboard estimate: CV mean ± σ predicts holdout before submission
  7. Explain: permutation importance to identify the top features

40. The competition pipeline recipe

Pattern

  1. EDA: distributions, correlation table, missing value map, IQR outlier check
  2. Feature engineering: add interactions/polynomials only if CV score improves
  3. Baseline: Ridge or GBM; scale inputs; run 5-fold CV R²
  4. Model selection: compare linear/tree/neural families by CV, never training score
  5. Ensemble: stack diverse models (Ridge + RF + GBM); meta-learner = Ridge
  6. Leaderboard estimate: CV mean ± σ predicts holdout before submission
  7. Explain: permutation importance to identify the top features

41. Where does it stop working: The competition pipeline recipe

Edge cases

Discussion prompt

The competition pipeline recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. EDA: distributions, correlation table, missing value map, IQR outlier check
  2. Feature engineering: add interactions/polynomials only if CV score improves
  3. Baseline: Ridge or GBM; scale inputs; run 5-fold CV R²
  4. Model selection: compare linear/tree/neural families by CV, never training score
  5. Ensemble: stack diverse models (Ridge + RF + GBM); meta-learner = Ridge
  6. Leaderboard estimate: CV mean ± σ predicts holdout before submission
  7. Explain: permutation importance to identify the top features

42. Rule out three: Check yourself — feature evaluation

Elimination

Eliminate the wrong options

You add a bmi×s5 interaction feature. Training R² goes from 0.52 to 0.55, but CV R² stays at 0.479. What should you do?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Drop the interaction — CV says it adds no generalization value
  • B. Keep it — the training score improved, so it must help
  • C. Use a larger model to take advantage of the new feature
  • D. Add more polynomial terms to extract more signal from bmi and s5

Survives elimination: A

Why: CV R² is the unbiased estimate of generalization. The training R² increase is pure overfitting — the model is fitting noise in the interaction column. Drop it.

43. Check yourself — feature evaluation

Check

Work through this before clicking.

Check your understanding

You add a bmi×s5 interaction feature. Training R² goes from 0.52 to 0.55, but CV R² stays at 0.479. What should you do?

  • A. Drop the interaction — CV says it adds no generalization value (correct)
  • B. Keep it — the training score improved, so it must help
  • C. Use a larger model to take advantage of the new feature
  • D. Add more polynomial terms to extract more signal from bmi and s5

Answer: A

Why: CV R² is the unbiased estimate of generalization. The training R² increase is pure overfitting — the model is fitting noise in the interaction column. Drop it.

Why B tempts people
Training R² increases for any feature, even pure noise — it is not a valid selection criterion. Only CV or holdout score tells you about generalization.
Why C tempts people
A larger model will overfit even more — the CV score of the interaction feature already shows it adds no signal worth capturing.
Why D tempts people
More polynomial terms would explode the feature space further while CV is already flat, adding variance without bias reduction.

44. Answer it before you see the options: Check yourself — model selection

Prediction

Predict first

On a 442-sample regression dataset, MLP (64 units) achieves training R²=0.95 but CV R²=0.023. What does this indicate?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: High variance overfitting — MLP memorizes the small training set but cannot generalize

Why: Training R²=0.95 vs CV R²=0.023 is the hallmark of high variance overfitting. The MLP memorizes 442 training points but generalizes poorly — exactly the bias-variance regime (Lesson 32) where flexible models fail on small data.

45. Check yourself — model selection

Check

Classic small-data trap.

Check your understanding

On a 442-sample regression dataset, MLP (64 units) achieves training R²=0.95 but CV R²=0.023. What does this indicate?

  • A. High variance overfitting — MLP memorizes the small training set but cannot generalize (correct)
  • B. High bias underfitting — the MLP capacity is too low for the data
  • C. The CV splits are stratified incorrectly, causing data leakage
  • D. The scaler was not applied to the MLP inputs

Answer: A

Why: Training R²=0.95 vs CV R²=0.023 is the hallmark of high variance overfitting. The MLP memorizes 442 training points but generalizes poorly — exactly the bias-variance regime (Lesson 32) where flexible models fail on small data.

Why B tempts people
High bias gives LOW training R² (the model also fails on training data). A 64-unit hidden layer is plenty of capacity; the problem is too much capacity, not too little.
Why C tempts people
Leakage would inflate CV R², not crush it. A leaky model shows CV ≈ train ≈ suspiciously high. Here CV is near 0, the opposite.
Why D tempts people
Scaling is important but an unscaled MLP would underfit uniformly — both training and CV would suffer. The huge train/CV gap points to memorization, not a preprocessing error.

46. Rule out three: Check yourself — stacking

Elimination

Eliminate the wrong options

In StackingRegressor(cv=5), what do the '5-fold' splits inside the stacker do?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Generate out-of-fold predictions to train the meta-learner without data leakage
  • B. Select the best of the 5 base models to be the final predictor
  • C. Evaluate the meta-learner's generalization on 5 test folds
  • D. Average the predictions of all 5 base model checkpoints

Survives elimination: A

Why: The inner cv=5 in StackingRegressor creates out-of-fold (OOF) predictions: for each fold, base learners train on the other 4 folds and predict on the held-out fold. The OOF predictions form the level-1 feature matrix the meta-learner trains on — this prevents the meta-learner from seeing base model outputs computed on the same data it trains on.

47. Check yourself — stacking

Check

Think about what stacking actually does.

Check your understanding

In StackingRegressor(cv=5), what do the '5-fold' splits inside the stacker do?

  • A. Generate out-of-fold predictions to train the meta-learner without data leakage (correct)
  • B. Select the best of the 5 base models to be the final predictor
  • C. Evaluate the meta-learner's generalization on 5 test folds
  • D. Average the predictions of all 5 base model checkpoints

Answer: A

Why: The inner cv=5 in StackingRegressor creates out-of-fold (OOF) predictions: for each fold, base learners train on the other 4 folds and predict on the held-out fold. The OOF predictions form the level-1 feature matrix the meta-learner trains on — this prevents the meta-learner from seeing base model outputs computed on the same data it trains on.

Why B tempts people
All base estimators are always used; none are eliminated. A stacker combines them, it does not select among them.
Why C tempts people
Those folds evaluate the base learners, not the meta-learner. The outer cross_val_score call (if any) evaluates the full stack.
Why D tempts people
StackingRegressor uses a learned meta-learner (e.g., Ridge), not simple averaging. Averaging is a separate technique called blending or voting.

48. Your turn: build the pipeline

Section

Project

49. Project: diabetes competition pipeline

Concept

Replicate and extend the full pipeline on load_diabetes: EDA → feature engineering → model selection → stacking → leaderboard estimate.

#milestonekey tool
1EDA: correlations + outliersnp.corrcoef, IQR
2Feature engineering: test an interactioncross_val_score before/after
3Model selection: 5-fold CV tableRidge, RF, GBM, MLP
4Stacking + holdout estimateStackingRegressor, train_test_split

Build rule: write a single pipeline script you could hand to a competition teammate. Every claim — 'this feature helps', 'this model wins' — must be backed by a CV number.

50. Break it if you can: Project: diabetes competition pipeline

Counterexample

Discussion prompt

Replicate and extend the full pipeline on load_diabetes: EDA → feature engineering → model selection → stacking → leaderboard estimate.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Build rule: write a single pipeline script you could hand to a competition teammate. Every claim — 'this feature helps', 'this model wins' — must be backed by a CV number.

51. Milestone 1 — EDA report

Worked example

Your turn: print (a) top-3 correlated features, (b) number of missing values, (c) count of BMI outliers by IQR. Predict the results before running.

Hint: np.corrcoef(X[:,i], y)[0,1] for each column; np.isnan(X).sum(); IQR fence at Q1-1.5·IQR and Q3+1.5·IQR.

from sklearn.datasets import load_diabetes
import numpy as np
data = load_diabetes(); X, y = data.data, data.target
names = list(data.feature_names)
corrs = [np.corrcoef(X[:,i],y)[0,1] for i in range(10)]
top3 = sorted(range(10), key=lambda i:abs(corrs[i]), reverse=True)[:3]
print('top3:', [(names[i], round(corrs[i],3)) for i in top3])
print('missing:', np.isnan(X).sum())
bmi = X[:, names.index('bmi')]
q1,q3 = np.percentile(bmi,[25,75]); iqr=q3-q1
print('bmi outliers:', ((bmi<q1-1.5*iqr)|(bmi>q3+1.5*iqr)).sum())
checkresult
top featuresbmi(0.586), s5(0.566), bp(0.441)
missing values0
BMI outliers (IQR)3

52. Milestone 2 — model selection table

Worked example

Your turn: run 5-fold CV R² for all five model families and predict which model wins. Predict why MLP will trail.

Hint: StandardScaler().fit_transform(X) first; KFold(5, shuffle=True, random_state=42); pass scoring='r2' to cross_val_score.

from sklearn.datasets import load_diabetes
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import Ridge
from sklearn.tree import DecisionTreeRegressor
from sklearn.ensemble import RandomForestRegressor, GradientBoostingRegressor
from sklearn.neural_network import MLPRegressor
from sklearn.model_selection import cross_val_score, KFold
X, y = load_diabetes(return_X_y=True)
X = StandardScaler().fit_transform(X)
cv = KFold(5, shuffle=True, random_state=42)
for nm, m in [('Ridge', Ridge(1.0)),
              ('DTree', DecisionTreeRegressor(max_depth=3,random_state=42)),
              ('RF',    RandomForestRegressor(50,max_depth=3,random_state=42)),
              ('GBM',  GradientBoostingRegressor(50,max_depth=2,random_state=42)),
              ('MLP',  MLPRegressor((64,),max_iter=500,random_state=42))]:
    s = cross_val_score(m, X, y, cv=cv, scoring='r2')
    print(f'{nm}: {s.mean():.3f} +/- {s.std():.3f}')
modelCV R²±std
Ridge0.4790.083
GBM0.4580.078
RF0.4440.084
DTree0.3280.077
MLP0.0230.101

53. What each one costs: Milestone 2 — model selection table

Trade off

Comparison matrix

From Milestone 2 — model selection table: every row here is a choice with a cost. Fill the CV R² column, then say which row you would actually pick and what you give up for it.

modelCV R²±std
Ridge0.4790.083
GBM0.4580.078
RF0.4440.084
DTree0.3280.077
MLP0.0230.101

54. Milestone 3 — stacking + holdout

Worked example

Your turn: build the stacking ensemble, estimate holdout R², and compare the CV estimate to the actual holdout score.

Hint: StackingRegressor(estimators=[...], final_estimator=Ridge(0.5), cv=5); 80/20 split with random_state=42; r2_score on the held-out 20%.

from sklearn.datasets import load_diabetes
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import Ridge
from sklearn.ensemble import (RandomForestRegressor,
    GradientBoostingRegressor, StackingRegressor)
from sklearn.model_selection import cross_val_score, KFold, train_test_split
from sklearn.metrics import r2_score
X, y = load_diabetes(return_X_y=True)
X = StandardScaler().fit_transform(X)
base = [('r',Ridge(1.0)),('rf',RandomForestRegressor(50,max_depth=3,random_state=42)),
        ('gb',GradientBoostingRegressor(50,max_depth=2,random_state=42))]
stack = StackingRegressor(base, final_estimator=Ridge(0.5), cv=5)
print('CV:', cross_val_score(stack, X, y,
      cv=KFold(5,shuffle=True,random_state=42), scoring='r2').mean().round(3))
Xtr,Xte,ytr,yte = train_test_split(X,y,test_size=0.2,random_state=42)
stack.fit(Xtr,ytr)
print('holdout:', r2_score(yte, stack.predict(Xte)).round(3))
metricvalueinterpretation
CV R² (5-fold)0.480pre-submission estimate
holdout R²0.472simulated leaderboard
optimism gap0.008CV is well-calibrated here

55. Fill in: value for Milestone 3 — stacking + holdout

Comparison

Comparison matrix

From Milestone 3 — stacking + holdout: refill the value column from what you know. The rest of the table is as it appeared.

metricvalueinterpretation
CV R² (5-fold)0.480pre-submission estimate
holdout R²0.472simulated leaderboard
optimism gap0.008CV is well-calibrated here

56. Show it off

Concept

Slides closed: walk through your pipeline script as if presenting at a competition debrief. Cover (1) what EDA revealed about the top features, (2) why you selected Ridge as the best single model, (3) how the stacker was trained without leakage, and (4) how close your CV estimate was to the holdout score.

Stretch (homework): add permutation importance plots with sklearn.inspection.permutation_importance; implement ydata-profiling for automated EDA; write a SHAP summary plot for your best model. Next lesson: advanced ensembles and submission strategy.

57. Connect it up: Lesson 71: Kaggle Competition Pipeline

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — EDA: what the data is telling you · Feature engineering: building signal · Model selection: linear, tree, neural · Ensemble: stacking diverse models · Your turn: build the pipeline. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

58. What you can do now

Recap

stepthe one thing to remember
EDAcorrelations guide features; outliers inform imputation strategy
feature selectionCV score, not training score — always
model selectionMLP fails on small data; Ridge/GBM dominate ≤1k samples
stackingOOF predictions → meta-learner; diversity beats duplicates
CV estimateCV mean ≈ leaderboard; CV std is your confidence interval

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 71 (Phase 4 — Competition Pipeline) — Barron · USAAIO Round 2 Preparation, 2026
  2. sklearn diabetes dataset: EDA, feature engineering, Ridge/RF/GB/Stacking CV scores verified — numpy 2.2.6, scikit-learn, real execution, June 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108