USAAIO Lesson 71, from Phase 4, on the end-to-end competition workflow. It covers exploratory data analysis - distributions, correlations, missing values, and outliers - then interaction and polynomial feature engineering, model selection across the linear, tree, and neural families with 5-fold cross-validation, stacking ensembles, and using cross-validation scores to predict where you will land on the leaderboard. It was verified on the sklearn diabetes dataset with numpy 2.2.6 and scikit-learn. The lesson runs to 30 slides.
Subject: Machine Learning · 58 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 71 · Phase 4
End-to-end: EDA → feature engineering → model selection → ensemble → CV-based leaderboard estimate. The full workflow in one session.
Objectives
Section
Part 1 of 4
Concept
Competitions are lost to silent data problems — a skewed target, a near-duplicate feature, or outliers that drag every model toward one region of the space.
A 30-minute EDA on the diabetes dataset surfaces everything you need: 442 samples, 10 features, no missing values, and 3 BMI outliers by IQR.
Counterexample
Discussion prompt
Competitions are lost to silent data problems — a skewed target, a near-duplicate feature, or outliers that drag every model toward one region of the space.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
A 30-minute EDA on the diabetes dataset surfaces everything you need: 442 samples, 10 features, no missing values, and 3 BMI outliers by IQR.
Estimation
Predict first
Load load_diabetes, print shape and target summary, check for missing values, compute feature–target correlations, and flag BMI outliers.
Commit before you compute: what does EDA on the diabetes dataset come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Shape (442, 10); target 25–346, mean 152.1; 0 missing; top features bmi r=0.586, s5 r=0.566, bp r=0.441; 3 BMI outliers
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. No imputation needed here; bmi and s5 dominate the correlation ranking — they are the first candidates for feature engineering.
Worked example
Load load_diabetes, print shape and target summary, check for missing values, compute feature–target correlations, and flag BMI outliers.
from sklearn.datasets import load_diabetes
import numpy as np
data = load_diabetes()
X, y, names = data.data, data.target, list(data.feature_names)
print(X.shape, y.min(), y.max(), y.mean().round(1))
print('missing:', np.isnan(X).sum())
corrs = [np.corrcoef(X[:,i], y)[0,1] for i in range(10)]
top3 = sorted(range(10), key=lambda i: abs(corrs[i]), reverse=True)[:3]
for i in top3: print(f'{names[i]}: r={corrs[i]:.3f}')
bmi = X[:, names.index('bmi')]
q1,q3 = np.percentile(bmi,[25,75]); iqr=q3-q1
print('bmi outliers:', ((bmi<q1-1.5*iqr)|(bmi>q3+1.5*iqr)).sum())Shape (442, 10); target 25–346, mean 152.1; 0 missing; top features bmi r=0.586, s5 r=0.566, bp r=0.441; 3 BMI outliers
Why: No imputation needed here; bmi and s5 dominate the correlation ranking — they are the first candidates for feature engineering.
| feature | r with target | note |
|---|---|---|
| bmi | 0.586 | strongest predictor |
| s5 | 0.566 | log serum triglycerides |
| bp | 0.441 | blood pressure |
| s3 | −0.395 | negative correlation |
| other 6 | <0.22 | weaker signal |
Comparison
Comparison matrix
From EDA on the diabetes dataset: refill the r with target column from what you know. The rest of the table is as it appeared.
| feature | r with target | note |
|---|---|---|
| bmi | 0.586 | strongest predictor |
| s5 | 0.566 | log serum triglycerides |
| bp | 0.441 | blood pressure |
| s3 | −0.395 | negative correlation |
| other 6 | <0.22 | weaker signal |
Concept
| strategy | when to use | risk |
|---|---|---|
| drop row | outlier is a data-entry error | loses real examples |
| clip to fence (IQR±1.5) | genuine but extreme; tree models don't need it | distorts real extremes |
| log / power transform | right-skewed target or feature | changes scale, affects interpretability |
| encode as binary flag | missingness is informative (MNAR) | adds a feature; safe default |
For diabetes the 3 BMI outliers are genuine patients — leave them in and let tree-based models handle them naturally.
Trade off
Comparison matrix
From Outlier strategies: drop vs. clip vs. encode: every row here is a choice with a cost. Fill the when to use column, then say which row you would actually pick and what you give up for it.
| strategy | when to use | risk |
|---|---|---|
| drop row | outlier is a data-entry error | loses real examples |
| clip to fence (IQR±1.5) | genuine but extreme; tree models don't need it | distorts real extremes |
| log / power transform | right-skewed target or feature | changes scale, affects interpretability |
| encode as binary flag | missingness is informative (MNAR) | adds a feature; safe default |
Section
Part 2 of 4
Concept
bmi × bp captures a joint effect neither captures alonebmi², bmi·s5 lets a linear model fit a curvelog(s5) stabilizes a skewed lab value; age bins encode non-linearityRule of thumb: add a feature only if CV score improves. A raw feature matrix with 10 columns can explode to 65 degree-2 polynomial terms — every one adds variance and needs regularization.
Analogy
Discussion prompt
Explain Three feature engineering primitives by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Rule of thumb: add a feature only if CV score improves. A raw feature matrix with 10 columns can explode to 65 degree-2 polynomial terms — every one adds variance and needs regularization.
Missing information
Discussion prompt
Hypothesis: high BMI combined with high triglycerides (s5) predicts worse outcomes synergistically. Add bmi × s5 and re-run Ridge CV.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
bmi and s5 are already in the model; their product adds colinearity without new signal for a linear model. The feature is useless here — and CV told us that without touching the test set.
Worked example
Hypothesis: high BMI combined with high triglycerides (s5) predicts worse outcomes synergistically. Add bmi × s5 and re-run Ridge CV.
from sklearn.datasets import load_diabetes
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import Ridge
from sklearn.model_selection import cross_val_score, KFold
import numpy as np
data = load_diabetes(); X, y = data.data, data.target
names = list(data.feature_names)
bmi_s5 = X[:, names.index('bmi')] * X[:, names.index('s5')]
X_aug = np.column_stack([X, bmi_s5])
sc = StandardScaler(); cv = KFold(5, shuffle=True, random_state=42)
base = cross_val_score(Ridge(1.0), sc.fit_transform(X), y, cv=cv, scoring='r2')
aug = cross_val_score(Ridge(1.0), sc.fit_transform(X_aug), y, cv=cv, scoring='r2')
print('base:', base.mean().round(3), 'aug:', aug.mean().round(3))base R²=0.479; aug R²=0.479 — the interaction adds nothing for Ridge on this dataset
Why: bmi and s5 are already in the model; their product adds colinearity without new signal for a linear model. The feature is useless here — and CV told us that without touching the test set.
| model | features | CV R² |
|---|---|---|
| Ridge | raw 10 | 0.479 |
| Ridge | raw 10 + bmi×s5 | 0.479 |
| verdict | — | interaction ≈ no gain |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
base R²=0.479; aug R²=0.479 — the interaction adds nothing for Ridge on this dataset
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Hypothesis: high BMI combined with high triglycerides (s5) predicts worse outcomes synergistically. Add bmi × s5 and re-run Ridge CV.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Add bmi × s5, fit a Ridge on all 442 samples, get training R²=0.521 — the feature helped.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Training R² always weakly increases when you add features — the model just interpolates the noise.
Always measure feature value with held-out or CV score, not training score.
Why: Training R² always weakly increases when you add features — the model just interpolates the noise. This is pure overfitting detection failure.
Trap
Add bmi × s5, fit a Ridge on all 442 samples, get training R²=0.521 — the feature helped.
Report training R² as evidence the interaction improves the model
Why: Training R² always weakly increases when you add features — the model just interpolates the noise. This is pure overfitting detection failure.
Always measure feature value with held-out or CV score, not training score.
Run cross_val_score before and after adding the feature; CV R²=0.479 vs 0.479 — feature rejected
Why: CV measures generalization. The feature that looks useful on training data but flat on CV is adding variance without signal — drop it.
Break the constraint
Discussion prompt
The rule this trap just fixed:
CV measures generalization. The feature that looks useful on training data but flat on CV is adding variance without signal — drop it.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
Training R² always weakly increases when you add features — the model just interpolates the noise. This is pure overfitting detection failure.
Section
Part 3 of 4
Concept
| family | example | strength | weakness |
|---|---|---|---|
| linear | Ridge, Lasso | fast, interpretable, regularizable | can't model interactions without manual features |
| tree-based | RF, GradBoost | non-linear, handles outliers/interactions | can overfit small datasets |
| neural | MLP | flexible universal approximator | needs more data, hyperparameter-sensitive |
For 442 samples the tree-based and neural models face a high-variance regime — you need heavy regularization or ensemble averaging to beat Ridge (Lesson 32).
Explain it
Discussion prompt
Explain Three model families — their trade-offs to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
For 442 samples the tree-based and neural models face a high-variance regime — you need heavy regularization or ensemble averaging to beat Ridge (Lesson 32).
Estimation
Predict first
Scale the data, then run 5-fold CV R² for Ridge, DecisionTree (depth 3), RandomForest (50 trees), GradientBoosting, and MLP.
Commit before you compute: what does 5-fold CV comparison: all three families come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Ridge leads at 0.479; GBM second at 0.458; RF at 0.444; DTree at 0.328; MLP at 0.023 — the small dataset punishes MLP hard
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. MLP's 0.023 isn't a bug — it simply lacks enough samples to converge reliably with a 64-unit hidden layer.
Worked example
Scale the data, then run 5-fold CV R² for Ridge, DecisionTree (depth 3), RandomForest (50 trees), GradientBoosting, and MLP.
from sklearn.datasets import load_diabetes
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import Ridge
from sklearn.tree import DecisionTreeRegressor
from sklearn.ensemble import RandomForestRegressor, GradientBoostingRegressor
from sklearn.neural_network import MLPRegressor
from sklearn.model_selection import cross_val_score, KFold
data = load_diabetes(); X, y = data.data, data.target
X = StandardScaler().fit_transform(X)
cv = KFold(5, shuffle=True, random_state=42)
for name, m in [
('Ridge', Ridge(1.0)),
('DTree', DecisionTreeRegressor(max_depth=3, random_state=42)),
('RF', RandomForestRegressor(50, max_depth=3, random_state=42)),
('GBM', GradientBoostingRegressor(50, max_depth=2, random_state=42)),
('MLP', MLPRegressor((64,), max_iter=500, random_state=42)),
]:
s = cross_val_score(m, X, y, cv=cv, scoring='r2')
print(f'{name}: {s.mean():.3f} ± {s.std():.3f}')Ridge leads at 0.479; GBM second at 0.458; RF at 0.444; DTree at 0.328; MLP at 0.023 — the small dataset punishes MLP hard
Why: MLP's 0.023 isn't a bug — it simply lacks enough samples to converge reliably with a 64-unit hidden layer. GBM's depth-2 trees generalize better than depth-3 DTree because shallow trees regularize by construction.
| model | CV R² | CV std |
|---|---|---|
| Ridge | 0.479 | ±0.083 |
| GradBoost (d=2) | 0.458 | ±0.078 |
| RandomForest (d=3) | 0.444 | ±0.084 |
| DecisionTree (d=3) | 0.328 | ±0.077 |
| MLP (64,) | 0.023 | ±0.101 |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Ridge leads at 0.479; GBM second at 0.458; RF at 0.444; DTree at 0.328; MLP at 0.023 — the small dataset punishes MLP hard
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Scale the data, then run 5-fold CV R² for Ridge, DecisionTree (depth 3), RandomForest (50 trees), GradientBoosting, and MLP.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Fit all five models on the full dataset, pick the one with the highest training R² — that's the best model.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: MLP memorizes the small dataset completely (Lesson 32 bias-variance).
Select by CV score on held-out folds, never training score.
Why: MLP memorizes the small dataset completely (Lesson 32 bias-variance). Its CV R²=0.023 means it predicts near-baseline on new data.
Trap
Fit all five models on the full dataset, pick the one with the highest training R² — that's the best model.
Select MLP because it achieves training R²≈0.95 on 442 samples
Why: MLP memorizes the small dataset completely (Lesson 32 bias-variance). Its CV R²=0.023 means it predicts near-baseline on new data.
Select by CV score on held-out folds, never training score.
Select Ridge (CV R²=0.479) — lower training R² but far better generalization than MLP's CV R²=0.023
Why: CV measures expected leaderboard performance. Training R² measures memorization. They diverge most for flexible models on small data.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
diabetes the 3 BMI outliers are genuine patients — leave them in and let tree-based models handle them naturally.; For 442 samples the tree-based and neural models face a high-variance regime — you need heavy regularization or ensemble averaging to beat Ridge (Lesson 32).bmi × s5, fit a Ridge on all 442 samples, get training R²=0.521 — the feature helped.; Fit all five models on the full dataset, pick the one with the highest training R² — that's the best model.Section
Part 4 of 4
Concept
Stacking trains a meta-learner to combine out-of-fold predictions from diverse base learners. The meta-learner learns when to trust each base model.
yDiversity is essential — stacking Ridge + RF + GBM helps because each model captures different structure. Stacking three copies of Ridge gains nothing.
Explain it
Discussion prompt
Explain Why stacking works to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Diversity is essential — stacking Ridge + RF + GBM helps because each model captures different structure. Stacking three copies of Ridge gains nothing.
Estimation
Predict first
Build a stacking ensemble with Ridge, RandomForest, and GradientBoosting as base learners; use Ridge as the meta-learner. Measure CV and holdout R².
Commit before you compute: what does StackingRegressor: Ridge + RF + GBM come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: CV R²=0.480; holdout R²=0.472 — stacking matches best single model (Ridge at 0.479) and slightly beats GB+RF alone
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. On 442 samples, the base models are correlated enough that stacking gain is modest.
Worked example
Build a stacking ensemble with Ridge, RandomForest, and GradientBoosting as base learners; use Ridge as the meta-learner. Measure CV and holdout R².
from sklearn.datasets import load_diabetes
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import Ridge
from sklearn.ensemble import RandomForestRegressor, GradientBoostingRegressor, StackingRegressor
from sklearn.model_selection import cross_val_score, KFold, train_test_split
from sklearn.metrics import r2_score
import numpy as np
data = load_diabetes(); X, y = data.data, data.target
X = StandardScaler().fit_transform(X)
base = [('ridge', Ridge(1.0)),
('rf', RandomForestRegressor(50, max_depth=3, random_state=42)),
('gb', GradientBoostingRegressor(50, max_depth=2, random_state=42))]
stack = StackingRegressor(base, final_estimator=Ridge(0.5), cv=5)
cv5 = KFold(5, shuffle=True, random_state=42)
print('CV:', cross_val_score(stack, X, y, cv=cv5, scoring='r2').mean().round(3))
Xtr,Xte,ytr,yte = train_test_split(X, y, test_size=0.2, random_state=42)
stack.fit(Xtr, ytr)
print('holdout R²:', r2_score(yte, stack.predict(Xte)).round(3))CV R²=0.480; holdout R²=0.472 — stacking matches best single model (Ridge at 0.479) and slightly beats GB+RF alone
Why: On 442 samples, the base models are correlated enough that stacking gain is modest. The holdout score (0.472) slightly below CV (0.480) is the typical small optimism gap.
| model | CV R² | holdout R² |
|---|---|---|
| Ridge (best single) | 0.479 | 0.454 |
| GradBoost | 0.458 | 0.485 |
| RandomForest | 0.444 | 0.459 |
| Stacking (all 3) | 0.480 | 0.472 |
Comparison
Comparison matrix
From StackingRegressor: Ridge + RF + GBM: refill the holdout R² column from what you know. The rest of the table is as it appeared.
| model | CV R² | holdout R² |
|---|---|---|
| Ridge (best single) | 0.479 | 0.454 |
| GradBoost | 0.458 | 0.485 |
| RandomForest | 0.444 | 0.459 |
| Stacking (all 3) | 0.480 | 0.472 |
Concept
Before submitting, estimate your public leaderboard score from CV: the K-fold mean is your unbiased estimate; the std tells you variance across folds.
\[ \widehat{\text{score}}_{\text{leaderboard}} \approx \overline{\text{CV}} \pm 1\text{-}2\,\sigma_{\text{fold}} \]
Our stacking model: CV 0.480 ± 0.090 → predict leaderboard near 0.39–0.57. Holdout truth was 0.472 — within one fold std. CV is calibrated.
Analogy
Discussion prompt
Explain CV score as leaderboard estimate by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Before submitting, estimate your public leaderboard score from CV: the K-fold mean is your unbiased estimate; the std tells you variance across folds.
Concept
After selecting a model, identify which features drive it: permutation importance shuffles each feature in turn and measures R² drop (Lesson 55). Works on any model — unlike tree split-count importance.
| feature | mean R² drop | std |
|---|---|---|
| bmi | 0.3176 | ±0.040 |
| s5 | 0.2778 | ±0.026 |
| bp | 0.0313 | ±0.005 |
| others | <0.03 | — |
The permutation ranking matches the raw correlation ranking (bmi, s5, bp) — consistent signal, no spurious interactions are being exploited.
Pattern
Step through it
Step through Permutation feature importance one row at a time. What is driving the change, and what would the row after the last one be?
Constraint
Discussion prompt
Run The competition pipeline recipe with this step confiscated:
Model selection: compare linear/tree/neural families by CV, never training score
Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.
Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.
Answer:
Pattern
Edge cases
Discussion prompt
The competition pipeline recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
Elimination
Eliminate the wrong options
You add a bmi×s5 interaction feature. Training R² goes from 0.52 to 0.55, but CV R² stays at 0.479. What should you do?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: CV R² is the unbiased estimate of generalization. The training R² increase is pure overfitting — the model is fitting noise in the interaction column. Drop it.
Check
Work through this before clicking.
Check your understanding
You add a bmi×s5 interaction feature. Training R² goes from 0.52 to 0.55, but CV R² stays at 0.479. What should you do?
Answer: A
Why: CV R² is the unbiased estimate of generalization. The training R² increase is pure overfitting — the model is fitting noise in the interaction column. Drop it.
Prediction
Predict first
On a 442-sample regression dataset, MLP (64 units) achieves training R²=0.95 but CV R²=0.023. What does this indicate?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: High variance overfitting — MLP memorizes the small training set but cannot generalize
Why: Training R²=0.95 vs CV R²=0.023 is the hallmark of high variance overfitting. The MLP memorizes 442 training points but generalizes poorly — exactly the bias-variance regime (Lesson 32) where flexible models fail on small data.
Check
Classic small-data trap.
Check your understanding
On a 442-sample regression dataset, MLP (64 units) achieves training R²=0.95 but CV R²=0.023. What does this indicate?
Answer: A
Why: Training R²=0.95 vs CV R²=0.023 is the hallmark of high variance overfitting. The MLP memorizes 442 training points but generalizes poorly — exactly the bias-variance regime (Lesson 32) where flexible models fail on small data.
Elimination
Eliminate the wrong options
In StackingRegressor(cv=5), what do the '5-fold' splits inside the stacker do?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: The inner cv=5 in StackingRegressor creates out-of-fold (OOF) predictions: for each fold, base learners train on the other 4 folds and predict on the held-out fold. The OOF predictions form the level-1 feature matrix the meta-learner trains on — this prevents the meta-learner from seeing base model outputs computed on the same data it trains on.
Check
Think about what stacking actually does.
Check your understanding
In StackingRegressor(cv=5), what do the '5-fold' splits inside the stacker do?
Answer: A
Why: The inner cv=5 in StackingRegressor creates out-of-fold (OOF) predictions: for each fold, base learners train on the other 4 folds and predict on the held-out fold. The OOF predictions form the level-1 feature matrix the meta-learner trains on — this prevents the meta-learner from seeing base model outputs computed on the same data it trains on.
Section
Project
Concept
Replicate and extend the full pipeline on load_diabetes: EDA → feature engineering → model selection → stacking → leaderboard estimate.
| # | milestone | key tool |
|---|---|---|
| 1 | EDA: correlations + outliers | np.corrcoef, IQR |
| 2 | Feature engineering: test an interaction | cross_val_score before/after |
| 3 | Model selection: 5-fold CV table | Ridge, RF, GBM, MLP |
| 4 | Stacking + holdout estimate | StackingRegressor, train_test_split |
Build rule: write a single pipeline script you could hand to a competition teammate. Every claim — 'this feature helps', 'this model wins' — must be backed by a CV number.
Counterexample
Discussion prompt
Replicate and extend the full pipeline on load_diabetes: EDA → feature engineering → model selection → stacking → leaderboard estimate.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Build rule: write a single pipeline script you could hand to a competition teammate. Every claim — 'this feature helps', 'this model wins' — must be backed by a CV number.
Worked example
Your turn: print (a) top-3 correlated features, (b) number of missing values, (c) count of BMI outliers by IQR. Predict the results before running.
Hint: np.corrcoef(X[:,i], y)[0,1] for each column; np.isnan(X).sum(); IQR fence at Q1-1.5·IQR and Q3+1.5·IQR.
from sklearn.datasets import load_diabetes
import numpy as np
data = load_diabetes(); X, y = data.data, data.target
names = list(data.feature_names)
corrs = [np.corrcoef(X[:,i],y)[0,1] for i in range(10)]
top3 = sorted(range(10), key=lambda i:abs(corrs[i]), reverse=True)[:3]
print('top3:', [(names[i], round(corrs[i],3)) for i in top3])
print('missing:', np.isnan(X).sum())
bmi = X[:, names.index('bmi')]
q1,q3 = np.percentile(bmi,[25,75]); iqr=q3-q1
print('bmi outliers:', ((bmi<q1-1.5*iqr)|(bmi>q3+1.5*iqr)).sum())| check | result |
|---|---|
| top features | bmi(0.586), s5(0.566), bp(0.441) |
| missing values | 0 |
| BMI outliers (IQR) | 3 |
Worked example
Your turn: run 5-fold CV R² for all five model families and predict which model wins. Predict why MLP will trail.
Hint: StandardScaler().fit_transform(X) first; KFold(5, shuffle=True, random_state=42); pass scoring='r2' to cross_val_score.
from sklearn.datasets import load_diabetes
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import Ridge
from sklearn.tree import DecisionTreeRegressor
from sklearn.ensemble import RandomForestRegressor, GradientBoostingRegressor
from sklearn.neural_network import MLPRegressor
from sklearn.model_selection import cross_val_score, KFold
X, y = load_diabetes(return_X_y=True)
X = StandardScaler().fit_transform(X)
cv = KFold(5, shuffle=True, random_state=42)
for nm, m in [('Ridge', Ridge(1.0)),
('DTree', DecisionTreeRegressor(max_depth=3,random_state=42)),
('RF', RandomForestRegressor(50,max_depth=3,random_state=42)),
('GBM', GradientBoostingRegressor(50,max_depth=2,random_state=42)),
('MLP', MLPRegressor((64,),max_iter=500,random_state=42))]:
s = cross_val_score(m, X, y, cv=cv, scoring='r2')
print(f'{nm}: {s.mean():.3f} +/- {s.std():.3f}')| model | CV R² | ±std |
|---|---|---|
| Ridge | 0.479 | 0.083 |
| GBM | 0.458 | 0.078 |
| RF | 0.444 | 0.084 |
| DTree | 0.328 | 0.077 |
| MLP | 0.023 | 0.101 |
Trade off
Comparison matrix
From Milestone 2 — model selection table: every row here is a choice with a cost. Fill the CV R² column, then say which row you would actually pick and what you give up for it.
| model | CV R² | ±std |
|---|---|---|
| Ridge | 0.479 | 0.083 |
| GBM | 0.458 | 0.078 |
| RF | 0.444 | 0.084 |
| DTree | 0.328 | 0.077 |
| MLP | 0.023 | 0.101 |
Worked example
Your turn: build the stacking ensemble, estimate holdout R², and compare the CV estimate to the actual holdout score.
Hint: StackingRegressor(estimators=[...], final_estimator=Ridge(0.5), cv=5); 80/20 split with random_state=42; r2_score on the held-out 20%.
from sklearn.datasets import load_diabetes
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import Ridge
from sklearn.ensemble import (RandomForestRegressor,
GradientBoostingRegressor, StackingRegressor)
from sklearn.model_selection import cross_val_score, KFold, train_test_split
from sklearn.metrics import r2_score
X, y = load_diabetes(return_X_y=True)
X = StandardScaler().fit_transform(X)
base = [('r',Ridge(1.0)),('rf',RandomForestRegressor(50,max_depth=3,random_state=42)),
('gb',GradientBoostingRegressor(50,max_depth=2,random_state=42))]
stack = StackingRegressor(base, final_estimator=Ridge(0.5), cv=5)
print('CV:', cross_val_score(stack, X, y,
cv=KFold(5,shuffle=True,random_state=42), scoring='r2').mean().round(3))
Xtr,Xte,ytr,yte = train_test_split(X,y,test_size=0.2,random_state=42)
stack.fit(Xtr,ytr)
print('holdout:', r2_score(yte, stack.predict(Xte)).round(3))| metric | value | interpretation |
|---|---|---|
| CV R² (5-fold) | 0.480 | pre-submission estimate |
| holdout R² | 0.472 | simulated leaderboard |
| optimism gap | 0.008 | CV is well-calibrated here |
Comparison
Comparison matrix
From Milestone 3 — stacking + holdout: refill the value column from what you know. The rest of the table is as it appeared.
| metric | value | interpretation |
|---|---|---|
| CV R² (5-fold) | 0.480 | pre-submission estimate |
| holdout R² | 0.472 | simulated leaderboard |
| optimism gap | 0.008 | CV is well-calibrated here |
Concept
Slides closed: walk through your pipeline script as if presenting at a competition debrief. Cover (1) what EDA revealed about the top features, (2) why you selected Ridge as the best single model, (3) how the stacker was trained without leakage, and (4) how close your CV estimate was to the holdout score.
Stretch (homework): add permutation importance plots with sklearn.inspection.permutation_importance; implement ydata-profiling for automated EDA; write a SHAP summary plot for your best model. Next lesson: advanced ensembles and submission strategy.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — EDA: what the data is telling you · Feature engineering: building signal · Model selection: linear, tree, neural · Ensemble: stacking diverse models · Your turn: build the pipeline. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
| step | the one thing to remember |
|---|---|
| EDA | correlations guide features; outliers inform imputation strategy |
| feature selection | CV score, not training score — always |
| model selection | MLP fails on small data; Ridge/GBM dominate ≤1k samples |
| stacking | OOF predictions → meta-learner; diversity beats duplicates |
| CV estimate | CV mean ≈ leaderboard; CV std is your confidence interval |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.