USAAIO Lesson 60, from Week 21. An sklearn Pipeline chains estimators together and prevents leakage; ColumnTransformer applies different transforms to different feature types; and FunctionTransformer wraps any function. Combining a Pipeline with GridSearchCV lets you tune the preprocessing and the model hyperparameters jointly, and joblib saves and loads versioned pipelines. You build a production-ready pipeline that imputes, encodes, scales, and fits a logistic regression, tune it end to end, and serialize it. The lesson runs to 31 slides.
Subject: Machine Learning · 61 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 60 · Week 21
Chain estimators, prevent leakage, tune everything end-to-end, and ship a versioned model file. The plumbing every competition submission needs.
Objectives
Pipeline prevents data leakage and chain fit/transform/predict into one objectColumnTransformer that applies different preprocessing to numeric vs categorical columnsFunctionTransformerPipeline to GridSearchCV and tune both preprocessing and model hyperparameters togetherjoblib.dump / joblib.loadWarm-up
Discussion prompt
Before we open Lesson 60: sklearn Pipeline, ColumnTransformer & GridSearchCV: without looking back, what was the main idea of Hyperparameter Optimization, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
grid search (exhaustive, exponential cost), random search (log-scale sampling, matches or beats grid at a fraction of the evaluations), Gaussian Process Bayesian optimization (surrogate + expected-improvement acquisition), successive halving and Hyperband (early elimination), and practical log-scale tips. All numbers verified with sklearn 1.x on digits, June 2026.
Section
Part 1 of 4
Concept
Pipeline([('name', step), ...]) wraps a sequence of transformers ending in an estimator. Calling fit(X_train, y) on the pipeline fits every step on training data only.
Without a pipeline it is easy to accidentally call scaler.fit(X_all) before the split — fitting on test data poisons every downstream metric. The pipeline's fit enforces the split boundary.
| call | what happens internally |
|---|---|
| pipe.fit(X_train) | each step: fit on train → transform train → pass to next |
| pipe.predict(X_test) | each step: transform only (no re-fit) → final estimator predicts |
| pipe.score(X_test, y_test) | transform + predict + metric in one line |
Counterexample
Discussion prompt
Pipeline([('name', step), ...]) wraps a sequence of transformers ending in an estimator. Calling fit(X_train, y) on the pipeline fits every step on training data only.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Without a pipeline it is easy to accidentally call scaler.fit(X_all) before the split — fitting on test data poisons every downstream metric. The pipeline's fit enforces the split boundary.
Intuition
Picture a factory assembly line: raw parts enter at one end, each station does exactly one job (impute → scale → encode → classify), and the finished product exits at the other end.
At training time, each station calibrates its own settings (the imputer learns its fill values, the scaler learns its mean/std) using only the train parts. At serving time, every station applies its calibrated settings to the new part — no re-calibration.
Data leakage is when a downstream station peeks at the test parts during calibration. The pipeline physically prevents this — transform (not fit) is the only operation allowed at prediction time.
Analogy
Discussion prompt
Explain The assembly-line mental model by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Picture a factory assembly line: raw parts enter at one end, each station does exactly one job (impute → scale → encode → classify), and the finished product exits at the other end.
Estimation
Predict first
600 samples, 3 numeric features with 30 injected NaNs. Build an impute→scale pipeline and verify the scaler was fitted on training data only (Lesson 40 pattern).
Commit before you compute: what does Impute → Scale pipeline (numeric data) come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: imputer fill = 0.5336, scaler mean = 0.5336 — both computed from X_train only
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The imputed column's mean IS the scaler's input mean.
Worked example
600 samples, 3 numeric features with 30 injected NaNs. Build an impute→scale pipeline and verify the scaler was fitted on training data only (Lesson 40 pattern).
import numpy as np
from sklearn.pipeline import Pipeline
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import StandardScaler
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
rng = np.random.default_rng(42)
X, y = make_classification(n_samples=600, n_features=3, n_informative=3,
n_redundant=0, random_state=42)
X[rng.integers(0, 600, 30), 0] = np.nan # inject 30 NaNs in col 0
X_train, X_test = train_test_split(X, test_size=0.2, random_state=42)
num_pipe = Pipeline([('imp', SimpleImputer(strategy='mean')),
('scl', StandardScaler())])
num_pipe.fit(X_train)
train_mean = num_pipe.named_steps['imp'].statistics_[0]
scaler_mean = num_pipe.named_steps['scl'].mean_[0]
print(f'imputer fill = {train_mean:.4f}, scaler mean = {scaler_mean:.4f}')imputer fill = 0.5336, scaler mean = 0.5336 — both computed from X_train only
Why: The imputed column's mean IS the scaler's input mean. Both statistics come solely from training rows — the 120 test rows were never seen during fit.
| row | raw col0 | after impute | after scale |
|---|---|---|---|
| 0 | 0.3692 | 0.3692 | -0.1173 |
| 1 | 1.4882 | 1.4882 | 0.6812 |
| 2 | 0.2582 | 0.2582 | -0.1965 |
| 3 | 1.9805 | 1.9805 | 1.0325 |
| 4 | 2.0951 | 2.0951 | 1.1143 |
Comparison
Comparison matrix
From Impute → Scale pipeline (numeric data): refill the after impute column from what you know. The rest of the table is as it appeared.
| row | raw col0 | after impute | after scale |
|---|---|---|---|
| 0 | 0.3692 | 0.3692 | -0.1173 |
| 1 | 1.4882 | 1.4882 | 0.6812 |
| 2 | 0.2582 | 0.2582 | -0.1965 |
| 3 | 1.9805 | 1.9805 | 1.0325 |
| 4 | 2.0951 | 2.0951 | 1.1143 |
Anomaly
Predict first
A student writes this, and it looks reasonable:
Scale everything, THEN split — feels fine because the scaler is 'just normalization'.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: The scaler's mean and std now encode information from the test rows.
Split first, then put the scaler INSIDE the pipeline.
Why: The scaler's mean and std now encode information from the test rows. The test set is no longer held-out — metrics are optimistically biased.
Trap
Scale everything, THEN split — feels fine because the scaler is 'just normalization'.
scaler.fit_transform(X_all) → then train_test_split
Why: The scaler's mean and std now encode information from the test rows. The test set is no longer held-out — metrics are optimistically biased.
Split first, then put the scaler INSIDE the pipeline.
train_test_split → Pipeline([scaler, model]).fit(X_train)
Why: The pipeline calls scaler.fit only on X_train. X_test passes through transform (apply learned stats) — no leakage, honest metrics.
Break the constraint
Discussion prompt
The rule this trap just fixed:
The pipeline calls scaler.fit only on X_train. X_test passes through transform (apply learned stats) — no leakage, honest metrics.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
The scaler's mean and std now encode information from the test rows. The test set is no longer held-out — metrics are optimistically biased.
Section
Part 2 of 4
Concept
ColumnTransformer([(name, transformer, columns), ...]) applies a different sub-pipeline to each subset of columns and concatenates the results.
| column type | typical pipeline | output cols |
|---|---|---|
| numeric (0,1,2) | SimpleImputer → StandardScaler | 3 |
| categorical (3) | SimpleImputer → OneHotEncoder | 3 (one per category) |
| combined | hstack of both outputs | 6 total |
This keeps each transformation semantically correct: StandardScaler on text strings is nonsense; OneHotEncoder on floats is equally wrong. ColumnTransformer routes each feature type to the right transformer.
Trade off
Comparison matrix
From ColumnTransformer: every row here is a choice with a cost. Fill the output cols column, then say which row you would actually pick and what you give up for it.
| column type | typical pipeline | output cols |
|---|---|---|
| numeric (0,1,2) | SimpleImputer → StandardScaler | 3 |
| categorical (3) | SimpleImputer → OneHotEncoder | 3 (one per category) |
| combined | hstack of both outputs | 6 total |
Concept
FunctionTransformer(func) wraps any numpy-compatible function as a sklearn transformer, so it plugs into a Pipeline or ColumnTransformer.
Common uses: np.log1p for right-skewed features, np.sqrt for count data, or a custom clip. The transformer is stateless — fit is a no-op, transform applies func each time.
| input | log1p output |
|---|---|
| 1.2000 | 0.7885 |
| 0.5000 | 0.4055 |
| 2.7000 | 1.3083 |
| 0.1000 | 0.0953 |
| 3.8000 | 1.5686 |
Pattern
Step through it
Step through FunctionTransformer one row at a time. What is driving the change, and what would the row after the last one be?
Missing information
Discussion prompt
Mixed dataset: 3 numeric columns (with NaNs) + 1 categorical column ('low'/'mid'/'high'). Chain ColumnTransformer into LogisticRegression.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
The pipeline's single fit() propagates through every stage in order; predict() re-runs only transform (not fit) at each stage. All preprocessing statistics come from X_train only.
Worked example
Mixed dataset: 3 numeric columns (with NaNs) + 1 categorical column ('low'/'mid'/'high'). Chain ColumnTransformer into LogisticRegression.
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import OneHotEncoder
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score
numeric_pipe = Pipeline([('imp', SimpleImputer(strategy='mean')),
('scl', StandardScaler())])
cat_pipe = Pipeline([('imp', SimpleImputer(strategy='most_frequent')),
('enc', OneHotEncoder(handle_unknown='ignore',
sparse_output=False))])
prep = ColumnTransformer([('num', numeric_pipe, [0,1,2]),
('cat', cat_pipe, [3])])
clf_pipe = Pipeline([('prep', prep),
('clf', LogisticRegression(max_iter=500, random_state=0))])
clf_pipe.fit(X_train, y_train)
print(accuracy_score(y_test, clf_pipe.predict(X_test)))Test accuracy = 0.7333 — one fit call handles imputation, encoding, scaling, and classification
Why: The pipeline's single fit() propagates through every stage in order; predict() re-runs only transform (not fit) at each stage. All preprocessing statistics come from X_train only.
| stage | input shape | output shape | key operation |
|---|---|---|---|
| numeric branch | (480, 3) | (480, 3) | fill NaN → z-score |
| cat branch | (480, 1) | (480, 3) | fill mode → OHE |
| hstack | (480, 3)+(480, 3) | (480, 6) | concatenate |
| LogReg | (480, 6) | predictions | fit coefficients |
Discrimination
Sort into buckets
Sort these by output shape, from memory, without looking back at Full Pipeline: impute→encode→scale→LogReg. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Section
Part 3 of 4
Concept
Because Pipeline exposes every step by name, GridSearchCV can tune preprocessing hyperparameters and model hyperparameters in a single grid search.
Hyperparameter keys use double-underscore chaining: 'clf__C' sets C on the step named 'clf'; 'prep__num__imputer__strategy' drills into the 'num' sub-pipeline's 'imputer' step.
Crucially, each cross-validation fold fits the entire pipeline (including preprocessing) on its training split only — so the imputer and scaler cannot leak test-fold statistics into the validation score.
Explain it
Discussion prompt
Explain Tuning inside the pipeline to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Because Pipeline exposes every step by name, GridSearchCV can tune preprocessing hyperparameters and model hyperparameters in a single grid search.
Estimation
Predict first
Tune clf__C (LogReg regularization) AND prep__num__imputer__strategy jointly over 5-fold CV.
Commit before you compute: what does GridSearchCV over Pipeline come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: best params: C=0.1, strategy='mean'; best CV=0.7083; test accuracy=0.7333
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The grid explored 3×2=6 parameter combinations, each evaluated over 5 folds.
Worked example
Tune clf__C (LogReg regularization) AND prep__num__imputer__strategy jointly over 5-fold CV.
from sklearn.model_selection import GridSearchCV
param_grid = {
'clf__C': [0.1, 1.0, 10.0],
'prep__num__imputer__strategy': ['mean', 'median'],
}
gs = GridSearchCV(clf_pipe, param_grid, cv=5, scoring='accuracy', n_jobs=-1)
gs.fit(X_train, y_train)
print('best params :', gs.best_params_)
print('best CV :', round(gs.best_score_, 4))
print('test acc :', round(gs.score(X_test, y_test), 4))best params: C=0.1, strategy='mean'; best CV=0.7083; test accuracy=0.7333
Why: The grid explored 3×2=6 parameter combinations, each evaluated over 5 folds. The best estimator (refitted on full train) reaches 0.7333 on the held-out test set — same as the default because C=0.1 already regularizes gently enough for this data.
| rank | C | impute strategy | mean CV acc |
|---|---|---|---|
| 1 | 0.1 | mean | 0.7083 |
| 1 | 0.1 | median | 0.7083 |
| 3 | 1.0 | mean | 0.7021 |
Pattern
Step through it
Step through GridSearchCV over Pipeline one row at a time. What is driving the change, and what would the row after the last one be?
Anomaly
Predict first
A student writes this, and it looks reasonable:
Fit the scaler once on X_train, then pass the pre-scaled X_train to GridSearchCV — 'that's efficient: the scaler runs only once'.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Each CV fold's validation split was already transformed using all 480 training rows (including that fold's validation portion).
Pass the entire Pipeline to GridSearchCV so each fold re-fits preprocessing.
Why: Each CV fold's validation split was already transformed using all 480 training rows (including that fold's validation portion). The scaler statistics leak validation data into every fold — CV scores are optimistically biased.
Trap
Fit the scaler once on X_train, then pass the pre-scaled X_train to GridSearchCV — 'that's efficient: the scaler runs only once'.
scaler.fit_transform(X_train) → GridSearchCV(LogReg, ...).fit(X_scaled_train)
Why: Each CV fold's validation split was already transformed using all 480 training rows (including that fold's validation portion). The scaler statistics leak validation data into every fold — CV scores are optimistically biased.
Pass the entire Pipeline to GridSearchCV so each fold re-fits preprocessing.
GridSearchCV(Pipeline([scaler, model]), ...).fit(X_train)
Why: Now each fold fits the scaler on its 384-row train split and transforms its 96-row val split. No leakage, honest CV estimate. This is the only correct way to do preprocessing + hyperparam search.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
Pipeline([('name', step), ...]) wraps a sequence of transformers ending in an estimator. Calling fit(X_train, y) on the pipeline fits every step on training data only.; ColumnTransformer([(name, transformer, columns), ...]) applies a different sub-pipeline to each subset of columns and concatenates the results.; FunctionTransformer(func) wraps any numpy-compatible function as a sklearn transformer, so it plugs into a Pipeline or ColumnTransformer.Section
Part 4 of 4
Concept
joblib.dump(pipeline, 'model_v1.joblib') serializes the entire fitted object — imputer statistics, scaler parameters, model weights — to disk.
joblib.load('model_v1.joblib') deserializes it. The reloaded pipeline calls predict directly — no re-fit.
| tool | use case | note |
|---|---|---|
| joblib.dump / load | sklearn pipelines, numpy arrays | preferred for sklearn; faster than pickle for large arrays |
| torch.save / load | PyTorch models (Lesson 40) | saves state_dict or full module |
| pickle | general Python objects | slower for arrays; avoid for large models |
Explain it
Discussion prompt
Explain Serializing a fitted pipeline to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
joblib.dump(pipeline, 'model_v1.joblib') serializes the entire fitted object — imputer statistics, scaler parameters, model weights — to disk.
Concept
A versioned filename encodes the model's identity: pipeline_logreg_v1_2026-06.joblib. When you retrain on new data, bump the version — never overwrite the previous file.
{'model': pipe, 'sklearn_version': sklearn.__version__})Analogy
Discussion prompt
Explain Model versioning by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
A versioned filename encodes the model's identity: pipeline_logreg_v1_2026-06.joblib. When you retrain on new data, bump the version — never overwrite the previous file.
Estimation
Predict first
Serialize clf_pipe, reload it, and confirm the reloaded version produces identical predictions.
Commit before you compute: what does dump, load, verify come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: size=3354 bytes; original=0.7333; reloaded=0.7333; match=True
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. joblib stores the entire pipeline object graph — all fitted parameters are preserved exactly.
Worked example
Serialize clf_pipe, reload it, and confirm the reloaded version produces identical predictions.
import joblib, tempfile, os
tmp = tempfile.NamedTemporaryFile(suffix='.joblib', delete=False)
tmp.close()
joblib.dump(clf_pipe, tmp.name)
size = os.path.getsize(tmp.name)
loaded = joblib.load(tmp.name)
os.unlink(tmp.name)
original_acc = accuracy_score(y_test, clf_pipe.predict(X_test))
reloaded_acc = accuracy_score(y_test, loaded.predict(X_test))
print(f'size: {size} bytes')
print(f'original: {original_acc:.4f}')
print(f'reloaded: {reloaded_acc:.4f}')
print(f'match: {original_acc == reloaded_acc}')size=3354 bytes; original=0.7333; reloaded=0.7333; match=True
Why: joblib stores the entire pipeline object graph — all fitted parameters are preserved exactly. The reloaded pipeline produces bit-for-bit identical predictions without any re-fit.
| metric | original | reloaded |
|---|---|---|
| accuracy | 0.7333 | 0.7333 |
| predictions match | — | True |
| file size | — | 3 354 bytes |
Comparison
Comparison matrix
From dump, load, verify: refill the reloaded column from what you know. The rest of the table is as it appeared.
| metric | original | reloaded |
|---|---|---|
| accuracy | 0.7333 | 0.7333 |
| predictions match | — | True |
| file size | — | 3 354 bytes |
Ranking
Put in order
These are the steps of The production pipeline recipe, scrambled. Put them back in order before the next slide shows you.
train_test_split before any fittingfit(X_train) onceWhy: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
train_test_split before any fittingfit(X_train) onceEdge cases
Discussion prompt
The production pipeline recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
train_test_split before any fittingfit(X_train) onceElimination
Eliminate the wrong options
You call scaler.fit_transform(X_all) on the full 600-row dataset, then do train_test_split. Why is the test accuracy reported by model.score(X_test_scaled, y_test) optimistically biased?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: The scaler fitted on X_all used 120 test-set rows to compute its mean/std. The 'test' rows were implicitly used in preprocessing, so their statistics leak into what looks like a held-out evaluation — the metric is an overly optimistic estimate of true generalization.
Check
Think before clicking.
Check your understanding
You call scaler.fit_transform(X_all) on the full 600-row dataset, then do train_test_split. Why is the test accuracy reported by model.score(X_test_scaled, y_test) optimistically biased?
Answer: A
Why: The scaler fitted on X_all used 120 test-set rows to compute its mean/std. The 'test' rows were implicitly used in preprocessing, so their statistics leak into what looks like a held-out evaluation — the metric is an overly optimistic estimate of true generalization.
Prediction
Predict first
In GridSearchCV(Pipeline([('prep', scaler), ('clf', model)]), param_grid, cv=5), what happens inside each of the 5 CV folds?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: The scaler is re-fitted on the fold's train split; the fold's val split is only transformed
Why: GridSearchCV calls pipeline.fit(X_fold_train) for each fold, which triggers scaler.fit on that fold's training portion. The held-out val rows only pass through scaler.transform — exactly the leakage-free behavior you want.
Check
Reason about the plumbing.
Check your understanding
In GridSearchCV(Pipeline([('prep', scaler), ('clf', model)]), param_grid, cv=5), what happens inside each of the 5 CV folds?
Answer: A
Why: GridSearchCV calls pipeline.fit(X_fold_train) for each fold, which triggers scaler.fit on that fold's training portion. The held-out val rows only pass through scaler.transform — exactly the leakage-free behavior you want.
strategy or n_components.Elimination
Eliminate the wrong options
After joblib.dump(clf_pipe, 'v1.joblib'), you load it on a different machine. Which statement is true?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: joblib serializes the entire Python object graph — imputer statistics (fill values), scaler mean/std, and model coefficients. The reloaded pipeline calls predict() directly on new data with no re-fit. Our test confirmed accuracy 0.7333 before and after reload.
Check
What gets serialized?
Check your understanding
After joblib.dump(clf_pipe, 'v1.joblib'), you load it on a different machine. Which statement is true?
Answer: A
Why: joblib serializes the entire Python object graph — imputer statistics (fill values), scaler mean/std, and model coefficients. The reloaded pipeline calls predict() directly on new data with no re-fit. Our test confirmed accuracy 0.7333 before and after reload.
Section
Project
Concept
Build a complete sklearn Pipeline for a mixed-type dataset, tune it with GridSearchCV, evaluate it on a held-out test set, and serialize the winner to disk.
| # | requirement | tool |
|---|---|---|
| 1 | impute NaNs → scale numerics, impute mode → OHE categoricals | ColumnTransformer + Pipeline |
| 2 | add log1p step for one column | FunctionTransformer |
| 3 | tune C and imputer strategy jointly | GridSearchCV(pipeline, param_grid) |
| 4 | serialize winner, reload, verify accuracy matches | joblib.dump / load |
Build rule: split before any fitting. Every preprocessing step lives inside the pipeline passed to GridSearchCV. Never call fit on test data.
Counterexample
Discussion prompt
Build a complete sklearn Pipeline for a mixed-type dataset, tune it with GridSearchCV, evaluate it on a held-out test set, and serialize the winner to disk.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Build rule: split before any fitting. Every preprocessing step lives inside the pipeline passed to GridSearchCV. Never call fit on test data.
Worked example
Your turn: Build numeric_pipe = Pipeline([imputer, scaler]) and cat_pipe = Pipeline([imputer, OHE]). Wrap both in a ColumnTransformer. Predict the output shape after fit_transform(X_train).
Hint: 3 numeric cols → 3 scaled cols; 1 categorical with 3 levels → 3 OHE cols. Total output = 6 columns.
from sklearn.pipeline import Pipeline
from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import StandardScaler, OneHotEncoder
numeric_pipe = Pipeline([('imp', SimpleImputer(strategy='mean')),
('scl', StandardScaler())])
cat_pipe = Pipeline([('imp', SimpleImputer(strategy='most_frequent')),
('enc', OneHotEncoder(handle_unknown='ignore',
sparse_output=False))])
prep = ColumnTransformer([('num', numeric_pipe, [0,1,2]),
('cat', cat_pipe, [3])])
out = prep.fit_transform(X_train)
print(f'output shape: {out.shape}') # expect (480, 6)| branch | input cols | output cols |
|---|---|---|
| numeric | 0, 1, 2 | 3 (scaled) |
| categorical | 3 | 3 (OHE) |
| combined | — | 6 total |
Pattern
Step through it
Step through Milestone 1 — ColumnTransformer pipeline one row at a time. What is driving the change, and what would the row after the last one be?
Worked example
Your turn: wrap prep + LogReg in a Pipeline. Define a param_grid for clf__C and prep__num__imputer__strategy. Run 5-fold CV. Predict which C wins.
Hint: key syntax is 'stepname__paramname' (double underscore). Sub-pipeline drilling: 'prep__num__imp__strategy'.
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import GridSearchCV
clf_pipe = Pipeline([('prep', prep),
('clf', LogisticRegression(max_iter=500, random_state=0))])
param_grid = {
'clf__C': [0.1, 1.0, 10.0],
'prep__num__imp__strategy': ['mean', 'median'],
}
gs = GridSearchCV(clf_pipe, param_grid, cv=5, scoring='accuracy', n_jobs=-1)
gs.fit(X_train, y_train)
print(gs.best_params_)
print(round(gs.best_score_, 4), round(gs.score(X_test, y_test), 4))| best C | best strategy | CV acc | test acc |
|---|---|---|---|
| 0.1 | mean | 0.7083 | 0.7333 |
Worked example
Your turn: dump gs.best_estimator_ to a .joblib file, reload it, and verify the reloaded accuracy matches. Predict whether fit() is needed after loading.
Hint: joblib.dump(obj, path) / joblib.load(path). The best estimator is already refitted on full X_train by GridSearchCV (refit=True by default).
import joblib, tempfile, os
from sklearn.metrics import accuracy_score
path = tempfile.mktemp(suffix='.joblib')
joblib.dump(gs.best_estimator_, path)
loaded = joblib.load(path)
os.unlink(path)
acc_orig = accuracy_score(y_test, gs.best_estimator_.predict(X_test))
acc_reload = accuracy_score(y_test, loaded.predict(X_test))
print(f'original {acc_orig:.4f} reloaded {acc_reload:.4f} match {acc_orig==acc_reload}')| pipeline | accuracy | needs fit()? |
|---|---|---|
| best_estimator_ (original) | 0.7333 | no — already fitted |
| loaded from .joblib | 0.7333 | no — state preserved |
| match | True | — |
Trade off
Comparison matrix
From Milestone 3 — serialize and reload: every row here is a choice with a cost. Fill the accuracy column, then say which row you would actually pick and what you give up for it.
| pipeline | accuracy | needs fit()? |
|---|---|---|
| best_estimator_ (original) | 0.7333 | no — already fitted |
| loaded from .joblib | 0.7333 | no — state preserved |
| match | True | — |
Concept
import numpy as np, joblib, tempfile, os
from sklearn.pipeline import Pipeline
from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import GridSearchCV, train_test_split, make_classification
from sklearn.metrics import accuracy_score
rng = np.random.default_rng(42)
X, y = make_classification(n_samples=600, n_features=4, n_informative=3,
n_redundant=0, random_state=42)
X = X.astype(object)
X[rng.integers(0,600,30), 0] = np.nan
cats = np.array(['low','mid','high'])
X[:, 3] = cats[np.floor(np.abs(X[:, 3].astype(float))).astype(int) % 3]
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
prep = ColumnTransformer([
('num', Pipeline([('imp', SimpleImputer(strategy='mean')),
('scl', StandardScaler())]), [0,1,2]),
('cat', Pipeline([('imp', SimpleImputer(strategy='most_frequent')),
('enc', OneHotEncoder(handle_unknown='ignore',
sparse_output=False))]), [3])
])
pipe = Pipeline([('prep', prep),
('clf', LogisticRegression(max_iter=500, random_state=0))])
gs = GridSearchCV(pipe, {'clf__C': [0.1,1.0,10.0],
'prep__num__imp__strategy': ['mean','median']},
cv=5, scoring='accuracy', n_jobs=-1)
gs.fit(X_train, y_train)
path = tempfile.mktemp(suffix='.joblib')
joblib.dump(gs.best_estimator_, path)
loaded = joblib.load(path); os.unlink(path)
print('best params:', gs.best_params_)
print('test acc :', round(gs.score(X_test, y_test), 4))
print('reload acc :', round(accuracy_score(y_test, loaded.predict(X_test)), 4))| output | value |
|---|---|
| best params | C=0.1, strategy=mean |
| test acc | 0.7333 |
| reload acc | 0.7333 |
If your reloaded pipeline reproduces the test accuracy with no re-fit — you've built the production pattern every competition submission needs.
Comparison
Comparison matrix
From The full program: refill the value column from what you know. The rest of the table is as it appeared.
| output | value |
|---|---|
| best params | C=0.1, strategy=mean |
| test acc | 0.7333 |
| reload acc | 0.7333 |
Concept
Out loud, slides closed: explain (1) how Pipeline prevents leakage by splitting fit vs transform, (2) what double-underscore notation does in a GridSearchCV param_grid, and (3) what joblib.dump preserves.
Stretch (homework): add a FunctionTransformer(np.log1p) stage for a skewed numeric column, write a unit test that asserts the pipeline's output shape, and implement a version-tagged metadata dict alongside the joblib file. Next up: multiclass classification and learning rate schedules (Lesson 61).
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Pipeline: chain without leaking · ColumnTransformer: mix-typed features · GridSearchCV over the full pipeline · joblib: save, load, version · Your turn: production pipeline. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
Pipeline chains estimators and enforces no-leakage fit: preprocessing statistics come from training data onlyColumnTransformer routes numeric and categorical columns to different sub-pipelinesFunctionTransformer wraps any stateless function (e.g., np.log1p) into a pipeline stepGridSearchCV(pipeline, param_grid) tunes preprocessing + model hyperparameters per CV fold — the only correct wayjoblib.dump/load serializes the entire fitted pipeline; no re-fit needed on reload| concept | the one thing to remember |
|---|---|
| Pipeline | fit on train only — transform is all test ever sees |
| ColumnTransformer | routes by column index; output is hstacked |
| GridSearchCV + Pipeline | preprocessing re-fits per fold — no leakage in CV |
| joblib | serializes all fitted parameters; reload → predict directly |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.