Lesson 60: sklearn Pipeline, ColumnTransformer & GridSearchCV

USAAIO Lesson 60, from Week 21. An sklearn Pipeline chains estimators together and prevents leakage; ColumnTransformer applies different transforms to different feature types; and FunctionTransformer wraps any function. Combining a Pipeline with GridSearchCV lets you tune the preprocessing and the model hyperparameters jointly, and joblib saves and loads versioned pipelines. You build a production-ready pipeline that imputes, encodes, scales, and fits a logistic regression, tune it end to end, and serialize it. The lesson runs to 31 slides.

Subject: Machine Learning · 61 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. sklearn Pipeline, ColumnTransformer & GridSearchCV

Title

USAAIO · Lesson 60 · Week 21

Chain estimators, prevent leakage, tune everything end-to-end, and ship a versioned model file. The plumbing every competition submission needs.

2. By the end of this lesson you can

Objectives

  1. Explain why Pipeline prevents data leakage and chain fit/transform/predict into one object
  2. Build a ColumnTransformer that applies different preprocessing to numeric vs categorical columns
  3. Wrap any numpy function as a sklearn step with FunctionTransformer
  4. Pass a Pipeline to GridSearchCV and tune both preprocessing and model hyperparameters together
  5. Serialize and reload a fitted pipeline with joblib.dump / joblib.load

3. What survived from Hyperparameter Optimization?

Warm-up

Discussion prompt

Before we open Lesson 60: sklearn Pipeline, ColumnTransformer & GridSearchCV: without looking back, what was the main idea of Hyperparameter Optimization, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

grid search (exhaustive, exponential cost), random search (log-scale sampling, matches or beats grid at a fraction of the evaluations), Gaussian Process Bayesian optimization (surrogate + expected-improvement acquisition), successive halving and Hyperband (early elimination), and practical log-scale tips. All numbers verified with sklearn 1.x on digits, June 2026.

4. Pipeline: chain without leaking

Section

Part 1 of 4

5. What Pipeline does

Concept

Pipeline([('name', step), ...]) wraps a sequence of transformers ending in an estimator. Calling fit(X_train, y) on the pipeline fits every step on training data only.

Without a pipeline it is easy to accidentally call scaler.fit(X_all) before the split — fitting on test data poisons every downstream metric. The pipeline's fit enforces the split boundary.

callwhat happens internally
pipe.fit(X_train)each step: fit on train → transform train → pass to next
pipe.predict(X_test)each step: transform only (no re-fit) → final estimator predicts
pipe.score(X_test, y_test)transform + predict + metric in one line

6. Break it if you can: What Pipeline does

Counterexample

Discussion prompt

Pipeline([('name', step), ...]) wraps a sequence of transformers ending in an estimator. Calling fit(X_train, y) on the pipeline fits every step on training data only.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Without a pipeline it is easy to accidentally call scaler.fit(X_all) before the split — fitting on test data poisons every downstream metric. The pipeline's fit enforces the split boundary.

7. The assembly-line mental model

Intuition

Picture a factory assembly line: raw parts enter at one end, each station does exactly one job (impute → scale → encode → classify), and the finished product exits at the other end.

At training time, each station calibrates its own settings (the imputer learns its fill values, the scaler learns its mean/std) using only the train parts. At serving time, every station applies its calibrated settings to the new part — no re-calibration.

Data leakage is when a downstream station peeks at the test parts during calibration. The pipeline physically prevents this — transform (not fit) is the only operation allowed at prediction time.

8. By analogy: The assembly-line mental model

Analogy

Discussion prompt

Explain The assembly-line mental model by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Picture a factory assembly line: raw parts enter at one end, each station does exactly one job (impute → scale → encode → classify), and the finished product exits at the other end.

9. Guess the shape of the answer: Impute → Scale pipeline (numeric data)

Estimation

Predict first

600 samples, 3 numeric features with 30 injected NaNs. Build an impute→scale pipeline and verify the scaler was fitted on training data only (Lesson 40 pattern).

Commit before you compute: what does Impute → Scale pipeline (numeric data) come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: imputer fill = 0.5336, scaler mean = 0.5336 — both computed from X_train only

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The imputed column's mean IS the scaler's input mean.

10. Impute → Scale pipeline (numeric data)

Worked example

600 samples, 3 numeric features with 30 injected NaNs. Build an impute→scale pipeline and verify the scaler was fitted on training data only (Lesson 40 pattern).

import numpy as np
from sklearn.pipeline import Pipeline
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import StandardScaler
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split

rng = np.random.default_rng(42)
X, y = make_classification(n_samples=600, n_features=3, n_informative=3,
                            n_redundant=0, random_state=42)
X[rng.integers(0, 600, 30), 0] = np.nan          # inject 30 NaNs in col 0
X_train, X_test = train_test_split(X, test_size=0.2, random_state=42)

num_pipe = Pipeline([('imp', SimpleImputer(strategy='mean')),
                     ('scl', StandardScaler())])
num_pipe.fit(X_train)
train_mean = num_pipe.named_steps['imp'].statistics_[0]
scaler_mean = num_pipe.named_steps['scl'].mean_[0]
print(f'imputer fill = {train_mean:.4f}, scaler mean = {scaler_mean:.4f}')

imputer fill = 0.5336, scaler mean = 0.5336 — both computed from X_train only

Why: The imputed column's mean IS the scaler's input mean. Both statistics come solely from training rows — the 120 test rows were never seen during fit.

rowraw col0after imputeafter scale
00.36920.3692-0.1173
11.48821.48820.6812
20.25820.2582-0.1965
31.98051.98051.0325
42.09512.09511.1143

11. Fill in: after impute for Impute → Scale pipeline (numeric data)

Comparison

Comparison matrix

From Impute → Scale pipeline (numeric data): refill the after impute column from what you know. The rest of the table is as it appeared.

rowraw col0after imputeafter scale
00.36920.3692-0.1173
11.48821.48820.6812
20.25820.2582-0.1965
31.98051.98051.0325
42.09512.09511.1143

12. Something is wrong here: fit the scaler before the split

Anomaly

Predict first

A student writes this, and it looks reasonable:

Scale everything, THEN split — feels fine because the scaler is 'just normalization'.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: The scaler's mean and std now encode information from the test rows.

Split first, then put the scaler INSIDE the pipeline.

Why: The scaler's mean and std now encode information from the test rows. The test set is no longer held-out — metrics are optimistically biased.

13. Trap: fit the scaler before the split

Trap

The trap

Scale everything, THEN split — feels fine because the scaler is 'just normalization'.

scaler.fit_transform(X_all) → then train_test_split

Why: The scaler's mean and std now encode information from the test rows. The test set is no longer held-out — metrics are optimistically biased.

The fix

Split first, then put the scaler INSIDE the pipeline.

train_test_split → Pipeline([scaler, model]).fit(X_train)

Why: The pipeline calls scaler.fit only on X_train. X_test passes through transform (apply learned stats) — no leakage, honest metrics.

14. Break it on purpose: fit the scaler before the split

Break the constraint

Discussion prompt

The rule this trap just fixed:

The pipeline calls scaler.fit only on X_train. X_test passes through transform (apply learned stats) — no leakage, honest metrics.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

The scaler's mean and std now encode information from the test rows. The test set is no longer held-out — metrics are optimistically biased.

15. ColumnTransformer: mix-typed features

Section

Part 2 of 4

16. ColumnTransformer

Concept

ColumnTransformer([(name, transformer, columns), ...]) applies a different sub-pipeline to each subset of columns and concatenates the results.

column typetypical pipelineoutput cols
numeric (0,1,2)SimpleImputer → StandardScaler3
categorical (3)SimpleImputer → OneHotEncoder3 (one per category)
combinedhstack of both outputs6 total

This keeps each transformation semantically correct: StandardScaler on text strings is nonsense; OneHotEncoder on floats is equally wrong. ColumnTransformer routes each feature type to the right transformer.

17. What each one costs: ColumnTransformer

Trade off

Comparison matrix

From ColumnTransformer: every row here is a choice with a cost. Fill the output cols column, then say which row you would actually pick and what you give up for it.

column typetypical pipelineoutput cols
numeric (0,1,2)SimpleImputer → StandardScaler3
categorical (3)SimpleImputer → OneHotEncoder3 (one per category)
combinedhstack of both outputs6 total

18. FunctionTransformer

Concept

FunctionTransformer(func) wraps any numpy-compatible function as a sklearn transformer, so it plugs into a Pipeline or ColumnTransformer.

Common uses: np.log1p for right-skewed features, np.sqrt for count data, or a custom clip. The transformer is stateless — fit is a no-op, transform applies func each time.

inputlog1p output
1.20000.7885
0.50000.4055
2.70001.3083
0.10000.0953
3.80001.5686

19. Watch it run: FunctionTransformer

Pattern

Step through it

Step through FunctionTransformer one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: input is 1.2000
  2. Step 2: input is 0.5000
  3. Step 3: input is 2.7000
  4. Step 4: input is 0.1000
  5. Step 5: input is 3.8000

20. What has to be given first: Full Pipeline: impute→encode→scale→LogReg

Missing information

Discussion prompt

Mixed dataset: 3 numeric columns (with NaNs) + 1 categorical column ('low'/'mid'/'high'). Chain ColumnTransformer into LogisticRegression.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

The pipeline's single fit() propagates through every stage in order; predict() re-runs only transform (not fit) at each stage. All preprocessing statistics come from X_train only.

21. Full Pipeline: impute→encode→scale→LogReg

Worked example

Mixed dataset: 3 numeric columns (with NaNs) + 1 categorical column ('low'/'mid'/'high'). Chain ColumnTransformer into LogisticRegression.

from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import OneHotEncoder
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score

numeric_pipe = Pipeline([('imp', SimpleImputer(strategy='mean')),
                          ('scl', StandardScaler())])
cat_pipe = Pipeline([('imp', SimpleImputer(strategy='most_frequent')),
                     ('enc', OneHotEncoder(handle_unknown='ignore',
                                           sparse_output=False))])
prep = ColumnTransformer([('num', numeric_pipe, [0,1,2]),
                           ('cat', cat_pipe,    [3])])
clf_pipe = Pipeline([('prep', prep),
                     ('clf', LogisticRegression(max_iter=500, random_state=0))])
clf_pipe.fit(X_train, y_train)
print(accuracy_score(y_test, clf_pipe.predict(X_test)))

Test accuracy = 0.7333 — one fit call handles imputation, encoding, scaling, and classification

Why: The pipeline's single fit() propagates through every stage in order; predict() re-runs only transform (not fit) at each stage. All preprocessing statistics come from X_train only.

stageinput shapeoutput shapekey operation
numeric branch(480, 3)(480, 3)fill NaN → z-score
cat branch(480, 1)(480, 3)fill mode → OHE
hstack(480, 3)+(480, 3)(480, 6)concatenate
LogReg(480, 6)predictionsfit coefficients

22. Which is which, by output shape

Discrimination

Sort into buckets

Sort these by output shape, from memory, without looking back at Full Pipeline: impute→encode→scale→LogReg. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

(480, 3)
numeric branch; cat branch
(480, 6)
hstack
predictions
LogReg
g1
output shape is "(480, 3)" for numeric branch, cat branch — that is what the table on "Full Pipeline: impute→encode→scale→LogRe…" records, and it is the single property separating this group from the rest.
g2
output shape is "(480, 6)" for hstack — that is what the table on "Full Pipeline: impute→encode→scale→LogRe…" records, and it is the single property separating this group from the rest.
g3
output shape is "predictions" for LogReg — that is what the table on "Full Pipeline: impute→encode→scale→LogRe…" records, and it is the single property separating this group from the rest.

23. GridSearchCV over the full pipeline

Section

Part 3 of 4

24. Tuning inside the pipeline

Concept

Because Pipeline exposes every step by name, GridSearchCV can tune preprocessing hyperparameters and model hyperparameters in a single grid search.

Hyperparameter keys use double-underscore chaining: 'clf__C' sets C on the step named 'clf'; 'prep__num__imputer__strategy' drills into the 'num' sub-pipeline's 'imputer' step.

Crucially, each cross-validation fold fits the entire pipeline (including preprocessing) on its training split only — so the imputer and scaler cannot leak test-fold statistics into the validation score.

25. Teach it back: Tuning inside the pipeline

Explain it

Discussion prompt

Explain Tuning inside the pipeline to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Because Pipeline exposes every step by name, GridSearchCV can tune preprocessing hyperparameters and model hyperparameters in a single grid search.

26. Guess the shape of the answer: GridSearchCV over Pipeline

Estimation

Predict first

Tune clf__C (LogReg regularization) AND prep__num__imputer__strategy jointly over 5-fold CV.

Commit before you compute: what does GridSearchCV over Pipeline come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: best params: C=0.1, strategy='mean'; best CV=0.7083; test accuracy=0.7333

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The grid explored 3×2=6 parameter combinations, each evaluated over 5 folds.

27. GridSearchCV over Pipeline

Worked example

Tune clf__C (LogReg regularization) AND prep__num__imputer__strategy jointly over 5-fold CV.

from sklearn.model_selection import GridSearchCV

param_grid = {
    'clf__C':                         [0.1, 1.0, 10.0],
    'prep__num__imputer__strategy':   ['mean', 'median'],
}
gs = GridSearchCV(clf_pipe, param_grid, cv=5, scoring='accuracy', n_jobs=-1)
gs.fit(X_train, y_train)
print('best params :', gs.best_params_)
print('best CV    :', round(gs.best_score_, 4))
print('test acc   :', round(gs.score(X_test, y_test), 4))

best params: C=0.1, strategy='mean'; best CV=0.7083; test accuracy=0.7333

Why: The grid explored 3×2=6 parameter combinations, each evaluated over 5 folds. The best estimator (refitted on full train) reaches 0.7333 on the held-out test set — same as the default because C=0.1 already regularizes gently enough for this data.

rankCimpute strategymean CV acc
10.1mean0.7083
10.1median0.7083
31.0mean0.7021

28. Watch it run: GridSearchCV over Pipeline

Pattern

Step through it

Step through GridSearchCV over Pipeline one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: rank is 1
  2. Step 2: rank is 1
  3. Step 3: rank is 3

29. Something is wrong here: fitting preprocessing outside GridSearchCV

Anomaly

Predict first

A student writes this, and it looks reasonable:

Fit the scaler once on X_train, then pass the pre-scaled X_train to GridSearchCV — 'that's efficient: the scaler runs only once'.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Each CV fold's validation split was already transformed using all 480 training rows (including that fold's validation portion).

Pass the entire Pipeline to GridSearchCV so each fold re-fits preprocessing.

Why: Each CV fold's validation split was already transformed using all 480 training rows (including that fold's validation portion). The scaler statistics leak validation data into every fold — CV scores are optimistically biased.

30. Trap: fitting preprocessing outside GridSearchCV

Trap

The trap

Fit the scaler once on X_train, then pass the pre-scaled X_train to GridSearchCV — 'that's efficient: the scaler runs only once'.

scaler.fit_transform(X_train) → GridSearchCV(LogReg, ...).fit(X_scaled_train)

Why: Each CV fold's validation split was already transformed using all 480 training rows (including that fold's validation portion). The scaler statistics leak validation data into every fold — CV scores are optimistically biased.

The fix

Pass the entire Pipeline to GridSearchCV so each fold re-fits preprocessing.

GridSearchCV(Pipeline([scaler, model]), ...).fit(X_train)

Why: Now each fold fits the scaler on its 384-row train split and transforms its 96-row val split. No leakage, honest CV estimate. This is the only correct way to do preprocessing + hyperparam search.

31. Which of these survive contact with Lesson 60: sklearn Pipeline…?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Pipeline([('name', step), ...]) wraps a sequence of transformers ending in an estimator. Calling fit(X_train, y) on the pipeline fits every step on training data only.; ColumnTransformer([(name, transformer, columns), ...]) applies a different sub-pipeline to each subset of columns and concatenates the results.; FunctionTransformer(func) wraps any numpy-compatible function as a sklearn transformer, so it plugs into a Pipeline or ColumnTransformer.
Breaks
Scale everything, THEN split — feels fine because the scaler is 'just normalization'.; Fit the scaler once on X_train, then pass the pre-scaled X_train to GridSearchCV — 'that's efficient: the scaler runs only once'.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 60: sklearn Pipeline, ColumnTransformer & GridSearchCV puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

32. joblib: save, load, version

Section

Part 4 of 4

33. Serializing a fitted pipeline

Concept

joblib.dump(pipeline, 'model_v1.joblib') serializes the entire fitted object — imputer statistics, scaler parameters, model weights — to disk.

joblib.load('model_v1.joblib') deserializes it. The reloaded pipeline calls predict directly — no re-fit.

tooluse casenote
joblib.dump / loadsklearn pipelines, numpy arrayspreferred for sklearn; faster than pickle for large arrays
torch.save / loadPyTorch models (Lesson 40)saves state_dict or full module
picklegeneral Python objectsslower for arrays; avoid for large models

34. Teach it back: Serializing a fitted pipeline

Explain it

Discussion prompt

Explain Serializing a fitted pipeline to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

joblib.dump(pipeline, 'model_v1.joblib') serializes the entire fitted object — imputer statistics, scaler parameters, model weights — to disk.

35. Model versioning

Concept

A versioned filename encodes the model's identity: pipeline_logreg_v1_2026-06.joblib. When you retrain on new data, bump the version — never overwrite the previous file.

36. By analogy: Model versioning

Analogy

Discussion prompt

Explain Model versioning by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

A versioned filename encodes the model's identity: pipeline_logreg_v1_2026-06.joblib. When you retrain on new data, bump the version — never overwrite the previous file.

37. Guess the shape of the answer: dump, load, verify

Estimation

Predict first

Serialize clf_pipe, reload it, and confirm the reloaded version produces identical predictions.

Commit before you compute: what does dump, load, verify come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: size=3354 bytes; original=0.7333; reloaded=0.7333; match=True

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. joblib stores the entire pipeline object graph — all fitted parameters are preserved exactly.

38. dump, load, verify

Worked example

Serialize clf_pipe, reload it, and confirm the reloaded version produces identical predictions.

import joblib, tempfile, os

tmp = tempfile.NamedTemporaryFile(suffix='.joblib', delete=False)
tmp.close()
joblib.dump(clf_pipe, tmp.name)
size = os.path.getsize(tmp.name)
loaded = joblib.load(tmp.name)
os.unlink(tmp.name)

original_acc = accuracy_score(y_test, clf_pipe.predict(X_test))
reloaded_acc = accuracy_score(y_test, loaded.predict(X_test))
print(f'size:     {size} bytes')
print(f'original: {original_acc:.4f}')
print(f'reloaded: {reloaded_acc:.4f}')
print(f'match:    {original_acc == reloaded_acc}')

size=3354 bytes; original=0.7333; reloaded=0.7333; match=True

Why: joblib stores the entire pipeline object graph — all fitted parameters are preserved exactly. The reloaded pipeline produces bit-for-bit identical predictions without any re-fit.

metricoriginalreloaded
accuracy0.73330.7333
predictions match—True
file size—3 354 bytes

39. Fill in: reloaded for dump, load, verify

Comparison

Comparison matrix

From dump, load, verify: refill the reloaded column from what you know. The rest of the table is as it appeared.

metricoriginalreloaded
accuracy0.73330.7333
predictions match—True
file size—3 354 bytes

40. Rebuild the recipe: The production pipeline recipe

Ranking

Put in order

These are the steps of The production pipeline recipe, scrambled. Put them back in order before the next slide shows you.

  1. Split first: train_test_split before any fitting
  2. ColumnTransformer routes each column type to its own sub-pipeline
  3. Pipeline chains ColumnTransformer → model; call fit(X_train) once
  4. GridSearchCV(Pipeline, param_grid) tunes preprocessing + model in one search — preprocessing re-fits per fold
  5. joblib.dump saves the fitted pipeline; include version + sklearn version in filename or metadata

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

41. The production pipeline recipe

Pattern

  1. Split first: train_test_split before any fitting
  2. ColumnTransformer routes each column type to its own sub-pipeline
  3. Pipeline chains ColumnTransformer → model; call fit(X_train) once
  4. GridSearchCV(Pipeline, param_grid) tunes preprocessing + model in one search — preprocessing re-fits per fold
  5. joblib.dump saves the fitted pipeline; include version + sklearn version in filename or metadata

42. Where does it stop working: The production pipeline recipe

Edge cases

Discussion prompt

The production pipeline recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Split first: train_test_split before any fitting
  2. ColumnTransformer routes each column type to its own sub-pipeline
  3. Pipeline chains ColumnTransformer → model; call fit(X_train) once
  4. GridSearchCV(Pipeline, param_grid) tunes preprocessing + model in one search — preprocessing re-fits per fold
  5. joblib.dump saves the fitted pipeline; include version + sklearn version in filename or metadata

43. Rule out three: Check yourself — leakage

Elimination

Eliminate the wrong options

You call scaler.fit_transform(X_all) on the full 600-row dataset, then do train_test_split. Why is the test accuracy reported by model.score(X_test_scaled, y_test) optimistically biased?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. The scaler's mean and std were computed from test rows, so X_test was already 'seen' during preprocessing
  • B. StandardScaler changes the label distribution, inflating accuracy
  • C. fit_transform returns a different dtype than transform, confusing the model
  • D. The bias comes from using too many features, not the split order

Survives elimination: A

Why: The scaler fitted on X_all used 120 test-set rows to compute its mean/std. The 'test' rows were implicitly used in preprocessing, so their statistics leak into what looks like a held-out evaluation — the metric is an overly optimistic estimate of true generalization.

44. Check yourself — leakage

Check

Think before clicking.

Check your understanding

You call scaler.fit_transform(X_all) on the full 600-row dataset, then do train_test_split. Why is the test accuracy reported by model.score(X_test_scaled, y_test) optimistically biased?

  • A. The scaler's mean and std were computed from test rows, so X_test was already 'seen' during preprocessing (correct)
  • B. StandardScaler changes the label distribution, inflating accuracy
  • C. fit_transform returns a different dtype than transform, confusing the model
  • D. The bias comes from using too many features, not the split order

Answer: A

Why: The scaler fitted on X_all used 120 test-set rows to compute its mean/std. The 'test' rows were implicitly used in preprocessing, so their statistics leak into what looks like a held-out evaluation — the metric is an overly optimistic estimate of true generalization.

Why B tempts people
StandardScaler is a feature transform; it never touches labels. Label distribution is unaffected.
Why C tempts people
Both fit_transform and transform return the same numpy float64 array — dtype is not the issue.
Why D tempts people
Leakage is about data flow across the train/test boundary, not feature count. Even a single-feature model leaks if the scaler was fitted on all data.

45. Answer it before you see the options: Check yourself — GridSearchCV + Pipeline

Prediction

Predict first

In GridSearchCV(Pipeline([('prep', scaler), ('clf', model)]), param_grid, cv=5), what happens inside each of the 5 CV folds?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: The scaler is re-fitted on the fold's train split; the fold's val split is only transformed

Why: GridSearchCV calls pipeline.fit(X_fold_train) for each fold, which triggers scaler.fit on that fold's training portion. The held-out val rows only pass through scaler.transform — exactly the leakage-free behavior you want.

46. Check yourself — GridSearchCV + Pipeline

Check

Reason about the plumbing.

Check your understanding

In GridSearchCV(Pipeline([('prep', scaler), ('clf', model)]), param_grid, cv=5), what happens inside each of the 5 CV folds?

  • A. The scaler is re-fitted on the fold's train split; the fold's val split is only transformed (correct)
  • B. The scaler is fitted once on all of X_train and then reused across all folds
  • C. The scaler is not called inside GridSearchCV — it runs before the loop
  • D. GridSearchCV removes preprocessing steps and only tunes the final estimator

Answer: A

Why: GridSearchCV calls pipeline.fit(X_fold_train) for each fold, which triggers scaler.fit on that fold's training portion. The held-out val rows only pass through scaler.transform — exactly the leakage-free behavior you want.

Why B tempts people
This is the leakage bug GridSearchCV+Pipeline avoids. Each fold gets its own scaler fit — not a shared one.
Why C tempts people
If preprocessing ran before the loop it would leak validation data into every fold's scaler statistics.
Why D tempts people
GridSearchCV tunes ALL named parameters in the pipeline, including preprocessing hyperparameters like strategy or n_components.

47. Rule out three: Check yourself — joblib

Elimination

Eliminate the wrong options

After joblib.dump(clf_pipe, 'v1.joblib'), you load it on a different machine. Which statement is true?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. The reloaded pipeline contains all fitted parameters; calling predict() needs no X_train
  • B. You must call fit() again after loading because joblib only stores the pipeline structure
  • C. joblib stores only the model weights, not the scaler statistics — you must re-fit the scaler
  • D. The loaded pipeline's accuracy will be higher because joblib optimizes the coefficients during serialization

Survives elimination: A

Why: joblib serializes the entire Python object graph — imputer statistics (fill values), scaler mean/std, and model coefficients. The reloaded pipeline calls predict() directly on new data with no re-fit. Our test confirmed accuracy 0.7333 before and after reload.

48. Check yourself — joblib

Check

What gets serialized?

Check your understanding

After joblib.dump(clf_pipe, 'v1.joblib'), you load it on a different machine. Which statement is true?

  • A. The reloaded pipeline contains all fitted parameters; calling predict() needs no X_train (correct)
  • B. You must call fit() again after loading because joblib only stores the pipeline structure
  • C. joblib stores only the model weights, not the scaler statistics — you must re-fit the scaler
  • D. The loaded pipeline's accuracy will be higher because joblib optimizes the coefficients during serialization

Answer: A

Why: joblib serializes the entire Python object graph — imputer statistics (fill values), scaler mean/std, and model coefficients. The reloaded pipeline calls predict() directly on new data with no re-fit. Our test confirmed accuracy 0.7333 before and after reload.

Why B tempts people
This confuses joblib with saving only a model architecture (as in some deep learning frameworks that save state_dict separately from layer definitions). joblib always saves the full fitted object.
Why C tempts people
joblib serializes ALL attributes, including preprocessing statistics. The scaler's mean_ and scale_ arrays are part of the dumped object.
Why D tempts people
joblib is a serialization library, not an optimizer. It faithfully stores and restores the object state — no transformation of coefficients occurs.

49. Your turn: production pipeline

Section

Project

50. Project: end-to-end production ML system

Concept

Build a complete sklearn Pipeline for a mixed-type dataset, tune it with GridSearchCV, evaluate it on a held-out test set, and serialize the winner to disk.

#requirementtool
1impute NaNs → scale numerics, impute mode → OHE categoricalsColumnTransformer + Pipeline
2add log1p step for one columnFunctionTransformer
3tune C and imputer strategy jointlyGridSearchCV(pipeline, param_grid)
4serialize winner, reload, verify accuracy matchesjoblib.dump / load

Build rule: split before any fitting. Every preprocessing step lives inside the pipeline passed to GridSearchCV. Never call fit on test data.

51. Break it if you can: Project: end-to-end production ML system

Counterexample

Discussion prompt

Build a complete sklearn Pipeline for a mixed-type dataset, tune it with GridSearchCV, evaluate it on a held-out test set, and serialize the winner to disk.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Build rule: split before any fitting. Every preprocessing step lives inside the pipeline passed to GridSearchCV. Never call fit on test data.

52. Milestone 1 — ColumnTransformer pipeline

Worked example

Your turn: Build numeric_pipe = Pipeline([imputer, scaler]) and cat_pipe = Pipeline([imputer, OHE]). Wrap both in a ColumnTransformer. Predict the output shape after fit_transform(X_train).

Hint: 3 numeric cols → 3 scaled cols; 1 categorical with 3 levels → 3 OHE cols. Total output = 6 columns.

from sklearn.pipeline import Pipeline
from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import StandardScaler, OneHotEncoder

numeric_pipe = Pipeline([('imp', SimpleImputer(strategy='mean')),
                          ('scl', StandardScaler())])
cat_pipe = Pipeline([('imp', SimpleImputer(strategy='most_frequent')),
                     ('enc', OneHotEncoder(handle_unknown='ignore',
                                           sparse_output=False))])
prep = ColumnTransformer([('num', numeric_pipe, [0,1,2]),
                           ('cat', cat_pipe,    [3])])
out = prep.fit_transform(X_train)
print(f'output shape: {out.shape}')  # expect (480, 6)
branchinput colsoutput cols
numeric0, 1, 23 (scaled)
categorical33 (OHE)
combined—6 total

53. Watch it run: Milestone 1 — ColumnTransformer pipeline

Pattern

Step through it

Step through Milestone 1 — ColumnTransformer pipeline one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: branch is numeric
  2. Step 2: branch is categorical
  3. Step 3: branch is combined

54. Milestone 2 — GridSearchCV over the pipeline

Worked example

Your turn: wrap prep + LogReg in a Pipeline. Define a param_grid for clf__C and prep__num__imputer__strategy. Run 5-fold CV. Predict which C wins.

Hint: key syntax is 'stepname__paramname' (double underscore). Sub-pipeline drilling: 'prep__num__imp__strategy'.

from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import GridSearchCV

clf_pipe = Pipeline([('prep', prep),
                     ('clf', LogisticRegression(max_iter=500, random_state=0))])
param_grid = {
    'clf__C':                        [0.1, 1.0, 10.0],
    'prep__num__imp__strategy':      ['mean', 'median'],
}
gs = GridSearchCV(clf_pipe, param_grid, cv=5, scoring='accuracy', n_jobs=-1)
gs.fit(X_train, y_train)
print(gs.best_params_)
print(round(gs.best_score_, 4), round(gs.score(X_test, y_test), 4))
best Cbest strategyCV acctest acc
0.1mean0.70830.7333

55. Milestone 3 — serialize and reload

Worked example

Your turn: dump gs.best_estimator_ to a .joblib file, reload it, and verify the reloaded accuracy matches. Predict whether fit() is needed after loading.

Hint: joblib.dump(obj, path) / joblib.load(path). The best estimator is already refitted on full X_train by GridSearchCV (refit=True by default).

import joblib, tempfile, os
from sklearn.metrics import accuracy_score

path = tempfile.mktemp(suffix='.joblib')
joblib.dump(gs.best_estimator_, path)
loaded = joblib.load(path)
os.unlink(path)
acc_orig   = accuracy_score(y_test, gs.best_estimator_.predict(X_test))
acc_reload = accuracy_score(y_test, loaded.predict(X_test))
print(f'original {acc_orig:.4f}  reloaded {acc_reload:.4f}  match {acc_orig==acc_reload}')
pipelineaccuracyneeds fit()?
best_estimator_ (original)0.7333no — already fitted
loaded from .joblib0.7333no — state preserved
matchTrue—

56. What each one costs: Milestone 3 — serialize and reload

Trade off

Comparison matrix

From Milestone 3 — serialize and reload: every row here is a choice with a cost. Fill the accuracy column, then say which row you would actually pick and what you give up for it.

pipelineaccuracyneeds fit()?
best_estimator_ (original)0.7333no — already fitted
loaded from .joblib0.7333no — state preserved
matchTrue—

57. The full program

Concept

import numpy as np, joblib, tempfile, os
from sklearn.pipeline import Pipeline
from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import GridSearchCV, train_test_split, make_classification
from sklearn.metrics import accuracy_score

rng = np.random.default_rng(42)
X, y = make_classification(n_samples=600, n_features=4, n_informative=3,
                            n_redundant=0, random_state=42)
X = X.astype(object)
X[rng.integers(0,600,30), 0] = np.nan
cats = np.array(['low','mid','high'])
X[:, 3] = cats[np.floor(np.abs(X[:, 3].astype(float))).astype(int) % 3]
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

prep = ColumnTransformer([
    ('num', Pipeline([('imp', SimpleImputer(strategy='mean')),
                      ('scl', StandardScaler())]), [0,1,2]),
    ('cat', Pipeline([('imp', SimpleImputer(strategy='most_frequent')),
                      ('enc', OneHotEncoder(handle_unknown='ignore',
                                            sparse_output=False))]), [3])
])
pipe = Pipeline([('prep', prep),
                 ('clf', LogisticRegression(max_iter=500, random_state=0))])
gs = GridSearchCV(pipe, {'clf__C': [0.1,1.0,10.0],
                         'prep__num__imp__strategy': ['mean','median']},
                  cv=5, scoring='accuracy', n_jobs=-1)
gs.fit(X_train, y_train)
path = tempfile.mktemp(suffix='.joblib')
joblib.dump(gs.best_estimator_, path)
loaded = joblib.load(path); os.unlink(path)
print('best params:', gs.best_params_)
print('test acc   :', round(gs.score(X_test, y_test), 4))
print('reload acc :', round(accuracy_score(y_test, loaded.predict(X_test)), 4))
outputvalue
best paramsC=0.1, strategy=mean
test acc0.7333
reload acc0.7333

If your reloaded pipeline reproduces the test accuracy with no re-fit — you've built the production pattern every competition submission needs.

58. Fill in: value for The full program

Comparison

Comparison matrix

From The full program: refill the value column from what you know. The rest of the table is as it appeared.

outputvalue
best paramsC=0.1, strategy=mean
test acc0.7333
reload acc0.7333

59. Show it off

Concept

Out loud, slides closed: explain (1) how Pipeline prevents leakage by splitting fit vs transform, (2) what double-underscore notation does in a GridSearchCV param_grid, and (3) what joblib.dump preserves.

Stretch (homework): add a FunctionTransformer(np.log1p) stage for a skewed numeric column, write a unit test that asserts the pipeline's output shape, and implement a version-tagged metadata dict alongside the joblib file. Next up: multiclass classification and learning rate schedules (Lesson 61).

60. Connect it up: Lesson 60: sklearn Pipeline, ColumnTransformer & GridSearchCV

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Pipeline: chain without leaking · ColumnTransformer: mix-typed features · GridSearchCV over the full pipeline · joblib: save, load, version · Your turn: production pipeline. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

61. What you can do now

Recap

conceptthe one thing to remember
Pipelinefit on train only — transform is all test ever sees
ColumnTransformerroutes by column index; output is hstacked
GridSearchCV + Pipelinepreprocessing re-fits per fold — no leakage in CV
joblibserializes all fitted parameters; reload → predict directly

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 60 (Week 21 — Multiclass + Pipelines) — Barron · USAAIO Round 2 Preparation, 2026
  2. sklearn Pipeline, ColumnTransformer, GridSearchCV, and joblib — all snippets run and output verified — scikit-learn 1.x + numpy 2.2.6, real execution, June 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108