Lesson 59: Hyperparameter Optimization

USAAIO Lesson 59, from Phase 3. It covers grid search, which is exhaustive and exponentially costly; random search, which samples on a log scale and matches or beats grid search at a fraction of the evaluations; Gaussian-process Bayesian optimization, with a surrogate model and an expected-improvement acquisition function; and successive halving with Hyperband, which eliminates candidates early. It closes with practical tips for working on a log scale. All the numbers were verified with sklearn 1.x on the digits dataset in June 2026. The lesson runs to 29 slides.

Subject: Machine Learning · 55 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Hyperparameter Optimization

Title

USAAIO · Lesson 59 · Phase 3

Grid search · random search · Bayesian optimization · successive halving. Finding the best knobs without burning all your compute.

2. By the end of this lesson you can

Objectives

  1. Explain why grid search is exponential in the number of hyperparameters
  2. Implement random search and explain why it matches grid performance with fewer evals
  3. Describe the Gaussian Process surrogate in Bayesian optimization and the role of the acquisition function
  4. Trace successive halving: which configs survive and how it reduces total compute
  5. Apply the log-scale rule for learning rate and weight decay, and integer-scale for hidden sizes

3. What survived from Nonlinear Dimensionality Reduction (Kernel PCA, t-SNE, UMAP)?

Warm-up

Discussion prompt

Before we open Lesson 59: Hyperparameter Optimization: without looking back, what was the main idea of Nonlinear Dimensionality Reduction (Kernel PCA, t-SNE, UMAP), and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

Kernel PCA — algorithm from scratch (kernel matrix, centering, eigendecomposition, projection); RBF vs polynomial kernel choice; t-SNE mechanics (perplexity, learning_rate, KL divergence, why new-point projection fails); UMAP theory (fuzzy simplicial sets, cross-entropy optimization, inductive mapping); method comparison on Swiss roll.

4. Grid search: exhaustive but exponential

Section

Part 1 of 4

5. Grid search mechanics

Concept

Grid search enumerates every combination from a finite set of values per hyperparameter. With k choices per param and d params the eval count is k^d — exponential in d.

params (d)values each (k=5)evaluations
155
2525
35125
45625
553 125

Each evaluation is typically a full k-fold cross-validation run. 3 params × 3 values = 27 CV runs — already heavy for slow models.

6. Fill in: values each (k=5) for Grid search mechanics

Comparison

Comparison matrix

From Grid search mechanics: refill the values each (k=5) column from what you know. The rest of the table is as it appeared.

params (d)values each (k=5)evaluations
155
2525
35125
45625
553 125

7. Guess the shape of the answer: Manual grid search — SVC on digits

Estimation

Predict first

Grid: C ∈ {0.1, 1, 10}, gamma ∈ {0.01, 0.001} → 6 configs, 5-fold CV each. Run every combo with nested loops.

Commit before you compute: what does Manual grid search — SVC on digits come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Best config: C=1, gamma=0.001, CV=0.9910 (test acc=0.9917 with sklearn GridSearchCV)

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. 6 CV runs exhausted the grid; grid search guarantees you evaluated every combo, but scales poorly when you add a third hyperparameter (then 18 runs, 54 with 4 values, etc.).

8. Manual grid search — SVC on digits

Worked example

Grid: C ∈ {0.1, 1, 10}, gamma ∈ {0.01, 0.001} → 6 configs, 5-fold CV each. Run every combo with nested loops.

from sklearn.datasets import load_digits
from sklearn.svm import SVC
from sklearn.model_selection import train_test_split, cross_val_score
from itertools import product

X, y = load_digits(return_X_y=True)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, random_state=0)

best, best_p = -1, {}
for C, gamma in product([0.1, 1, 10], [0.01, 0.001]):
    sc = cross_val_score(SVC(C=C, gamma=gamma), X_tr, y_tr, cv=5).mean()
    if sc > best: best, best_p = sc, {'C': C, 'gamma': gamma}

print(best_p, round(best, 4))
Cgamma5-fold CV acc
0.10.010.1072
0.10.0010.9610
10.010.7689
10.0010.9910
100.010.7849
100.0010.9910

Best config: C=1, gamma=0.001, CV=0.9910 (test acc=0.9917 with sklearn GridSearchCV)

Why: 6 CV runs exhausted the grid; grid search guarantees you evaluated every combo, but scales poorly when you add a third hyperparameter (then 18 runs, 54 with 4 values, etc.).

9. What each one costs: Manual grid search — SVC on digits

Trade off

Comparison matrix

From Manual grid search — SVC on digits: every row here is a choice with a cost. Fill the gamma column, then say which row you would actually pick and what you give up for it.

Cgamma5-fold CV acc
0.10.010.1072
0.10.0010.9610
10.010.7689
10.0010.9910
100.010.7849
100.0010.9910

10. Something is wrong here: searching learning rate on a linear grid

Anomaly

Predict first

A student writes this, and it looks reasonable:

lr ∈ {0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0} — 10 equally spaced values.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Learning rate spans multiple orders of magnitude.

lr = 10^u where u ~ Uniform(−5, 0) — log-uniform sampling.

Why: Learning rate spans multiple orders of magnitude. On a linear grid the entire range [1e-5, 0.09] — where most effective LRs live — gets zero candidates. Result: best CV = 0.9708 at C=0.4.

11. Trap: searching learning rate on a linear grid

Trap

The trap

lr ∈ {0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0} — 10 equally spaced values.

All 10 values lie in [0.1, 1.0]; nothing below 0.1 is ever tried

Why: Learning rate spans multiple orders of magnitude. On a linear grid the entire range [1e-5, 0.09] — where most effective LRs live — gets zero candidates. Result: best CV = 0.9708 at C=0.4.

The fix

lr = 10^u where u ~ Uniform(−5, 0) — log-uniform sampling.

70% of log-uniform samples fall below 0.1, covering the useful low-LR range

Why: Verified: linear sampling gives 0% of 10 samples below 0.1; log sampling gives 70%. Same number of evals, log scale finds C=0.3 with CV=0.9701 while linear misses it.

12. Break it on purpose: searching learning rate on a linear grid

Break the constraint

Discussion prompt

The rule this trap just fixed:

lr = 10^u where u ~ Uniform(−5, 0) — log-uniform sampling.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

Learning rate spans multiple orders of magnitude. On a linear grid the entire range [1e-5, 0.09] — where most effective LRs live — gets zero candidates. Result: best CV = 0.9708 at C=0.4.

13. Random search: same performance, fewer evaluations

Section

Part 2 of 4

14. Why random search works (Bergstra & Bengio 2012)

Concept

Most hyperparameters are not equally important. If performance depends strongly on 2 of 5 params, grid search wastes most of its budget on the 3 irrelevant ones.

Random search samples each param independently. For any 2-param subspace it covers many distinct values — while grid search covers only its fixed grid points. With 20 random trials you explore 20 distinct points in every 2D subspace; a 4×4×4 grid gives only 4 distinct values per axis.

method5-param problemevaluationstest acc (MLP)
grid (2 values each)2^5 = 32 combos320.8500
random searchn_iter=20200.8450

15. Break it if you can: Why random search works (Bergstra & Bengio 2012)

Counterexample

Discussion prompt

Most hyperparameters are not equally important. If performance depends strongly on 2 of 5 params, grid search wastes most of its budget on the 3 irrelevant ones.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

16. What has to be given first: Manual random search — log-uniform sampling

Missing information

Discussion prompt

Sample C = 10^u, u ~ U(−1, 2) and gamma = 10^v, v ~ U(−4, −1). Run 10 random configs, rank by 5-fold CV.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

Log sampling densely covers small gamma values (1e-4 to 1e-3) where SVC is sensitive, whereas a coarse linear grid misses them entirely.

17. Manual random search — log-uniform sampling

Worked example

Sample C = 10^u, u ~ U(−1, 2) and gamma = 10^v, v ~ U(−4, −1). Run 10 random configs, rank by 5-fold CV.

import numpy as np
from sklearn.svm import SVC
from sklearn.model_selection import cross_val_score

np.random.seed(7)
results = []
for _ in range(10):
    C = 10 ** np.random.uniform(-1, 2)
    gamma = 10 ** np.random.uniform(-4, -1)
    sc = cross_val_score(SVC(C=C, gamma=gamma), X_tr, y_tr, cv=5).mean()
    results.append((round(C,4), round(gamma,6), round(sc,4)))

results.sort(key=lambda x: x[2], reverse=True)
for row in results[:5]: print(row)
Cgamma5-fold CV
20.22750.0008750.9910
4.38230.0008520.9910
0.83950.0006190.9882
62.17530.0001190.9875
17.83320.0101630.7703

10 random configs reached CV=0.9910 — matching the 6-config grid — with no gaps in the low-gamma regime

Why: Log sampling densely covers small gamma values (1e-4 to 1e-3) where SVC is sensitive, whereas a coarse linear grid misses them entirely.

18. Watch it run: Manual random search — log-uniform sampling

Pattern

Step through it

Step through Manual random search — log-uniform sampling one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: C is 20.2275
  2. Step 2: C is 4.3823
  3. Step 3: C is 0.8395
  4. Step 4: C is 62.1753
  5. Step 5: C is 17.8332

19. Bayesian optimization: learn the landscape

Section

Part 3 of 4

20. The surrogate model idea

Concept

Bayesian optimization fits a surrogate model (typically a Gaussian Process) to the observed (hyperparams → metric) points. The GP predicts the mean and uncertainty of the metric at every untried point.

An acquisition function (e.g. Expected Improvement) balances exploitation (go where the mean is high) and exploration (go where uncertainty is high). The next config to evaluate is the argmax of the acquisition function.

  1. Evaluate a few random points to seed the GP
  2. Fit GP to all observed (x, metric) pairs
  3. Find next x = argmax Expected Improvement
  4. Evaluate the real model at x; add to observations
  5. Repeat until budget exhausted

21. Gaussian Process surrogate — the math

Concept

A GP defines a distribution over functions. Given observations (X_obs, y_obs), the predictive distribution at a new point x* is Gaussian:

\[ p(f(x^*) \mid X_{\text{obs}}, y_{\text{obs}}) = \mathcal{N}(\mu(x^*),\, \sigma^2(x^*)) \]

\[ \text{EI}(x) = (\mu(x) - f^* - \xi)\,\Phi(Z) + \sigma(x)\,\phi(Z), \quad Z = \frac{\mu(x) - f^* - \xi}{\sigma(x)} \]

Here f* is the current best observed value, Φ is the normal CDF, φ is its PDF, and ξ > 0 is an exploration jitter. High EI means: good predicted value OR high uncertainty (unexplored region).

22. By analogy: Gaussian Process surrogate — the math

Analogy

Discussion prompt

Explain Gaussian Process surrogate — the math by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

A GP defines a distribution over functions. Given observations (X_obs, y_obs), the predictive distribution at a new point x* is Gaussian:

23. Guess the shape of the answer: GP surrogate homes in on the optimum

Estimation

Predict first

Toy 1-D: maximize f(x) = −(x−1.5)² + 2.25 on [0, 3]. True peak at x=1.5, f=2.25. Start from 3 random observations.

Commit before you compute: what does GP surrogate homes in on the optimum come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: After 8 total evaluations the GP located the peak at x≈1.64 (true 1.5), best f=2.32 vs true 2.25

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The GP built a model of the landscape from 3 seeds and directed subsequent evaluations toward high-EI points.

24. GP surrogate homes in on the optimum

Worked example

Toy 1-D: maximize f(x) = −(x−1.5)² + 2.25 on [0, 3]. True peak at x=1.5, f=2.25. Start from 3 random observations.

import numpy as np
from sklearn.gaussian_process import GaussianProcessRegressor
from sklearn.gaussian_process.kernels import Matern
from scipy.stats import norm

def ei(X_new, X_obs, y_obs, gp, xi=0.01):
    mu, sig = gp.predict(X_new.reshape(-1,1), return_std=True)
    Z = (mu - y_obs.max() - xi) / (sig + 1e-9)
    return (mu - y_obs.max() - xi)*norm.cdf(Z) + sig*norm.pdf(Z)

X_obs = np.array([0.2, 0.8, 2.5])
y_obs = -(X_obs-1.5)**2 + 2.25 + np.random.default_rng(3).normal(0, 0.1, 3)
gp = GaussianProcessRegressor(Matern(nu=2.5), alpha=0.01, random_state=0)
gp.fit(X_obs.reshape(-1,1), y_obs)
X_cand = np.linspace(0, 3, 100)
next_x = X_cand[ei(X_cand, X_obs, y_obs, gp).argmax()]
print(f'next_x={next_x:.4f}  true_opt=1.5')
evalx chosen (EI argmax)f(x) observednotes
seed 10.22.0285random init
seed 20.81.9975random init
seed 32.51.9775random init
BO step 11.48482.1946EI peak; near true opt
BO step 21.63642.2331tightening
BO step 51.78792.2019best obs = 2.3227

After 8 total evaluations the GP located the peak at x≈1.64 (true 1.5), best f=2.32 vs true 2.25

Why: The GP built a model of the landscape from 3 seeds and directed subsequent evaluations toward high-EI points. A random search over 8 points would find this by chance ~30% of the time.

25. Fill in: f(x) observed for GP surrogate homes in on the optimum

Comparison

Comparison matrix

From GP surrogate homes in on the optimum: refill the f(x) observed column from what you know. The rest of the table is as it appeared.

evalx chosen (EI argmax)f(x) observednotes
seed 10.22.0285random init
seed 20.81.9975random init
seed 32.51.9775random init
BO step 11.48482.1946EI peak; near true opt
BO step 21.63642.2331tightening
BO step 51.78792.2019best obs = 2.3227

26. Something is wrong here: treating Bayesian opt as free (ignoring surrogate cost)

Anomaly

Predict first

A student writes this, and it looks reasonable:

BO minimizes expensive function evaluations, so always use Bayesian opt — even for 3 hyperparameters and a 1-second model.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Fitting and querying the GP itself costs O(n³) in the number of observations.

Use Bayesian optimization when each evaluation is expensive (minutes to hours of training).

Why: Fitting and querying the GP itself costs O(n³) in the number of observations. For cheap models (seconds per fit) or small grids (≤3 params × 3 vals = 27 evals), simple random search has no surrogate overhead and often performs equally well.

27. Trap: treating Bayesian opt as free (ignoring surrogate cost)

Trap

The trap

BO minimizes expensive function evaluations, so always use Bayesian opt — even for 3 hyperparameters and a 1-second model.

Apply a GP surrogate to every tuning job regardless of evaluation cost

Why: Fitting and querying the GP itself costs O(n³) in the number of observations. For cheap models (seconds per fit) or small grids (≤3 params × 3 vals = 27 evals), simple random search has no surrogate overhead and often performs equally well.

The fix

Use Bayesian optimization when each evaluation is expensive (minutes to hours of training).

Match method to evaluation cost: cheap → random/grid; expensive → Bayesian (Optuna, Hyperopt, SMAC)

Why: The GP overhead is worthwhile only when it saves more evaluations than it costs. Optuna's default sampler uses Tree-structured Parzen Estimators (TPE), a faster approximate alternative to full GPs.

28. Which of these survive contact with Lesson 59: Hyperparameter Optimization?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Each evaluation is typically a full k-fold cross-validation run. 3 params × 3 values = 27 CV runs — already heavy for slow models.; Most hyperparameters are not equally important. If performance depends strongly on 2 of 5 params, grid search wastes most of its budget on the 3 irrelevant ones.; A GP defines a distribution over functions. Given observations (X_obs, y_obs), the predictive distribution at a new point x* is Gaussian:
Breaks
lr ∈ {0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0} — 10 equally spaced values.; BO minimizes expensive function evaluations, so always use Bayesian opt — even for 3 hyperparameters and a 1-second model.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 59: Hyperparameter Optimization puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

29. Successive halving & Hyperband

Section

Part 4 of 4

30. Successive halving: eliminate early

Concept

Start with n random configs but train each for only a small budget (epochs, data fraction, or CV folds). Keep the top half; double the budget; repeat. Most bad configs are eliminated cheap.

roundconfigs alivebudget eachtotal CV fits
0 (seed)8cv=324
1 (halve)4cv=520
2 (halve)2cv=1020
total——64 vs 80 (full)

On the digits SVC experiment: round 0 best CV=0.9896, round 1 best CV=0.9923, final winner CV=0.9903 — saves 20% of fits vs fully evaluating all 8 configs.

31. Which is which, by total CV fits

Discrimination

Sort into buckets

Sort these by total CV fits, from memory, without looking back at Successive halving: eliminate early. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

24
0 (seed)
20
1 (halve); 2 (halve)
64 vs 80 (full)
total
g1
total CV fits is "24" for 0 (seed) — that is what the table on "Successive halving: eliminate early" records, and it is the single property separating this group from the rest.
g2
total CV fits is "20" for 1 (halve), 2 (halve) — that is what the table on "Successive halving: eliminate early" records, and it is the single property separating this group from the rest.
g3
total CV fits is "64 vs 80 (full)" for total — that is what the table on "Successive halving: eliminate early" records, and it is the single property separating this group from the rest.

32. Hyperband: run many successive halvings

Concept

Hyperband runs successive halving multiple times with different starting budgets (small budget → many configs; large budget → few configs) and takes the best result across brackets. It fixes the bias-variance tradeoff of choosing the initial budget in SH.

\[ \text{total budget} = B \cdot (\lfloor\log_{\eta} n\rfloor + 1) \]

Here η is the halving rate (often 3) and n is configs per bracket. Hyperband is the default early-stopping strategy in Optuna and Ray Tune.

33. Rebuild the recipe: Hyperparameter optimization recipe

Ranking

Put in order

These are the steps of Hyperparameter optimization recipe, scrambled. Put them back in order before the next slide shows you.

  1. Identify which params matter (LR, weight decay: log-scale; hidden units: integer; batch size: powers of 2)
  2. Start cheap: grid for ≤2 params; random search (n=20–50) for 3+
  3. Log-scale LR and weight decay always; never search them linearly
  4. Bayesian opt (Optuna/SMAC) when each trial costs minutes or more
  5. Successive halving / Hyperband when budget is limited and training can be checkpointed
  6. Validate honestly: use a held-out test set, never re-use it to select hyperparams

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

34. Hyperparameter optimization recipe

Pattern

  1. Identify which params matter (LR, weight decay: log-scale; hidden units: integer; batch size: powers of 2)
  2. Start cheap: grid for ≤2 params; random search (n=20–50) for 3+
  3. Log-scale LR and weight decay always; never search them linearly
  4. Bayesian opt (Optuna/SMAC) when each trial costs minutes or more
  5. Successive halving / Hyperband when budget is limited and training can be checkpointed
  6. Validate honestly: use a held-out test set, never re-use it to select hyperparams

35. Where does it stop working: Hyperparameter optimization recipe

Edge cases

Discussion prompt

Hyperparameter optimization recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Identify which params matter (LR, weight decay: log-scale; hidden units: integer; batch size: powers of 2)
  2. Start cheap: grid for ≤2 params; random search (n=20–50) for 3+
  3. Log-scale LR and weight decay always; never search them linearly
  4. Bayesian opt (Optuna/SMAC) when each trial costs minutes or more
  5. Successive halving / Hyperband when budget is limited and training can be checkpointed
  6. Validate honestly: use a held-out test set, never re-use it to select hyperparams

36. Rule out three: Check yourself — grid search scaling

Elimination

Eliminate the wrong options

You add a third hyperparameter (3 values) to an existing 3×3 grid. How many additional evaluations does this require?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. 27 — the grid triples from 9 to 27
  • B. 9 — one row per new value
  • C. 3 — one evaluation per new value
  • D. 18 — 9 extra per new value

Survives elimination: A

Why: Original 3×3 = 9 evals. Adding a 3-value third param gives 3×3×3 = 27 evals — triple the original. Grid search cost is multiplicative: k^d, so each new param multiplies by k.

37. Check yourself — grid search scaling

Check

Think before clicking.

Check your understanding

You add a third hyperparameter (3 values) to an existing 3×3 grid. How many additional evaluations does this require?

  • A. 27 — the grid triples from 9 to 27 (correct)
  • B. 9 — one row per new value
  • C. 3 — one evaluation per new value
  • D. 18 — 9 extra per new value

Answer: A

Why: Original 3×3 = 9 evals. Adding a 3-value third param gives 3×3×3 = 27 evals — triple the original. Grid search cost is multiplicative: k^d, so each new param multiplies by k.

Why B tempts people
This would be true if you only added one slice (3 combos of the new param with fixed old params), but grid search evaluates ALL combinations — 9 old configs × 3 new values = 27 total.
Why C tempts people
3 evals would mean fixing the existing params at one point each — that's not a grid, that's a 1-D sweep ignoring all existing combos.
Why D tempts people
18 would be 9 + 9 — off by one doubling. The multiplication is 9×3=27, not 9+9.

38. Answer it before you see the options: Check yourself — random vs grid coverage

Prediction

Predict first

Random search with the SAME number of evaluations as grid search tends to perform better when:

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: Only a few hyperparameters strongly affect performance (others are irrelevant)

Why: When only 2 of 5 params matter, grid search wastes most of its budget covering the 3 irrelevant ones at fixed points, giving few distinct values in the important 2D subspace. Random search samples each axis independently, covering the important subspace more densely.

39. Check yourself — random vs grid coverage

Check

Key insight from Bergstra & Bengio.

Check your understanding

Random search with the SAME number of evaluations as grid search tends to perform better when:

  • A. Only a few hyperparameters strongly affect performance (others are irrelevant) (correct)
  • B. All hyperparameters are equally important
  • C. The search space is very small (≤4 total configs)
  • D. The model is deterministic (no random seed)

Answer: A

Why: When only 2 of 5 params matter, grid search wastes most of its budget covering the 3 irrelevant ones at fixed points, giving few distinct values in the important 2D subspace. Random search samples each axis independently, covering the important subspace more densely.

Why B tempts people
If all params are equally important, grid search can be competitive — random search's advantage comes precisely from the low-effective-dimensionality structure of most real problems.
Why C tempts people
With ≤4 total configs you can just evaluate everything — the question of random vs grid doesn't arise for exhaustible spaces.
Why D tempts people
Model determinism is irrelevant to the sampling strategy. Noise in evaluation metrics affects which method wins on a single run but not the structural coverage argument.

40. Rule out three: Check yourself — Bayesian optimization

Elimination

Eliminate the wrong options

In Bayesian optimization, the acquisition function (e.g. Expected Improvement) is maximized to choose the next config. Its two terms balance:

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Exploitation (high predicted mean) vs exploration (high GP uncertainty)
  • B. Training loss vs validation loss
  • C. Learning rate vs batch size
  • D. Grid coverage vs random coverage

Survives elimination: A

Why: EI = (μ−f−ξ)·Φ(Z) + σ·φ(Z). The first term is exploitation (go where predicted metric μ exceeds the best seen f); the second term is exploration (go where σ is large — uncertain, unexplored region). The GP verified: after 3 seeds it directed the next eval to x=1.4848, near the true optimum 1.5.

41. Check yourself — Bayesian optimization

Check

Acquisition functions.

Check your understanding

In Bayesian optimization, the acquisition function (e.g. Expected Improvement) is maximized to choose the next config. Its two terms balance:

  • A. Exploitation (high predicted mean) vs exploration (high GP uncertainty) (correct)
  • B. Training loss vs validation loss
  • C. Learning rate vs batch size
  • D. Grid coverage vs random coverage

Answer: A

Why: EI = (μ−f−ξ)·Φ(Z) + σ·φ(Z). The first term is exploitation (go where predicted metric μ exceeds the best seen f); the second term is exploration (go where σ is large — uncertain, unexplored region). The GP verified: after 3 seeds it directed the next eval to x=1.4848, near the true optimum 1.5.

Why B tempts people
Train vs val loss describes model evaluation, not the acquisition function. The GP models the hyperparameter→metric mapping, not the loss decomposition.
Why C tempts people
LR vs batch size are hyperparameters being searched, not the axes of the acquisition function tradeoff.
Why D tempts people
Grid vs random coverage is a comparison of search strategies, not the internal mechanism of Bayesian optimization.

42. Your turn: tune it

Section

Project

43. Project: Grid → Random → Successive Halving

Concept

Tune an SVC on load_digits. Run grid search, then random search with fewer evaluations, then a successive halving loop. Compare best CV scores and total CV fits.

#tasktool
1grid search: C∈{0.1,1,10}, gamma∈{0.01,0.001,0.0001}nested loops + cross_val_score
2random search (10 evals): C, gamma log-uniformnp.random.uniform on log scale
3successive halving: 8 configs, halve twicerank by cv=3, keep top 4; then cv=5, keep top 2

Build rule: always use random_state=0 for reproducibility; report best CV score and total number of cross_val_score calls per method.

44. Break it if you can: Project: Grid → Random → Successive Halving

Counterexample

Discussion prompt

Tune an SVC on load_digits. Run grid search, then random search with fewer evaluations, then a successive halving loop. Compare best CV scores and total CV fits.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Build rule: always use random_state=0 for reproducibility; report best CV score and total number of cross_val_score calls per method.

45. Milestone 1 — grid search

Worked example

Your turn: implement the nested-loop grid search. Predict: how many CV calls total? What's the best config?

Hint: from itertools import product; 3 × 3 = 9 combos × 5-fold = 45 CV fits.

from sklearn.datasets import load_digits
from sklearn.svm import SVC
from sklearn.model_selection import train_test_split, cross_val_score
from itertools import product

X, y = load_digits(return_X_y=True)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, random_state=0)

best, best_p = -1, {}
for C, gamma in product([0.1,1,10], [0.01,0.001,0.0001]):
    sc = cross_val_score(SVC(C=C, gamma=gamma), X_tr, y_tr, cv=5).mean()
    if sc > best: best, best_p = sc, {'C':C,'gamma':gamma}
print(best_p, round(best, 4))
best Cbest gamma5-fold CVtotal CV fits
10.0010.991045 (9 configs × 5)

46. Guess the shape of the answer: Milestone 2 — random search (log-uniform)

Estimation

Predict first

Your turn: draw 10 random configs on log-scale. Predict whether 10 random evals can match the 9-config grid.

Commit before you compute: what does Milestone 2 — random search (log-uniform) come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: 10 random log-uniform evals match grid CV=0.9910 — and cover C values (4.4, 20.2) the discrete grid never touched

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Log-uniform sampling densely explores the sensitive low-gamma, mid-C region the coarse grid misses, finding equivalent performance.

47. Milestone 2 — random search (log-uniform)

Worked example

Your turn: draw 10 random configs on log-scale. Predict whether 10 random evals can match the 9-config grid.

Hint: C = 10**np.random.uniform(-1, 2), gamma = 10**np.random.uniform(-4, -1), np.random.seed(7).

import numpy as np

np.random.seed(7)
results = []
for _ in range(10):
    C = 10 ** np.random.uniform(-1, 2)
    gamma = 10 ** np.random.uniform(-4, -1)
    sc = cross_val_score(SVC(C=C, gamma=gamma), X_tr, y_tr, cv=5).mean()
    results.append((round(C,4), round(gamma,6), round(sc,4)))

results.sort(key=lambda r: r[2], reverse=True)
print(results[0])
Cgamma5-fold CVtotal CV fits
20.22750.0008750.991050 (10 configs × 5)

10 random log-uniform evals match grid CV=0.9910 — and cover C values (4.4, 20.2) the discrete grid never touched

Why: Log-uniform sampling densely explores the sensitive low-gamma, mid-C region the coarse grid misses, finding equivalent performance.

48. Work backwards from the answer: Milestone 2 — random search (log-uniform)

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

10 random log-uniform evals match grid CV=0.9910 — and cover C values (4.4, 20.2) the discrete grid never touched

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Your turn: draw 10 random configs on log-scale. Predict whether 10 random evals can match the 9-config grid.

49. Milestone 3 — successive halving

Worked example

Your turn: seed 8 random configs; evaluate cheaply (cv=3); keep top 4; raise budget (cv=5); keep top 2; final eval (cv=10). Count total CV fits.

Hint: sort by score after each round; slice survivors[:n//2]; total fits = 8×3 + 4×5 + 2×10 = 64 vs 8×10=80 for full evaluation.

import numpy as np

np.random.seed(42)
configs = [(10**np.random.uniform(-1,2), 10**np.random.uniform(-4,-1)) for _ in range(8)]

def eval_round(cfgs, cv):
    return sorted(
        [(C, g, cross_val_score(SVC(C=round(C,4), gamma=round(g,6)),
                                X_tr, y_tr, cv=cv).mean())
         for C, g in cfgs],
        key=lambda r: r[2], reverse=True)

r0 = eval_round(configs, cv=3)[:4]           # keep top 4
r1 = eval_round([(c,g) for c,g,_ in r0], cv=5)[:2]  # keep top 2
r2 = eval_round([(c,g) for c,g,_ in r1], cv=10)     # final
print(f'winner: C={r2[0][0]:.4f}, cv10={r2[0][2]:.4f}')
roundconfigscv foldsCV fitsbest score
083240.9896
1 (top 4)45200.9923
2 (top 2)210200.9903
total——64vs 80 (full)

50. What each one costs: Milestone 3 — successive halving

Trade off

Comparison matrix

From Milestone 3 — successive halving: every row here is a choice with a cost. Fill the CV fits column, then say which row you would actually pick and what you give up for it.

roundconfigscv foldsCV fitsbest score
083240.9896
1 (top 4)45200.9923
2 (top 2)210200.9903
total——64vs 80 (full)

51. Full program — put it together

Concept

from sklearn.datasets import load_digits
from sklearn.svm import SVC
from sklearn.model_selection import train_test_split, cross_val_score, GridSearchCV, RandomizedSearchCV
from scipy.stats import loguniform
import numpy as np

X, y = load_digits(return_X_y=True)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, random_state=0)

# Grid
gs = GridSearchCV(SVC(), {'C':[0.1,1,10],'gamma':[0.01,0.001,0.0001]}, cv=5)
gs.fit(X_tr, y_tr)
print('grid  best_cv:', round(gs.best_score_,4), 'test:', round(gs.score(X_te,y_te),4))

# Random (same evals)
rs = RandomizedSearchCV(SVC(), {'C':loguniform(0.01,100),'gamma':loguniform(1e-5,0.1)},
                         n_iter=9, cv=5, random_state=42)
rs.fit(X_tr, y_tr)
print('random best_cv:', round(rs.best_score_,4), 'test:', round(rs.score(X_te,y_te),4))
methodevalsbest CVtest acc
grid (3×3)90.99100.9917
random (n=9)90.98960.9889

Grid and random reach near-identical results with the same evaluation budget. Increase n_iter to 20+ on a 5-param problem and random search wins (as shown in Part 2).

52. Fill in: best CV for Full program — put it together

Comparison

Comparison matrix

From Full program — put it together: refill the best CV column from what you know. The rest of the table is as it appeared.

methodevalsbest CVtest acc
grid (3×3)90.99100.9917
random (n=9)90.98960.9889

53. Show it off

Concept

Out loud, slides closed: (1) why is grid search exponential, (2) why does log-scale sampling matter for learning rate, (3) what does the GP surrogate do in Bayesian opt, (4) what is the halving loop in successive halving.

Stretch (homework): use Optuna for Bayesian hyperparameter optimization on an MLP. Log the study and plot the contour of the objective. Compare Optuna's TPE sampler vs grid on a 5-param MLP problem.

54. Connect it up: Lesson 59: Hyperparameter Optimization

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Grid search: exhaustive but exponential · Random search: same performance, fewer evaluations · Bayesian optimization: learn the landscape · Successive halving & Hyperband · Your turn: tune it. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

55. What you can do now

Recap

methodbest forkey cost
grid1–2 paramsk^d evaluations
random3+ params, cheap modeln_iter evaluations
Bayesian (Optuna)expensive model (mins)GP/TPE overhead
successive halvinglimited budget, checkpointableearly rounds noise risk

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 59 — Hyperparameter Optimization — Barron · USAAIO Round 2 Preparation, 2026
  2. Grid vs random search verified on sklearn digits dataset (SVC, LogisticRegression, MLPClassifier) — sklearn 1.x + numpy 2.2.6, real execution, June 2026
  3. Bergstra & Bengio (2012): Random Search for Hyper-Parameter Optimization — JMLR 13, 281-305

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108