USAAIO Lesson 59, from Phase 3. It covers grid search, which is exhaustive and exponentially costly; random search, which samples on a log scale and matches or beats grid search at a fraction of the evaluations; Gaussian-process Bayesian optimization, with a surrogate model and an expected-improvement acquisition function; and successive halving with Hyperband, which eliminates candidates early. It closes with practical tips for working on a log scale. All the numbers were verified with sklearn 1.x on the digits dataset in June 2026. The lesson runs to 29 slides.
Subject: Machine Learning · 55 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 59 · Phase 3
Grid search · random search · Bayesian optimization · successive halving. Finding the best knobs without burning all your compute.
Objectives
Warm-up
Discussion prompt
Before we open Lesson 59: Hyperparameter Optimization: without looking back, what was the main idea of Nonlinear Dimensionality Reduction (Kernel PCA, t-SNE, UMAP), and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
Kernel PCA — algorithm from scratch (kernel matrix, centering, eigendecomposition, projection); RBF vs polynomial kernel choice; t-SNE mechanics (perplexity, learning_rate, KL divergence, why new-point projection fails); UMAP theory (fuzzy simplicial sets, cross-entropy optimization, inductive mapping); method comparison on Swiss roll.
Section
Part 1 of 4
Concept
Grid search enumerates every combination from a finite set of values per hyperparameter. With k choices per param and d params the eval count is k^d — exponential in d.
| params (d) | values each (k=5) | evaluations |
|---|---|---|
| 1 | 5 | 5 |
| 2 | 5 | 25 |
| 3 | 5 | 125 |
| 4 | 5 | 625 |
| 5 | 5 | 3 125 |
Each evaluation is typically a full k-fold cross-validation run. 3 params × 3 values = 27 CV runs — already heavy for slow models.
Comparison
Comparison matrix
From Grid search mechanics: refill the values each (k=5) column from what you know. The rest of the table is as it appeared.
| params (d) | values each (k=5) | evaluations |
|---|---|---|
| 1 | 5 | 5 |
| 2 | 5 | 25 |
| 3 | 5 | 125 |
| 4 | 5 | 625 |
| 5 | 5 | 3 125 |
Estimation
Predict first
Grid: C ∈ {0.1, 1, 10}, gamma ∈ {0.01, 0.001} → 6 configs, 5-fold CV each. Run every combo with nested loops.
Commit before you compute: what does Manual grid search — SVC on digits come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Best config: C=1, gamma=0.001, CV=0.9910 (test acc=0.9917 with sklearn GridSearchCV)
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. 6 CV runs exhausted the grid; grid search guarantees you evaluated every combo, but scales poorly when you add a third hyperparameter (then 18 runs, 54 with 4 values, etc.).
Worked example
Grid: C ∈ {0.1, 1, 10}, gamma ∈ {0.01, 0.001} → 6 configs, 5-fold CV each. Run every combo with nested loops.
from sklearn.datasets import load_digits
from sklearn.svm import SVC
from sklearn.model_selection import train_test_split, cross_val_score
from itertools import product
X, y = load_digits(return_X_y=True)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, random_state=0)
best, best_p = -1, {}
for C, gamma in product([0.1, 1, 10], [0.01, 0.001]):
sc = cross_val_score(SVC(C=C, gamma=gamma), X_tr, y_tr, cv=5).mean()
if sc > best: best, best_p = sc, {'C': C, 'gamma': gamma}
print(best_p, round(best, 4))| C | gamma | 5-fold CV acc |
|---|---|---|
| 0.1 | 0.01 | 0.1072 |
| 0.1 | 0.001 | 0.9610 |
| 1 | 0.01 | 0.7689 |
| 1 | 0.001 | 0.9910 |
| 10 | 0.01 | 0.7849 |
| 10 | 0.001 | 0.9910 |
Best config: C=1, gamma=0.001, CV=0.9910 (test acc=0.9917 with sklearn GridSearchCV)
Why: 6 CV runs exhausted the grid; grid search guarantees you evaluated every combo, but scales poorly when you add a third hyperparameter (then 18 runs, 54 with 4 values, etc.).
Trade off
Comparison matrix
From Manual grid search — SVC on digits: every row here is a choice with a cost. Fill the gamma column, then say which row you would actually pick and what you give up for it.
| C | gamma | 5-fold CV acc |
|---|---|---|
| 0.1 | 0.01 | 0.1072 |
| 0.1 | 0.001 | 0.9610 |
| 1 | 0.01 | 0.7689 |
| 1 | 0.001 | 0.9910 |
| 10 | 0.01 | 0.7849 |
| 10 | 0.001 | 0.9910 |
Anomaly
Predict first
A student writes this, and it looks reasonable:
lr ∈ {0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0} — 10 equally spaced values.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Learning rate spans multiple orders of magnitude.
lr = 10^u where u ~ Uniform(−5, 0) — log-uniform sampling.
Why: Learning rate spans multiple orders of magnitude. On a linear grid the entire range [1e-5, 0.09] — where most effective LRs live — gets zero candidates. Result: best CV = 0.9708 at C=0.4.
Trap
lr ∈ {0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0} — 10 equally spaced values.
All 10 values lie in [0.1, 1.0]; nothing below 0.1 is ever tried
Why: Learning rate spans multiple orders of magnitude. On a linear grid the entire range [1e-5, 0.09] — where most effective LRs live — gets zero candidates. Result: best CV = 0.9708 at C=0.4.
lr = 10^u where u ~ Uniform(−5, 0) — log-uniform sampling.
70% of log-uniform samples fall below 0.1, covering the useful low-LR range
Why: Verified: linear sampling gives 0% of 10 samples below 0.1; log sampling gives 70%. Same number of evals, log scale finds C=0.3 with CV=0.9701 while linear misses it.
Break the constraint
Discussion prompt
The rule this trap just fixed:lr = 10^u where u ~ Uniform(−5, 0) — log-uniform sampling.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
Learning rate spans multiple orders of magnitude. On a linear grid the entire range [1e-5, 0.09] — where most effective LRs live — gets zero candidates. Result: best CV = 0.9708 at C=0.4.
Section
Part 2 of 4
Concept
Most hyperparameters are not equally important. If performance depends strongly on 2 of 5 params, grid search wastes most of its budget on the 3 irrelevant ones.
Random search samples each param independently. For any 2-param subspace it covers many distinct values — while grid search covers only its fixed grid points. With 20 random trials you explore 20 distinct points in every 2D subspace; a 4×4×4 grid gives only 4 distinct values per axis.
| method | 5-param problem | evaluations | test acc (MLP) |
|---|---|---|---|
| grid (2 values each) | 2^5 = 32 combos | 32 | 0.8500 |
| random search | n_iter=20 | 20 | 0.8450 |
Counterexample
Discussion prompt
Most hyperparameters are not equally important. If performance depends strongly on 2 of 5 params, grid search wastes most of its budget on the 3 irrelevant ones.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Missing information
Discussion prompt
Sample C = 10^u, u ~ U(−1, 2) and gamma = 10^v, v ~ U(−4, −1). Run 10 random configs, rank by 5-fold CV.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Log sampling densely covers small gamma values (1e-4 to 1e-3) where SVC is sensitive, whereas a coarse linear grid misses them entirely.
Worked example
Sample C = 10^u, u ~ U(−1, 2) and gamma = 10^v, v ~ U(−4, −1). Run 10 random configs, rank by 5-fold CV.
import numpy as np
from sklearn.svm import SVC
from sklearn.model_selection import cross_val_score
np.random.seed(7)
results = []
for _ in range(10):
C = 10 ** np.random.uniform(-1, 2)
gamma = 10 ** np.random.uniform(-4, -1)
sc = cross_val_score(SVC(C=C, gamma=gamma), X_tr, y_tr, cv=5).mean()
results.append((round(C,4), round(gamma,6), round(sc,4)))
results.sort(key=lambda x: x[2], reverse=True)
for row in results[:5]: print(row)| C | gamma | 5-fold CV |
|---|---|---|
| 20.2275 | 0.000875 | 0.9910 |
| 4.3823 | 0.000852 | 0.9910 |
| 0.8395 | 0.000619 | 0.9882 |
| 62.1753 | 0.000119 | 0.9875 |
| 17.8332 | 0.010163 | 0.7703 |
10 random configs reached CV=0.9910 — matching the 6-config grid — with no gaps in the low-gamma regime
Why: Log sampling densely covers small gamma values (1e-4 to 1e-3) where SVC is sensitive, whereas a coarse linear grid misses them entirely.
Pattern
Step through it
Step through Manual random search — log-uniform sampling one row at a time. What is driving the change, and what would the row after the last one be?
Section
Part 3 of 4
Concept
Bayesian optimization fits a surrogate model (typically a Gaussian Process) to the observed (hyperparams → metric) points. The GP predicts the mean and uncertainty of the metric at every untried point.
An acquisition function (e.g. Expected Improvement) balances exploitation (go where the mean is high) and exploration (go where uncertainty is high). The next config to evaluate is the argmax of the acquisition function.
Concept
A GP defines a distribution over functions. Given observations (X_obs, y_obs), the predictive distribution at a new point x* is Gaussian:
\[ p(f(x^*) \mid X_{\text{obs}}, y_{\text{obs}}) = \mathcal{N}(\mu(x^*),\, \sigma^2(x^*)) \]
\[ \text{EI}(x) = (\mu(x) - f^* - \xi)\,\Phi(Z) + \sigma(x)\,\phi(Z), \quad Z = \frac{\mu(x) - f^* - \xi}{\sigma(x)} \]
Here f* is the current best observed value, Φ is the normal CDF, φ is its PDF, and ξ > 0 is an exploration jitter. High EI means: good predicted value OR high uncertainty (unexplored region).
Analogy
Discussion prompt
Explain Gaussian Process surrogate — the math by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
A GP defines a distribution over functions. Given observations (X_obs, y_obs), the predictive distribution at a new point x* is Gaussian:
Estimation
Predict first
Toy 1-D: maximize f(x) = −(x−1.5)² + 2.25 on [0, 3]. True peak at x=1.5, f=2.25. Start from 3 random observations.
Commit before you compute: what does GP surrogate homes in on the optimum come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: After 8 total evaluations the GP located the peak at x≈1.64 (true 1.5), best f=2.32 vs true 2.25
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The GP built a model of the landscape from 3 seeds and directed subsequent evaluations toward high-EI points.
Worked example
Toy 1-D: maximize f(x) = −(x−1.5)² + 2.25 on [0, 3]. True peak at x=1.5, f=2.25. Start from 3 random observations.
import numpy as np
from sklearn.gaussian_process import GaussianProcessRegressor
from sklearn.gaussian_process.kernels import Matern
from scipy.stats import norm
def ei(X_new, X_obs, y_obs, gp, xi=0.01):
mu, sig = gp.predict(X_new.reshape(-1,1), return_std=True)
Z = (mu - y_obs.max() - xi) / (sig + 1e-9)
return (mu - y_obs.max() - xi)*norm.cdf(Z) + sig*norm.pdf(Z)
X_obs = np.array([0.2, 0.8, 2.5])
y_obs = -(X_obs-1.5)**2 + 2.25 + np.random.default_rng(3).normal(0, 0.1, 3)
gp = GaussianProcessRegressor(Matern(nu=2.5), alpha=0.01, random_state=0)
gp.fit(X_obs.reshape(-1,1), y_obs)
X_cand = np.linspace(0, 3, 100)
next_x = X_cand[ei(X_cand, X_obs, y_obs, gp).argmax()]
print(f'next_x={next_x:.4f} true_opt=1.5')| eval | x chosen (EI argmax) | f(x) observed | notes |
|---|---|---|---|
| seed 1 | 0.2 | 2.0285 | random init |
| seed 2 | 0.8 | 1.9975 | random init |
| seed 3 | 2.5 | 1.9775 | random init |
| BO step 1 | 1.4848 | 2.1946 | EI peak; near true opt |
| BO step 2 | 1.6364 | 2.2331 | tightening |
| BO step 5 | 1.7879 | 2.2019 | best obs = 2.3227 |
After 8 total evaluations the GP located the peak at x≈1.64 (true 1.5), best f=2.32 vs true 2.25
Why: The GP built a model of the landscape from 3 seeds and directed subsequent evaluations toward high-EI points. A random search over 8 points would find this by chance ~30% of the time.
Comparison
Comparison matrix
From GP surrogate homes in on the optimum: refill the f(x) observed column from what you know. The rest of the table is as it appeared.
| eval | x chosen (EI argmax) | f(x) observed | notes |
|---|---|---|---|
| seed 1 | 0.2 | 2.0285 | random init |
| seed 2 | 0.8 | 1.9975 | random init |
| seed 3 | 2.5 | 1.9775 | random init |
| BO step 1 | 1.4848 | 2.1946 | EI peak; near true opt |
| BO step 2 | 1.6364 | 2.2331 | tightening |
| BO step 5 | 1.7879 | 2.2019 | best obs = 2.3227 |
Anomaly
Predict first
A student writes this, and it looks reasonable:
BO minimizes expensive function evaluations, so always use Bayesian opt — even for 3 hyperparameters and a 1-second model.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Fitting and querying the GP itself costs O(n³) in the number of observations.
Use Bayesian optimization when each evaluation is expensive (minutes to hours of training).
Why: Fitting and querying the GP itself costs O(n³) in the number of observations. For cheap models (seconds per fit) or small grids (≤3 params × 3 vals = 27 evals), simple random search has no surrogate overhead and often performs equally well.
Trap
BO minimizes expensive function evaluations, so always use Bayesian opt — even for 3 hyperparameters and a 1-second model.
Apply a GP surrogate to every tuning job regardless of evaluation cost
Why: Fitting and querying the GP itself costs O(n³) in the number of observations. For cheap models (seconds per fit) or small grids (≤3 params × 3 vals = 27 evals), simple random search has no surrogate overhead and often performs equally well.
Use Bayesian optimization when each evaluation is expensive (minutes to hours of training).
Match method to evaluation cost: cheap → random/grid; expensive → Bayesian (Optuna, Hyperopt, SMAC)
Why: The GP overhead is worthwhile only when it saves more evaluations than it costs. Optuna's default sampler uses Tree-structured Parzen Estimators (TPE), a faster approximate alternative to full GPs.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
(X_obs, y_obs), the predictive distribution at a new point x* is Gaussian:lr ∈ {0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0} — 10 equally spaced values.; BO minimizes expensive function evaluations, so always use Bayesian opt — even for 3 hyperparameters and a 1-second model.Section
Part 4 of 4
Concept
Start with n random configs but train each for only a small budget (epochs, data fraction, or CV folds). Keep the top half; double the budget; repeat. Most bad configs are eliminated cheap.
| round | configs alive | budget each | total CV fits |
|---|---|---|---|
| 0 (seed) | 8 | cv=3 | 24 |
| 1 (halve) | 4 | cv=5 | 20 |
| 2 (halve) | 2 | cv=10 | 20 |
| total | — | — | 64 vs 80 (full) |
On the digits SVC experiment: round 0 best CV=0.9896, round 1 best CV=0.9923, final winner CV=0.9903 — saves 20% of fits vs fully evaluating all 8 configs.
Discrimination
Sort into buckets
Sort these by total CV fits, from memory, without looking back at Successive halving: eliminate early. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Concept
Hyperband runs successive halving multiple times with different starting budgets (small budget → many configs; large budget → few configs) and takes the best result across brackets. It fixes the bias-variance tradeoff of choosing the initial budget in SH.
\[ \text{total budget} = B \cdot (\lfloor\log_{\eta} n\rfloor + 1) \]
Here η is the halving rate (often 3) and n is configs per bracket. Hyperband is the default early-stopping strategy in Optuna and Ray Tune.
Ranking
Put in order
These are the steps of Hyperparameter optimization recipe, scrambled. Put them back in order before the next slide shows you.
Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
Edge cases
Discussion prompt
Hyperparameter optimization recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
Elimination
Eliminate the wrong options
You add a third hyperparameter (3 values) to an existing 3×3 grid. How many additional evaluations does this require?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Original 3×3 = 9 evals. Adding a 3-value third param gives 3×3×3 = 27 evals — triple the original. Grid search cost is multiplicative: k^d, so each new param multiplies by k.
Check
Think before clicking.
Check your understanding
You add a third hyperparameter (3 values) to an existing 3×3 grid. How many additional evaluations does this require?
Answer: A
Why: Original 3×3 = 9 evals. Adding a 3-value third param gives 3×3×3 = 27 evals — triple the original. Grid search cost is multiplicative: k^d, so each new param multiplies by k.
Prediction
Predict first
Random search with the SAME number of evaluations as grid search tends to perform better when:
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Only a few hyperparameters strongly affect performance (others are irrelevant)
Why: When only 2 of 5 params matter, grid search wastes most of its budget covering the 3 irrelevant ones at fixed points, giving few distinct values in the important 2D subspace. Random search samples each axis independently, covering the important subspace more densely.
Check
Key insight from Bergstra & Bengio.
Check your understanding
Random search with the SAME number of evaluations as grid search tends to perform better when:
Answer: A
Why: When only 2 of 5 params matter, grid search wastes most of its budget covering the 3 irrelevant ones at fixed points, giving few distinct values in the important 2D subspace. Random search samples each axis independently, covering the important subspace more densely.
Elimination
Eliminate the wrong options
In Bayesian optimization, the acquisition function (e.g. Expected Improvement) is maximized to choose the next config. Its two terms balance:
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: EI = (μ−f−ξ)·Φ(Z) + σ·φ(Z). The first term is exploitation (go where predicted metric μ exceeds the best seen f); the second term is exploration (go where σ is large — uncertain, unexplored region). The GP verified: after 3 seeds it directed the next eval to x=1.4848, near the true optimum 1.5.
Check
Acquisition functions.
Check your understanding
In Bayesian optimization, the acquisition function (e.g. Expected Improvement) is maximized to choose the next config. Its two terms balance:
Answer: A
Why: EI = (μ−f−ξ)·Φ(Z) + σ·φ(Z). The first term is exploitation (go where predicted metric μ exceeds the best seen f); the second term is exploration (go where σ is large — uncertain, unexplored region). The GP verified: after 3 seeds it directed the next eval to x=1.4848, near the true optimum 1.5.
Section
Project
Concept
Tune an SVC on load_digits. Run grid search, then random search with fewer evaluations, then a successive halving loop. Compare best CV scores and total CV fits.
| # | task | tool |
|---|---|---|
| 1 | grid search: C∈{0.1,1,10}, gamma∈{0.01,0.001,0.0001} | nested loops + cross_val_score |
| 2 | random search (10 evals): C, gamma log-uniform | np.random.uniform on log scale |
| 3 | successive halving: 8 configs, halve twice | rank by cv=3, keep top 4; then cv=5, keep top 2 |
Build rule: always use random_state=0 for reproducibility; report best CV score and total number of cross_val_score calls per method.
Counterexample
Discussion prompt
Tune an SVC on load_digits. Run grid search, then random search with fewer evaluations, then a successive halving loop. Compare best CV scores and total CV fits.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Build rule: always use random_state=0 for reproducibility; report best CV score and total number of cross_val_score calls per method.
Worked example
Your turn: implement the nested-loop grid search. Predict: how many CV calls total? What's the best config?
Hint: from itertools import product; 3 × 3 = 9 combos × 5-fold = 45 CV fits.
from sklearn.datasets import load_digits
from sklearn.svm import SVC
from sklearn.model_selection import train_test_split, cross_val_score
from itertools import product
X, y = load_digits(return_X_y=True)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, random_state=0)
best, best_p = -1, {}
for C, gamma in product([0.1,1,10], [0.01,0.001,0.0001]):
sc = cross_val_score(SVC(C=C, gamma=gamma), X_tr, y_tr, cv=5).mean()
if sc > best: best, best_p = sc, {'C':C,'gamma':gamma}
print(best_p, round(best, 4))| best C | best gamma | 5-fold CV | total CV fits |
|---|---|---|---|
| 1 | 0.001 | 0.9910 | 45 (9 configs × 5) |
Estimation
Predict first
Your turn: draw 10 random configs on log-scale. Predict whether 10 random evals can match the 9-config grid.
Commit before you compute: what does Milestone 2 — random search (log-uniform) come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: 10 random log-uniform evals match grid CV=0.9910 — and cover C values (4.4, 20.2) the discrete grid never touched
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Log-uniform sampling densely explores the sensitive low-gamma, mid-C region the coarse grid misses, finding equivalent performance.
Worked example
Your turn: draw 10 random configs on log-scale. Predict whether 10 random evals can match the 9-config grid.
Hint: C = 10**np.random.uniform(-1, 2), gamma = 10**np.random.uniform(-4, -1), np.random.seed(7).
import numpy as np
np.random.seed(7)
results = []
for _ in range(10):
C = 10 ** np.random.uniform(-1, 2)
gamma = 10 ** np.random.uniform(-4, -1)
sc = cross_val_score(SVC(C=C, gamma=gamma), X_tr, y_tr, cv=5).mean()
results.append((round(C,4), round(gamma,6), round(sc,4)))
results.sort(key=lambda r: r[2], reverse=True)
print(results[0])| C | gamma | 5-fold CV | total CV fits |
|---|---|---|---|
| 20.2275 | 0.000875 | 0.9910 | 50 (10 configs × 5) |
10 random log-uniform evals match grid CV=0.9910 — and cover C values (4.4, 20.2) the discrete grid never touched
Why: Log-uniform sampling densely explores the sensitive low-gamma, mid-C region the coarse grid misses, finding equivalent performance.
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
10 random log-uniform evals match grid CV=0.9910 — and cover C values (4.4, 20.2) the discrete grid never touched
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Your turn: draw 10 random configs on log-scale. Predict whether 10 random evals can match the 9-config grid.
Worked example
Your turn: seed 8 random configs; evaluate cheaply (cv=3); keep top 4; raise budget (cv=5); keep top 2; final eval (cv=10). Count total CV fits.
Hint: sort by score after each round; slice survivors[:n//2]; total fits = 8×3 + 4×5 + 2×10 = 64 vs 8×10=80 for full evaluation.
import numpy as np
np.random.seed(42)
configs = [(10**np.random.uniform(-1,2), 10**np.random.uniform(-4,-1)) for _ in range(8)]
def eval_round(cfgs, cv):
return sorted(
[(C, g, cross_val_score(SVC(C=round(C,4), gamma=round(g,6)),
X_tr, y_tr, cv=cv).mean())
for C, g in cfgs],
key=lambda r: r[2], reverse=True)
r0 = eval_round(configs, cv=3)[:4] # keep top 4
r1 = eval_round([(c,g) for c,g,_ in r0], cv=5)[:2] # keep top 2
r2 = eval_round([(c,g) for c,g,_ in r1], cv=10) # final
print(f'winner: C={r2[0][0]:.4f}, cv10={r2[0][2]:.4f}')| round | configs | cv folds | CV fits | best score |
|---|---|---|---|---|
| 0 | 8 | 3 | 24 | 0.9896 |
| 1 (top 4) | 4 | 5 | 20 | 0.9923 |
| 2 (top 2) | 2 | 10 | 20 | 0.9903 |
| total | — | — | 64 | vs 80 (full) |
Trade off
Comparison matrix
From Milestone 3 — successive halving: every row here is a choice with a cost. Fill the CV fits column, then say which row you would actually pick and what you give up for it.
| round | configs | cv folds | CV fits | best score |
|---|---|---|---|---|
| 0 | 8 | 3 | 24 | 0.9896 |
| 1 (top 4) | 4 | 5 | 20 | 0.9923 |
| 2 (top 2) | 2 | 10 | 20 | 0.9903 |
| total | — | — | 64 | vs 80 (full) |
Concept
from sklearn.datasets import load_digits
from sklearn.svm import SVC
from sklearn.model_selection import train_test_split, cross_val_score, GridSearchCV, RandomizedSearchCV
from scipy.stats import loguniform
import numpy as np
X, y = load_digits(return_X_y=True)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, random_state=0)
# Grid
gs = GridSearchCV(SVC(), {'C':[0.1,1,10],'gamma':[0.01,0.001,0.0001]}, cv=5)
gs.fit(X_tr, y_tr)
print('grid best_cv:', round(gs.best_score_,4), 'test:', round(gs.score(X_te,y_te),4))
# Random (same evals)
rs = RandomizedSearchCV(SVC(), {'C':loguniform(0.01,100),'gamma':loguniform(1e-5,0.1)},
n_iter=9, cv=5, random_state=42)
rs.fit(X_tr, y_tr)
print('random best_cv:', round(rs.best_score_,4), 'test:', round(rs.score(X_te,y_te),4))| method | evals | best CV | test acc |
|---|---|---|---|
| grid (3×3) | 9 | 0.9910 | 0.9917 |
| random (n=9) | 9 | 0.9896 | 0.9889 |
Grid and random reach near-identical results with the same evaluation budget. Increase n_iter to 20+ on a 5-param problem and random search wins (as shown in Part 2).
Comparison
Comparison matrix
From Full program — put it together: refill the best CV column from what you know. The rest of the table is as it appeared.
| method | evals | best CV | test acc |
|---|---|---|---|
| grid (3×3) | 9 | 0.9910 | 0.9917 |
| random (n=9) | 9 | 0.9896 | 0.9889 |
Concept
Out loud, slides closed: (1) why is grid search exponential, (2) why does log-scale sampling matter for learning rate, (3) what does the GP surrogate do in Bayesian opt, (4) what is the halving loop in successive halving.
Stretch (homework): use Optuna for Bayesian hyperparameter optimization on an MLP. Log the study and plot the contour of the objective. Compare Optuna's TPE sampler vs grid on a 5-param MLP problem.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Grid search: exhaustive but exponential · Random search: same performance, fewer evaluations · Bayesian optimization: learn the landscape · Successive halving & Hyperband · Your turn: tune it. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
| method | best for | key cost |
|---|---|---|
| grid | 1–2 params | k^d evaluations |
| random | 3+ params, cheap model | n_iter evaluations |
| Bayesian (Optuna) | expensive model (mins) | GP/TPE overhead |
| successive halving | limited budget, checkpointable | early rounds noise risk |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.