Lesson 76: Algorithm Review & Selection

USAAIO Lesson 76, a Phase 2 review. It is a flash review of seven core classifiers - logistic regression, the SVM, decision trees, random forests, GBMs, kNN, and Naive Bayes - covering their loss functions, gradients, hyperparameters, and bias-variance profiles, with whiteboard-style derivations. It also covers the USAAIO exam question patterns and gives an algorithm-selection framework. You implement and benchmark all seven on a synthetic dataset. The lesson runs to 28 slides.

Subject: Machine Learning · 56 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Algorithm Review & Selection

Title

USAAIO · Lesson 76 · Phase 2 Review

All 7 core classifiers in one session: loss, gradient, hyperparameters, bias-variance, and a framework for choosing the right algorithm under exam conditions.

2. By the end of this lesson you can

Objectives

  1. State the loss function and gradient for each of the 7 core algorithms
  2. Identify each algorithm's key hyperparameters and bias-variance profile
  3. Derive from scratch the gradient of logistic loss, Gini impurity, and the GBM pseudo-residual update
  4. Answer USAAIO exam patterns: 'derive gradient of X', 'prove property Y', 'implement Z'
  5. Apply the algorithm-selection framework: match problem structure to the right model

3. What survived from Mock Exam — Phase 2 Full Review?

Warm-up

Discussion prompt

Before we open Lesson 76: Algorithm Review & Selection: without looking back, what was the main idea of Mock Exam — Phase 2 Full Review, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

a timed Phase 2 mock exam — 2-hour theory exam on ensemble methods, normalization, optimization, and evaluation; 2-hour coding exam implementing random forest, GMM, and CNN training loop from scratch; full answer-key walkthrough and gap analysis before Phase 3 (transformers).

4. The 7 algorithms at a glance

Section

Section 1

5. Loss functions & optimization — the master table

Concept

algorithmlossoptimizationkey hyperparams
Logistic Regressionlog-loss (cross-entropy)gradient descent (L-BFGS/SGD)C, solver, max_iter
SVMhinge loss + C·slackquadratic programming (SMO)C, kernel, γ (RBF)
Decision TreeGini / entropy impuritygreedy split searchmax_depth, min_samples
Random Forestavg Gini across treesbootstrap + greedyn_estimators, max_features
GBMdeviance (log-loss/MSE)gradient boosting (additive)n_estimators, lr, max_depth
kNNnone (lazy learner)none (distance at predict)k, distance metric
Naive BayesMLE per class/featureclosed form (analytically)var_smoothing

6. Fill in: loss for Loss functions & optimization — the master…

Comparison

Comparison matrix

From Loss functions & optimization — the master table: refill the loss column from what you know. The rest of the table is as it appeared.

algorithmlossoptimizationkey hyperparams
Logistic Regressionlog-loss (cross-entropy)gradient descent (L-BFGS/SGD)C, solver, max_iter
SVMhinge loss + C·slackquadratic programming (SMO)C, kernel, γ (RBF)
Decision TreeGini / entropy impuritygreedy split searchmax_depth, min_samples
Random Forestavg Gini across treesbootstrap + greedyn_estimators, max_features
GBMdeviance (log-loss/MSE)gradient boosting (additive)n_estimators, lr, max_depth
kNNnone (lazy learner)none (distance at predict)k, distance metric
Naive BayesMLE per class/featureclosed form (analytically)var_smoothing

7. Bias-variance profiles

Concept

algorithmbiasvarianceoverfits?needs scaling?
Logistic Regressionmoderate (linear boundary)lowrarelyyes
SVM (RBF)lowmoderateif C too highyes
Decision Treelow (deep) / high (shallow)high (deep)easilyno
Random Forestlowreduced vs treeless than DTno
GBMlowlow–moderateif lr high + many treesno
kNNlowhigh (small k)small kyes
Naive Bayeshigh (independence)very lowrarelyno

Benchmark on make_classification(n=500, n_features=10, random_state=42): RF 0.93, GBM 0.92, SVM 0.91, kNN 0.86, NB 0.85, DT 0.84, LR 0.82.

8. Which is which, by needs scaling?

Discrimination

Sort into buckets

Sort these by needs scaling?, from memory, without looking back at Bias-variance profiles. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

yes
Logistic Regression; SVM (RBF); kNN
no
Decision Tree; Random Forest; GBM; Naive Bayes
g1
needs scaling? is "yes" for Logistic Regression, SVM (RBF), kNN — that is what the table on "Bias-variance profiles" records, and it is the single property separating this group from the rest.
g2
needs scaling? is "no" for Decision Tree, Random Forest, GBM, Naive Bayes — that is what the table on "Bias-variance profiles" records, and it is the single property separating this group from the rest.

9. Derivations — logistic & SVM

Section

Section 2

10. Guess the shape of the answer: Derive: gradient of logistic loss

Estimation

Predict first

USAAIO pattern: 'derive the gradient of logistic loss'. Start from the definition of cross-entropy loss.

Commit before you compute: what does Derive: gradient of logistic loss come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Differentiate with respect to w using the chain rule

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. σ'(z) = σ(z)(1−σ(z)); applying the chain rule through log collapses the expression cleanly.

11. Derive: gradient of logistic loss

Worked example

USAAIO pattern: 'derive the gradient of logistic loss'. Start from the definition of cross-entropy loss.

\[ \mathcal{L}(w) = -\tfrac{1}{n}\sum_{i=1}^{n}\bigl[y_i\log\sigma(x_i^\top w) + (1-y_i)\log(1-\sigma(x_i^\top w))\bigr] \]

Differentiate with respect to w using the chain rule

Why: σ'(z) = σ(z)(1−σ(z)); applying the chain rule through log collapses the expression cleanly.

\[ \nabla_w \mathcal{L} = \tfrac{1}{n} X^\top (\sigma(Xw) - y) \]

import numpy as np
def sigmoid(z): return 1 / (1 + np.exp(-z))
X = np.array([[1.,0.5],[2.,1.],[3.,1.5],[-1.,-0.5]])
y = np.array([0.,0.,1.,0.])
w = np.zeros(2); lr = 0.5
for step in range(3):
    p = sigmoid(X @ w)
    loss = -np.mean(y*np.log(p+1e-10)+(1-y)*np.log(1-p+1e-10))
    grad = X.T @ (p - y) / len(y)
    w = w - lr * grad
    print(f'step={step}: loss={loss:.4f}, w={w.round(4)}')
steplossw[0]w[1]
00.69310.06250.0312
10.68620.08850.0443
20.68500.09950.0497

12. What each one costs: Derive: gradient of logistic loss

Trade off

Comparison matrix

From Derive: gradient of logistic loss: every row here is a choice with a cost. Fill the loss column, then say which row you would actually pick and what you give up for it.

steplossw[0]w[1]
00.69310.06250.0312
10.68620.08850.0443
20.68500.09950.0497

13. SVM: hinge loss and the margin

Concept

SVM finds the maximum-margin separating hyperplane. The primal objective balances margin width against slack violations.

\[ \min_{w,b,\xi}\;\tfrac{1}{2}\lVert w\rVert^2 + C\sum_i \xi_i \quad\text{s.t.}\; y_i(w^\top x_i+b)\geq 1-\xi_i,\;\xi_i\geq 0 \]

C controls the bias-variance tradeoff: large C → narrow margin (low bias, high variance); small C → wide margin (high bias, low variance).

Cn_svmargintest acc (iris 2-class)
0.1281.44471.0000
1.0100.78001.0000
10.040.43031.0000

14. What stays fixed: SVM: hinge loss and the margin

Invariant

Step through it

Step through SVM: hinge loss and the margin one row at a time. One of these columns never changes — find it, and say why it cannot.

  1. Step 1: C is 0.1
  2. Step 2: C is 1.0
  3. Step 3: C is 10.0

15. Something is wrong here: confusing SVM's C with regularization strength

Anomaly

Predict first

A student writes this, and it looks reasonable:

Large C means stronger regularization — the model is more conservative.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: This is the L2 intuition from Ridge/logistic (Lesson 12).

Large C means LESS regularization — the model tolerates fewer margin violations.

Why: This is the L2 intuition from Ridge/logistic (Lesson 12). But SVM's C is the INVERSE of regularization strength — the penalty is on slack, not on w.

16. Trap: confusing SVM's C with regularization strength

Trap

The trap

Large C means stronger regularization — the model is more conservative.

Increase C to reduce overfitting

Why: This is the L2 intuition from Ridge/logistic (Lesson 12). But SVM's C is the INVERSE of regularization strength — the penalty is on slack, not on w.

The fix

Large C means LESS regularization — the model tolerates fewer margin violations.

Decrease C to reduce overfitting in SVM

Why: Small C allows a wider margin (larger ‖w‖ penalty implicit), accepting some misclassifications for a smoother boundary. Think: C = cost of violations, not cost of complexity.

17. Trees, forests, and boosting

Section

Section 3

18. What has to be given first: Derive: Gini impurity and information gain

Missing information

Discussion prompt

USAAIO pattern: 'prove that this split maximizes information gain'. Start from the Gini definition.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

p₀ = 0.6, p₁ = 0.4 → Gini = 1 − (0.36 + 0.16) = 0.48

19. Derive: Gini impurity and information gain

Worked example

USAAIO pattern: 'prove that this split maximizes information gain'. Start from the Gini definition.

\[ \text{Gini}(t) = 1 - \sum_{k} p_k^2 \]

Compute Gini at the root: 10 samples, 6 class-0, 4 class-1

Why: p₀ = 0.6, p₁ = 0.4 → Gini = 1 − (0.36 + 0.16) = 0.48

Evaluate a split: left (7 samples, 6/1), right (3 samples, 0/3)

Why: Gini_left = 1 − (6/7)² − (1/7)² = 0.2449; Gini_right = 0 (pure). Weighted = 7/10 × 0.2449 + 3/10 × 0 = 0.1714.

p0, p1 = 0.6, 0.4
gini_root = 1 - (p0**2 + p1**2)
p0L, p1L = 6/7, 1/7
giniL = 1 - (p0L**2 + p1L**2)
giniR = 0.0  # pure right node
gini_split = 7/10 * giniL + 3/10 * giniR
print(f'root={gini_root:.4f}, split={gini_split:.4f}, gain={gini_root-gini_split:.4f}')
nodesamplesp(class-0)Gini
root100.600.4800
left after split70.8570.2449
right after split30.0000.0000
weighted after split10—0.1714
information gain——0.3086

20. Decode the notation: Derive: Gini impurity and information gain

Notation

Annotate

From Derive: Gini impurity and information gain — read this one piece at a time. What is each part doing?

On: \( \text{Gini}(t) = 1 - \sum_{k} p_k^2 \)

  • p₀ = 0.6, p₁ = 0.4 → Gini = 1 − (0.36 + 0.16) = 0.48
  • Gini_left = 1 − (6/7)² − (1/7)² = 0.2449; Gini_right = 0 (pure). Weighted = 7/10 × 0.2449 + 3/10 × 0 = 0.1714.

21. Guess the shape of the answer: Derive: GBM pseudo-residual update

Estimation

Predict first

GBM adds trees that fit the pseudo-residuals — the gradient of the loss w.r.t. the current prediction. For L2 loss these are just residuals; for log-loss they are yᵢ − σ(Fₜ(xᵢ)).

Commit before you compute: what does Derive: GBM pseudo-residual update come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Compute r₀ = y − F₀ = [−1.25, 0.75, 2.75, −2.25]

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. These are the residuals. For L2 loss the pseudo-residual IS the residual: −∂(y−F)²/∂F = (y−F).

22. Derive: GBM pseudo-residual update

Worked example

GBM adds trees that fit the pseudo-residuals — the gradient of the loss w.r.t. the current prediction. For L2 loss these are just residuals; for log-loss they are yᵢ − σ(Fₜ(xᵢ)).

\[ r_i^{(t)} = -\left[\frac{\partial \mathcal{L}(y_i, F(x_i))}{\partial F(x_i)}\right]_{F=F_{t-1}} \]

Fit a 4-sample dataset: y = [3, 5, 7, 2], F₀ = mean(y) = 4.25

Why: The initial model predicts the constant that minimizes the loss — the mean for L2, log-odds for log-loss.

Compute r₀ = y − F₀ = [−1.25, 0.75, 2.75, −2.25]

Why: These are the residuals. For L2 loss the pseudo-residual IS the residual: −∂(y−F)²/∂F = (y−F).

import numpy as np
from sklearn.tree import DecisionTreeRegressor
y = np.array([3.,5.,7.,2.])
X = np.array([[1.],[2.],[3.],[0.5]])
F0 = y.mean()
r0 = y - F0
h1 = DecisionTreeRegressor(max_depth=1).fit(X, r0)
F1 = F0 + 0.5 * h1.predict(X)
print('F0:', round(F0,4))
print('r0:', r0)
print('F1:', F1.round(4))
print('r1:', (y-F1).round(4))
sampleyF₀r₀F₁ (lr=0.5)r₁
03.04.25-1.253.375-0.375
15.04.250.755.125-0.125
27.04.252.755.1251.875
32.04.25-2.253.375-1.375

23. Decode the notation: Derive: GBM pseudo-residual update

Notation

Annotate

From Derive: GBM pseudo-residual update — read this one piece at a time. What is each part doing?

On: \( r_i^{(t)} = -\left[\frac{\partial \mathcal{L}(y_i, F(x_i))}{\partial F(x_i)}\right]_{F=F_{t-1}} \)

  • The initial model predicts the constant that minimizes the loss — the mean for L2, log-odds for log-loss.
  • These are the residuals. For L2 loss the pseudo-residual IS the residual: −∂(y−F)²/∂F = (y−F).

24. Something is wrong here: treating GBM trees as independent models

Anomaly

Predict first

A student writes this, and it looks reasonable:

GBM's n_estimators=200 means 200 separate trees each trained on the full dataset — like a random forest.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Each GBM tree targets the RESIDUALS of the running prediction, not the original labels.

GBM trees are sequential: tree t+1 fits the pseudo-residuals from tree t.

Why: Each GBM tree targets the RESIDUALS of the running prediction, not the original labels. Tree t+1 sees the errors of trees 1..t — it depends on all previous trees.

25. Trap: treating GBM trees as independent models

Trap

The trap

GBM's n_estimators=200 means 200 separate trees each trained on the full dataset — like a random forest.

Parallelize GBM tree fitting the same way as random forest

Why: Each GBM tree targets the RESIDUALS of the running prediction, not the original labels. Tree t+1 sees the errors of trees 1..t — it depends on all previous trees.

The fix

GBM trees are sequential: tree t+1 fits the pseudo-residuals from tree t.

Fit tree t on r_i^(t), then update F_t = F_{t-1} + lr · h_t(x)

Why: This sequential dependency is why GBM cannot be trivially parallelized (unlike RF's independently bootstrapped trees). Histogram-based GBMs (LightGBM) parallelize the split-finding, not the boosting.

26. Break it on purpose: treating GBM trees as independent models

Break the constraint

Discussion prompt

The rule this trap just fixed:

GBM trees are sequential: tree t+1 fits the pseudo-residuals from tree t.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

Each GBM tree targets the RESIDUALS of the running prediction, not the original labels. Tree t+1 sees the errors of trees 1..t — it depends on all previous trees.

27. kNN and Naive Bayes

Section

Section 4

28. What has to be given first: kNN: distance computation and vote

Missing information

Discussion prompt

kNN stores all training points. At predict time: compute distances, take the k nearest, return majority class.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

Sort Euclidean distances to find the k nearest neighbors.

29. kNN: distance computation and vote

Worked example

kNN stores all training points. At predict time: compute distances, take the k nearest, return majority class.

Query point x=4.0; training set: {1.0→0, 3.5→0, 6.0→1, 8.0→1, 2.5→0}

Why: Sort Euclidean distances to find the k nearest neighbors.

import numpy as np
X_tr = np.array([[1.],[3.5],[6.],[8.],[2.5]])
y_tr = np.array([0,0,1,1,0])
q = 4.0
dists = np.abs(X_tr.flatten() - q)
sorted_i = np.argsort(dists)
for r, i in enumerate(sorted_i[:3]):
    print(f'rank {r+1}: x={X_tr[i,0]}, dist={dists[i]:.4f}, cls={y_tr[i]}')
print('k=3 vote:', np.bincount(y_tr[sorted_i[:3]]).argmax())
rankx_traindistclass
13.50.50000
22.51.50000
36.02.00001

30. Watch it run: kNN: distance computation and vote

Pattern

Step through it

Step through kNN: distance computation and vote one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: rank is 1
  2. Step 2: rank is 2
  3. Step 3: rank is 3

31. Guess the shape of the answer: Naive Bayes: conditional probability

Estimation

Predict first

Naive Bayes applies Bayes' theorem, assuming feature independence: P(C|x) ∝ P(C) · ∏ P(xⱼ|C).

Commit before you compute: what does Naive Bayes: conditional probability come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Classify x=[5.0, 18.0] with priors P(0)=0.7, P(1)=0.3; Gaussian likelihoods

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. C=0: μ=[2,10], σ=[1,3]; C=1: μ=[8,30], σ=[2,5].

32. Naive Bayes: conditional probability

Worked example

Naive Bayes applies Bayes' theorem, assuming feature independence: P(C|x) ∝ P(C) · ∏ P(xⱼ|C).

\[ \hat{y} = \arg\max_c \; P(C=c)\prod_{j=1}^{d} P(x_j \mid C=c) \]

Classify x=[5.0, 18.0] with priors P(0)=0.7, P(1)=0.3; Gaussian likelihoods

Why: C=0: μ=[2,10], σ=[1,3]; C=1: μ=[8,30], σ=[2,5]. Feature independence means we multiply the two Gaussian PDFs.

from scipy.stats import norm
import numpy as np
mu0,s0 = [2.,10.],[1.,3.]; mu1,s1 = [8.,30.],[2.,5.]
x = [5.,18.]; prior = [0.7, 0.3]
p0 = prior[0]*norm.pdf(x[0],mu0[0],s0[0])*norm.pdf(x[1],mu0[1],s0[1])
p1 = prior[1]*norm.pdf(x[0],mu1[0],s1[0])*norm.pdf(x[1],mu1[1],s1[1])
total = p0 + p1
print(f'P(0|x)={p0/total:.4f}, P(1|x)={p1/total:.4f}')
print('predict:', 0 if p0 > p1 else 1)
classpriorP(x|C)unnormalizedP(C|x)
00.700.00001700.00001190.1193
10.300.00028970.00008690.8807

33. Fill in: P(C|x) for Naive Bayes: conditional probability

Comparison

Comparison matrix

From Naive Bayes: conditional probability: refill the P(C|x) column from what you know. The rest of the table is as it appeared.

classpriorP(x|C)unnormalizedP(C|x)
00.700.00001700.00001190.1193
10.300.00028970.00008690.8807

34. Something is wrong here: using Naive Bayes when features are correlated

Anomaly

Predict first

A student writes this, and it looks reasonable:

Naive Bayes is a fast probabilistic classifier — apply it whenever you need a probability estimate.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: NB assumes P(x₁, x₂|C) = P(x₁|C)·P(x₂|C).

Naive Bayes probability rankings are often correct even when magnitudes are wrong — but don't use raw NB probabilities for calibration.

Why: NB assumes P(x₁, x₂|C) = P(x₁|C)·P(x₂|C). When features are correlated, the likelihood is over-counted, making the model over-confident. Probability outputs are miscalibrated.

35. Trap: using Naive Bayes when features are correlated

Trap

The trap

Naive Bayes is a fast probabilistic classifier — apply it whenever you need a probability estimate.

Use NB on text features that are strongly correlated (e.g. 'buy' and 'purchase')

Why: NB assumes P(x₁, x₂|C) = P(x₁|C)·P(x₂|C). When features are correlated, the likelihood is over-counted, making the model over-confident. Probability outputs are miscalibrated.

The fix

Naive Bayes probability rankings are often correct even when magnitudes are wrong — but don't use raw NB probabilities for calibration.

Use NB for fast, high-bias baselines or text with near-independent n-grams; calibrate with Platt scaling if you need probabilities

Why: NB decision boundaries are often good even with violations of independence. The 'naive' assumption hurts calibration more than classification accuracy.

36. Which of these survive contact with Lesson 76: Algorithm Review & Selection?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Benchmark on make_classification(n=500, n_features=10, random_state=42): RF 0.93, GBM 0.92, SVM 0.91, kNN 0.86, NB 0.85, DT 0.84, LR 0.82.; SVM finds the maximum-margin separating hyperplane. The primal objective balances margin width against slack violations.; Build rule: implement each algorithm from scratch in a fresh cell; compare test accuracies in a single summary table at the end.
Breaks
Large C means stronger regularization — the model is more conservative.; GBM's n_estimators=200 means 200 separate trees each trained on the full dataset — like a random forest.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 76: Algorithm Review & Selection puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

37. Rebuild the recipe: Algorithm selection framework

Ranking

Put in order

These are the steps of Algorithm selection framework, scrambled. Put them back in order before the next slide shows you.

  1. Data size & speed: large n → GBM/RF (parallelizable) or SGD-LR; very small → NB/LR
  2. Linearity: linear boundary likely → LR or linear SVM; nonlinear → RBF-SVM, GBM, RF
  3. Interpretability needed: DT (human-readable rules); otherwise GBM or RF for accuracy
  4. Few labeled examples: kNN or NB (no gradient, no overfitting risk from optimization)
  5. Feature engineering done: LR; noisy high-dim features: RF/GBM handle them natively
  6. Probability output required (calibrated): LR; post-hoc calibrate SVM/RF/GBM if needed

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

38. Algorithm selection framework

Pattern

  1. Data size & speed: large n → GBM/RF (parallelizable) or SGD-LR; very small → NB/LR
  2. Linearity: linear boundary likely → LR or linear SVM; nonlinear → RBF-SVM, GBM, RF
  3. Interpretability needed: DT (human-readable rules); otherwise GBM or RF for accuracy
  4. Few labeled examples: kNN or NB (no gradient, no overfitting risk from optimization)
  5. Feature engineering done: LR; noisy high-dim features: RF/GBM handle them natively
  6. Probability output required (calibrated): LR; post-hoc calibrate SVM/RF/GBM if needed
problem typefirst tryif needs more power
tabular, mix of feature typesRFGBM
text classificationNB baselineLogistic + TF-IDF
small structured datasetLRSVM (RBF)
anomaly / density estimationkNNIsolation Forest (L60)
exam: 'derive gradient'LR or GBM—
exam: 'prove margin property'SVM dual—

39. Where does it stop working: Algorithm selection framework

Edge cases

Discussion prompt

Algorithm selection framework works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Data size & speed: large n → GBM/RF (parallelizable) or SGD-LR; very small → NB/LR
  2. Linearity: linear boundary likely → LR or linear SVM; nonlinear → RBF-SVM, GBM, RF
  3. Interpretability needed: DT (human-readable rules); otherwise GBM or RF for accuracy
  4. Few labeled examples: kNN or NB (no gradient, no overfitting risk from optimization)
  5. Feature engineering done: LR; noisy high-dim features: RF/GBM handle them natively
  6. Probability output required (calibrated): LR; post-hoc calibrate SVM/RF/GBM if needed

40. Rule out three: Check yourself — GBM vs Random Forest

Elimination

Eliminate the wrong options

Which statement correctly distinguishes GBM from Random Forest?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. GBM trains trees sequentially on pseudo-residuals; RF trains trees independently on bootstrap samples
  • B. GBM uses bootstrap sampling; RF uses additive sequential trees
  • C. Both use identical tree-fitting procedures; they differ only in final aggregation
  • D. Random Forest minimizes a differentiable loss; GBM does not use a loss function

Survives elimination: A

Why: GBM is a boosting method — each tree fits the gradient of the loss (pseudo-residuals) of the ensemble so far, making trees sequential and dependent. RF is a bagging method — each tree trains independently on a bootstrap sample and random feature subset. RF is easily parallelized; GBM is not.

41. Check yourself — GBM vs Random Forest

Check

Know what makes boosting and bagging fundamentally different.

Check your understanding

Which statement correctly distinguishes GBM from Random Forest?

  • A. GBM trains trees sequentially on pseudo-residuals; RF trains trees independently on bootstrap samples (correct)
  • B. GBM uses bootstrap sampling; RF uses additive sequential trees
  • C. Both use identical tree-fitting procedures; they differ only in final aggregation
  • D. Random Forest minimizes a differentiable loss; GBM does not use a loss function

Answer: A

Why: GBM is a boosting method — each tree fits the gradient of the loss (pseudo-residuals) of the ensemble so far, making trees sequential and dependent. RF is a bagging method — each tree trains independently on a bootstrap sample and random feature subset. RF is easily parallelized; GBM is not.

Why B tempts people
These descriptions are swapped. GBM is sequential/pseudo-residuals; RF uses bootstrap sampling.
Why C tempts people
The tree-fitting procedure IS different: GBM trees target pseudo-residuals, RF trees target original labels on bootstrap samples.
Why D tempts people
GBM explicitly minimizes a differentiable loss function via functional gradient descent — it is the most loss-centric ensemble. RF has no explicit global loss.

42. Answer it before you see the options: Check yourself — SVM hyperparameter C

Prediction

Predict first

An SVM trained with C=0.1 has margin=1.44 and 28 support vectors. You retrain with C=10. What do you expect?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: Narrower margin, fewer support vectors (harder boundary)

Why: Large C makes the SVM penalize margin violations heavily, so it squeezes the margin tighter to fit more points correctly. Result: narrower margin, fewer support vectors, higher risk of overfitting. Verified: C=0.1 → 28 SV / margin 1.44; C=10 → 4 SV / margin 0.43.

43. Check yourself — SVM hyperparameter C

Check

Trace the effect of C on margin width.

Check your understanding

An SVM trained with C=0.1 has margin=1.44 and 28 support vectors. You retrain with C=10. What do you expect?

  • A. Narrower margin, fewer support vectors (harder boundary) (correct)
  • B. Wider margin, more support vectors (softer boundary)
  • C. Same margin; only the kernel changes
  • D. Narrower margin, more support vectors (softer boundary)

Answer: A

Why: Large C makes the SVM penalize margin violations heavily, so it squeezes the margin tighter to fit more points correctly. Result: narrower margin, fewer support vectors, higher risk of overfitting. Verified: C=0.1 → 28 SV / margin 1.44; C=10 → 4 SV / margin 0.43.

Why B tempts people
A wider margin with more SVs is the SMALL-C behavior (soft margin). Large C narrows the margin.
Why C tempts people
C and the kernel are orthogonal hyperparameters; changing C changes the margin directly.
Why D tempts people
The margin does narrow (correct), but larger C means FEWER violations tolerated → fewer support vectors, not more.

44. Rule out three: Check yourself — kNN complexity

Elimination

Eliminate the wrong options

For kNN with n training points, d features, and a single query: what is the prediction-time complexity?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. O(n·d) — compute all pairwise distances
  • B. O(log n) — like a binary search
  • C. O(k·d) — only the k nearest matter
  • D. O(n²·d) — brute force pairwise

Survives elimination: A

Why: Naive kNN computes the distance from the query to all n training points, each taking O(d) time → O(n·d) per query. Training is O(1) (just store the data). With KD-trees or ball trees the average drops to O(d·log n), but the worst case is still O(n·d).

45. Check yourself — kNN complexity

Check

Know the train and predict complexity.

Check your understanding

For kNN with n training points, d features, and a single query: what is the prediction-time complexity?

  • A. O(n·d) — compute all pairwise distances (correct)
  • B. O(log n) — like a binary search
  • C. O(k·d) — only the k nearest matter
  • D. O(n²·d) — brute force pairwise

Answer: A

Why: Naive kNN computes the distance from the query to all n training points, each taking O(d) time → O(n·d) per query. Training is O(1) (just store the data). With KD-trees or ball trees the average drops to O(d·log n), but the worst case is still O(n·d).

Why B tempts people
O(log n) applies to sorted structures (KD-tree average case), not naive kNN — and even KD-tree degrades to O(n) in high dimensions.
Why C tempts people
O(k·d) would require knowing the k nearest in advance; finding them requires scanning all n points first.
Why D tempts people
O(n²·d) is the complexity of computing ALL pairwise distances in the training set — useful for offline pre-computation but not a single query.

46. Your turn: implement all 7

Section

Project

47. Project: 7-algorithm benchmark

Concept

Implement and evaluate all 7 algorithms on the same synthetic dataset. Then, for each, answer from memory: loss function, gradient, two key hyperparameters, and bias-variance profile.

#milestonefocus
1Logistic regression + manual gradient tracederive and verify one GD step
2SVM: compare C=0.1 vs C=10margin width and support vector count
3Decision tree: compute one Gini split manuallyinformation gain derivation
4RF + GBM: feature importance + pseudo-residual stepensemble mechanics
5kNN: distance table; NB: posterior calculationlazy learning + independence assumption

Build rule: implement each algorithm from scratch in a fresh cell; compare test accuracies in a single summary table at the end.

48. Break it if you can: Project: 7-algorithm benchmark

Counterexample

Discussion prompt

Implement and evaluate all 7 algorithms on the same synthetic dataset. Then, for each, answer from memory: loss function, gradient, two key hyperparameters, and bias-variance profile.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Build rule: implement each algorithm from scratch in a fresh cell; compare test accuracies in a single summary table at the end.

49. Milestone 1 — Logistic regression gradient step

Worked example

Your turn: load make_classification(n=500, random_state=42), fit LogisticRegression, then manually compute one gradient step on the first mini-batch (n=5) and verify the gradient formula.

Hint: p = sigmoid(X @ w), grad = X.T @ (p - y) / n, w_new = w - lr * grad. Compare your grad to lr.coef_.

import numpy as np
from sklearn.datasets import make_classification
from sklearn.preprocessing import StandardScaler
def sigmoid(z): return 1/(1+np.exp(-z))
X, y = make_classification(n_samples=500, n_features=10, random_state=42)
X = StandardScaler().fit_transform(X)
Xm, ym = X[:5,:2], y[:5].astype(float)
w = np.zeros(2)
p = sigmoid(Xm @ w)
grad = Xm.T @ (p - ym) / 5
print('grad:', grad.round(6))
print('w after step (lr=0.5):', (w - 0.5*grad).round(6))
quantityvalue
p at w=0 (all 0.5)0.5, 0.5, 0.5, 0.5, 0.5
grad[0]-0.125 (approx, depends on mini-batch)
w[0] after 1 step (lr=0.5)≈ 0.0625

50. Milestone 2 — SVM margin vs C

Worked example

Your turn: fit SVC(kernel='linear') for C ∈ {0.1, 1.0, 10.0} and print margin = 2/‖w‖ and n_support_vectors.

Hint: svm.coef_[0] gives w; len(svm.support_vectors_) gives the count. Expect margin to shrink and SV count to drop as C grows.

from sklearn.svm import SVC
from sklearn.datasets import load_iris
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split
import numpy as np
X, y = load_iris(return_X_y=True)
X, y = X[y<2,:2], y[y<2]
Xtr,Xte,ytr,yte = train_test_split(X, y, test_size=0.2, random_state=0)
sc = StandardScaler(); Xtr = sc.fit_transform(Xtr); Xte = sc.transform(Xte)
for C in [0.1, 1.0, 10.0]:
    m = SVC(kernel='linear', C=C).fit(Xtr, ytr)
    print(f'C={C}: n_sv={len(m.support_vectors_)}, margin={2/np.linalg.norm(m.coef_[0]):.4f}')
Cn_svmargin
0.1281.4447
1.0100.7800
10.040.4303

51. What each one costs: Milestone 2 — SVM margin vs C

Trade off

Comparison matrix

From Milestone 2 — SVM margin vs C: every row here is a choice with a cost. Fill the n_sv column, then say which row you would actually pick and what you give up for it.

Cn_svmargin
0.1281.4447
1.0100.7800
10.040.4303

52. Milestones 3-5 — DT / RF / GBM / kNN / NB

Worked example

Your turn: fit all remaining algorithms on make_classification(n=500, random_state=42) and record test accuracy for each.

Hint: scale features for LR, SVM, kNN. Don't scale for DT, RF, GBM, NB. Use accuracy_score(y_test, model.predict(X_test)).

from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.tree import DecisionTreeClassifier
from sklearn.ensemble import RandomForestClassifier, GradientBoostingClassifier
from sklearn.neighbors import KNeighborsClassifier
from sklearn.naive_bayes import GaussianNB
from sklearn.metrics import accuracy_score
X, y = make_classification(n_samples=500, n_features=10, random_state=42)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.2, random_state=42)
sc = StandardScaler(); Xsc_tr = sc.fit_transform(Xtr); Xsc_te = sc.transform(Xte)
models = {'DT':DecisionTreeClassifier(max_depth=3,random_state=42),
          'RF':RandomForestClassifier(n_estimators=100,random_state=42),
          'GBM':GradientBoostingClassifier(n_estimators=100,random_state=42),
          'kNN':KNeighborsClassifier(n_neighbors=5),
          'NB':GaussianNB()}
for name, m in models.items():
    Xf = Xsc_tr if name in ('kNN',) else Xtr
    Xt = Xsc_te if name in ('kNN',) else Xte
    m.fit(Xf, ytr)
    print(f'{name}: {accuracy_score(yte, m.predict(Xt)):.4f}')
algorithmtest accuracy
Decision Tree (max_depth=3)0.8400
Random Forest (100 trees)0.9300
GBM (100 trees, lr=0.1)0.9200
kNN (k=5)0.8600
Naive Bayes (Gaussian)0.8500

53. Fill in: test accuracy for Milestones 3-5 — DT / RF / GBM / kNN / NB

Comparison

Comparison matrix

From Milestones 3-5 — DT / RF / GBM / kNN / NB: refill the test accuracy column from what you know. The rest of the table is as it appeared.

algorithmtest accuracy
Decision Tree (max_depth=3)0.8400
Random Forest (100 trees)0.9300
GBM (100 trees, lr=0.1)0.9200
kNN (k=5)0.8600
Naive Bayes (Gaussian)0.8500

54. Show it off

Concept

Close these slides. For each of the 7 algorithms, state out loud: (1) the loss function, (2) the gradient or update rule, (3) two key hyperparameters, (4) whether it scales with feature space, and (5) when you would NOT choose it.

Stretch (homework): implement the USAAIO patterns cold — (a) derive logistic gradient from scratch on paper; (b) derive Gini gain for a 3-class node; (c) write a kNN classifier in <20 lines; (d) choose the best algorithm for: 1M rows tabular, high dimensionality text, 50-sample medical dataset, real-time streaming.

55. Connect it up: Lesson 76: Algorithm Review & Selection

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — The 7 algorithms at a glance · Derivations — logistic & SVM · Trees, forests, and boosting · kNN and Naive Bayes · Your turn: implement all 7. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

56. What you can do now

Recap

algorithmone-line memory hook
Logistic Regressionsigmoid(Xw); gradient = Xᵀ(σ−y)/n; regularize with C
SVMmax margin = 2/‖w‖; large C → narrow margin, fewer SVs
Decision Treeminimize Gini; gain = Gini_parent − weighted Gini_children
Random Forestbootstrap + random features; reduces variance over one tree
GBMfit each tree on pseudo-residuals; sequential, not parallel
kNNlazy; O(n·d) predict; small k = high variance
Naive Bayes∝ prior × product of likelihoods; assumes feature independence

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 76 (Phase 2 — Algorithm Review & Selection) — Barron · USAAIO Round 2 Preparation, 2026
  2. All accuracy values, gradient traces, Gini computations, kNN distances, and NB probabilities verified with sklearn 1.x + numpy 2.2.6, real execution, June 2026 — python verify_lesson76.py + verify_lesson76b.py, June 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108