USAAIO Lesson 76, a Phase 2 review. It is a flash review of seven core classifiers - logistic regression, the SVM, decision trees, random forests, GBMs, kNN, and Naive Bayes - covering their loss functions, gradients, hyperparameters, and bias-variance profiles, with whiteboard-style derivations. It also covers the USAAIO exam question patterns and gives an algorithm-selection framework. You implement and benchmark all seven on a synthetic dataset. The lesson runs to 28 slides.
Subject: Machine Learning · 56 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 76 · Phase 2 Review
All 7 core classifiers in one session: loss, gradient, hyperparameters, bias-variance, and a framework for choosing the right algorithm under exam conditions.
Objectives
Warm-up
Discussion prompt
Before we open Lesson 76: Algorithm Review & Selection: without looking back, what was the main idea of Mock Exam — Phase 2 Full Review, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
a timed Phase 2 mock exam — 2-hour theory exam on ensemble methods, normalization, optimization, and evaluation; 2-hour coding exam implementing random forest, GMM, and CNN training loop from scratch; full answer-key walkthrough and gap analysis before Phase 3 (transformers).
Section
Section 1
Concept
| algorithm | loss | optimization | key hyperparams |
|---|---|---|---|
| Logistic Regression | log-loss (cross-entropy) | gradient descent (L-BFGS/SGD) | C, solver, max_iter |
| SVM | hinge loss + C·slack | quadratic programming (SMO) | C, kernel, γ (RBF) |
| Decision Tree | Gini / entropy impurity | greedy split search | max_depth, min_samples |
| Random Forest | avg Gini across trees | bootstrap + greedy | n_estimators, max_features |
| GBM | deviance (log-loss/MSE) | gradient boosting (additive) | n_estimators, lr, max_depth |
| kNN | none (lazy learner) | none (distance at predict) | k, distance metric |
| Naive Bayes | MLE per class/feature | closed form (analytically) | var_smoothing |
Comparison
Comparison matrix
From Loss functions & optimization — the master table: refill the loss column from what you know. The rest of the table is as it appeared.
| algorithm | loss | optimization | key hyperparams |
|---|---|---|---|
| Logistic Regression | log-loss (cross-entropy) | gradient descent (L-BFGS/SGD) | C, solver, max_iter |
| SVM | hinge loss + C·slack | quadratic programming (SMO) | C, kernel, γ (RBF) |
| Decision Tree | Gini / entropy impurity | greedy split search | max_depth, min_samples |
| Random Forest | avg Gini across trees | bootstrap + greedy | n_estimators, max_features |
| GBM | deviance (log-loss/MSE) | gradient boosting (additive) | n_estimators, lr, max_depth |
| kNN | none (lazy learner) | none (distance at predict) | k, distance metric |
| Naive Bayes | MLE per class/feature | closed form (analytically) | var_smoothing |
Concept
| algorithm | bias | variance | overfits? | needs scaling? |
|---|---|---|---|---|
| Logistic Regression | moderate (linear boundary) | low | rarely | yes |
| SVM (RBF) | low | moderate | if C too high | yes |
| Decision Tree | low (deep) / high (shallow) | high (deep) | easily | no |
| Random Forest | low | reduced vs tree | less than DT | no |
| GBM | low | low–moderate | if lr high + many trees | no |
| kNN | low | high (small k) | small k | yes |
| Naive Bayes | high (independence) | very low | rarely | no |
Benchmark on make_classification(n=500, n_features=10, random_state=42): RF 0.93, GBM 0.92, SVM 0.91, kNN 0.86, NB 0.85, DT 0.84, LR 0.82.
Discrimination
Sort into buckets
Sort these by needs scaling?, from memory, without looking back at Bias-variance profiles. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Section
Section 2
Estimation
Predict first
USAAIO pattern: 'derive the gradient of logistic loss'. Start from the definition of cross-entropy loss.
Commit before you compute: what does Derive: gradient of logistic loss come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Differentiate with respect to w using the chain rule
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. σ'(z) = σ(z)(1−σ(z)); applying the chain rule through log collapses the expression cleanly.
Worked example
USAAIO pattern: 'derive the gradient of logistic loss'. Start from the definition of cross-entropy loss.
\[ \mathcal{L}(w) = -\tfrac{1}{n}\sum_{i=1}^{n}\bigl[y_i\log\sigma(x_i^\top w) + (1-y_i)\log(1-\sigma(x_i^\top w))\bigr] \]
Differentiate with respect to w using the chain rule
Why: σ'(z) = σ(z)(1−σ(z)); applying the chain rule through log collapses the expression cleanly.
\[ \nabla_w \mathcal{L} = \tfrac{1}{n} X^\top (\sigma(Xw) - y) \]
import numpy as np
def sigmoid(z): return 1 / (1 + np.exp(-z))
X = np.array([[1.,0.5],[2.,1.],[3.,1.5],[-1.,-0.5]])
y = np.array([0.,0.,1.,0.])
w = np.zeros(2); lr = 0.5
for step in range(3):
p = sigmoid(X @ w)
loss = -np.mean(y*np.log(p+1e-10)+(1-y)*np.log(1-p+1e-10))
grad = X.T @ (p - y) / len(y)
w = w - lr * grad
print(f'step={step}: loss={loss:.4f}, w={w.round(4)}')| step | loss | w[0] | w[1] |
|---|---|---|---|
| 0 | 0.6931 | 0.0625 | 0.0312 |
| 1 | 0.6862 | 0.0885 | 0.0443 |
| 2 | 0.6850 | 0.0995 | 0.0497 |
Trade off
Comparison matrix
From Derive: gradient of logistic loss: every row here is a choice with a cost. Fill the loss column, then say which row you would actually pick and what you give up for it.
| step | loss | w[0] | w[1] |
|---|---|---|---|
| 0 | 0.6931 | 0.0625 | 0.0312 |
| 1 | 0.6862 | 0.0885 | 0.0443 |
| 2 | 0.6850 | 0.0995 | 0.0497 |
Concept
SVM finds the maximum-margin separating hyperplane. The primal objective balances margin width against slack violations.
\[ \min_{w,b,\xi}\;\tfrac{1}{2}\lVert w\rVert^2 + C\sum_i \xi_i \quad\text{s.t.}\; y_i(w^\top x_i+b)\geq 1-\xi_i,\;\xi_i\geq 0 \]
C controls the bias-variance tradeoff: large C → narrow margin (low bias, high variance); small C → wide margin (high bias, low variance).
| C | n_sv | margin | test acc (iris 2-class) |
|---|---|---|---|
| 0.1 | 28 | 1.4447 | 1.0000 |
| 1.0 | 10 | 0.7800 | 1.0000 |
| 10.0 | 4 | 0.4303 | 1.0000 |
Invariant
Step through it
Step through SVM: hinge loss and the margin one row at a time. One of these columns never changes — find it, and say why it cannot.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Large C means stronger regularization — the model is more conservative.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: This is the L2 intuition from Ridge/logistic (Lesson 12).
Large C means LESS regularization — the model tolerates fewer margin violations.
Why: This is the L2 intuition from Ridge/logistic (Lesson 12). But SVM's C is the INVERSE of regularization strength — the penalty is on slack, not on w.
Trap
Large C means stronger regularization — the model is more conservative.
Increase C to reduce overfitting
Why: This is the L2 intuition from Ridge/logistic (Lesson 12). But SVM's C is the INVERSE of regularization strength — the penalty is on slack, not on w.
Large C means LESS regularization — the model tolerates fewer margin violations.
Decrease C to reduce overfitting in SVM
Why: Small C allows a wider margin (larger ‖w‖ penalty implicit), accepting some misclassifications for a smoother boundary. Think: C = cost of violations, not cost of complexity.
Section
Section 3
Missing information
Discussion prompt
USAAIO pattern: 'prove that this split maximizes information gain'. Start from the Gini definition.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
p₀ = 0.6, p₁ = 0.4 → Gini = 1 − (0.36 + 0.16) = 0.48
Worked example
USAAIO pattern: 'prove that this split maximizes information gain'. Start from the Gini definition.
\[ \text{Gini}(t) = 1 - \sum_{k} p_k^2 \]
Compute Gini at the root: 10 samples, 6 class-0, 4 class-1
Why: p₀ = 0.6, p₁ = 0.4 → Gini = 1 − (0.36 + 0.16) = 0.48
Evaluate a split: left (7 samples, 6/1), right (3 samples, 0/3)
Why: Gini_left = 1 − (6/7)² − (1/7)² = 0.2449; Gini_right = 0 (pure). Weighted = 7/10 × 0.2449 + 3/10 × 0 = 0.1714.
p0, p1 = 0.6, 0.4
gini_root = 1 - (p0**2 + p1**2)
p0L, p1L = 6/7, 1/7
giniL = 1 - (p0L**2 + p1L**2)
giniR = 0.0 # pure right node
gini_split = 7/10 * giniL + 3/10 * giniR
print(f'root={gini_root:.4f}, split={gini_split:.4f}, gain={gini_root-gini_split:.4f}')| node | samples | p(class-0) | Gini |
|---|---|---|---|
| root | 10 | 0.60 | 0.4800 |
| left after split | 7 | 0.857 | 0.2449 |
| right after split | 3 | 0.000 | 0.0000 |
| weighted after split | 10 | — | 0.1714 |
| information gain | — | — | 0.3086 |
Notation
Annotate
From Derive: Gini impurity and information gain — read this one piece at a time. What is each part doing?
On: \( \text{Gini}(t) = 1 - \sum_{k} p_k^2 \)
Estimation
Predict first
GBM adds trees that fit the pseudo-residuals — the gradient of the loss w.r.t. the current prediction. For L2 loss these are just residuals; for log-loss they are yᵢ − σ(Fₜ(xᵢ)).
Commit before you compute: what does Derive: GBM pseudo-residual update come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Compute r₀ = y − F₀ = [−1.25, 0.75, 2.75, −2.25]
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. These are the residuals. For L2 loss the pseudo-residual IS the residual: −∂(y−F)²/∂F = (y−F).
Worked example
GBM adds trees that fit the pseudo-residuals — the gradient of the loss w.r.t. the current prediction. For L2 loss these are just residuals; for log-loss they are yᵢ − σ(Fₜ(xᵢ)).
\[ r_i^{(t)} = -\left[\frac{\partial \mathcal{L}(y_i, F(x_i))}{\partial F(x_i)}\right]_{F=F_{t-1}} \]
Fit a 4-sample dataset: y = [3, 5, 7, 2], F₀ = mean(y) = 4.25
Why: The initial model predicts the constant that minimizes the loss — the mean for L2, log-odds for log-loss.
Compute r₀ = y − F₀ = [−1.25, 0.75, 2.75, −2.25]
Why: These are the residuals. For L2 loss the pseudo-residual IS the residual: −∂(y−F)²/∂F = (y−F).
import numpy as np
from sklearn.tree import DecisionTreeRegressor
y = np.array([3.,5.,7.,2.])
X = np.array([[1.],[2.],[3.],[0.5]])
F0 = y.mean()
r0 = y - F0
h1 = DecisionTreeRegressor(max_depth=1).fit(X, r0)
F1 = F0 + 0.5 * h1.predict(X)
print('F0:', round(F0,4))
print('r0:', r0)
print('F1:', F1.round(4))
print('r1:', (y-F1).round(4))| sample | y | F₀ | r₀ | F₁ (lr=0.5) | r₁ |
|---|---|---|---|---|---|
| 0 | 3.0 | 4.25 | -1.25 | 3.375 | -0.375 |
| 1 | 5.0 | 4.25 | 0.75 | 5.125 | -0.125 |
| 2 | 7.0 | 4.25 | 2.75 | 5.125 | 1.875 |
| 3 | 2.0 | 4.25 | -2.25 | 3.375 | -1.375 |
Notation
Annotate
From Derive: GBM pseudo-residual update — read this one piece at a time. What is each part doing?
On: \( r_i^{(t)} = -\left[\frac{\partial \mathcal{L}(y_i, F(x_i))}{\partial F(x_i)}\right]_{F=F_{t-1}} \)
Anomaly
Predict first
A student writes this, and it looks reasonable:
GBM's n_estimators=200 means 200 separate trees each trained on the full dataset — like a random forest.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Each GBM tree targets the RESIDUALS of the running prediction, not the original labels.
GBM trees are sequential: tree t+1 fits the pseudo-residuals from tree t.
Why: Each GBM tree targets the RESIDUALS of the running prediction, not the original labels. Tree t+1 sees the errors of trees 1..t — it depends on all previous trees.
Trap
GBM's n_estimators=200 means 200 separate trees each trained on the full dataset — like a random forest.
Parallelize GBM tree fitting the same way as random forest
Why: Each GBM tree targets the RESIDUALS of the running prediction, not the original labels. Tree t+1 sees the errors of trees 1..t — it depends on all previous trees.
GBM trees are sequential: tree t+1 fits the pseudo-residuals from tree t.
Fit tree t on r_i^(t), then update F_t = F_{t-1} + lr · h_t(x)
Why: This sequential dependency is why GBM cannot be trivially parallelized (unlike RF's independently bootstrapped trees). Histogram-based GBMs (LightGBM) parallelize the split-finding, not the boosting.
Break the constraint
Discussion prompt
The rule this trap just fixed:
GBM trees are sequential: tree t+1 fits the pseudo-residuals from tree t.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
Each GBM tree targets the RESIDUALS of the running prediction, not the original labels. Tree t+1 sees the errors of trees 1..t — it depends on all previous trees.
Section
Section 4
Missing information
Discussion prompt
kNN stores all training points. At predict time: compute distances, take the k nearest, return majority class.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Sort Euclidean distances to find the k nearest neighbors.
Worked example
kNN stores all training points. At predict time: compute distances, take the k nearest, return majority class.
Query point x=4.0; training set: {1.0→0, 3.5→0, 6.0→1, 8.0→1, 2.5→0}
Why: Sort Euclidean distances to find the k nearest neighbors.
import numpy as np
X_tr = np.array([[1.],[3.5],[6.],[8.],[2.5]])
y_tr = np.array([0,0,1,1,0])
q = 4.0
dists = np.abs(X_tr.flatten() - q)
sorted_i = np.argsort(dists)
for r, i in enumerate(sorted_i[:3]):
print(f'rank {r+1}: x={X_tr[i,0]}, dist={dists[i]:.4f}, cls={y_tr[i]}')
print('k=3 vote:', np.bincount(y_tr[sorted_i[:3]]).argmax())| rank | x_train | dist | class |
|---|---|---|---|
| 1 | 3.5 | 0.5000 | 0 |
| 2 | 2.5 | 1.5000 | 0 |
| 3 | 6.0 | 2.0000 | 1 |
Pattern
Step through it
Step through kNN: distance computation and vote one row at a time. What is driving the change, and what would the row after the last one be?
Estimation
Predict first
Naive Bayes applies Bayes' theorem, assuming feature independence: P(C|x) ∝ P(C) · ∏ P(xⱼ|C).
Commit before you compute: what does Naive Bayes: conditional probability come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Classify x=[5.0, 18.0] with priors P(0)=0.7, P(1)=0.3; Gaussian likelihoods
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. C=0: μ=[2,10], σ=[1,3]; C=1: μ=[8,30], σ=[2,5].
Worked example
Naive Bayes applies Bayes' theorem, assuming feature independence: P(C|x) ∝ P(C) · ∏ P(xⱼ|C).
\[ \hat{y} = \arg\max_c \; P(C=c)\prod_{j=1}^{d} P(x_j \mid C=c) \]
Classify x=[5.0, 18.0] with priors P(0)=0.7, P(1)=0.3; Gaussian likelihoods
Why: C=0: μ=[2,10], σ=[1,3]; C=1: μ=[8,30], σ=[2,5]. Feature independence means we multiply the two Gaussian PDFs.
from scipy.stats import norm
import numpy as np
mu0,s0 = [2.,10.],[1.,3.]; mu1,s1 = [8.,30.],[2.,5.]
x = [5.,18.]; prior = [0.7, 0.3]
p0 = prior[0]*norm.pdf(x[0],mu0[0],s0[0])*norm.pdf(x[1],mu0[1],s0[1])
p1 = prior[1]*norm.pdf(x[0],mu1[0],s1[0])*norm.pdf(x[1],mu1[1],s1[1])
total = p0 + p1
print(f'P(0|x)={p0/total:.4f}, P(1|x)={p1/total:.4f}')
print('predict:', 0 if p0 > p1 else 1)| class | prior | P(x|C) | unnormalized | P(C|x) |
|---|---|---|---|---|
| 0 | 0.70 | 0.0000170 | 0.0000119 | 0.1193 |
| 1 | 0.30 | 0.0002897 | 0.0000869 | 0.8807 |
Comparison
Comparison matrix
From Naive Bayes: conditional probability: refill the P(C|x) column from what you know. The rest of the table is as it appeared.
| class | prior | P(x|C) | unnormalized | P(C|x) |
|---|---|---|---|---|
| 0 | 0.70 | 0.0000170 | 0.0000119 | 0.1193 |
| 1 | 0.30 | 0.0002897 | 0.0000869 | 0.8807 |
Anomaly
Predict first
A student writes this, and it looks reasonable:
Naive Bayes is a fast probabilistic classifier — apply it whenever you need a probability estimate.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: NB assumes P(x₁, x₂|C) = P(x₁|C)·P(x₂|C).
Naive Bayes probability rankings are often correct even when magnitudes are wrong — but don't use raw NB probabilities for calibration.
Why: NB assumes P(x₁, x₂|C) = P(x₁|C)·P(x₂|C). When features are correlated, the likelihood is over-counted, making the model over-confident. Probability outputs are miscalibrated.
Trap
Naive Bayes is a fast probabilistic classifier — apply it whenever you need a probability estimate.
Use NB on text features that are strongly correlated (e.g. 'buy' and 'purchase')
Why: NB assumes P(x₁, x₂|C) = P(x₁|C)·P(x₂|C). When features are correlated, the likelihood is over-counted, making the model over-confident. Probability outputs are miscalibrated.
Naive Bayes probability rankings are often correct even when magnitudes are wrong — but don't use raw NB probabilities for calibration.
Use NB for fast, high-bias baselines or text with near-independent n-grams; calibrate with Platt scaling if you need probabilities
Why: NB decision boundaries are often good even with violations of independence. The 'naive' assumption hurts calibration more than classification accuracy.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
make_classification(n=500, n_features=10, random_state=42): RF 0.93, GBM 0.92, SVM 0.91, kNN 0.86, NB 0.85, DT 0.84, LR 0.82.; SVM finds the maximum-margin separating hyperplane. The primal objective balances margin width against slack violations.; Build rule: implement each algorithm from scratch in a fresh cell; compare test accuracies in a single summary table at the end.Ranking
Put in order
These are the steps of Algorithm selection framework, scrambled. Put them back in order before the next slide shows you.
Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
| problem type | first try | if needs more power |
|---|---|---|
| tabular, mix of feature types | RF | GBM |
| text classification | NB baseline | Logistic + TF-IDF |
| small structured dataset | LR | SVM (RBF) |
| anomaly / density estimation | kNN | Isolation Forest (L60) |
| exam: 'derive gradient' | LR or GBM | — |
| exam: 'prove margin property' | SVM dual | — |
Edge cases
Discussion prompt
Algorithm selection framework works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
Elimination
Eliminate the wrong options
Which statement correctly distinguishes GBM from Random Forest?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: GBM is a boosting method — each tree fits the gradient of the loss (pseudo-residuals) of the ensemble so far, making trees sequential and dependent. RF is a bagging method — each tree trains independently on a bootstrap sample and random feature subset. RF is easily parallelized; GBM is not.
Check
Know what makes boosting and bagging fundamentally different.
Check your understanding
Which statement correctly distinguishes GBM from Random Forest?
Answer: A
Why: GBM is a boosting method — each tree fits the gradient of the loss (pseudo-residuals) of the ensemble so far, making trees sequential and dependent. RF is a bagging method — each tree trains independently on a bootstrap sample and random feature subset. RF is easily parallelized; GBM is not.
Prediction
Predict first
An SVM trained with C=0.1 has margin=1.44 and 28 support vectors. You retrain with C=10. What do you expect?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Narrower margin, fewer support vectors (harder boundary)
Why: Large C makes the SVM penalize margin violations heavily, so it squeezes the margin tighter to fit more points correctly. Result: narrower margin, fewer support vectors, higher risk of overfitting. Verified: C=0.1 → 28 SV / margin 1.44; C=10 → 4 SV / margin 0.43.
Check
Trace the effect of C on margin width.
Check your understanding
An SVM trained with C=0.1 has margin=1.44 and 28 support vectors. You retrain with C=10. What do you expect?
Answer: A
Why: Large C makes the SVM penalize margin violations heavily, so it squeezes the margin tighter to fit more points correctly. Result: narrower margin, fewer support vectors, higher risk of overfitting. Verified: C=0.1 → 28 SV / margin 1.44; C=10 → 4 SV / margin 0.43.
Elimination
Eliminate the wrong options
For kNN with n training points, d features, and a single query: what is the prediction-time complexity?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Naive kNN computes the distance from the query to all n training points, each taking O(d) time → O(n·d) per query. Training is O(1) (just store the data). With KD-trees or ball trees the average drops to O(d·log n), but the worst case is still O(n·d).
Check
Know the train and predict complexity.
Check your understanding
For kNN with n training points, d features, and a single query: what is the prediction-time complexity?
Answer: A
Why: Naive kNN computes the distance from the query to all n training points, each taking O(d) time → O(n·d) per query. Training is O(1) (just store the data). With KD-trees or ball trees the average drops to O(d·log n), but the worst case is still O(n·d).
Section
Project
Concept
Implement and evaluate all 7 algorithms on the same synthetic dataset. Then, for each, answer from memory: loss function, gradient, two key hyperparameters, and bias-variance profile.
| # | milestone | focus |
|---|---|---|
| 1 | Logistic regression + manual gradient trace | derive and verify one GD step |
| 2 | SVM: compare C=0.1 vs C=10 | margin width and support vector count |
| 3 | Decision tree: compute one Gini split manually | information gain derivation |
| 4 | RF + GBM: feature importance + pseudo-residual step | ensemble mechanics |
| 5 | kNN: distance table; NB: posterior calculation | lazy learning + independence assumption |
Build rule: implement each algorithm from scratch in a fresh cell; compare test accuracies in a single summary table at the end.
Counterexample
Discussion prompt
Implement and evaluate all 7 algorithms on the same synthetic dataset. Then, for each, answer from memory: loss function, gradient, two key hyperparameters, and bias-variance profile.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Build rule: implement each algorithm from scratch in a fresh cell; compare test accuracies in a single summary table at the end.
Worked example
Your turn: load make_classification(n=500, random_state=42), fit LogisticRegression, then manually compute one gradient step on the first mini-batch (n=5) and verify the gradient formula.
Hint: p = sigmoid(X @ w), grad = X.T @ (p - y) / n, w_new = w - lr * grad. Compare your grad to lr.coef_.
import numpy as np
from sklearn.datasets import make_classification
from sklearn.preprocessing import StandardScaler
def sigmoid(z): return 1/(1+np.exp(-z))
X, y = make_classification(n_samples=500, n_features=10, random_state=42)
X = StandardScaler().fit_transform(X)
Xm, ym = X[:5,:2], y[:5].astype(float)
w = np.zeros(2)
p = sigmoid(Xm @ w)
grad = Xm.T @ (p - ym) / 5
print('grad:', grad.round(6))
print('w after step (lr=0.5):', (w - 0.5*grad).round(6))| quantity | value |
|---|---|
| p at w=0 (all 0.5) | 0.5, 0.5, 0.5, 0.5, 0.5 |
| grad[0] | -0.125 (approx, depends on mini-batch) |
| w[0] after 1 step (lr=0.5) | ≈ 0.0625 |
Worked example
Your turn: fit SVC(kernel='linear') for C ∈ {0.1, 1.0, 10.0} and print margin = 2/‖w‖ and n_support_vectors.
Hint: svm.coef_[0] gives w; len(svm.support_vectors_) gives the count. Expect margin to shrink and SV count to drop as C grows.
from sklearn.svm import SVC
from sklearn.datasets import load_iris
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split
import numpy as np
X, y = load_iris(return_X_y=True)
X, y = X[y<2,:2], y[y<2]
Xtr,Xte,ytr,yte = train_test_split(X, y, test_size=0.2, random_state=0)
sc = StandardScaler(); Xtr = sc.fit_transform(Xtr); Xte = sc.transform(Xte)
for C in [0.1, 1.0, 10.0]:
m = SVC(kernel='linear', C=C).fit(Xtr, ytr)
print(f'C={C}: n_sv={len(m.support_vectors_)}, margin={2/np.linalg.norm(m.coef_[0]):.4f}')| C | n_sv | margin |
|---|---|---|
| 0.1 | 28 | 1.4447 |
| 1.0 | 10 | 0.7800 |
| 10.0 | 4 | 0.4303 |
Trade off
Comparison matrix
From Milestone 2 — SVM margin vs C: every row here is a choice with a cost. Fill the n_sv column, then say which row you would actually pick and what you give up for it.
| C | n_sv | margin |
|---|---|---|
| 0.1 | 28 | 1.4447 |
| 1.0 | 10 | 0.7800 |
| 10.0 | 4 | 0.4303 |
Worked example
Your turn: fit all remaining algorithms on make_classification(n=500, random_state=42) and record test accuracy for each.
Hint: scale features for LR, SVM, kNN. Don't scale for DT, RF, GBM, NB. Use accuracy_score(y_test, model.predict(X_test)).
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.tree import DecisionTreeClassifier
from sklearn.ensemble import RandomForestClassifier, GradientBoostingClassifier
from sklearn.neighbors import KNeighborsClassifier
from sklearn.naive_bayes import GaussianNB
from sklearn.metrics import accuracy_score
X, y = make_classification(n_samples=500, n_features=10, random_state=42)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.2, random_state=42)
sc = StandardScaler(); Xsc_tr = sc.fit_transform(Xtr); Xsc_te = sc.transform(Xte)
models = {'DT':DecisionTreeClassifier(max_depth=3,random_state=42),
'RF':RandomForestClassifier(n_estimators=100,random_state=42),
'GBM':GradientBoostingClassifier(n_estimators=100,random_state=42),
'kNN':KNeighborsClassifier(n_neighbors=5),
'NB':GaussianNB()}
for name, m in models.items():
Xf = Xsc_tr if name in ('kNN',) else Xtr
Xt = Xsc_te if name in ('kNN',) else Xte
m.fit(Xf, ytr)
print(f'{name}: {accuracy_score(yte, m.predict(Xt)):.4f}')| algorithm | test accuracy |
|---|---|
| Decision Tree (max_depth=3) | 0.8400 |
| Random Forest (100 trees) | 0.9300 |
| GBM (100 trees, lr=0.1) | 0.9200 |
| kNN (k=5) | 0.8600 |
| Naive Bayes (Gaussian) | 0.8500 |
Comparison
Comparison matrix
From Milestones 3-5 — DT / RF / GBM / kNN / NB: refill the test accuracy column from what you know. The rest of the table is as it appeared.
| algorithm | test accuracy |
|---|---|
| Decision Tree (max_depth=3) | 0.8400 |
| Random Forest (100 trees) | 0.9300 |
| GBM (100 trees, lr=0.1) | 0.9200 |
| kNN (k=5) | 0.8600 |
| Naive Bayes (Gaussian) | 0.8500 |
Concept
Close these slides. For each of the 7 algorithms, state out loud: (1) the loss function, (2) the gradient or update rule, (3) two key hyperparameters, (4) whether it scales with feature space, and (5) when you would NOT choose it.
Stretch (homework): implement the USAAIO patterns cold — (a) derive logistic gradient from scratch on paper; (b) derive Gini gain for a 3-class node; (c) write a kNN classifier in <20 lines; (d) choose the best algorithm for: 1M rows tabular, high dimensionality text, 50-sample medical dataset, real-time streaming.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — The 7 algorithms at a glance · Derivations — logistic & SVM · Trees, forests, and boosting · kNN and Naive Bayes · Your turn: implement all 7. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
| algorithm | one-line memory hook |
|---|---|
| Logistic Regression | sigmoid(Xw); gradient = Xᵀ(σ−y)/n; regularize with C |
| SVM | max margin = 2/‖w‖; large C → narrow margin, fewer SVs |
| Decision Tree | minimize Gini; gain = Gini_parent − weighted Gini_children |
| Random Forest | bootstrap + random features; reduces variance over one tree |
| GBM | fit each tree on pseudo-residuals; sequential, not parallel |
| kNN | lazy; O(n·d) predict; small k = high variance |
| Naive Bayes | ∝ prior × product of likelihoods; assumes feature independence |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.