USAAIO Lesson 72, from Week 25, consolidating Phase 2. It is a timed review of every classification and regression algorithm, the ensemble methods of bagging, boosting, and stacking, and the clustering methods k-Means, GMM, DBSCAN, and hierarchical clustering, ending with a 30-minute PyTorch mastery check. Everything was verified on sklearn datasets with real execution numbers. The lesson runs to 30 slides.
Subject: Machine Learning · 59 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 72 · Week 25 (Phase 2 end)
Every classification and regression algorithm. Every ensemble. Every clustering method. And a timed PyTorch mastery check — implement any architecture from scratch in 30 minutes.
Objectives
Warm-up
Discussion prompt
Before we open Lesson 72: Phase 2 Algorithm Review: without looking back, what was the main idea of Kaggle Competition Pipeline, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
end-to-end competition workflow — EDA (distributions, correlations, missing values, outliers), interaction and polynomial feature engineering, model selection across linear/tree/neural families with 5-fold CV, stacking ensembles, and using CV scores to predict leaderboard position.
Section
Part 1 of 4
Concept
Both criteria measure impurity at a node — Gini is a bit faster; entropy gives more balanced splits. In practice they produce nearly identical trees.
| criterion | formula | range | min at pure node |
|---|---|---|---|
| Gini | 1 − Σ pᵢ² | 0 to 0.5 | 0 |
| Entropy | −Σ pᵢ log₂ pᵢ | 0 to log₂k | 0 |
| Information gain | H(parent) − weighted H(children) | ≥ 0 | 0 (no split needed) |
The split that maximizes information gain (or minimizes weighted Gini) is chosen at each node — greedy and locally optimal, not globally.
Comparison
Comparison matrix
From Splitting criteria: Gini vs entropy: refill the formula column from what you know. The rest of the table is as it appeared.
| criterion | formula | range | min at pure node |
|---|---|---|---|
| Gini | 1 − Σ pᵢ² | 0 to 0.5 | 0 |
| Entropy | −Σ pᵢ log₂ pᵢ | 0 to log₂k | 0 |
| Information gain | H(parent) − weighted H(children) | ≥ 0 | 0 (no split needed) |
Estimation
Predict first
Fit an unpruned tree and a max_depth=3 tree on iris. Predict: does pruning hurt accuracy on this 3-class, 4-feature dataset?
Commit before you compute: what does Decision tree on iris — depth vs accuracy come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: depth=4, acc=0.9778 vs max_depth=3, acc=0.9778
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The tree reaches its best split at depth 4 anyway; imposing max_depth=3 is equivalent here.
Worked example
Fit an unpruned tree and a max_depth=3 tree on iris. Predict: does pruning hurt accuracy on this 3-class, 4-feature dataset?
from sklearn.datasets import load_iris
from sklearn.tree import DecisionTreeClassifier
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score
X, y = load_iris(return_X_y=True)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.3, random_state=0)
dt_full = DecisionTreeClassifier(random_state=0).fit(X_tr, y_tr)
dt_pruned = DecisionTreeClassifier(max_depth=3, random_state=0).fit(X_tr, y_tr)
print(dt_full.get_depth(), accuracy_score(y_te, dt_full.predict(X_te)))
print(3, accuracy_score(y_te, dt_pruned.predict(X_te)))Iris is linearly separable enough that depth-4 is already optimal — pruning to depth 3 costs nothing.
depth=4, acc=0.9778 vs max_depth=3, acc=0.9778
Why: The tree reaches its best split at depth 4 anyway; imposing max_depth=3 is equivalent here. On noisier data, shallow trees generalize better — this is the bias-variance tradeoff from Lesson 32.
| model | depth | test acc |
|---|---|---|
| unpruned (gini) | 4 | 0.9778 |
| max_depth=3 (gini) | 3 | 0.9778 |
| unpruned (entropy) | 4 | 0.9778 |
Trade off
Comparison matrix
From Decision tree on iris — depth vs accuracy: every row here is a choice with a cost. Fill the depth column, then say which row you would actually pick and what you give up for it.
| model | depth | test acc |
|---|---|---|
| unpruned (gini) | 4 | 0.9778 |
| max_depth=3 (gini) | 3 | 0.9778 |
| unpruned (entropy) | 4 | 0.9778 |
Concept
| advantage | disadvantage |
|---|---|
| interpretable (feature importances, rules) | high variance — tiny data changes → different tree |
| no feature scaling required | greedy splits → not globally optimal |
| handles mixed data types | tends to overfit without pruning / min_samples |
| fast inference: O(depth) | struggles with linear or smooth decision boundaries |
The high variance is exactly why ensembles exist: average many trees to cancel out their individual noise.
Counterexample
Discussion prompt
The high variance is exactly why ensembles exist: average many trees to cancel out their individual noise.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Section
Part 2 of 4
Concept
| strategy | how it trains | error it fixes | canonical algorithm |
|---|---|---|---|
| Bagging | bootstrap resample + parallel trees + vote/avg | variance | Random Forest |
| Boosting | sequential; each learner corrects last's residuals | bias | GBM / AdaBoost / XGBoost |
| Stacking | train diverse base models, meta-learner on their outputs | both | StackingClassifier |
Bagging when your single model overfits (high variance). Boosting when it underfits (high bias). Stacking to squeeze out the last few percent on Kaggle-style benchmarks.
Analogy
Discussion prompt
Explain Three ensemble strategies by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Bagging when your single model overfits (high variance). Boosting when it underfits (high bias). Stacking to squeeze out the last few percent on Kaggle-style benchmarks.
Missing information
Discussion prompt
Fit a single DT, Random Forest, GBM, and a stacking meta-classifier on the binary digits task. Predict the ranking.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
100 bootstrapped trees vote out the single-tree variance entirely. GBM matches DT here because the base error is already low; boosting gains more on harder (high-bias) tasks.
Worked example
Fit a single DT, Random Forest, GBM, and a stacking meta-classifier on the binary digits task. Predict the ranking.
from sklearn.datasets import load_digits
from sklearn.tree import DecisionTreeClassifier
from sklearn.ensemble import RandomForestClassifier, GradientBoostingClassifier, StackingClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score
digits = load_digits(); mask = digits.target < 2
Xd, yd = digits.data[mask], digits.target[mask]
Xd_tr, Xd_te, yd_tr, yd_te = train_test_split(Xd, yd, test_size=0.3, random_state=42)
models = [
('DT', DecisionTreeClassifier(random_state=0)),
('RF', RandomForestClassifier(n_estimators=100, random_state=0)),
('GBM', GradientBoostingClassifier(n_estimators=100, random_state=0)),
]
for name, m in models:
m.fit(Xd_tr, yd_tr)
print(name, accuracy_score(yd_te, m.predict(Xd_te)))DT=0.9907, GBM=0.9907, RF=1.0000 — RF perfect on this near-separable binary task
Why: 100 bootstrapped trees vote out the single-tree variance entirely. GBM matches DT here because the base error is already low; boosting gains more on harder (high-bias) tasks.
| algorithm | test acc | note |
|---|---|---|
| Decision Tree | 0.9907 | single tree |
| GBM (boosting) | 0.9907 | sequential correction |
| Random Forest (bagging) | 1.0000 | variance cancelled |
| Stacking (DT+RF+LR) | 0.9907 | meta overhead on easy task |
Discrimination
Sort into buckets
Sort these by test acc, from memory, without looking back at Ensemble comparison on digits (0 vs 1). Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Validation loss is much higher than training loss. Add a GBM with 500 estimators to push accuracy higher.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Boosting reduces bias by adding complexity at each stage — it makes overfitting worse, not better.
Diagnose: if train acc >> test acc, the error is variance — use bagging or regularization first.
Why: Boosting reduces bias by adding complexity at each stage — it makes overfitting worse, not better. If train << test, you have a variance problem; boosting amplifies it.
Trap
Validation loss is much higher than training loss. Add a GBM with 500 estimators to push accuracy higher.
Apply boosting to an already high-variance model
Why: Boosting reduces bias by adding complexity at each stage — it makes overfitting worse, not better. If train << test, you have a variance problem; boosting amplifies it.
Diagnose: if train acc >> test acc, the error is variance — use bagging or regularization first.
High variance → Random Forest / reduce depth / dropout; high bias → GBM / deeper model
Why: Boosting corrects the residuals of previous weak learners; it is fundamentally a bias reducer. Bagging averages independent noisy trees to reduce variance (Lesson 32's bias-variance decomposition).
Break the constraint
Discussion prompt
The rule this trap just fixed:
Boosting corrects the residuals of previous weak learners; it is fundamentally a bias reducer. Bagging averages independent noisy trees to reduce variance (Lesson 32's bias-variance decomposition).
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
Boosting reduces bias by adding complexity at each stage — it makes overfitting worse, not better. If train << test, you have a variance problem; boosting amplifies it.
Section
Part 3 of 4
Concept
| algorithm | cluster shape | needs k? | handles noise? | key hyperparams |
|---|---|---|---|---|
| k-Means | spherical / convex | yes | no | k, n_init |
| GMM | elliptical / soft | yes | weakly | n_components, covariance_type |
| DBSCAN | arbitrary | no | yes (noise label) | eps, min_samples |
| Hierarchical (Ward) | various | post-hoc | no | linkage, n_clusters |
k-Means minimizes inertia (sum of squared distances to centroids). GMM maximizes likelihood — gives soft assignments and handles elongated clusters. DBSCAN groups density-connected points; outliers become noise (label = −1).
Estimation
Predict first
Four clusters, cluster_std=0.8, 300 points. All four algorithms see the same data — predict which ones find k=4 correctly.
Commit before you compute: what does Clustering comparison on make_blobs come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: k-Means=1.0, GMM=1.0, DBSCAN=1.0, Ward=1.0 — all perfect on separable blobs
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. With well-separated, roughly equal-sized, spherical clusters all four methods agree.
Worked example
Four clusters, cluster_std=0.8, 300 points. All four algorithms see the same data — predict which ones find k=4 correctly.
import numpy as np
from sklearn.datasets import make_blobs
from sklearn.cluster import KMeans, DBSCAN, AgglomerativeClustering
from sklearn.mixture import GaussianMixture
from sklearn.preprocessing import StandardScaler
from sklearn.metrics import adjusted_rand_score, silhouette_score
Xb, yb = make_blobs(300, centers=4, cluster_std=0.8, random_state=7)
Xbs = StandardScaler().fit_transform(Xb)
km = KMeans(4, random_state=0, n_init='auto').fit(Xb)
gmm = GaussianMixture(4, random_state=0).fit(Xb)
db = DBSCAN(eps=0.5, min_samples=5).fit(Xbs)
agg = AgglomerativeClustering(4).fit(Xb)
for name, labels in [('k-Means',km.labels_),('GMM',gmm.predict(Xb)),
('DBSCAN',db.labels_),('Ward',agg.labels_)]:
print(name, adjusted_rand_score(yb, labels).round(4))k-Means=1.0, GMM=1.0, DBSCAN=1.0, Ward=1.0 — all perfect on separable blobs
Why: With well-separated, roughly equal-sized, spherical clusters all four methods agree. Differences emerge on non-convex shapes (DBSCAN shines), elliptical clusters (GMM shines), or when k is unknown (DBSCAN infers it from density).
| algorithm | ARI | silhouette | note |
|---|---|---|---|
| k-Means | 1.0000 | 0.846 | needs k=4 given |
| GMM | 1.0000 | 0.846 | soft; EM converged |
| DBSCAN (eps=0.5) | 1.0000 | 0.839 | inferred k from density |
| Hierarchical (Ward) | 1.0000 | 0.846 | dendrogram cut at k=4 |
Discrimination
Sort into buckets
Sort these by silhouette, from memory, without looking back at Clustering comparison on make_blobs. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Concept
Run k-Means for k=2…6 and plot inertia. The elbow — where adding k stops dropping inertia sharply — is the right k. Complement with silhouette score (higher is better).
| k | inertia | sharp drop? |
|---|---|---|
| 2 | 12942.7 | yes |
| 3 | 2313.8 | yes |
| 4 | 365.6 | yes — elbow here |
| 5 | 331.5 | small drop |
| 6 | 293.0 | flat |
The true k=4 is unmistakable from the elbow. On messier data, GMM's BIC/AIC provides a formal model-selection criterion instead.
Sorting
Sort into buckets
These are the pieces of Lesson 72: Phase 2 Algorithm Review, out of order. Put each one back under the part of the lesson it belongs to.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Set eps=0.5 directly on raw feature data — 0.5 sounds small and reasonable.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: eps is in the feature-space distance.
Scale first, then tune eps on the normalized data.
Why: eps is in the feature-space distance. If features have different scales (e.g. age in 0-100, income in 0-100000), eps=0.5 is either too tiny (all noise) or too huge (one cluster) depending on the dominant feature. Always StandardScaler before DBSCAN.
Trap
Set eps=0.5 directly on raw feature data — 0.5 sounds small and reasonable.
DBSCAN(eps=2.0) on un-scaled blobs → 1 cluster (everything merged)
Why: eps is in the feature-space distance. If features have different scales (e.g. age in 0-100, income in 0-100000), eps=0.5 is either too tiny (all noise) or too huge (one cluster) depending on the dominant feature. Always StandardScaler before DBSCAN.
Scale first, then tune eps on the normalized data.
Xbs = StandardScaler().fit_transform(Xb); DBSCAN(eps=0.5).fit(Xbs) → 4 clusters, ARI=1.0
Why: After scaling, eps=0.5 is 0.5 standard deviations in every dimension — a meaningful radius. Use a k-distance plot (sort the k-th nearest neighbor distances) to pick eps systematically.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
eps=0.5 directly on raw feature data — 0.5 sounds small and reasonable.Section
Part 4 of 4
Concept
Lesson 72's benchmark: build a 3-layer MLP for any classification task from scratch in 30 minutes. This means the loop, the architecture, and the evaluation are second nature.
Linear → ReLU → Linear → ReLU → Linearzero_grad → forward → CrossEntropyLoss → backward → stepAdam(lr=1e-3) — the default for classificationExplain it
Discussion prompt
Explain The mastery standard to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Lesson 72's benchmark: build a 3-layer MLP for any classification task from scratch in 30 minutes. This means the loop, the architecture, and the evaluation are second nature.
Estimation
Predict first
Predict: how many epochs to reach < 0.1 loss? What test accuracy does a 64→128→64→10 MLP achieve?
Commit before you compute: what does 3-layer MLP on load_digits — full trace come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Epoch 1: loss=2.3134 → epoch 100: 0.0752 → epoch 300: 0.0045; test acc=0.9750
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Adam on a normalized input converges fast.
Worked example
Predict: how many epochs to reach < 0.1 loss? What test accuracy does a 64→128→64→10 MLP achieve?
import numpy as np, torch
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
X, y = load_digits(return_X_y=True)
X = X.astype(np.float32)
X = (X - X.mean(0)) / (X.std(0) + 1e-8)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, random_state=0)
torch.manual_seed(0)
Xt, yt = torch.tensor(X_tr), torch.tensor(y_tr, dtype=torch.long)
Xte, yte = torch.tensor(X_te), torch.tensor(y_te, dtype=torch.long)
model = torch.nn.Sequential(
torch.nn.Linear(64,128), torch.nn.ReLU(),
torch.nn.Linear(128,64), torch.nn.ReLU(),
torch.nn.Linear(64,10))
opt = torch.optim.Adam(model.parameters(), lr=1e-3)
loss_fn = torch.nn.CrossEntropyLoss()
for epoch in range(300):
opt.zero_grad()
loss = loss_fn(model(Xt), yt)
loss.backward(); opt.step()
with torch.no_grad():
acc = (model(Xte).argmax(1)==yte).float().mean().item()
print(f'test acc {acc:.4f}')Epoch 1: loss=2.3134 → epoch 100: 0.0752 → epoch 300: 0.0045; test acc=0.9750
Why: Adam on a normalized input converges fast. Cross-entropy loss below 0.1 by epoch 100; near-zero by 300. 97.5% accuracy on 10-class digits from scratch — this is the 30-minute mastery benchmark.
| epoch | CE loss | phase |
|---|---|---|
| 1 | 2.3134 | random init |
| 10 | 2.0562 | early learning |
| 50 | 0.4056 | rapid descent |
| 100 | 0.0752 | < 0.1 threshold |
| 200 | 0.0122 | converging |
| 300 | 0.0045 | converged; test acc 0.9750 |
Comparison
Comparison matrix
From 3-layer MLP on load_digits — full trace: refill the CE loss column from what you know. The rest of the table is as it appeared.
| epoch | CE loss | phase |
|---|---|---|
| 1 | 2.3134 | random init |
| 10 | 2.0562 | early learning |
| 50 | 0.4056 | rapid descent |
| 100 | 0.0752 | < 0.1 threshold |
| 200 | 0.0122 | converging |
| 300 | 0.0045 | converged; test acc 0.9750 |
Constraint
Discussion prompt
Run Phase 2 algorithm-selection recipe with this step confiscated:
Pick a clusterer: k-Means / GMM when k is known; DBSCAN when shape is arbitrary / noise exists — always scale first
Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.
Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.
Answer:
Pattern
Edge cases
Discussion prompt
Phase 2 algorithm-selection recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
Elimination
Eliminate the wrong options
A decision tree node contains 40 class-A and 60 class-B samples. Which value is closest to its Gini impurity?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Gini = 1 − (0.4² + 0.6²) = 1 − (0.16 + 0.36) = 1 − 0.52 = 0.48. Maximum Gini for two classes is 0.5 (at 50/50), so 0.48 is near-maximum — this is an impure node worth splitting.
Check
Theory question — no code needed.
Check your understanding
A decision tree node contains 40 class-A and 60 class-B samples. Which value is closest to its Gini impurity?
Answer: A
Why: Gini = 1 − (0.4² + 0.6²) = 1 − (0.16 + 0.36) = 1 − 0.52 = 0.48. Maximum Gini for two classes is 0.5 (at 50/50), so 0.48 is near-maximum — this is an impure node worth splitting.
Prediction
Predict first
Your gradient-boosted model has train acc = 0.99 and val acc = 0.72. The FIRST intervention should be:
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: reduce n_estimators or add shrinkage (learning_rate < 1)
Why: Train 0.99 vs val 0.72 is a variance (overfitting) problem. Boosting with many rounds memorizes training data — reduce rounds or add learning_rate shrinkage (= L2 toward weaker ensemble) to regularize.
Check
Scenario: the model generalizes poorly.
Check your understanding
Your gradient-boosted model has train acc = 0.99 and val acc = 0.72. The FIRST intervention should be:
Answer: A
Why: Train 0.99 vs val 0.72 is a variance (overfitting) problem. Boosting with many rounds memorizes training data — reduce rounds or add learning_rate shrinkage (= L2 toward weaker ensemble) to regularize.
Elimination
Eliminate the wrong options
After running DBSCAN on a dataset you get label −1 for 30% of points. This most likely means:
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Label −1 in DBSCAN means 'noise' — points with fewer than min_samples neighbors within eps. 30% noise typically means eps is too small (too tight a radius), or the data truly has many outliers. Increase eps or decrease min_samples and recheck.
Check
An exam-favorite edge case.
Check your understanding
After running DBSCAN on a dataset you get label −1 for 30% of points. This most likely means:
Answer: A
Why: Label −1 in DBSCAN means 'noise' — points with fewer than min_samples neighbors within eps. 30% noise typically means eps is too small (too tight a radius), or the data truly has many outliers. Increase eps or decrease min_samples and recheck.
Section
Project
Concept
The objective is speed and accuracy together: implement a working, generalizing MLP from a blank file in 30 minutes.
| # | milestone | tool |
|---|---|---|
| 1 | Load + normalize load_digits | np, train_test_split |
| 2 | Define 3-layer Sequential MLP | nn.Sequential, Linear, ReLU |
| 3 | Training loop: 300 epochs, Adam | zero_grad → loss → backward → step |
| 4 | Evaluate — target > 97% test acc | argmax, float().mean() |
Build rules: no torch.nn.utils shortcuts, no sklearn wrappers — type every line. Set torch.manual_seed(0) for reproducibility.
Analogy
Discussion prompt
Explain Project: MLP from scratch under the clock by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
The objective is speed and accuracy together: implement a working, generalizing MLP from a blank file in 30 minutes.
Worked example
Your turn: load digits, cast to float32, normalize per-feature, split 80/20. What shape should X_tr be?
Hint: load_digits() returns 1797×64. After split (test_size=0.2, random_state=0): X_tr is (1437, 64).
import numpy as np
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
X, y = load_digits(return_X_y=True)
X = X.astype(np.float32)
X = (X - X.mean(0)) / (X.std(0) + 1e-8)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, random_state=0)
print(X_tr.shape, X_te.shape)| split | shape | classes |
|---|---|---|
| X_tr | (1437, 64) | 10 (0–9) |
| X_te | (360, 64) | 10 (0–9) |
Worked example
Your turn: build the Sequential MLP and run 300 epochs with Adam. Predict the loss at epoch 100.
Hint: torch.manual_seed(0) first; .unsqueeze not needed — inputs are already 2D; loop = zero_grad → forward → CELoss → backward → step.
import torch
torch.manual_seed(0)
Xt, yt = torch.tensor(X_tr), torch.tensor(y_tr, dtype=torch.long)
Xte, yte = torch.tensor(X_te), torch.tensor(y_te, dtype=torch.long)
model = torch.nn.Sequential(
torch.nn.Linear(64,128), torch.nn.ReLU(),
torch.nn.Linear(128,64), torch.nn.ReLU(),
torch.nn.Linear(64,10))
opt = torch.optim.Adam(model.parameters(), lr=1e-3)
loss_fn = torch.nn.CrossEntropyLoss()
for ep in range(300):
opt.zero_grad()
loss = loss_fn(model(Xt), yt)
loss.backward(); opt.step()
if ep+1 in (1,100,300): print(ep+1, round(loss.item(),4))| epoch | CE loss |
|---|---|
| 1 | 2.3134 |
| 100 | 0.0752 |
| 300 | 0.0045 |
Pattern
Step through it
Step through Milestone 2 — define architecture and run loop one row at a time. What is driving the change, and what would the row after the last one be?
Worked example
Your turn: evaluate test accuracy with argmax. Predict whether you beat 97%.
Hint: model(Xte).argmax(1) gives predicted class indices; compare element-wise to yte, cast to float, take mean.
with torch.no_grad():
preds = model(Xte).argmax(1)
acc = (preds == yte).float().mean().item()
print(f'test acc: {acc:.4f}')| metric | value | target |
|---|---|---|
| test accuracy | 0.9750 | > 0.97 ✓ |
| CE loss (epoch 300) | 0.0045 | < 0.01 ✓ |
Trade off
Comparison matrix
From Milestone 3 — evaluate accuracy: every row here is a choice with a cost. Fill the target column, then say which row you would actually pick and what you give up for it.
| metric | value | target |
|---|---|---|
| test accuracy | 0.9750 | > 0.97 ✓ |
| CE loss (epoch 300) | 0.0045 | < 0.01 ✓ |
Concept
import numpy as np, torch
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
X, y = load_digits(return_X_y=True)
X = X.astype(np.float32)
X = (X - X.mean(0)) / (X.std(0) + 1e-8)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, random_state=0)
torch.manual_seed(0)
Xt, yt = torch.tensor(X_tr), torch.tensor(y_tr, dtype=torch.long)
Xte, yte = torch.tensor(X_te), torch.tensor(y_te, dtype=torch.long)
model = torch.nn.Sequential(
torch.nn.Linear(64,128), torch.nn.ReLU(),
torch.nn.Linear(128,64), torch.nn.ReLU(),
torch.nn.Linear(64,10))
opt = torch.optim.Adam(model.parameters(), lr=1e-3)
for _ in range(300):
opt.zero_grad()
torch.nn.CrossEntropyLoss()(model(Xt),yt).backward(); opt.step()
with torch.no_grad():
print((model(Xte).argmax(1)==yte).float().mean().item())| output | value |
|---|---|
| test accuracy | 0.9750 |
| 30-min target | passed ✓ |
If you can reproduce this from memory in under 30 minutes, you've cleared the Phase 2 PyTorch mastery benchmark.
Comparison
Comparison matrix
From The full program: refill the value column from what you know. The rest of the table is as it appeared.
| output | value |
|---|---|
| test accuracy | 0.9750 |
| 30-min target | passed ✓ |
Concept
Out loud, slides closed: (1) explain when you'd pick GBM over Random Forest and why; (2) state what DBSCAN label −1 means and how to fix it; (3) recite the 5-step PyTorch training loop.
Homework drill: create a one-page summary sheet listing every Phase 2 algorithm, its splitting/training criterion, its top hyperparameter, and its canonical failure mode. Then run the speed drill — k-Means, logistic regression, 3-layer CNN — in 30-minute timed blocks.
Counterexample
Discussion prompt
Out loud, slides closed: (1) explain when you'd pick GBM over Random Forest and why; (2) state what DBSCAN label −1 means and how to fix it; (3) recite the 5-step PyTorch training loop.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Decision trees — criteria, pruning, tradeoffs · Ensembles — bagging, boosting, stacking · Clustering — k-Means, GMM, DBSCAN, hierarchical · PyTorch mastery check — 30-minute MLP · Your turn: 30-min MLP mastery drill. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
| algorithm | the one thing to remember |
|---|---|
| Decision Tree | greedy splits; high variance without pruning |
| Random Forest | bagging ↓ variance; use when single DT overfits |
| GBM | boosting ↓ bias; use when model underfits |
| k-Means / GMM | needs k; GMM is soft + elliptical |
| DBSCAN | scale first; eps too large → all one cluster |
| PyTorch MLP | zero_grad every step; Adam + CE for classification |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.