Lesson 72: Phase 2 Algorithm Review

USAAIO Lesson 72, from Week 25, consolidating Phase 2. It is a timed review of every classification and regression algorithm, the ensemble methods of bagging, boosting, and stacking, and the clustering methods k-Means, GMM, DBSCAN, and hierarchical clustering, ending with a 30-minute PyTorch mastery check. Everything was verified on sklearn datasets with real execution numbers. The lesson runs to 30 slides.

Subject: Machine Learning · 59 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Phase 2 Algorithm Review

Title

USAAIO · Lesson 72 · Week 25 (Phase 2 end)

Every classification and regression algorithm. Every ensemble. Every clustering method. And a timed PyTorch mastery check — implement any architecture from scratch in 30 minutes.

2. By the end of this lesson you can

Objectives

  1. State splitting criteria, pruning tradeoffs, and exam-level gotchas for decision trees
  2. Choose between bagging, boosting, and stacking given a bias-variance diagnosis
  3. Compare k-Means, GMM, DBSCAN, and hierarchical clustering and their evaluation metrics
  4. Implement a PyTorch MLP for any architecture from scratch within 30 minutes
  5. Produce a one-page algorithm summary sheet anchoring every Phase 2 model to a use case

3. What survived from Kaggle Competition Pipeline?

Warm-up

Discussion prompt

Before we open Lesson 72: Phase 2 Algorithm Review: without looking back, what was the main idea of Kaggle Competition Pipeline, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

end-to-end competition workflow — EDA (distributions, correlations, missing values, outliers), interaction and polynomial feature engineering, model selection across linear/tree/neural families with 5-fold CV, stacking ensembles, and using CV scores to predict leaderboard position.

4. Decision trees — criteria, pruning, tradeoffs

Section

Part 1 of 4

5. Splitting criteria: Gini vs entropy

Concept

Both criteria measure impurity at a node — Gini is a bit faster; entropy gives more balanced splits. In practice they produce nearly identical trees.

criterionformularangemin at pure node
Gini1 − Σ pᵢ²0 to 0.50
Entropy−Σ pᵢ log₂ pᵢ0 to log₂k0
Information gainH(parent) − weighted H(children)≥ 00 (no split needed)

The split that maximizes information gain (or minimizes weighted Gini) is chosen at each node — greedy and locally optimal, not globally.

6. Fill in: formula for Splitting criteria: Gini vs entropy

Comparison

Comparison matrix

From Splitting criteria: Gini vs entropy: refill the formula column from what you know. The rest of the table is as it appeared.

criterionformularangemin at pure node
Gini1 − Σ pᵢ²0 to 0.50
Entropy−Σ pᵢ log₂ pᵢ0 to log₂k0
Information gainH(parent) − weighted H(children)≥ 00 (no split needed)

7. Guess the shape of the answer: Decision tree on iris — depth vs accuracy

Estimation

Predict first

Fit an unpruned tree and a max_depth=3 tree on iris. Predict: does pruning hurt accuracy on this 3-class, 4-feature dataset?

Commit before you compute: what does Decision tree on iris — depth vs accuracy come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: depth=4, acc=0.9778 vs max_depth=3, acc=0.9778

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The tree reaches its best split at depth 4 anyway; imposing max_depth=3 is equivalent here.

8. Decision tree on iris — depth vs accuracy

Worked example

Fit an unpruned tree and a max_depth=3 tree on iris. Predict: does pruning hurt accuracy on this 3-class, 4-feature dataset?

from sklearn.datasets import load_iris
from sklearn.tree import DecisionTreeClassifier
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score

X, y = load_iris(return_X_y=True)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.3, random_state=0)
dt_full   = DecisionTreeClassifier(random_state=0).fit(X_tr, y_tr)
dt_pruned = DecisionTreeClassifier(max_depth=3, random_state=0).fit(X_tr, y_tr)
print(dt_full.get_depth(), accuracy_score(y_te, dt_full.predict(X_te)))
print(3,                   accuracy_score(y_te, dt_pruned.predict(X_te)))

Iris is linearly separable enough that depth-4 is already optimal — pruning to depth 3 costs nothing.

depth=4, acc=0.9778 vs max_depth=3, acc=0.9778

Why: The tree reaches its best split at depth 4 anyway; imposing max_depth=3 is equivalent here. On noisier data, shallow trees generalize better — this is the bias-variance tradeoff from Lesson 32.

modeldepthtest acc
unpruned (gini)40.9778
max_depth=3 (gini)30.9778
unpruned (entropy)40.9778

9. What each one costs: Decision tree on iris — depth vs accuracy

Trade off

Comparison matrix

From Decision tree on iris — depth vs accuracy: every row here is a choice with a cost. Fill the depth column, then say which row you would actually pick and what you give up for it.

modeldepthtest acc
unpruned (gini)40.9778
max_depth=3 (gini)30.9778
unpruned (entropy)40.9778

10. Decision tree advantages and disadvantages

Concept

advantagedisadvantage
interpretable (feature importances, rules)high variance — tiny data changes → different tree
no feature scaling requiredgreedy splits → not globally optimal
handles mixed data typestends to overfit without pruning / min_samples
fast inference: O(depth)struggles with linear or smooth decision boundaries

The high variance is exactly why ensembles exist: average many trees to cancel out their individual noise.

11. Break it if you can: Decision tree advantages and disadvantages

Counterexample

Discussion prompt

The high variance is exactly why ensembles exist: average many trees to cancel out their individual noise.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

12. Ensembles — bagging, boosting, stacking

Section

Part 2 of 4

13. Three ensemble strategies

Concept

strategyhow it trainserror it fixescanonical algorithm
Baggingbootstrap resample + parallel trees + vote/avgvarianceRandom Forest
Boostingsequential; each learner corrects last's residualsbiasGBM / AdaBoost / XGBoost
Stackingtrain diverse base models, meta-learner on their outputsbothStackingClassifier

Bagging when your single model overfits (high variance). Boosting when it underfits (high bias). Stacking to squeeze out the last few percent on Kaggle-style benchmarks.

14. By analogy: Three ensemble strategies

Analogy

Discussion prompt

Explain Three ensemble strategies by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Bagging when your single model overfits (high variance). Boosting when it underfits (high bias). Stacking to squeeze out the last few percent on Kaggle-style benchmarks.

15. What has to be given first: Ensemble comparison on digits (0 vs 1)

Missing information

Discussion prompt

Fit a single DT, Random Forest, GBM, and a stacking meta-classifier on the binary digits task. Predict the ranking.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

100 bootstrapped trees vote out the single-tree variance entirely. GBM matches DT here because the base error is already low; boosting gains more on harder (high-bias) tasks.

16. Ensemble comparison on digits (0 vs 1)

Worked example

Fit a single DT, Random Forest, GBM, and a stacking meta-classifier on the binary digits task. Predict the ranking.

from sklearn.datasets import load_digits
from sklearn.tree import DecisionTreeClassifier
from sklearn.ensemble import RandomForestClassifier, GradientBoostingClassifier, StackingClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score

digits = load_digits(); mask = digits.target < 2
Xd, yd = digits.data[mask], digits.target[mask]
Xd_tr, Xd_te, yd_tr, yd_te = train_test_split(Xd, yd, test_size=0.3, random_state=42)

models = [
    ('DT',  DecisionTreeClassifier(random_state=0)),
    ('RF',  RandomForestClassifier(n_estimators=100, random_state=0)),
    ('GBM', GradientBoostingClassifier(n_estimators=100, random_state=0)),
]
for name, m in models:
    m.fit(Xd_tr, yd_tr)
    print(name, accuracy_score(yd_te, m.predict(Xd_te)))

DT=0.9907, GBM=0.9907, RF=1.0000 — RF perfect on this near-separable binary task

Why: 100 bootstrapped trees vote out the single-tree variance entirely. GBM matches DT here because the base error is already low; boosting gains more on harder (high-bias) tasks.

algorithmtest accnote
Decision Tree0.9907single tree
GBM (boosting)0.9907sequential correction
Random Forest (bagging)1.0000variance cancelled
Stacking (DT+RF+LR)0.9907meta overhead on easy task

17. Which is which, by test acc

Discrimination

Sort into buckets

Sort these by test acc, from memory, without looking back at Ensemble comparison on digits (0 vs 1). Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

0.9907
Decision Tree; GBM (boosting); Stacking (DT+RF+LR)
1.0000
Random Forest (bagging)
g1
test acc is "0.9907" for Decision Tree, GBM (boosting), Stacking (DT+RF+LR) — that is what the table on "Ensemble comparison on digits (0 vs 1)" records, and it is the single property separating this group from the rest.
g2
test acc is "1.0000" for Random Forest (bagging) — that is what the table on "Ensemble comparison on digits (0 vs 1)" records, and it is the single property separating this group from the rest.

18. Something is wrong here: choosing boosting when you're already overfitting

Anomaly

Predict first

A student writes this, and it looks reasonable:

Validation loss is much higher than training loss. Add a GBM with 500 estimators to push accuracy higher.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Boosting reduces bias by adding complexity at each stage — it makes overfitting worse, not better.

Diagnose: if train acc >> test acc, the error is variance — use bagging or regularization first.

Why: Boosting reduces bias by adding complexity at each stage — it makes overfitting worse, not better. If train << test, you have a variance problem; boosting amplifies it.

19. Trap: choosing boosting when you're already overfitting

Trap

The trap

Validation loss is much higher than training loss. Add a GBM with 500 estimators to push accuracy higher.

Apply boosting to an already high-variance model

Why: Boosting reduces bias by adding complexity at each stage — it makes overfitting worse, not better. If train << test, you have a variance problem; boosting amplifies it.

The fix

Diagnose: if train acc >> test acc, the error is variance — use bagging or regularization first.

High variance → Random Forest / reduce depth / dropout; high bias → GBM / deeper model

Why: Boosting corrects the residuals of previous weak learners; it is fundamentally a bias reducer. Bagging averages independent noisy trees to reduce variance (Lesson 32's bias-variance decomposition).

20. Break it on purpose: choosing boosting when you're already…

Break the constraint

Discussion prompt

The rule this trap just fixed:

Boosting corrects the residuals of previous weak learners; it is fundamentally a bias reducer. Bagging averages independent noisy trees to reduce variance (Lesson 32's bias-variance decomposition).

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

Boosting reduces bias by adding complexity at each stage — it makes overfitting worse, not better. If train << test, you have a variance problem; boosting amplifies it.

21. Clustering — k-Means, GMM, DBSCAN, hierarchical

Section

Part 3 of 4

22. Four clustering algorithms side-by-side

Concept

algorithmcluster shapeneeds k?handles noise?key hyperparams
k-Meansspherical / convexyesnok, n_init
GMMelliptical / softyesweaklyn_components, covariance_type
DBSCANarbitrarynoyes (noise label)eps, min_samples
Hierarchical (Ward)variouspost-hocnolinkage, n_clusters

k-Means minimizes inertia (sum of squared distances to centroids). GMM maximizes likelihood — gives soft assignments and handles elongated clusters. DBSCAN groups density-connected points; outliers become noise (label = −1).

23. Guess the shape of the answer: Clustering comparison on make_blobs

Estimation

Predict first

Four clusters, cluster_std=0.8, 300 points. All four algorithms see the same data — predict which ones find k=4 correctly.

Commit before you compute: what does Clustering comparison on make_blobs come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: k-Means=1.0, GMM=1.0, DBSCAN=1.0, Ward=1.0 — all perfect on separable blobs

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. With well-separated, roughly equal-sized, spherical clusters all four methods agree.

24. Clustering comparison on make_blobs

Worked example

Four clusters, cluster_std=0.8, 300 points. All four algorithms see the same data — predict which ones find k=4 correctly.

import numpy as np
from sklearn.datasets import make_blobs
from sklearn.cluster import KMeans, DBSCAN, AgglomerativeClustering
from sklearn.mixture import GaussianMixture
from sklearn.preprocessing import StandardScaler
from sklearn.metrics import adjusted_rand_score, silhouette_score

Xb, yb = make_blobs(300, centers=4, cluster_std=0.8, random_state=7)
Xbs = StandardScaler().fit_transform(Xb)

km  = KMeans(4, random_state=0, n_init='auto').fit(Xb)
gmm = GaussianMixture(4, random_state=0).fit(Xb)
db  = DBSCAN(eps=0.5, min_samples=5).fit(Xbs)
agg = AgglomerativeClustering(4).fit(Xb)

for name, labels in [('k-Means',km.labels_),('GMM',gmm.predict(Xb)),
                     ('DBSCAN',db.labels_),('Ward',agg.labels_)]:
    print(name, adjusted_rand_score(yb, labels).round(4))

k-Means=1.0, GMM=1.0, DBSCAN=1.0, Ward=1.0 — all perfect on separable blobs

Why: With well-separated, roughly equal-sized, spherical clusters all four methods agree. Differences emerge on non-convex shapes (DBSCAN shines), elliptical clusters (GMM shines), or when k is unknown (DBSCAN infers it from density).

algorithmARIsilhouettenote
k-Means1.00000.846needs k=4 given
GMM1.00000.846soft; EM converged
DBSCAN (eps=0.5)1.00000.839inferred k from density
Hierarchical (Ward)1.00000.846dendrogram cut at k=4

25. Which is which, by silhouette

Discrimination

Sort into buckets

Sort these by silhouette, from memory, without looking back at Clustering comparison on make_blobs. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

0.846
k-Means; GMM; Hierarchical (Ward)
0.839
DBSCAN (eps=0.5)
g1
silhouette is "0.846" for k-Means, GMM, Hierarchical (Ward) — that is what the table on "Clustering comparison on make_blobs" records, and it is the single property separating this group from the rest.
g2
silhouette is "0.839" for DBSCAN (eps=0.5) — that is what the table on "Clustering comparison on make_blobs" records, and it is the single property separating this group from the rest.

26. Choosing k — the inertia elbow

Concept

Run k-Means for k=2…6 and plot inertia. The elbow — where adding k stops dropping inertia sharply — is the right k. Complement with silhouette score (higher is better).

kinertiasharp drop?
212942.7yes
32313.8yes
4365.6yes — elbow here
5331.5small drop
6293.0flat

The true k=4 is unmistakable from the elbow. On messier data, GMM's BIC/AIC provides a formal model-selection criterion instead.

27. Where does each piece belong: Lesson 72: Phase 2 Algorithm Review

Sorting

Sort into buckets

These are the pieces of Lesson 72: Phase 2 Algorithm Review, out of order. Put each one back under the part of the lesson it belongs to.

Decision trees — criteria, pruning, tradeoffs
Splitting criteria: Gini vs entropy; Decision tree on iris — depth vs accuracy; Decision tree advantages and disadvantages
Ensembles — bagging, boosting, stacking
Three ensemble strategies; Ensemble comparison on digits (0 vs 1)
Clustering — k-Means, GMM, DBSCAN, hierarchical
Four clustering algorithms side-by-side; Clustering comparison on make_blobs; Choosing k — the inertia elbow
s1
Decision trees — criteria, pruning, tradeoffs is where Lesson 72: Phase 2 Algorithm Review puts Splitting criteria: Gini vs entropy, Decision tree on iris — depth vs accuracy, Decision tree advantages and disadvantages. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
Ensembles — bagging, boosting, stacking is where Lesson 72: Phase 2 Algorithm Review puts Three ensemble strategies, Ensemble comparison on digits (0 vs 1). Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
Clustering — k-Means, GMM, DBSCAN, hierarchical is where Lesson 72: Phase 2 Algorithm Review puts Four clustering algorithms side-by-side, Clustering comparison on make_blobs, Choosing k — the inertia elbow. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

28. Something is wrong here: running DBSCAN without scaling — eps is meaningless in…

Anomaly

Predict first

A student writes this, and it looks reasonable:

Set eps=0.5 directly on raw feature data — 0.5 sounds small and reasonable.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: eps is in the feature-space distance.

Scale first, then tune eps on the normalized data.

Why: eps is in the feature-space distance. If features have different scales (e.g. age in 0-100, income in 0-100000), eps=0.5 is either too tiny (all noise) or too huge (one cluster) depending on the dominant feature. Always StandardScaler before DBSCAN.

29. Trap: running DBSCAN without scaling — eps is meaningless in raw feature space

Trap

The trap

Set eps=0.5 directly on raw feature data — 0.5 sounds small and reasonable.

DBSCAN(eps=2.0) on un-scaled blobs → 1 cluster (everything merged)

Why: eps is in the feature-space distance. If features have different scales (e.g. age in 0-100, income in 0-100000), eps=0.5 is either too tiny (all noise) or too huge (one cluster) depending on the dominant feature. Always StandardScaler before DBSCAN.

The fix

Scale first, then tune eps on the normalized data.

Xbs = StandardScaler().fit_transform(Xb); DBSCAN(eps=0.5).fit(Xbs) → 4 clusters, ARI=1.0

Why: After scaling, eps=0.5 is 0.5 standard deviations in every dimension — a meaningful radius. Use a k-distance plot (sort the k-th nearest neighbor distances) to pick eps systematically.

30. Which of these survive contact with Lesson 72: Phase 2 Algorithm Review?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Both criteria measure impurity at a node — Gini is a bit faster; entropy gives more balanced splits. In practice they produce nearly identical trees.; The high variance is exactly why ensembles exist: average many trees to cancel out their individual noise.; Bagging when your single model overfits (high variance). Boosting when it underfits (high bias). Stacking to squeeze out the last few percent on Kaggle-style benchmarks.
Breaks
Validation loss is much higher than training loss. Add a GBM with 500 estimators to push accuracy higher.; Set eps=0.5 directly on raw feature data — 0.5 sounds small and reasonable.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 72: Phase 2 Algorithm Review puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

31. PyTorch mastery check — 30-minute MLP

Section

Part 4 of 4

32. The mastery standard

Concept

Lesson 72's benchmark: build a 3-layer MLP for any classification task from scratch in 30 minutes. This means the loop, the architecture, and the evaluation are second nature.

33. Teach it back: The mastery standard

Explain it

Discussion prompt

Explain The mastery standard to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Lesson 72's benchmark: build a 3-layer MLP for any classification task from scratch in 30 minutes. This means the loop, the architecture, and the evaluation are second nature.

34. Guess the shape of the answer: 3-layer MLP on load_digits — full trace

Estimation

Predict first

Predict: how many epochs to reach < 0.1 loss? What test accuracy does a 64→128→64→10 MLP achieve?

Commit before you compute: what does 3-layer MLP on load_digits — full trace come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Epoch 1: loss=2.3134 → epoch 100: 0.0752 → epoch 300: 0.0045; test acc=0.9750

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Adam on a normalized input converges fast.

35. 3-layer MLP on load_digits — full trace

Worked example

Predict: how many epochs to reach < 0.1 loss? What test accuracy does a 64→128→64→10 MLP achieve?

import numpy as np, torch
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split

X, y = load_digits(return_X_y=True)
X = X.astype(np.float32)
X = (X - X.mean(0)) / (X.std(0) + 1e-8)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, random_state=0)

torch.manual_seed(0)
Xt, yt   = torch.tensor(X_tr), torch.tensor(y_tr, dtype=torch.long)
Xte, yte = torch.tensor(X_te), torch.tensor(y_te, dtype=torch.long)

model = torch.nn.Sequential(
    torch.nn.Linear(64,128), torch.nn.ReLU(),
    torch.nn.Linear(128,64), torch.nn.ReLU(),
    torch.nn.Linear(64,10))
opt  = torch.optim.Adam(model.parameters(), lr=1e-3)
loss_fn = torch.nn.CrossEntropyLoss()

for epoch in range(300):
    opt.zero_grad()
    loss = loss_fn(model(Xt), yt)
    loss.backward(); opt.step()

with torch.no_grad():
    acc = (model(Xte).argmax(1)==yte).float().mean().item()
    print(f'test acc {acc:.4f}')

Epoch 1: loss=2.3134 → epoch 100: 0.0752 → epoch 300: 0.0045; test acc=0.9750

Why: Adam on a normalized input converges fast. Cross-entropy loss below 0.1 by epoch 100; near-zero by 300. 97.5% accuracy on 10-class digits from scratch — this is the 30-minute mastery benchmark.

epochCE lossphase
12.3134random init
102.0562early learning
500.4056rapid descent
1000.0752< 0.1 threshold
2000.0122converging
3000.0045converged; test acc 0.9750

36. Fill in: CE loss for 3-layer MLP on load_digits — full trace

Comparison

Comparison matrix

From 3-layer MLP on load_digits — full trace: refill the CE loss column from what you know. The rest of the table is as it appeared.

epochCE lossphase
12.3134random init
102.0562early learning
500.4056rapid descent
1000.0752< 0.1 threshold
2000.0122converging
3000.0045converged; test acc 0.9750

37. Without one step: Phase 2 algorithm-selection recipe

Constraint

Discussion prompt

Run Phase 2 algorithm-selection recipe with this step confiscated:

Pick a clusterer: k-Means / GMM when k is known; DBSCAN when shape is arbitrary / noise exists — always scale first

Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.

Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.

Answer:

  1. Diagnose error: high bias → boosting / deeper model; high variance → bagging / regularization
  2. Pick a classifier: DT for interpretability; RF when variance is the enemy; GBM when bias is; MLP for complex non-linear patterns
  3. Pick a clusterer: k-Means / GMM when k is known; DBSCAN when shape is arbitrary / noise exists — always scale first
  4. Choose k: elbow on inertia + silhouette score; GMM → BIC/AIC for formal comparison
  5. Evaluate: classification → acc/F1/ROC-AUC; clustering → ARI (if ground truth), silhouette (if not)

38. Phase 2 algorithm-selection recipe

Pattern

  1. Diagnose error: high bias → boosting / deeper model; high variance → bagging / regularization
  2. Pick a classifier: DT for interpretability; RF when variance is the enemy; GBM when bias is; MLP for complex non-linear patterns
  3. Pick a clusterer: k-Means / GMM when k is known; DBSCAN when shape is arbitrary / noise exists — always scale first
  4. Choose k: elbow on inertia + silhouette score; GMM → BIC/AIC for formal comparison
  5. Evaluate: classification → acc/F1/ROC-AUC; clustering → ARI (if ground truth), silhouette (if not)

39. Where does it stop working: Phase 2 algorithm-selection recipe

Edge cases

Discussion prompt

Phase 2 algorithm-selection recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Diagnose error: high bias → boosting / deeper model; high variance → bagging / regularization
  2. Pick a classifier: DT for interpretability; RF when variance is the enemy; GBM when bias is; MLP for complex non-linear patterns
  3. Pick a clusterer: k-Means / GMM when k is known; DBSCAN when shape is arbitrary / noise exists — always scale first
  4. Choose k: elbow on inertia + silhouette score; GMM → BIC/AIC for formal comparison
  5. Evaluate: classification → acc/F1/ROC-AUC; clustering → ARI (if ground truth), silhouette (if not)

40. Rule out three: Check yourself — splitting criteria

Elimination

Eliminate the wrong options

A decision tree node contains 40 class-A and 60 class-B samples. Which value is closest to its Gini impurity?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. 0.48
  • B. 0.97
  • C. 0.24
  • D. 0.50

Survives elimination: A

Why: Gini = 1 − (0.4² + 0.6²) = 1 − (0.16 + 0.36) = 1 − 0.52 = 0.48. Maximum Gini for two classes is 0.5 (at 50/50), so 0.48 is near-maximum — this is an impure node worth splitting.

41. Check yourself — splitting criteria

Check

Theory question — no code needed.

Check your understanding

A decision tree node contains 40 class-A and 60 class-B samples. Which value is closest to its Gini impurity?

  • A. 0.48 (correct)
  • B. 0.97
  • C. 0.24
  • D. 0.50

Answer: A

Why: Gini = 1 − (0.4² + 0.6²) = 1 − (0.16 + 0.36) = 1 − 0.52 = 0.48. Maximum Gini for two classes is 0.5 (at 50/50), so 0.48 is near-maximum — this is an impure node worth splitting.

Why B tempts people
0.97 is near the maximum binary entropy (−0.4 log 0.4 − 0.6 log 0.6 ≈ 0.971), not Gini — the two formulas are easy to confuse.
Why C tempts people
0.24 would correspond to Gini with p = 0.2/0.8 (1 − (0.04 + 0.64) = 0.32, still not 0.24). Squaring the wrong term gives this.
Why D tempts people
0.5 is the Gini at exactly 50/50 — the 40/60 split is slightly more pure, so Gini < 0.5.

42. Answer it before you see the options: Check yourself — ensemble choice

Prediction

Predict first

Your gradient-boosted model has train acc = 0.99 and val acc = 0.72. The FIRST intervention should be:

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: reduce n_estimators or add shrinkage (learning_rate < 1)

Why: Train 0.99 vs val 0.72 is a variance (overfitting) problem. Boosting with many rounds memorizes training data — reduce rounds or add learning_rate shrinkage (= L2 toward weaker ensemble) to regularize.

43. Check yourself — ensemble choice

Check

Scenario: the model generalizes poorly.

Check your understanding

Your gradient-boosted model has train acc = 0.99 and val acc = 0.72. The FIRST intervention should be:

  • A. reduce n_estimators or add shrinkage (learning_rate < 1) (correct)
  • B. switch to a deeper boosting model
  • C. add more boosting rounds
  • D. switch to a higher-capacity neural network

Answer: A

Why: Train 0.99 vs val 0.72 is a variance (overfitting) problem. Boosting with many rounds memorizes training data — reduce rounds or add learning_rate shrinkage (= L2 toward weaker ensemble) to regularize.

Why B tempts people
Deeper boosting adds capacity and worsens overfitting when train >> val; it would make the gap larger.
Why C tempts people
More rounds = more boosting iterations = more fitting to training noise = worse generalization.
Why D tempts people
A higher-capacity neural network would overfit even more severely than a GBM already overfitting — capacity is the wrong direction.

44. Rule out three: Check yourself — DBSCAN

Elimination

Eliminate the wrong options

After running DBSCAN on a dataset you get label −1 for 30% of points. This most likely means:

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. eps is too small — those points are genuine outliers or eps misses their neighbors
  • B. you forgot to call fit_predict
  • C. DBSCAN found k+1 clusters and labelled the extra one −1
  • D. the data has only one cluster

Survives elimination: A

Why: Label −1 in DBSCAN means 'noise' — points with fewer than min_samples neighbors within eps. 30% noise typically means eps is too small (too tight a radius), or the data truly has many outliers. Increase eps or decrease min_samples and recheck.

45. Check yourself — DBSCAN

Check

An exam-favorite edge case.

Check your understanding

After running DBSCAN on a dataset you get label −1 for 30% of points. This most likely means:

  • A. eps is too small — those points are genuine outliers or eps misses their neighbors (correct)
  • B. you forgot to call fit_predict
  • C. DBSCAN found k+1 clusters and labelled the extra one −1
  • D. the data has only one cluster

Answer: A

Why: Label −1 in DBSCAN means 'noise' — points with fewer than min_samples neighbors within eps. 30% noise typically means eps is too small (too tight a radius), or the data truly has many outliers. Increase eps or decrease min_samples and recheck.

Why B tempts people
fit_predict and fit+labels_ are equivalent — calling the wrong method doesn't produce −1 labels.
Why C tempts people
DBSCAN doesn't label an extra cluster −1; −1 is always noise, never a real cluster id.
Why D tempts people
One cluster would produce label 0 for all core points and −1 only for isolated outliers — not 30% noise.

46. Your turn: 30-min MLP mastery drill

Section

Project

47. Project: MLP from scratch under the clock

Concept

The objective is speed and accuracy together: implement a working, generalizing MLP from a blank file in 30 minutes.

#milestonetool
1Load + normalize load_digitsnp, train_test_split
2Define 3-layer Sequential MLPnn.Sequential, Linear, ReLU
3Training loop: 300 epochs, Adamzero_grad → loss → backward → step
4Evaluate — target > 97% test accargmax, float().mean()

Build rules: no torch.nn.utils shortcuts, no sklearn wrappers — type every line. Set torch.manual_seed(0) for reproducibility.

48. By analogy: Project: MLP from scratch under the clock

Analogy

Discussion prompt

Explain Project: MLP from scratch under the clock by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

The objective is speed and accuracy together: implement a working, generalizing MLP from a blank file in 30 minutes.

49. Milestone 1 — load and normalize

Worked example

Your turn: load digits, cast to float32, normalize per-feature, split 80/20. What shape should X_tr be?

Hint: load_digits() returns 1797×64. After split (test_size=0.2, random_state=0): X_tr is (1437, 64).

import numpy as np
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split

X, y = load_digits(return_X_y=True)
X = X.astype(np.float32)
X = (X - X.mean(0)) / (X.std(0) + 1e-8)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, random_state=0)
print(X_tr.shape, X_te.shape)
splitshapeclasses
X_tr(1437, 64)10 (0–9)
X_te(360, 64)10 (0–9)

50. Milestone 2 — define architecture and run loop

Worked example

Your turn: build the Sequential MLP and run 300 epochs with Adam. Predict the loss at epoch 100.

Hint: torch.manual_seed(0) first; .unsqueeze not needed — inputs are already 2D; loop = zero_grad → forward → CELoss → backward → step.

import torch
torch.manual_seed(0)
Xt, yt   = torch.tensor(X_tr), torch.tensor(y_tr, dtype=torch.long)
Xte, yte = torch.tensor(X_te), torch.tensor(y_te, dtype=torch.long)

model = torch.nn.Sequential(
    torch.nn.Linear(64,128), torch.nn.ReLU(),
    torch.nn.Linear(128,64), torch.nn.ReLU(),
    torch.nn.Linear(64,10))
opt     = torch.optim.Adam(model.parameters(), lr=1e-3)
loss_fn = torch.nn.CrossEntropyLoss()
for ep in range(300):
    opt.zero_grad()
    loss = loss_fn(model(Xt), yt)
    loss.backward(); opt.step()
    if ep+1 in (1,100,300): print(ep+1, round(loss.item(),4))
epochCE loss
12.3134
1000.0752
3000.0045

51. Watch it run: Milestone 2 — define architecture and run loop

Pattern

Step through it

Step through Milestone 2 — define architecture and run loop one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: epoch is 1
  2. Step 2: epoch is 100
  3. Step 3: epoch is 300

52. Milestone 3 — evaluate accuracy

Worked example

Your turn: evaluate test accuracy with argmax. Predict whether you beat 97%.

Hint: model(Xte).argmax(1) gives predicted class indices; compare element-wise to yte, cast to float, take mean.

with torch.no_grad():
    preds = model(Xte).argmax(1)
    acc = (preds == yte).float().mean().item()
print(f'test acc: {acc:.4f}')
metricvaluetarget
test accuracy0.9750> 0.97 ✓
CE loss (epoch 300)0.0045< 0.01 ✓

53. What each one costs: Milestone 3 — evaluate accuracy

Trade off

Comparison matrix

From Milestone 3 — evaluate accuracy: every row here is a choice with a cost. Fill the target column, then say which row you would actually pick and what you give up for it.

metricvaluetarget
test accuracy0.9750> 0.97 ✓
CE loss (epoch 300)0.0045< 0.01 ✓

54. The full program

Concept

import numpy as np, torch
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split

X, y = load_digits(return_X_y=True)
X = X.astype(np.float32)
X = (X - X.mean(0)) / (X.std(0) + 1e-8)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, random_state=0)
torch.manual_seed(0)
Xt, yt   = torch.tensor(X_tr), torch.tensor(y_tr, dtype=torch.long)
Xte, yte = torch.tensor(X_te), torch.tensor(y_te, dtype=torch.long)
model = torch.nn.Sequential(
    torch.nn.Linear(64,128), torch.nn.ReLU(),
    torch.nn.Linear(128,64), torch.nn.ReLU(),
    torch.nn.Linear(64,10))
opt = torch.optim.Adam(model.parameters(), lr=1e-3)
for _ in range(300):
    opt.zero_grad()
    torch.nn.CrossEntropyLoss()(model(Xt),yt).backward(); opt.step()
with torch.no_grad():
    print((model(Xte).argmax(1)==yte).float().mean().item())
outputvalue
test accuracy0.9750
30-min targetpassed ✓

If you can reproduce this from memory in under 30 minutes, you've cleared the Phase 2 PyTorch mastery benchmark.

55. Fill in: value for The full program

Comparison

Comparison matrix

From The full program: refill the value column from what you know. The rest of the table is as it appeared.

outputvalue
test accuracy0.9750
30-min targetpassed ✓

56. Show it off

Concept

Out loud, slides closed: (1) explain when you'd pick GBM over Random Forest and why; (2) state what DBSCAN label −1 means and how to fix it; (3) recite the 5-step PyTorch training loop.

Homework drill: create a one-page summary sheet listing every Phase 2 algorithm, its splitting/training criterion, its top hyperparameter, and its canonical failure mode. Then run the speed drill — k-Means, logistic regression, 3-layer CNN — in 30-minute timed blocks.

57. Break it if you can: Show it off

Counterexample

Discussion prompt

Out loud, slides closed: (1) explain when you'd pick GBM over Random Forest and why; (2) state what DBSCAN label −1 means and how to fix it; (3) recite the 5-step PyTorch training loop.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

58. Connect it up: Lesson 72: Phase 2 Algorithm Review

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Decision trees — criteria, pruning, tradeoffs · Ensembles — bagging, boosting, stacking · Clustering — k-Means, GMM, DBSCAN, hierarchical · PyTorch mastery check — 30-minute MLP · Your turn: 30-min MLP mastery drill. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

59. Phase 2 algorithm review — what you can do now

Recap

algorithmthe one thing to remember
Decision Treegreedy splits; high variance without pruning
Random Forestbagging ↓ variance; use when single DT overfits
GBMboosting ↓ bias; use when model underfits
k-Means / GMMneeds k; GMM is soft + elliptical
DBSCANscale first; eps too large → all one cluster
PyTorch MLPzero_grad every step; Adam + CE for classification

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 72 (Week 25 — Phase 2 Consolidation) — Barron · USAAIO Round 2 Preparation, 2026
  2. Decision tree, ensemble, and clustering metrics verified on iris/digits/make_blobs; PyTorch MLP verified on load_digits — torch 2.7.1 + sklearn + numpy 2.2.6, real execution, June 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108