USAAIO Lesson 75, from Week 26: a timed Phase 2 mock exam. The two-hour theory paper covers ensemble methods, normalization, optimization, and evaluation, and the two-hour coding paper asks you to implement a random forest, a GMM, and a CNN training loop from scratch. It comes with a full answer-key walkthrough and a gap analysis to run before Phase 3 on transformers. All the numbers were verified against torch 2.7.1, sklearn, and numpy 2.2.6. The lesson runs to 32 slides.
Subject: Machine Learning · 57 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 75 · Week 26
2-hour theory + 2-hour coding: ensemble methods, normalization, optimization, evaluation — implement random forest, GMM, and CNN from scratch. The Phase-2 coding milestone.
Objectives
Warm-up
Discussion prompt
Before we open Lesson 75: Mock Exam — Phase 2 Full Review: without looking back, what was the main idea of Neural Network Debugging Checklist, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
a systematic 5-step debugging protocol for broken training loops — overfitting a single batch, checking gradient flow with per-layer norms, verifying initial loss against log(K), checking initial accuracy against 1/K, and choosing literature hyperparameters first. Includes an automated checklist function and a bug-injection exercise.
Section
Part 1 of 4
Concept
USAAIO Phase 2 exam: a 2-hour theory block then a 2-hour coding block in Colab. Theory answers guide the code — always do theory first.
| block | time | content |
|---|---|---|
| Theory | 2 hr | ensemble methods, normalization, optimization, evaluation |
| Coding | 2 hr | RF, GMM, CNN — from scratch in NumPy / PyTorch |
| Post-exam | — | solution walkthrough + gap analysis |
Comparison
Comparison matrix
From Two-block structure: refill the time column from what you know. The rest of the table is as it appeared.
| block | time | content |
|---|---|---|
| Theory | 2 hr | ensemble methods, normalization, optimization, evaluation |
| Coding | 2 hr | RF, GMM, CNN — from scratch in NumPy / PyTorch |
| Post-exam | — | solution walkthrough + gap analysis |
Ranking
Put in order
These are the steps of Per-problem protocol, scrambled. Put them back in order before the next slide shows you.
zero_grad)Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
zero_grad)Section
Part 2 of 4
Concept
Random forests combine bagging (bootstrap aggregating) with random feature subsets. Each tree sees a different bootstrapped sample; majority vote (classification) or mean (regression) reduces variance without increasing bias.
| concept | key formula / fact |
|---|---|
| Gini impurity | G = 1 − Σ pₖ²; pure node → 0, equal split (binary) → 0.5 |
| Bootstrap retention | ≈ 63.2% of samples in-bag (1 − 1/n)ⁿ → 1 − 1/e |
| OOB error | validate each tree on its ~36.8% out-of-bag samples |
| Gradient boosting | fits additive residuals with learning rate η; high variance ↔ RF |
Concept
Normalization prevents gradient-magnitude differences from dominating optimization. Iris sepal-length: mean = 5.84, std = 0.83, range [4.3, 7.9].
| method | formula | result on feature 0 |
|---|---|---|
| StandardScaler | (x − μ)/σ | mean → 0, std → 1 |
| MinMaxScaler | (x − min)/(max − min) | range → [0, 1] |
| BatchNorm (neural) | normalize per-batch, learnable γ, β | stable layer inputs |
Trade off
Comparison matrix
From Theory topic 2 — Normalization: every row here is a choice with a cost. Fill the result on feature 0 column, then say which row you would actually pick and what you give up for it.
| method | formula | result on feature 0 |
|---|---|---|
| StandardScaler | (x − μ)/σ | mean → 0, std → 1 |
| MinMaxScaler | (x − min)/(max − min) | range → [0, 1] |
| BatchNorm (neural) | normalize per-batch, learnable γ, β | stable layer inputs |
Concept
SGD update (L40): w ← w − η·∇L. Adam (L56) uses adaptive per-parameter rates via first and second moment estimates.
| optimizer | update rule | hyperparams |
|---|---|---|
| SGD | w − η·g | η (lr) |
| SGD + momentum | w − η·(β·m + g) | η, β≈0.9 |
| Adam | w − η·m̂/(√v̂ + ε) | η, β₁=0.9, β₂=0.999, ε=1e-8 |
| weight_decay | adds λ‖w‖² to loss (= L2/ridge) | λ |
Counterexample
Discussion prompt
SGD update (L40): w ← w − η·∇L. Adam (L56) uses adaptive per-parameter rates via first and second moment estimates.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Concept
On 8 predictions (TP=3, FP=1, FN=1, TN=3): Precision = 0.75, Recall = 0.75, F1 = 0.75. AUC measures ranking quality — 1.0 = perfect separation.
| metric | formula | value (example) |
|---|---|---|
| Precision | TP / (TP + FP) | 3/4 = 0.75 |
| Recall | TP / (TP + FN) | 3/4 = 0.75 |
| F1 | 2·P·R / (P + R) | 0.75 |
| AUC-ROC | area under TPR vs FPR curve | RF train AUC ≈ 1.0 |
Analogy
Discussion prompt
Explain Theory topic 4 — Evaluation by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
On 8 predictions (TP=3, FP=1, FN=1, TN=3): Precision = 0.75, Recall = 0.75, F1 = 0.75. AUC measures ranking quality — 1.0 = perfect separation.
Anomaly
Predict first
A student writes this, and it looks reasonable:
A node with 50 class-A and 50 class-B is balanced — Gini = 0 (pure).
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Gini = 0 only when one class owns the whole node (pure leaf).
Impurity is maximized at the equal split, minimized at purity.
Why: Gini = 0 only when one class owns the whole node (pure leaf). G = 1 − Σpₖ² with p=[0.5,0.5] gives 1 − 0.25 − 0.25 = 0.5, the worst possible case.
Trap
A node with 50 class-A and 50 class-B is balanced — Gini = 0 (pure).
Write Gini = 0 for [0.5, 0.5]
Why: Gini = 0 only when one class owns the whole node (pure leaf). G = 1 − Σpₖ² with p=[0.5,0.5] gives 1 − 0.25 − 0.25 = 0.5, the worst possible case.
Impurity is maximized at the equal split, minimized at purity.
Gini([0.5, 0.5]) = 1 − 2·(0.5²) = 0.50; Gini([1.0, 0.0]) = 0
Why: The split that does no discrimination has maximum Gini. A pure node — where the algorithm stops splitting — has G = 0.
Section
Part 3 of 4
Worked example
Theory sub-q: why does bagging reduce variance without increasing bias? Then: implement 3 bootstrapped stumps + majority vote on iris; report OOB accuracy per tree and ensemble accuracy.
Hint: np.random.choice(n, size=n, replace=True) for bootstrap; np.setdiff1d for OOB indices; majority class via np.bincount(...).argmax(). Predict ensemble ≥ any individual tree.
import numpy as np
from sklearn.datasets import load_iris
from sklearn.tree import DecisionTreeClassifier
X, y = load_iris(return_X_y=True)
n = len(X)
stumps = []
for seed in [42, 7, 13]:
rng = np.random.default_rng(seed)
idx = rng.choice(n, size=n, replace=True)
oob = np.setdiff1d(np.arange(n), idx)
dt = DecisionTreeClassifier(max_depth=2, random_state=seed)
dt.fit(X[idx], y[idx])
stumps.append(dt)
print(f'seed={seed} OOB={dt.score(X[oob],y[oob]):.4f}')
preds = np.stack([s.predict(X) for s in stumps], axis=1)
votes = np.apply_along_axis(lambda r: np.bincount(r,minlength=3).argmax(),1,preds)
print('ensemble acc:', (votes==y).mean().round(4))| tree (seed) | OOB accuracy | in-bag size |
|---|---|---|
| 42 | 0.9492 | 150 |
| 7 | 0.9574 | 150 |
| 13 | 0.9322 | 150 |
| 3-stump ensemble | 0.9600 | — |
Worked example
Theory sub-q: in the E-step, what do the responsibilities rᵢₖ represent, and why do they sum to 1 across k? Then: compute responsibilities for x ∈ {−3, −1, 0, 2, 4} given μ₁=−2, μ₂=3, σ=1, π₁=π₂=0.5.
Hint: rᵢₖ = πₖ·N(xᵢ|μₖ,σ²) / Σⱼ πⱼ·N(xᵢ|μⱼ,σ²). The denominator is the marginal p(xᵢ); dividing by it normalizes over components.
import numpy as np
from scipy.stats import norm
mus=[-2.,3.]; sig=1.; pis=[.5,.5]
X=[-3.,-1.,0.,2.,4.]
for xi in X:
p=[pis[k]*norm.pdf(xi,mus[k],sig) for k in range(2)]
s=sum(p)
r=[round(pk/s,4) for pk in p]
print(f'x={xi:4}: r0={r[0]}, r1={r[1]}')| x | r(k=0) | r(k=1) | assigned |
|---|---|---|---|
| -3.0 | 1.0000 | 0.0000 | k=0 |
| -1.0 | 0.9994 | 0.0006 | k=0 |
| 0.0 | 0.9241 | 0.0759 | k=0 |
| 2.0 | 0.0006 | 0.9994 | k=1 |
| 4.0 | 0.0000 | 1.0000 | k=1 |
Discrimination
Sort into buckets
Sort these by assigned, from memory, without looking back at Problem 2 — GMM E-step from scratch. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Worked example
Theory sub-q: what is the output shape after Conv2d(1, 16, kernel_size=3, padding=1) on a (N, 1, 8, 8) input? Then: build a SmallCNN (1 conv + 1 linear) on digits, train 20 epochs with Adam, print loss trajectory and accuracy.
Hint: shape formula H_out = (H_in + 2p − k)/s + 1; with p=1, k=3, s=1 → 8. Conv params = out_ch × in_ch × k × k + out_ch = 16×1×3×3 + 16 = 160. Train loop: zero_grad → forward → CrossEntropyLoss → backward → step.
import torch, torch.nn as nn
from sklearn.datasets import load_digits
import numpy as np
X,y=load_digits(return_X_y=True)
Xt=torch.tensor(X.reshape(-1,1,8,8)/16.,dtype=torch.float32)
yt=torch.tensor(y)
torch.manual_seed(42)
class CNN(nn.Module):
def __init__(self):
super().__init__()
self.conv=nn.Conv2d(1,16,3,padding=1)
self.fc=nn.Linear(16*8*8,10)
def forward(self,x):
return self.fc(torch.relu(self.conv(x)).view(x.size(0),-1))
model=CNN(); crit=nn.CrossEntropyLoss()
opt=torch.optim.Adam(model.parameters(),lr=5e-3)
for ep in range(20):
opt.zero_grad(); loss=crit(model(Xt),yt)
loss.backward(); opt.step()
if ep in [0,4,9,19]: print(f'ep {ep+1} loss={loss.item():.4f}')
acc=(model(Xt).argmax(1)==yt).float().mean()
print(f'train acc={acc:.4f}')| epoch | loss | note |
|---|---|---|
| 1 | 2.2961 | near random (log 10 ≈ 2.30) |
| 5 | 1.6783 | falling steadily |
| 10 | 0.9558 | below 1.0 |
| 20 | 0.3140 | converged; train acc = 0.9349 |
Comparison
Comparison matrix
From Problem 3 — CNN training loop from scratch: refill the note column from what you know. The rest of the table is as it appeared.
| epoch | loss | note |
|---|---|---|
| 1 | 2.2961 | near random (log 10 ≈ 2.30) |
| 5 | 1.6783 | falling steadily |
| 10 | 0.9558 | below 1.0 |
| 20 | 0.3140 | converged; train acc = 0.9349 |
Anomaly
Predict first
A student writes this, and it looks reasonable:
Compute rᵢₖ = πₖ · N(xᵢ | μₖ, σ²) and use it directly — it's already the 'soft assignment'.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: The unnormalized product is just a weighted density, not a probability.
Always divide by the sum over all K components.
Why: The unnormalized product is just a weighted density, not a probability. Without the denominator it may sum to 0.98 or 1.3 — the M-step will then compute wrong means and the algorithm diverges.
Trap
Compute rᵢₖ = πₖ · N(xᵢ | μₖ, σ²) and use it directly — it's already the 'soft assignment'.
Skip normalizing by Σⱼ πⱼ·N(xᵢ|μⱼ,σ²)
Why: The unnormalized product is just a weighted density, not a probability. Without the denominator it may sum to 0.98 or 1.3 — the M-step will then compute wrong means and the algorithm diverges.
Always divide by the sum over all K components.
rᵢₖ = πₖ·N(xᵢ|μₖ,σ²) / Σⱼ πⱼ·N(xᵢ|μⱼ,σ²)
Why: This is Bayes' theorem: the denominator is p(xᵢ), marginalizing out the component. Each row (over k) sums exactly to 1 — which the test verifies by printing sum(r).
Break the constraint
Discussion prompt
The rule this trap just fixed:
Always divide by the sum over all K components.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
The unnormalized product is just a weighted density, not a probability. Without the denominator it may sum to 0.98 or 1.3 — the M-step will then compute wrong means and the algorithm diverges.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Input is (N, 1, 8, 8). Conv2d(1, 16, kernel_size=3) with NO padding → output is (N, 16, 8, 8).
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Without padding, each spatial dim shrinks: H_out = (8 − 3)/1 + 1 = 6.
Apply H_out = (H_in + 2p − k)/s + 1; add padding=1 to preserve spatial size.
Why: Without padding, each spatial dim shrinks: H_out = (8 − 3)/1 + 1 = 6. The output is (N, 16, 6, 6). Passing 16×6×6=576 to a Linear(16*8*8, 10) crashes with a shape mismatch.
Trap
Input is (N, 1, 8, 8). Conv2d(1, 16, kernel_size=3) with NO padding → output is (N, 16, 8, 8).
Write H_out = 8 without applying the formula
Why: Without padding, each spatial dim shrinks: H_out = (8 − 3)/1 + 1 = 6. The output is (N, 16, 6, 6). Passing 16×6×6=576 to a Linear(16*8*8, 10) crashes with a shape mismatch.
Apply H_out = (H_in + 2p − k)/s + 1; add padding=1 to preserve spatial size.
Conv2d(1, 16, 3, padding=1) → H_out = (8 + 2 − 3)/1 + 1 = 8; output (N, 16, 8, 8)
Why: Padding=1 pads one zero on each edge, preserving the 8×8 spatial size. The fc layer then receives 16×8×8 = 1024 features — as coded.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
w ← w − η·∇L. Adam (L56) uses adaptive per-parameter rates via first and second moment estimates.Ranking
Put in order
These are the steps of Phase 2 exam recipe, scrambled. Put them back in order before the next slide shows you.
zero_grad → loss → backward → stepWhy: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
zero_grad → loss → backward → stepEdge cases
Discussion prompt
Phase 2 exam recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
zero_grad → loss → backward → stepElimination
Eliminate the wrong options
A decision tree node contains 80 class-A and 20 class-B samples (p = [0.8, 0.2]). Its Gini impurity is:
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Gini = 1 − (0.8² + 0.2²) = 1 − (0.64 + 0.04) = 1 − 0.68 = 0.32. Equal split gives max 0.50; pure node gives 0.
Check
Evaluate before clicking.
Check your understanding
A decision tree node contains 80 class-A and 20 class-B samples (p = [0.8, 0.2]). Its Gini impurity is:
Answer: A
Why: Gini = 1 − (0.8² + 0.2²) = 1 − (0.64 + 0.04) = 1 − 0.68 = 0.32. Equal split gives max 0.50; pure node gives 0.
Prediction
Predict first
Starting at w = 3.0 with f(w) = w², after ONE Adam step (lr=0.1, β₁=0.9, β₂=0.999, ε=1e-8) the new w is approximately:
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: 2.90
Why: g=6; m̂=6 (bias-corrected after step 1); v̂=36; update = 0.1·6/√36 = 0.1·1 = 0.1; w_new = 3.0 − 0.1 = 2.90. Adam normalizes by the gradient magnitude so the effective step is exactly lr on the first iteration.
Check
Recall from Lesson 56.
Check your understanding
Starting at w = 3.0 with f(w) = w², after ONE Adam step (lr=0.1, β₁=0.9, β₂=0.999, ε=1e-8) the new w is approximately:
Answer: A
Why: g=6; m̂=6 (bias-corrected after step 1); v̂=36; update = 0.1·6/√36 = 0.1·1 = 0.1; w_new = 3.0 − 0.1 = 2.90. Adam normalizes by the gradient magnitude so the effective step is exactly lr on the first iteration.
Elimination
Eliminate the wrong options
In the GMM E-step the responsibility r(i,k) is computed for sample xᵢ and component k. What must be true of the responsibilities over all K components for a single sample xᵢ?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Each sample's responsibilities sum to 1 across components (they are posterior probabilities P(k|xᵢ)). This is enforced by the normalizing denominator in Bayes' theorem. The M-step uses these as soft weights to re-estimate means, covariances, and mixing proportions.
Check
Core EM concept.
Check your understanding
In the GMM E-step the responsibility r(i,k) is computed for sample xᵢ and component k. What must be true of the responsibilities over all K components for a single sample xᵢ?
Answer: A
Why: Each sample's responsibilities sum to 1 across components (they are posterior probabilities P(k|xᵢ)). This is enforced by the normalizing denominator in Bayes' theorem. The M-step uses these as soft weights to re-estimate means, covariances, and mixing proportions.
Section
Part 4 of 4
Concept
After the exam: for every mistake, assign it to one Phase 2 topic. The distribution tells you exactly what to reinforce before Phase 3.
| topic area | Lessons | common error type |
|---|---|---|
| Training loop mechanics | L40–L42 | zero_grad order, shape mismatch |
| Backprop / MLP | L43–L45 | chain rule application, weight init |
| CNNs | L46–L50 | output shape formula, padding vs stride |
| Regularization | L51–L55 | L1 vs L2, when dropout hurts |
| Optimization | L56–L60 | Adam bias correction, LR schedule choice |
| Ensembles | L61–L65 | OOB vs validation, Gini formula |
| Normalization + Eval | L66–L70 | StandardScaler leakage, AUC interpretation |
| GMM / EM | L71–L74 | E-step normalization, convergence criterion |
Explain it
Discussion prompt
Explain Error analysis — categorize by topic to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
After the exam: for every mistake, assign it to one Phase 2 topic. The distribution tells you exactly what to reinforce before Phase 3.
Concept
Phase 2 coding milestone: you can now implement and train — without a reference — random forest, GMM, and a CNN training loop in PyTorch. If any of the three still needed a hint, that topic is the priority before Phase 3.
Verify each implementation against the library (RF vs sklearn RF, GMM E-step vs GaussianMixture, CNN loss trajectory vs known cross-entropy at random ≈ 2.30). Library agreement = coding milestone passed.
Concept
Phase 3 begins immediately after this exam. The pivot: from convolution (local receptive fields) to self-attention (global, content-based). The transformer training loop is the same PyTorch skeleton — only the model changes.
torch.cuda if available; Phase 3 models are largerSection
Project
Concept
Re-implement all three problems cleanly — no hints — then produce verified output matching the answer key. This is the Phase 2 coding milestone submission.
| # | implement | verify against |
|---|---|---|
| 1 | 3-stump RF on iris | sklearn RF; OOB ≈ 0.94, ensemble acc ≈ 0.96 |
| 2 | GMM E-step (5 points) | scipy.stats.norm; r(x=0)=(0.9241, 0.0759) |
| 3 | CNN on digits, 20 ep | loss[1]≈2.30, loss[20]≈0.31, acc≈0.93 |
Build rule: type every line from scratch. No copy-paste from earlier decks. The act of re-typing exposes every shape assumption.
Analogy
Discussion prompt
Explain Project brief by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Re-implement all three problems cleanly — no hints — then produce verified output matching the answer key. This is the Phase 2 coding milestone submission.
Worked example
Your turn: implement 3-stump bootstrap ensemble on iris. Predict that ensemble accuracy ≥ each individual tree's OOB accuracy.
Hint: rng.choice(n, n, replace=True) for bootstrap; np.setdiff1d for OOB; stack predictions and take np.bincount(...).argmax() majority vote.
import numpy as np
from sklearn.datasets import load_iris
from sklearn.tree import DecisionTreeClassifier
X,y=load_iris(return_X_y=True); n=len(X)
stumps=[]
for seed in [42,7,13]:
rng=np.random.default_rng(seed)
idx=rng.choice(n,n,replace=True)
oob=np.setdiff1d(np.arange(n),idx)
dt=DecisionTreeClassifier(max_depth=2,random_state=seed)
dt.fit(X[idx],y[idx]); stumps.append(dt)
print(f'seed={seed} OOB={dt.score(X[oob],y[oob]):.4f}')
preds=np.stack([s.predict(X) for s in stumps],axis=1)
votes=np.apply_along_axis(lambda r:np.bincount(r,3).argmax(),1,preds)
print('ensemble acc:',(votes==y).mean().round(4))| tree | OOB acc |
|---|---|
| seed=42 | 0.9492 |
| seed=7 | 0.9574 |
| seed=13 | 0.9322 |
| ensemble | 0.9600 |
Pattern
Step through it
Step through Milestone 1 — RF with OOB verification one row at a time. What is driving the change, and what would the row after the last one be?
Worked example
Your turn: compute rᵢₖ for x ∈ {−3,−1,0,2,4} with μ₁=−2, μ₂=3, σ=1, π₁=π₂=0.5. Predict that x=0 stays with k=0 (closer to −2 than to 3).
Hint: import scipy.stats.norm; for each x compute numerator list, divide by sum. Verify sum(r) == 1.0 for every row.
from scipy.stats import norm
mus=[-2.,3.]; sig=1.; pis=[.5,.5]
for xi in [-3.,-1.,0.,2.,4.]:
p=[pis[k]*norm.pdf(xi,mus[k],sig) for k in range(2)]
s=sum(p)
r=[round(pk/s,4) for pk in p]
print(f'x={xi:5}: r0={r[0]}, r1={r[1]}, sum={sum(r)}')| x | r0 | r1 | sum |
|---|---|---|---|
| -3.0 | 1.0000 | 0.0000 | 1.0 |
| -1.0 | 0.9994 | 0.0006 | 1.0 |
| 0.0 | 0.9241 | 0.0759 | 1.0 |
| 2.0 | 0.0006 | 0.9994 | 1.0 |
| 4.0 | 0.0000 | 1.0000 | 1.0 |
Trade off
Comparison matrix
From Milestone 2 — GMM E-step verification: every row here is a choice with a cost. Fill the r0 column, then say which row you would actually pick and what you give up for it.
| x | r0 | r1 | sum |
|---|---|---|---|
| -3.0 | 1.0000 | 0.0000 | 1.0 |
| -1.0 | 0.9994 | 0.0006 | 1.0 |
| 0.0 | 0.9241 | 0.0759 | 1.0 |
| 2.0 | 0.0006 | 0.9994 | 1.0 |
| 4.0 | 0.0000 | 1.0000 | 1.0 |
Worked example
Your turn: build DigitCNN (Conv2d(1,16,3,pad=1) → ReLU → Linear(1024,10)), train 20 epochs with Adam(lr=5e-3), print losses at epochs 1/5/10/20 and final train accuracy.
Hint: reshape digits to (N,1,8,8) and divide by 16.0 to normalize; loss at epoch 1 should be near log(10) ≈ 2.30 (near-random initialization); after 20 epochs train acc should be ≥ 0.90.
import torch, torch.nn as nn
from sklearn.datasets import load_digits
X,y=load_digits(return_X_y=True)
Xt=torch.tensor(X.reshape(-1,1,8,8)/16.,dtype=torch.float32)
yt=torch.tensor(y)
torch.manual_seed(42)
class CNN(nn.Module):
def __init__(self):
super().__init__()
self.conv=nn.Conv2d(1,16,3,padding=1)
self.fc=nn.Linear(16*8*8,10)
def forward(self,x):
return self.fc(torch.relu(self.conv(x)).view(x.size(0),-1))
model=CNN(); crit=nn.CrossEntropyLoss()
opt=torch.optim.Adam(model.parameters(),lr=5e-3)
for ep in range(20):
opt.zero_grad(); loss=crit(model(Xt),yt)
loss.backward(); opt.step()
if ep in [0,4,9,19]: print(f'ep{ep+1} loss={loss.item():.4f}')
print('acc=',(model(Xt).argmax(1)==yt).float().mean().item())| epoch | loss |
|---|---|
| 1 | 2.2961 |
| 5 | 1.6783 |
| 10 | 0.9558 |
| 20 | 0.3140 |
| train acc | 0.9349 |
Comparison
Comparison matrix
From Milestone 3 — CNN answer key: refill the loss column from what you know. The rest of the table is as it appeared.
| epoch | loss |
|---|---|
| 1 | 2.2961 |
| 5 | 1.6783 |
| 10 | 0.9558 |
| 20 | 0.3140 |
| train acc | 0.9349 |
Concept
Slides closed, out loud: (1) explain why bagging reduces variance but not bias; (2) walk the GMM E-step for x=0 to r₀=0.92 step by step; (3) state the five steps of the PyTorch training loop and what breaks if you skip zero_grad.
Stretch: add gradient boosting to Problem 1 and compare with RF; run more GMM EM iterations manually to watch μ converge; add BatchNorm between the CNN conv and relu and compare convergence.
Counterexample
Discussion prompt
Stretch: add gradient boosting to Problem 1 and compare with RF; run more GMM EM iterations manually to watch μ converge; add BatchNorm between the CNN conv and relu and compare convergence.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Exam format & strategy · Theory block: four topics · Coding block: three problems · Gap analysis & Phase 3 prep · Your turn: full answer key. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
| algorithm | the one gotcha |
|---|---|
| Random forest | Gini([0.5,0.5]) = 0.50, NOT 0 — max impurity |
| GMM E-step | always divide by the marginal: Σⱼ πⱼN(x|μⱼ,σⱼ) |
| CNN shapes | H_out = (H + 2p − k)/s + 1; padding=1 preserves 8×8 |
| Adam step 1 | m̂/√v̂ = g/|g| = ±1 → effective lr = lr exactly |
| PyTorch loop | zero_grad FIRST — accumulated grads diverge the model |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.