Lesson 75: Mock Exam — Phase 2 Full Review

USAAIO Lesson 75, from Week 26: a timed Phase 2 mock exam. The two-hour theory paper covers ensemble methods, normalization, optimization, and evaluation, and the two-hour coding paper asks you to implement a random forest, a GMM, and a CNN training loop from scratch. It comes with a full answer-key walkthrough and a gap analysis to run before Phase 3 on transformers. All the numbers were verified against torch 2.7.1, sklearn, and numpy 2.2.6. The lesson runs to 32 slides.

Subject: Machine Learning · 57 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Mock Exam Phase 2 Full Review

Title

USAAIO · Lesson 75 · Week 26

2-hour theory + 2-hour coding: ensemble methods, normalization, optimization, evaluation — implement random forest, GMM, and CNN from scratch. The Phase-2 coding milestone.

2. Exam goals

Objectives

  1. Answer theory sub-questions on ensembles, normalization, optimization, and evaluation before coding
  2. Implement random forest (bootstrap + majority vote) from scratch and verify OOB accuracy
  3. Implement the GMM E-step by hand and confirm responsibilities match sklearn
  4. Build a CNN training loop (Conv2d → ReLU → Linear) in PyTorch and watch loss fall
  5. Perform a gap analysis: categorize Phase 2 errors by topic before Phase 3

3. What survived from Neural Network Debugging Checklist?

Warm-up

Discussion prompt

Before we open Lesson 75: Mock Exam — Phase 2 Full Review: without looking back, what was the main idea of Neural Network Debugging Checklist, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

a systematic 5-step debugging protocol for broken training loops — overfitting a single batch, checking gradient flow with per-layer norms, verifying initial loss against log(K), checking initial accuracy against 1/K, and choosing literature hyperparameters first. Includes an automated checklist function and a bug-injection exercise.

4. Exam format & strategy

Section

Part 1 of 4

5. Two-block structure

Concept

USAAIO Phase 2 exam: a 2-hour theory block then a 2-hour coding block in Colab. Theory answers guide the code — always do theory first.

blocktimecontent
Theory2 hrensemble methods, normalization, optimization, evaluation
Coding2 hrRF, GMM, CNN — from scratch in NumPy / PyTorch
Post-exam—solution walkthrough + gap analysis

6. Fill in: time for Two-block structure

Comparison

Comparison matrix

From Two-block structure: refill the time column from what you know. The rest of the table is as it appeared.

blocktimecontent
Theory2 hrensemble methods, normalization, optimization, evaluation
Coding2 hrRF, GMM, CNN — from scratch in NumPy / PyTorch
Post-exam—solution walkthrough + gap analysis

7. Rebuild the recipe: Per-problem protocol

Ranking

Put in order

These are the steps of Per-problem protocol, scrambled. Put them back in order before the next slide shows you.

  1. Theory sub-question first — it forces you to write the algorithm before coding it
  2. Predict the expected output (shape, sign, magnitude) before running
  3. Implement from scratch — NumPy for classical; PyTorch for neural models
  4. Verify against sklearn / known analytic value
  5. Review: edge cases (bootstrap size, responsibilities sum to 1, zero_grad)

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

8. Per-problem protocol

Pattern

  1. Theory sub-question first — it forces you to write the algorithm before coding it
  2. Predict the expected output (shape, sign, magnitude) before running
  3. Implement from scratch — NumPy for classical; PyTorch for neural models
  4. Verify against sklearn / known analytic value
  5. Review: edge cases (bootstrap size, responsibilities sum to 1, zero_grad)

9. Theory block: four topics

Section

Part 2 of 4

10. Theory topic 1 — Ensemble methods

Concept

Random forests combine bagging (bootstrap aggregating) with random feature subsets. Each tree sees a different bootstrapped sample; majority vote (classification) or mean (regression) reduces variance without increasing bias.

conceptkey formula / fact
Gini impurityG = 1 − Σ pₖ²; pure node → 0, equal split (binary) → 0.5
Bootstrap retention≈ 63.2% of samples in-bag (1 − 1/n)ⁿ → 1 − 1/e
OOB errorvalidate each tree on its ~36.8% out-of-bag samples
Gradient boostingfits additive residuals with learning rate η; high variance ↔ RF

11. Theory topic 2 — Normalization

Concept

Normalization prevents gradient-magnitude differences from dominating optimization. Iris sepal-length: mean = 5.84, std = 0.83, range [4.3, 7.9].

methodformularesult on feature 0
StandardScaler(x − μ)/σmean → 0, std → 1
MinMaxScaler(x − min)/(max − min)range → [0, 1]
BatchNorm (neural)normalize per-batch, learnable γ, βstable layer inputs

12. What each one costs: Theory topic 2 — Normalization

Trade off

Comparison matrix

From Theory topic 2 — Normalization: every row here is a choice with a cost. Fill the result on feature 0 column, then say which row you would actually pick and what you give up for it.

methodformularesult on feature 0
StandardScaler(x − μ)/σmean → 0, std → 1
MinMaxScaler(x − min)/(max − min)range → [0, 1]
BatchNorm (neural)normalize per-batch, learnable γ, βstable layer inputs

13. Theory topic 3 — Optimization

Concept

SGD update (L40): w ← w − η·∇L. Adam (L56) uses adaptive per-parameter rates via first and second moment estimates.

optimizerupdate rulehyperparams
SGDw − η·gη (lr)
SGD + momentumw − η·(β·m + g)η, β≈0.9
Adamw − η·m̂/(√v̂ + ε)η, β₁=0.9, β₂=0.999, ε=1e-8
weight_decayadds λ‖w‖² to loss (= L2/ridge)λ

14. Break it if you can: Theory topic 3 — Optimization

Counterexample

Discussion prompt

SGD update (L40): w ← w − η·∇L. Adam (L56) uses adaptive per-parameter rates via first and second moment estimates.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

15. Theory topic 4 — Evaluation

Concept

On 8 predictions (TP=3, FP=1, FN=1, TN=3): Precision = 0.75, Recall = 0.75, F1 = 0.75. AUC measures ranking quality — 1.0 = perfect separation.

metricformulavalue (example)
PrecisionTP / (TP + FP)3/4 = 0.75
RecallTP / (TP + FN)3/4 = 0.75
F12·P·R / (P + R)0.75
AUC-ROCarea under TPR vs FPR curveRF train AUC ≈ 1.0

16. By analogy: Theory topic 4 — Evaluation

Analogy

Discussion prompt

Explain Theory topic 4 — Evaluation by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

On 8 predictions (TP=3, FP=1, FN=1, TN=3): Precision = 0.75, Recall = 0.75, F1 = 0.75. AUC measures ranking quality — 1.0 = perfect separation.

17. Something is wrong here: Gini impurity at a 50/50 split is NOT 0

Anomaly

Predict first

A student writes this, and it looks reasonable:

A node with 50 class-A and 50 class-B is balanced — Gini = 0 (pure).

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Gini = 0 only when one class owns the whole node (pure leaf).

Impurity is maximized at the equal split, minimized at purity.

Why: Gini = 0 only when one class owns the whole node (pure leaf). G = 1 − Σpₖ² with p=[0.5,0.5] gives 1 − 0.25 − 0.25 = 0.5, the worst possible case.

18. Trap: Gini impurity at a 50/50 split is NOT 0

Trap

The trap

A node with 50 class-A and 50 class-B is balanced — Gini = 0 (pure).

Write Gini = 0 for [0.5, 0.5]

Why: Gini = 0 only when one class owns the whole node (pure leaf). G = 1 − Σpₖ² with p=[0.5,0.5] gives 1 − 0.25 − 0.25 = 0.5, the worst possible case.

The fix

Impurity is maximized at the equal split, minimized at purity.

Gini([0.5, 0.5]) = 1 − 2·(0.5²) = 0.50; Gini([1.0, 0.0]) = 0

Why: The split that does no discrimination has maximum Gini. A pure node — where the algorithm stops splitting — has G = 0.

19. Coding block: three problems

Section

Part 3 of 4

20. Problem 1 — Random forest from scratch

Worked example

Theory sub-q: why does bagging reduce variance without increasing bias? Then: implement 3 bootstrapped stumps + majority vote on iris; report OOB accuracy per tree and ensemble accuracy.

Hint: np.random.choice(n, size=n, replace=True) for bootstrap; np.setdiff1d for OOB indices; majority class via np.bincount(...).argmax(). Predict ensemble ≥ any individual tree.

import numpy as np
from sklearn.datasets import load_iris
from sklearn.tree import DecisionTreeClassifier
X, y = load_iris(return_X_y=True)
n = len(X)
stumps = []
for seed in [42, 7, 13]:
    rng = np.random.default_rng(seed)
    idx = rng.choice(n, size=n, replace=True)
    oob = np.setdiff1d(np.arange(n), idx)
    dt = DecisionTreeClassifier(max_depth=2, random_state=seed)
    dt.fit(X[idx], y[idx])
    stumps.append(dt)
    print(f'seed={seed} OOB={dt.score(X[oob],y[oob]):.4f}')
preds = np.stack([s.predict(X) for s in stumps], axis=1)
votes = np.apply_along_axis(lambda r: np.bincount(r,minlength=3).argmax(),1,preds)
print('ensemble acc:', (votes==y).mean().round(4))
tree (seed)OOB accuracyin-bag size
420.9492150
70.9574150
130.9322150
3-stump ensemble0.9600—

21. Problem 2 — GMM E-step from scratch

Worked example

Theory sub-q: in the E-step, what do the responsibilities rᵢₖ represent, and why do they sum to 1 across k? Then: compute responsibilities for x ∈ {−3, −1, 0, 2, 4} given μ₁=−2, μ₂=3, σ=1, π₁=π₂=0.5.

Hint: rᵢₖ = πₖ·N(xᵢ|μₖ,σ²) / Σⱼ πⱼ·N(xᵢ|μⱼ,σ²). The denominator is the marginal p(xᵢ); dividing by it normalizes over components.

import numpy as np
from scipy.stats import norm
mus=[-2.,3.]; sig=1.; pis=[.5,.5]
X=[-3.,-1.,0.,2.,4.]
for xi in X:
    p=[pis[k]*norm.pdf(xi,mus[k],sig) for k in range(2)]
    s=sum(p)
    r=[round(pk/s,4) for pk in p]
    print(f'x={xi:4}: r0={r[0]}, r1={r[1]}')
xr(k=0)r(k=1)assigned
-3.01.00000.0000k=0
-1.00.99940.0006k=0
0.00.92410.0759k=0
2.00.00060.9994k=1
4.00.00001.0000k=1

22. Which is which, by assigned

Discrimination

Sort into buckets

Sort these by assigned, from memory, without looking back at Problem 2 — GMM E-step from scratch. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

k=0
-3.0; -1.0; 0.0
k=1
2.0; 4.0
g1
assigned is "k=0" for -3.0, -1.0, 0.0 — that is what the table on "Problem 2 — GMM E-step from scratch" records, and it is the single property separating this group from the rest.
g2
assigned is "k=1" for 2.0, 4.0 — that is what the table on "Problem 2 — GMM E-step from scratch" records, and it is the single property separating this group from the rest.

23. Problem 3 — CNN training loop from scratch

Worked example

Theory sub-q: what is the output shape after Conv2d(1, 16, kernel_size=3, padding=1) on a (N, 1, 8, 8) input? Then: build a SmallCNN (1 conv + 1 linear) on digits, train 20 epochs with Adam, print loss trajectory and accuracy.

Hint: shape formula H_out = (H_in + 2p − k)/s + 1; with p=1, k=3, s=1 → 8. Conv params = out_ch × in_ch × k × k + out_ch = 16×1×3×3 + 16 = 160. Train loop: zero_grad → forward → CrossEntropyLoss → backward → step.

import torch, torch.nn as nn
from sklearn.datasets import load_digits
import numpy as np
X,y=load_digits(return_X_y=True)
Xt=torch.tensor(X.reshape(-1,1,8,8)/16.,dtype=torch.float32)
yt=torch.tensor(y)
torch.manual_seed(42)
class CNN(nn.Module):
    def __init__(self):
        super().__init__()
        self.conv=nn.Conv2d(1,16,3,padding=1)
        self.fc=nn.Linear(16*8*8,10)
    def forward(self,x):
        return self.fc(torch.relu(self.conv(x)).view(x.size(0),-1))
model=CNN(); crit=nn.CrossEntropyLoss()
opt=torch.optim.Adam(model.parameters(),lr=5e-3)
for ep in range(20):
    opt.zero_grad(); loss=crit(model(Xt),yt)
    loss.backward(); opt.step()
    if ep in [0,4,9,19]: print(f'ep {ep+1} loss={loss.item():.4f}')
acc=(model(Xt).argmax(1)==yt).float().mean()
print(f'train acc={acc:.4f}')
epochlossnote
12.2961near random (log 10 ≈ 2.30)
51.6783falling steadily
100.9558below 1.0
200.3140converged; train acc = 0.9349

24. Fill in: note for Problem 3 — CNN training loop from scratch

Comparison

Comparison matrix

From Problem 3 — CNN training loop from scratch: refill the note column from what you know. The rest of the table is as it appeared.

epochlossnote
12.2961near random (log 10 ≈ 2.30)
51.6783falling steadily
100.9558below 1.0
200.3140converged; train acc = 0.9349

25. Something is wrong here: responsibilities in the GMM E-step don't sum to 1

Anomaly

Predict first

A student writes this, and it looks reasonable:

Compute rᵢₖ = πₖ · N(xᵢ | μₖ, σ²) and use it directly — it's already the 'soft assignment'.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: The unnormalized product is just a weighted density, not a probability.

Always divide by the sum over all K components.

Why: The unnormalized product is just a weighted density, not a probability. Without the denominator it may sum to 0.98 or 1.3 — the M-step will then compute wrong means and the algorithm diverges.

26. Trap: responsibilities in the GMM E-step don't sum to 1

Trap

The trap

Compute rᵢₖ = πₖ · N(xᵢ | μₖ, σ²) and use it directly — it's already the 'soft assignment'.

Skip normalizing by Σⱼ πⱼ·N(xᵢ|μⱼ,σ²)

Why: The unnormalized product is just a weighted density, not a probability. Without the denominator it may sum to 0.98 or 1.3 — the M-step will then compute wrong means and the algorithm diverges.

The fix

Always divide by the sum over all K components.

rᵢₖ = πₖ·N(xᵢ|μₖ,σ²) / Σⱼ πⱼ·N(xᵢ|μⱼ,σ²)

Why: This is Bayes' theorem: the denominator is p(xᵢ), marginalizing out the component. Each row (over k) sums exactly to 1 — which the test verifies by printing sum(r).

27. Break it on purpose: responsibilities in the GMM E-step don't sum…

Break the constraint

Discussion prompt

The rule this trap just fixed:

Always divide by the sum over all K components.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

The unnormalized product is just a weighted density, not a probability. Without the denominator it may sum to 0.98 or 1.3 — the M-step will then compute wrong means and the algorithm diverges.

28. Something is wrong here: Conv2d output shape miscounted

Anomaly

Predict first

A student writes this, and it looks reasonable:

Input is (N, 1, 8, 8). Conv2d(1, 16, kernel_size=3) with NO padding → output is (N, 16, 8, 8).

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Without padding, each spatial dim shrinks: H_out = (8 − 3)/1 + 1 = 6.

Apply H_out = (H_in + 2p − k)/s + 1; add padding=1 to preserve spatial size.

Why: Without padding, each spatial dim shrinks: H_out = (8 − 3)/1 + 1 = 6. The output is (N, 16, 6, 6). Passing 16×6×6=576 to a Linear(16*8*8, 10) crashes with a shape mismatch.

29. Trap: Conv2d output shape miscounted

Trap

The trap

Input is (N, 1, 8, 8). Conv2d(1, 16, kernel_size=3) with NO padding → output is (N, 16, 8, 8).

Write H_out = 8 without applying the formula

Why: Without padding, each spatial dim shrinks: H_out = (8 − 3)/1 + 1 = 6. The output is (N, 16, 6, 6). Passing 16×6×6=576 to a Linear(16*8*8, 10) crashes with a shape mismatch.

The fix

Apply H_out = (H_in + 2p − k)/s + 1; add padding=1 to preserve spatial size.

Conv2d(1, 16, 3, padding=1) → H_out = (8 + 2 − 3)/1 + 1 = 8; output (N, 16, 8, 8)

Why: Padding=1 pads one zero on each edge, preserving the 8×8 spatial size. The fc layer then receives 16×8×8 = 1024 features — as coded.

30. Which of these survive contact with Lesson 75: Mock Exam — Phase 2 Full Review?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
USAAIO Phase 2 exam: a 2-hour theory block then a 2-hour coding block in Colab. Theory answers guide the code — always do theory first.; Normalization prevents gradient-magnitude differences from dominating optimization. Iris sepal-length: mean = 5.84, std = 0.83, range [4.3, 7.9].; SGD update (L40): w ← w − η·∇L. Adam (L56) uses adaptive per-parameter rates via first and second moment estimates.
Breaks
A node with 50 class-A and 50 class-B is balanced — Gini = 0 (pure).; Compute rᵢₖ = πₖ · N(xᵢ | μₖ, σ²) and use it directly — it's already the 'soft assignment'.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 75: Mock Exam — Phase 2 Full Review puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

31. Rebuild the recipe: Phase 2 exam recipe

Ranking

Put in order

These are the steps of Phase 2 exam recipe, scrambled. Put them back in order before the next slide shows you.

  1. Ensemble: Gini = 1 − Σpₖ²; bootstrap 63.2% in-bag; OOB on ~36.8%; majority vote
  2. GMM E-step: rᵢₖ = πₖN(x|μₖ,σₖ) / Σⱼ; normalize per sample (sums to 1)
  3. CNN loop: Conv(in, out, k, pad) → ReLU → flatten → Linear; zero_grad → loss → backward → step
  4. Normalization: StandardScaler for GD-based models; MinMaxScaler for distance models
  5. Evaluation: Precision/Recall/F1 from confusion matrix; AUC for ranking quality

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

32. Phase 2 exam recipe

Pattern

  1. Ensemble: Gini = 1 − Σpₖ²; bootstrap 63.2% in-bag; OOB on ~36.8%; majority vote
  2. GMM E-step: rᵢₖ = πₖN(x|μₖ,σₖ) / Σⱼ; normalize per sample (sums to 1)
  3. CNN loop: Conv(in, out, k, pad) → ReLU → flatten → Linear; zero_grad → loss → backward → step
  4. Normalization: StandardScaler for GD-based models; MinMaxScaler for distance models
  5. Evaluation: Precision/Recall/F1 from confusion matrix; AUC for ranking quality

33. Where does it stop working: Phase 2 exam recipe

Edge cases

Discussion prompt

Phase 2 exam recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Ensemble: Gini = 1 − Σpₖ²; bootstrap 63.2% in-bag; OOB on ~36.8%; majority vote
  2. GMM E-step: rᵢₖ = πₖN(x|μₖ,σₖ) / Σⱼ; normalize per sample (sums to 1)
  3. CNN loop: Conv(in, out, k, pad) → ReLU → flatten → Linear; zero_grad → loss → backward → step
  4. Normalization: StandardScaler for GD-based models; MinMaxScaler for distance models
  5. Evaluation: Precision/Recall/F1 from confusion matrix; AUC for ranking quality

34. Rule out three: Theory check — Gini impurity

Elimination

Eliminate the wrong options

A decision tree node contains 80 class-A and 20 class-B samples (p = [0.8, 0.2]). Its Gini impurity is:

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. 0.32
  • B. 0.50
  • C. 0.16
  • D. 0.80

Survives elimination: A

Why: Gini = 1 − (0.8² + 0.2²) = 1 − (0.64 + 0.04) = 1 − 0.68 = 0.32. Equal split gives max 0.50; pure node gives 0.

35. Theory check — Gini impurity

Check

Evaluate before clicking.

Check your understanding

A decision tree node contains 80 class-A and 20 class-B samples (p = [0.8, 0.2]). Its Gini impurity is:

  • A. 0.32 (correct)
  • B. 0.50
  • C. 0.16
  • D. 0.80

Answer: A

Why: Gini = 1 − (0.8² + 0.2²) = 1 − (0.64 + 0.04) = 1 − 0.68 = 0.32. Equal split gives max 0.50; pure node gives 0.

Why B tempts people
0.50 is the maximum Gini (equal 50/50 split), not the value at [0.8, 0.2].
Why C tempts people
0.16 confuses Gini with one squared probability (0.4² = 0.16) rather than 1 − Σpₖ².
Why D tempts people
0.80 is simply the majority class proportion p₁ — plugging p₁ directly without the formula.

36. Answer it before you see the options: Theory check — Adam vs SGD

Prediction

Predict first

Starting at w = 3.0 with f(w) = w², after ONE Adam step (lr=0.1, β₁=0.9, β₂=0.999, ε=1e-8) the new w is approximately:

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: 2.90

Why: g=6; m̂=6 (bias-corrected after step 1); v̂=36; update = 0.1·6/√36 = 0.1·1 = 0.1; w_new = 3.0 − 0.1 = 2.90. Adam normalizes by the gradient magnitude so the effective step is exactly lr on the first iteration.

37. Theory check — Adam vs SGD

Check

Recall from Lesson 56.

Check your understanding

Starting at w = 3.0 with f(w) = w², after ONE Adam step (lr=0.1, β₁=0.9, β₂=0.999, ε=1e-8) the new w is approximately:

  • A. 2.90 (correct)
  • B. 1.80
  • C. 2.40
  • D. 3.10

Answer: A

Why: g=6; m̂=6 (bias-corrected after step 1); v̂=36; update = 0.1·6/√36 = 0.1·1 = 0.1; w_new = 3.0 − 0.1 = 2.90. Adam normalizes by the gradient magnitude so the effective step is exactly lr on the first iteration.

Why B tempts people
1.80 would require a step size of 1.2 — that's using raw gradient×lr (0.1×6 = 0.6) incorrectly scaled, or 10× the lr.
Why C tempts people
2.40 corresponds to a step of 0.6, which is what plain SGD with lr=0.1 would give (0.1×6=0.6) — mistaking Adam for SGD.
Why D tempts people
3.10 is an increase; gradient descent on a convex f always decreases w toward the minimum.

38. Rule out three: Theory check — GMM E-step

Elimination

Eliminate the wrong options

In the GMM E-step the responsibility r(i,k) is computed for sample xᵢ and component k. What must be true of the responsibilities over all K components for a single sample xᵢ?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Σₖ r(i,k) = 1 for every sample i
  • B. Σᵢ r(i,k) = 1 for every component k
  • C. r(i,k) = 0 or 1 (hard assignment)
  • D. r(i,k) equals the mixing weight πₖ for all i

Survives elimination: A

Why: Each sample's responsibilities sum to 1 across components (they are posterior probabilities P(k|xᵢ)). This is enforced by the normalizing denominator in Bayes' theorem. The M-step uses these as soft weights to re-estimate means, covariances, and mixing proportions.

39. Theory check — GMM E-step

Check

Core EM concept.

Check your understanding

In the GMM E-step the responsibility r(i,k) is computed for sample xᵢ and component k. What must be true of the responsibilities over all K components for a single sample xᵢ?

  • A. Σₖ r(i,k) = 1 for every sample i (correct)
  • B. Σᵢ r(i,k) = 1 for every component k
  • C. r(i,k) = 0 or 1 (hard assignment)
  • D. r(i,k) equals the mixing weight πₖ for all i

Answer: A

Why: Each sample's responsibilities sum to 1 across components (they are posterior probabilities P(k|xᵢ)). This is enforced by the normalizing denominator in Bayes' theorem. The M-step uses these as soft weights to re-estimate means, covariances, and mixing proportions.

Why B tempts people
Summing over samples i (not components k) gives the effective count Nₖ, not 1 — that's a different quantity used in the M-step.
Why C tempts people
Hard assignment (0 or 1) is k-means, not EM/GMM. GMM uses soft assignments to allow ambiguous points to contribute to multiple components.
Why D tempts people
r(i,k) equals πₖ only when p(xᵢ|k) is the same for all k — i.e. when the data gives no information about the component. In general it depends on xᵢ.

40. Gap analysis & Phase 3 prep

Section

Part 4 of 4

41. Error analysis — categorize by topic

Concept

After the exam: for every mistake, assign it to one Phase 2 topic. The distribution tells you exactly what to reinforce before Phase 3.

topic areaLessonscommon error type
Training loop mechanicsL40–L42zero_grad order, shape mismatch
Backprop / MLPL43–L45chain rule application, weight init
CNNsL46–L50output shape formula, padding vs stride
RegularizationL51–L55L1 vs L2, when dropout hurts
OptimizationL56–L60Adam bias correction, LR schedule choice
EnsemblesL61–L65OOB vs validation, Gini formula
Normalization + EvalL66–L70StandardScaler leakage, AUC interpretation
GMM / EML71–L74E-step normalization, convergence criterion

42. Teach it back: Error analysis — categorize by topic

Explain it

Discussion prompt

Explain Error analysis — categorize by topic to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

After the exam: for every mistake, assign it to one Phase 2 topic. The distribution tells you exactly what to reinforce before Phase 3.

43. Milestone: can you implement all three from scratch?

Concept

Phase 2 coding milestone: you can now implement and train — without a reference — random forest, GMM, and a CNN training loop in PyTorch. If any of the three still needed a hint, that topic is the priority before Phase 3.

Verify each implementation against the library (RF vs sklearn RF, GMM E-step vs GaussianMixture, CNN loss trajectory vs known cross-entropy at random ≈ 2.30). Library agreement = coding milestone passed.

44. Phase 3 preview: attention & transformers

Concept

Phase 3 begins immediately after this exam. The pivot: from convolution (local receptive fields) to self-attention (global, content-based). The transformer training loop is the same PyTorch skeleton — only the model changes.

45. Your turn: full answer key

Section

Project

46. Project brief

Concept

Re-implement all three problems cleanly — no hints — then produce verified output matching the answer key. This is the Phase 2 coding milestone submission.

#implementverify against
13-stump RF on irissklearn RF; OOB ≈ 0.94, ensemble acc ≈ 0.96
2GMM E-step (5 points)scipy.stats.norm; r(x=0)=(0.9241, 0.0759)
3CNN on digits, 20 eploss[1]≈2.30, loss[20]≈0.31, acc≈0.93

Build rule: type every line from scratch. No copy-paste from earlier decks. The act of re-typing exposes every shape assumption.

47. By analogy: Project brief

Analogy

Discussion prompt

Explain Project brief by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Re-implement all three problems cleanly — no hints — then produce verified output matching the answer key. This is the Phase 2 coding milestone submission.

48. Milestone 1 — RF with OOB verification

Worked example

Your turn: implement 3-stump bootstrap ensemble on iris. Predict that ensemble accuracy ≥ each individual tree's OOB accuracy.

Hint: rng.choice(n, n, replace=True) for bootstrap; np.setdiff1d for OOB; stack predictions and take np.bincount(...).argmax() majority vote.

import numpy as np
from sklearn.datasets import load_iris
from sklearn.tree import DecisionTreeClassifier
X,y=load_iris(return_X_y=True); n=len(X)
stumps=[]
for seed in [42,7,13]:
    rng=np.random.default_rng(seed)
    idx=rng.choice(n,n,replace=True)
    oob=np.setdiff1d(np.arange(n),idx)
    dt=DecisionTreeClassifier(max_depth=2,random_state=seed)
    dt.fit(X[idx],y[idx]); stumps.append(dt)
    print(f'seed={seed} OOB={dt.score(X[oob],y[oob]):.4f}')
preds=np.stack([s.predict(X) for s in stumps],axis=1)
votes=np.apply_along_axis(lambda r:np.bincount(r,3).argmax(),1,preds)
print('ensemble acc:',(votes==y).mean().round(4))
treeOOB acc
seed=420.9492
seed=70.9574
seed=130.9322
ensemble0.9600

49. Watch it run: Milestone 1 — RF with OOB verification

Pattern

Step through it

Step through Milestone 1 — RF with OOB verification one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: tree is seed=42
  2. Step 2: tree is seed=7
  3. Step 3: tree is seed=13
  4. Step 4: tree is ensemble

50. Milestone 2 — GMM E-step verification

Worked example

Your turn: compute rᵢₖ for x ∈ {−3,−1,0,2,4} with μ₁=−2, μ₂=3, σ=1, π₁=π₂=0.5. Predict that x=0 stays with k=0 (closer to −2 than to 3).

Hint: import scipy.stats.norm; for each x compute numerator list, divide by sum. Verify sum(r) == 1.0 for every row.

from scipy.stats import norm
mus=[-2.,3.]; sig=1.; pis=[.5,.5]
for xi in [-3.,-1.,0.,2.,4.]:
    p=[pis[k]*norm.pdf(xi,mus[k],sig) for k in range(2)]
    s=sum(p)
    r=[round(pk/s,4) for pk in p]
    print(f'x={xi:5}: r0={r[0]}, r1={r[1]}, sum={sum(r)}')
xr0r1sum
-3.01.00000.00001.0
-1.00.99940.00061.0
0.00.92410.07591.0
2.00.00060.99941.0
4.00.00001.00001.0

51. What each one costs: Milestone 2 — GMM E-step verification

Trade off

Comparison matrix

From Milestone 2 — GMM E-step verification: every row here is a choice with a cost. Fill the r0 column, then say which row you would actually pick and what you give up for it.

xr0r1sum
-3.01.00000.00001.0
-1.00.99940.00061.0
0.00.92410.07591.0
2.00.00060.99941.0
4.00.00001.00001.0

52. Milestone 3 — CNN answer key

Worked example

Your turn: build DigitCNN (Conv2d(1,16,3,pad=1) → ReLU → Linear(1024,10)), train 20 epochs with Adam(lr=5e-3), print losses at epochs 1/5/10/20 and final train accuracy.

Hint: reshape digits to (N,1,8,8) and divide by 16.0 to normalize; loss at epoch 1 should be near log(10) ≈ 2.30 (near-random initialization); after 20 epochs train acc should be ≥ 0.90.

import torch, torch.nn as nn
from sklearn.datasets import load_digits
X,y=load_digits(return_X_y=True)
Xt=torch.tensor(X.reshape(-1,1,8,8)/16.,dtype=torch.float32)
yt=torch.tensor(y)
torch.manual_seed(42)
class CNN(nn.Module):
    def __init__(self):
        super().__init__()
        self.conv=nn.Conv2d(1,16,3,padding=1)
        self.fc=nn.Linear(16*8*8,10)
    def forward(self,x):
        return self.fc(torch.relu(self.conv(x)).view(x.size(0),-1))
model=CNN(); crit=nn.CrossEntropyLoss()
opt=torch.optim.Adam(model.parameters(),lr=5e-3)
for ep in range(20):
    opt.zero_grad(); loss=crit(model(Xt),yt)
    loss.backward(); opt.step()
    if ep in [0,4,9,19]: print(f'ep{ep+1} loss={loss.item():.4f}')
print('acc=',(model(Xt).argmax(1)==yt).float().mean().item())
epochloss
12.2961
51.6783
100.9558
200.3140
train acc0.9349

53. Fill in: loss for Milestone 3 — CNN answer key

Comparison

Comparison matrix

From Milestone 3 — CNN answer key: refill the loss column from what you know. The rest of the table is as it appeared.

epochloss
12.2961
51.6783
100.9558
200.3140
train acc0.9349

54. Show it off

Concept

Slides closed, out loud: (1) explain why bagging reduces variance but not bias; (2) walk the GMM E-step for x=0 to r₀=0.92 step by step; (3) state the five steps of the PyTorch training loop and what breaks if you skip zero_grad.

Stretch: add gradient boosting to Problem 1 and compare with RF; run more GMM EM iterations manually to watch μ converge; add BatchNorm between the CNN conv and relu and compare convergence.

55. Break it if you can: Show it off

Counterexample

Discussion prompt

Stretch: add gradient boosting to Problem 1 and compare with RF; run more GMM EM iterations manually to watch μ converge; add BatchNorm between the CNN conv and relu and compare convergence.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

56. Connect it up: Lesson 75: Mock Exam — Phase 2 Full Review

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Exam format & strategy · Theory block: four topics · Coding block: three problems · Gap analysis & Phase 3 prep · Your turn: full answer key. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

57. Phase 2 milestone — what you can do now

Recap

algorithmthe one gotcha
Random forestGini([0.5,0.5]) = 0.50, NOT 0 — max impurity
GMM E-stepalways divide by the marginal: Σⱼ πⱼN(x|μⱼ,σⱼ)
CNN shapesH_out = (H + 2p − k)/s + 1; padding=1 preserves 8×8
Adam step 1m̂/√v̂ = g/|g| = ±1 → effective lr = lr exactly
PyTorch loopzero_grad FIRST — accumulated grads diverge the model

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 75 (Week 26 — Phase 2 Review Week) — Barron · USAAIO Round 2 Preparation, 2026
  2. RF bootstrap, GMM E-step, CNN training loop, Adam trace, evaluation metrics verified — torch 2.7.1 + numpy 2.2.6 + scikit-learn, real execution, June 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108