Lesson 78: Mock Exam — Phase 2 (Theory + Coding)

USAAIO Lesson 78, from Week 26: a three-hour combined Phase 2 mock exam. It has 30 theory questions spanning every Phase 2 topic, plus two timed coding problems - logistic regression from scratch, and a CNN in PyTorch - each with a theory sub-question and a verified answer key. It is the Phase 2 coding and theory milestone. The lesson runs to 27 slides.

Subject: Machine Learning · 49 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Mock Exam Phase 2: Classical ML + Deep Learning

Title

USAAIO · Lesson 78 · Week 26

3-hour combined exam: 1.5 hr theory (30 questions, all Phase 2 topics) + 1.5 hr coding (2 problems — classical ML and deep learning). Timed, then full debrief.

2. Exam goals

Objectives

  1. Score >75% on the 30-question theory section (≥ 23 correct)
  2. Implement logistic regression from scratch in NumPy and verify against sklearn
  3. Implement and train a CNN in PyTorch on load_digits within time
  4. Pass a code review: correct gradients, model.eval() at inference, no data-leakage
  5. Debrief every error and close all Phase 2 gaps before Phase 3 (Transformers)

3. What survived from Deep Learning Foundations?

Warm-up

Discussion prompt

Before we open Lesson 78: Mock Exam — Phase 2 (Theory + Coding): without looking back, what was the main idea of Deep Learning Foundations, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

MLP, CNN, and ResNet architectures; ReLU/sigmoid/tanh activations and vanishing-gradient mechanics; BatchNorm, Dropout, and Xavier/Kaiming initialization; the PyTorch training loop line-by-line; backpropagation via chain rule with a concrete 2-layer trace; SGD/Momentum/Adam/AdamW optimizers and LR scheduling; and Dataset/DataLoader for mini-batch pipelines. Build a complete digits classifier with every technique applied.

4. Exam format & rules

Section

Part 1 of 4

5. Structure of the 3-hour exam

Concept

blocktimecontent
Theory1.5 hr30 MCQs — all Phase 2 topics (L40–L77)
Coding Problem 145 minClassical ML from scratch (NumPy)
Coding Problem 245 minDeep learning in PyTorch

6. Fill in: time for Structure of the 3-hour exam

Comparison

Comparison matrix

From Structure of the 3-hour exam: refill the time column from what you know. The rest of the table is as it appeared.

blocktimecontent
Theory1.5 hr30 MCQs — all Phase 2 topics (L40–L77)
Coding Problem 145 minClassical ML from scratch (NumPy)
Coding Problem 245 minDeep learning in PyTorch

7. Rebuild the recipe: Per-problem exam protocol

Ranking

Put in order

These are the steps of Per-problem exam protocol, scrambled. Put them back in order before the next slide shows you.

  1. Answer the theory sub-question first — it forces you to recall the math before coding
  2. Predict the output on paper before running
  3. Implement from scratch, vectorized where possible
  4. Verify against the library or a known value
  5. Review: correct gradient formula, shapes, train/eval modes, leakage

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

8. Per-problem exam protocol

Pattern

  1. Answer the theory sub-question first — it forces you to recall the math before coding
  2. Predict the output on paper before running
  3. Implement from scratch, vectorized where possible
  4. Verify against the library or a known value
  5. Review: correct gradient formula, shapes, train/eval modes, leakage

9. Where does it stop working: Per-problem exam protocol

Edge cases

Discussion prompt

Per-problem exam protocol works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Answer the theory sub-question first — it forces you to recall the math before coding
  2. Predict the output on paper before running
  3. Implement from scratch, vectorized where possible
  4. Verify against the library or a known value
  5. Review: correct gradient formula, shapes, train/eval modes, leakage

10. Theory section — 30 questions

Section

Part 2 of 4

11. Phase 2 topic distribution (30 questions)

Concept

topic areaapprox Qslessons
Linear & logistic regression4L40–L43
SVMs, kernels, decision trees4L44–L49
Ensemble methods (RF, XGBoost)3L50–L52
Unsupervised (k-means, DBSCAN, PCA)4L53–L57
Neural networks & backprop5L58–L63
CNNs (conv, pooling, ResNet)4L64–L68
Regularization, optimization, batch norm4L69–L72
Hyperparameter tuning & pipelines2L73–L77

Milestone check: >75% = ≥ 23 correct. Use the distribution above to spot your weak area — every column with ≤ 50% is a Phase 3 liability.

12. What each one costs: Phase 2 topic distribution (30 questions)

Trade off

Comparison matrix

From Phase 2 topic distribution (30 questions): every row here is a choice with a cost. Fill the approx Qs column, then say which row you would actually pick and what you give up for it.

topic areaapprox Qslessons
Linear & logistic regression4L40–L43
SVMs, kernels, decision trees4L44–L49
Ensemble methods (RF, XGBoost)3L50–L52
Unsupervised (k-means, DBSCAN, PCA)4L53–L57
Neural networks & backprop5L58–L63
CNNs (conv, pooling, ResNet)4L64–L68
Regularization, optimization, batch norm4L69–L72
Hyperparameter tuning & pipelines2L73–L77

13. Rule out three: Theory warm-up — sigmoid gradient

Elimination

Eliminate the wrong options

The derivative of σ(z) = 1/(1+e^{−z}) with respect to z equals:

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. σ(z)(1 − σ(z))
  • B. σ(z)²
  • C. 1 − σ(z)
  • D. σ(z) / (1 − σ(z))

Survives elimination: A

Why: σ'(z) = σ(z)(1−σ(z)). At z=0: σ(0)=0.5 so σ'(0)=0.25 — the maximum slope. This identity is used directly in the logistic regression gradient and in sigmoid-activated neural networks (Lesson 41).

14. Theory warm-up — sigmoid gradient

Check

Foundation for logistic regression's update rule (Lesson 41).

Check your understanding

The derivative of σ(z) = 1/(1+e^{−z}) with respect to z equals:

  • A. σ(z)(1 − σ(z)) (correct)
  • B. σ(z)²
  • C. 1 − σ(z)
  • D. σ(z) / (1 − σ(z))

Answer: A

Why: σ'(z) = σ(z)(1−σ(z)). At z=0: σ(0)=0.5 so σ'(0)=0.25 — the maximum slope. This identity is used directly in the logistic regression gradient and in sigmoid-activated neural networks (Lesson 41).

Why B tempts people
σ(z)² omits the (1−σ) factor — it's the square of the output, not its derivative.
Why C tempts people
1−σ(z) is the probability of the negative class, not the derivative.
Why D tempts people
σ(z)/(1−σ(z)) is the odds ratio, a separate concept in logistic regression, not the sigmoid's derivative.

15. Answer it before you see the options: Theory warm-up — hinge loss

Prediction

Predict first

For a correctly classified SVM point with decision score +1.5 (true label +1), the hinge loss is:

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: 0.0

Why: Hinge loss = max(0, 1 − y·score) = max(0, 1 − 1×1.5) = max(0, −0.5) = 0. The point is on the correct side with margin > 1, so it incurs no loss.

16. Theory warm-up — hinge loss

Check

Phase 2 SVM theory (Lesson 44).

Check your understanding

For a correctly classified SVM point with decision score +1.5 (true label +1), the hinge loss is:

  • A. 0.0 (correct)
  • B. 0.5
  • C. 1.5
  • D. 2.5

Answer: A

Why: Hinge loss = max(0, 1 − y·score) = max(0, 1 − 1×1.5) = max(0, −0.5) = 0. The point is on the correct side with margin > 1, so it incurs no loss.

Why B tempts people
0.5 would result from score = 0.5 (on the margin boundary), not 1.5.
Why C tempts people
1.5 would be the loss if the score were −0.5 (wrong side, magnitude 1.5), not +1.5.
Why D tempts people
2.5 would come from score = −1.5, a strongly misclassified point.

17. Coding problem 1 — Logistic Regression

Section

Part 3 of 4

18. Problem 1 — logistic regression from scratch

Worked example

Theory sub-q: write the binary cross-entropy gradient ∂L/∂w. Then: implement fit() and predict() using only NumPy on iris (setosa vs versicolor).

Hint: gradient = X.T @ (σ(Xw+b) − y) / n. The answer at z=0 is σ=0.5, so loss epoch 0 ≈ log(2) ≈ 0.6931. Predict train accuracy after 500 epochs.

import numpy as np
from sklearn.datasets import load_iris
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split

iris = load_iris()
X = StandardScaler().fit_transform(iris.data[:100])
y = iris.target[:100].astype(float)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, random_state=42)

def sigmoid(z): return 1 / (1 + np.exp(-z))

w, b = np.zeros(4), 0.0
for _ in range(500):
    p = sigmoid(X_tr @ w + b)
    w -= 0.5 * X_tr.T @ (p - y_tr) / len(X_tr)
    b -= 0.5 * np.mean(p - y_tr)

acc = ((sigmoid(X_te @ w + b) >= 0.5) == y_te).mean()
print(round(acc, 4))
metricvalue
loss epoch 00.6931 (= ln 2)
loss epoch 4990.0027
train accuracy1.0000
test accuracy1.0000
sklearn match1.0000

19. Which is which, by value

Discrimination

Sort into buckets

Sort these by value, from memory, without looking back at Problem 1 — logistic regression from scratch. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

0.6931 (= ln 2)
loss epoch 0
0.0027
loss epoch 499
1.0000
train accuracy; test accuracy; sklearn match
g1
value is "0.6931 (= ln 2)" for loss epoch 0 — that is what the table on "Problem 1 — logistic regression from…" records, and it is the single property separating this group from the rest.
g2
value is "0.0027" for loss epoch 499 — that is what the table on "Problem 1 — logistic regression from…" records, and it is the single property separating this group from the rest.
g3
value is "1.0000" for train accuracy, test accuracy, sklearn match — that is what the table on "Problem 1 — logistic regression from…" records, and it is the single property separating this group from the rest.

20. Something is wrong here: dividing gradient by wrong denominator

Anomaly

Predict first

A student writes this, and it looks reasonable:

The gradient is X.T @ (p − y) — no division, just sum over the batch.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Summing (not averaging) the gradient makes the effective learning rate scale with dataset size n.

Divide by n to average: gradient = X.T @ (p − y) / n.

Why: Summing (not averaging) the gradient makes the effective learning rate scale with dataset size n. On 80 training points lr=0.5 becomes effectively lr=40 — the loss explodes or oscillates instead of converging.

21. Trap: dividing gradient by wrong denominator

Trap

The trap

The gradient is X.T @ (p − y) — no division, just sum over the batch.

Update: w -= lr * X.T @ (p - y)

Why: Summing (not averaging) the gradient makes the effective learning rate scale with dataset size n. On 80 training points lr=0.5 becomes effectively lr=40 — the loss explodes or oscillates instead of converging.

The fix

Divide by n to average: gradient = X.T @ (p − y) / n.

Update: w -= lr * X.T @ (p - y) / len(X_tr)

Why: Averaging makes the update scale-independent of n, so the same lr=0.5 works on 80 or 8,000 samples. This is the source of the standard formula and is checked explicitly on the rubric.

22. Break it on purpose: dividing gradient by wrong denominator

Break the constraint

Discussion prompt

The rule this trap just fixed:

Averaging makes the update scale-independent of n, so the same lr=0.5 works on 80 or 8,000 samples. This is the source of the standard formula and is checked explicitly on the rubric.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

Summing (not averaging) the gradient makes the effective learning rate scale with dataset size n. On 80 training points lr=0.5 becomes effectively lr=40 — the loss explodes or oscillates instead of converging.

23. Gradient descent trace — 3 epochs

Concept

epochloss|dw|db
00.69310.8327-0.0250
10.40840.5412-0.0173
20.28340.3876-0.0142

Both |dw| and |db| shrink each step — the model is converging. Loss dropping from 0.6931 → 0.4084 (−41%) in one epoch is typical for well-scaled features (iris is z-scored).

24. Watch it run: Gradient descent trace — 3 epochs

Pattern

Step through it

Step through Gradient descent trace — 3 epochs one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: epoch is 0
  2. Step 2: epoch is 1
  3. Step 3: epoch is 2

25. Coding problem 2 — CNN in PyTorch

Section

Part 4 of 4

26. Problem 2 — CNN on load_digits

Worked example

Theory sub-q: what is the receptive field of two stacked 3×3 convolutions? Then: build a two-conv + pool + two-FC CNN, train for 15 epochs with Adam, and report test accuracy.

Hint: receptive field after two 3×3 convs (stride 1) = 5×5. Reshape digits (N,64) to (N,1,8,8). Predict: accuracy should exceed 0.97 after 15 epochs.

import torch, torch.nn as nn
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
from torch.utils.data import TensorDataset, DataLoader

digits = load_digits()
X = digits.data.astype('float32') / 16.0
y = digits.target
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.2, random_state=42)
Xtr_t = torch.tensor(Xtr.reshape(-1,1,8,8))
Xte_t = torch.tensor(Xte.reshape(-1,1,8,8))
ytr_t = torch.tensor(ytr, dtype=torch.long)
yte_t = torch.tensor(yte, dtype=torch.long)

class CNN(nn.Module):
    def __init__(self):
        super().__init__()
        self.net = nn.Sequential(
            nn.Conv2d(1,16,3,padding=1), nn.ReLU(),
            nn.Conv2d(16,32,3,padding=1), nn.ReLU(),
            nn.MaxPool2d(2), nn.Flatten(),
            nn.Linear(32*4*4,64), nn.ReLU(),
            nn.Linear(64,10))
    def forward(self,x): return self.net(x)

torch.manual_seed(0)
model = CNN()
opt = torch.optim.Adam(model.parameters(), lr=1e-3)
loss_fn = nn.CrossEntropyLoss()
for _ in range(15):
    model.train()
    for xb,yb in DataLoader(TensorDataset(Xtr_t,ytr_t),32,shuffle=True):
        opt.zero_grad(); loss_fn(model(xb),yb).backward(); opt.step()
model.eval()
with torch.no_grad():
    acc = (model(Xte_t).argmax(1)==yte_t).float().mean().item()
print(round(acc, 4))
epochtrain losstrain acc
12.16950.5915
21.07480.8608
30.40260.9283
40.25520.9478
50.18230.9576

27. Watch it run: Problem 2 — CNN on load_digits

Pattern

Step through it

Step through Problem 2 — CNN on load_digits one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: epoch is 1
  2. Step 2: epoch is 2
  3. Step 3: epoch is 3
  4. Step 4: epoch is 4
  5. Step 5: epoch is 5

28. CNN architecture: receptive field & param count

Concept

layeroutput shapeparams
Conv2d(1,16,3,p=1) + ReLU(N,16,8,8)160
Conv2d(16,32,3,p=1) + ReLU(N,32,8,8)4640
MaxPool2d(2) + Flatten(N,512)0
Linear(512,64) + ReLU(N,64)32832
Linear(64,10)(N,10)650

Total: 38,282 parameters. Receptive field of two stacked 3×3 convolutions (stride 1): 5×5. After the pooling layer, spatial resolution halves from 8×8 to 4×4.

29. Break it if you can: CNN architecture: receptive field & param count

Counterexample

Discussion prompt

Total: 38,282 parameters. Receptive field of two stacked 3×3 convolutions (stride 1): 5×5. After the pooling layer, spatial resolution halves from 8×8 to 4×4.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

30. Something is wrong here: skipping model.eval() at inference

Anomaly

Predict first

A student writes this, and it looks reasonable:

After training, just run model(X_test) — the model was trained, so it's ready.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Dropout randomly zeroes activations at every forward pass.

Always call model.eval() (and torch.no_grad()) before any inference.

Why: Dropout randomly zeroes activations at every forward pass. In train mode, two calls on the SAME input produce DIFFERENT logits. The test accuracy appears to jitter and is systematically too low — an exam rubric failure.

31. Trap: skipping model.eval() at inference

Trap

The trap

After training, just run model(X_test) — the model was trained, so it's ready.

Evaluate while model is still in train mode

Why: Dropout randomly zeroes activations at every forward pass. In train mode, two calls on the SAME input produce DIFFERENT logits. The test accuracy appears to jitter and is systematically too low — an exam rubric failure.

The fix

Always call model.eval() (and torch.no_grad()) before any inference.

model.eval() → torch.no_grad() → model(X_test)

Why: eval() disables dropout (and fixes BatchNorm stats). Identical inputs now produce identical outputs. The verified CNN test accuracy is 0.9806 — consistent across runs.

32. Which of these survive contact with Lesson 78: Mock Exam — Phase 2 (Theory +…?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Total: 38,282 parameters. Receptive field of two stacked 3×3 convolutions (stride 1): 5×5. After the pooling layer, spatial resolution halves from 8×8 to 4×4.; Work each problem to completion before looking at the answer key. Time yourself: ≤ 45 min per coding problem.
Breaks
The gradient is X.T @ (p − y) — no division, just sum over the batch.; After training, just run model(X_test) — the model was trained, so it's ready.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 78: Mock Exam — Phase 2 (Theory + Coding) puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

33. Full answer key — both coding problems

Concept

# KEY VALUES — all verified by real execution
# Problem 1: Logistic Regression from scratch on iris (binary)
#   loss epoch 0   : 0.6931   (= ln 2, since p=0.5 everywhere)
#   loss epoch 499 : 0.0027   (near-zero: linearly separable)
#   test accuracy  : 1.0000   (matches sklearn LogReg)
#   gradient formula: dL/dw = X.T @ (sigmoid(Xw+b) - y) / n
#
# Problem 2: CNN (2 conv + pool + 2 FC) on load_digits, 15 epochs Adam
#   epoch 1 loss   : 2.1695,  train acc: 0.5915
#   epoch 5 loss   : 0.1823,  train acc: 0.9576
#   test accuracy  : 0.9806
#   total params   : 38,282
#   receptive field (2x 3x3 conv, stride=1): 5x5
print('verified')
problemkey resultverified vs
Logistic Regressiontest acc = 1.0000sklearn LogisticRegression
Cross-entropy gradientX.T @ (p − y) / nanalytic derivation
CNN 15 epochstest acc = 0.9806deterministic (torch.manual_seed(0))
CNN receptive field5×5 after 2 conv(3×3)formula: 1+(k-1)*L
CNN param count38,282 totalmanual layer sum

If either result differs, the most common causes are: (1) no /n in the gradient, (2) no model.eval() before CNN inference, (3) wrong reshape (N,64) instead of (N,1,8,8).

34. Fill in: key result for Full answer key — both coding problems

Comparison

Comparison matrix

From Full answer key — both coding problems: refill the key result column from what you know. The rest of the table is as it appeared.

problemkey resultverified vs
Logistic Regressiontest acc = 1.0000sklearn LogisticRegression
Cross-entropy gradientX.T @ (p − y) / nanalytic derivation
CNN 15 epochstest acc = 0.9806deterministic (torch.manual_seed(0))
CNN receptive field5×5 after 2 conv(3×3)formula: 1+(k-1)*L
CNN param count38,282 totalmanual layer sum

35. Rule out three: Code review — logistic gradient

Elimination

Eliminate the wrong options

A student implements: w -= lr * X.T @ (sigmoid(X@w+b) - y) (no /n). What goes wrong?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. The effective learning rate scales with n, causing instability on large datasets
  • B. The gradient direction is wrong — it should use (y − p)
  • C. The result is unaffected — summing and averaging give the same optimum
  • D. It increases the bias term instead of decreasing it

Survives elimination: A

Why: Without /n the step = lr × (sum of n gradients) instead of lr × (mean). On 80 points the effective lr is 80×; on 8,000 it's 8,000×. The model either oscillates or explodes. Dividing by n keeps the learning rate dataset-size-independent.

36. Code review — logistic gradient

Check

Catch the subtle bug before it costs points.

Check your understanding

A student implements: w -= lr * X.T @ (sigmoid(X@w+b) - y) (no /n). What goes wrong?

  • A. The effective learning rate scales with n, causing instability on large datasets (correct)
  • B. The gradient direction is wrong — it should use (y − p)
  • C. The result is unaffected — summing and averaging give the same optimum
  • D. It increases the bias term instead of decreasing it

Answer: A

Why: Without /n the step = lr × (sum of n gradients) instead of lr × (mean). On 80 points the effective lr is 80×; on 8,000 it's 8,000×. The model either oscillates or explodes. Dividing by n keeps the learning rate dataset-size-independent.

Why B tempts people
The direction is correct: gradient of −log-likelihood w.r.t. w is X.T(p−y)/n, not (y−p). The direction is right, only the magnitude is wrong.
Why C tempts people
Same direction, but magnitude scales with n. On any non-trivial dataset this causes divergence rather than convergence at the same point.
Why D tempts people
The bias update is a scalar mean, so it's affected by the same /n issue — not a directional error.

37. Your turn: build both

Section

Project

38. Project: Phase 2 exam under exam conditions

Concept

Work each problem to completion before looking at the answer key. Time yourself: ≤ 45 min per coding problem.

#taskverification
1Logistic regression from scratch on iris binarytest acc = 1.0 vs sklearn
2Cross-entropy gradient — write it out analytically firstX.T@(p−y)/n
3CNN 2-conv+pool+2-FC on load_digits, 15 epochs Adamtest acc ≥ 0.97
4Apply model.eval() + no_grad before reporting accuracyconsistent result

Build rules: type every line, verify shapes at each layer (x.shape), and don't peek at the answer key until your code runs.

39. Break it if you can: Project: Phase 2 exam under exam conditions

Counterexample

Discussion prompt

Work each problem to completion before looking at the answer key. Time yourself: ≤ 45 min per coding problem.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Build rules: type every line, verify shapes at each layer (x.shape), and don't peek at the answer key until your code runs.

40. Milestone 1 — logistic regression skeleton

Worked example

Your turn: implement sigmoid, the gradient, and the update loop. Predict the epoch-0 loss before running.

Hint: at initialization w=0, b=0, so p=0.5 for every point, and binary cross-entropy = −log(0.5) = log(2) ≈ 0.6931.

import numpy as np
from sklearn.datasets import load_iris
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split

iris = load_iris()
X = StandardScaler().fit_transform(iris.data[:100])
y = iris.target[:100].astype(float)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, random_state=42)

def sigmoid(z): return 1 / (1 + np.exp(-z))

w, b = np.zeros(4), 0.0
for ep in range(3):  # first 3 epochs to see trace
    p = sigmoid(X_tr @ w + b)
    loss = -np.mean(y_tr*np.log(p+1e-12)+(1-y_tr)*np.log(1-p+1e-12))
    w -= 0.5 * X_tr.T @ (p - y_tr) / len(X_tr)
    b -= 0.5 * np.mean(p - y_tr)
    print(f'epoch {ep}: loss={loss:.4f}')
epochloss
00.6931
10.4084
20.2834

41. Watch it run: Milestone 1 — logistic regression skeleton

Pattern

Step through it

Step through Milestone 1 — logistic regression skeleton one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: epoch is 0
  2. Step 2: epoch is 1
  3. Step 3: epoch is 2

42. Milestone 2 — CNN architecture

Worked example

Your turn: build the CNN class. Before coding, write out each layer's output shape on paper.

Hint: Conv2d(1,16,3,padding=1) preserves 8×8. MaxPool2d(2) halves to 4×4. After flatten: 32×4×4 = 512. Then Linear(512,64), Linear(64,10).

import torch, torch.nn as nn

class CNN(nn.Module):
    def __init__(self):
        super().__init__()
        self.net = nn.Sequential(
            nn.Conv2d(1, 16, 3, padding=1), nn.ReLU(),
            nn.Conv2d(16, 32, 3, padding=1), nn.ReLU(),
            nn.MaxPool2d(2), nn.Flatten(),
            nn.Linear(32*4*4, 64), nn.ReLU(),
            nn.Linear(64, 10))
    def forward(self, x): return self.net(x)

# verify architecture
model = CNN()
print(sum(p.numel() for p in model.parameters()))  # 38282
x_probe = torch.zeros(1, 1, 8, 8)
print(model(x_probe).shape)  # torch.Size([1, 10])
checkvalue
total params38282
output shapetorch.Size([1, 10])
receptive field (2 conv 3x3)5x5

43. Milestone 3 — training loop & eval

Worked example

Your turn: write the DataLoader loop (15 epochs, Adam, CrossEntropyLoss) and finish with model.eval() before reporting accuracy.

Hint: reshape digits to (N,1,8,8), normalize by dividing by 16. After 15 epochs test accuracy should exceed 0.97.

from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
from torch.utils.data import TensorDataset, DataLoader

digits = load_digits()
X = digits.data.astype('float32') / 16.0
y = digits.target
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.2, random_state=42)
Xtr_t = torch.tensor(Xtr.reshape(-1,1,8,8))
Xte_t = torch.tensor(Xte.reshape(-1,1,8,8))
ytr_t = torch.tensor(ytr, dtype=torch.long)
yte_t = torch.tensor(yte, dtype=torch.long)

torch.manual_seed(0)
model = CNN()
opt = torch.optim.Adam(model.parameters(), lr=1e-3)
for _ in range(15):
    model.train()
    for xb, yb in DataLoader(TensorDataset(Xtr_t,ytr_t), 32, shuffle=True):
        opt.zero_grad(); nn.CrossEntropyLoss()(model(xb),yb).backward(); opt.step()
model.eval()
with torch.no_grad():
    acc = (model(Xte_t).argmax(1)==yte_t).float().mean().item()
print(round(acc, 4))  # 0.9806
epochtrain losstrain acc
12.16950.5915
30.40260.9283
50.18230.9576
100.0785≈0.98
15 (test)0.04660.9806

44. What each one costs: Milestone 3 — training loop & eval

Trade off

Comparison matrix

From Milestone 3 — training loop & eval: every row here is a choice with a cost. Fill the train acc column, then say which row you would actually pick and what you give up for it.

epochtrain losstrain acc
12.16950.5915
30.40260.9283
50.18230.9576
100.0785≈0.98
15 (test)0.04660.9806

45. The full program — both problems

Concept

# Full Phase 2 exam submission — both problems
import numpy as np, torch, torch.nn as nn
from sklearn.datasets import load_iris, load_digits
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split
from torch.utils.data import TensorDataset, DataLoader

# --- Problem 1: Logistic Regression ---
iris = load_iris()
X = StandardScaler().fit_transform(iris.data[:100])
y = iris.target[:100].astype(float)
X_tr,X_te,y_tr,y_te = train_test_split(X,y,test_size=0.2,random_state=42)
def sigmoid(z): return 1/(1+np.exp(-z))
w,b = np.zeros(4), 0.0
for _ in range(500):
    p = sigmoid(X_tr@w+b)
    w -= 0.5*X_tr.T@(p-y_tr)/len(X_tr); b -= 0.5*np.mean(p-y_tr)
print('LR test acc:', ((sigmoid(X_te@w+b)>=0.5)==y_te).mean())

# --- Problem 2: CNN ---
d = load_digits()
Xtr,Xte,ytr,yte = train_test_split(d.data.astype('f4')/16,d.target,test_size=.2,random_state=42)
Xtr_t,Xte_t = torch.tensor(Xtr.reshape(-1,1,8,8)),torch.tensor(Xte.reshape(-1,1,8,8))
ytr_t,yte_t = torch.tensor(ytr,dtype=torch.long),torch.tensor(yte,dtype=torch.long)
class CNN(nn.Module):
    def __init__(self):
        super().__init__()
        self.net=nn.Sequential(nn.Conv2d(1,16,3,padding=1),nn.ReLU(),nn.Conv2d(16,32,3,padding=1),nn.ReLU(),nn.MaxPool2d(2),nn.Flatten(),nn.Linear(512,64),nn.ReLU(),nn.Linear(64,10))
    def forward(self,x): return self.net(x)
torch.manual_seed(0); model=CNN(); opt=torch.optim.Adam(model.parameters(),lr=1e-3)
for _ in range(15):
    model.train()
    for xb,yb in DataLoader(TensorDataset(Xtr_t,ytr_t),32,shuffle=True):
        opt.zero_grad(); nn.CrossEntropyLoss()(model(xb),yb).backward(); opt.step()
model.eval()
with torch.no_grad(): print('CNN test acc:', round((model(Xte_t).argmax(1)==yte_t).float().mean().item(),4))
output linevalue
LR test acc:1.0
CNN test acc:0.9806

Both results match the answer key exactly. If yours differs, check (1) /len(X_tr) in the gradient, (2) model.eval() before CNN inference, (3) torch.manual_seed(0) before CNN init.

46. Fill in: value for The full program — both problems

Comparison

Comparison matrix

From The full program — both problems: refill the value column from what you know. The rest of the table is as it appeared.

output linevalue
LR test acc:1.0
CNN test acc:0.9806

47. Show it off & Phase 3 readiness

Concept

Out loud, slides closed: (1) derive the logistic regression gradient from the cross-entropy loss, (2) explain why the /n matters for learning-rate stability, (3) describe why model.eval() is required before any CNN inference.

Grade your theory section: identify which topic area had the most misses and revisit those slides before Lesson 79. Phase 3 starts with attention mechanisms — the most heavily weighted USAAIO topic.

48. Connect it up: Lesson 78: Mock Exam — Phase 2 (Theory + Coding)

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Exam format & rules · Theory section — 30 questions · Coding problem 1 — Logistic Regression · Coding problem 2 — CNN in PyTorch · Your turn: build both. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

49. Phase 2 exam debrief

Recap

conceptthe one thing to remember
logistic gradientX.T @ (p − y) / n — divide by n
cross-entropy initloss = ln(2) ≈ 0.6931 when p = 0.5
CNN inferencemodel.eval() + torch.no_grad() always
receptive field2× conv(3×3, stride 1) → 5×5
Phase 2 milestone>75% theory + both coding problems complete

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 78 (Week 26 — Phase 2 Mock Exam) — Barron · USAAIO Round 2 Preparation, 2026
  2. Logistic regression from scratch and CNN on load_digits verified vs sklearn, torch 2.7.1 — numpy 2.2.6 + scikit-learn + torch 2.7.1+cpu, real execution, June 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108