USAAIO Lesson 78, from Week 26: a three-hour combined Phase 2 mock exam. It has 30 theory questions spanning every Phase 2 topic, plus two timed coding problems - logistic regression from scratch, and a CNN in PyTorch - each with a theory sub-question and a verified answer key. It is the Phase 2 coding and theory milestone. The lesson runs to 27 slides.
Subject: Machine Learning · 49 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 78 · Week 26
3-hour combined exam: 1.5 hr theory (30 questions, all Phase 2 topics) + 1.5 hr coding (2 problems — classical ML and deep learning). Timed, then full debrief.
Objectives
model.eval() at inference, no data-leakageWarm-up
Discussion prompt
Before we open Lesson 78: Mock Exam — Phase 2 (Theory + Coding): without looking back, what was the main idea of Deep Learning Foundations, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
MLP, CNN, and ResNet architectures; ReLU/sigmoid/tanh activations and vanishing-gradient mechanics; BatchNorm, Dropout, and Xavier/Kaiming initialization; the PyTorch training loop line-by-line; backpropagation via chain rule with a concrete 2-layer trace; SGD/Momentum/Adam/AdamW optimizers and LR scheduling; and Dataset/DataLoader for mini-batch pipelines. Build a complete digits classifier with every technique applied.
Section
Part 1 of 4
Concept
| block | time | content |
|---|---|---|
| Theory | 1.5 hr | 30 MCQs — all Phase 2 topics (L40–L77) |
| Coding Problem 1 | 45 min | Classical ML from scratch (NumPy) |
| Coding Problem 2 | 45 min | Deep learning in PyTorch |
Comparison
Comparison matrix
From Structure of the 3-hour exam: refill the time column from what you know. The rest of the table is as it appeared.
| block | time | content |
|---|---|---|
| Theory | 1.5 hr | 30 MCQs — all Phase 2 topics (L40–L77) |
| Coding Problem 1 | 45 min | Classical ML from scratch (NumPy) |
| Coding Problem 2 | 45 min | Deep learning in PyTorch |
Ranking
Put in order
These are the steps of Per-problem exam protocol, scrambled. Put them back in order before the next slide shows you.
train/eval modes, leakageWhy: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
train/eval modes, leakageEdge cases
Discussion prompt
Per-problem exam protocol works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
train/eval modes, leakageSection
Part 2 of 4
Concept
| topic area | approx Qs | lessons |
|---|---|---|
| Linear & logistic regression | 4 | L40–L43 |
| SVMs, kernels, decision trees | 4 | L44–L49 |
| Ensemble methods (RF, XGBoost) | 3 | L50–L52 |
| Unsupervised (k-means, DBSCAN, PCA) | 4 | L53–L57 |
| Neural networks & backprop | 5 | L58–L63 |
| CNNs (conv, pooling, ResNet) | 4 | L64–L68 |
| Regularization, optimization, batch norm | 4 | L69–L72 |
| Hyperparameter tuning & pipelines | 2 | L73–L77 |
Milestone check: >75% = ≥ 23 correct. Use the distribution above to spot your weak area — every column with ≤ 50% is a Phase 3 liability.
Trade off
Comparison matrix
From Phase 2 topic distribution (30 questions): every row here is a choice with a cost. Fill the approx Qs column, then say which row you would actually pick and what you give up for it.
| topic area | approx Qs | lessons |
|---|---|---|
| Linear & logistic regression | 4 | L40–L43 |
| SVMs, kernels, decision trees | 4 | L44–L49 |
| Ensemble methods (RF, XGBoost) | 3 | L50–L52 |
| Unsupervised (k-means, DBSCAN, PCA) | 4 | L53–L57 |
| Neural networks & backprop | 5 | L58–L63 |
| CNNs (conv, pooling, ResNet) | 4 | L64–L68 |
| Regularization, optimization, batch norm | 4 | L69–L72 |
| Hyperparameter tuning & pipelines | 2 | L73–L77 |
Elimination
Eliminate the wrong options
The derivative of σ(z) = 1/(1+e^{−z}) with respect to z equals:
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: σ'(z) = σ(z)(1−σ(z)). At z=0: σ(0)=0.5 so σ'(0)=0.25 — the maximum slope. This identity is used directly in the logistic regression gradient and in sigmoid-activated neural networks (Lesson 41).
Check
Foundation for logistic regression's update rule (Lesson 41).
Check your understanding
The derivative of σ(z) = 1/(1+e^{−z}) with respect to z equals:
Answer: A
Why: σ'(z) = σ(z)(1−σ(z)). At z=0: σ(0)=0.5 so σ'(0)=0.25 — the maximum slope. This identity is used directly in the logistic regression gradient and in sigmoid-activated neural networks (Lesson 41).
Prediction
Predict first
For a correctly classified SVM point with decision score +1.5 (true label +1), the hinge loss is:
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: 0.0
Why: Hinge loss = max(0, 1 − y·score) = max(0, 1 − 1×1.5) = max(0, −0.5) = 0. The point is on the correct side with margin > 1, so it incurs no loss.
Check
Phase 2 SVM theory (Lesson 44).
Check your understanding
For a correctly classified SVM point with decision score +1.5 (true label +1), the hinge loss is:
Answer: A
Why: Hinge loss = max(0, 1 − y·score) = max(0, 1 − 1×1.5) = max(0, −0.5) = 0. The point is on the correct side with margin > 1, so it incurs no loss.
Section
Part 3 of 4
Worked example
Theory sub-q: write the binary cross-entropy gradient ∂L/∂w. Then: implement fit() and predict() using only NumPy on iris (setosa vs versicolor).
Hint: gradient = X.T @ (σ(Xw+b) − y) / n. The answer at z=0 is σ=0.5, so loss epoch 0 ≈ log(2) ≈ 0.6931. Predict train accuracy after 500 epochs.
import numpy as np
from sklearn.datasets import load_iris
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split
iris = load_iris()
X = StandardScaler().fit_transform(iris.data[:100])
y = iris.target[:100].astype(float)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, random_state=42)
def sigmoid(z): return 1 / (1 + np.exp(-z))
w, b = np.zeros(4), 0.0
for _ in range(500):
p = sigmoid(X_tr @ w + b)
w -= 0.5 * X_tr.T @ (p - y_tr) / len(X_tr)
b -= 0.5 * np.mean(p - y_tr)
acc = ((sigmoid(X_te @ w + b) >= 0.5) == y_te).mean()
print(round(acc, 4))| metric | value |
|---|---|
| loss epoch 0 | 0.6931 (= ln 2) |
| loss epoch 499 | 0.0027 |
| train accuracy | 1.0000 |
| test accuracy | 1.0000 |
| sklearn match | 1.0000 |
Discrimination
Sort into buckets
Sort these by value, from memory, without looking back at Problem 1 — logistic regression from scratch. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Anomaly
Predict first
A student writes this, and it looks reasonable:
The gradient is X.T @ (p − y) — no division, just sum over the batch.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Summing (not averaging) the gradient makes the effective learning rate scale with dataset size n.
Divide by n to average: gradient = X.T @ (p − y) / n.
Why: Summing (not averaging) the gradient makes the effective learning rate scale with dataset size n. On 80 training points lr=0.5 becomes effectively lr=40 — the loss explodes or oscillates instead of converging.
Trap
The gradient is X.T @ (p − y) — no division, just sum over the batch.
Update: w -= lr * X.T @ (p - y)
Why: Summing (not averaging) the gradient makes the effective learning rate scale with dataset size n. On 80 training points lr=0.5 becomes effectively lr=40 — the loss explodes or oscillates instead of converging.
Divide by n to average: gradient = X.T @ (p − y) / n.
Update: w -= lr * X.T @ (p - y) / len(X_tr)
Why: Averaging makes the update scale-independent of n, so the same lr=0.5 works on 80 or 8,000 samples. This is the source of the standard formula and is checked explicitly on the rubric.
Break the constraint
Discussion prompt
The rule this trap just fixed:
Averaging makes the update scale-independent of n, so the same lr=0.5 works on 80 or 8,000 samples. This is the source of the standard formula and is checked explicitly on the rubric.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
Summing (not averaging) the gradient makes the effective learning rate scale with dataset size n. On 80 training points lr=0.5 becomes effectively lr=40 — the loss explodes or oscillates instead of converging.
Concept
| epoch | loss | |dw| | db |
|---|---|---|---|
| 0 | 0.6931 | 0.8327 | -0.0250 |
| 1 | 0.4084 | 0.5412 | -0.0173 |
| 2 | 0.2834 | 0.3876 | -0.0142 |
Both |dw| and |db| shrink each step — the model is converging. Loss dropping from 0.6931 → 0.4084 (−41%) in one epoch is typical for well-scaled features (iris is z-scored).
Pattern
Step through it
Step through Gradient descent trace — 3 epochs one row at a time. What is driving the change, and what would the row after the last one be?
Section
Part 4 of 4
Worked example
Theory sub-q: what is the receptive field of two stacked 3×3 convolutions? Then: build a two-conv + pool + two-FC CNN, train for 15 epochs with Adam, and report test accuracy.
Hint: receptive field after two 3×3 convs (stride 1) = 5×5. Reshape digits (N,64) to (N,1,8,8). Predict: accuracy should exceed 0.97 after 15 epochs.
import torch, torch.nn as nn
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
from torch.utils.data import TensorDataset, DataLoader
digits = load_digits()
X = digits.data.astype('float32') / 16.0
y = digits.target
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.2, random_state=42)
Xtr_t = torch.tensor(Xtr.reshape(-1,1,8,8))
Xte_t = torch.tensor(Xte.reshape(-1,1,8,8))
ytr_t = torch.tensor(ytr, dtype=torch.long)
yte_t = torch.tensor(yte, dtype=torch.long)
class CNN(nn.Module):
def __init__(self):
super().__init__()
self.net = nn.Sequential(
nn.Conv2d(1,16,3,padding=1), nn.ReLU(),
nn.Conv2d(16,32,3,padding=1), nn.ReLU(),
nn.MaxPool2d(2), nn.Flatten(),
nn.Linear(32*4*4,64), nn.ReLU(),
nn.Linear(64,10))
def forward(self,x): return self.net(x)
torch.manual_seed(0)
model = CNN()
opt = torch.optim.Adam(model.parameters(), lr=1e-3)
loss_fn = nn.CrossEntropyLoss()
for _ in range(15):
model.train()
for xb,yb in DataLoader(TensorDataset(Xtr_t,ytr_t),32,shuffle=True):
opt.zero_grad(); loss_fn(model(xb),yb).backward(); opt.step()
model.eval()
with torch.no_grad():
acc = (model(Xte_t).argmax(1)==yte_t).float().mean().item()
print(round(acc, 4))| epoch | train loss | train acc |
|---|---|---|
| 1 | 2.1695 | 0.5915 |
| 2 | 1.0748 | 0.8608 |
| 3 | 0.4026 | 0.9283 |
| 4 | 0.2552 | 0.9478 |
| 5 | 0.1823 | 0.9576 |
Pattern
Step through it
Step through Problem 2 — CNN on load_digits one row at a time. What is driving the change, and what would the row after the last one be?
Concept
| layer | output shape | params |
|---|---|---|
| Conv2d(1,16,3,p=1) + ReLU | (N,16,8,8) | 160 |
| Conv2d(16,32,3,p=1) + ReLU | (N,32,8,8) | 4640 |
| MaxPool2d(2) + Flatten | (N,512) | 0 |
| Linear(512,64) + ReLU | (N,64) | 32832 |
| Linear(64,10) | (N,10) | 650 |
Total: 38,282 parameters. Receptive field of two stacked 3×3 convolutions (stride 1): 5×5. After the pooling layer, spatial resolution halves from 8×8 to 4×4.
Counterexample
Discussion prompt
Total: 38,282 parameters. Receptive field of two stacked 3×3 convolutions (stride 1): 5×5. After the pooling layer, spatial resolution halves from 8×8 to 4×4.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Anomaly
Predict first
A student writes this, and it looks reasonable:
After training, just run model(X_test) — the model was trained, so it's ready.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Dropout randomly zeroes activations at every forward pass.
Always call model.eval() (and torch.no_grad()) before any inference.
Why: Dropout randomly zeroes activations at every forward pass. In train mode, two calls on the SAME input produce DIFFERENT logits. The test accuracy appears to jitter and is systematically too low — an exam rubric failure.
Trap
After training, just run model(X_test) — the model was trained, so it's ready.
Evaluate while model is still in train mode
Why: Dropout randomly zeroes activations at every forward pass. In train mode, two calls on the SAME input produce DIFFERENT logits. The test accuracy appears to jitter and is systematically too low — an exam rubric failure.
Always call model.eval() (and torch.no_grad()) before any inference.
model.eval() → torch.no_grad() → model(X_test)
Why: eval() disables dropout (and fixes BatchNorm stats). Identical inputs now produce identical outputs. The verified CNN test accuracy is 0.9806 — consistent across runs.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
X.T @ (p − y) — no division, just sum over the batch.; After training, just run model(X_test) — the model was trained, so it's ready.Concept
# KEY VALUES — all verified by real execution
# Problem 1: Logistic Regression from scratch on iris (binary)
# loss epoch 0 : 0.6931 (= ln 2, since p=0.5 everywhere)
# loss epoch 499 : 0.0027 (near-zero: linearly separable)
# test accuracy : 1.0000 (matches sklearn LogReg)
# gradient formula: dL/dw = X.T @ (sigmoid(Xw+b) - y) / n
#
# Problem 2: CNN (2 conv + pool + 2 FC) on load_digits, 15 epochs Adam
# epoch 1 loss : 2.1695, train acc: 0.5915
# epoch 5 loss : 0.1823, train acc: 0.9576
# test accuracy : 0.9806
# total params : 38,282
# receptive field (2x 3x3 conv, stride=1): 5x5
print('verified')| problem | key result | verified vs |
|---|---|---|
| Logistic Regression | test acc = 1.0000 | sklearn LogisticRegression |
| Cross-entropy gradient | X.T @ (p − y) / n | analytic derivation |
| CNN 15 epochs | test acc = 0.9806 | deterministic (torch.manual_seed(0)) |
| CNN receptive field | 5×5 after 2 conv(3×3) | formula: 1+(k-1)*L |
| CNN param count | 38,282 total | manual layer sum |
If either result differs, the most common causes are: (1) no /n in the gradient, (2) no model.eval() before CNN inference, (3) wrong reshape (N,64) instead of (N,1,8,8).
Comparison
Comparison matrix
From Full answer key — both coding problems: refill the key result column from what you know. The rest of the table is as it appeared.
| problem | key result | verified vs |
|---|---|---|
| Logistic Regression | test acc = 1.0000 | sklearn LogisticRegression |
| Cross-entropy gradient | X.T @ (p − y) / n | analytic derivation |
| CNN 15 epochs | test acc = 0.9806 | deterministic (torch.manual_seed(0)) |
| CNN receptive field | 5×5 after 2 conv(3×3) | formula: 1+(k-1)*L |
| CNN param count | 38,282 total | manual layer sum |
Elimination
Eliminate the wrong options
A student implements: w -= lr * X.T @ (sigmoid(X@w+b) - y) (no /n). What goes wrong?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Without /n the step = lr × (sum of n gradients) instead of lr × (mean). On 80 points the effective lr is 80×; on 8,000 it's 8,000×. The model either oscillates or explodes. Dividing by n keeps the learning rate dataset-size-independent.
Check
Catch the subtle bug before it costs points.
Check your understanding
A student implements: w -= lr * X.T @ (sigmoid(X@w+b) - y) (no /n). What goes wrong?
Answer: A
Why: Without /n the step = lr × (sum of n gradients) instead of lr × (mean). On 80 points the effective lr is 80×; on 8,000 it's 8,000×. The model either oscillates or explodes. Dividing by n keeps the learning rate dataset-size-independent.
Section
Project
Concept
Work each problem to completion before looking at the answer key. Time yourself: ≤ 45 min per coding problem.
| # | task | verification |
|---|---|---|
| 1 | Logistic regression from scratch on iris binary | test acc = 1.0 vs sklearn |
| 2 | Cross-entropy gradient — write it out analytically first | X.T@(p−y)/n |
| 3 | CNN 2-conv+pool+2-FC on load_digits, 15 epochs Adam | test acc ≥ 0.97 |
| 4 | Apply model.eval() + no_grad before reporting accuracy | consistent result |
Build rules: type every line, verify shapes at each layer (x.shape), and don't peek at the answer key until your code runs.
Counterexample
Discussion prompt
Work each problem to completion before looking at the answer key. Time yourself: ≤ 45 min per coding problem.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Build rules: type every line, verify shapes at each layer (x.shape), and don't peek at the answer key until your code runs.
Worked example
Your turn: implement sigmoid, the gradient, and the update loop. Predict the epoch-0 loss before running.
Hint: at initialization w=0, b=0, so p=0.5 for every point, and binary cross-entropy = −log(0.5) = log(2) ≈ 0.6931.
import numpy as np
from sklearn.datasets import load_iris
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split
iris = load_iris()
X = StandardScaler().fit_transform(iris.data[:100])
y = iris.target[:100].astype(float)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, random_state=42)
def sigmoid(z): return 1 / (1 + np.exp(-z))
w, b = np.zeros(4), 0.0
for ep in range(3): # first 3 epochs to see trace
p = sigmoid(X_tr @ w + b)
loss = -np.mean(y_tr*np.log(p+1e-12)+(1-y_tr)*np.log(1-p+1e-12))
w -= 0.5 * X_tr.T @ (p - y_tr) / len(X_tr)
b -= 0.5 * np.mean(p - y_tr)
print(f'epoch {ep}: loss={loss:.4f}')| epoch | loss |
|---|---|
| 0 | 0.6931 |
| 1 | 0.4084 |
| 2 | 0.2834 |
Pattern
Step through it
Step through Milestone 1 — logistic regression skeleton one row at a time. What is driving the change, and what would the row after the last one be?
Worked example
Your turn: build the CNN class. Before coding, write out each layer's output shape on paper.
Hint: Conv2d(1,16,3,padding=1) preserves 8×8. MaxPool2d(2) halves to 4×4. After flatten: 32×4×4 = 512. Then Linear(512,64), Linear(64,10).
import torch, torch.nn as nn
class CNN(nn.Module):
def __init__(self):
super().__init__()
self.net = nn.Sequential(
nn.Conv2d(1, 16, 3, padding=1), nn.ReLU(),
nn.Conv2d(16, 32, 3, padding=1), nn.ReLU(),
nn.MaxPool2d(2), nn.Flatten(),
nn.Linear(32*4*4, 64), nn.ReLU(),
nn.Linear(64, 10))
def forward(self, x): return self.net(x)
# verify architecture
model = CNN()
print(sum(p.numel() for p in model.parameters())) # 38282
x_probe = torch.zeros(1, 1, 8, 8)
print(model(x_probe).shape) # torch.Size([1, 10])| check | value |
|---|---|
| total params | 38282 |
| output shape | torch.Size([1, 10]) |
| receptive field (2 conv 3x3) | 5x5 |
Worked example
Your turn: write the DataLoader loop (15 epochs, Adam, CrossEntropyLoss) and finish with model.eval() before reporting accuracy.
Hint: reshape digits to (N,1,8,8), normalize by dividing by 16. After 15 epochs test accuracy should exceed 0.97.
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
from torch.utils.data import TensorDataset, DataLoader
digits = load_digits()
X = digits.data.astype('float32') / 16.0
y = digits.target
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.2, random_state=42)
Xtr_t = torch.tensor(Xtr.reshape(-1,1,8,8))
Xte_t = torch.tensor(Xte.reshape(-1,1,8,8))
ytr_t = torch.tensor(ytr, dtype=torch.long)
yte_t = torch.tensor(yte, dtype=torch.long)
torch.manual_seed(0)
model = CNN()
opt = torch.optim.Adam(model.parameters(), lr=1e-3)
for _ in range(15):
model.train()
for xb, yb in DataLoader(TensorDataset(Xtr_t,ytr_t), 32, shuffle=True):
opt.zero_grad(); nn.CrossEntropyLoss()(model(xb),yb).backward(); opt.step()
model.eval()
with torch.no_grad():
acc = (model(Xte_t).argmax(1)==yte_t).float().mean().item()
print(round(acc, 4)) # 0.9806| epoch | train loss | train acc |
|---|---|---|
| 1 | 2.1695 | 0.5915 |
| 3 | 0.4026 | 0.9283 |
| 5 | 0.1823 | 0.9576 |
| 10 | 0.0785 | ≈0.98 |
| 15 (test) | 0.0466 | 0.9806 |
Trade off
Comparison matrix
From Milestone 3 — training loop & eval: every row here is a choice with a cost. Fill the train acc column, then say which row you would actually pick and what you give up for it.
| epoch | train loss | train acc |
|---|---|---|
| 1 | 2.1695 | 0.5915 |
| 3 | 0.4026 | 0.9283 |
| 5 | 0.1823 | 0.9576 |
| 10 | 0.0785 | ≈0.98 |
| 15 (test) | 0.0466 | 0.9806 |
Concept
# Full Phase 2 exam submission — both problems
import numpy as np, torch, torch.nn as nn
from sklearn.datasets import load_iris, load_digits
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split
from torch.utils.data import TensorDataset, DataLoader
# --- Problem 1: Logistic Regression ---
iris = load_iris()
X = StandardScaler().fit_transform(iris.data[:100])
y = iris.target[:100].astype(float)
X_tr,X_te,y_tr,y_te = train_test_split(X,y,test_size=0.2,random_state=42)
def sigmoid(z): return 1/(1+np.exp(-z))
w,b = np.zeros(4), 0.0
for _ in range(500):
p = sigmoid(X_tr@w+b)
w -= 0.5*X_tr.T@(p-y_tr)/len(X_tr); b -= 0.5*np.mean(p-y_tr)
print('LR test acc:', ((sigmoid(X_te@w+b)>=0.5)==y_te).mean())
# --- Problem 2: CNN ---
d = load_digits()
Xtr,Xte,ytr,yte = train_test_split(d.data.astype('f4')/16,d.target,test_size=.2,random_state=42)
Xtr_t,Xte_t = torch.tensor(Xtr.reshape(-1,1,8,8)),torch.tensor(Xte.reshape(-1,1,8,8))
ytr_t,yte_t = torch.tensor(ytr,dtype=torch.long),torch.tensor(yte,dtype=torch.long)
class CNN(nn.Module):
def __init__(self):
super().__init__()
self.net=nn.Sequential(nn.Conv2d(1,16,3,padding=1),nn.ReLU(),nn.Conv2d(16,32,3,padding=1),nn.ReLU(),nn.MaxPool2d(2),nn.Flatten(),nn.Linear(512,64),nn.ReLU(),nn.Linear(64,10))
def forward(self,x): return self.net(x)
torch.manual_seed(0); model=CNN(); opt=torch.optim.Adam(model.parameters(),lr=1e-3)
for _ in range(15):
model.train()
for xb,yb in DataLoader(TensorDataset(Xtr_t,ytr_t),32,shuffle=True):
opt.zero_grad(); nn.CrossEntropyLoss()(model(xb),yb).backward(); opt.step()
model.eval()
with torch.no_grad(): print('CNN test acc:', round((model(Xte_t).argmax(1)==yte_t).float().mean().item(),4))| output line | value |
|---|---|
| LR test acc: | 1.0 |
| CNN test acc: | 0.9806 |
Both results match the answer key exactly. If yours differs, check (1) /len(X_tr) in the gradient, (2) model.eval() before CNN inference, (3) torch.manual_seed(0) before CNN init.
Comparison
Comparison matrix
From The full program — both problems: refill the value column from what you know. The rest of the table is as it appeared.
| output line | value |
|---|---|
| LR test acc: | 1.0 |
| CNN test acc: | 0.9806 |
Concept
Out loud, slides closed: (1) derive the logistic regression gradient from the cross-entropy loss, (2) explain why the /n matters for learning-rate stability, (3) describe why model.eval() is required before any CNN inference.
Grade your theory section: identify which topic area had the most misses and revisit those slides before Lesson 79. Phase 3 starts with attention mechanisms — the most heavily weighted USAAIO topic.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Exam format & rules · Theory section — 30 questions · Coding problem 1 — Logistic Regression · Coding problem 2 — CNN in PyTorch · Your turn: build both. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
/n normalization| concept | the one thing to remember |
|---|---|
| logistic gradient | X.T @ (p − y) / n — divide by n |
| cross-entropy init | loss = ln(2) ≈ 0.6931 when p = 0.5 |
| CNN inference | model.eval() + torch.no_grad() always |
| receptive field | 2× conv(3×3, stride 1) → 5×5 |
| Phase 2 milestone | >75% theory + both coding problems complete |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.