USAAIO Lesson 74, from Phase 4. It gives a systematic five-step debugging protocol for a broken training loop: overfit a single batch, check gradient flow with per-layer norms, verify the initial loss against log(K), check the initial accuracy against 1/K, and start from hyperparameters taken from the literature. It includes an automated checklist function and a bug-injection exercise. The lesson runs to 30 slides.
Subject: Machine Learning · 60 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 74 · Phase 4
Five ordered checks that diagnose almost every broken training loop before you reach for hyperparameter tuning.
Objectives
Warm-up
Discussion prompt
Before we open Lesson 74: Neural Network Debugging Checklist: without looking back, what was the main idea of Metric Learning, Siamese Networks & Triplet Loss, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
learn embeddings where distance reflects semantic similarity — metric learning, Siamese networks, triplet loss (margin-based), contrastive loss, online hard negative mining, and the connection to attention. Implements a Siamese EmbeddingNet in PyTorch; hard mining improves inter/intra distance ratio from 5.60× to 8.17× on a 3-identity benchmark.
Section
Section 1 of 5
Concept
Before tuning anything, verify your model can learn. A healthy network should drive training loss to ~0 on a batch of 8–32 samples within a few hundred steps.
If it cannot: the bug is in the architecture, loss, or data pipeline — not in the learning rate. Tuning hyperparameters on a broken model is wasted effort.
Counterexample
Discussion prompt
Before tuning anything, verify your model can learn. A healthy network should drive training loss to ~0 on a batch of 8–32 samples within a few hundred steps.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
If it cannot: the bug is in the architecture, loss, or data pipeline — not in the learning rate. Tuning hyperparameters on a broken model is wasted effort.
Estimation
Predict first
4-class network, batch of 8. Predict: will loss reach ~0?
Commit before you compute: what does Overfit a batch of 8 — loss trace come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Loss falls from 1.41 to ~0 by epoch 20; batch accuracy hits 1.0 by epoch 5
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The network can learn the data — no structural bug.
Worked example
4-class network, batch of 8. Predict: will loss reach ~0?
import torch, torch.nn as nn, torch.optim as optim
torch.manual_seed(0)
X = torch.randn(8, 64) # fixed batch
y = torch.tensor([0,1,2,3,0,1,2,3])
class Net(nn.Module):
def __init__(self):
super().__init__()
self.fc1 = nn.Linear(64, 32)
self.fc2 = nn.Linear(32, 4)
def forward(self, x):
return self.fc2(torch.relu(self.fc1(x)))
net = Net()
opt = optim.Adam(net.parameters(), lr=1e-2)
lossfn = nn.CrossEntropyLoss()
for ep in range(101):
opt.zero_grad()
loss = lossfn(net(X), y)
loss.backward(); opt.step()
if ep in [0,5,20,50,100]:
acc = (net(X).argmax(1)==y).float().mean()
print(f'ep={ep:3d} loss={loss.item():.4f} acc={acc:.2f}')| epoch | loss | batch acc |
|---|---|---|
| 0 | 1.4115 | 0.375 |
| 5 | 0.4055 | 1.000 |
| 20 | 0.0007 | 1.000 |
| 50 | ~0.000 | 1.000 |
| 100 | ~0.000 | 1.000 |
Loss falls from 1.41 to ~0 by epoch 20; batch accuracy hits 1.0 by epoch 5
Why: The network can learn the data — no structural bug. If this fails, debug the architecture before touching the learning rate.
Comparison
Comparison matrix
From Overfit a batch of 8 — loss trace: refill the loss column from what you know. The rest of the table is as it appeared.
| epoch | loss | batch acc |
|---|---|---|
| 0 | 1.4115 | 0.375 |
| 5 | 0.4055 | 1.000 |
| 20 | 0.0007 | 1.000 |
| 50 | ~0.000 | 1.000 |
| 100 | ~0.000 | 1.000 |
Anomaly
Predict first
A student writes this, and it looks reasonable:
Training loss is stuck at 1.5 after 1000 steps. The first move: lower the learning rate.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: A stuck loss may have nothing to do with the learning rate.
Training loss is stuck at 1.5. First: overfit a single batch.
Why: A stuck loss may have nothing to do with the learning rate. Lowering it on a broken model wastes GPU hours and produces no signal.
Trap
Training loss is stuck at 1.5 after 1000 steps. The first move: lower the learning rate.
Try lr=0.001, 0.0001, 0.00001 in sequence
Why: A stuck loss may have nothing to do with the learning rate. Lowering it on a broken model wastes GPU hours and produces no signal.
Training loss is stuck at 1.5. First: overfit a single batch.
Pass 8 fixed samples with a large lr — can the network reach loss ~0?
Why: If not, the bug is structural (wrong loss, bad init, missing activation). Fix the bug first. Tuning hyperparameters is Step 5, not Step 1.
Break the constraint
Discussion prompt
The rule this trap just fixed:
If not, the bug is structural (wrong loss, bad init, missing activation). Fix the bug first. Tuning hyperparameters is Step 5, not Step 1.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
A stuck loss may have nothing to do with the learning rate. Lowering it on a broken model wastes GPU hours and produces no signal.
Section
Section 2 of 5
Concept
After one backward pass, compute param.grad.norm() for every named layer. A healthy network shows similar-order norms throughout the stack.
Vanishing: norms shrink exponentially toward the input (early layers see near-zero gradient — they do not learn). Exploding: norms grow, causing NaN loss and divergence.
| symptom | grad norm pattern | likely cause |
|---|---|---|
| vanishing | near 0 in early layers, normal in last | deep Sigmoid/Tanh, bad init |
| exploding | >1e3 anywhere | no gradient clipping, bad init |
| healthy | same order of magnitude throughout | ReLU/BatchNorm, He/Xavier init |
Trade off
Comparison matrix
From Gradient norms reveal where learning stops: every row here is a choice with a cost. Fill the grad norm pattern column, then say which row you would actually pick and what you give up for it.
| symptom | grad norm pattern | likely cause |
|---|---|---|
| vanishing | near 0 in early layers, normal in last | deep Sigmoid/Tanh, bad init |
| exploding | >1e3 anywhere | no gradient clipping, bad init |
| healthy | same order of magnitude throughout | ReLU/BatchNorm, He/Xavier init |
Missing information
Discussion prompt
Two architectures: 2-layer ReLU (healthy) vs 3-layer Sigmoid (vanishing). Same data, one backward pass.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Each Sigmoid gate saturates and multiplies the gradient by σ'≤0.25. Three such gates multiply the early gradient by ≤0.016, matching the observed ~50× shrinkage.
Worked example
Two architectures: 2-layer ReLU (healthy) vs 3-layer Sigmoid (vanishing). Same data, one backward pass.
import torch, torch.nn as nn
torch.manual_seed(0)
X = torch.randn(8, 64)
y = torch.tensor([0,1,2,3,0,1,2,3])
lossfn = nn.CrossEntropyLoss()
# Healthy: ReLU 2-layer
class ReLUNet(nn.Module):
def __init__(self):
super().__init__()
self.fc1 = nn.Linear(64,32)
self.fc2 = nn.Linear(32,4)
def forward(self,x): return self.fc2(torch.relu(self.fc1(x)))
# Broken: Sigmoid 3-layer (vanishing gradients)
class SigNet(nn.Module):
def __init__(self):
super().__init__()
self.layers = nn.Sequential(
nn.Linear(64,64), nn.Sigmoid(),
nn.Linear(64,64), nn.Sigmoid(),
nn.Linear(64,64), nn.Sigmoid(),
nn.Linear(64,4))
def forward(self,x): return self.layers(x)
for name, net in [('ReLU', ReLUNet()), ('Sigmoid', SigNet())]:
torch.manual_seed(0); net = net.__class__()
lossfn(net(X), y).backward()
print(f'--- {name} ---')
for n, p in net.named_parameters():
if p.grad is not None:
print(f' {n}: {p.grad.norm():.6f}')| architecture | layer | grad norm |
|---|---|---|
| ReLU (healthy) | fc1.weight | 0.931085 |
| ReLU (healthy) | fc2.weight | 0.527611 |
| Sigmoid (broken) | layers.0.weight | 0.006976 |
| Sigmoid (broken) | layers.2.weight | 0.010387 |
| Sigmoid (broken) | layers.4.weight | 0.060878 |
| Sigmoid (broken) | layers.6.weight | 0.366669 |
Sigmoid early layers: norms ~0.007–0.011, 50× smaller than the final layer; ReLU layers: norms stay in the same order throughout
Why: Each Sigmoid gate saturates and multiplies the gradient by σ'≤0.25. Three such gates multiply the early gradient by ≤0.016, matching the observed ~50× shrinkage.
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Sigmoid early layers: norms ~0.007–0.011, 50× smaller than the final layer; ReLU layers: norms stay in the same order throughout
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Two architectures: 2-layer ReLU (healthy) vs 3-layer Sigmoid (vanishing). Same data, one backward pass.
Section
Section 3 of 5
Concept
A randomly initialized network (near-zero weights → uniform logits → uniform softmax) should output probability 1/K for every class. The expected cross-entropy is therefore log(K).
\[ \mathcal{L}_\text{init} \approx -\log\frac{1}{K} = \log K \]
| K (classes) | expected initial CE loss |
|---|---|
| 2 | 0.6931 |
| 4 | 1.3863 |
| 5 | 1.6094 |
| 10 | 2.3026 |
| 100 | 4.6052 |
If initial loss is far above log(K): bad weight initialization (huge weights push logits to extremes, causing large CE). If far below: something is leaking label information into initialization.
Pattern
Step through it
Step through Random-classifier baseline: loss = log(K) one row at a time. What is driving the change, and what would the row after the last one be?
Estimation
Predict first
K=10, 100 samples. Compare near-zero init (standard) vs huge init (σ=10.0).
Commit before you compute: what does Verifying initial loss — near-zero init vs bad init come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Standard init → loss 2.298 ≈ log(10) = 2.303; huge init → loss 5051 (>1000× too high)
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Large random weights produce extreme logits; softmax saturates; CE can be enormous.
Worked example
K=10, 100 samples. Compare near-zero init (standard) vs huge init (σ=10.0).
import torch, torch.nn as nn, numpy as np
torch.manual_seed(42)
K = 10
X = torch.randn(100, 64)
y = torch.randint(0, K, (100,))
def check_init_loss(model, X, y, K):
model.eval()
with torch.no_grad():
loss = nn.CrossEntropyLoss()(model(X), y).item()
acc = (model(X).argmax(1) == y).float().mean().item()
expected = np.log(K)
ok = abs(loss - expected) < 0.5
return {'loss': round(loss,4), 'expected': round(expected,4),
'acc': round(acc,4), 'expected_acc': round(1/K,4), 'ok': ok}
net_good = nn.Linear(64, K) # default init (~near-zero)
net_bad = nn.Linear(64, K)
with torch.no_grad(): # huge init
nn.init.normal_(net_bad.weight, 0, 10.0)
nn.init.normal_(net_bad.bias, 0, 10.0)
for name, net in [('good init', net_good), ('bad init', net_bad)]:
r = check_init_loss(net, X, y, K)
print(f'{name}: loss={r["loss"]} expected={r["expected"]} ok={r["ok"]}')| init | actual loss | expected (log 10) | ok? |
|---|---|---|---|
| near-zero (standard) | 2.2980 | 2.3026 | True |
| huge (σ=10) | 5051.48 | 2.3026 | False |
Standard init → loss 2.298 ≈ log(10) = 2.303; huge init → loss 5051 (>1000× too high)
Why: Large random weights produce extreme logits; softmax saturates; CE can be enormous. This is a sign of bad initialization or a forgotten normalization layer.
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Standard init → loss 2.298 ≈ log(10) = 2.303; huge init → loss 5051 (>1000× too high)
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
K=10, 100 samples. Compare near-zero init (standard) vs huge init (σ=10.0).
Section
Section 4 of 5
Concept
Paired with the loss check: at initialization a random classifier should achieve accuracy ~1/K (random chance). If initial accuracy is already 0.8 on your test set, the model or data pipeline may be broken (label leak, duplicated samples, etc.).
| K | expected accuracy (1/K) | expected CE loss (log K) |
|---|---|---|
| 2 (binary) | 0.500 | 0.693 |
| 4 | 0.250 | 1.386 |
| 10 | 0.100 | 2.303 |
| 100 | 0.010 | 4.605 |
Check both together: loss and accuracy at initialization form a self-consistent pair. A discrepancy (correct loss, wrong accuracy) points to an output shape bug or a metric computation error.
Pattern
Step through it
Step through Initial accuracy baseline: 1/K one row at a time. What is driving the change, and what would the row after the last one be?
Anomaly
Predict first
A student writes this, and it looks reasonable:
4-class classification. Use MSELoss(logits, y) where y contains class indices 0–3 — training loss goes down, so it must be learning.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: MSE minimizes numeric distance to class indices, not probability calibration.
For K-class classification: use CrossEntropyLoss(logits, class_indices).
Why: MSE minimizes numeric distance to class indices, not probability calibration. The gradients point in a meaningless direction. Loss falling is not evidence of correct learning.
Trap
4-class classification. Use MSELoss(logits, y) where y contains class indices 0–3 — training loss goes down, so it must be learning.
Train with MSELoss(output, class_index_tensor)
Why: MSE minimizes numeric distance to class indices, not probability calibration. The gradients point in a meaningless direction. Loss falling is not evidence of correct learning.
For K-class classification: use CrossEntropyLoss(logits, class_indices).
Check initial CE loss ≈ log(K) and initial acc ≈ 1/K to confirm the loss is set up correctly
Why: CrossEntropyLoss expects raw logits and integer labels, produces calibrated probabilities, and obeys the log(K) baseline at init — the canary that confirms the loss is wired right.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
param.grad.norm() for every named layer. A healthy network shows similar-order norms throughout the stack.; Hyperparameter tuning is the last step, not the first. Use proven defaults from the literature while debugging — they eliminate hyperparameters as a confound.MSELoss(logits, y) where y contains class indices 0–3 — training loss goes down, so it must be learning.Section
Section 5 of 5
Concept
Hyperparameter tuning is the last step, not the first. Use proven defaults from the literature while debugging — they eliminate hyperparameters as a confound.
| setting | reliable default |
|---|---|
| optimizer | Adam with lr=1e-3 (SGD with lr=0.1 for vision) |
| weight init | He (ReLU nets) / Xavier (tanh/sigmoid nets) |
| batch size | 32–256 for most tasks |
| loss | CrossEntropyLoss (classification) / MSELoss (regression) |
| activations | ReLU or GELU in hidden layers; no activation at output |
Once the model trains correctly with defaults, then tune. Changing hyperparameters on a broken model gives you noise, not signal.
Comparison
Comparison matrix
From Start from literature defaults, not a random search: refill the reliable default column from what you know. The rest of the table is as it appeared.
| setting | reliable default |
|---|---|
| optimizer | Adam with lr=1e-3 (SGD with lr=0.1 for vision) |
| weight init | He (ReLU nets) / Xavier (tanh/sigmoid nets) |
| batch size | 32–256 for most tasks |
| loss | CrossEntropyLoss (classification) / MSELoss (regression) |
| activations | ReLU or GELU in hidden layers; no activation at output |
Missing information
Discussion prompt
Encode all five checks as a single Python function. Run it before every new experiment.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
The checklist isolates architectural bugs from tuning problems. Every flag that fails points to a specific fix (bad init, wrong loss, dead activation) rather than 'tune the lr'.
Worked example
Encode all five checks as a single Python function. Run it before every new experiment.
import torch, torch.nn as nn, torch.optim as optim, numpy as np
def debug_checklist(model, X_batch, y_batch, X_full, y_full, K, lr=1e-2, steps=100):
results = {}
# Step 3 & 4: initial loss and accuracy
model.eval()
with torch.no_grad():
logits = model(X_full)
results['init_loss'] = round(nn.CrossEntropyLoss()(logits, y_full).item(), 4)
results['expected_loss'] = round(float(np.log(K)), 4)
results['init_acc'] = round((logits.argmax(1)==y_full).float().mean().item(), 4)
results['loss_ok'] = abs(results['init_loss'] - results['expected_loss']) < 0.5
# Step 2: gradient norms
model.train()
opt = optim.SGD(model.parameters(), lr=0.01)
opt.zero_grad()
nn.CrossEntropyLoss()(model(X_batch), y_batch).backward()
norms = {n: round(p.grad.norm().item(),6)
for n,p in model.named_parameters() if p.grad is not None}
results['grad_norms'] = norms
results['vanishing'] = any(v < 1e-4 for v in norms.values())
# Step 1: overfit a single batch
opt2 = optim.Adam(model.parameters(), lr=lr)
for _ in range(steps):
opt2.zero_grad()
nn.CrossEntropyLoss()(model(X_batch), y_batch).backward(); opt2.step()
with torch.no_grad():
final_loss = nn.CrossEntropyLoss()(model(X_batch), y_batch).item()
results['overfit_loss'] = round(final_loss, 6)
results['can_overfit'] = final_loss < 0.01
return results| check | key result | healthy? |
|---|---|---|
| init_loss vs log(K=4) | 1.397 vs 1.386 | True |
| init_acc vs 1/4 | 0.180 vs 0.250 | True (within 0.15) |
| vanishing grads | all norms > 0.028 | False (no vanishing) |
| can_overfit (ep=100) | 0.000014 | True |
All four flags healthy — the network is structurally sound; proceed to tuning
Why: The checklist isolates architectural bugs from tuning problems. Every flag that fails points to a specific fix (bad init, wrong loss, dead activation) rather than 'tune the lr'.
Discrimination
Sort into buckets
Sort these by healthy?, from memory, without looking back at The automated debugging checklist function. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Ranking
Put in order
These are the steps of The 5-step debugging protocol, scrambled. Put them back in order before the next slide shows you.
Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
| step | tool | healthy signal |
|---|---|---|
| 1 overfit | Adam on 8 samples, 100 steps | loss → ~0, acc → 1.0 |
| 2 grad flow | param.grad.norm() per layer | same order throughout stack |
| 3 loss scale | CE at init vs log(K) | within ±0.5 of log(K) |
| 4 metric | accuracy at init vs 1/K | within ±0.15 of 1/K |
| 5 defaults | lr=1e-3 Adam, He init, ReLU | trains correctly; then tune |
Edge cases
Discussion prompt
The 5-step debugging protocol works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
Elimination
Eliminate the wrong options
Your 10-class classifier shows initial cross-entropy loss of 4.8. What does this most likely indicate?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: A randomly initialized network with near-zero weights produces uniform logits → uniform softmax → CE ≈ log(10) = 2.303. A loss of 4.8 (≈ log(121)) means logits are highly non-uniform at init, which points to large random weights (bad initialization) or a forgotten input normalization.
Check
Work it out before clicking.
Check your understanding
Your 10-class classifier shows initial cross-entropy loss of 4.8. What does this most likely indicate?
Answer: A
Why: A randomly initialized network with near-zero weights produces uniform logits → uniform softmax → CE ≈ log(10) = 2.303. A loss of 4.8 (≈ log(121)) means logits are highly non-uniform at init, which points to large random weights (bad initialization) or a forgotten input normalization.
Prediction
Predict first
After one backward pass, grad norms are: layer1=0.007, layer2=0.009, layer3=0.058, layer4=0.37. What is the most likely cause and fix?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Vanishing gradients — early layers barely update; switch to ReLU or add batch normalization
Why: Norms of 0.007–0.009 in early layers vs 0.37 in the last layer indicate vanishing gradients — each Sigmoid gate multiplies the backward signal by σ'≤0.25, shrinking it ~50× across three layers. Early layers barely learn. Fix: replace Sigmoid with ReLU, or add BatchNorm to keep activations in the linear regime.
Check
Read the gradient norms and diagnose.
Check your understanding
After one backward pass, grad norms are: layer1=0.007, layer2=0.009, layer3=0.058, layer4=0.37. What is the most likely cause and fix?
Answer: A
Why: Norms of 0.007–0.009 in early layers vs 0.37 in the last layer indicate vanishing gradients — each Sigmoid gate multiplies the backward signal by σ'≤0.25, shrinking it ~50× across three layers. Early layers barely learn. Fix: replace Sigmoid with ReLU, or add BatchNorm to keep activations in the linear regime.
Elimination
Eliminate the wrong options
A network's training loss decreases smoothly but validation accuracy stays at chance (1/K) throughout training. Which bug is most consistent with this pattern?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Decreasing training loss with validation accuracy stuck at 1/K is the signature of a train/validation data pipeline bug. Most commonly: train and validation sets share samples (so the model memorizes train labels), or train labels were accidentally permuted before splitting while validation retains the real labels. The model learns something on train but that information doesn't transfer to clean validation data.
Check
Match symptom to bug.
Check your understanding
A network's training loss decreases smoothly but validation accuracy stays at chance (1/K) throughout training. Which bug is most consistent with this pattern?
Answer: A
Why: Decreasing training loss with validation accuracy stuck at 1/K is the signature of a train/validation data pipeline bug. Most commonly: train and validation sets share samples (so the model memorizes train labels), or train labels were accidentally permuted before splitting while validation retains the real labels. The model learns something on train but that information doesn't transfer to clean validation data.
Section
Project
Concept
You will (1) implement debug_checklist() from scratch, (2) inject 5 classic bugs into a working network and use your function to find them all.
| # | milestone | technique |
|---|---|---|
| 1 | Implement debug_checklist() | init loss/acc, grad norms, overfit test |
| 2 | Inject and find wrong loss (MSE→CE) | check init_loss vs log(K) |
| 3 | Inject and find missing activation | overfit test + manual forward trace |
| 4 | Inject and find bad init (huge σ) | init_loss >> log(K) flag |
| 5 | Inject and find vanishing grad (Sigmoid) | grad_norm < 1e-4 in early layers |
| 6 | Full run: fix all 5 bugs in < 10 min | timed exercise |
Build rule: run debug_checklist() after each injected bug and record which flag fires. A bug that doesn't trigger any flag means your checklist needs a new check.
Analogy
Discussion prompt
Explain Project: build and stress-test the debugging checklist by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
You will (1) implement debug_checklist() from scratch, (2) inject 5 classic bugs into a working network and use your function to find them all.
Worked example
Your turn: write debug_checklist(model, X_batch, y_batch, X_full, y_full, K) that returns a dict with at least: init_loss, loss_ok, init_acc, grad_norms, vanishing, overfit_loss, can_overfit.
Hint: check init loss/acc before any training step; do one backward to get .grad; then overfit X_batch for 100 steps with Adam lr=1e-2.
import torch, torch.nn as nn, torch.optim as optim, numpy as np
torch.manual_seed(0)
X_b = torch.randn(8, 64)
y_b = torch.tensor([0,1,2,3,0,1,2,3])
X_f = torch.randn(100, 64)
y_f = torch.randint(0, 4, (100,))
class Net(nn.Module):
def __init__(self):
super().__init__()
self.fc1 = nn.Linear(64,32)
self.fc2 = nn.Linear(32,4)
def forward(self,x): return self.fc2(torch.relu(self.fc1(x)))
net = Net()
r = debug_checklist(net, X_b, y_b, X_f, y_f, K=4)
print(r)| key | value | healthy? |
|---|---|---|
| init_loss | 1.3970 | True (≈ log 4 = 1.386) |
| init_acc | 0.1800 | True (1/4 = 0.250) |
| vanishing | False | True |
| can_overfit | True | True |
Worked example
Your turn: create a net, then overwrite its weights with nn.init.normal_(p, 0, 10.0). Run debug_checklist. Which flag fires?
Hint: init_loss will be far above log(K). The loss_ok flag should become False.
torch.manual_seed(0)
bad_net = Net()
with torch.no_grad():
for p in bad_net.parameters():
nn.init.normal_(p, 0, 10.0) # inject: huge init
r_bad = debug_checklist(bad_net, X_b, y_b, X_f, y_f, K=4)
print('init_loss:', r_bad['init_loss'])
print('loss_ok: ', r_bad['loss_ok'])| flag | healthy net | bad init net |
|---|---|---|
| init_loss | 1.397 | ≫ 1.386 (e.g. 2000+) |
| loss_ok | True | False (bug caught!) |
| can_overfit | True | may still be True |
Trade off
Comparison matrix
From Milestone 2 — inject and catch bad initialization: every row here is a choice with a cost. Fill the healthy net column, then say which row you would actually pick and what you give up for it.
| flag | healthy net | bad init net |
|---|---|---|
| init_loss | 1.397 | ≫ 1.386 (e.g. 2000+) |
| loss_ok | True | False (bug caught!) |
| can_overfit | True | may still be True |
Estimation
Predict first
Your turn: build a 3-layer Sigmoid network. Run debug_checklist. Which flags fire?
Commit before you compute: what does Milestone 3 — inject and catch vanishing gradients come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Early layers: grad_norm < 0.01, 50× below last layer; vanishing=True flag fires
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Fix: replace nn.Sigmoid() with nn.ReLU() throughout the hidden stack.
Worked example
Your turn: build a 3-layer Sigmoid network. Run debug_checklist. Which flags fire?
Hint: vanishing should be True because early-layer grad norms will be < 1e-4 after one backward pass.
class SigNet(nn.Module):
def __init__(self):
super().__init__()
self.layers = nn.Sequential(
nn.Linear(64,64), nn.Sigmoid(),
nn.Linear(64,64), nn.Sigmoid(),
nn.Linear(64,64), nn.Sigmoid(),
nn.Linear(64,4))
def forward(self,x): return self.layers(x)
torch.manual_seed(0)
sig_net = SigNet()
r_sig = debug_checklist(sig_net, X_b, y_b, X_f, y_f, K=4)
print('vanishing:', r_sig['vanishing'])
for k,v in r_sig['grad_norms'].items():
print(f' {k}: {v}')| layer | grad norm | flag |
|---|---|---|
| layers.0.weight (early) | 0.006976 | vanishing=True |
| layers.2.weight (mid) | 0.010387 | vanishing=True |
| layers.4.weight (mid) | 0.060878 | borderline |
| layers.6.weight (last) | 0.366669 | healthy |
Early layers: grad_norm < 0.01, 50× below last layer; vanishing=True flag fires
Why: Fix: replace nn.Sigmoid() with nn.ReLU() throughout the hidden stack. With ReLU the gradient passes through unsaturated neurons unchanged, keeping early-layer norms in the same order as the last layer.
Comparison
Comparison matrix
From Milestone 3 — inject and catch vanishing gradients: refill the grad norm column from what you know. The rest of the table is as it appeared.
| layer | grad norm | flag |
|---|---|---|
| layers.0.weight (early) | 0.006976 | vanishing=True |
| layers.2.weight (mid) | 0.010387 | vanishing=True |
| layers.4.weight (mid) | 0.060878 | borderline |
| layers.6.weight (last) | 0.366669 | healthy |
Concept
import torch, torch.nn as nn, torch.optim as optim, numpy as np
torch.manual_seed(0)
X_b = torch.randn(8, 64); y_b = torch.tensor([0,1,2,3,0,1,2,3])
X_f = torch.randn(100, 64); y_f = torch.randint(0, 4, (100,))
class Net(nn.Module):
def __init__(self):
super().__init__()
self.fc1=nn.Linear(64,32); self.fc2=nn.Linear(32,4)
def forward(self,x): return self.fc2(torch.relu(self.fc1(x)))
def debug_checklist(model, X_batch, y_batch, X_full, y_full, K, lr=1e-2, steps=100):
r = {}
model.eval()
with torch.no_grad():
logits = model(X_full)
r['init_loss'] = round(nn.CrossEntropyLoss()(logits, y_full).item(),4)
r['expected_loss'] = round(float(np.log(K)),4)
r['init_acc'] = round((logits.argmax(1)==y_full).float().mean().item(),4)
r['loss_ok'] = abs(r['init_loss']-r['expected_loss']) < 0.5
model.train()
opt = optim.SGD(model.parameters(), lr=0.01)
opt.zero_grad()
nn.CrossEntropyLoss()(model(X_batch),y_batch).backward()
norms = {n:round(p.grad.norm().item(),6)
for n,p in model.named_parameters() if p.grad is not None}
r['grad_norms'] = norms; r['vanishing'] = any(v<1e-4 for v in norms.values())
opt2 = optim.Adam(model.parameters(), lr=lr)
for _ in range(steps):
opt2.zero_grad()
nn.CrossEntropyLoss()(model(X_batch),y_batch).backward(); opt2.step()
with torch.no_grad():
fl = nn.CrossEntropyLoss()(model(X_batch),y_batch).item()
r['overfit_loss']=round(fl,6); r['can_overfit']=(fl<0.01)
return r
net = Net()
report = debug_checklist(net, X_b, y_b, X_f, y_f, K=4)
for k,v in report.items(): print(f'{k}: {v}')| flag | value | interpretation |
|---|---|---|
| init_loss / expected_loss | 1.397 / 1.386 | initialization OK |
| loss_ok | True | no bad-init bug |
| vanishing | False | no gradient flow bug |
| can_overfit | True | architecture sound — proceed to tuning |
All four flags passing means the network is structurally sound. Step 5: now switch to literature hyperparameters and tune from there.
Discrimination
Sort into buckets
Sort these by value, from memory, without looking back at The full debugging program. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Concept
Out loud, slides closed: walk through the five steps in order and explain what each one catches. Then inject one bug of your choice, run the checklist, and narrate which flag fires and why.
Stretch (homework): extend debug_checklist to also (a) plot gradient norms as a bar chart per layer, (b) detect dead ReLU neurons (fraction of zero activations), and (c) flag if training loss increases in the first 5 steps. Next: hyperparameter tuning and learning rate schedules (L75).
Counterexample
Discussion prompt
Out loud, slides closed: walk through the five steps in order and explain what each one catches. Then inject one bug of your choice, run the checklist, and narrate which flag fires and why.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Step 1: Overfit a single batch · Step 2: Gradient flow · Step 3: Check loss scale · Step 4: Check metrics · Step 5: Known hyperparameters first · Your turn: debug it. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
debug_checklist() that encodes all five checks in one function| step | what it catches | healthy signal |
|---|---|---|
| 1 overfit batch | architecture / loss / data-pipeline bugs | loss ~0, acc 1.0 in <100 steps |
| 2 grad norms | vanishing / exploding gradients | same order of magnitude layer-to-layer |
| 3 init loss | bad init, missing normalization | within ±0.5 of log(K) |
| 4 init acc | label leak, shape bugs | within ±0.15 of 1/K |
| 5 defaults first | confounded hyperparameter search | train correctly, then tune |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.