Lesson 74: Neural Network Debugging Checklist

USAAIO Lesson 74, from Phase 4. It gives a systematic five-step debugging protocol for a broken training loop: overfit a single batch, check gradient flow with per-layer norms, verify the initial loss against log(K), check the initial accuracy against 1/K, and start from hyperparameters taken from the literature. It includes an automated checklist function and a bug-injection exercise. The lesson runs to 30 slides.

Subject: Machine Learning · 60 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Neural Network Debugging Checklist

Title

USAAIO · Lesson 74 · Phase 4

Five ordered checks that diagnose almost every broken training loop before you reach for hyperparameter tuning.

2. By the end of this lesson you can

Objectives

  1. Overfit a single batch as the first sanity check — if you can't, there is a bug not a tuning problem
  2. Plot per-layer gradient norms and identify vanishing or exploding gradients by name
  3. Verify initial loss against log(K) and initial accuracy against 1/K for a random classifier baseline
  4. Run an automated debugging checklist that flags all five failure modes in one pass
  5. Inject and find five classic bugs (wrong loss, missing activation, bad init, etc.) in under 10 minutes

3. What survived from Metric Learning, Siamese Networks & Triplet Loss?

Warm-up

Discussion prompt

Before we open Lesson 74: Neural Network Debugging Checklist: without looking back, what was the main idea of Metric Learning, Siamese Networks & Triplet Loss, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

learn embeddings where distance reflects semantic similarity — metric learning, Siamese networks, triplet loss (margin-based), contrastive loss, online hard negative mining, and the connection to attention. Implements a Siamese EmbeddingNet in PyTorch; hard mining improves inter/intra distance ratio from 5.60× to 8.17× on a 3-identity benchmark.

4. Step 1: Overfit a single batch

Section

Section 1 of 5

5. Why overfitting a batch is the first test

Concept

Before tuning anything, verify your model can learn. A healthy network should drive training loss to ~0 on a batch of 8–32 samples within a few hundred steps.

If it cannot: the bug is in the architecture, loss, or data pipeline — not in the learning rate. Tuning hyperparameters on a broken model is wasted effort.

6. Break it if you can: Why overfitting a batch is the first test

Counterexample

Discussion prompt

Before tuning anything, verify your model can learn. A healthy network should drive training loss to ~0 on a batch of 8–32 samples within a few hundred steps.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

If it cannot: the bug is in the architecture, loss, or data pipeline — not in the learning rate. Tuning hyperparameters on a broken model is wasted effort.

7. Guess the shape of the answer: Overfit a batch of 8 — loss trace

Estimation

Predict first

4-class network, batch of 8. Predict: will loss reach ~0?

Commit before you compute: what does Overfit a batch of 8 — loss trace come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Loss falls from 1.41 to ~0 by epoch 20; batch accuracy hits 1.0 by epoch 5

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The network can learn the data — no structural bug.

8. Overfit a batch of 8 — loss trace

Worked example

4-class network, batch of 8. Predict: will loss reach ~0?

import torch, torch.nn as nn, torch.optim as optim
torch.manual_seed(0)

X = torch.randn(8, 64)          # fixed batch
y = torch.tensor([0,1,2,3,0,1,2,3])

class Net(nn.Module):
    def __init__(self):
        super().__init__()
        self.fc1 = nn.Linear(64, 32)
        self.fc2 = nn.Linear(32, 4)
    def forward(self, x):
        return self.fc2(torch.relu(self.fc1(x)))

net = Net()
opt = optim.Adam(net.parameters(), lr=1e-2)
lossfn = nn.CrossEntropyLoss()
for ep in range(101):
    opt.zero_grad()
    loss = lossfn(net(X), y)
    loss.backward(); opt.step()
    if ep in [0,5,20,50,100]:
        acc = (net(X).argmax(1)==y).float().mean()
        print(f'ep={ep:3d} loss={loss.item():.4f} acc={acc:.2f}')
epochlossbatch acc
01.41150.375
50.40551.000
200.00071.000
50~0.0001.000
100~0.0001.000

Loss falls from 1.41 to ~0 by epoch 20; batch accuracy hits 1.0 by epoch 5

Why: The network can learn the data — no structural bug. If this fails, debug the architecture before touching the learning rate.

9. Fill in: loss for Overfit a batch of 8 — loss trace

Comparison

Comparison matrix

From Overfit a batch of 8 — loss trace: refill the loss column from what you know. The rest of the table is as it appeared.

epochlossbatch acc
01.41150.375
50.40551.000
200.00071.000
50~0.0001.000
100~0.0001.000

10. Something is wrong here: tuning hyperparameters before overfitting a batch

Anomaly

Predict first

A student writes this, and it looks reasonable:

Training loss is stuck at 1.5 after 1000 steps. The first move: lower the learning rate.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: A stuck loss may have nothing to do with the learning rate.

Training loss is stuck at 1.5. First: overfit a single batch.

Why: A stuck loss may have nothing to do with the learning rate. Lowering it on a broken model wastes GPU hours and produces no signal.

11. Trap: tuning hyperparameters before overfitting a batch

Trap

The trap

Training loss is stuck at 1.5 after 1000 steps. The first move: lower the learning rate.

Try lr=0.001, 0.0001, 0.00001 in sequence

Why: A stuck loss may have nothing to do with the learning rate. Lowering it on a broken model wastes GPU hours and produces no signal.

The fix

Training loss is stuck at 1.5. First: overfit a single batch.

Pass 8 fixed samples with a large lr — can the network reach loss ~0?

Why: If not, the bug is structural (wrong loss, bad init, missing activation). Fix the bug first. Tuning hyperparameters is Step 5, not Step 1.

12. Break it on purpose: tuning hyperparameters before overfitting a…

Break the constraint

Discussion prompt

The rule this trap just fixed:

If not, the bug is structural (wrong loss, bad init, missing activation). Fix the bug first. Tuning hyperparameters is Step 5, not Step 1.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

A stuck loss may have nothing to do with the learning rate. Lowering it on a broken model wastes GPU hours and produces no signal.

13. Step 2: Gradient flow

Section

Section 2 of 5

14. Gradient norms reveal where learning stops

Concept

After one backward pass, compute param.grad.norm() for every named layer. A healthy network shows similar-order norms throughout the stack.

Vanishing: norms shrink exponentially toward the input (early layers see near-zero gradient — they do not learn). Exploding: norms grow, causing NaN loss and divergence.

symptomgrad norm patternlikely cause
vanishingnear 0 in early layers, normal in lastdeep Sigmoid/Tanh, bad init
exploding>1e3 anywhereno gradient clipping, bad init
healthysame order of magnitude throughoutReLU/BatchNorm, He/Xavier init

15. What each one costs: Gradient norms reveal where learning stops

Trade off

Comparison matrix

From Gradient norms reveal where learning stops: every row here is a choice with a cost. Fill the grad norm pattern column, then say which row you would actually pick and what you give up for it.

symptomgrad norm patternlikely cause
vanishingnear 0 in early layers, normal in lastdeep Sigmoid/Tanh, bad init
exploding>1e3 anywhereno gradient clipping, bad init
healthysame order of magnitude throughoutReLU/BatchNorm, He/Xavier init

16. What has to be given first: Healthy vs vanishing gradient norms

Missing information

Discussion prompt

Two architectures: 2-layer ReLU (healthy) vs 3-layer Sigmoid (vanishing). Same data, one backward pass.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

Each Sigmoid gate saturates and multiplies the gradient by σ'≤0.25. Three such gates multiply the early gradient by ≤0.016, matching the observed ~50× shrinkage.

17. Healthy vs vanishing gradient norms

Worked example

Two architectures: 2-layer ReLU (healthy) vs 3-layer Sigmoid (vanishing). Same data, one backward pass.

import torch, torch.nn as nn
torch.manual_seed(0)
X = torch.randn(8, 64)
y = torch.tensor([0,1,2,3,0,1,2,3])
lossfn = nn.CrossEntropyLoss()

# Healthy: ReLU 2-layer
class ReLUNet(nn.Module):
    def __init__(self):
        super().__init__()
        self.fc1 = nn.Linear(64,32)
        self.fc2 = nn.Linear(32,4)
    def forward(self,x): return self.fc2(torch.relu(self.fc1(x)))

# Broken: Sigmoid 3-layer (vanishing gradients)
class SigNet(nn.Module):
    def __init__(self):
        super().__init__()
        self.layers = nn.Sequential(
            nn.Linear(64,64), nn.Sigmoid(),
            nn.Linear(64,64), nn.Sigmoid(),
            nn.Linear(64,64), nn.Sigmoid(),
            nn.Linear(64,4))
    def forward(self,x): return self.layers(x)

for name, net in [('ReLU', ReLUNet()), ('Sigmoid', SigNet())]:
    torch.manual_seed(0); net = net.__class__()
    lossfn(net(X), y).backward()
    print(f'--- {name} ---')
    for n, p in net.named_parameters():
        if p.grad is not None:
            print(f'  {n}: {p.grad.norm():.6f}')
architecturelayergrad norm
ReLU (healthy)fc1.weight0.931085
ReLU (healthy)fc2.weight0.527611
Sigmoid (broken)layers.0.weight0.006976
Sigmoid (broken)layers.2.weight0.010387
Sigmoid (broken)layers.4.weight0.060878
Sigmoid (broken)layers.6.weight0.366669

Sigmoid early layers: norms ~0.007–0.011, 50× smaller than the final layer; ReLU layers: norms stay in the same order throughout

Why: Each Sigmoid gate saturates and multiplies the gradient by σ'≤0.25. Three such gates multiply the early gradient by ≤0.016, matching the observed ~50× shrinkage.

18. Work backwards from the answer: Healthy vs vanishing gradient norms

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Sigmoid early layers: norms ~0.007–0.011, 50× smaller than the final layer; ReLU layers: norms stay in the same order throughout

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Two architectures: 2-layer ReLU (healthy) vs 3-layer Sigmoid (vanishing). Same data, one backward pass.

19. Step 3: Check loss scale

Section

Section 3 of 5

20. Random-classifier baseline: loss = log(K)

Concept

A randomly initialized network (near-zero weights → uniform logits → uniform softmax) should output probability 1/K for every class. The expected cross-entropy is therefore log(K).

\[ \mathcal{L}_\text{init} \approx -\log\frac{1}{K} = \log K \]

K (classes)expected initial CE loss
20.6931
41.3863
51.6094
102.3026
1004.6052

If initial loss is far above log(K): bad weight initialization (huge weights push logits to extremes, causing large CE). If far below: something is leaking label information into initialization.

21. Watch it run: Random-classifier baseline: loss = log(K)

Pattern

Step through it

Step through Random-classifier baseline: loss = log(K) one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: K (classes) is 2
  2. Step 2: K (classes) is 4
  3. Step 3: K (classes) is 5
  4. Step 4: K (classes) is 10
  5. Step 5: K (classes) is 100

22. Guess the shape of the answer: Verifying initial loss — near-zero init vs…

Estimation

Predict first

K=10, 100 samples. Compare near-zero init (standard) vs huge init (σ=10.0).

Commit before you compute: what does Verifying initial loss — near-zero init vs bad init come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Standard init → loss 2.298 ≈ log(10) = 2.303; huge init → loss 5051 (>1000× too high)

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Large random weights produce extreme logits; softmax saturates; CE can be enormous.

23. Verifying initial loss — near-zero init vs bad init

Worked example

K=10, 100 samples. Compare near-zero init (standard) vs huge init (σ=10.0).

import torch, torch.nn as nn, numpy as np
torch.manual_seed(42)

K = 10
X = torch.randn(100, 64)
y = torch.randint(0, K, (100,))

def check_init_loss(model, X, y, K):
    model.eval()
    with torch.no_grad():
        loss = nn.CrossEntropyLoss()(model(X), y).item()
        acc  = (model(X).argmax(1) == y).float().mean().item()
    expected = np.log(K)
    ok = abs(loss - expected) < 0.5
    return {'loss': round(loss,4), 'expected': round(expected,4),
            'acc': round(acc,4), 'expected_acc': round(1/K,4), 'ok': ok}

net_good = nn.Linear(64, K)          # default init (~near-zero)
net_bad  = nn.Linear(64, K)
with torch.no_grad():                 # huge init
    nn.init.normal_(net_bad.weight, 0, 10.0)
    nn.init.normal_(net_bad.bias,   0, 10.0)

for name, net in [('good init', net_good), ('bad init', net_bad)]:
    r = check_init_loss(net, X, y, K)
    print(f'{name}: loss={r["loss"]} expected={r["expected"]} ok={r["ok"]}')
initactual lossexpected (log 10)ok?
near-zero (standard)2.29802.3026True
huge (σ=10)5051.482.3026False

Standard init → loss 2.298 ≈ log(10) = 2.303; huge init → loss 5051 (>1000× too high)

Why: Large random weights produce extreme logits; softmax saturates; CE can be enormous. This is a sign of bad initialization or a forgotten normalization layer.

24. Work backwards from the answer: Verifying initial loss — near-zero init vs…

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Standard init → loss 2.298 ≈ log(10) = 2.303; huge init → loss 5051 (>1000× too high)

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

K=10, 100 samples. Compare near-zero init (standard) vs huge init (σ=10.0).

25. Step 4: Check metrics

Section

Section 4 of 5

26. Initial accuracy baseline: 1/K

Concept

Paired with the loss check: at initialization a random classifier should achieve accuracy ~1/K (random chance). If initial accuracy is already 0.8 on your test set, the model or data pipeline may be broken (label leak, duplicated samples, etc.).

Kexpected accuracy (1/K)expected CE loss (log K)
2 (binary)0.5000.693
40.2501.386
100.1002.303
1000.0104.605

Check both together: loss and accuracy at initialization form a self-consistent pair. A discrepancy (correct loss, wrong accuracy) points to an output shape bug or a metric computation error.

27. Watch it run: Initial accuracy baseline: 1/K

Pattern

Step through it

Step through Initial accuracy baseline: 1/K one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: K is 2 (binary)
  2. Step 2: K is 4
  3. Step 3: K is 10
  4. Step 4: K is 100

28. Something is wrong here: wrong loss function (MSE on class indices)

Anomaly

Predict first

A student writes this, and it looks reasonable:

4-class classification. Use MSELoss(logits, y) where y contains class indices 0–3 — training loss goes down, so it must be learning.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: MSE minimizes numeric distance to class indices, not probability calibration.

For K-class classification: use CrossEntropyLoss(logits, class_indices).

Why: MSE minimizes numeric distance to class indices, not probability calibration. The gradients point in a meaningless direction. Loss falling is not evidence of correct learning.

29. Trap: wrong loss function (MSE on class indices)

Trap

The trap

4-class classification. Use MSELoss(logits, y) where y contains class indices 0–3 — training loss goes down, so it must be learning.

Train with MSELoss(output, class_index_tensor)

Why: MSE minimizes numeric distance to class indices, not probability calibration. The gradients point in a meaningless direction. Loss falling is not evidence of correct learning.

The fix

For K-class classification: use CrossEntropyLoss(logits, class_indices).

Check initial CE loss ≈ log(K) and initial acc ≈ 1/K to confirm the loss is set up correctly

Why: CrossEntropyLoss expects raw logits and integer labels, produces calibrated probabilities, and obeys the log(K) baseline at init — the canary that confirms the loss is wired right.

30. Which of these survive contact with Lesson 74: Neural Network Debugging Checklist?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Before tuning anything, verify your model can learn. A healthy network should drive training loss to ~0 on a batch of 8–32 samples within a few hundred steps.; After one backward pass, compute param.grad.norm() for every named layer. A healthy network shows similar-order norms throughout the stack.; Hyperparameter tuning is the last step, not the first. Use proven defaults from the literature while debugging — they eliminate hyperparameters as a confound.
Breaks
Training loss is stuck at 1.5 after 1000 steps. The first move: lower the learning rate.; 4-class classification. Use MSELoss(logits, y) where y contains class indices 0–3 — training loss goes down, so it must be learning.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 74: Neural Network Debugging Checklist puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

31. Step 5: Known hyperparameters first

Section

Section 5 of 5

32. Start from literature defaults, not a random search

Concept

Hyperparameter tuning is the last step, not the first. Use proven defaults from the literature while debugging — they eliminate hyperparameters as a confound.

settingreliable default
optimizerAdam with lr=1e-3 (SGD with lr=0.1 for vision)
weight initHe (ReLU nets) / Xavier (tanh/sigmoid nets)
batch size32–256 for most tasks
lossCrossEntropyLoss (classification) / MSELoss (regression)
activationsReLU or GELU in hidden layers; no activation at output

Once the model trains correctly with defaults, then tune. Changing hyperparameters on a broken model gives you noise, not signal.

33. Fill in: reliable default for Start from literature defaults, not a random…

Comparison

Comparison matrix

From Start from literature defaults, not a random search: refill the reliable default column from what you know. The rest of the table is as it appeared.

settingreliable default
optimizerAdam with lr=1e-3 (SGD with lr=0.1 for vision)
weight initHe (ReLU nets) / Xavier (tanh/sigmoid nets)
batch size32–256 for most tasks
lossCrossEntropyLoss (classification) / MSELoss (regression)
activationsReLU or GELU in hidden layers; no activation at output

34. What has to be given first: The automated debugging checklist function

Missing information

Discussion prompt

Encode all five checks as a single Python function. Run it before every new experiment.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

The checklist isolates architectural bugs from tuning problems. Every flag that fails points to a specific fix (bad init, wrong loss, dead activation) rather than 'tune the lr'.

35. The automated debugging checklist function

Worked example

Encode all five checks as a single Python function. Run it before every new experiment.

import torch, torch.nn as nn, torch.optim as optim, numpy as np

def debug_checklist(model, X_batch, y_batch, X_full, y_full, K, lr=1e-2, steps=100):
    results = {}

    # Step 3 & 4: initial loss and accuracy
    model.eval()
    with torch.no_grad():
        logits = model(X_full)
        results['init_loss'] = round(nn.CrossEntropyLoss()(logits, y_full).item(), 4)
        results['expected_loss'] = round(float(np.log(K)), 4)
        results['init_acc']  = round((logits.argmax(1)==y_full).float().mean().item(), 4)
        results['loss_ok']   = abs(results['init_loss'] - results['expected_loss']) < 0.5

    # Step 2: gradient norms
    model.train()
    opt = optim.SGD(model.parameters(), lr=0.01)
    opt.zero_grad()
    nn.CrossEntropyLoss()(model(X_batch), y_batch).backward()
    norms = {n: round(p.grad.norm().item(),6)
             for n,p in model.named_parameters() if p.grad is not None}
    results['grad_norms']  = norms
    results['vanishing']   = any(v < 1e-4 for v in norms.values())

    # Step 1: overfit a single batch
    opt2 = optim.Adam(model.parameters(), lr=lr)
    for _ in range(steps):
        opt2.zero_grad()
        nn.CrossEntropyLoss()(model(X_batch), y_batch).backward(); opt2.step()
    with torch.no_grad():
        final_loss = nn.CrossEntropyLoss()(model(X_batch), y_batch).item()
    results['overfit_loss'] = round(final_loss, 6)
    results['can_overfit']  = final_loss < 0.01
    return results
checkkey resulthealthy?
init_loss vs log(K=4)1.397 vs 1.386True
init_acc vs 1/40.180 vs 0.250True (within 0.15)
vanishing gradsall norms > 0.028False (no vanishing)
can_overfit (ep=100)0.000014True

All four flags healthy — the network is structurally sound; proceed to tuning

Why: The checklist isolates architectural bugs from tuning problems. Every flag that fails points to a specific fix (bad init, wrong loss, dead activation) rather than 'tune the lr'.

36. Which is which, by healthy?

Discrimination

Sort into buckets

Sort these by healthy?, from memory, without looking back at The automated debugging checklist function. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

True
init_loss vs log(K=4); can_overfit (ep=100)
True (within 0.15)
init_acc vs 1/4
False (no vanishing)
vanishing grads
g1
healthy? is "True" for init_loss vs log(K=4), can_overfit (ep=100) — that is what the table on "The automated debugging checklist…" records, and it is the single property separating this group from the rest.
g2
healthy? is "True (within 0.15)" for init_acc vs 1/4 — that is what the table on "The automated debugging checklist…" records, and it is the single property separating this group from the rest.
g3
healthy? is "False (no vanishing)" for vanishing grads — that is what the table on "The automated debugging checklist…" records, and it is the single property separating this group from the rest.

37. Rebuild the recipe: The 5-step debugging protocol

Ranking

Put in order

These are the steps of The 5-step debugging protocol, scrambled. Put them back in order before the next slide shows you.

  1. Overfit a single batch — if loss can't reach ~0, fix architecture before anything else
  2. Check gradient norms per layer — near-zero early norms = vanishing; astronomic = exploding
  3. Check initial loss against log(K) — far above log(K) = bad init or missing normalization
  4. Check initial accuracy against 1/K — inconsistent pair = shape bug or label leak
  5. Use literature defaults first, then tune — hyperparameter search on a broken model yields noise

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

38. The 5-step debugging protocol

Pattern

  1. Overfit a single batch — if loss can't reach ~0, fix architecture before anything else
  2. Check gradient norms per layer — near-zero early norms = vanishing; astronomic = exploding
  3. Check initial loss against log(K) — far above log(K) = bad init or missing normalization
  4. Check initial accuracy against 1/K — inconsistent pair = shape bug or label leak
  5. Use literature defaults first, then tune — hyperparameter search on a broken model yields noise
steptoolhealthy signal
1 overfitAdam on 8 samples, 100 stepsloss → ~0, acc → 1.0
2 grad flowparam.grad.norm() per layersame order throughout stack
3 loss scaleCE at init vs log(K)within ±0.5 of log(K)
4 metricaccuracy at init vs 1/Kwithin ±0.15 of 1/K
5 defaultslr=1e-3 Adam, He init, ReLUtrains correctly; then tune

39. Where does it stop working: The 5-step debugging protocol

Edge cases

Discussion prompt

The 5-step debugging protocol works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Overfit a single batch — if loss can't reach ~0, fix architecture before anything else
  2. Check gradient norms per layer — near-zero early norms = vanishing; astronomic = exploding
  3. Check initial loss against log(K) — far above log(K) = bad init or missing normalization
  4. Check initial accuracy against 1/K — inconsistent pair = shape bug or label leak
  5. Use literature defaults first, then tune — hyperparameter search on a broken model yields noise

40. Rule out three: Check yourself — initial loss baseline

Elimination

Eliminate the wrong options

Your 10-class classifier shows initial cross-entropy loss of 4.8. What does this most likely indicate?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Bad weight initialization — logits are far from uniform, pushing loss well above log(10) = 2.30
  • B. The model is already learning fast — high initial loss means strong early gradients
  • C. Correct behavior — cross-entropy for 10 classes is always around 4–5
  • D. The learning rate is too high, causing the optimizer to overshoot before the first evaluation

Survives elimination: A

Why: A randomly initialized network with near-zero weights produces uniform logits → uniform softmax → CE ≈ log(10) = 2.303. A loss of 4.8 (≈ log(121)) means logits are highly non-uniform at init, which points to large random weights (bad initialization) or a forgotten input normalization.

41. Check yourself — initial loss baseline

Check

Work it out before clicking.

Check your understanding

Your 10-class classifier shows initial cross-entropy loss of 4.8. What does this most likely indicate?

  • A. Bad weight initialization — logits are far from uniform, pushing loss well above log(10) = 2.30 (correct)
  • B. The model is already learning fast — high initial loss means strong early gradients
  • C. Correct behavior — cross-entropy for 10 classes is always around 4–5
  • D. The learning rate is too high, causing the optimizer to overshoot before the first evaluation

Answer: A

Why: A randomly initialized network with near-zero weights produces uniform logits → uniform softmax → CE ≈ log(10) = 2.303. A loss of 4.8 (≈ log(121)) means logits are highly non-uniform at init, which points to large random weights (bad initialization) or a forgotten input normalization.

Why B tempts people
High initial loss does produce large gradients, but that is a consequence of the bug, not a sign of healthy learning. The network hasn't taken a step yet — there is no 'learning fast'.
Why C tempts people
The correct baseline for K=10 is log(10) = 2.303, not 4–5. An answer of 4.8 is more than 2× above the expected value and is a clear red flag.
Why D tempts people
Initial loss is measured before any optimizer step, so the learning rate cannot have affected it yet.

42. Answer it before you see the options: Check yourself — gradient flow

Prediction

Predict first

After one backward pass, grad norms are: layer1=0.007, layer2=0.009, layer3=0.058, layer4=0.37. What is the most likely cause and fix?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: Vanishing gradients — early layers barely update; switch to ReLU or add batch normalization

Why: Norms of 0.007–0.009 in early layers vs 0.37 in the last layer indicate vanishing gradients — each Sigmoid gate multiplies the backward signal by σ'≤0.25, shrinking it ~50× across three layers. Early layers barely learn. Fix: replace Sigmoid with ReLU, or add BatchNorm to keep activations in the linear regime.

43. Check yourself — gradient flow

Check

Read the gradient norms and diagnose.

Check your understanding

After one backward pass, grad norms are: layer1=0.007, layer2=0.009, layer3=0.058, layer4=0.37. What is the most likely cause and fix?

  • A. Vanishing gradients — early layers barely update; switch to ReLU or add batch normalization (correct)
  • B. Exploding gradients — norms are too large; add gradient clipping
  • C. Normal behavior — gradients are always smaller in earlier layers
  • D. The learning rate is too low — raise lr to increase gradient magnitude

Answer: A

Why: Norms of 0.007–0.009 in early layers vs 0.37 in the last layer indicate vanishing gradients — each Sigmoid gate multiplies the backward signal by σ'≤0.25, shrinking it ~50× across three layers. Early layers barely learn. Fix: replace Sigmoid with ReLU, or add BatchNorm to keep activations in the linear regime.

Why B tempts people
Exploding gradients would show norms >> 1 (often 1e3–1e6 and NaN loss). Values of 0.007–0.37 are uniformly small — the problem is vanishing, not exploding.
Why C tempts people
Gradients being slightly smaller in earlier layers is normal due to the chain rule, but a 50× ratio (0.007 vs 0.37) is pathological, not typical. A healthy ReLU net shows same-order norms throughout.
Why D tempts people
Gradient magnitude is set by the loss landscape and architecture, not by learning rate. Raising lr scales the update step but doesn't change the gradient norm computed by backward().

44. Rule out three: Check yourself — bug taxonomy

Elimination

Eliminate the wrong options

A network's training loss decreases smoothly but validation accuracy stays at chance (1/K) throughout training. Which bug is most consistent with this pattern?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Label leak in the training set — the model memorizes train labels but the validation labels are correct and unseen
  • B. Vanishing gradients — early layers don't update, so the network can't generalize
  • C. Missing ReLU activation — without nonlinearity the network has no expressive power
  • D. Wrong loss function (MSE instead of CE) — the model trains on a corrupted signal

Survives elimination: A

Why: Decreasing training loss with validation accuracy stuck at 1/K is the signature of a train/validation data pipeline bug. Most commonly: train and validation sets share samples (so the model memorizes train labels), or train labels were accidentally permuted before splitting while validation retains the real labels. The model learns something on train but that information doesn't transfer to clean validation data.

45. Check yourself — bug taxonomy

Check

Match symptom to bug.

Check your understanding

A network's training loss decreases smoothly but validation accuracy stays at chance (1/K) throughout training. Which bug is most consistent with this pattern?

  • A. Label leak in the training set — the model memorizes train labels but the validation labels are correct and unseen (correct)
  • B. Vanishing gradients — early layers don't update, so the network can't generalize
  • C. Missing ReLU activation — without nonlinearity the network has no expressive power
  • D. Wrong loss function (MSE instead of CE) — the model trains on a corrupted signal

Answer: A

Why: Decreasing training loss with validation accuracy stuck at 1/K is the signature of a train/validation data pipeline bug. Most commonly: train and validation sets share samples (so the model memorizes train labels), or train labels were accidentally permuted before splitting while validation retains the real labels. The model learns something on train but that information doesn't transfer to clean validation data.

Why B tempts people
Vanishing gradients would cause training loss to also plateau (early layers don't update → limited model capacity → training accuracy also stagnates). A smoothly decreasing training loss rules out severe gradient problems.
Why C tempts people
A linear-only network (no activation) still has representational power for linearly separable data. It would typically show both training and validation loss decreasing, just more slowly — not validation stuck at 1/K.
Why D tempts people
MSE on class indices still provides a gradient signal that often correlates with the correct class ordering. Both train and validation accuracy would likely show some above-chance improvement, not validation pinned at 1/K.

46. Your turn: debug it

Section

Project

47. Project: build and stress-test the debugging checklist

Concept

You will (1) implement debug_checklist() from scratch, (2) inject 5 classic bugs into a working network and use your function to find them all.

#milestonetechnique
1Implement debug_checklist()init loss/acc, grad norms, overfit test
2Inject and find wrong loss (MSE→CE)check init_loss vs log(K)
3Inject and find missing activationoverfit test + manual forward trace
4Inject and find bad init (huge σ)init_loss >> log(K) flag
5Inject and find vanishing grad (Sigmoid)grad_norm < 1e-4 in early layers
6Full run: fix all 5 bugs in < 10 mintimed exercise

Build rule: run debug_checklist() after each injected bug and record which flag fires. A bug that doesn't trigger any flag means your checklist needs a new check.

48. By analogy: Project: build and stress-test the debugging checklist

Analogy

Discussion prompt

Explain Project: build and stress-test the debugging checklist by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

You will (1) implement debug_checklist() from scratch, (2) inject 5 classic bugs into a working network and use your function to find them all.

49. Milestone 1 — implement debug_checklist()

Worked example

Your turn: write debug_checklist(model, X_batch, y_batch, X_full, y_full, K) that returns a dict with at least: init_loss, loss_ok, init_acc, grad_norms, vanishing, overfit_loss, can_overfit.

Hint: check init loss/acc before any training step; do one backward to get .grad; then overfit X_batch for 100 steps with Adam lr=1e-2.

import torch, torch.nn as nn, torch.optim as optim, numpy as np
torch.manual_seed(0)

X_b = torch.randn(8, 64)
y_b = torch.tensor([0,1,2,3,0,1,2,3])
X_f = torch.randn(100, 64)
y_f = torch.randint(0, 4, (100,))

class Net(nn.Module):
    def __init__(self):
        super().__init__()
        self.fc1 = nn.Linear(64,32)
        self.fc2 = nn.Linear(32,4)
    def forward(self,x): return self.fc2(torch.relu(self.fc1(x)))

net = Net()
r = debug_checklist(net, X_b, y_b, X_f, y_f, K=4)
print(r)
keyvaluehealthy?
init_loss1.3970True (≈ log 4 = 1.386)
init_acc0.1800True (1/4 = 0.250)
vanishingFalseTrue
can_overfitTrueTrue

50. Milestone 2 — inject and catch bad initialization

Worked example

Your turn: create a net, then overwrite its weights with nn.init.normal_(p, 0, 10.0). Run debug_checklist. Which flag fires?

Hint: init_loss will be far above log(K). The loss_ok flag should become False.

torch.manual_seed(0)
bad_net = Net()
with torch.no_grad():
    for p in bad_net.parameters():
        nn.init.normal_(p, 0, 10.0)   # inject: huge init

r_bad = debug_checklist(bad_net, X_b, y_b, X_f, y_f, K=4)
print('init_loss:', r_bad['init_loss'])
print('loss_ok:  ', r_bad['loss_ok'])
flaghealthy netbad init net
init_loss1.397≫ 1.386 (e.g. 2000+)
loss_okTrueFalse (bug caught!)
can_overfitTruemay still be True

51. What each one costs: Milestone 2 — inject and catch bad initialization

Trade off

Comparison matrix

From Milestone 2 — inject and catch bad initialization: every row here is a choice with a cost. Fill the healthy net column, then say which row you would actually pick and what you give up for it.

flaghealthy netbad init net
init_loss1.397≫ 1.386 (e.g. 2000+)
loss_okTrueFalse (bug caught!)
can_overfitTruemay still be True

52. Guess the shape of the answer: Milestone 3 — inject and catch vanishing…

Estimation

Predict first

Your turn: build a 3-layer Sigmoid network. Run debug_checklist. Which flags fire?

Commit before you compute: what does Milestone 3 — inject and catch vanishing gradients come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Early layers: grad_norm < 0.01, 50× below last layer; vanishing=True flag fires

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Fix: replace nn.Sigmoid() with nn.ReLU() throughout the hidden stack.

53. Milestone 3 — inject and catch vanishing gradients

Worked example

Your turn: build a 3-layer Sigmoid network. Run debug_checklist. Which flags fire?

Hint: vanishing should be True because early-layer grad norms will be < 1e-4 after one backward pass.

class SigNet(nn.Module):
    def __init__(self):
        super().__init__()
        self.layers = nn.Sequential(
            nn.Linear(64,64), nn.Sigmoid(),
            nn.Linear(64,64), nn.Sigmoid(),
            nn.Linear(64,64), nn.Sigmoid(),
            nn.Linear(64,4))
    def forward(self,x): return self.layers(x)

torch.manual_seed(0)
sig_net = SigNet()
r_sig = debug_checklist(sig_net, X_b, y_b, X_f, y_f, K=4)
print('vanishing:', r_sig['vanishing'])
for k,v in r_sig['grad_norms'].items():
    print(f'  {k}: {v}')
layergrad normflag
layers.0.weight (early)0.006976vanishing=True
layers.2.weight (mid)0.010387vanishing=True
layers.4.weight (mid)0.060878borderline
layers.6.weight (last)0.366669healthy

Early layers: grad_norm < 0.01, 50× below last layer; vanishing=True flag fires

Why: Fix: replace nn.Sigmoid() with nn.ReLU() throughout the hidden stack. With ReLU the gradient passes through unsaturated neurons unchanged, keeping early-layer norms in the same order as the last layer.

54. Fill in: grad norm for Milestone 3 — inject and catch vanishing…

Comparison

Comparison matrix

From Milestone 3 — inject and catch vanishing gradients: refill the grad norm column from what you know. The rest of the table is as it appeared.

layergrad normflag
layers.0.weight (early)0.006976vanishing=True
layers.2.weight (mid)0.010387vanishing=True
layers.4.weight (mid)0.060878borderline
layers.6.weight (last)0.366669healthy

55. The full debugging program

Concept

import torch, torch.nn as nn, torch.optim as optim, numpy as np
torch.manual_seed(0)

X_b = torch.randn(8, 64); y_b = torch.tensor([0,1,2,3,0,1,2,3])
X_f = torch.randn(100, 64); y_f = torch.randint(0, 4, (100,))

class Net(nn.Module):
    def __init__(self):
        super().__init__()
        self.fc1=nn.Linear(64,32); self.fc2=nn.Linear(32,4)
    def forward(self,x): return self.fc2(torch.relu(self.fc1(x)))

def debug_checklist(model, X_batch, y_batch, X_full, y_full, K, lr=1e-2, steps=100):
    r = {}
    model.eval()
    with torch.no_grad():
        logits = model(X_full)
        r['init_loss'] = round(nn.CrossEntropyLoss()(logits, y_full).item(),4)
        r['expected_loss'] = round(float(np.log(K)),4)
        r['init_acc']  = round((logits.argmax(1)==y_full).float().mean().item(),4)
        r['loss_ok']   = abs(r['init_loss']-r['expected_loss']) < 0.5
    model.train()
    opt = optim.SGD(model.parameters(), lr=0.01)
    opt.zero_grad()
    nn.CrossEntropyLoss()(model(X_batch),y_batch).backward()
    norms = {n:round(p.grad.norm().item(),6)
             for n,p in model.named_parameters() if p.grad is not None}
    r['grad_norms'] = norms; r['vanishing'] = any(v<1e-4 for v in norms.values())
    opt2 = optim.Adam(model.parameters(), lr=lr)
    for _ in range(steps):
        opt2.zero_grad()
        nn.CrossEntropyLoss()(model(X_batch),y_batch).backward(); opt2.step()
    with torch.no_grad():
        fl = nn.CrossEntropyLoss()(model(X_batch),y_batch).item()
    r['overfit_loss']=round(fl,6); r['can_overfit']=(fl<0.01)
    return r

net = Net()
report = debug_checklist(net, X_b, y_b, X_f, y_f, K=4)
for k,v in report.items(): print(f'{k}: {v}')
flagvalueinterpretation
init_loss / expected_loss1.397 / 1.386initialization OK
loss_okTrueno bad-init bug
vanishingFalseno gradient flow bug
can_overfitTruearchitecture sound — proceed to tuning

All four flags passing means the network is structurally sound. Step 5: now switch to literature hyperparameters and tune from there.

56. Which is which, by value

Discrimination

Sort into buckets

Sort these by value, from memory, without looking back at The full debugging program. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

1.397 / 1.386
init_loss / expected_loss
True
loss_ok; can_overfit
False
vanishing
g1
value is "1.397 / 1.386" for init_loss / expected_loss — that is what the table on "The full debugging program" records, and it is the single property separating this group from the rest.
g2
value is "True" for loss_ok, can_overfit — that is what the table on "The full debugging program" records, and it is the single property separating this group from the rest.
g3
value is "False" for vanishing — that is what the table on "The full debugging program" records, and it is the single property separating this group from the rest.

57. Show it off

Concept

Out loud, slides closed: walk through the five steps in order and explain what each one catches. Then inject one bug of your choice, run the checklist, and narrate which flag fires and why.

Stretch (homework): extend debug_checklist to also (a) plot gradient norms as a bar chart per layer, (b) detect dead ReLU neurons (fraction of zero activations), and (c) flag if training loss increases in the first 5 steps. Next: hyperparameter tuning and learning rate schedules (L75).

58. Break it if you can: Show it off

Counterexample

Discussion prompt

Out loud, slides closed: walk through the five steps in order and explain what each one catches. Then inject one bug of your choice, run the checklist, and narrate which flag fires and why.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

59. Connect it up: Lesson 74: Neural Network Debugging Checklist

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Step 1: Overfit a single batch · Step 2: Gradient flow · Step 3: Check loss scale · Step 4: Check metrics · Step 5: Known hyperparameters first · Your turn: debug it. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

60. What you can do now

Recap

stepwhat it catcheshealthy signal
1 overfit batcharchitecture / loss / data-pipeline bugsloss ~0, acc 1.0 in <100 steps
2 grad normsvanishing / exploding gradientssame order of magnitude layer-to-layer
3 init lossbad init, missing normalizationwithin ±0.5 of log(K)
4 init acclabel leak, shape bugswithin ±0.15 of 1/K
5 defaults firstconfounded hyperparameter searchtrain correctly, then tune

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 74 — Neural Network Debugging — Barron · USAAIO Round 2 Preparation, 2026
  2. All loss/accuracy/gradient-norm values verified with torch 2.7.1+cpu and numpy 2.2.6 — Real execution, June 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108