USAAIO Lesson 42, from Week 15. It covers writing a custom Dataset with __len__ and __getitem__, the DataLoader with its batching, shuffling, and num_workers, manual feature transforms, inspecting an nn.Module's parameters with .parameters() and .named_parameters(), and model serialization - saving a state_dict against saving the whole model, and saving and resuming from a checkpoint. You then build a complete pipeline: a 120-sample tabular dataset, a TwoLayerNet going 4→8→1, trained with Adam, saving a checkpoint at epoch 5, then resuming and reaching a loss of 0.1666 by epoch 10. The lesson runs to 31 slides.
Subject: Machine Learning · 62 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 42 · Week 15
Dataset → DataLoader → transforms → nn.Module internals → state_dict → checkpoint resume. The plumbing every real training run needs.
Objectives
Dataset with __len__ and __getitem__DataLoader for batching, shuffling, and parallel loadingDatasetnn.Module and verify total countstate_dict (preferred) vs full-model save, and resume training from a checkpointWarm-up
Discussion prompt
Before we open Lesson 42: PyTorch Data Pipeline & Model Serialization: without looking back, what was the main idea of Logistic Regression, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
derive the logistic regression loss from scratch and prove the gradient dL/dw = X^T(p−y); implement binary logistic regression two ways (NumPy manual backprop and PyTorch nn.Linear→sigmoid→BCELoss); BCEWithLogitsLoss and why it's numerically preferred; softmax regression for multi-class with CrossEntropyLoss; decision boundary intuition and effect of L2 regularization on boundary sharpness; comparison to linear SVM.
Section
Part 1 of 4
Concept
Any Python object that implements __len__ and __getitem__ is a valid PyTorch Dataset. The DataLoader calls these — nothing else.
__len__(self) → integer number of samples__getitem__(self, idx) → one (input, label) tensor pairtorch.utils.data.Dataset to satisfy the type contractThis single abstraction handles images, tabular CSVs, audio, text, and anything else — the training loop never sees the format.
Counterexample
Discussion prompt
Any Python object that implements __len__ and __getitem__ is a valid PyTorch Dataset. The DataLoader calls these — nothing else.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Estimation
Predict first
Wrap a 120×4 numpy array in a Dataset. Verify len() and the shape of one sample.
Commit before you compute: what does Custom Dataset for tabular data come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Output: 120 torch.Size([4]) torch.Size([1])
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. len() returns 120. ds[0] returns (x_tensor shape [4], y_tensor shape [1]) — the exact shapes nn.Linear(4, …) and MSELoss expect.
Worked example
Wrap a 120×4 numpy array in a Dataset. Verify len() and the shape of one sample.
import torch
from torch.utils.data import Dataset
import numpy as np
class TabularDataset(Dataset):
def __init__(self, X, y):
self.X = torch.tensor(X, dtype=torch.float32)
self.y = torch.tensor(y, dtype=torch.float32).unsqueeze(1)
def __len__(self):
return len(self.X)
def __getitem__(self, idx):
return self.X[idx], self.y[idx]
rng = np.random.default_rng(42)
X_np = rng.standard_normal((120, 4))
y_np = X_np[:, 0] + 0.5*X_np[:, 1] + 0.1*rng.standard_normal(120)
ds = TabularDataset(X_np, y_np)
print(len(ds), ds[0][0].shape, ds[0][1].shape)Output: 120 torch.Size([4]) torch.Size([1])
Why: len() returns 120. ds[0] returns (x_tensor shape [4], y_tensor shape [1]) — the exact shapes nn.Linear(4, …) and MSELoss expect.
| call | result | meaning |
|---|---|---|
| len(ds) | 120 | number of samples |
| ds[0][0].shape | [4] | one feature vector |
| ds[0][1].shape | [1] | one scalar target |
| ds[0][0][:4] | [0.3047, -1.04, 0.7505, 0.9406] | first sample values (rng seed 42) |
Comparison
Comparison matrix
From Custom Dataset for tabular data: refill the result column from what you know. The rest of the table is as it appeared.
| call | result | meaning |
|---|---|---|
| len(ds) | 120 | number of samples |
| ds[0][0].shape | [4] | one feature vector |
| ds[0][1].shape | [1] | one scalar target |
| ds[0][0][:4] | [0.3047, -1.04, 0.7505, 0.9406] | first sample values (rng seed 42) |
Section
Part 2 of 4
Concept
DataLoader calls __getitem__ repeatedly, collates results into batch tensors, and exposes an iterator over the whole dataset (one epoch).
| argument | effect | typical value |
|---|---|---|
| batch_size | samples per batch | 32 or 64 |
| shuffle | re-randomize each epoch | True for train, False for val |
| num_workers | parallel prefetch processes | 0 (Windows/Colab), 4+ (Linux) |
| pin_memory | page-lock CPU RAM → faster GPU transfer | True when CUDA available |
With n=120 and batch_size=32: batches of 32, 32, 32, 24 — the last batch is smaller (drop_last=False by default).
Trade off
Comparison matrix
From DataLoader wraps Dataset for training: every row here is a choice with a cost. Fill the effect column, then say which row you would actually pick and what you give up for it.
| argument | effect | typical value |
|---|---|---|
| batch_size | samples per batch | 32 or 64 |
| shuffle | re-randomize each epoch | True for train, False for val |
| num_workers | parallel prefetch processes | 0 (Windows/Colab), 4+ (Linux) |
| pin_memory | page-lock CPU RAM → faster GPU transfer | True when CUDA available |
Missing information
Discussion prompt
Create a DataLoader on the 120-sample dataset and inspect the batch structure.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
ceil(120/32) = 4. The last mini-batch gets the remaining 24 samples. The training loop sees all 120 samples each epoch regardless.
Worked example
Create a DataLoader on the 120-sample dataset and inspect the batch structure.
from torch.utils.data import DataLoader
loader = DataLoader(ds, batch_size=32, shuffle=True, num_workers=0)
print('batches per epoch:', len(loader))
xb, yb = next(iter(loader))
print('batch X shape:', xb.shape) # (32, 4)
print('batch y shape:', yb.shape) # (32, 1)
batch_sizes = [len(b[0]) for b in loader]
print('all batch sizes:', batch_sizes)4 batches: [32, 32, 32, 24] — 3 full + 1 remainder
Why: ceil(120/32) = 4. The last mini-batch gets the remaining 24 samples. The training loop sees all 120 samples each epoch regardless.
| batch | size | shape (X) | shape (y) |
|---|---|---|---|
| 1 | 32 | (32, 4) | (32, 1) |
| 2 | 32 | (32, 4) | (32, 1) |
| 3 | 32 | (32, 4) | (32, 1) |
| 4 (last) | 24 | (24, 4) | (24, 1) |
Pattern
Step through it
Step through Iterating a DataLoader one row at a time. What is driving the change, and what would the row after the last one be?
Anomaly
Predict first
A student writes this, and it looks reasonable:
Use shuffle=True on both train and validation DataLoaders — consistent is cleaner.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Shuffling the validation set changes nothing about the loss/accuracy (every sample is seen), but it invalidates deterministic debugging: the 'same' epoch-2 batch now contains different samples, making loss curves hard…
Only shuffle the training loader; keep shuffle=False on validation and test.
Why: Shuffling the validation set changes nothing about the loss/accuracy (every sample is seen), but it invalidates deterministic debugging: the 'same' epoch-2 batch now contains different samples, making loss curves hard to reproduce.
Trap
Use shuffle=True on both train and validation DataLoaders — consistent is cleaner.
DataLoader(val_ds, batch_size=64, shuffle=True)
Why: Shuffling the validation set changes nothing about the loss/accuracy (every sample is seen), but it invalidates deterministic debugging: the 'same' epoch-2 batch now contains different samples, making loss curves hard to reproduce.
Only shuffle the training loader; keep shuffle=False on validation and test.
DataLoader(train_ds, shuffle=True) / DataLoader(val_ds, shuffle=False)
Why: Training benefits from randomized order (escapes local correlations in minibatches). Validation needs deterministic, reproducible ordering for fair comparison across runs.
Break the constraint
Discussion prompt
The rule this trap just fixed:
Training benefits from randomized order (escapes local correlations in minibatches). Validation needs deterministic, reproducible ordering for fair comparison across runs.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
Shuffling the validation set changes nothing about the loss/accuracy (every sample is seen), but it invalidates deterministic debugging: the 'same' epoch-2 batch now contains different samples, making loss curves hard to reproduce.
Section
Part 3 of 4
Concept
Transforms pre-process each sample inside __getitem__. They are callables applied before returning a tensor. torchvision.transforms ships image ops; for tabular data you write your own.
__init__ on the full array__getitem__ applies the transform before returning each sampleThis keeps the transform logic co-located with the data, invisible to the training loop — same interface whether you add augmentation or not.
Analogy
Discussion prompt
Explain Where transforms live by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Transforms pre-process each sample inside __getitem__. They are callables applied before returning a tensor. torchvision.transforms ships image ops; for tabular data you write your own.
Pattern
Predict first
The table runs: feat 0 | ≈ 0.0 | 1.0 · feat 1 | ≈ 0.0 | 1.0 · feat 2 | ≈ 0.0 | 1.0
In NormalizingDataset — transform in __init__, given the rows so far: what is the next one — the row where feature is feat 3?
Correct: feat 3 | ≈ 0.0 | 1.0
| feature | mean after norm | std after norm |
|---|---|---|
| feat 0 | ≈ 0.0 | 1.0 |
| feat 1 | ≈ 0.0 | 1.0 |
| feat 2 | ≈ 0.0 | 1.0 |
| feat 3 | ≈ 0.0 | 1.0 |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. z-score normalization (Lesson 8 feature scaling) ensures all input dimensions compete fairly during gradient descent.
Worked example
Compute z-score statistics on construction and apply them on every __getitem__ call. Verify the output statistics.
class NormalizingDataset(Dataset):
def __init__(self, X, y):
X_t = torch.tensor(X, dtype=torch.float32)
self.mean = X_t.mean(0)
self.std = X_t.std(0) + 1e-8
self.X = (X_t - self.mean) / self.std # normalize all at once
self.y = torch.tensor(y, dtype=torch.float32).unsqueeze(1)
def __len__(self): return len(self.X)
def __getitem__(self, i): return self.X[i], self.y[i]
ds_norm = NormalizingDataset(X_np, y_np)
print(ds_norm.X.mean(0).numpy().round(4)) # [-0. 0. -0. 0.]
print(ds_norm.X.std(0).numpy().round(4)) # [1. 1. 1. 1.]Feature means ≈ [-0.0, 0.0, -0.0, 0.0], stds ≈ [1.0, 1.0, 1.0, 1.0]
Why: z-score normalization (Lesson 8 feature scaling) ensures all input dimensions compete fairly during gradient descent. The training loop code is unchanged — the Dataset handles it.
| feature | mean after norm | std after norm |
|---|---|---|
| feat 0 | ≈ 0.0 | 1.0 |
| feat 1 | ≈ 0.0 | 1.0 |
| feat 2 | ≈ 0.0 | 1.0 |
| feat 3 | ≈ 0.0 | 1.0 |
Invariant
Step through it
Step through NormalizingDataset — transform in __init__ one row at a time. One of these columns never changes — find it, and say why it cannot.
Section
Part 4 of 4
Concept
When you assign an nn.Linear (or any nn.Module) as an attribute of another Module, PyTorch registers its Parameter tensors automatically.
.parameters() → flat iterator over every Parameter (used by the optimizer).named_parameters() → same with names like 'fc1.weight', 'fc1.bias'p.numel() → number of scalars in parameter tensor psum(p.numel() for p in model.parameters()) → total trainable scalar countAlways verify parameter counts before training. A wrong architecture (mismatched dims) shows up here before the forward pass crashes.
Explain it
Discussion prompt
Explain How nn.Module tracks parameters to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
When you assign an nn.Linear (or any nn.Module) as an attribute of another Module, PyTorch registers its Parameter tensors automatically.
Estimation
Predict first
Build TwoLayerNet(4, 8, 1) and count every trainable parameter. Predict the total before running.
Commit before you compute: what does Counting parameters in TwoLayerNet come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: fc1.weight (8×4)=32, fc1.bias (8)=8, fc2.weight (1×8)=8, fc2.bias (1)=1 → total 49
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. nn.Linear(in, out) has (in × out) weights plus out biases.
Worked example
Build TwoLayerNet(4, 8, 1) and count every trainable parameter. Predict the total before running.
import torch.nn as nn
class TwoLayerNet(nn.Module):
def __init__(self, in_f, h, out_f):
super().__init__()
self.fc1 = nn.Linear(in_f, h)
self.relu = nn.ReLU()
self.fc2 = nn.Linear(h, out_f)
def forward(self, x):
return self.fc2(self.relu(self.fc1(x)))
net = TwoLayerNet(4, 8, 1)
for name, p in net.named_parameters():
print(f'{name}: {tuple(p.shape)} = {p.numel()}')
print('Total:', sum(p.numel() for p in net.parameters()))fc1.weight (8×4)=32, fc1.bias (8)=8, fc2.weight (1×8)=8, fc2.bias (1)=1 → total 49
Why: nn.Linear(in, out) has (in × out) weights plus out biases. Summing across both layers: 32+8+8+1=49. nn.ReLU has no learnable parameters.
| tensor | shape | numel |
|---|---|---|
| fc1.weight | (8, 4) | 32 |
| fc1.bias | (8,) | 8 |
| fc2.weight | (1, 8) | 8 |
| fc2.bias | (1,) | 1 |
| total | — | 49 |
Pattern
Step through it
Step through Counting parameters in TwoLayerNet one row at a time. What is driving the change, and what would the row after the last one be?
Anomaly
Predict first
A student writes this, and it looks reasonable:
Save the entire model: torch.save(model, 'model.pt') — simple, one line, loads back with torch.load.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: The full-model save pickles the Python class name.
Save only the state_dict — a dict of tensor names to tensors, independent of Python class structure.
Why: The full-model save pickles the Python class name. If you rename or move TwoLayerNet, or load in a different environment, unpickling fails with an AttributeError. Resume training is also fragile.
Trap
Save the entire model: torch.save(model, 'model.pt') — simple, one line, loads back with torch.load.
torch.save(model, 'full.pt') # 3583 bytes, requires the class definition at load time
Why: The full-model save pickles the Python class name. If you rename or move TwoLayerNet, or load in a different environment, unpickling fails with an AttributeError. Resume training is also fragile.
Save only the state_dict — a dict of tensor names to tensors, independent of Python class structure.
torch.save(model.state_dict(), 'sd.pt') # 2525 bytes; load with model.load_state_dict(...)
Why: state_dict is portable: the class definition lives in your code, the file holds only weights. Standard practice for checkpoints, deployment, and sharing. 33% smaller file vs full save.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
__len__ and __getitem__ is a valid PyTorch Dataset. The DataLoader calls these — nothing else.; DataLoader calls __getitem__ repeatedly, collates results into batch tensors, and exposes an iterator over the whole dataset (one epoch).; When you assign an nn.Linear (or any nn.Module) as an attribute of another Module, PyTorch registers its Parameter tensors automatically.shuffle=True on both train and validation DataLoaders — consistent is cleaner.; Save the entire model: torch.save(model, 'model.pt') — simple, one line, loads back with torch.load.Missing information
Discussion prompt
Save model.state_dict(), create a fresh TwoLayerNet, call load_state_dict(), and verify all weights are identical.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
The four tensors are restored exactly. weights_only=True (PyTorch 2+ default) prevents arbitrary code execution during unpickling — always pass it when loading a state_dict.
Worked example
Save model.state_dict(), create a fresh TwoLayerNet, call load_state_dict(), and verify all weights are identical.
import torch, os, tempfile
tmpdir = tempfile.mkdtemp()
sd_path = os.path.join(tmpdir, 'sd.pt')
torch.save(net.state_dict(), sd_path)
restored = TwoLayerNet(4, 8, 1)
restored.load_state_dict(torch.load(sd_path, weights_only=True))
restored.eval()
match = all(torch.allclose(a, b)
for a, b in zip(net.parameters(), restored.parameters()))
print('keys :', list(net.state_dict().keys()))
print('match:', match)keys: ['fc1.weight', 'fc1.bias', 'fc2.weight', 'fc2.bias'] match: True
Why: The four tensors are restored exactly. weights_only=True (PyTorch 2+ default) prevents arbitrary code execution during unpickling — always pass it when loading a state_dict.
| save style | file size | portable? | weights_only |
|---|---|---|---|
| state_dict only | 2525 bytes | yes | True (safe) |
| full model | 3583 bytes | fragile | False needed |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
keys: ['fc1.weight', 'fc1.bias', 'fc2.weight', 'fc2.bias'] match: True
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Save model.state_dict(), create a fresh TwoLayerNet, call load_state_dict(), and verify all weights are identical.
Concept
To resume training you need more than just model weights — the optimizer's momentum buffers matter too. Pack all state into one dict.
'epoch' — last completed epoch (resume from epoch+1)'model_state_dict' — model weights'optimizer_state_dict' — momentum buffers, adaptive LR accumulators'loss' — last recorded validation loss for auditingMissing the optimizer state is the silent bug: the model weights are right but momentum is cold-started, so the first few resumed epochs underfit.
Analogy
Discussion prompt
Explain Checkpoint dict: model + optimizer + epoch by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
To resume training you need more than just model weights — the optimizer's momentum buffers matter too. Pack all state into one dict.
Estimation
Predict first
Train 5 epochs, save a checkpoint, restore and continue epochs 6-10. Verify the loss trajectory is identical to a fresh 10-epoch run.
Commit before you compute: what does Checkpoint save and resume come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Resumed loss at epoch 6: 0.5569 → epoch 10: 0.1666 — identical to uninterrupted run
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Because we restored opt_state, Adam's per-parameter accumulators are warm.
Worked example
Train 5 epochs, save a checkpoint, restore and continue epochs 6-10. Verify the loss trajectory is identical to a fresh 10-epoch run.
model = TwoLayerNet(4, 8, 1); opt = torch.optim.Adam(model.parameters(), lr=1e-2)
loader = DataLoader(ds_norm, batch_size=32, shuffle=False, num_workers=0)
loss_fn = nn.MSELoss()
for epoch in range(5):
for xb, yb in loader:
opt.zero_grad(); loss_fn(model(xb), yb).backward(); opt.step()
torch.save({'epoch': 4,
'model_state': model.state_dict(),
'opt_state': opt.state_dict(),
'loss': 0.6849}, ckpt_path)
# --- resume ---
m2 = TwoLayerNet(4, 8, 1); o2 = torch.optim.Adam(m2.parameters(), lr=1e-2)
ck = torch.load(ckpt_path, weights_only=False)
m2.load_state_dict(ck['model_state']); o2.load_state_dict(ck['opt_state'])
for epoch in range(ck['epoch']+1, 10):
for xb, yb in loader:
o2.zero_grad(); loss_fn(m2(xb), yb).backward(); o2.step()Resumed loss at epoch 6: 0.5569 → epoch 10: 0.1666 — identical to uninterrupted run
Why: Because we restored opt_state, Adam's per-parameter accumulators are warm. The resumed loss curve is bit-for-bit identical to training straight through — the checkpoint is transparent.
| epoch | loss (no ckpt) | loss (resumed) |
|---|---|---|
| 1 | 1.1912 | 1.1912 |
| 5 | 0.6849 | 0.6849 → ckpt |
| 6 | 0.5569 | 0.5569 |
| 8 | 0.3118 | 0.3118 |
| 10 | 0.1666 | 0.1666 |
Comparison
Comparison matrix
From Checkpoint save and resume: refill the loss (no ckpt) column from what you know. The rest of the table is as it appeared.
| epoch | loss (no ckpt) | loss (resumed) |
|---|---|---|
| 1 | 1.1912 | 1.1912 |
| 5 | 0.6849 | 0.6849 → ckpt |
| 6 | 0.5569 | 0.5569 |
| 8 | 0.3118 | 0.3118 |
| 10 | 0.1666 | 0.1666 |
Ranking
Put in order
These are the steps of The data-pipeline + serialization recipe, scrambled. Put them back in order before the next slide shows you.
Dataset, implement __len__ and __getitem__; apply transforms in __init__ or __getitem__shuffle=True (train only); num_workers=0 (Windows), 4+ (Linux); pin_memory=True when using CUDAsum(p.numel() for p in model.parameters()) before first training runstate_dict over full model; pass weights_only=True on load{epoch, model_state_dict, optimizer_state_dict, loss}; restore all four before resumingWhy: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
Dataset, implement __len__ and __getitem__; apply transforms in __init__ or __getitem__shuffle=True (train only); num_workers=0 (Windows), 4+ (Linux); pin_memory=True when using CUDAsum(p.numel() for p in model.parameters()) before first training runstate_dict over full model; pass weights_only=True on load{epoch, model_state_dict, optimizer_state_dict, loss}; restore all four before resumingEdge cases
Discussion prompt
The data-pipeline + serialization recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
Dataset, implement __len__ and __getitem__; apply transforms in __init__ or __getitem__shuffle=True (train only); num_workers=0 (Windows), 4+ (Linux); pin_memory=True when using CUDAsum(p.numel() for p in model.parameters()) before first training runstate_dict over full model; pass weights_only=True on load{epoch, model_state_dict, optimizer_state_dict, loss}; restore all four before resumingElimination
Eliminate the wrong options
A dataset has 120 samples. DataLoader(ds, batch_size=32, drop_last=False). How many batches per epoch and what are their sizes?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: ceil(120/32) = 4 batches. The first three contain 32 samples; the final batch has the remaining 120 − 3×32 = 24. drop_last=False (default) keeps the partial batch.
Check
Predict first, then verify.
Check your understanding
A dataset has 120 samples. DataLoader(ds, batch_size=32, drop_last=False). How many batches per epoch and what are their sizes?
Answer: A
Why: ceil(120/32) = 4 batches. The first three contain 32 samples; the final batch has the remaining 120 − 3×32 = 24. drop_last=False (default) keeps the partial batch.
Check
Do the arithmetic on paper first.
Check your understanding
TwoLayerNet has nn.Linear(4, 8) then nn.Linear(8, 1) (with bias). What is the total trainable parameter count?
Answer: A
Why: nn.Linear(4,8): 4×8=32 weights + 8 biases = 40. nn.Linear(8,1): 8×1=8 weights + 1 bias = 9. Total = 40+9 = 49. Verified: sum(p.numel() for p in net.parameters()) = 49.
Elimination
Eliminate the wrong options
You save only model.state_dict() at epoch 5 and restore it before resuming. Compared to an uninterrupted run, what is most likely wrong?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Adam maintains per-parameter running mean (m) and variance (v) buffers — these live in the optimizer state_dict, not the model state_dict. Without them, Adam cold-starts its accumulators and behaves like the first step of training for several iterations.
Check
Identify the critical missing piece.
Check your understanding
You save only model.state_dict() at epoch 5 and restore it before resuming. Compared to an uninterrupted run, what is most likely wrong?
Answer: A
Why: Adam maintains per-parameter running mean (m) and variance (v) buffers — these live in the optimizer state_dict, not the model state_dict. Without them, Adam cold-starts its accumulators and behaves like the first step of training for several iterations.
Section
Project
Concept
Build a complete, checkpointed training run on synthetic tabular data: custom Dataset with normalization, DataLoader, TwoLayerNet, Adam, checkpoint at epoch 5, resume to epoch 10.
| # | requirement | key tool |
|---|---|---|
| 1 | NormalizingDataset(X, y) | Dataset.__init__, __getitem__ |
| 2 | DataLoader, inspect batch shapes | DataLoader, iter/next |
| 3 | count parameters before training | .named_parameters(), .numel() |
| 4 | train 5 epochs, save checkpoint | torch.save({epoch, model_state, opt_state}) |
| 5 | resume epochs 6-10, verify loss matches | load_state_dict, opt.load_state_dict |
Target: loss ≈ 0.1666 at epoch 10, matching an uninterrupted run. If your resumed loss diverges, check that you saved AND loaded the optimizer state.
Counterexample
Discussion prompt
Build a complete, checkpointed training run on synthetic tabular data: custom Dataset with normalization, DataLoader, TwoLayerNet, Adam, checkpoint at epoch 5, resume to epoch 10.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Target: loss ≈ 0.1666 at epoch 10, matching an uninterrupted run. If your resumed loss diverges, check that you saved AND loaded the optimizer state.
Worked example
Your turn: implement NormalizingDataset and create a DataLoader. Predict how many batches for n=120, batch_size=32.
Hint: compute mean and std in __init__, normalize all at once with (X_t - mean) / std. Add 1e-8 to std to prevent div-by-zero.
ds_norm = NormalizingDataset(X_np, y_np)
loader = DataLoader(ds_norm, batch_size=32, shuffle=False, num_workers=0)
print('batches:', len(loader)) # 4
xb, yb = next(iter(loader))
print('X batch:', xb.shape, 'y batch:', yb.shape)| check | expected | your output |
|---|---|---|
| len(loader) | 4 | 4 |
| xb.shape | (32, 4) | (32, 4) |
| yb.shape | (32, 1) | (32, 1) |
| ds_norm.X.std(0) ≈ | [1.0, 1.0, 1.0, 1.0] | match if normalized |
Worked example
Your turn: instantiate TwoLayerNet(4, 8, 1) and enumerate its parameters. Predict the total count before running.
Hint: nn.Linear(in, out) has in × out weights + out biases. Add up both layers.
net = TwoLayerNet(4, 8, 1)
for name, p in net.named_parameters():
print(f'{name}: {tuple(p.shape)} -> {p.numel()}')
print('Total:', sum(p.numel() for p in net.parameters()))| tensor | shape | count |
|---|---|---|
| fc1.weight | (8, 4) | 32 |
| fc1.bias | (8,) | 8 |
| fc2.weight | (1, 8) | 8 |
| fc2.bias | (1,) | 1 |
| TOTAL | — | 49 |
Pattern
Step through it
Step through Milestone 2 — parameter audit one row at a time. What is driving the change, and what would the row after the last one be?
Worked example
Your turn: train 5 epochs, save {epoch, model_state, opt_state, loss}, then restore and continue epochs 6-10. Predict: does the resumed loss match?
Hint: torch.save(ckpt_dict, path) / torch.load(path, weights_only=False). Restore both model_state and opt_state before continuing.
torch.manual_seed(0)
model = TwoLayerNet(4, 8, 1); opt = torch.optim.Adam(model.parameters(), lr=1e-2)
for epoch in range(5):
for xb, yb in loader:
opt.zero_grad(); loss_fn(model(xb), yb).backward(); opt.step()
torch.save({'epoch':4,'model_state':model.state_dict(),'opt_state':opt.state_dict(),'loss':0.6849}, 'ckpt.pt')
m2 = TwoLayerNet(4, 8, 1); o2 = torch.optim.Adam(m2.parameters(), lr=1e-2)
ck = torch.load('ckpt.pt', weights_only=False)
m2.load_state_dict(ck['model_state']); o2.load_state_dict(ck['opt_state'])
for epoch in range(ck['epoch']+1, 10):
for xb, yb in loader:
o2.zero_grad(); loss_fn(m2(xb), yb).backward(); o2.step()| epoch | expected loss |
|---|---|
| 5 (ckpt saved) | 0.6849 |
| 6 (resumed) | 0.5569 |
| 8 | 0.3118 |
| 10 | 0.1666 |
Trade off
Comparison matrix
From Milestone 3 — checkpoint save and resume: every row here is a choice with a cost. Fill the expected loss column, then say which row you would actually pick and what you give up for it.
| epoch | expected loss |
|---|---|
| 5 (ckpt saved) | 0.6849 |
| 6 (resumed) | 0.5569 |
| 8 | 0.3118 |
| 10 | 0.1666 |
Concept
import torch, torch.nn as nn, numpy as np
from torch.utils.data import Dataset, DataLoader
class NormalizingDataset(Dataset):
def __init__(self, X, y):
X_t = torch.tensor(X, dtype=torch.float32)
self.X = (X_t - X_t.mean(0)) / (X_t.std(0) + 1e-8)
self.y = torch.tensor(y, dtype=torch.float32).unsqueeze(1)
def __len__(self): return len(self.X)
def __getitem__(self, i): return self.X[i], self.y[i]
class TwoLayerNet(nn.Module):
def __init__(self, in_f, h, out_f):
super().__init__()
self.fc1 = nn.Linear(in_f, h); self.relu = nn.ReLU()
self.fc2 = nn.Linear(h, out_f)
def forward(self, x): return self.fc2(self.relu(self.fc1(x)))
rng = np.random.default_rng(42)
X = rng.standard_normal((120, 4))
y = X[:, 0] + 0.5*X[:, 1] + 0.1*rng.standard_normal(120)
ds = NormalizingDataset(X, y)
loader = DataLoader(ds, batch_size=32, shuffle=False)
print('params:', sum(p.numel() for p in TwoLayerNet(4,8,1).parameters())) # 49| output | value |
|---|---|
| Total params | 49 |
| Batches/epoch | 4 ([32,32,32,24]) |
| Loss @ epoch 5 | 0.6849 |
| Loss @ epoch 10 (resumed) | 0.1666 |
If your resumed epoch-10 loss equals the uninterrupted run's — you've built the full pipeline. The Dataset + DataLoader + checkpoint pattern is reused for every model from here through the competition.
Comparison
Comparison matrix
From The full program: refill the value column from what you know. The rest of the table is as it appeared.
| output | value |
|---|---|
| Total params | 49 |
| Batches/epoch | 4 ([32,32,32,24]) |
| Loss @ epoch 5 | 0.6849 |
| Loss @ epoch 10 (resumed) | 0.1666 |
Concept
Out loud, slides closed: (1) explain what __len__ and __getitem__ must return and why the DataLoader needs them; (2) why shuffle=False on a validation loader; (3) what goes wrong if you skip saving optimizer_state_dict.
Stretch (homework): profile num_workers=0 vs num_workers=2 on a larger dataset; implement augmentation in __getitem__ (add random Gaussian noise to features only during training); add a pin_memory=True branch for CUDA. Next up: SVM + MLP Architecture (Lesson 43).
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Dataset: wrapping your data · DataLoader: batching, shuffling, workers · Transforms: normalisation inside Dataset · nn.Module internals + serialization · Your turn: build the pipeline. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
__len__ / __getitem__ and embed transforms.named_parameters() + .numel()state_dict (weights_only=True) — not full model{epoch, model_state, opt_state} and resume without trajectory loss| concept | the one thing to remember |
|---|---|
| Dataset contract | __len__ + __getitem__ — that's all DataLoader needs |
| DataLoader shuffle | True for train, False for val/test |
| parameter count | sum(p.numel() for p in model.parameters()) |
| save preference | state_dict > full model (portable, smaller) |
| checkpoint | save opt state too — momentum is not in model weights |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.