Lesson 42: PyTorch Data Pipeline & Model Serialization

USAAIO Lesson 42, from Week 15. It covers writing a custom Dataset with __len__ and __getitem__, the DataLoader with its batching, shuffling, and num_workers, manual feature transforms, inspecting an nn.Module's parameters with .parameters() and .named_parameters(), and model serialization - saving a state_dict against saving the whole model, and saving and resuming from a checkpoint. You then build a complete pipeline: a 120-sample tabular dataset, a TwoLayerNet going 4→8→1, trained with Adam, saving a checkpoint at epoch 5, then resuming and reaching a loss of 0.1666 by epoch 10. The lesson runs to 31 slides.

Subject: Machine Learning · 62 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. PyTorch Data Pipeline & Model Serialization

Title

USAAIO · Lesson 42 · Week 15

Dataset → DataLoader → transforms → nn.Module internals → state_dict → checkpoint resume. The plumbing every real training run needs.

2. By the end of this lesson you can

Objectives

  1. Implement a custom Dataset with __len__ and __getitem__
  2. Configure a DataLoader for batching, shuffling, and parallel loading
  3. Embed a normalizing transform directly inside a Dataset
  4. Inspect all parameters of an nn.Module and verify total count
  5. Save and load models with state_dict (preferred) vs full-model save, and resume training from a checkpoint

3. What survived from Logistic Regression?

Warm-up

Discussion prompt

Before we open Lesson 42: PyTorch Data Pipeline & Model Serialization: without looking back, what was the main idea of Logistic Regression, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

derive the logistic regression loss from scratch and prove the gradient dL/dw = X^T(p−y); implement binary logistic regression two ways (NumPy manual backprop and PyTorch nn.Linear→sigmoid→BCELoss); BCEWithLogitsLoss and why it's numerically preferred; softmax regression for multi-class with CrossEntropyLoss; decision boundary intuition and effect of L2 regularization on boundary sharpness; comparison to linear SVM.

4. Dataset: wrapping your data

Section

Part 1 of 4

5. The Dataset contract

Concept

Any Python object that implements __len__ and __getitem__ is a valid PyTorch Dataset. The DataLoader calls these — nothing else.

This single abstraction handles images, tabular CSVs, audio, text, and anything else — the training loop never sees the format.

6. Break it if you can: The Dataset contract

Counterexample

Discussion prompt

Any Python object that implements __len__ and __getitem__ is a valid PyTorch Dataset. The DataLoader calls these — nothing else.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

7. Guess the shape of the answer: Custom Dataset for tabular data

Estimation

Predict first

Wrap a 120×4 numpy array in a Dataset. Verify len() and the shape of one sample.

Commit before you compute: what does Custom Dataset for tabular data come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Output: 120 torch.Size([4]) torch.Size([1])

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. len() returns 120. ds[0] returns (x_tensor shape [4], y_tensor shape [1]) — the exact shapes nn.Linear(4, …) and MSELoss expect.

8. Custom Dataset for tabular data

Worked example

Wrap a 120×4 numpy array in a Dataset. Verify len() and the shape of one sample.

import torch
from torch.utils.data import Dataset
import numpy as np

class TabularDataset(Dataset):
    def __init__(self, X, y):
        self.X = torch.tensor(X, dtype=torch.float32)
        self.y = torch.tensor(y, dtype=torch.float32).unsqueeze(1)
    def __len__(self):
        return len(self.X)
    def __getitem__(self, idx):
        return self.X[idx], self.y[idx]

rng = np.random.default_rng(42)
X_np = rng.standard_normal((120, 4))
y_np = X_np[:, 0] + 0.5*X_np[:, 1] + 0.1*rng.standard_normal(120)
ds = TabularDataset(X_np, y_np)
print(len(ds), ds[0][0].shape, ds[0][1].shape)

Output: 120 torch.Size([4]) torch.Size([1])

Why: len() returns 120. ds[0] returns (x_tensor shape [4], y_tensor shape [1]) — the exact shapes nn.Linear(4, …) and MSELoss expect.

callresultmeaning
len(ds)120number of samples
ds[0][0].shape[4]one feature vector
ds[0][1].shape[1]one scalar target
ds[0][0][:4][0.3047, -1.04, 0.7505, 0.9406]first sample values (rng seed 42)

9. Fill in: result for Custom Dataset for tabular data

Comparison

Comparison matrix

From Custom Dataset for tabular data: refill the result column from what you know. The rest of the table is as it appeared.

callresultmeaning
len(ds)120number of samples
ds[0][0].shape[4]one feature vector
ds[0][1].shape[1]one scalar target
ds[0][0][:4][0.3047, -1.04, 0.7505, 0.9406]first sample values (rng seed 42)

10. DataLoader: batching, shuffling, workers

Section

Part 2 of 4

11. DataLoader wraps Dataset for training

Concept

DataLoader calls __getitem__ repeatedly, collates results into batch tensors, and exposes an iterator over the whole dataset (one epoch).

argumenteffecttypical value
batch_sizesamples per batch32 or 64
shufflere-randomize each epochTrue for train, False for val
num_workersparallel prefetch processes0 (Windows/Colab), 4+ (Linux)
pin_memorypage-lock CPU RAM → faster GPU transferTrue when CUDA available

With n=120 and batch_size=32: batches of 32, 32, 32, 24 — the last batch is smaller (drop_last=False by default).

12. What each one costs: DataLoader wraps Dataset for training

Trade off

Comparison matrix

From DataLoader wraps Dataset for training: every row here is a choice with a cost. Fill the effect column, then say which row you would actually pick and what you give up for it.

argumenteffecttypical value
batch_sizesamples per batch32 or 64
shufflere-randomize each epochTrue for train, False for val
num_workersparallel prefetch processes0 (Windows/Colab), 4+ (Linux)
pin_memorypage-lock CPU RAM → faster GPU transferTrue when CUDA available

13. What has to be given first: Iterating a DataLoader

Missing information

Discussion prompt

Create a DataLoader on the 120-sample dataset and inspect the batch structure.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

ceil(120/32) = 4. The last mini-batch gets the remaining 24 samples. The training loop sees all 120 samples each epoch regardless.

14. Iterating a DataLoader

Worked example

Create a DataLoader on the 120-sample dataset and inspect the batch structure.

from torch.utils.data import DataLoader

loader = DataLoader(ds, batch_size=32, shuffle=True, num_workers=0)
print('batches per epoch:', len(loader))
xb, yb = next(iter(loader))
print('batch X shape:', xb.shape)   # (32, 4)
print('batch y shape:', yb.shape)   # (32, 1)
batch_sizes = [len(b[0]) for b in loader]
print('all batch sizes:', batch_sizes)

4 batches: [32, 32, 32, 24] — 3 full + 1 remainder

Why: ceil(120/32) = 4. The last mini-batch gets the remaining 24 samples. The training loop sees all 120 samples each epoch regardless.

batchsizeshape (X)shape (y)
132(32, 4)(32, 1)
232(32, 4)(32, 1)
332(32, 4)(32, 1)
4 (last)24(24, 4)(24, 1)

15. Watch it run: Iterating a DataLoader

Pattern

Step through it

Step through Iterating a DataLoader one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: batch is 1
  2. Step 2: batch is 2
  3. Step 3: batch is 3
  4. Step 4: batch is 4 (last)

16. Something is wrong here: shuffle=True on the validation loader

Anomaly

Predict first

A student writes this, and it looks reasonable:

Use shuffle=True on both train and validation DataLoaders — consistent is cleaner.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Shuffling the validation set changes nothing about the loss/accuracy (every sample is seen), but it invalidates deterministic debugging: the 'same' epoch-2 batch now contains different samples, making loss curves hard…

Only shuffle the training loader; keep shuffle=False on validation and test.

Why: Shuffling the validation set changes nothing about the loss/accuracy (every sample is seen), but it invalidates deterministic debugging: the 'same' epoch-2 batch now contains different samples, making loss curves hard to reproduce.

17. Trap: shuffle=True on the validation loader

Trap

The trap

Use shuffle=True on both train and validation DataLoaders — consistent is cleaner.

DataLoader(val_ds, batch_size=64, shuffle=True)

Why: Shuffling the validation set changes nothing about the loss/accuracy (every sample is seen), but it invalidates deterministic debugging: the 'same' epoch-2 batch now contains different samples, making loss curves hard to reproduce.

The fix

Only shuffle the training loader; keep shuffle=False on validation and test.

DataLoader(train_ds, shuffle=True) / DataLoader(val_ds, shuffle=False)

Why: Training benefits from randomized order (escapes local correlations in minibatches). Validation needs deterministic, reproducible ordering for fair comparison across runs.

18. Break it on purpose: shuffle=True on the validation loader

Break the constraint

Discussion prompt

The rule this trap just fixed:

Training benefits from randomized order (escapes local correlations in minibatches). Validation needs deterministic, reproducible ordering for fair comparison across runs.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

Shuffling the validation set changes nothing about the loss/accuracy (every sample is seen), but it invalidates deterministic debugging: the 'same' epoch-2 batch now contains different samples, making loss curves hard to reproduce.

19. Transforms: normalisation inside Dataset

Section

Part 3 of 4

20. Where transforms live

Concept

Transforms pre-process each sample inside __getitem__. They are callables applied before returning a tensor. torchvision.transforms ships image ops; for tabular data you write your own.

This keeps the transform logic co-located with the data, invisible to the training loop — same interface whether you add augmentation or not.

21. By analogy: Where transforms live

Analogy

Discussion prompt

Explain Where transforms live by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Transforms pre-process each sample inside __getitem__. They are callables applied before returning a tensor. torchvision.transforms ships image ops; for tabular data you write your own.

22. Predict the next row: NormalizingDataset — transform in __init__

Pattern

Predict first

The table runs: feat 0 | ≈ 0.0 | 1.0 · feat 1 | ≈ 0.0 | 1.0 · feat 2 | ≈ 0.0 | 1.0

In NormalizingDataset — transform in __init__, given the rows so far: what is the next one — the row where feature is feat 3?

Correct: feat 3 | ≈ 0.0 | 1.0

featuremean after normstd after norm
feat 0≈ 0.01.0
feat 1≈ 0.01.0
feat 2≈ 0.01.0
feat 3≈ 0.01.0

Why: The relationship between the columns, not the individual numbers, is what generates the next row. z-score normalization (Lesson 8 feature scaling) ensures all input dimensions compete fairly during gradient descent.

23. NormalizingDataset — transform in __init__

Worked example

Compute z-score statistics on construction and apply them on every __getitem__ call. Verify the output statistics.

class NormalizingDataset(Dataset):
    def __init__(self, X, y):
        X_t = torch.tensor(X, dtype=torch.float32)
        self.mean = X_t.mean(0)
        self.std  = X_t.std(0) + 1e-8
        self.X = (X_t - self.mean) / self.std  # normalize all at once
        self.y = torch.tensor(y, dtype=torch.float32).unsqueeze(1)
    def __len__(self): return len(self.X)
    def __getitem__(self, i): return self.X[i], self.y[i]

ds_norm = NormalizingDataset(X_np, y_np)
print(ds_norm.X.mean(0).numpy().round(4))   # [-0. 0. -0. 0.]
print(ds_norm.X.std(0).numpy().round(4))    # [1. 1. 1. 1.]

Feature means ≈ [-0.0, 0.0, -0.0, 0.0], stds ≈ [1.0, 1.0, 1.0, 1.0]

Why: z-score normalization (Lesson 8 feature scaling) ensures all input dimensions compete fairly during gradient descent. The training loop code is unchanged — the Dataset handles it.

featuremean after normstd after norm
feat 0≈ 0.01.0
feat 1≈ 0.01.0
feat 2≈ 0.01.0
feat 3≈ 0.01.0

24. What stays fixed: NormalizingDataset — transform in __init__

Invariant

Step through it

Step through NormalizingDataset — transform in __init__ one row at a time. One of these columns never changes — find it, and say why it cannot.

  1. Step 1: feature is feat 0
  2. Step 2: feature is feat 1
  3. Step 3: feature is feat 2
  4. Step 4: feature is feat 3

25. nn.Module internals + serialization

Section

Part 4 of 4

26. How nn.Module tracks parameters

Concept

When you assign an nn.Linear (or any nn.Module) as an attribute of another Module, PyTorch registers its Parameter tensors automatically.

Always verify parameter counts before training. A wrong architecture (mismatched dims) shows up here before the forward pass crashes.

27. Teach it back: How nn.Module tracks parameters

Explain it

Discussion prompt

Explain How nn.Module tracks parameters to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

When you assign an nn.Linear (or any nn.Module) as an attribute of another Module, PyTorch registers its Parameter tensors automatically.

28. Guess the shape of the answer: Counting parameters in TwoLayerNet

Estimation

Predict first

Build TwoLayerNet(4, 8, 1) and count every trainable parameter. Predict the total before running.

Commit before you compute: what does Counting parameters in TwoLayerNet come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: fc1.weight (8×4)=32, fc1.bias (8)=8, fc2.weight (1×8)=8, fc2.bias (1)=1 → total 49

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. nn.Linear(in, out) has (in × out) weights plus out biases.

29. Counting parameters in TwoLayerNet

Worked example

Build TwoLayerNet(4, 8, 1) and count every trainable parameter. Predict the total before running.

import torch.nn as nn

class TwoLayerNet(nn.Module):
    def __init__(self, in_f, h, out_f):
        super().__init__()
        self.fc1 = nn.Linear(in_f, h)
        self.relu = nn.ReLU()
        self.fc2 = nn.Linear(h, out_f)
    def forward(self, x):
        return self.fc2(self.relu(self.fc1(x)))

net = TwoLayerNet(4, 8, 1)
for name, p in net.named_parameters():
    print(f'{name}: {tuple(p.shape)} = {p.numel()}')
print('Total:', sum(p.numel() for p in net.parameters()))

fc1.weight (8×4)=32, fc1.bias (8)=8, fc2.weight (1×8)=8, fc2.bias (1)=1 → total 49

Why: nn.Linear(in, out) has (in × out) weights plus out biases. Summing across both layers: 32+8+8+1=49. nn.ReLU has no learnable parameters.

tensorshapenumel
fc1.weight(8, 4)32
fc1.bias(8,)8
fc2.weight(1, 8)8
fc2.bias(1,)1
total—49

30. Watch it run: Counting parameters in TwoLayerNet

Pattern

Step through it

Step through Counting parameters in TwoLayerNet one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: tensor is fc1.weight
  2. Step 2: tensor is fc1.bias
  3. Step 3: tensor is fc2.weight
  4. Step 4: tensor is fc2.bias
  5. Step 5: tensor is total

31. Something is wrong here: saving the full model object

Anomaly

Predict first

A student writes this, and it looks reasonable:

Save the entire model: torch.save(model, 'model.pt') — simple, one line, loads back with torch.load.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: The full-model save pickles the Python class name.

Save only the state_dict — a dict of tensor names to tensors, independent of Python class structure.

Why: The full-model save pickles the Python class name. If you rename or move TwoLayerNet, or load in a different environment, unpickling fails with an AttributeError. Resume training is also fragile.

32. Trap: saving the full model object

Trap

The trap

Save the entire model: torch.save(model, 'model.pt') — simple, one line, loads back with torch.load.

torch.save(model, 'full.pt') # 3583 bytes, requires the class definition at load time

Why: The full-model save pickles the Python class name. If you rename or move TwoLayerNet, or load in a different environment, unpickling fails with an AttributeError. Resume training is also fragile.

The fix

Save only the state_dict — a dict of tensor names to tensors, independent of Python class structure.

torch.save(model.state_dict(), 'sd.pt') # 2525 bytes; load with model.load_state_dict(...)

Why: state_dict is portable: the class definition lives in your code, the file holds only weights. Standard practice for checkpoints, deployment, and sharing. 33% smaller file vs full save.

33. Which of these survive contact with Lesson 42: PyTorch Data Pipeline & Model…?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Any Python object that implements __len__ and __getitem__ is a valid PyTorch Dataset. The DataLoader calls these — nothing else.; DataLoader calls __getitem__ repeatedly, collates results into batch tensors, and exposes an iterator over the whole dataset (one epoch).; When you assign an nn.Linear (or any nn.Module) as an attribute of another Module, PyTorch registers its Parameter tensors automatically.
Breaks
Use shuffle=True on both train and validation DataLoaders — consistent is cleaner.; Save the entire model: torch.save(model, 'model.pt') — simple, one line, loads back with torch.load.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 42: PyTorch Data Pipeline & Model Serialization puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

34. What has to be given first: state_dict save and reload

Missing information

Discussion prompt

Save model.state_dict(), create a fresh TwoLayerNet, call load_state_dict(), and verify all weights are identical.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

The four tensors are restored exactly. weights_only=True (PyTorch 2+ default) prevents arbitrary code execution during unpickling — always pass it when loading a state_dict.

35. state_dict save and reload

Worked example

Save model.state_dict(), create a fresh TwoLayerNet, call load_state_dict(), and verify all weights are identical.

import torch, os, tempfile

tmpdir = tempfile.mkdtemp()
sd_path = os.path.join(tmpdir, 'sd.pt')
torch.save(net.state_dict(), sd_path)

restored = TwoLayerNet(4, 8, 1)
restored.load_state_dict(torch.load(sd_path, weights_only=True))
restored.eval()

match = all(torch.allclose(a, b)
            for a, b in zip(net.parameters(), restored.parameters()))
print('keys :', list(net.state_dict().keys()))
print('match:', match)

keys: ['fc1.weight', 'fc1.bias', 'fc2.weight', 'fc2.bias'] match: True

Why: The four tensors are restored exactly. weights_only=True (PyTorch 2+ default) prevents arbitrary code execution during unpickling — always pass it when loading a state_dict.

save stylefile sizeportable?weights_only
state_dict only2525 bytesyesTrue (safe)
full model3583 bytesfragileFalse needed

36. Work backwards from the answer: state_dict save and reload

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

keys: ['fc1.weight', 'fc1.bias', 'fc2.weight', 'fc2.bias'] match: True

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Save model.state_dict(), create a fresh TwoLayerNet, call load_state_dict(), and verify all weights are identical.

37. Checkpoint dict: model + optimizer + epoch

Concept

To resume training you need more than just model weights — the optimizer's momentum buffers matter too. Pack all state into one dict.

Missing the optimizer state is the silent bug: the model weights are right but momentum is cold-started, so the first few resumed epochs underfit.

38. By analogy: Checkpoint dict: model + optimizer + epoch

Analogy

Discussion prompt

Explain Checkpoint dict: model + optimizer + epoch by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

To resume training you need more than just model weights — the optimizer's momentum buffers matter too. Pack all state into one dict.

39. Guess the shape of the answer: Checkpoint save and resume

Estimation

Predict first

Train 5 epochs, save a checkpoint, restore and continue epochs 6-10. Verify the loss trajectory is identical to a fresh 10-epoch run.

Commit before you compute: what does Checkpoint save and resume come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Resumed loss at epoch 6: 0.5569 → epoch 10: 0.1666 — identical to uninterrupted run

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Because we restored opt_state, Adam's per-parameter accumulators are warm.

40. Checkpoint save and resume

Worked example

Train 5 epochs, save a checkpoint, restore and continue epochs 6-10. Verify the loss trajectory is identical to a fresh 10-epoch run.

model = TwoLayerNet(4, 8, 1); opt = torch.optim.Adam(model.parameters(), lr=1e-2)
loader = DataLoader(ds_norm, batch_size=32, shuffle=False, num_workers=0)
loss_fn = nn.MSELoss()

for epoch in range(5):
    for xb, yb in loader:
        opt.zero_grad(); loss_fn(model(xb), yb).backward(); opt.step()

torch.save({'epoch': 4,
            'model_state': model.state_dict(),
            'opt_state':   opt.state_dict(),
            'loss': 0.6849}, ckpt_path)

# --- resume ---
m2 = TwoLayerNet(4, 8, 1); o2 = torch.optim.Adam(m2.parameters(), lr=1e-2)
ck = torch.load(ckpt_path, weights_only=False)
m2.load_state_dict(ck['model_state']); o2.load_state_dict(ck['opt_state'])
for epoch in range(ck['epoch']+1, 10):
    for xb, yb in loader:
        o2.zero_grad(); loss_fn(m2(xb), yb).backward(); o2.step()

Resumed loss at epoch 6: 0.5569 → epoch 10: 0.1666 — identical to uninterrupted run

Why: Because we restored opt_state, Adam's per-parameter accumulators are warm. The resumed loss curve is bit-for-bit identical to training straight through — the checkpoint is transparent.

epochloss (no ckpt)loss (resumed)
11.19121.1912
50.68490.6849 → ckpt
60.55690.5569
80.31180.3118
100.16660.1666

41. Fill in: loss (no ckpt) for Checkpoint save and resume

Comparison

Comparison matrix

From Checkpoint save and resume: refill the loss (no ckpt) column from what you know. The rest of the table is as it appeared.

epochloss (no ckpt)loss (resumed)
11.19121.1912
50.68490.6849 → ckpt
60.55690.5569
80.31180.3118
100.16660.1666

42. Rebuild the recipe: The data-pipeline + serialization recipe

Ranking

Put in order

These are the steps of The data-pipeline + serialization recipe, scrambled. Put them back in order before the next slide shows you.

  1. Dataset: subclass Dataset, implement __len__ and __getitem__; apply transforms in __init__ or __getitem__
  2. DataLoader: shuffle=True (train only); num_workers=0 (Windows), 4+ (Linux); pin_memory=True when using CUDA
  3. Parameter audit: sum(p.numel() for p in model.parameters()) before first training run
  4. Serialization: always state_dict over full model; pass weights_only=True on load
  5. Checkpoint: save {epoch, model_state_dict, optimizer_state_dict, loss}; restore all four before resuming

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

43. The data-pipeline + serialization recipe

Pattern

  1. Dataset: subclass Dataset, implement __len__ and __getitem__; apply transforms in __init__ or __getitem__
  2. DataLoader: shuffle=True (train only); num_workers=0 (Windows), 4+ (Linux); pin_memory=True when using CUDA
  3. Parameter audit: sum(p.numel() for p in model.parameters()) before first training run
  4. Serialization: always state_dict over full model; pass weights_only=True on load
  5. Checkpoint: save {epoch, model_state_dict, optimizer_state_dict, loss}; restore all four before resuming

44. Where does it stop working: The data-pipeline + serialization recipe

Edge cases

Discussion prompt

The data-pipeline + serialization recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Dataset: subclass Dataset, implement __len__ and __getitem__; apply transforms in __init__ or __getitem__
  2. DataLoader: shuffle=True (train only); num_workers=0 (Windows), 4+ (Linux); pin_memory=True when using CUDA
  3. Parameter audit: sum(p.numel() for p in model.parameters()) before first training run
  4. Serialization: always state_dict over full model; pass weights_only=True on load
  5. Checkpoint: save {epoch, model_state_dict, optimizer_state_dict, loss}; restore all four before resuming

45. Rule out three: Check yourself — DataLoader batches

Elimination

Eliminate the wrong options

A dataset has 120 samples. DataLoader(ds, batch_size=32, drop_last=False). How many batches per epoch and what are their sizes?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. 4 batches: [32, 32, 32, 24]
  • B. 3 batches: [32, 32, 32]
  • C. 4 batches: [30, 30, 30, 30]
  • D. 5 batches: [24, 24, 24, 24, 24]

Survives elimination: A

Why: ceil(120/32) = 4 batches. The first three contain 32 samples; the final batch has the remaining 120 − 3×32 = 24. drop_last=False (default) keeps the partial batch.

46. Check yourself — DataLoader batches

Check

Predict first, then verify.

Check your understanding

A dataset has 120 samples. DataLoader(ds, batch_size=32, drop_last=False). How many batches per epoch and what are their sizes?

  • A. 4 batches: [32, 32, 32, 24] (correct)
  • B. 3 batches: [32, 32, 32]
  • C. 4 batches: [30, 30, 30, 30]
  • D. 5 batches: [24, 24, 24, 24, 24]

Answer: A

Why: ceil(120/32) = 4 batches. The first three contain 32 samples; the final batch has the remaining 120 − 3×32 = 24. drop_last=False (default) keeps the partial batch.

Why B tempts people
floor(120/32)=3 is what drop_last=True gives — it discards the 24-sample remainder. The default keeps it.
Why C tempts people
120/4=30 is evenly-distributed, but DataLoader doesn't spread remainder samples — it fills full batches first, partial last.
Why D tempts people
120/5=24 batches of 24 only if batch_size=24, not 32.

47. Check yourself — parameter count

Check

Do the arithmetic on paper first.

Check your understanding

TwoLayerNet has nn.Linear(4, 8) then nn.Linear(8, 1) (with bias). What is the total trainable parameter count?

  • A. 49 (correct)
  • B. 40
  • C. 33
  • D. 48

Answer: A

Why: nn.Linear(4,8): 4×8=32 weights + 8 biases = 40. nn.Linear(8,1): 8×1=8 weights + 1 bias = 9. Total = 40+9 = 49. Verified: sum(p.numel() for p in net.parameters()) = 49.

Why B tempts people
40 counts only the first layer (32 weights + 8 biases = 40) and forgets the second layer entirely.
Why C tempts people
33 = 32 (fc1 weights) + 1 (fc2 bias) — omits both bias vectors and the fc2 weight matrix.
Why D tempts people
48 = 4×8 + 8×1 = 32+8 = 40? No — 48 would mean ignoring one bias term. The correct breakdown is 32+8+8+1=49.

48. Rule out three: Check yourself — checkpoint resume

Elimination

Eliminate the wrong options

You save only model.state_dict() at epoch 5 and restore it before resuming. Compared to an uninterrupted run, what is most likely wrong?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Adam's momentum accumulators are cold-started, causing slow or erratic loss for the first resumed epochs
  • B. The model weights are corrupted — you need the full-model save to get valid weights
  • C. Nothing; model.state_dict() is sufficient for a perfect resume
  • D. The learning rate resets to its default, overwriting any scheduler changes

Survives elimination: A

Why: Adam maintains per-parameter running mean (m) and variance (v) buffers — these live in the optimizer state_dict, not the model state_dict. Without them, Adam cold-starts its accumulators and behaves like the first step of training for several iterations.

49. Check yourself — checkpoint resume

Check

Identify the critical missing piece.

Check your understanding

You save only model.state_dict() at epoch 5 and restore it before resuming. Compared to an uninterrupted run, what is most likely wrong?

  • A. Adam's momentum accumulators are cold-started, causing slow or erratic loss for the first resumed epochs (correct)
  • B. The model weights are corrupted — you need the full-model save to get valid weights
  • C. Nothing; model.state_dict() is sufficient for a perfect resume
  • D. The learning rate resets to its default, overwriting any scheduler changes

Answer: A

Why: Adam maintains per-parameter running mean (m) and variance (v) buffers — these live in the optimizer state_dict, not the model state_dict. Without them, Adam cold-starts its accumulators and behaves like the first step of training for several iterations.

Why B tempts people
model.state_dict() saves all weight and bias tensors exactly — it is the correct and portable way to save weights. The full-model save is NOT needed for weight correctness.
Why C tempts people
Sufficient only for SGD without momentum. For Adam (and SGD with momentum=0.9), missing optimizer_state_dict means the resumed trajectory diverges from the uninterrupted one.
Why D tempts people
The learning rate is part of the optimizer state_dict and would also be lost — but the primary accuracy loss is the momentum/variance accumulators, not LR alone.

50. Your turn: build the pipeline

Section

Project

51. Project: end-to-end data pipeline

Concept

Build a complete, checkpointed training run on synthetic tabular data: custom Dataset with normalization, DataLoader, TwoLayerNet, Adam, checkpoint at epoch 5, resume to epoch 10.

#requirementkey tool
1NormalizingDataset(X, y)Dataset.__init__, __getitem__
2DataLoader, inspect batch shapesDataLoader, iter/next
3count parameters before training.named_parameters(), .numel()
4train 5 epochs, save checkpointtorch.save({epoch, model_state, opt_state})
5resume epochs 6-10, verify loss matchesload_state_dict, opt.load_state_dict

Target: loss ≈ 0.1666 at epoch 10, matching an uninterrupted run. If your resumed loss diverges, check that you saved AND loaded the optimizer state.

52. Break it if you can: Project: end-to-end data pipeline

Counterexample

Discussion prompt

Build a complete, checkpointed training run on synthetic tabular data: custom Dataset with normalization, DataLoader, TwoLayerNet, Adam, checkpoint at epoch 5, resume to epoch 10.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Target: loss ≈ 0.1666 at epoch 10, matching an uninterrupted run. If your resumed loss diverges, check that you saved AND loaded the optimizer state.

53. Milestone 1 — NormalizingDataset + DataLoader

Worked example

Your turn: implement NormalizingDataset and create a DataLoader. Predict how many batches for n=120, batch_size=32.

Hint: compute mean and std in __init__, normalize all at once with (X_t - mean) / std. Add 1e-8 to std to prevent div-by-zero.

ds_norm = NormalizingDataset(X_np, y_np)
loader  = DataLoader(ds_norm, batch_size=32, shuffle=False, num_workers=0)
print('batches:', len(loader))          # 4
xb, yb = next(iter(loader))
print('X batch:', xb.shape, 'y batch:', yb.shape)
checkexpectedyour output
len(loader)44
xb.shape(32, 4)(32, 4)
yb.shape(32, 1)(32, 1)
ds_norm.X.std(0) ≈[1.0, 1.0, 1.0, 1.0]match if normalized

54. Milestone 2 — parameter audit

Worked example

Your turn: instantiate TwoLayerNet(4, 8, 1) and enumerate its parameters. Predict the total count before running.

Hint: nn.Linear(in, out) has in × out weights + out biases. Add up both layers.

net = TwoLayerNet(4, 8, 1)
for name, p in net.named_parameters():
    print(f'{name}: {tuple(p.shape)} -> {p.numel()}')
print('Total:', sum(p.numel() for p in net.parameters()))
tensorshapecount
fc1.weight(8, 4)32
fc1.bias(8,)8
fc2.weight(1, 8)8
fc2.bias(1,)1
TOTAL—49

55. Watch it run: Milestone 2 — parameter audit

Pattern

Step through it

Step through Milestone 2 — parameter audit one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: tensor is fc1.weight
  2. Step 2: tensor is fc1.bias
  3. Step 3: tensor is fc2.weight
  4. Step 4: tensor is fc2.bias
  5. Step 5: tensor is TOTAL

56. Milestone 3 — checkpoint save and resume

Worked example

Your turn: train 5 epochs, save {epoch, model_state, opt_state, loss}, then restore and continue epochs 6-10. Predict: does the resumed loss match?

Hint: torch.save(ckpt_dict, path) / torch.load(path, weights_only=False). Restore both model_state and opt_state before continuing.

torch.manual_seed(0)
model = TwoLayerNet(4, 8, 1); opt = torch.optim.Adam(model.parameters(), lr=1e-2)
for epoch in range(5):
    for xb, yb in loader:
        opt.zero_grad(); loss_fn(model(xb), yb).backward(); opt.step()
torch.save({'epoch':4,'model_state':model.state_dict(),'opt_state':opt.state_dict(),'loss':0.6849}, 'ckpt.pt')

m2 = TwoLayerNet(4, 8, 1); o2 = torch.optim.Adam(m2.parameters(), lr=1e-2)
ck = torch.load('ckpt.pt', weights_only=False)
m2.load_state_dict(ck['model_state']); o2.load_state_dict(ck['opt_state'])
for epoch in range(ck['epoch']+1, 10):
    for xb, yb in loader:
        o2.zero_grad(); loss_fn(m2(xb), yb).backward(); o2.step()
epochexpected loss
5 (ckpt saved)0.6849
6 (resumed)0.5569
80.3118
100.1666

57. What each one costs: Milestone 3 — checkpoint save and resume

Trade off

Comparison matrix

From Milestone 3 — checkpoint save and resume: every row here is a choice with a cost. Fill the expected loss column, then say which row you would actually pick and what you give up for it.

epochexpected loss
5 (ckpt saved)0.6849
6 (resumed)0.5569
80.3118
100.1666

58. The full program

Concept

import torch, torch.nn as nn, numpy as np
from torch.utils.data import Dataset, DataLoader

class NormalizingDataset(Dataset):
    def __init__(self, X, y):
        X_t = torch.tensor(X, dtype=torch.float32)
        self.X = (X_t - X_t.mean(0)) / (X_t.std(0) + 1e-8)
        self.y = torch.tensor(y, dtype=torch.float32).unsqueeze(1)
    def __len__(self): return len(self.X)
    def __getitem__(self, i): return self.X[i], self.y[i]

class TwoLayerNet(nn.Module):
    def __init__(self, in_f, h, out_f):
        super().__init__()
        self.fc1 = nn.Linear(in_f, h); self.relu = nn.ReLU()
        self.fc2 = nn.Linear(h, out_f)
    def forward(self, x): return self.fc2(self.relu(self.fc1(x)))

rng = np.random.default_rng(42)
X = rng.standard_normal((120, 4))
y = X[:, 0] + 0.5*X[:, 1] + 0.1*rng.standard_normal(120)
ds = NormalizingDataset(X, y)
loader = DataLoader(ds, batch_size=32, shuffle=False)
print('params:', sum(p.numel() for p in TwoLayerNet(4,8,1).parameters()))  # 49
outputvalue
Total params49
Batches/epoch4 ([32,32,32,24])
Loss @ epoch 50.6849
Loss @ epoch 10 (resumed)0.1666

If your resumed epoch-10 loss equals the uninterrupted run's — you've built the full pipeline. The Dataset + DataLoader + checkpoint pattern is reused for every model from here through the competition.

59. Fill in: value for The full program

Comparison

Comparison matrix

From The full program: refill the value column from what you know. The rest of the table is as it appeared.

outputvalue
Total params49
Batches/epoch4 ([32,32,32,24])
Loss @ epoch 50.6849
Loss @ epoch 10 (resumed)0.1666

60. Show it off

Concept

Out loud, slides closed: (1) explain what __len__ and __getitem__ must return and why the DataLoader needs them; (2) why shuffle=False on a validation loader; (3) what goes wrong if you skip saving optimizer_state_dict.

Stretch (homework): profile num_workers=0 vs num_workers=2 on a larger dataset; implement augmentation in __getitem__ (add random Gaussian noise to features only during training); add a pin_memory=True branch for CUDA. Next up: SVM + MLP Architecture (Lesson 43).

61. Connect it up: Lesson 42: PyTorch Data Pipeline & Model Serialization

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Dataset: wrapping your data · DataLoader: batching, shuffling, workers · Transforms: normalisation inside Dataset · nn.Module internals + serialization · Your turn: build the pipeline. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

62. What you can do now

Recap

conceptthe one thing to remember
Dataset contract__len__ + __getitem__ — that's all DataLoader needs
DataLoader shuffleTrue for train, False for val/test
parameter countsum(p.numel() for p in model.parameters())
save preferencestate_dict > full model (portable, smaller)
checkpointsave opt state too — momentum is not in model weights

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 42 (Week 15 — SVM + MLP Architecture) — Barron · USAAIO Round 2 Preparation, 2026
  2. PyTorch 2.7 docs: torch.utils.data, nn.Module, torch.save/load
  3. Dataset, DataLoader, parameter counting, checkpoint resume — all numbers verified with torch 2.7.1+cpu, numpy 2.2.6, June 2026 — Real execution, synthetic tabular data, seed 42

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108