USAAIO Lesson 48, from Week 17 of Phase 2. It covers the theory and implementation of inverted dropout, the difference between training and evaluation mode, dropout read as averaging over an exponential ensemble, MC Dropout for Bayesian uncertainty estimation, and label smoothing from scratch, then compares several dropout rates on a classification network. All the numbers were verified with torch 2.7.1 and sklearn in June 2026. The lesson runs to 29 slides.
Subject: Machine Learning · 59 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 48 · Week 17 (Phase 2)
Inverted dropout, MC Dropout uncertainty, label smoothing, mixup — the regularizers that keep deep nets from memorizing the training set.
Objectives
nn.CrossEntropyLoss(label_smoothing=eps)Warm-up
Discussion prompt
Before we open Lesson 48: Dropout & Modern Regularizers: without looking back, what was the main idea of Decision Trees — CART, Information Gain, Gini, and Pruning, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
information gain from entropy, Gini impurity, the CART binary-split induction algorithm, stopping criteria, and cost-complexity pruning (ccp_alpha). Build a DecisionTreeClassifier from raw splits to a pruned final model on iris and make_classification.
Section
Part 1 of 3
Concept
At each forward pass during training, each neuron is independently zeroed with probability p. The surviving neurons are rescaled so the expected activation is unchanged.
\[ \tilde{h}_i = \frac{m_i \cdot h_i}{1-p}, \quad m_i \sim \text{Bernoulli}(1-p) \]
At test time no neurons are dropped — the full network runs. The /(1-p) scaling (inverted dropout) is already absorbed into weights, so inference needs no correction.
Counterexample
Discussion prompt
At each forward pass during training, each neuron is independently zeroed with probability p. The surviving neurons are rescaled so the expected activation is unchanged.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
At test time no neurons are dropped — the full network runs. The /(1-p) scaling (inverted dropout) is already absorbed into weights, so inference needs no correction.
Concept
A neuron can't rely on a specific neighbor always being present. Each neuron must learn a feature that is useful on its own, not in concert with a fixed set of collaborators.
This forces the network to learn redundant, distributed representations of each concept — exactly what makes it generalize. Deep overfitting often traces to neurons that specialize to co-dependencies in the training batch.
Analogy
Discussion prompt
Explain Why dropout prevents co-adaptation by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
A neuron can't rely on a specific neighbor always being present. Each neuron must learn a feature that is useful on its own, not in concert with a fixed set of collaborators.
Intuition
A network with n neurons has 2^n possible thinned sub-networks. Each training step samples and trains one of them.
At test time, using the full network with no dropout approximates averaging the predictions of all 2^n sub-networks — a geometric mean over an exponentially large ensemble (Srivastava et al., 2014).
p=0.5 maximizes the number of distinct thinned networks, which is why it is the canonical starting point for hidden layers (p=0.1–0.2 is common for inputs to avoid discarding too much signal).
Explain it
Discussion prompt
Explain Dropout as an exponential ensemble to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
A network with n neurons has 2^n possible thinned sub-networks. Each training step samples and trains one of them.
Estimation
Predict first
Implement inverted dropout manually. Seed 0, p=0.5, four input neurons — trace the mask, the scaled output in train mode, and the unchanged output in eval mode.
Commit before you compute: what does Inverted dropout from scratch come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: mask=[0,1,0,0]; neuron 1 survives and is scaled from 2.0 to 4.0; others zeroed
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Multiplying by 1/(1-0.5)=2 compensates for the expected 50% zeroing, preserving the expected activation magnitude at each layer.
Worked example
Implement inverted dropout manually. Seed 0, p=0.5, four input neurons — trace the mask, the scaled output in train mode, and the unchanged output in eval mode.
import torch
torch.manual_seed(0)
x = torch.tensor([1.0, 2.0, 3.0, 4.0])
p = 0.5
# train mode: sample mask, scale by 1/(1-p)
mask = (torch.rand_like(x) > p).float()
x_train = x * mask / (1.0 - p)
# eval mode: no dropout, no scaling
x_eval = x
print('mask :', mask.tolist())
print('train :', x_train.tolist())
print('eval :', x_eval.tolist())mask=[0,1,0,0]; neuron 1 survives and is scaled from 2.0 to 4.0; others zeroed
Why: Multiplying by 1/(1-0.5)=2 compensates for the expected 50% zeroing, preserving the expected activation magnitude at each layer.
| neuron | x | mask | train output | eval output |
|---|---|---|---|---|
| 0 | 1.0 | 0 | 0.0 | 1.0 |
| 1 | 2.0 | 1 | 4.0 | 2.0 |
| 2 | 3.0 | 0 | 0.0 | 3.0 |
| 3 | 4.0 | 0 | 0.0 | 4.0 |
Comparison
Comparison matrix
From Inverted dropout from scratch: refill the x column from what you know. The rest of the table is as it appeared.
| neuron | x | mask | train output | eval output |
|---|---|---|---|---|
| 0 | 1.0 | 0 | 0.0 | 1.0 |
| 1 | 2.0 | 1 | 4.0 | 2.0 |
| 2 | 3.0 | 0 | 0.0 | 3.0 |
| 3 | 4.0 | 0 | 0.0 | 4.0 |
Anomaly
Predict first
A student writes this, and it looks reasonable:
The model trained fine, so just call model(x_test) directly.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: PyTorch's nn.Dropout checks self.training.
Call model.eval() before any inference, and model.train() to resume training.
Why: PyTorch's nn.Dropout checks self.training. If the model is still in train mode, it randomly zeros neurons during inference — predictions become stochastic and wrong every call.
Trap
The model trained fine, so just call model(x_test) directly.
Run inference without calling model.eval()
Why: PyTorch's nn.Dropout checks self.training. If the model is still in train mode, it randomly zeros neurons during inference — predictions become stochastic and wrong every call.
Call model.eval() before any inference, and model.train() to resume training.
Call model.eval() then wrap inference in torch.no_grad()
Why: eval() disables dropout (and batch norm running stats). no_grad() saves memory. Both are required for correct, efficient inference.
Break the constraint
Discussion prompt
The rule this trap just fixed:eval() disables dropout (and batch norm running stats). no_grad() saves memory. Both are required for correct, efficient inference.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
PyTorch's nn.Dropout checks self.training. If the model is still in train mode, it randomly zeros neurons during inference — predictions become stochastic and wrong every call.
Missing information
Discussion prompt
Train a two-hidden-layer ReLU net (64 → 64 → 2) on make_classification (2000 samples, 20 features) with dropout rates 0.0 through 0.7. Observe the train/test accuracy gap.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
No dropout memorizes training data. Light dropout (p=0.1) regularizes without crippling capacity. Very high p (0.7) under-trains the network too aggressively, hurting both.
Worked example
Train a two-hidden-layer ReLU net (64 → 64 → 2) on make_classification (2000 samples, 20 features) with dropout rates 0.0 through 0.7. Observe the train/test accuracy gap.
import torch, torch.nn as nn
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
X, y = make_classification(n_samples=2000, n_features=20,
n_informative=10, random_state=0)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.3, random_state=0)
Xtr = torch.tensor(X_tr, dtype=torch.float32)
Xte = torch.tensor(X_te, dtype=torch.float32)
ytr = torch.tensor(y_tr, dtype=torch.long)
yte = torch.tensor(y_te, dtype=torch.long)
results = {}
for p in [0.0, 0.1, 0.3, 0.5, 0.7]:
torch.manual_seed(0)
net = nn.Sequential(nn.Linear(20,64), nn.ReLU(), nn.Dropout(p),
nn.Linear(64,64), nn.ReLU(), nn.Dropout(p),
nn.Linear(64,2))
opt = torch.optim.Adam(net.parameters(), lr=1e-3)
for _ in range(200):
net.train(); opt.zero_grad()
nn.CrossEntropyLoss()(net(Xtr), ytr).backward(); opt.step()
net.eval()
with torch.no_grad():
tr = (net(Xtr).argmax(1)==ytr).float().mean().item()
te = (net(Xte).argmax(1)==yte).float().mean().item()
results[p] = (round(tr,4), round(te,4))
print(f'p={p}: train={tr:.4f} test={te:.4f}')p=0.0 overfits: train=0.9964 vs test=0.9383; p=0.1 closes the gap: train=0.9836, test=0.9500 (best test)
Why: No dropout memorizes training data. Light dropout (p=0.1) regularizes without crippling capacity. Very high p (0.7) under-trains the network too aggressively, hurting both.
| p | train acc | test acc | gap |
|---|---|---|---|
| 0.0 | 0.9964 | 0.9383 | 0.0581 |
| 0.1 | 0.9836 | 0.9500 | 0.0336 |
| 0.3 | 0.9614 | 0.9383 | 0.0231 |
| 0.5 | 0.9329 | 0.9150 | 0.0179 |
| 0.7 | 0.9029 | 0.8950 | 0.0079 |
Trade off
Comparison matrix
From Dropout rates vs accuracy: every row here is a choice with a cost. Fill the gap column, then say which row you would actually pick and what you give up for it.
| p | train acc | test acc | gap |
|---|---|---|---|
| 0.0 | 0.9964 | 0.9383 | 0.0581 |
| 0.1 | 0.9836 | 0.9500 | 0.0336 |
| 0.3 | 0.9614 | 0.9383 | 0.0231 |
| 0.5 | 0.9329 | 0.9150 | 0.0179 |
| 0.7 | 0.9029 | 0.8950 | 0.0079 |
Section
Part 2 of 3
Concept
Gal & Ghahramani (2016) showed that a neural network with dropout approximates a Bayesian neural network — each forward pass draws a sample from the approximate posterior over weights.
\[ p(y^* \mid x^*, X, Y) \approx \frac{1}{T}\sum_{t=1}^{T} p(y^* \mid x^*, \hat{W}_t) \]
Run T forward passes with dropout active at test time. The mean gives the prediction; the variance gives the model's epistemic uncertainty (how confident it is about this region of input space).
Analogy
Discussion prompt
Explain Dropout as approximate Bayesian inference by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Gal & Ghahramani (2016) showed that a neural network with dropout approximates a Bayesian neural network — each forward pass draws a sample from the approximate posterior over weights.
Estimation
Predict first
Keep dropout active at test time (model.train()), run 100 forward passes, and report mean and std of the class-1 probability. High std = uncertain.
Commit before you compute: what does MC Dropout uncertainty estimation come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Points 0 and 4 (mean ~0.01-0.00, std ~0.02-0.01) are confidently class 0; point 2 (mean 0.816, std 0.126) is near a decision boundary
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. High std signals the model is uncertain — these are regions where the ensemble of thinned networks disagrees, indicating the model needs more data or a different feature near this boundary.
Worked example
Keep dropout active at test time (model.train()), run 100 forward passes, and report mean and std of the class-1 probability. High std = uncertain.
import torch, torch.nn as nn, torch.nn.functional as F
class MCNet(nn.Module):
def __init__(self, p=0.3):
super().__init__()
self.fc1 = nn.Linear(20, 64)
self.drop = nn.Dropout(p)
self.fc2 = nn.Linear(64, 2)
def forward(self, x):
return self.fc2(self.drop(F.relu(self.fc1(x))))
torch.manual_seed(0)
net = MCNet(p=0.3)
opt = torch.optim.Adam(net.parameters(), lr=1e-3)
for _ in range(200):
net.train(); opt.zero_grad()
nn.CrossEntropyLoss()(net(Xtr), ytr).backward(); opt.step()
# MC inference: keep train mode so dropout stays ON
net.train()
T = 100; torch.manual_seed(5); sample = Xte[:5]
probs = torch.stack([F.softmax(net(sample),dim=-1)[:,1]
for _ in range(T)])
print('mean:', probs.mean(0).round(decimals=4).tolist())
print('std: ', probs.std(0).round(decimals=4).tolist())Points 0 and 4 (mean ~0.01-0.00, std ~0.02-0.01) are confidently class 0; point 2 (mean 0.816, std 0.126) is near a decision boundary
Why: High std signals the model is uncertain — these are regions where the ensemble of thinned networks disagrees, indicating the model needs more data or a different feature near this boundary.
| point | mean P(class=1) | std | interpretation |
|---|---|---|---|
| 0 | 0.0139 | 0.0170 | confident class 0 |
| 1 | 0.8909 | 0.1010 | fairly confident class 1 |
| 2 | 0.8159 | 0.1256 | uncertain (near boundary) |
| 3 | 0.2017 | 0.0787 | leans class 0 |
| 4 | 0.0015 | 0.0076 | very confident class 0 |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Points 0 and 4 (mean ~0.01-0.00, std ~0.02-0.01) are confidently class 0; point 2 (mean 0.816, std 0.126) is near a decision boundary
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Keep dropout active at test time (model.train()), run 100 forward passes, and report mean and std of the class-1 probability. High std = uncertain.
Anomaly
Predict first
A student writes this, and it looks reasonable:
To get predictions, always call model.eval() first — then run the T forward passes.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: eval() disables dropout, making every forward pass identical.
For MC Dropout, keep the model in train() mode (or use a custom flag) so dropout stays active during the T inference passes.
Why: eval() disables dropout, making every forward pass identical. You compute the same deterministic output T times — all variance is zero, uncertainty is useless.
Trap
To get predictions, always call model.eval() first — then run the T forward passes.
Call model.eval() then loop T forward passes expecting stochastic outputs
Why: eval() disables dropout, making every forward pass identical. You compute the same deterministic output T times — all variance is zero, uncertainty is useless.
For MC Dropout, keep the model in train() mode (or use a custom flag) so dropout stays active during the T inference passes.
Keep model.train() during the T-pass loop, then call model.eval() for normal inference afterward
Why: Dropout is the source of stochasticity. eval() disables it. MC Dropout deliberately exploits that stochasticity to sample from the weight posterior — you must leave it on.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
p. The surviving neurons are rescaled so the expected activation is unchanged.; A network with n neurons has 2^n possible thinned sub-networks. Each training step samples and trains one of them.; Standard one-hot targets push the model to assign zero probability to all non-true classes, driving logits to large magnitudes and producing an overconfident model.model(x_test) directly.; To get predictions, always call model.eval() first — then run the T forward passes.Section
Part 3 of 3
Concept
Standard one-hot targets push the model to assign zero probability to all non-true classes, driving logits to large magnitudes and producing an overconfident model.
\[ y_k^{\text{smooth}} = \begin{cases} 1 - \varepsilon & k = \text{true class} \\ \dfrac{\varepsilon}{K-1} & k \neq \text{true class} \end{cases} \]
With K=10 and eps=0.1: the true-class target becomes 0.9, and each other class gets 0.0111. The model is penalized for being too certain — it learns calibrated probabilities.
Counterexample
Discussion prompt
Standard one-hot targets push the model to assign zero probability to all non-true classes, driving logits to large magnitudes and producing an overconfident model.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
With K=10 and eps=0.1: the true-class target becomes 0.9, and each other class gets 0.0111. The model is penalized for being too certain — it learns calibrated probabilities.
Estimation
Predict first
Implement label smoothing using log_softmax and a convex combination of NLL loss and uniform loss. Verify it matches PyTorch's built-in label_smoothing argument.
Commit before you compute: what does Label smoothing from scratch come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: std CE=1.7088, smooth CE (scratch)=1.7369, builtin=1.7369, match=True
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Smooth CE > standard CE on this batch because the model is forced to spread probability mass — the loss is always slightly higher.
Worked example
Implement label smoothing using log_softmax and a convex combination of NLL loss and uniform loss. Verify it matches PyTorch's built-in label_smoothing argument.
import torch, torch.nn as nn, torch.nn.functional as F
torch.manual_seed(7)
logits = torch.randn(4, 5) # 4 samples, 5 classes
targets = torch.tensor([0, 2, 1, 4])
eps = 0.1; K = 5
def smooth_ce(logits, targets, eps, K):
log_p = F.log_softmax(logits, dim=-1)
n = logits.shape[0]
nll = -log_p[range(n), targets] # standard NLL per sample
uniform = -log_p.mean(dim=-1) # uniform term
return ((1-eps)*nll + eps*uniform).mean()
std = nn.CrossEntropyLoss()(logits, targets)
scratch = smooth_ce(logits, targets, eps, K)
builtin = nn.CrossEntropyLoss(label_smoothing=0.1)(logits, targets)
print(f'std CE: {std.item():.4f}')
print(f'scratch: {scratch.item():.4f}')
print(f'builtin: {builtin.item():.4f}')
print(f'match: {abs(scratch.item()-builtin.item()) < 1e-5}')std CE=1.7088, smooth CE (scratch)=1.7369, builtin=1.7369, match=True
Why: Smooth CE > standard CE on this batch because the model is forced to spread probability mass — the loss is always slightly higher. The exact equality confirms the derivation is correct.
| loss variant | value | notes |
|---|---|---|
| standard CE | 1.7088 | one-hot targets |
| label smooth (scratch) | 1.7369 | (1-eps)NLL + epsuniform |
| label smooth (builtin) | 1.7369 | nn.CrossEntropyLoss(label_smoothing=0.1) |
| match | True | abs diff < 1e-5 |
Discrimination
Sort into buckets
Sort these by value, from memory, without looking back at Label smoothing from scratch. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Concept
Mixup (Zhang et al., 2018): linearly blend two training examples and their labels. The model must predict a soft mixture — a stronger form of label smoothing applied in input space.
\[ \tilde{x} = \lambda x_i + (1-\lambda) x_j, \quad \tilde{y} = \lambda y_i + (1-\lambda) y_j, \quad \lambda \sim \text{Beta}(\alpha,\alpha) \]
CutMix (Yun et al., 2019): paste a rectangular crop of one image into another, mixing labels proportionally to patch area. Both are dominant data augmentation strategies in vision competitions.
Sorting
Sort into buckets
These are the pieces of Lesson 48: Dropout & Modern Regularizers, out of order. Put each one back under the part of the lesson it belongs to.
Ranking
Put in order
These are the steps of The dropout & modern-regularizer recipe, scrambled. Put them back in order before the next slide shows you.
mask*(x/(1-p)) in train; x unchanged in eval — use model.train() / model.eval()model.train() at inference, run T passes, report mean and std of softmaxnn.CrossEntropyLoss(label_smoothing=eps) or (1-eps)*NLL + eps*uniform; eps=0.1 is typicalWhy: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
mask*(x/(1-p)) in train; x unchanged in eval — use model.train() / model.eval()model.train() at inference, run T passes, report mean and std of softmaxnn.CrossEntropyLoss(label_smoothing=eps) or (1-eps)*NLL + eps*uniform; eps=0.1 is typicalEdge cases
Discussion prompt
The dropout & modern-regularizer recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
mask*(x/(1-p)) in train; x unchanged in eval — use model.train() / model.eval()model.train() at inference, run T passes, report mean and std of softmaxnn.CrossEntropyLoss(label_smoothing=eps) or (1-eps)*NLL + eps*uniform; eps=0.1 is typicalElimination
Eliminate the wrong options
Inverted dropout multiplies surviving activations by 1/(1-p) during training. What is the purpose of this scaling?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: With p=0.5, on average half the neurons are zeroed — without rescaling, the expected activation magnitude drops by half. The 1/(1-p) factor restores E[h_tilde]=E[h], so the full network at test time sees the same expected magnitude and needs no separate scaling step.
Check
Why scale by 1/(1-p) during training?
Check your understanding
Inverted dropout multiplies surviving activations by 1/(1-p) during training. What is the purpose of this scaling?
Answer: A
Why: With p=0.5, on average half the neurons are zeroed — without rescaling, the expected activation magnitude drops by half. The 1/(1-p) factor restores E[h_tilde]=E[h], so the full network at test time sees the same expected magnitude and needs no separate scaling step.
Prediction
Predict first
In MC Dropout with T=100 forward passes, a test point has mean P(class=1)=0.52 and std=0.24. What does the high std indicate?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: The model is epistemically uncertain about this input — it lies near a decision boundary or in a low-data region
Why: High std across MC Dropout passes means the thinned sub-networks disagree — a proxy for epistemic (model) uncertainty. It indicates the input lies in a region where the model is poorly constrained, often near a decision boundary or outside the training distribution. The prediction could still be correct.
Check
What does high std across T passes tell you?
Check your understanding
In MC Dropout with T=100 forward passes, a test point has mean P(class=1)=0.52 and std=0.24. What does the high std indicate?
Answer: A
Why: High std across MC Dropout passes means the thinned sub-networks disagree — a proxy for epistemic (model) uncertainty. It indicates the input lies in a region where the model is poorly constrained, often near a decision boundary or outside the training distribution. The prediction could still be correct.
Elimination
Eliminate the wrong options
What problem does label smoothing directly address in standard cross-entropy training?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Standard cross-entropy with one-hot targets is minimized by driving the logit of the true class to +infinity — the model never gains from being less than maximally confident. Label smoothing replaces 1 with 1-eps, capping the reward for extreme logits and producing calibrated probabilities.
Check
What does label smoothing prevent?
Check your understanding
What problem does label smoothing directly address in standard cross-entropy training?
Answer: A
Why: Standard cross-entropy with one-hot targets is minimized by driving the logit of the true class to +infinity — the model never gains from being less than maximally confident. Label smoothing replaces 1 with 1-eps, capping the reward for extreme logits and producing calibrated probabilities.
Section
Project
Concept
Build three things from scratch: (1) an inverted-dropout function, (2) MC Dropout inference with uncertainty reporting, and (3) a label smoothing loss. Then measure how dropout rate affects the train/test gap.
| # | milestone | deliverable |
|---|---|---|
| 1 | Inverted dropout function | manual mask + scale; verify expected value ~1.0 |
| 2 | MC Dropout inference | T=100 passes, report mean+std per test point |
| 3 | Label smoothing loss | scratch impl == nn.CrossEntropyLoss(label_smoothing=0.1) |
Rules: implement dropout yourself before using nn.Dropout. Call model.eval() only after you understand when NOT to call it (MC Dropout). Type every line — no copy-paste.
Comparison
Comparison matrix
From Project: Dropout Lab: refill the milestone column from what you know. The rest of the table is as it appeared.
| # | milestone | deliverable |
|---|---|---|
| 1 | Inverted dropout function | manual mask + scale; verify expected value ~1.0 |
| 2 | MC Dropout inference | T=100 passes, report mean+std per test point |
| 3 | Label smoothing loss | scratch impl == nn.CrossEntropyLoss(label_smoothing=0.1) |
Worked example
Your turn: write inverted_dropout(x, p, training) that applies the mask in train mode and is a no-op in eval mode. Test that the mean is preserved.
Hint: mask = (torch.rand_like(x) > p).float(). In train mode multiply by mask / (1-p); in eval mode return x unchanged.
import torch
def inverted_dropout(x, p, training):
if not training or p == 0.0:
return x
mask = (torch.rand_like(x) > p).float()
return x * mask / (1.0 - p)
torch.manual_seed(1)
x_big = torch.ones(100000)
out = inverted_dropout(x_big, p=0.3, training=True)
print(f'mean (train, p=0.3): {out.mean().item():.4f}') # ~1.0
print(f'mean (eval): {inverted_dropout(x_big, 0.3, False).mean().item():.4f}')| mode | p | mean output |
|---|---|---|
| train | 0.3 | ~1.0020 |
| eval | 0.3 | 1.0000 |
| train | 0.5 | ~1.0027 |
Pattern
Step through it
Step through Milestone 1 — inverted dropout one row at a time. What is driving the change, and what would the row after the last one be?
Worked example
Your turn: train a small net on make_classification, then run 100 forward passes with dropout active. Report mean and std for 5 test points. Predict which points will have high std.
Hint: after training call model.train() (NOT eval) before the T-pass loop. Stack softmax outputs into a (T, n_points) tensor.
import torch, torch.nn as nn, torch.nn.functional as F
# (Assumes Xtr/Xte/ytr/yte tensors from earlier)
class MCNet(nn.Module):
def __init__(self, p=0.3):
super().__init__()
self.net = nn.Sequential(nn.Linear(20,64), nn.ReLU(),
nn.Dropout(p), nn.Linear(64,2))
def forward(self, x): return self.net(x)
torch.manual_seed(0)
net = MCNet(); opt = torch.optim.Adam(net.parameters(), lr=1e-3)
for _ in range(200):
net.train(); opt.zero_grad()
nn.CrossEntropyLoss()(net(Xtr), ytr).backward(); opt.step()
net.train() # keep dropout ON for MC inference
torch.manual_seed(5); pts = Xte[:5]
probs = torch.stack([F.softmax(net(pts),dim=-1)[:,1] for _ in range(100)])
print('mean:', probs.mean(0).tolist())
print('std: ', probs.std(0).tolist())| point | mean P(1) | std | verdict |
|---|---|---|---|
| 0 | 0.0139 | 0.0170 | confident class 0 |
| 1 | 0.8909 | 0.1010 | fairly confident class 1 |
| 2 | 0.8159 | 0.1256 | uncertain |
| 3 | 0.2017 | 0.0787 | leans class 0 |
| 4 | 0.0015 | 0.0076 | very confident class 0 |
Trade off
Comparison matrix
From Milestone 2 — MC Dropout uncertainty: every row here is a choice with a cost. Fill the std column, then say which row you would actually pick and what you give up for it.
| point | mean P(1) | std | verdict |
|---|---|---|---|
| 0 | 0.0139 | 0.0170 | confident class 0 |
| 1 | 0.8909 | 0.1010 | fairly confident class 1 |
| 2 | 0.8159 | 0.1256 | uncertain |
| 3 | 0.2017 | 0.0787 | leans class 0 |
| 4 | 0.0015 | 0.0076 | very confident class 0 |
Worked example
Your turn: implement smooth_ce(logits, targets, eps, K) using log_softmax. Predict whether it equals PyTorch's builtin before running.
Hint: loss = (1-eps)*NLL + eps*(-log_p.mean(dim=-1)).mean(). The eps*uniform term spreads probability mass.
import torch, torch.nn as nn, torch.nn.functional as F
torch.manual_seed(7)
logits = torch.randn(4, 5)
targets = torch.tensor([0, 2, 1, 4])
eps = 0.1; K = 5
def smooth_ce(logits, targets, eps, K):
log_p = F.log_softmax(logits, dim=-1)
n = logits.shape[0]
nll = -log_p[range(n), targets]
uniform = -log_p.mean(dim=-1)
return ((1-eps)*nll + eps*uniform).mean()
print(f'std CE: {nn.CrossEntropyLoss()(logits, targets).item():.4f}')
print(f'scratch: {smooth_ce(logits, targets, eps, K).item():.4f}')
print(f'builtin: {nn.CrossEntropyLoss(label_smoothing=0.1)(logits, targets).item():.4f}')| loss | value |
|---|---|
| std CrossEntropyLoss | 1.7088 |
| smooth_ce (scratch) | 1.7369 |
| CrossEntropyLoss(label_smoothing=0.1) | 1.7369 |
| scratch == builtin | True |
Discrimination
Sort into buckets
Sort these by value, from memory, without looking back at Milestone 3 — label smoothing loss. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Concept
import torch, torch.nn as nn, torch.nn.functional as F
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
X, y = make_classification(n_samples=2000, n_features=20,
n_informative=10, random_state=0)
X_tr,X_te,y_tr,y_te = train_test_split(X, y, test_size=0.3, random_state=0)
Xtr = torch.tensor(X_tr, dtype=torch.float32)
Xte = torch.tensor(X_te, dtype=torch.float32)
ytr = torch.tensor(y_tr, dtype=torch.long)
yte = torch.tensor(y_te, dtype=torch.long)
# Compare p=0.0 vs p=0.1
for p in [0.0, 0.1]:
torch.manual_seed(0)
net = nn.Sequential(nn.Linear(20,64),nn.ReLU(),nn.Dropout(p),
nn.Linear(64,64),nn.ReLU(),nn.Dropout(p),nn.Linear(64,2))
opt = torch.optim.Adam(net.parameters(), lr=1e-3)
for _ in range(200):
net.train(); opt.zero_grad()
nn.CrossEntropyLoss(label_smoothing=0.1)(net(Xtr),ytr).backward(); opt.step()
net.eval()
with torch.no_grad():
tr = (net(Xtr).argmax(1)==ytr).float().mean().item()
te = (net(Xte).argmax(1)==yte).float().mean().item()
print(f'p={p}: train={tr:.4f}, test={te:.4f}')| p | train acc | test acc | gap |
|---|---|---|---|
| 0.0 | 0.9964 | 0.9383 | 0.0581 |
| 0.1 | 0.9836 | 0.9500 | 0.0336 |
If your p=0.1 net closes the generalization gap to ~0.034 while maintaining test accuracy — you have shipped a regularized network with calibrated labels.
Comparison
Comparison matrix
From The full program: refill the gap column from what you know. The rest of the table is as it appeared.
| p | train acc | test acc | gap |
|---|---|---|---|
| 0.0 | 0.9964 | 0.9383 | 0.0581 |
| 0.1 | 0.9836 | 0.9500 | 0.0336 |
Concept
Out loud, slides closed: explain (1) why inverted dropout needs no test-time scaling, (2) why MC Dropout uncertainty requires leaving train mode ON, and (3) how label smoothing prevents overconfident logits.
Stretch (homework): implement MC Dropout calibration — compute expected calibration error (ECE) with and without T-pass averaging. Next up: ensemble methods (bagging, boosting) — the ensemble insight from this lesson will connect directly.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Dropout: theory · MC Dropout: uncertainty · Label smoothing & other regularizers · Your turn: build it. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
mask*(x/(1-p)) train, x unchanged evalmodel.train() / model.eval() correctly — and deliberately break the rule for MC Dropoutnn.CrossEntropyLoss(label_smoothing=eps)| concept | the one thing to remember |
|---|---|
| inverted dropout | scale by 1/(1-p) during training; no scaling at test |
| train vs eval mode | eval() disables dropout — don't call it before MC passes |
| MC Dropout | std across T passes = epistemic uncertainty |
| label smoothing | (1-eps)NLL + epsuniform; eps=0.1 is typical |
| dropout rate | p=0.1-0.3 for hidden layers; validate on test, not train |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.