Lesson 48: Dropout & Modern Regularizers

USAAIO Lesson 48, from Week 17 of Phase 2. It covers the theory and implementation of inverted dropout, the difference between training and evaluation mode, dropout read as averaging over an exponential ensemble, MC Dropout for Bayesian uncertainty estimation, and label smoothing from scratch, then compares several dropout rates on a classification network. All the numbers were verified with torch 2.7.1 and sklearn in June 2026. The lesson runs to 29 slides.

Subject: Machine Learning · 59 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Dropout & Modern Regularizers

Title

USAAIO · Lesson 48 · Week 17 (Phase 2)

Inverted dropout, MC Dropout uncertainty, label smoothing, mixup — the regularizers that keep deep nets from memorizing the training set.

2. By the end of this lesson you can

Objectives

  1. Implement inverted dropout from scratch, handling train vs eval mode correctly
  2. Explain dropout as exponential ensemble averaging over thinned networks
  3. Run MC Dropout inference (T forward passes) and interpret prediction variance as uncertainty
  4. Implement label smoothing from scratch and verify it matches nn.CrossEntropyLoss(label_smoothing=eps)
  5. Choose a dropout rate by experiment and explain why p=0.5 is the theoretical sweet spot

3. What survived from Decision Trees — CART, Information Gain, Gini, and Pruning?

Warm-up

Discussion prompt

Before we open Lesson 48: Dropout & Modern Regularizers: without looking back, what was the main idea of Decision Trees — CART, Information Gain, Gini, and Pruning, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

information gain from entropy, Gini impurity, the CART binary-split induction algorithm, stopping criteria, and cost-complexity pruning (ccp_alpha). Build a DecisionTreeClassifier from raw splits to a pruned final model on iris and make_classification.

4. Dropout: theory

Section

Part 1 of 3

5. Dropout: what it does

Concept

At each forward pass during training, each neuron is independently zeroed with probability p. The surviving neurons are rescaled so the expected activation is unchanged.

\[ \tilde{h}_i = \frac{m_i \cdot h_i}{1-p}, \quad m_i \sim \text{Bernoulli}(1-p) \]

At test time no neurons are dropped — the full network runs. The /(1-p) scaling (inverted dropout) is already absorbed into weights, so inference needs no correction.

6. Break it if you can: Dropout: what it does

Counterexample

Discussion prompt

At each forward pass during training, each neuron is independently zeroed with probability p. The surviving neurons are rescaled so the expected activation is unchanged.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

At test time no neurons are dropped — the full network runs. The /(1-p) scaling (inverted dropout) is already absorbed into weights, so inference needs no correction.

7. Why dropout prevents co-adaptation

Concept

A neuron can't rely on a specific neighbor always being present. Each neuron must learn a feature that is useful on its own, not in concert with a fixed set of collaborators.

This forces the network to learn redundant, distributed representations of each concept — exactly what makes it generalize. Deep overfitting often traces to neurons that specialize to co-dependencies in the training batch.

8. By analogy: Why dropout prevents co-adaptation

Analogy

Discussion prompt

Explain Why dropout prevents co-adaptation by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

A neuron can't rely on a specific neighbor always being present. Each neuron must learn a feature that is useful on its own, not in concert with a fixed set of collaborators.

9. Dropout as an exponential ensemble

Intuition

A network with n neurons has 2^n possible thinned sub-networks. Each training step samples and trains one of them.

At test time, using the full network with no dropout approximates averaging the predictions of all 2^n sub-networks — a geometric mean over an exponentially large ensemble (Srivastava et al., 2014).

p=0.5 maximizes the number of distinct thinned networks, which is why it is the canonical starting point for hidden layers (p=0.1–0.2 is common for inputs to avoid discarding too much signal).

10. Teach it back: Dropout as an exponential ensemble

Explain it

Discussion prompt

Explain Dropout as an exponential ensemble to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

A network with n neurons has 2^n possible thinned sub-networks. Each training step samples and trains one of them.

11. Guess the shape of the answer: Inverted dropout from scratch

Estimation

Predict first

Implement inverted dropout manually. Seed 0, p=0.5, four input neurons — trace the mask, the scaled output in train mode, and the unchanged output in eval mode.

Commit before you compute: what does Inverted dropout from scratch come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: mask=[0,1,0,0]; neuron 1 survives and is scaled from 2.0 to 4.0; others zeroed

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Multiplying by 1/(1-0.5)=2 compensates for the expected 50% zeroing, preserving the expected activation magnitude at each layer.

12. Inverted dropout from scratch

Worked example

Implement inverted dropout manually. Seed 0, p=0.5, four input neurons — trace the mask, the scaled output in train mode, and the unchanged output in eval mode.

import torch
torch.manual_seed(0)
x = torch.tensor([1.0, 2.0, 3.0, 4.0])
p = 0.5
# train mode: sample mask, scale by 1/(1-p)
mask = (torch.rand_like(x) > p).float()
x_train = x * mask / (1.0 - p)
# eval mode: no dropout, no scaling
x_eval = x
print('mask  :', mask.tolist())
print('train :', x_train.tolist())
print('eval  :', x_eval.tolist())

mask=[0,1,0,0]; neuron 1 survives and is scaled from 2.0 to 4.0; others zeroed

Why: Multiplying by 1/(1-0.5)=2 compensates for the expected 50% zeroing, preserving the expected activation magnitude at each layer.

neuronxmasktrain outputeval output
01.000.01.0
12.014.02.0
23.000.03.0
34.000.04.0

13. Fill in: x for Inverted dropout from scratch

Comparison

Comparison matrix

From Inverted dropout from scratch: refill the x column from what you know. The rest of the table is as it appeared.

neuronxmasktrain outputeval output
01.000.01.0
12.014.02.0
23.000.03.0
34.000.04.0

14. Something is wrong here: not switching to eval() at inference

Anomaly

Predict first

A student writes this, and it looks reasonable:

The model trained fine, so just call model(x_test) directly.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: PyTorch's nn.Dropout checks self.training.

Call model.eval() before any inference, and model.train() to resume training.

Why: PyTorch's nn.Dropout checks self.training. If the model is still in train mode, it randomly zeros neurons during inference — predictions become stochastic and wrong every call.

15. Trap: not switching to eval() at inference

Trap

The trap

The model trained fine, so just call model(x_test) directly.

Run inference without calling model.eval()

Why: PyTorch's nn.Dropout checks self.training. If the model is still in train mode, it randomly zeros neurons during inference — predictions become stochastic and wrong every call.

The fix

Call model.eval() before any inference, and model.train() to resume training.

Call model.eval() then wrap inference in torch.no_grad()

Why: eval() disables dropout (and batch norm running stats). no_grad() saves memory. Both are required for correct, efficient inference.

16. Break it on purpose: not switching to eval() at inference

Break the constraint

Discussion prompt

The rule this trap just fixed:

eval() disables dropout (and batch norm running stats). no_grad() saves memory. Both are required for correct, efficient inference.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

PyTorch's nn.Dropout checks self.training. If the model is still in train mode, it randomly zeros neurons during inference — predictions become stochastic and wrong every call.

17. What has to be given first: Dropout rates vs accuracy

Missing information

Discussion prompt

Train a two-hidden-layer ReLU net (64 → 64 → 2) on make_classification (2000 samples, 20 features) with dropout rates 0.0 through 0.7. Observe the train/test accuracy gap.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

No dropout memorizes training data. Light dropout (p=0.1) regularizes without crippling capacity. Very high p (0.7) under-trains the network too aggressively, hurting both.

18. Dropout rates vs accuracy

Worked example

Train a two-hidden-layer ReLU net (64 → 64 → 2) on make_classification (2000 samples, 20 features) with dropout rates 0.0 through 0.7. Observe the train/test accuracy gap.

import torch, torch.nn as nn
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split

X, y = make_classification(n_samples=2000, n_features=20,
                            n_informative=10, random_state=0)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.3, random_state=0)
Xtr = torch.tensor(X_tr, dtype=torch.float32)
Xte = torch.tensor(X_te, dtype=torch.float32)
ytr = torch.tensor(y_tr, dtype=torch.long)
yte = torch.tensor(y_te, dtype=torch.long)

results = {}
for p in [0.0, 0.1, 0.3, 0.5, 0.7]:
    torch.manual_seed(0)
    net = nn.Sequential(nn.Linear(20,64), nn.ReLU(), nn.Dropout(p),
                        nn.Linear(64,64), nn.ReLU(), nn.Dropout(p),
                        nn.Linear(64,2))
    opt = torch.optim.Adam(net.parameters(), lr=1e-3)
    for _ in range(200):
        net.train(); opt.zero_grad()
        nn.CrossEntropyLoss()(net(Xtr), ytr).backward(); opt.step()
    net.eval()
    with torch.no_grad():
        tr = (net(Xtr).argmax(1)==ytr).float().mean().item()
        te = (net(Xte).argmax(1)==yte).float().mean().item()
    results[p] = (round(tr,4), round(te,4))
    print(f'p={p}: train={tr:.4f} test={te:.4f}')

p=0.0 overfits: train=0.9964 vs test=0.9383; p=0.1 closes the gap: train=0.9836, test=0.9500 (best test)

Why: No dropout memorizes training data. Light dropout (p=0.1) regularizes without crippling capacity. Very high p (0.7) under-trains the network too aggressively, hurting both.

ptrain acctest accgap
0.00.99640.93830.0581
0.10.98360.95000.0336
0.30.96140.93830.0231
0.50.93290.91500.0179
0.70.90290.89500.0079

19. What each one costs: Dropout rates vs accuracy

Trade off

Comparison matrix

From Dropout rates vs accuracy: every row here is a choice with a cost. Fill the gap column, then say which row you would actually pick and what you give up for it.

ptrain acctest accgap
0.00.99640.93830.0581
0.10.98360.95000.0336
0.30.96140.93830.0231
0.50.93290.91500.0179
0.70.90290.89500.0079

20. MC Dropout: uncertainty

Section

Part 2 of 3

21. Dropout as approximate Bayesian inference

Concept

Gal & Ghahramani (2016) showed that a neural network with dropout approximates a Bayesian neural network — each forward pass draws a sample from the approximate posterior over weights.

\[ p(y^* \mid x^*, X, Y) \approx \frac{1}{T}\sum_{t=1}^{T} p(y^* \mid x^*, \hat{W}_t) \]

Run T forward passes with dropout active at test time. The mean gives the prediction; the variance gives the model's epistemic uncertainty (how confident it is about this region of input space).

22. By analogy: Dropout as approximate Bayesian inference

Analogy

Discussion prompt

Explain Dropout as approximate Bayesian inference by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Gal & Ghahramani (2016) showed that a neural network with dropout approximates a Bayesian neural network — each forward pass draws a sample from the approximate posterior over weights.

23. Guess the shape of the answer: MC Dropout uncertainty estimation

Estimation

Predict first

Keep dropout active at test time (model.train()), run 100 forward passes, and report mean and std of the class-1 probability. High std = uncertain.

Commit before you compute: what does MC Dropout uncertainty estimation come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Points 0 and 4 (mean ~0.01-0.00, std ~0.02-0.01) are confidently class 0; point 2 (mean 0.816, std 0.126) is near a decision boundary

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. High std signals the model is uncertain — these are regions where the ensemble of thinned networks disagrees, indicating the model needs more data or a different feature near this boundary.

24. MC Dropout uncertainty estimation

Worked example

Keep dropout active at test time (model.train()), run 100 forward passes, and report mean and std of the class-1 probability. High std = uncertain.

import torch, torch.nn as nn, torch.nn.functional as F

class MCNet(nn.Module):
    def __init__(self, p=0.3):
        super().__init__()
        self.fc1 = nn.Linear(20, 64)
        self.drop = nn.Dropout(p)
        self.fc2 = nn.Linear(64, 2)
    def forward(self, x):
        return self.fc2(self.drop(F.relu(self.fc1(x))))

torch.manual_seed(0)
net = MCNet(p=0.3)
opt = torch.optim.Adam(net.parameters(), lr=1e-3)
for _ in range(200):
    net.train(); opt.zero_grad()
    nn.CrossEntropyLoss()(net(Xtr), ytr).backward(); opt.step()

# MC inference: keep train mode so dropout stays ON
net.train()
T = 100; torch.manual_seed(5); sample = Xte[:5]
probs = torch.stack([F.softmax(net(sample),dim=-1)[:,1]
                     for _ in range(T)])
print('mean:', probs.mean(0).round(decimals=4).tolist())
print('std: ', probs.std(0).round(decimals=4).tolist())

Points 0 and 4 (mean ~0.01-0.00, std ~0.02-0.01) are confidently class 0; point 2 (mean 0.816, std 0.126) is near a decision boundary

Why: High std signals the model is uncertain — these are regions where the ensemble of thinned networks disagrees, indicating the model needs more data or a different feature near this boundary.

pointmean P(class=1)stdinterpretation
00.01390.0170confident class 0
10.89090.1010fairly confident class 1
20.81590.1256uncertain (near boundary)
30.20170.0787leans class 0
40.00150.0076very confident class 0

25. Work backwards from the answer: MC Dropout uncertainty estimation

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Points 0 and 4 (mean ~0.01-0.00, std ~0.02-0.01) are confidently class 0; point 2 (mean 0.816, std 0.126) is near a decision boundary

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Keep dropout active at test time (model.train()), run 100 forward passes, and report mean and std of the class-1 probability. High std = uncertain.

26. Something is wrong here: calling eval() before MC Dropout

Anomaly

Predict first

A student writes this, and it looks reasonable:

To get predictions, always call model.eval() first — then run the T forward passes.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: eval() disables dropout, making every forward pass identical.

For MC Dropout, keep the model in train() mode (or use a custom flag) so dropout stays active during the T inference passes.

Why: eval() disables dropout, making every forward pass identical. You compute the same deterministic output T times — all variance is zero, uncertainty is useless.

27. Trap: calling eval() before MC Dropout

Trap

The trap

To get predictions, always call model.eval() first — then run the T forward passes.

Call model.eval() then loop T forward passes expecting stochastic outputs

Why: eval() disables dropout, making every forward pass identical. You compute the same deterministic output T times — all variance is zero, uncertainty is useless.

The fix

For MC Dropout, keep the model in train() mode (or use a custom flag) so dropout stays active during the T inference passes.

Keep model.train() during the T-pass loop, then call model.eval() for normal inference afterward

Why: Dropout is the source of stochasticity. eval() disables it. MC Dropout deliberately exploits that stochasticity to sample from the weight posterior — you must leave it on.

28. Which of these survive contact with Lesson 48: Dropout & Modern Regularizers?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
At each forward pass during training, each neuron is independently zeroed with probability p. The surviving neurons are rescaled so the expected activation is unchanged.; A network with n neurons has 2^n possible thinned sub-networks. Each training step samples and trains one of them.; Standard one-hot targets push the model to assign zero probability to all non-true classes, driving logits to large magnitudes and producing an overconfident model.
Breaks
The model trained fine, so just call model(x_test) directly.; To get predictions, always call model.eval() first — then run the T forward passes.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 48: Dropout & Modern Regularizers puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

29. Label smoothing & other regularizers

Section

Part 3 of 3

30. Label smoothing: softer targets

Concept

Standard one-hot targets push the model to assign zero probability to all non-true classes, driving logits to large magnitudes and producing an overconfident model.

\[ y_k^{\text{smooth}} = \begin{cases} 1 - \varepsilon & k = \text{true class} \\ \dfrac{\varepsilon}{K-1} & k \neq \text{true class} \end{cases} \]

With K=10 and eps=0.1: the true-class target becomes 0.9, and each other class gets 0.0111. The model is penalized for being too certain — it learns calibrated probabilities.

31. Break it if you can: Label smoothing: softer targets

Counterexample

Discussion prompt

Standard one-hot targets push the model to assign zero probability to all non-true classes, driving logits to large magnitudes and producing an overconfident model.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

With K=10 and eps=0.1: the true-class target becomes 0.9, and each other class gets 0.0111. The model is penalized for being too certain — it learns calibrated probabilities.

32. Guess the shape of the answer: Label smoothing from scratch

Estimation

Predict first

Implement label smoothing using log_softmax and a convex combination of NLL loss and uniform loss. Verify it matches PyTorch's built-in label_smoothing argument.

Commit before you compute: what does Label smoothing from scratch come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: std CE=1.7088, smooth CE (scratch)=1.7369, builtin=1.7369, match=True

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Smooth CE > standard CE on this batch because the model is forced to spread probability mass — the loss is always slightly higher.

33. Label smoothing from scratch

Worked example

Implement label smoothing using log_softmax and a convex combination of NLL loss and uniform loss. Verify it matches PyTorch's built-in label_smoothing argument.

import torch, torch.nn as nn, torch.nn.functional as F
torch.manual_seed(7)
logits = torch.randn(4, 5)          # 4 samples, 5 classes
targets = torch.tensor([0, 2, 1, 4])
eps = 0.1; K = 5

def smooth_ce(logits, targets, eps, K):
    log_p = F.log_softmax(logits, dim=-1)
    n = logits.shape[0]
    nll     = -log_p[range(n), targets]      # standard NLL per sample
    uniform = -log_p.mean(dim=-1)            # uniform term
    return ((1-eps)*nll + eps*uniform).mean()

std   = nn.CrossEntropyLoss()(logits, targets)
scratch = smooth_ce(logits, targets, eps, K)
builtin = nn.CrossEntropyLoss(label_smoothing=0.1)(logits, targets)
print(f'std CE:    {std.item():.4f}')
print(f'scratch:   {scratch.item():.4f}')
print(f'builtin:   {builtin.item():.4f}')
print(f'match: {abs(scratch.item()-builtin.item()) < 1e-5}')

std CE=1.7088, smooth CE (scratch)=1.7369, builtin=1.7369, match=True

Why: Smooth CE > standard CE on this batch because the model is forced to spread probability mass — the loss is always slightly higher. The exact equality confirms the derivation is correct.

loss variantvaluenotes
standard CE1.7088one-hot targets
label smooth (scratch)1.7369(1-eps)NLL + epsuniform
label smooth (builtin)1.7369nn.CrossEntropyLoss(label_smoothing=0.1)
matchTrueabs diff < 1e-5

34. Which is which, by value

Discrimination

Sort into buckets

Sort these by value, from memory, without looking back at Label smoothing from scratch. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

1.7088
standard CE
1.7369
label smooth (scratch); label smooth (builtin)
True
match
g1
value is "1.7088" for standard CE — that is what the table on "Label smoothing from scratch" records, and it is the single property separating this group from the rest.
g2
value is "1.7369" for label smooth (scratch), label smooth (builtin) — that is what the table on "Label smoothing from scratch" records, and it is the single property separating this group from the rest.
g3
value is "True" for match — that is what the table on "Label smoothing from scratch" records, and it is the single property separating this group from the rest.

35. Mixup and CutMix (brief)

Concept

Mixup (Zhang et al., 2018): linearly blend two training examples and their labels. The model must predict a soft mixture — a stronger form of label smoothing applied in input space.

\[ \tilde{x} = \lambda x_i + (1-\lambda) x_j, \quad \tilde{y} = \lambda y_i + (1-\lambda) y_j, \quad \lambda \sim \text{Beta}(\alpha,\alpha) \]

CutMix (Yun et al., 2019): paste a rectangular crop of one image into another, mixing labels proportionally to patch area. Both are dominant data augmentation strategies in vision competitions.

36. Where does each piece belong: Lesson 48: Dropout & Modern Regularizers

Sorting

Sort into buckets

These are the pieces of Lesson 48: Dropout & Modern Regularizers, out of order. Put each one back under the part of the lesson it belongs to.

Dropout: theory
Dropout: what it does; Why dropout prevents co-adaptation; Dropout as an exponential ensemble
MC Dropout: uncertainty
Dropout as approximate Bayesian inference; MC Dropout uncertainty estimation
Label smoothing & other regularizers
Label smoothing: softer targets; Label smoothing from scratch; Mixup and CutMix (brief)
s1
Dropout: theory is where Lesson 48: Dropout & Modern Regularizers puts Dropout: what it does, Why dropout prevents co-adaptation, Dropout as an exponential ensemble. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
MC Dropout: uncertainty is where Lesson 48: Dropout & Modern Regularizers puts Dropout as approximate Bayesian inference, MC Dropout uncertainty estimation. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
Label smoothing & other regularizers is where Lesson 48: Dropout & Modern Regularizers puts Label smoothing: softer targets, Label smoothing from scratch, Mixup and CutMix (brief). Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

37. Rebuild the recipe: The dropout & modern-regularizer recipe

Ranking

Put in order

These are the steps of The dropout & modern-regularizer recipe, scrambled. Put them back in order before the next slide shows you.

  1. Inverted dropout: mask*(x/(1-p)) in train; x unchanged in eval — use model.train() / model.eval()
  2. Rate search: start p=0.5 for hidden, p=0.1–0.2 for inputs; validate on held-out set, not training set
  3. MC Dropout uncertainty: keep model.train() at inference, run T passes, report mean and std of softmax
  4. Label smoothing: nn.CrossEntropyLoss(label_smoothing=eps) or (1-eps)*NLL + eps*uniform; eps=0.1 is typical
  5. Mixup/CutMix: blend two samples and their labels with a Beta(alpha,alpha) mixing coefficient

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

38. The dropout & modern-regularizer recipe

Pattern

  1. Inverted dropout: mask*(x/(1-p)) in train; x unchanged in eval — use model.train() / model.eval()
  2. Rate search: start p=0.5 for hidden, p=0.1–0.2 for inputs; validate on held-out set, not training set
  3. MC Dropout uncertainty: keep model.train() at inference, run T passes, report mean and std of softmax
  4. Label smoothing: nn.CrossEntropyLoss(label_smoothing=eps) or (1-eps)*NLL + eps*uniform; eps=0.1 is typical
  5. Mixup/CutMix: blend two samples and their labels with a Beta(alpha,alpha) mixing coefficient

39. Where does it stop working: The dropout & modern-regularizer recipe

Edge cases

Discussion prompt

The dropout & modern-regularizer recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Inverted dropout: mask*(x/(1-p)) in train; x unchanged in eval — use model.train() / model.eval()
  2. Rate search: start p=0.5 for hidden, p=0.1–0.2 for inputs; validate on held-out set, not training set
  3. MC Dropout uncertainty: keep model.train() at inference, run T passes, report mean and std of softmax
  4. Label smoothing: nn.CrossEntropyLoss(label_smoothing=eps) or (1-eps)*NLL + eps*uniform; eps=0.1 is typical
  5. Mixup/CutMix: blend two samples and their labels with a Beta(alpha,alpha) mixing coefficient

40. Rule out three: Check yourself — inverted dropout

Elimination

Eliminate the wrong options

Inverted dropout multiplies surviving activations by 1/(1-p) during training. What is the purpose of this scaling?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. It preserves the expected value of each activation, so test-time inference needs no correction
  • B. It makes the gradient larger so training converges faster
  • C. It ensures the mask values are binary (0 or 1)
  • D. It compensates for the reduced learning rate caused by zeroing neurons

Survives elimination: A

Why: With p=0.5, on average half the neurons are zeroed — without rescaling, the expected activation magnitude drops by half. The 1/(1-p) factor restores E[h_tilde]=E[h], so the full network at test time sees the same expected magnitude and needs no separate scaling step.

41. Check yourself — inverted dropout

Check

Why scale by 1/(1-p) during training?

Check your understanding

Inverted dropout multiplies surviving activations by 1/(1-p) during training. What is the purpose of this scaling?

  • A. It preserves the expected value of each activation, so test-time inference needs no correction (correct)
  • B. It makes the gradient larger so training converges faster
  • C. It ensures the mask values are binary (0 or 1)
  • D. It compensates for the reduced learning rate caused by zeroing neurons

Answer: A

Why: With p=0.5, on average half the neurons are zeroed — without rescaling, the expected activation magnitude drops by half. The 1/(1-p) factor restores E[h_tilde]=E[h], so the full network at test time sees the same expected magnitude and needs no separate scaling step.

Why B tempts people
The scaling is on activations, not on gradients. It does not directly increase the gradient magnitude.
Why C tempts people
The mask is already binary (Bernoulli draws). The 1/(1-p) term is purely an amplitude correction.
Why D tempts people
The learning rate is set separately in the optimizer; inverted dropout has no interaction with it.

42. Answer it before you see the options: Check yourself — MC Dropout

Prediction

Predict first

In MC Dropout with T=100 forward passes, a test point has mean P(class=1)=0.52 and std=0.24. What does the high std indicate?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: The model is epistemically uncertain about this input — it lies near a decision boundary or in a low-data region

Why: High std across MC Dropout passes means the thinned sub-networks disagree — a proxy for epistemic (model) uncertainty. It indicates the input lies in a region where the model is poorly constrained, often near a decision boundary or outside the training distribution. The prediction could still be correct.

43. Check yourself — MC Dropout

Check

What does high std across T passes tell you?

Check your understanding

In MC Dropout with T=100 forward passes, a test point has mean P(class=1)=0.52 and std=0.24. What does the high std indicate?

  • A. The model is epistemically uncertain about this input — it lies near a decision boundary or in a low-data region (correct)
  • B. The model's weights diverged during training
  • C. The dropout rate p is too low, causing variance in the mask
  • D. The prediction is wrong; high std always indicates a misclassification

Answer: A

Why: High std across MC Dropout passes means the thinned sub-networks disagree — a proxy for epistemic (model) uncertainty. It indicates the input lies in a region where the model is poorly constrained, often near a decision boundary or outside the training distribution. The prediction could still be correct.

Why B tempts people
Diverged weights would cause exploding logits uniformly, not localized variance on specific test points.
Why C tempts people
p only controls the per-neuron drop probability, not whether std is high or low for a given test point.
Why D tempts people
A point near the boundary can have mean ~0.52 and be correctly classified as class 1 while still being uncertain. Uncertainty and correctness are orthogonal.

44. Rule out three: Check yourself — label smoothing

Elimination

Eliminate the wrong options

What problem does label smoothing directly address in standard cross-entropy training?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Overconfidence: the model learns to push logits to ±infinity to perfectly satisfy one-hot targets
  • B. Class imbalance: smooth targets assign more weight to rare classes
  • C. Vanishing gradients: smoothed targets keep gradients large in deep networks
  • D. Overfitting to dropout masks: smoothing prevents the same neurons from being always selected

Survives elimination: A

Why: Standard cross-entropy with one-hot targets is minimized by driving the logit of the true class to +infinity — the model never gains from being less than maximally confident. Label smoothing replaces 1 with 1-eps, capping the reward for extreme logits and producing calibrated probabilities.

45. Check yourself — label smoothing

Check

What does label smoothing prevent?

Check your understanding

What problem does label smoothing directly address in standard cross-entropy training?

  • A. Overconfidence: the model learns to push logits to ±infinity to perfectly satisfy one-hot targets (correct)
  • B. Class imbalance: smooth targets assign more weight to rare classes
  • C. Vanishing gradients: smoothed targets keep gradients large in deep networks
  • D. Overfitting to dropout masks: smoothing prevents the same neurons from being always selected

Answer: A

Why: Standard cross-entropy with one-hot targets is minimized by driving the logit of the true class to +infinity — the model never gains from being less than maximally confident. Label smoothing replaces 1 with 1-eps, capping the reward for extreme logits and producing calibrated probabilities.

Why B tempts people
Label smoothing distributes eps equally across all non-true classes regardless of their frequency; it is not a class-imbalance fix.
Why C tempts people
Label smoothing targets overconfidence in outputs, not gradient magnitudes in intermediate layers.
Why D tempts people
Dropout and label smoothing are independent mechanisms; smoothing operates on the loss targets, not on the mask sampling.

46. Your turn: build it

Section

Project

47. Project: Dropout Lab

Concept

Build three things from scratch: (1) an inverted-dropout function, (2) MC Dropout inference with uncertainty reporting, and (3) a label smoothing loss. Then measure how dropout rate affects the train/test gap.

#milestonedeliverable
1Inverted dropout functionmanual mask + scale; verify expected value ~1.0
2MC Dropout inferenceT=100 passes, report mean+std per test point
3Label smoothing lossscratch impl == nn.CrossEntropyLoss(label_smoothing=0.1)

Rules: implement dropout yourself before using nn.Dropout. Call model.eval() only after you understand when NOT to call it (MC Dropout). Type every line — no copy-paste.

48. Fill in: milestone for Project: Dropout Lab

Comparison

Comparison matrix

From Project: Dropout Lab: refill the milestone column from what you know. The rest of the table is as it appeared.

#milestonedeliverable
1Inverted dropout functionmanual mask + scale; verify expected value ~1.0
2MC Dropout inferenceT=100 passes, report mean+std per test point
3Label smoothing lossscratch impl == nn.CrossEntropyLoss(label_smoothing=0.1)

49. Milestone 1 — inverted dropout

Worked example

Your turn: write inverted_dropout(x, p, training) that applies the mask in train mode and is a no-op in eval mode. Test that the mean is preserved.

Hint: mask = (torch.rand_like(x) > p).float(). In train mode multiply by mask / (1-p); in eval mode return x unchanged.

import torch
def inverted_dropout(x, p, training):
    if not training or p == 0.0:
        return x
    mask = (torch.rand_like(x) > p).float()
    return x * mask / (1.0 - p)

torch.manual_seed(1)
x_big = torch.ones(100000)
out = inverted_dropout(x_big, p=0.3, training=True)
print(f'mean (train, p=0.3): {out.mean().item():.4f}')  # ~1.0
print(f'mean (eval):         {inverted_dropout(x_big, 0.3, False).mean().item():.4f}')
modepmean output
train0.3~1.0020
eval0.31.0000
train0.5~1.0027

50. Watch it run: Milestone 1 — inverted dropout

Pattern

Step through it

Step through Milestone 1 — inverted dropout one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: mode is train
  2. Step 2: mode is eval
  3. Step 3: mode is train

51. Milestone 2 — MC Dropout uncertainty

Worked example

Your turn: train a small net on make_classification, then run 100 forward passes with dropout active. Report mean and std for 5 test points. Predict which points will have high std.

Hint: after training call model.train() (NOT eval) before the T-pass loop. Stack softmax outputs into a (T, n_points) tensor.

import torch, torch.nn as nn, torch.nn.functional as F
# (Assumes Xtr/Xte/ytr/yte tensors from earlier)
class MCNet(nn.Module):
    def __init__(self, p=0.3):
        super().__init__()
        self.net = nn.Sequential(nn.Linear(20,64), nn.ReLU(),
                                 nn.Dropout(p), nn.Linear(64,2))
    def forward(self, x): return self.net(x)
torch.manual_seed(0)
net = MCNet(); opt = torch.optim.Adam(net.parameters(), lr=1e-3)
for _ in range(200):
    net.train(); opt.zero_grad()
    nn.CrossEntropyLoss()(net(Xtr), ytr).backward(); opt.step()
net.train()   # keep dropout ON for MC inference
torch.manual_seed(5); pts = Xte[:5]
probs = torch.stack([F.softmax(net(pts),dim=-1)[:,1] for _ in range(100)])
print('mean:', probs.mean(0).tolist())
print('std: ', probs.std(0).tolist())
pointmean P(1)stdverdict
00.01390.0170confident class 0
10.89090.1010fairly confident class 1
20.81590.1256uncertain
30.20170.0787leans class 0
40.00150.0076very confident class 0

52. What each one costs: Milestone 2 — MC Dropout uncertainty

Trade off

Comparison matrix

From Milestone 2 — MC Dropout uncertainty: every row here is a choice with a cost. Fill the std column, then say which row you would actually pick and what you give up for it.

pointmean P(1)stdverdict
00.01390.0170confident class 0
10.89090.1010fairly confident class 1
20.81590.1256uncertain
30.20170.0787leans class 0
40.00150.0076very confident class 0

53. Milestone 3 — label smoothing loss

Worked example

Your turn: implement smooth_ce(logits, targets, eps, K) using log_softmax. Predict whether it equals PyTorch's builtin before running.

Hint: loss = (1-eps)*NLL + eps*(-log_p.mean(dim=-1)).mean(). The eps*uniform term spreads probability mass.

import torch, torch.nn as nn, torch.nn.functional as F
torch.manual_seed(7)
logits = torch.randn(4, 5)
targets = torch.tensor([0, 2, 1, 4])
eps = 0.1; K = 5

def smooth_ce(logits, targets, eps, K):
    log_p = F.log_softmax(logits, dim=-1)
    n = logits.shape[0]
    nll     = -log_p[range(n), targets]
    uniform = -log_p.mean(dim=-1)
    return ((1-eps)*nll + eps*uniform).mean()

print(f'std CE:    {nn.CrossEntropyLoss()(logits, targets).item():.4f}')
print(f'scratch:   {smooth_ce(logits, targets, eps, K).item():.4f}')
print(f'builtin:   {nn.CrossEntropyLoss(label_smoothing=0.1)(logits, targets).item():.4f}')
lossvalue
std CrossEntropyLoss1.7088
smooth_ce (scratch)1.7369
CrossEntropyLoss(label_smoothing=0.1)1.7369
scratch == builtinTrue

54. Which is which, by value

Discrimination

Sort into buckets

Sort these by value, from memory, without looking back at Milestone 3 — label smoothing loss. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

1.7088
std CrossEntropyLoss
1.7369
smooth_ce (scratch); CrossEntropyLoss(label_smoothing=0.1)
True
scratch == builtin
g1
value is "1.7088" for std CrossEntropyLoss — that is what the table on "Milestone 3 — label smoothing loss" records, and it is the single property separating this group from the rest.
g2
value is "1.7369" for smooth_ce (scratch), CrossEntropyLoss(label_smoothing=0.1) — that is what the table on "Milestone 3 — label smoothing loss" records, and it is the single property separating this group from the rest.
g3
value is "True" for scratch == builtin — that is what the table on "Milestone 3 — label smoothing loss" records, and it is the single property separating this group from the rest.

55. The full program

Concept

import torch, torch.nn as nn, torch.nn.functional as F
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split

X, y = make_classification(n_samples=2000, n_features=20,
                            n_informative=10, random_state=0)
X_tr,X_te,y_tr,y_te = train_test_split(X, y, test_size=0.3, random_state=0)
Xtr = torch.tensor(X_tr, dtype=torch.float32)
Xte = torch.tensor(X_te, dtype=torch.float32)
ytr = torch.tensor(y_tr, dtype=torch.long)
yte = torch.tensor(y_te, dtype=torch.long)

# Compare p=0.0 vs p=0.1
for p in [0.0, 0.1]:
    torch.manual_seed(0)
    net = nn.Sequential(nn.Linear(20,64),nn.ReLU(),nn.Dropout(p),
                        nn.Linear(64,64),nn.ReLU(),nn.Dropout(p),nn.Linear(64,2))
    opt = torch.optim.Adam(net.parameters(), lr=1e-3)
    for _ in range(200):
        net.train(); opt.zero_grad()
        nn.CrossEntropyLoss(label_smoothing=0.1)(net(Xtr),ytr).backward(); opt.step()
    net.eval()
    with torch.no_grad():
        tr = (net(Xtr).argmax(1)==ytr).float().mean().item()
        te = (net(Xte).argmax(1)==yte).float().mean().item()
    print(f'p={p}: train={tr:.4f}, test={te:.4f}')
ptrain acctest accgap
0.00.99640.93830.0581
0.10.98360.95000.0336

If your p=0.1 net closes the generalization gap to ~0.034 while maintaining test accuracy — you have shipped a regularized network with calibrated labels.

56. Fill in: gap for The full program

Comparison

Comparison matrix

From The full program: refill the gap column from what you know. The rest of the table is as it appeared.

ptrain acctest accgap
0.00.99640.93830.0581
0.10.98360.95000.0336

57. Show it off

Concept

Out loud, slides closed: explain (1) why inverted dropout needs no test-time scaling, (2) why MC Dropout uncertainty requires leaving train mode ON, and (3) how label smoothing prevents overconfident logits.

Stretch (homework): implement MC Dropout calibration — compute expected calibration error (ECE) with and without T-pass averaging. Next up: ensemble methods (bagging, boosting) — the ensemble insight from this lesson will connect directly.

58. Connect it up: Lesson 48: Dropout & Modern Regularizers

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Dropout: theory · MC Dropout: uncertainty · Label smoothing & other regularizers · Your turn: build it. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

59. What you can do now

Recap

conceptthe one thing to remember
inverted dropoutscale by 1/(1-p) during training; no scaling at test
train vs eval modeeval() disables dropout — don't call it before MC passes
MC Dropoutstd across T passes = epistemic uncertainty
label smoothing(1-eps)NLL + epsuniform; eps=0.1 is typical
dropout ratep=0.1-0.3 for hidden layers; validate on test, not train

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 48 (Week 17 — Ensemble Methods & Custom PyTorch) — Barron · USAAIO Round 2 Preparation, 2026
  2. Gal & Ghahramani, Dropout as a Bayesian Approximation, ICML 2016 — Gal, Y. & Ghahramani, Z. (2016)
  3. Muller et al., When Does Label Smoothing Help?, NeurIPS 2019 — Muller, R., Kornblith, S., & Hinton, G. (2019)
  4. Inverted dropout, MC Dropout, and label smoothing implementation verified with torch 2.7.1 + sklearn, June 2026 — Real execution output, June 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108