USAAIO Lesson 118, from Week 40 of Phase 4 on generative AI. It presents the autoencoder as an encoder feeding a bottleneck latent z feeding a decoder, trained on reconstruction loss alone, and covers undercomplete compression. It then covers the denoising autoencoder, which corrupts the input and reconstructs the clean signal, and the sparse autoencoder, which puts an L1 penalty on the activations and is the workhorse of mechanistic interpretability. It covers interpolation in latent space and ends on the key limitation that motivates the VAE: a standard autoencoder's latent space has no prior, so you cannot sample from it. You build a convolutional autoencoder yourself. The lesson runs to 26 slides.
Subject: Machine Learning · 57 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 118 · Phase 4 (Generative AI)
Compress to a bottleneck and reconstruct — then meet the variants that learn robust and interpretable features, and the limitation that forces us to invent the VAE next lesson.
Objectives
Warm-up
Discussion prompt
Before we open Lesson 118: Autoencoders — Standard, Denoising & Sparse: without looking back, what was the main idea of Phase 3 Gap Analysis & Phase 4 Preview, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
structured gap analysis across all nine Phase 3 topic clusters (transformers, BERT/GPT, ViT, GNNs, NLP, CV, Seq2Seq, BLEU/ROUGE, positional encoding), a scored self-assessment, study-plan adjustment for weak areas, Phase 4 overview (VAE / ELBO, GAN minimax, DDPM noise schedule), and exam strategy (time allocation, derivation approach, mindset).
Section
Part 1 of 4
Concept
An autoencoder learns to copy its input to its output through a narrow middle. The encoder maps x to a low-dimensional code z; the decoder maps z back to a reconstruction x_hat.
\[ z = f_{\text{enc}}(x), \qquad \hat{x} = f_{\text{dec}}(z) \]
bottleneck (latent code) — The low-dimensional vector z. Because it is too small to hold a verbatim copy, the network must keep only what matters to reconstruct x.
Counterexample
Discussion prompt
An autoencoder learns to copy its input to its output through a narrow middle. The encoder maps x to a low-dimensional code z; the decoder maps z back to a reconstruction x_hat.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Concept
There are no labels. The target is the input — the network is trained on how well x_hat matches x.
\[ \mathcal{L}_{\text{MSE}} = \frac{1}{N}\sum_i \lVert x_i - \hat{x}_i\rVert^2 \qquad \mathcal{L}_{\text{BCE}} = -\sum_j x_j\log \hat{x}_j + (1-x_j)\log(1-\hat{x}_j) \]
Use MSE for real-valued inputs and BCE for inputs in [0,1] (e.g. normalized pixel intensities). That is the whole objective — self-supervised, no annotation needed.
Analogy
Discussion prompt
Explain The loss is reconstruction only by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Use MSE for real-valued inputs and BCE for inputs in [0,1] (e.g. normalized pixel intensities). That is the whole objective — self-supervised, no annotation needed.
Concept
The latent dimension d_z relative to the input dimension d_x decides what the AE learns.
| d_z vs d_x | name | what happens |
|---|---|---|
| d_z < d_x | undercomplete | forced compression — must discard redundancy, learns structure |
| d_z = d_x | complete | can learn identity — risk of learning nothing useful |
| d_z > d_x | overcomplete | can copy verbatim — needs a regularizer (sparsity/denoising) to be useful |
The undercomplete bottleneck is the classic case: the squeeze is the teacher.
Comparison
Comparison matrix
From Undercomplete vs overcomplete: refill the name column from what you know. The rest of the table is as it appeared.
| d_z vs d_x | name | what happens |
|---|---|---|
| d_z < d_x | undercomplete | forced compression — must discard redundancy, learns structure |
| d_z = d_x | complete | can learn identity — risk of learning nothing useful |
| d_z > d_x | overcomplete | can copy verbatim — needs a regularizer (sparsity/denoising) to be useful |
Estimation
Predict first
An undercomplete AE (d_x=4, d_z=2). Hand-set weights so we can read every number. Input x = [1, 0, 1, 0].
Commit before you compute: what does Forward-pass trace: 4 -> 2 -> 4 come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Decode: x_hat = W_dec z = [1, 0, 1, 0]
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The decoder copies z0 into outputs 0 and 2, z1 into outputs 1 and 3.
Worked example
An undercomplete AE (d_x=4, d_z=2). Hand-set weights so we can read every number. Input x = [1, 0, 1, 0].
import torch
x = torch.tensor([[1.0, 0.0, 1.0, 0.0]])
W_enc = torch.tensor([[0.5,0,0.5,0],[0,0.5,0,0.5]]) # (2,4)
z = torch.relu(x @ W_enc.T) # bottleneck
W_dec = torch.tensor([[1.,0],[0,1.],[1.,0],[0,1.]]) # (4,2)
x_hat = z @ W_dec.T
print('z =', z.tolist())
print('x_hat=', x_hat.tolist())
print('MSE =', ((x-x_hat)**2).mean().item())Encode: z = ReLU(W_enc x) = [0.5+0.5, 0] = [1.0, 0.0]
Why: The encoder averages features (0,2) into z0 and (1,3) into z1; here only z0 fires.
Decode: x_hat = W_dec z = [1, 0, 1, 0]
Why: The decoder copies z0 into outputs 0 and 2, z1 into outputs 1 and 3.
| quantity | value | note |
|---|---|---|
| z (latent) | [1.0, 0.0] | 2 numbers stand in for 4 |
| x_hat | [1.0, 0.0, 1.0, 0.0] | exact reconstruction |
| reconstruction MSE | 0.0 | this structured input is perfectly recoverable from 2 dims |
Trade off
Comparison matrix
From Forward-pass trace: 4 -> 2 -> 4: every row here is a choice with a cost. Fill the value column, then say which row you would actually pick and what you give up for it.
| quantity | value | note |
|---|---|---|
| z (latent) | [1.0, 0.0] | 2 numbers stand in for 4 |
| x_hat | [1.0, 0.0, 1.0, 0.0] | exact reconstruction |
| reconstruction MSE | 0.0 | this structured input is perfectly recoverable from 2 dims |
Intuition
The AE is lossy compression fitted to your data. MNIST digits live on a thin manifold inside 784-pixel space; a good encoder finds coordinates on that manifold instead of storing raw pixels.
That is why a 2-number code can perfectly rebuild our 4-pixel input: the input had structure (features came in matched pairs), and the bottleneck captured exactly that structure.
Explain it
Discussion prompt
Explain What the bottleneck really learns to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
The AE is lossy compression fitted to your data. MNIST digits live on a thin manifold inside 784-pixel space; a good encoder finds coordinates on that manifold instead of storing raw pixels.
Section
Part 2 of 4
Concept
Corrupt the input with noise, feed the corrupted version in, but score the reconstruction against the clean original. The network can't just copy — it must learn what the signal should be.
\[ \tilde{x} = x + \epsilon,\ \epsilon\sim\mathcal{N}(0,\sigma^2 I); \qquad \mathcal{L} = \lVert x - f_{\text{dec}}(f_{\text{enc}}(\tilde{x}))\rVert^2 \]
The target in the loss is x (clean), never the corrupted input. This forces robust features and is a stepping stone to diffusion models (which denoise repeatedly).
Socratic
Discussion prompt
Corrupt the input with noise, feed the corrupted version in, but score the reconstruction against the clean original. The network can't just copy — it must learn what the signal should be.
Suppose that were not true. What is the first thing in Lesson 118: Autoencoders — Standard, Denoising & Sparse that would stop working?
Hint: Follow it one step downstream. The answer is whatever was quietly relying on it.
Answer:
The target in the loss is x (clean), never the corrupted input. This forces robust features and is a stepping stone to diffusion models (which denoise repeatedly).
Missing information
Discussion prompt
Add Gaussian noise (sigma=0.5) to a clean vector. Note which signal the loss compares against.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Input is corrupted, target is clean. Targeting the corrupted vector would just teach the identity and defeat the point.
Worked example
Add Gaussian noise (sigma=0.5) to a clean vector. Note which signal the loss compares against.
import torch
torch.manual_seed(1)
clean = torch.tensor([[1.0, 0.0, 1.0, 0.0]])
corrupted = clean + 0.5 * torch.randn_like(clean)
print('clean =', clean.tolist())
print('corrupted=', [round(v,4) for v in corrupted.flatten().tolist()])
# encoder sees `corrupted`; loss compares reconstruction to `clean`| vector | values | role |
|---|---|---|
| clean x | [1.0, 0.0, 1.0, 0.0] | loss TARGET |
| noise eps | [0.3307, 0.1335, 0.0308, 0.3107] | sampled N(0, 0.25) |
| corrupted x~ | [1.3307, 0.1335, 1.0308, 0.3107] | encoder INPUT |
Loss = ||clean - decode(encode(corrupted))||^2
Why: Input is corrupted, target is clean. Targeting the corrupted vector would just teach the identity and defeat the point.
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Loss = ||clean - decode(encode(corrupted))||^2
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Add Gaussian noise (sigma=0.5) to a clean vector. Note which signal the loss compares against.
Concept
Add an L1 penalty on the activations (not the weights). The network reconstructs well while keeping most latent units at exactly zero — a sparse code.
\[ \mathcal{L} = \lVert x-\hat{x}\rVert^2 + \lambda \sum_k \lvert a_k\rvert \]
Each input then activates only a few units, each standing for one human-readable feature. This is the backbone of modern mechanistic interpretability (sparse features extracted from LLM activations).
Explain it
Discussion prompt
Explain Sparse autoencoder to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Add an L1 penalty on the activations (not the weights). The network reconstructs well while keeping most latent units at exactly zero — a sparse code.
Estimation
Predict first
Compare a sparse code with a dense code of equal reconstruction power. Read the penalty and the fraction of units that are non-zero.
Commit before you compute: what does Sparse trace: L1 penalty and the active fraction come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Minimizing recon + lambda*L1 drives small activations to exactly 0
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. L1's constant-magnitude gradient pushes a small a_k to zero (unlike L2, which only shrinks it).
Worked example
Compare a sparse code with a dense code of equal reconstruction power. Read the penalty and the fraction of units that are non-zero.
import torch
sparse = torch.tensor([0.0, 0.0, 2.0, 0.0, 1.0])
dense = torch.tensor([0.8, 0.7, 0.9, 0.6, 1.0])
for name, a in [('sparse', sparse), ('dense', dense)]:
print(name, 'L1=', a.abs().sum().item(),
'active=', (a != 0).float().mean().item())| code | L1 = sum|a| | fraction active | reading |
|---|---|---|---|
| sparse [0,0,2,0,1] | 3.0 | 0.40 | 2 of 5 units fire |
| dense [0.8,0.7,0.9,0.6,1.0] | 4.0 | 1.00 | every unit fires |
Minimizing recon + lambda*L1 drives small activations to exactly 0
Why: L1's constant-magnitude gradient pushes a small a_k to zero (unlike L2, which only shrinks it). Sparsity is measured by the active fraction, not by total magnitude alone.
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Minimizing recon + lambda*L1 drives small activations to exactly 0
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Compare a sparse code with a dense code of equal reconstruction power. Read the penalty and the fraction of units that are non-zero.
Anomaly
Predict first
A student writes this, and it looks reasonable:
I corrupted the input, so the network should learn to reproduce what it was given.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Target = corrupted input. The optimum is the identity map — the AE learns to copy noise and learns nothing robust.
The corruption is only on the input path. The target stays the clean signal.
Why: Target = corrupted input. The optimum is the identity map — the AE learns to copy noise and learns nothing robust.
Trap
I corrupted the input, so the network should learn to reproduce what it was given.
loss = ||corrupted - decode(encode(corrupted))||^2
Why: Target = corrupted input. The optimum is the identity map — the AE learns to copy noise and learns nothing robust.
The corruption is only on the input path. The target stays the clean signal.
loss = ||clean - decode(encode(corrupted))||^2
Why: Target = clean input. The network must undo the noise, which is exactly the robust feature-learning we want.
Break the constraint
Discussion prompt
The rule this trap just fixed:
Target = clean input. The network must undo the noise, which is exactly the robust feature-learning we want.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
Target = corrupted input. The optimum is the identity map — the AE learns to copy noise and learns nothing robust.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Sparsity = add an L1 weight-decay term, like Lasso regression.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: This sparsifies the connections, giving a smaller/pruned network — it does NOT make the per-input latent code sparse.
A sparse autoencoder penalizes the activations produced for each input.
Why: This sparsifies the connections, giving a smaller/pruned network — it does NOT make the per-input latent code sparse.
Trap
Sparsity = add an L1 weight-decay term, like Lasso regression.
penalty = lambda * sum(|W|) (weights)
Why: This sparsifies the connections, giving a smaller/pruned network — it does NOT make the per-input latent code sparse.
A sparse autoencoder penalizes the activations produced for each input.
penalty = lambda * sum(|a|) (activations a = encoder output)
Why: Now each input lights up only a few latent units — the interpretable, per-example sparsity that defines a sparse AE.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
d_z relative to the input dimension d_x decides what the AE learns.; The target in the loss is x (clean), never the corrupted input. This forces robust features and is a stepping stone to diffusion models (which denoise repeatedly).Section
Part 3 of 4
Estimation
Predict first
Walk a straight line between two codes z_A and z_B and decode each step — the AE's version of morphing one example into another.
Commit before you compute: what does Latent interpolation come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Decode each z to get a smooth morph between the two reconstructions
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Interpolation is well-defined ONLY where the decoder has seen data.
Worked example
Walk a straight line between two codes z_A and z_B and decode each step — the AE's version of morphing one example into another.
import torch
zA = torch.tensor([0.0, 2.0]); zB = torch.tensor([2.0, 0.0])
for a in [0.0, 0.25, 0.5, 0.75, 1.0]:
print(a, ((1-a)*zA + a*zB).tolist())| alpha | z = (1-a)zA + a zB |
|---|---|
| 0.00 | [0.0, 2.0] |
| 0.25 | [0.5, 1.5] |
| 0.50 | [1.0, 1.0] |
| 0.75 | [1.5, 0.5] |
| 1.00 | [2.0, 0.0] |
Decode each z to get a smooth morph between the two reconstructions
Why: Interpolation is well-defined ONLY where the decoder has seen data. Off the manifold, decoded points are meaningless.
Comparison
Comparison matrix
From Latent interpolation: refill the z = (1-a)zA + a zB column from what you know. The rest of the table is as it appeared.
| alpha | z = (1-a)zA + a zB |
|---|---|
| 0.00 | [0.0, 2.0] |
| 0.25 | [0.5, 1.5] |
| 0.50 | [1.0, 1.0] |
| 0.75 | [1.5, 0.5] |
| 1.00 | [2.0, 0.0] |
Concept
To generate, you would draw a fresh z and decode it. But a standard AE imposes no prior on z — the encoder is free to scatter codes anywhere, leaving holes the decoder never visited.
import torch, torch.nn as nn
torch.manual_seed(2)
enc = nn.Sequential(nn.Linear(4,8), nn.ReLU(), nn.Linear(8,2))
data = torch.randn(200,4) @ torch.diag(torch.tensor([2.,2.,.1,.1]))
Z = enc(data).detach()
print('latent mean=', [round(v,3) for v in Z.mean(0).tolist()])
print('latent std =', [round(v,3) for v in Z.std(0).tolist()])
sample = torch.randn(2)
print('N(0,1) sample dist to nearest real code=',
round((Z-sample).pow(2).sum(1).min().sqrt().item(),3))| measurement | value | implication |
|---|---|---|
| latent mean | [-0.055, 0.511] | not centered at 0 |
| latent std | [0.212, 0.197] | not unit variance |
| N(0,1) sample -> nearest code | 1.168 | a standard-normal draw lands far from any real code |
Decode that far-away z and you get garbage — the decoder was never trained there. The VAE (Lesson 119) fixes exactly this by forcing q(z|x) toward N(0, I), so sampling z ~ N(0, I) is finally valid.
Sorting
Sort into buckets
These are the pieces of Lesson 118: Autoencoders — Standard, Denoising & Sparse, out of order. Put each one back under the part of the lesson it belongs to.
Section
Part 4 of 4
Ranking
Put in order
These are the steps of Choosing an autoencoder, scrambled. Put them back in order before the next slide shows you.
Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
Edge cases
Discussion prompt
Choosing an autoencoder works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
Elimination
Eliminate the wrong options
An autoencoder has input dimension d_x = 784 and latent dimension d_z = 32. Which statement is TRUE?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: d_z (32) < d_x (784), so the AE is undercomplete: it cannot store a verbatim copy and must compress, learning the data's structure.
Check
Decide it before clicking.
Check your understanding
An autoencoder has input dimension d_x = 784 and latent dimension d_z = 32. Which statement is TRUE?
Answer: A
Why: d_z (32) < d_x (784), so the AE is undercomplete: it cannot store a verbatim copy and must compress, learning the data's structure.
Elimination
Eliminate the wrong options
You train a denoising autoencoder. Which loss is correct, and can you later generate new images by sampling z ~ N(0, I) and decoding?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Denoising AEs target the clean signal (the network must undo the noise). And a standard/denoising AE imposes no prior on the latent, so codes cluster unpredictably — sampling z ~ N(0, I) lands in empty regions the decoder never learned. The VAE fixes this.
Check
One subtlety from each variant.
Check your understanding
You train a denoising autoencoder. Which loss is correct, and can you later generate new images by sampling z ~ N(0, I) and decoding?
Answer: A
Why: Denoising AEs target the clean signal (the network must undo the noise). And a standard/denoising AE imposes no prior on the latent, so codes cluster unpredictably — sampling z ~ N(0, I) lands in empty regions the decoder never learned. The VAE fixes this.
Section
Project
Concept
Build it yourself, one piece at a time. Type every line, run after each step, and read errors — do not erase them.
| requirement | tool you'll use |
|---|---|
| Encoder: 28x28 -> latent | nn.Conv2d, nn.ReLU, nn.Flatten, nn.Linear |
| Decoder: latent -> 28x28 | nn.Linear, nn.Unflatten, nn.ConvTranspose2d |
| Reconstruction loss | nn.MSELoss (or BCEWithLogits on [0,1] pixels) |
| Denoising variant | add torch.randn_like(x)*sigma to the INPUT only |
Analogy
Discussion prompt
Explain Project: a convolutional autoencoder for MNIST by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Build it yourself, one piece at a time. Type every line, run after each step, and read errors — do not erase them.
Worked example
Your turn: write an encoder that takes a (B,1,28,28) batch down to a (B,16) latent. Predict the output shape before you run.
Hint: two stride-2 Conv2d layers quarter then quarter the spatial size (28 -> 14 -> 7), then Flatten and a Linear to 16.
import torch, torch.nn as nn
encoder = nn.Sequential(
nn.Conv2d(1, 8, 3, stride=2, padding=1), nn.ReLU(), # 28->14
nn.Conv2d(8, 16, 3, stride=2, padding=1), nn.ReLU(), # 14->7
nn.Flatten(),
nn.Linear(16*7*7, 16))
z = encoder(torch.randn(4, 1, 28, 28))
print(z.shape)| stage | shape |
|---|---|
| input | (4, 1, 28, 28) |
| after conv1 (stride 2) | (4, 8, 14, 14) |
| after conv2 (stride 2) | (4, 16, 7, 7) |
| latent z | (4, 16) |
Trade off
Comparison matrix
From Milestone 1 — the encoder: every row here is a choice with a cost. Fill the shape column, then say which row you would actually pick and what you give up for it.
| stage | shape |
|---|---|
| input | (4, 1, 28, 28) |
| after conv1 (stride 2) | (4, 8, 14, 14) |
| after conv2 (stride 2) | (4, 16, 7, 7) |
| latent z | (4, 16) |
Worked example
Your turn: mirror the encoder back up to (B,1,28,28), then run one MSE reconstruction step. Predict: should the loss go down after opt.step()?
Hint: Linear(16 -> 16*7*7), Unflatten to (16,7,7), then two ConvTranspose2d(stride=2) to climb 7 -> 14 -> 28.
import torch, torch.nn as nn
decoder = nn.Sequential(
nn.Linear(16, 16*7*7), nn.Unflatten(1, (16,7,7)),
nn.ConvTranspose2d(16, 8, 4, stride=2, padding=1), nn.ReLU(), # 7->14
nn.ConvTranspose2d(8, 1, 4, stride=2, padding=1)) # 14->28
x = torch.rand(4, 1, 28, 28)
ae = nn.Sequential(encoder, decoder)
opt = torch.optim.Adam(ae.parameters(), lr=1e-3)
l1 = nn.MSELoss()(ae(x), x); l1.backward(); opt.step()
l2 = nn.MSELoss()(ae(x), x)
print(round(l1.item(),4), '->', round(l2.item(),4), 'down?', l2.item() < l1.item())| check | expected |
|---|---|
| decoder output shape | (4, 1, 28, 28) — matches input |
| loss after one Adam step | lower than before (down? True) |
| denoising tweak | feed ae(x + 0.5*torch.randn_like(x)), keep target x |
Comparison
Comparison matrix
From Milestone 2 — the decoder & a training step: refill the expected column from what you know. The rest of the table is as it appeared.
| check | expected |
|---|---|
| decoder output shape | (4, 1, 28, 28) — matches input |
| loss after one Adam step | lower than before (down? True) |
| denoising tweak | feed ae(x + 0.5*torch.randn_like(x)), keep target x |
Concept
Explain your program out loud, in order: where is the bottleneck? Which line would you change to make it denoising? Which to make it sparse?
Linear(... , 16) at the end of the encoder.ae(...); keep the target clean.lambda * z.abs().sum() to the loss — penalize the activations.Counterexample
Discussion prompt
Explain your program out loud, in order: where is the bottleneck? Which line would you change to make it denoising? Which to make it sparse?
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — The standard autoencoder · Denoising & sparse variants · The latent space & its limit · Recipe & checks · Your turn: build it. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
You can now build and reason about the three workhorse autoencoders — and you know precisely why they are not yet generative.
| variant | twist | buys you |
|---|---|---|
| standard (undercomplete) | bottleneck d_z < d_x | compression / visualization |
| denoising | corrupt input, target clean | robust features (-> diffusion) |
| sparse | L1 on activations | interpretable units (-> mech interp) |
| limitation | no prior on z | cannot sample -> motivates the VAE (L119) |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.