Lesson 118: Autoencoders — Standard, Denoising & Sparse

USAAIO Lesson 118, from Week 40 of Phase 4 on generative AI. It presents the autoencoder as an encoder feeding a bottleneck latent z feeding a decoder, trained on reconstruction loss alone, and covers undercomplete compression. It then covers the denoising autoencoder, which corrupts the input and reconstructs the clean signal, and the sparse autoencoder, which puts an L1 penalty on the activations and is the workhorse of mechanistic interpretability. It covers interpolation in latent space and ends on the key limitation that motivates the VAE: a standard autoencoder's latent space has no prior, so you cannot sample from it. You build a convolutional autoencoder yourself. The lesson runs to 26 slides.

Subject: Machine Learning · 57 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Autoencoders: Standard, Denoising & Sparse

Title

USAAIO · Lesson 118 · Phase 4 (Generative AI)

Compress to a bottleneck and reconstruct — then meet the variants that learn robust and interpretable features, and the limitation that forces us to invent the VAE next lesson.

2. By the end of this lesson you can

Objectives

  1. Describe the encoder -> latent z -> decoder structure and write the reconstruction loss (MSE or BCE)
  2. Explain why an undercomplete bottleneck (d_z < d_x) forces meaningful compression
  3. Build a denoising autoencoder — corrupt the input, reconstruct the clean target
  4. Build a sparse autoencoder with an L1 activation penalty and read off the active fraction
  5. Interpolate in latent space, and explain why a standard AE cannot be sampled — the gap the VAE fills

3. What survived from Phase 3 Gap Analysis & Phase 4 Preview?

Warm-up

Discussion prompt

Before we open Lesson 118: Autoencoders — Standard, Denoising & Sparse: without looking back, what was the main idea of Phase 3 Gap Analysis & Phase 4 Preview, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

structured gap analysis across all nine Phase 3 topic clusters (transformers, BERT/GPT, ViT, GNNs, NLP, CV, Seq2Seq, BLEU/ROUGE, positional encoding), a scored self-assessment, study-plan adjustment for weak areas, Phase 4 overview (VAE / ELBO, GAN minimax, DDPM noise schedule), and exam strategy (time allocation, derivation approach, mindset).

4. The standard autoencoder

Section

Part 1 of 4

5. Encoder, bottleneck, decoder

Concept

An autoencoder learns to copy its input to its output through a narrow middle. The encoder maps x to a low-dimensional code z; the decoder maps z back to a reconstruction x_hat.

\[ z = f_{\text{enc}}(x), \qquad \hat{x} = f_{\text{dec}}(z) \]

bottleneck (latent code) — The low-dimensional vector z. Because it is too small to hold a verbatim copy, the network must keep only what matters to reconstruct x.

6. Break it if you can: Encoder, bottleneck, decoder

Counterexample

Discussion prompt

An autoencoder learns to copy its input to its output through a narrow middle. The encoder maps x to a low-dimensional code z; the decoder maps z back to a reconstruction x_hat.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

7. The loss is reconstruction only

Concept

There are no labels. The target is the input — the network is trained on how well x_hat matches x.

\[ \mathcal{L}_{\text{MSE}} = \frac{1}{N}\sum_i \lVert x_i - \hat{x}_i\rVert^2 \qquad \mathcal{L}_{\text{BCE}} = -\sum_j x_j\log \hat{x}_j + (1-x_j)\log(1-\hat{x}_j) \]

Use MSE for real-valued inputs and BCE for inputs in [0,1] (e.g. normalized pixel intensities). That is the whole objective — self-supervised, no annotation needed.

8. By analogy: The loss is reconstruction only

Analogy

Discussion prompt

Explain The loss is reconstruction only by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Use MSE for real-valued inputs and BCE for inputs in [0,1] (e.g. normalized pixel intensities). That is the whole objective — self-supervised, no annotation needed.

9. Undercomplete vs overcomplete

Concept

The latent dimension d_z relative to the input dimension d_x decides what the AE learns.

d_z vs d_xnamewhat happens
d_z < d_xundercompleteforced compression — must discard redundancy, learns structure
d_z = d_xcompletecan learn identity — risk of learning nothing useful
d_z > d_xovercompletecan copy verbatim — needs a regularizer (sparsity/denoising) to be useful

The undercomplete bottleneck is the classic case: the squeeze is the teacher.

10. Fill in: name for Undercomplete vs overcomplete

Comparison

Comparison matrix

From Undercomplete vs overcomplete: refill the name column from what you know. The rest of the table is as it appeared.

d_z vs d_xnamewhat happens
d_z < d_xundercompleteforced compression — must discard redundancy, learns structure
d_z = d_xcompletecan learn identity — risk of learning nothing useful
d_z > d_xovercompletecan copy verbatim — needs a regularizer (sparsity/denoising) to be useful

11. Guess the shape of the answer: Forward-pass trace: 4 -> 2 -> 4

Estimation

Predict first

An undercomplete AE (d_x=4, d_z=2). Hand-set weights so we can read every number. Input x = [1, 0, 1, 0].

Commit before you compute: what does Forward-pass trace: 4 -> 2 -> 4 come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Decode: x_hat = W_dec z = [1, 0, 1, 0]

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The decoder copies z0 into outputs 0 and 2, z1 into outputs 1 and 3.

12. Forward-pass trace: 4 -> 2 -> 4

Worked example

An undercomplete AE (d_x=4, d_z=2). Hand-set weights so we can read every number. Input x = [1, 0, 1, 0].

import torch
x = torch.tensor([[1.0, 0.0, 1.0, 0.0]])
W_enc = torch.tensor([[0.5,0,0.5,0],[0,0.5,0,0.5]])  # (2,4)
z = torch.relu(x @ W_enc.T)                          # bottleneck
W_dec = torch.tensor([[1.,0],[0,1.],[1.,0],[0,1.]])  # (4,2)
x_hat = z @ W_dec.T
print('z    =', z.tolist())
print('x_hat=', x_hat.tolist())
print('MSE  =', ((x-x_hat)**2).mean().item())

Encode: z = ReLU(W_enc x) = [0.5+0.5, 0] = [1.0, 0.0]

Why: The encoder averages features (0,2) into z0 and (1,3) into z1; here only z0 fires.

Decode: x_hat = W_dec z = [1, 0, 1, 0]

Why: The decoder copies z0 into outputs 0 and 2, z1 into outputs 1 and 3.

quantityvaluenote
z (latent)[1.0, 0.0]2 numbers stand in for 4
x_hat[1.0, 0.0, 1.0, 0.0]exact reconstruction
reconstruction MSE0.0this structured input is perfectly recoverable from 2 dims

13. What each one costs: Forward-pass trace: 4 -> 2 -> 4

Trade off

Comparison matrix

From Forward-pass trace: 4 -> 2 -> 4: every row here is a choice with a cost. Fill the value column, then say which row you would actually pick and what you give up for it.

quantityvaluenote
z (latent)[1.0, 0.0]2 numbers stand in for 4
x_hat[1.0, 0.0, 1.0, 0.0]exact reconstruction
reconstruction MSE0.0this structured input is perfectly recoverable from 2 dims

14. What the bottleneck really learns

Intuition

The AE is lossy compression fitted to your data. MNIST digits live on a thin manifold inside 784-pixel space; a good encoder finds coordinates on that manifold instead of storing raw pixels.

That is why a 2-number code can perfectly rebuild our 4-pixel input: the input had structure (features came in matched pairs), and the bottleneck captured exactly that structure.

15. Teach it back: What the bottleneck really learns

Explain it

Discussion prompt

Explain What the bottleneck really learns to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

The AE is lossy compression fitted to your data. MNIST digits live on a thin manifold inside 784-pixel space; a good encoder finds coordinates on that manifold instead of storing raw pixels.

16. Denoising & sparse variants

Section

Part 2 of 4

17. Denoising autoencoder

Concept

Corrupt the input with noise, feed the corrupted version in, but score the reconstruction against the clean original. The network can't just copy — it must learn what the signal should be.

\[ \tilde{x} = x + \epsilon,\ \epsilon\sim\mathcal{N}(0,\sigma^2 I); \qquad \mathcal{L} = \lVert x - f_{\text{dec}}(f_{\text{enc}}(\tilde{x}))\rVert^2 \]

The target in the loss is x (clean), never the corrupted input. This forces robust features and is a stepping stone to diffusion models (which denoise repeatedly).

18. What rests on this: Denoising autoencoder

Socratic

Discussion prompt

Corrupt the input with noise, feed the corrupted version in, but score the reconstruction against the clean original. The network can't just copy — it must learn what the signal should be.

Suppose that were not true. What is the first thing in Lesson 118: Autoencoders — Standard, Denoising & Sparse that would stop working?

Hint: Follow it one step downstream. The answer is whatever was quietly relying on it.

Answer:

The target in the loss is x (clean), never the corrupted input. This forces robust features and is a stepping stone to diffusion models (which denoise repeatedly).

19. What has to be given first: Denoising trace: corrupt, then target the…

Missing information

Discussion prompt

Add Gaussian noise (sigma=0.5) to a clean vector. Note which signal the loss compares against.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

Input is corrupted, target is clean. Targeting the corrupted vector would just teach the identity and defeat the point.

20. Denoising trace: corrupt, then target the clean signal

Worked example

Add Gaussian noise (sigma=0.5) to a clean vector. Note which signal the loss compares against.

import torch
torch.manual_seed(1)
clean = torch.tensor([[1.0, 0.0, 1.0, 0.0]])
corrupted = clean + 0.5 * torch.randn_like(clean)
print('clean    =', clean.tolist())
print('corrupted=', [round(v,4) for v in corrupted.flatten().tolist()])
# encoder sees `corrupted`; loss compares reconstruction to `clean`
vectorvaluesrole
clean x[1.0, 0.0, 1.0, 0.0]loss TARGET
noise eps[0.3307, 0.1335, 0.0308, 0.3107]sampled N(0, 0.25)
corrupted x~[1.3307, 0.1335, 1.0308, 0.3107]encoder INPUT

Loss = ||clean - decode(encode(corrupted))||^2

Why: Input is corrupted, target is clean. Targeting the corrupted vector would just teach the identity and defeat the point.

21. Work backwards from the answer: Denoising trace: corrupt, then target the…

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Loss = ||clean - decode(encode(corrupted))||^2

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Add Gaussian noise (sigma=0.5) to a clean vector. Note which signal the loss compares against.

22. Sparse autoencoder

Concept

Add an L1 penalty on the activations (not the weights). The network reconstructs well while keeping most latent units at exactly zero — a sparse code.

\[ \mathcal{L} = \lVert x-\hat{x}\rVert^2 + \lambda \sum_k \lvert a_k\rvert \]

Each input then activates only a few units, each standing for one human-readable feature. This is the backbone of modern mechanistic interpretability (sparse features extracted from LLM activations).

23. Teach it back: Sparse autoencoder

Explain it

Discussion prompt

Explain Sparse autoencoder to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Add an L1 penalty on the activations (not the weights). The network reconstructs well while keeping most latent units at exactly zero — a sparse code.

24. Guess the shape of the answer: Sparse trace: L1 penalty and the active…

Estimation

Predict first

Compare a sparse code with a dense code of equal reconstruction power. Read the penalty and the fraction of units that are non-zero.

Commit before you compute: what does Sparse trace: L1 penalty and the active fraction come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Minimizing recon + lambda*L1 drives small activations to exactly 0

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. L1's constant-magnitude gradient pushes a small a_k to zero (unlike L2, which only shrinks it).

25. Sparse trace: L1 penalty and the active fraction

Worked example

Compare a sparse code with a dense code of equal reconstruction power. Read the penalty and the fraction of units that are non-zero.

import torch
sparse = torch.tensor([0.0, 0.0, 2.0, 0.0, 1.0])
dense  = torch.tensor([0.8, 0.7, 0.9, 0.6, 1.0])
for name, a in [('sparse', sparse), ('dense', dense)]:
    print(name, 'L1=', a.abs().sum().item(),
          'active=', (a != 0).float().mean().item())
codeL1 = sum|a|fraction activereading
sparse [0,0,2,0,1]3.00.402 of 5 units fire
dense [0.8,0.7,0.9,0.6,1.0]4.01.00every unit fires

Minimizing recon + lambda*L1 drives small activations to exactly 0

Why: L1's constant-magnitude gradient pushes a small a_k to zero (unlike L2, which only shrinks it). Sparsity is measured by the active fraction, not by total magnitude alone.

26. Work backwards from the answer: Sparse trace: L1 penalty and the active…

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Minimizing recon + lambda*L1 drives small activations to exactly 0

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Compare a sparse code with a dense code of equal reconstruction power. Read the penalty and the fraction of units that are non-zero.

27. Something is wrong here: training a denoising AE against the corrupted input

Anomaly

Predict first

A student writes this, and it looks reasonable:

I corrupted the input, so the network should learn to reproduce what it was given.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Target = corrupted input. The optimum is the identity map — the AE learns to copy noise and learns nothing robust.

The corruption is only on the input path. The target stays the clean signal.

Why: Target = corrupted input. The optimum is the identity map — the AE learns to copy noise and learns nothing robust.

28. Trap: training a denoising AE against the corrupted input

Trap

The trap

I corrupted the input, so the network should learn to reproduce what it was given.

loss = ||corrupted - decode(encode(corrupted))||^2

Why: Target = corrupted input. The optimum is the identity map — the AE learns to copy noise and learns nothing robust.

The fix

The corruption is only on the input path. The target stays the clean signal.

loss = ||clean - decode(encode(corrupted))||^2

Why: Target = clean input. The network must undo the noise, which is exactly the robust feature-learning we want.

29. Break it on purpose: training a denoising AE against the…

Break the constraint

Discussion prompt

The rule this trap just fixed:

Target = clean input. The network must undo the noise, which is exactly the robust feature-learning we want.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

Target = corrupted input. The optimum is the identity map — the AE learns to copy noise and learns nothing robust.

30. Something is wrong here: L1 on the weights instead of the activations

Anomaly

Predict first

A student writes this, and it looks reasonable:

Sparsity = add an L1 weight-decay term, like Lasso regression.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: This sparsifies the connections, giving a smaller/pruned network — it does NOT make the per-input latent code sparse.

A sparse autoencoder penalizes the activations produced for each input.

Why: This sparsifies the connections, giving a smaller/pruned network — it does NOT make the per-input latent code sparse.

31. Trap: L1 on the weights instead of the activations

Trap

The trap

Sparsity = add an L1 weight-decay term, like Lasso regression.

penalty = lambda * sum(|W|) (weights)

Why: This sparsifies the connections, giving a smaller/pruned network — it does NOT make the per-input latent code sparse.

The fix

A sparse autoencoder penalizes the activations produced for each input.

penalty = lambda * sum(|a|) (activations a = encoder output)

Why: Now each input lights up only a few latent units — the interpretable, per-example sparsity that defines a sparse AE.

32. Which of these survive contact with Lesson 118: Autoencoders — Standard…?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Use MSE for real-valued inputs and BCE for inputs in [0,1] (e.g. normalized pixel intensities). That is the whole objective — self-supervised, no annotation needed.; The latent dimension d_z relative to the input dimension d_x decides what the AE learns.; The target in the loss is x (clean), never the corrupted input. This forces robust features and is a stepping stone to diffusion models (which denoise repeatedly).
Breaks
I corrupted the input, so the network should learn to reproduce what it was given.; Sparsity = add an L1 weight-decay term, like Lasso regression.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 118: Autoencoders — Standard, Denoising & Sparse puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

33. The latent space & its limit

Section

Part 3 of 4

34. Guess the shape of the answer: Latent interpolation

Estimation

Predict first

Walk a straight line between two codes z_A and z_B and decode each step — the AE's version of morphing one example into another.

Commit before you compute: what does Latent interpolation come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Decode each z to get a smooth morph between the two reconstructions

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Interpolation is well-defined ONLY where the decoder has seen data.

35. Latent interpolation

Worked example

Walk a straight line between two codes z_A and z_B and decode each step — the AE's version of morphing one example into another.

import torch
zA = torch.tensor([0.0, 2.0]); zB = torch.tensor([2.0, 0.0])
for a in [0.0, 0.25, 0.5, 0.75, 1.0]:
    print(a, ((1-a)*zA + a*zB).tolist())
alphaz = (1-a)zA + a zB
0.00[0.0, 2.0]
0.25[0.5, 1.5]
0.50[1.0, 1.0]
0.75[1.5, 0.5]
1.00[2.0, 0.0]

Decode each z to get a smooth morph between the two reconstructions

Why: Interpolation is well-defined ONLY where the decoder has seen data. Off the manifold, decoded points are meaningless.

36. Fill in: z = (1-a)zA + a zB for Latent interpolation

Comparison

Comparison matrix

From Latent interpolation: refill the z = (1-a)zA + a zB column from what you know. The rest of the table is as it appeared.

alphaz = (1-a)zA + a zB
0.00[0.0, 2.0]
0.25[0.5, 1.5]
0.50[1.0, 1.0]
0.75[1.5, 0.5]
1.00[2.0, 0.0]

37. Why you cannot SAMPLE a standard AE

Concept

To generate, you would draw a fresh z and decode it. But a standard AE imposes no prior on z — the encoder is free to scatter codes anywhere, leaving holes the decoder never visited.

import torch, torch.nn as nn
torch.manual_seed(2)
enc = nn.Sequential(nn.Linear(4,8), nn.ReLU(), nn.Linear(8,2))
data = torch.randn(200,4) @ torch.diag(torch.tensor([2.,2.,.1,.1]))
Z = enc(data).detach()
print('latent mean=', [round(v,3) for v in Z.mean(0).tolist()])
print('latent std =', [round(v,3) for v in Z.std(0).tolist()])
sample = torch.randn(2)
print('N(0,1) sample dist to nearest real code=',
      round((Z-sample).pow(2).sum(1).min().sqrt().item(),3))
measurementvalueimplication
latent mean[-0.055, 0.511]not centered at 0
latent std[0.212, 0.197]not unit variance
N(0,1) sample -> nearest code1.168a standard-normal draw lands far from any real code

Decode that far-away z and you get garbage — the decoder was never trained there. The VAE (Lesson 119) fixes exactly this by forcing q(z|x) toward N(0, I), so sampling z ~ N(0, I) is finally valid.

38. Where does each piece belong: Lesson 118: Autoencoders — Standard…

Sorting

Sort into buckets

These are the pieces of Lesson 118: Autoencoders — Standard, Denoising & Sparse, out of order. Put each one back under the part of the lesson it belongs to.

The standard autoencoder
Encoder, bottleneck, decoder; The loss is reconstruction only; Undercomplete vs overcomplete
Denoising & sparse variants
Denoising autoencoder; Denoising trace: corrupt, then target the clean signal; Sparse autoencoder
The latent space & its limit
Latent interpolation; Why you cannot SAMPLE a standard AE
s1
The standard autoencoder is where Lesson 118: Autoencoders — Standard, Denoising & Sparse puts Encoder, bottleneck, decoder, The loss is reconstruction only, Undercomplete vs overcomplete. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
Denoising & sparse variants is where Lesson 118: Autoencoders — Standard, Denoising & Sparse puts Denoising autoencoder, Denoising trace: corrupt, then target the clean signal, Sparse autoencoder. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
The latent space & its limit is where Lesson 118: Autoencoders — Standard, Denoising & Sparse puts Latent interpolation, Why you cannot SAMPLE a standard AE. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

39. Recipe & checks

Section

Part 4 of 4

40. Rebuild the recipe: Choosing an autoencoder

Ranking

Put in order

These are the steps of Choosing an autoencoder, scrambled. Put them back in order before the next slide shows you.

  1. Structure: encoder -> bottleneck z -> decoder; train on reconstruction loss (MSE for real, BCE for [0,1]).
  2. Compress / visualize: undercomplete (d_z < d_x). The squeeze forces it to learn structure.
  3. Robust features: denoising AE — corrupt the input, target the clean signal.
  4. Interpretable / disentangled units: sparse AE — L1 on activations, watch the active fraction.
  5. Generate new samples? A plain AE cannot — no prior on z. Reach for a VAE or diffusion model.

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

41. Choosing an autoencoder

Pattern

  1. Structure: encoder -> bottleneck z -> decoder; train on reconstruction loss (MSE for real, BCE for [0,1]).
  2. Compress / visualize: undercomplete (d_z < d_x). The squeeze forces it to learn structure.
  3. Robust features: denoising AE — corrupt the input, target the clean signal.
  4. Interpretable / disentangled units: sparse AE — L1 on activations, watch the active fraction.
  5. Generate new samples? A plain AE cannot — no prior on z. Reach for a VAE or diffusion model.

42. Where does it stop working: Choosing an autoencoder

Edge cases

Discussion prompt

Choosing an autoencoder works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Structure: encoder -> bottleneck z -> decoder; train on reconstruction loss (MSE for real, BCE for [0,1]).
  2. Compress / visualize: undercomplete (d_z < d_x). The squeeze forces it to learn structure.
  3. Robust features: denoising AE — corrupt the input, target the clean signal.
  4. Interpretable / disentangled units: sparse AE — L1 on activations, watch the active fraction.
  5. Generate new samples? A plain AE cannot — no prior on z. Reach for a VAE or diffusion model.

43. Rule out three: Check 1: the bottleneck

Elimination

Eliminate the wrong options

An autoencoder has input dimension d_x = 784 and latent dimension d_z = 32. Which statement is TRUE?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. It is undercomplete; the bottleneck forces it to learn a compressed representation.
  • B. It is overcomplete and will copy the input verbatim with no regularizer.
  • C. It needs labels for the 784 outputs to compute its loss.
  • D. With d_z = 32 it can perfectly reconstruct ANY 784-dim input.

Survives elimination: A

Why: d_z (32) < d_x (784), so the AE is undercomplete: it cannot store a verbatim copy and must compress, learning the data's structure.

44. Check 1: the bottleneck

Check

Decide it before clicking.

Check your understanding

An autoencoder has input dimension d_x = 784 and latent dimension d_z = 32. Which statement is TRUE?

  • A. It is undercomplete; the bottleneck forces it to learn a compressed representation. (correct)
  • B. It is overcomplete and will copy the input verbatim with no regularizer.
  • C. It needs labels for the 784 outputs to compute its loss.
  • D. With d_z = 32 it can perfectly reconstruct ANY 784-dim input.

Answer: A

Why: d_z (32) < d_x (784), so the AE is undercomplete: it cannot store a verbatim copy and must compress, learning the data's structure.

Why B tempts people
Overcomplete means d_z > d_x. Here d_z < d_x, so it is undercomplete.
Why C tempts people
Autoencoders are self-supervised — the target is the input itself; no external labels are used.
Why D tempts people
It reconstructs structured data drawn from a low-dimensional manifold well, not arbitrary 784-dim vectors; a 32-dim code cannot losslessly encode all of R^784.

45. Rule out three: Check 2: denoising target & sampling

Elimination

Eliminate the wrong options

You train a denoising autoencoder. Which loss is correct, and can you later generate new images by sampling z ~ N(0, I) and decoding?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Loss compares reconstruction to the CLEAN input; and no — a standard/denoising AE has no prior on z, so sampling z ~ N(0, I) is not valid.
  • B. Loss compares reconstruction to the CORRUPTED input; and yes — any AE can be sampled.
  • C. Loss compares reconstruction to the CLEAN input; and yes — sampling z ~ N(0, I) always works.
  • D. Loss compares reconstruction to the CORRUPTED input; and no — AEs can never generate.

Survives elimination: A

Why: Denoising AEs target the clean signal (the network must undo the noise). And a standard/denoising AE imposes no prior on the latent, so codes cluster unpredictably — sampling z ~ N(0, I) lands in empty regions the decoder never learned. The VAE fixes this.

46. Check 2: denoising target & sampling

Check

One subtlety from each variant.

Check your understanding

You train a denoising autoencoder. Which loss is correct, and can you later generate new images by sampling z ~ N(0, I) and decoding?

  • A. Loss compares reconstruction to the CLEAN input; and no — a standard/denoising AE has no prior on z, so sampling z ~ N(0, I) is not valid. (correct)
  • B. Loss compares reconstruction to the CORRUPTED input; and yes — any AE can be sampled.
  • C. Loss compares reconstruction to the CLEAN input; and yes — sampling z ~ N(0, I) always works.
  • D. Loss compares reconstruction to the CORRUPTED input; and no — AEs can never generate.

Answer: A

Why: Denoising AEs target the clean signal (the network must undo the noise). And a standard/denoising AE imposes no prior on the latent, so codes cluster unpredictably — sampling z ~ N(0, I) lands in empty regions the decoder never learned. The VAE fixes this.

Why B tempts people
Targeting the corrupted input makes the identity map optimal — no robust learning — and AEs without a latent prior cannot be reliably sampled.
Why C tempts people
The loss is right, but the sampling claim is wrong: a plain AE has no prior on z, so z ~ N(0, I) is off-distribution.
Why D tempts people
The loss is wrong (target is the clean signal), even though the no-sampling conclusion is right for the wrong reason.

47. Your turn: build it

Section

Project

48. Project: a convolutional autoencoder for MNIST

Concept

Build it yourself, one piece at a time. Type every line, run after each step, and read errors — do not erase them.

requirementtool you'll use
Encoder: 28x28 -> latentnn.Conv2d, nn.ReLU, nn.Flatten, nn.Linear
Decoder: latent -> 28x28nn.Linear, nn.Unflatten, nn.ConvTranspose2d
Reconstruction lossnn.MSELoss (or BCEWithLogits on [0,1] pixels)
Denoising variantadd torch.randn_like(x)*sigma to the INPUT only

49. By analogy: Project: a convolutional autoencoder for MNIST

Analogy

Discussion prompt

Explain Project: a convolutional autoencoder for MNIST by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Build it yourself, one piece at a time. Type every line, run after each step, and read errors — do not erase them.

50. Milestone 1 — the encoder

Worked example

Your turn: write an encoder that takes a (B,1,28,28) batch down to a (B,16) latent. Predict the output shape before you run.

Hint: two stride-2 Conv2d layers quarter then quarter the spatial size (28 -> 14 -> 7), then Flatten and a Linear to 16.

import torch, torch.nn as nn
encoder = nn.Sequential(
    nn.Conv2d(1, 8, 3, stride=2, padding=1), nn.ReLU(),   # 28->14
    nn.Conv2d(8, 16, 3, stride=2, padding=1), nn.ReLU(),  # 14->7
    nn.Flatten(),
    nn.Linear(16*7*7, 16))
z = encoder(torch.randn(4, 1, 28, 28))
print(z.shape)
stageshape
input(4, 1, 28, 28)
after conv1 (stride 2)(4, 8, 14, 14)
after conv2 (stride 2)(4, 16, 7, 7)
latent z(4, 16)

51. What each one costs: Milestone 1 — the encoder

Trade off

Comparison matrix

From Milestone 1 — the encoder: every row here is a choice with a cost. Fill the shape column, then say which row you would actually pick and what you give up for it.

stageshape
input(4, 1, 28, 28)
after conv1 (stride 2)(4, 8, 14, 14)
after conv2 (stride 2)(4, 16, 7, 7)
latent z(4, 16)

52. Milestone 2 — the decoder & a training step

Worked example

Your turn: mirror the encoder back up to (B,1,28,28), then run one MSE reconstruction step. Predict: should the loss go down after opt.step()?

Hint: Linear(16 -> 16*7*7), Unflatten to (16,7,7), then two ConvTranspose2d(stride=2) to climb 7 -> 14 -> 28.

import torch, torch.nn as nn
decoder = nn.Sequential(
    nn.Linear(16, 16*7*7), nn.Unflatten(1, (16,7,7)),
    nn.ConvTranspose2d(16, 8, 4, stride=2, padding=1), nn.ReLU(),  # 7->14
    nn.ConvTranspose2d(8, 1, 4, stride=2, padding=1))              # 14->28
x = torch.rand(4, 1, 28, 28)
ae = nn.Sequential(encoder, decoder)
opt = torch.optim.Adam(ae.parameters(), lr=1e-3)
l1 = nn.MSELoss()(ae(x), x); l1.backward(); opt.step()
l2 = nn.MSELoss()(ae(x), x)
print(round(l1.item(),4), '->', round(l2.item(),4), 'down?', l2.item() < l1.item())
checkexpected
decoder output shape(4, 1, 28, 28) — matches input
loss after one Adam steplower than before (down? True)
denoising tweakfeed ae(x + 0.5*torch.randn_like(x)), keep target x

53. Fill in: expected for Milestone 2 — the decoder & a training step

Comparison

Comparison matrix

From Milestone 2 — the decoder & a training step: refill the expected column from what you know. The rest of the table is as it appeared.

checkexpected
decoder output shape(4, 1, 28, 28) — matches input
loss after one Adam steplower than before (down? True)
denoising tweakfeed ae(x + 0.5*torch.randn_like(x)), keep target x

54. Show it off

Concept

Explain your program out loud, in order: where is the bottleneck? Which line would you change to make it denoising? Which to make it sparse?

55. Break it if you can: Show it off

Counterexample

Discussion prompt

Explain your program out loud, in order: where is the bottleneck? Which line would you change to make it denoising? Which to make it sparse?

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

56. Connect it up: Lesson 118: Autoencoders — Standard, Denoising & Sparse

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — The standard autoencoder · Denoising & sparse variants · The latent space & its limit · Recipe & checks · Your turn: build it. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

57. Recap: autoencoders

Recap

You can now build and reason about the three workhorse autoencoders — and you know precisely why they are not yet generative.

varianttwistbuys you
standard (undercomplete)bottleneck d_z < d_xcompression / visualization
denoisingcorrupt input, target cleanrobust features (-> diffusion)
sparseL1 on activationsinterpretable units (-> mech interp)
limitationno prior on zcannot sample -> motivates the VAE (L119)

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 118 (Autoencoders — standard, undercomplete, denoising, sparse; latent interpolation; why a standard AE is not generative) — Barron · USAAIO Round 2 Preparation, 2026
  2. Forward pass, denoising corruption, sparse L1 activation penalty, latent interpolation, and latent-distribution statistics all verified by real execution — torch 2.7.1+cpu + numpy 2.2.6, June 2026
  3. Vincent et al., Extracting and Composing Robust Features with Denoising Autoencoders (ICML 2008) — Denoising autoencoder origin

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108