Lesson 117: Phase 3 Gap Analysis & Phase 4 Preview

USAAIO Lesson 117, at the end of Phase 3. It is a structured gap analysis across all nine Phase 3 topic clusters - transformers, BERT and GPT, ViT, GNNs, NLP, computer vision, seq2seq, BLEU and ROUGE, and positional encoding - with a scored self-assessment and a way to adjust your study plan for the weak areas. It then previews Phase 4, covering the VAE and its ELBO, the GAN minimax objective, and the DDPM noise schedule, and closes on exam strategy: allocating your time, approaching a derivation, and mindset. All the math was verified by real Python execution with numpy 2.2.6 and torch 2.7.1+cpu. The lesson runs to 29 slides.

Subject: Machine Learning · 59 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Phase 3 Gap Analysis & Phase 4 Preview

Title

USAAIO · Lesson 117 · End of Phase 3

Locate your weakest Phase 3 concepts, recalibrate your Phase 4 study plan, and get a concrete preview of VAE, GAN, and diffusion. The exam is closer than it feels.

2. By the end of this lesson you can

Objectives

  1. Score yourself on all nine Phase 3 topic clusters and rank weaknesses
  2. Write a concrete Phase 4 study plan with weak-area re-allocation
  3. State the ELBO objective (VAE) and explain what each term does
  4. Write the GAN minimax objective and distinguish D's and G's gradients
  5. Describe the DDPM forward process q(x_t|x_0) and its closed form
  6. Apply the exam-strategy timing model to any USAAIO-style paper

3. What survived from NLP + CV Review & Exam Simulation?

Warm-up

Discussion prompt

Before we open Lesson 117: Phase 3 Gap Analysis & Phase 4 Preview: without looking back, what was the main idea of NLP + CV Review & Exam Simulation, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

comprehensive NLP+CV exam-simulation deck covering tokenization/word2vec/BERT/LoRA/seq2seq/beam-search/BLEU/ROUGE and convolution/ResNet/ViT/U-Net/YOLO/IoU/NMS/CLIP. Emphasizes cross-modal connections through transformers, pattern recognition for exam pace, and gap-spotting for Phase 4.

4. Phase 3 self-assessment — score the nine clusters

Section

Part 1 of 4

5. The nine Phase 3 clusters

Concept

Phase 3 (Lessons 87-116) covered transformers, NLP, and computer vision. Rate yourself 0-10 on each cluster before seeing the benchmarks.

#clusterkey lessonsyour score
1Transformer self-attention87-90 /10
2BERT / GPT fine-tuning91, 94 /10
3ViT patch embedding92 /10
4GNN message passing95-97 /10
5NER / sequence labeling99-101 /10
6Object detection (IoU/NMS)105-107 /10
7Seq2Seq + cross-attention88, 103 /10
8BLEU / ROUGE metrics102-104 /10
9Positional encoding variants89, 92 /10

Clusters with a score of 7 or below go on your Phase 4 re-study list. Clusters at 8+ are consolidate-and-move-on.

6. Break it if you can: The nine Phase 3 clusters

Counterexample

Discussion prompt

Phase 3 (Lessons 87-116) covered transformers, NLP, and computer vision. Rate yourself 0-10 on each cluster before seeing the benchmarks.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

7. Guess the shape of the answer: Scoring example — running the audit

Estimation

Predict first

Compute a weighted priority score: priority = (10 - score) * exam_weight. Here exam_weight is an estimate of how often the topic appears in USAAIO derivation problems.

Commit before you compute: what does Scoring example — running the audit come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Rank by gap * exam_weight to find where extra study hours yield the most exam points

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. A score-6 topic with high exam weight (like GNN or ObjDet) costs more than a score-7 on a low-weight topic.

8. Scoring example — running the audit

Worked example

Compute a weighted priority score: priority = (10 - score) * exam_weight. Here exam_weight is an estimate of how often the topic appears in USAAIO derivation problems.

import numpy as np
scores      = np.array([8, 7, 9, 6, 7, 6, 8, 5, 9])
wts         = np.array([3, 2, 2, 1, 1, 2, 2, 1, 2])
gap         = 10 - scores
priority    = gap * wts
topics = ['self-attn','BERT/GPT','ViT','GNN','NER',
          'ObjDet','Seq2Seq','BLEU','PosEnc']
order = np.argsort(-priority)
for i in order:
    print(f'{priority[i]:3.0f} pts  {topics[i]} (score {scores[i]})')

Rank by gap * exam_weight to find where extra study hours yield the most exam points

Why: A score-6 topic with high exam weight (like GNN or ObjDet) costs more than a score-7 on a low-weight topic. Priority = gap x weight makes the trade-off explicit.

priority ptsclusterscore /10
6BERT/GPT (wt=2)7
6ObjDet (wt=2)6
5BLEU (wt=1)5
4GNN (wt=1)6
4Seq2Seq (wt=2)8

9. Fill in: score /10 for Scoring example — running the audit

Comparison

Comparison matrix

From Scoring example — running the audit: refill the score /10 column from what you know. The rest of the table is as it appeared.

priority ptsclusterscore /10
6BERT/GPT (wt=2)7
6ObjDet (wt=2)6
5BLEU (wt=1)5
4GNN (wt=1)6
4Seq2Seq (wt=2)8

10. Phase 4 study plan re-allocation

Concept

Phase 4 (Weeks 40-52) has 39 lessons. The first pass covers new content (VAE, GAN, Diffusion). Lessons marked 'review' are your reclaimed hours.

phaselessonshoursstrategy
new content118-13026 hattend fully; build each model
review slots131-13510 hdrill top-3 gap clusters
mock exams136-15234 htimed, then targeted debrief
final sprint153-1568 hweak-area formula sheets only

Rule: if a mock-exam debrief reveals a cluster re-entering the 'weak' list, pull one review slot from Week 49 and re-drill it before the next mock.

11. What each one costs: Phase 4 study plan re-allocation

Trade off

Comparison matrix

From Phase 4 study plan re-allocation: every row here is a choice with a cost. Fill the lessons column, then say which row you would actually pick and what you give up for it.

phaselessonshoursstrategy
new content118-13026 hattend fully; build each model
review slots131-13510 hdrill top-3 gap clusters
mock exams136-15234 htimed, then targeted debrief
final sprint153-1568 hweak-area formula sheets only

12. Phase 4 preview — VAE, GAN, Diffusion

Section

Part 2 of 4

13. VAE — the ELBO objective

Concept

A VAE (Kingma & Welling 2014) learns a continuous, structured latent space by maximizing the Evidence Lower BOund (ELBO). Unlike a plain AE, the latent code is regularized to approximate N(0,I).

\[ \mathcal{L}(\theta,\phi;x) = \underbrace{\mathbb{E}_{q_\phi(z|x)}[\log p_\theta(x|z)]}_{\text{reconstruction}} - \underbrace{D_{\mathrm{KL}}(q_\phi(z|x)\,\|\,p(z))}_{\text{regularization}} \]

The KL term forces the encoder distribution q(z|x) toward the prior p(z)=N(0,I). The reconstruction term rewards the decoder for recovering x from the sampled z.

14. By analogy: VAE — the ELBO objective

Analogy

Discussion prompt

Explain VAE — the ELBO objective by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

A VAE (Kingma & Welling 2014) learns a continuous, structured latent space by maximizing the Evidence Lower BOund (ELBO). Unlike a plain AE, the latent code is regularized to approximate N(0,I).

15. What has to be given first: ELBO with real numbers

Missing information

Discussion prompt

Verify the ELBO for a single 1D datapoint: x=2.0, encoder outputs mu_z=1.5, log_var=-0.693 (sigma^2=0.5), decoder outputs x_recon=1.8.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

The KL term dominates here because mu_z=1.5 is far from the prior mean 0. Training maximizes ELBO, so the encoder is pushed to reduce KL while the decoder reduces reconstruction error.

16. ELBO with real numbers

Worked example

Verify the ELBO for a single 1D datapoint: x=2.0, encoder outputs mu_z=1.5, log_var=-0.693 (sigma^2=0.5), decoder outputs x_recon=1.8.

import numpy as np
x_obs, mu_z, log_var, x_recon = 2.0, 1.5, np.log(0.5), 1.8
# Reconstruction: log p(x|z) ~ -0.5*(x-x_recon)^2  (unit decoder var)
recon = -0.5 * (x_obs - x_recon)**2
# KL(N(mu,sigma^2) || N(0,1))
sigma2 = np.exp(log_var)
kl = 0.5 * (mu_z**2 + sigma2 - 1 - log_var)
elbo = recon - kl
print(f'recon = {recon:.4f}')
print(f'KL    = {kl:.4f}')
print(f'ELBO  = {elbo:.4f}')

recon = -0.0200, KL = 1.2216, ELBO = -1.2416

Why: The KL term dominates here because mu_z=1.5 is far from the prior mean 0. Training maximizes ELBO, so the encoder is pushed to reduce KL while the decoder reduces reconstruction error.

termformulavalue
reconstruction-0.5*(2.0-1.8)^2-0.0200
KL0.5*(1.5^2 + 0.5 - 1 - log(0.5))1.2216
ELBOrecon - KL-1.2416

17. Work backwards from the answer: ELBO with real numbers

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

recon = -0.0200, KL = 1.2216, ELBO = -1.2416

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Verify the ELBO for a single 1D datapoint: x=2.0, encoder outputs mu_z=1.5, log_var=-0.693 (sigma^2=0.5), decoder outputs x_recon=1.8.

18. Reparameterization trick

Concept

Sampling z ~ N(mu, sigma^2) is non-differentiable. The reparameterization trick writes z = mu + sigma * eps where eps ~ N(0,1), making the gradient flow through mu and sigma.

\[ z = \mu_\phi(x) + \sigma_\phi(x) \cdot \varepsilon, \quad \varepsilon \sim \mathcal{N}(0,I) \]

With eps=0.3, mu=1.5, sigma=0.7071: z = 1.5 + 0.7071*0.3 = 1.7121. The gradient of the loss w.r.t. mu flows through the addition, not the sampling.

19. Teach it back: Reparameterization trick

Explain it

Discussion prompt

Explain Reparameterization trick to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Sampling z ~ N(mu, sigma^2) is non-differentiable. The reparameterization trick writes z = mu + sigma * eps where eps ~ N(0,1), making the gradient flow through mu and sigma.

20. Something is wrong here: treating a plain AE as a generative model

Anomaly

Predict first

A student writes this, and it looks reasonable:

To generate a new sample, sample z from the AE's latent space and decode it.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: A plain AE's latent space has no enforced structure — latent codes cluster arbitrarily.

Use a VAE: the KL term regularizes the encoder to approximate N(0,I).

Why: A plain AE's latent space has no enforced structure — latent codes cluster arbitrarily. Sampling z from a uniform or Gaussian prior will almost always land in an 'empty' region and produce nonsense.

21. Trap: treating a plain AE as a generative model

Trap

The trap

To generate a new sample, sample z from the AE's latent space and decode it.

Sample z ~ Uniform[-1, 1] and pass through AE decoder

Why: A plain AE's latent space has no enforced structure — latent codes cluster arbitrarily. Sampling z from a uniform or Gaussian prior will almost always land in an 'empty' region and produce nonsense.

The fix

Use a VAE: the KL term regularizes the encoder to approximate N(0,I).

Sample z ~ N(0,I) and pass through the VAE decoder

Why: Because the VAE ELBO penalizes any encoder distribution that deviates from N(0,I), the decoder learns to handle samples from the prior — so at inference, sampling z~N(0,I) produces coherent outputs.

22. Break it on purpose: treating a plain AE as a generative model

Break the constraint

Discussion prompt

The rule this trap just fixed:

Because the VAE ELBO penalizes any encoder distribution that deviates from N(0,I), the decoder learns to handle samples from the prior — so at inference, sampling z~N(0,I) produces coherent outputs.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

A plain AE's latent space has no enforced structure — latent codes cluster arbitrarily. Sampling z from a uniform or Gaussian prior will almost always land in an 'empty' region and produce nonsense.

23. GAN — the minimax objective

Concept

A GAN (Goodfellow 2014) trains a discriminator D to distinguish real from fake, and a generator G to fool D. They play a zero-sum game.

\[ \min_G \max_D \;V(D,G) = \mathbb{E}_{x\sim p_{\mathrm{data}}}[\log D(x)] + \mathbb{E}_{z\sim p_z}[\log(1-D(G(z)))] \]

D maximizes V (pushes D(real) toward 1, D(fake) toward 0). G minimizes V (pushes D(G(z)) toward 1). At Nash equilibrium D(x) = 0.5 everywhere and V = -2ln(2) = -1.3863.

24. What rests on this: GAN — the minimax objective

Socratic

Discussion prompt

A GAN (Goodfellow 2014) trains a discriminator D to distinguish real from fake, and a generator G to fool D. They play a zero-sum game.

Suppose that were not true. What is the first thing in Lesson 117: Phase 3 Gap Analysis & Phase 4 Preview that would stop working?

Hint: Follow it one step downstream. The answer is whatever was quietly relying on it.

Answer:

D maximizes V (pushes D(real) toward 1, D(fake) toward 0). G minimizes V (pushes D(G(z)) toward 1). At Nash equilibrium D(x) = 0.5 everywhere and V = -2ln(2) = -1.3863.

25. Guess the shape of the answer: GAN objective — real numbers

Estimation

Predict first

Compute V(D,G) for a batch where D outputs 0.9, 0.85, 0.8 on real samples and 0.3, 0.2, 0.25 on generated samples.

Commit before you compute: what does GAN objective — real numbers come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: V(D,G) = -0.4528; at Nash equilibrium V = -1.3863

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. D is winning here (V > -1.3863 means G has not yet fooled D).

26. GAN objective — real numbers

Worked example

Compute V(D,G) for a batch where D outputs 0.9, 0.85, 0.8 on real samples and 0.3, 0.2, 0.25 on generated samples.

import numpy as np
D_real = np.array([0.9, 0.85, 0.8])
D_fake = np.array([0.3, 0.2,  0.25])
V = np.mean(np.log(D_real) + np.log(1 - D_fake))
G_loss_sat   = np.mean(np.log(1 - D_fake))
G_loss_nonsat = -np.mean(np.log(D_fake))
print(f'V(D,G)                 = {V:.4f}')
print(f'G saturating loss      = {G_loss_sat:.4f}')
print(f'G non-saturating loss  = {G_loss_nonsat:.4f}')
print(f'Nash V(D*,G*) = 2*log(0.5) = {2*np.log(0.5):.4f}')

V(D,G) = -0.4528; at Nash equilibrium V = -1.3863

Why: D is winning here (V > -1.3863 means G has not yet fooled D). Training continues until V approaches -1.3863, at which point D can't do better than random guessing.

quantityformulavalue
V(D,G)E[logD(x)] + E[log(1-D(G(z)))]-0.4528
G saturating lossE[log(1-D(G(z)))]-0.2892
G non-saturating loss-E[log D(G(z))]1.3999
Nash V(D,G)2*log(0.5)-1.3863

27. Work backwards from the answer: GAN objective — real numbers

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

V(D,G) = -0.4528; at Nash equilibrium V = -1.3863

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Compute V(D,G) for a batch where D outputs 0.9, 0.85, 0.8 on real samples and 0.3, 0.2, 0.25 on generated samples.

28. Diffusion — forward process q(x_t | x_0)

Concept

DDPM (Ho et al. 2020) defines a forward process that gradually corrupts a clean sample x_0 by adding Gaussian noise over T=1000 steps with a linear beta schedule.

\[ q(x_t \mid x_0) = \mathcal{N}\!\left(\sqrt{\bar{\alpha}_t}\,x_0,\;(1-\bar{\alpha}_t)I\right), \quad \bar{\alpha}_t = \prod_{s=1}^{t}(1-\beta_s) \]

The closed form means you can jump to any noise level t in one step — no need to run t sequential steps during training. The network learns to reverse each small step.

29. What rests on this: Diffusion — forward process q(x_t | x_0)

Socratic

Discussion prompt

DDPM (Ho et al. 2020) defines a forward process that gradually corrupts a clean sample x_0 by adding Gaussian noise over T=1000 steps with a linear beta schedule.

Suppose that were not true. What is the first thing in Lesson 117: Phase 3 Gap Analysis & Phase 4 Preview that would stop working?

Hint: Follow it one step downstream. The answer is whatever was quietly relying on it.

Answer:

The closed form means you can jump to any noise level t in one step — no need to run t sequential steps during training. The network learns to reverse each small step.

30. Guess the shape of the answer: DDPM noise schedule — real values

Estimation

Predict first

Compute the signal fraction sqrt(alpha_bar_t) and noise std at key timesteps for T=1000, beta linearly spaced from 1e-4 to 0.02.

Commit before you compute: what does DDPM noise schedule — real values come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: By t=500 signal_fraction is 0.2803 and noise_std is 0.9599; by t=1000 the sample is nearly pure noise (signal=0.0064)

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The forward process destroys all structure by T=1000.

31. DDPM noise schedule — real values

Worked example

Compute the signal fraction sqrt(alpha_bar_t) and noise std at key timesteps for T=1000, beta linearly spaced from 1e-4 to 0.02.

import numpy as np
T = 1000
betas = np.linspace(1e-4, 0.02, T)
alpha_bar = np.cumprod(1 - betas)
for t in [0, 49, 99, 249, 499, 749, 999]:
    sb = np.sqrt(alpha_bar[t])
    ns = np.sqrt(1 - alpha_bar[t])
    print(f't={t+1:4d}  signal={sb:.4f}  noise={ns:.4f}')

By t=500 signal_fraction is 0.2803 and noise_std is 0.9599; by t=1000 the sample is nearly pure noise (signal=0.0064)

Why: The forward process destroys all structure by T=1000. The reverse model (the neural network) learns to undo each small step. The closed form of q(x_t|x_0) means you can directly sample any t during training without simulating the Markov chain.

tsqrt(alpha_bar_t) — signalnoise std
10.99990.0100
500.98540.1702
1000.94710.3209
2500.72390.6899
5000.28030.9599
7500.05790.9983
10000.00641.0000

32. Fill in: noise std for DDPM noise schedule — real values

Comparison

Comparison matrix

From DDPM noise schedule — real values: refill the noise std column from what you know. The rest of the table is as it appeared.

tsqrt(alpha_bar_t) — signalnoise std
10.99990.0100
500.98540.1702
1000.94710.3209
2500.72390.6899
5000.28030.9599
7500.05790.9983
10000.00641.0000

33. Something is wrong here: confusing the saturating and non-saturating generator…

Anomaly

Predict first

A student writes this, and it looks reasonable:

Train G by minimizing E[log(1 - D(G(z)))] — the original minimax objective.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Early in training, when G is weak and D(G(z)) is near 0, the gradient of log(1-D(G(z))) is nearly flat — the generator receives almost no signal to improve.

In practice, minimize -E[log D(G(z))] (non-saturating) — same optimum, stronger gradients early.

Why: Early in training, when G is weak and D(G(z)) is near 0, the gradient of log(1-D(G(z))) is nearly flat — the generator receives almost no signal to improve. Training stalls.

34. Trap: confusing the saturating and non-saturating generator loss

Trap

The trap

Train G by minimizing E[log(1 - D(G(z)))] — the original minimax objective.

Use saturating G loss = E[log(1-D(G(z)))]

Why: Early in training, when G is weak and D(G(z)) is near 0, the gradient of log(1-D(G(z))) is nearly flat — the generator receives almost no signal to improve. Training stalls.

The fix

In practice, minimize -E[log D(G(z))] (non-saturating) — same optimum, stronger gradients early.

Use non-saturating G loss = -E[log D(G(z))] = 1.3999 (vs saturating -0.2892)

Why: The non-saturating loss has a large gradient when D(G(z)) is near 0 (early training). Same Nash equilibrium, but the generator gets a useful learning signal from the first iteration.

35. Which of these survive contact with Lesson 117: Phase 3 Gap Analysis & Phase 4…?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Phase 3 (Lessons 87-116) covered transformers, NLP, and computer vision. Rate yourself 0-10 on each cluster before seeing the benchmarks.; Phase 4 (Weeks 40-52) has 39 lessons. The first pass covers new content (VAE, GAN, Diffusion). Lessons marked 'review' are your reclaimed hours.; The KL term forces the encoder distribution q(z|x) toward the prior p(z)=N(0,I). The reconstruction term rewards the decoder for recovering x from the sampled z.
Breaks
To generate a new sample, sample z from the AE's latent space and decode it.; Train G by minimizing E[log(1 - D(G(z)))] — the original minimax objective.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 117: Phase 3 Gap Analysis & Phase 4 Preview puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

36. Exam strategy — time, derivations, mindset

Section

Part 3 of 4

37. Time allocation model

Concept

A typical USAAIO-style exam allocates roughly 180 minutes across three section types. Map your minutes to points-per-minute to decide where to spend effort.

sectiontime (min)pointspts/min
Multiple Choice (30 q)45300.667
Short Answer / Derivations (5 q)60400.667
Coding Problems (3 q)75300.400

Multiple Choice and Short Answer are equal in pts/min. Coding costs more time per point — do it last and prioritize partial-credit comments if you're short on time.

38. Teach it back: Time allocation model

Explain it

Discussion prompt

Explain Time allocation model to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

A typical USAAIO-style exam allocates roughly 180 minutes across three section types. Map your minutes to points-per-minute to decide where to spend effort.

39. Derivation approach

Concept

For derivation questions: write the goal first (what you're deriving), set up notation (define all symbols), then proceed line by line with one algebraic step per line and a brief justification at each dimension-change.

  1. State objective / what to prove (1-2 lines)
  2. Write the definition (e.g., ELBO = E[log p(x|z)] - KL)
  3. Expand KL for Gaussian encoder in closed form
  4. Simplify; check shapes and signs
  5. State the result clearly (boxed or underlined)

Partial credit is awarded for correct structure even if the algebra has a sign error at step 3. Never skip steps — an examiner can't award credit for invisible work.

40. By analogy: Derivation approach

Analogy

Discussion prompt

Explain Derivation approach by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Partial credit is awarded for correct structure even if the algebra has a sign error at step 3. Never skip steps — an examiner can't award credit for invisible work.

41. Without one step: The gap-analysis + Phase 4 ramp-up recipe

Constraint

Discussion prompt

Run The gap-analysis + Phase 4 ramp-up recipe with this step confiscated:

Exam: time budget — 0.667 pts/min on MC and derivations, 0.40 on coding; coding is last

Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.

Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.

Answer:

  1. Score each Phase 3 cluster (0-10), compute gap * exam_weight to rank priorities
  2. Re-allocate Phase 4 review slots to your top-3 weak clusters
  3. New content (Phase 4): VAE (ELBO + reparameterization) → GAN (minimax + non-saturating G loss) → DDPM (beta schedule, closed-form q(x_t|x_0))
  4. Exam: time budget — 0.667 pts/min on MC and derivations, 0.40 on coding; coding is last
  5. Derivation order: goal → notation → definition → line-by-line algebra → boxed result
  6. Weekly review: after each mock, re-run gap scoring; rotate review slots to new worst cluster

42. The gap-analysis + Phase 4 ramp-up recipe

Pattern

  1. Score each Phase 3 cluster (0-10), compute gap * exam_weight to rank priorities
  2. Re-allocate Phase 4 review slots to your top-3 weak clusters
  3. New content (Phase 4): VAE (ELBO + reparameterization) → GAN (minimax + non-saturating G loss) → DDPM (beta schedule, closed-form q(x_t|x_0))
  4. Exam: time budget — 0.667 pts/min on MC and derivations, 0.40 on coding; coding is last
  5. Derivation order: goal → notation → definition → line-by-line algebra → boxed result
  6. Weekly review: after each mock, re-run gap scoring; rotate review slots to new worst cluster

43. Where does each piece belong: Lesson 117: Phase 3 Gap Analysis & Phase 4…

Sorting

Sort into buckets

These are the pieces of Lesson 117: Phase 3 Gap Analysis & Phase 4 Preview, out of order. Put each one back under the part of the lesson it belongs to.

Phase 3 self-assessment — score the nine…
The nine Phase 3 clusters; Scoring example — running the audit; Phase 4 study plan re-allocation
Phase 4 preview — VAE, GAN, Diffusion
VAE — the ELBO objective; ELBO with real numbers; Reparameterization trick
Exam strategy — time, derivations, mindset
Time allocation model; Derivation approach; The gap-analysis + Phase 4 ramp-up recipe
s1
Phase 3 self-assessment — score the nine… is where Lesson 117: Phase 3 Gap Analysis & Phase 4 Preview puts The nine Phase 3 clusters, Scoring example — running the audit, Phase 4 study plan re-allocation. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
Phase 4 preview — VAE, GAN, Diffusion is where Lesson 117: Phase 3 Gap Analysis & Phase 4 Preview puts VAE — the ELBO objective, ELBO with real numbers, Reparameterization trick. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
Exam strategy — time, derivations, mindset is where Lesson 117: Phase 3 Gap Analysis & Phase 4 Preview puts Time allocation model, Derivation approach, The gap-analysis + Phase 4 ramp-up recipe. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

44. Rule out three: Check yourself — VAE ELBO

Elimination

Eliminate the wrong options

In the VAE ELBO = E[log p(x|z)] - KL(q(z|x)||p(z)), the KL term for a Gaussian encoder with mu=0, sigma^2=1 equals:

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. 0 — the encoder already matches the prior
  • B. 0.5*(0 + 1 - 1 - 0) = 0, confirmed
  • C. 1 — the KL is always at least 1 for Gaussian encoders
  • D. 0.5*(mu^2 + sigma^2) = 0.5

Survives elimination: B

Why: KL(N(mu,sigma^2)||N(0,1)) = 0.5(mu^2 + sigma^2 - 1 - log sigma^2). At mu=0, sigma^2=1: 0.5(0+1-1-log1) = 0.5*(0) = 0. Choices A and B both state this — B shows the algebra explicitly, confirming the formula.

45. Check yourself — VAE ELBO

Check

Compute mentally, then verify.

Check your understanding

In the VAE ELBO = E[log p(x|z)] - KL(q(z|x)||p(z)), the KL term for a Gaussian encoder with mu=0, sigma^2=1 equals:

  • A. 0 — the encoder already matches the prior
  • B. 0.5*(0 + 1 - 1 - 0) = 0, confirmed (correct)
  • C. 1 — the KL is always at least 1 for Gaussian encoders
  • D. 0.5*(mu^2 + sigma^2) = 0.5

Answer: B

Why: KL(N(mu,sigma^2)||N(0,1)) = 0.5(mu^2 + sigma^2 - 1 - log sigma^2). At mu=0, sigma^2=1: 0.5(0+1-1-log1) = 0.5*(0) = 0. Choices A and B both state this — B shows the algebra explicitly, confirming the formula.

Why A tempts people
Correct conclusion but shows no algebra — on a derivation question you'd lose partial credit for skipping the formula.
Why C tempts people
KL is non-negative but has no lower bound of 1; it reaches 0 when q matches p exactly.
Why D tempts people
This omits the -1 and -log(sigma^2) terms from the KL formula, giving the wrong value even when sigma^2=1.

46. Answer it before you see the options: Check yourself — DDPM forward process

Prediction

Predict first

In DDPM with T=1000 and a linear beta schedule (1e-4 to 0.02), the signal fraction sqrt(alpha_bar_t) at t=500 is approximately:

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: 0.28 (about 28% signal, 96% noise std)

Why: From real execution: sqrt(alpha_bar_500) = 0.2803, noise_std = 0.9599. The signal decays super-linearly (not linearly) because alpha_bar is a product — by t=500 less than 8% of the variance is signal.

47. Check yourself — DDPM forward process

Check

From the verified table: at t=500 (out of T=1000), what fraction of the original signal remains?

Check your understanding

In DDPM with T=1000 and a linear beta schedule (1e-4 to 0.02), the signal fraction sqrt(alpha_bar_t) at t=500 is approximately:

  • A. 0.28 (about 28% signal, 96% noise std) (correct)
  • B. 0.50 (halfway through, half the signal remains)
  • C. 0.72 (signal decays slowly, still mostly clean at t=500)
  • D. 0.00 (signal is destroyed well before t=T)

Answer: A

Why: From real execution: sqrt(alpha_bar_500) = 0.2803, noise_std = 0.9599. The signal decays super-linearly (not linearly) because alpha_bar is a product — by t=500 less than 8% of the variance is signal.

Why B tempts people
A linear intuition (halfway = half signal) ignores that alpha_bar is a cumulative product. Products of numbers less than 1 decay much faster than their arithmetic halfway point.
Why C tempts people
0.72 corresponds to t=250, not t=500. The schedule is non-linear — signal drops steeply in the second half.
Why D tempts people
At t=500 there is still 0.28 signal fraction; complete destruction (< 0.01) occurs only near t=1000.

48. Rule out three: Check yourself — GAN gradient saturation

Elimination

Eliminate the wrong options

When D(G(z)) is near 0, the saturating generator loss E[log(1-D(G(z)))] has a gradient that is:

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Nearly zero — log(1-x) is flat near x=0; G gets almost no signal
  • B. Large and negative — the log curve is steep near x=0
  • C. Constant — the gradient of a log is always 1/x
  • D. Undefined — the log is undefined at x=0

Survives elimination: A

Why: d/dx log(1-x) = -1/(1-x). Near x=0 this equals -1, which is small compared to the gradient of -log(x) = 1/x which blows up near x=0. The saturating loss provides weak gradients early — the non-saturating -E[log D(G(z))] has gradient 1/D(G(z)) -> large when D(G(z)) is small.

49. Check yourself — GAN gradient saturation

Check

Early training: D is strong, D(G(z)) ~ 0.05 for all generated samples.

Check your understanding

When D(G(z)) is near 0, the saturating generator loss E[log(1-D(G(z)))] has a gradient that is:

  • A. Nearly zero — log(1-x) is flat near x=0; G gets almost no signal (correct)
  • B. Large and negative — the log curve is steep near x=0
  • C. Constant — the gradient of a log is always 1/x
  • D. Undefined — the log is undefined at x=0

Answer: A

Why: d/dx log(1-x) = -1/(1-x). Near x=0 this equals -1, which is small compared to the gradient of -log(x) = 1/x which blows up near x=0. The saturating loss provides weak gradients early — the non-saturating -E[log D(G(z))] has gradient 1/D(G(z)) -> large when D(G(z)) is small.

Why B tempts people
log(1-x) is NOT steep near x=0; its slope is -1/(1-x) which equals -1 at x=0. It is -log(x) that is steep there — and that is the non-saturating loss, not the saturating one.
Why C tempts people
d/dx log(1-x) = -1/(1-x), not 1/x. The gradient depends on the argument; there is no universal constant.
Why D tempts people
log(1-x) is perfectly defined at x=0 (log(1)=0). It is log(x) itself that is undefined at x=0.

50. Your turn: Gap audit + Phase 4 plan

Section

Project

51. Project: audit and plan

Concept

Three deliverables: a scored Phase 3 gap table, a Phase 4 re-allocation plan (with hours per cluster), and a short derivation of the KL term for a Gaussian encoder.

#deliverabletool / method
1Phase 3 gap table (9 clusters, scored + prioritized)numpy priority scoring
2Phase 4 hour allocation (new content + review + mocks)written plan, ~half page
3Derive KL(N(mu,sigma^2) || N(0,1)) in closed formpen and paper derivation

Build rules: score honestly (not what you wish), show every algebra step in the KL derivation, and re-check your plan against the verified timing model (0.667 pts/min for MC/derivations).

52. Break it if you can: Project: audit and plan

Counterexample

Discussion prompt

Three deliverables: a scored Phase 3 gap table, a Phase 4 re-allocation plan (with hours per cluster), and a short derivation of the KL term for a Gaussian encoder.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Build rules: score honestly (not what you wish), show every algebra step in the KL derivation, and re-check your plan against the verified timing model (0.667 pts/min for MC/derivations).

53. Milestone 1 — run the priority scorer

Worked example

Your turn: fill in your own scores (0-10) for each cluster and run the priority computation. Which cluster tops your list?

Hint: exam weights for the nine clusters (self-attn=3, BERT=2, ViT=2, GNN=1, NER=1, ObjDet=2, Seq2Seq=2, BLEU=1, PosEnc=2) reflect typical USAAIO derivation frequency.

import numpy as np
# Replace with YOUR scores:
scores = np.array([8, 7, 9, 6, 7, 6, 8, 5, 9])
wts    = np.array([3, 2, 2, 1, 1, 2, 2, 1, 2])
topics = ['self-attn','BERT/GPT','ViT','GNN','NER',
          'ObjDet','Seq2Seq','BLEU','PosEnc']
priority = (10 - scores) * wts
for i in np.argsort(-priority):
    print(f'{priority[i]:3.0f}  {topics[i]:12s}  score={scores[i]}')
priority ptsclusterscore
6BERT/GPT7
6ObjDet6
5BLEU5
4GNN6
4Seq2Seq8

54. What each one costs: Milestone 1 — run the priority scorer

Trade off

Comparison matrix

From Milestone 1 — run the priority scorer: every row here is a choice with a cost. Fill the score column, then say which row you would actually pick and what you give up for it.

priority ptsclusterscore
6BERT/GPT7
6ObjDet6
5BLEU5
4GNN6
4Seq2Seq8

55. Milestone 2 — derive the Gaussian KL

Worked example

Your turn: derive KL(N(mu, sigma^2) || N(0,1)) in closed form. Start from the definition of KL divergence between two Gaussians.

Hint: integrate log[p(z)/q(z|x)] under q(z|x). Use E_q[z^2] = mu^2 + sigma^2 and E_q[log sigma^2] = log sigma^2.

import numpy as np
# Verify closed-form KL for three (mu, sigma^2) pairs
test_cases = [(0.0, 1.0), (1.5, 0.5), (0.0, 2.0)]
print(f'{"mu":>6} {"sigma^2":>8} {"KL":>10} {"expected":>12}')
for mu, s2 in test_cases:
    lv = np.log(s2)
    kl = 0.5 * (mu**2 + s2 - 1 - lv)
    print(f'{mu:>6.1f} {s2:>8.2f} {kl:>10.4f}')
musigma^2KL(q||p)
0.01.00.0000 (matches prior)
1.50.51.2216 (far from prior)
0.02.00.1534 (wider than prior)

56. Fill in: KL(q||p) for Milestone 2 — derive the Gaussian KL

Comparison

Comparison matrix

From Milestone 2 — derive the Gaussian KL: refill the KL(q||p) column from what you know. The rest of the table is as it appeared.

musigma^2KL(q||p)
0.01.00.0000 (matches prior)
1.50.51.2216 (far from prior)
0.02.00.1534 (wider than prior)

57. Show it off

Concept

With slides closed, explain: (1) which Phase 3 cluster is your personal top priority and why the gap*weight formula puts it there; (2) the ELBO in one sentence; (3) why GAN training uses the non-saturating G loss.

Stretch (homework, per syllabus): write your Phase 4 study priorities ordered by importance; re-read the USAAIO syllabus and check off each confident topic; download and read the DDPM paper (Ho et al. 2020) abstract and intro. Take one full day off before Phase 4 starts — consolidation happens during rest.

58. Connect it up: Lesson 117: Phase 3 Gap Analysis & Phase 4 Preview

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Phase 3 self-assessment — score the nine clusters · Phase 4 preview — VAE, GAN, Diffusion · Exam strategy — time, derivations, mindset · Your turn: Gap audit + Phase 4 plan. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

59. What you can do now

Recap

topicthe one thing to remember
VAE ELBO= reconstruction - KL; KL=0 when encoder = prior
reparameterizationz = mu + sigma * eps; makes sampling differentiable
GAN minimaxNash equilibrium at D*=0.5, V=-1.3863
non-saturating G loss-E[log D(G(z))]; large gradient when D(G(z)) near 0
DDPM q(x_t|x_0)closed form: N(sqrt(alpha_bar_t)x_0, (1-alpha_bar_t)I)
exam timing0.667 pts/min MC+derivation; coding last

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 117 — Phase 3 Gap Analysis & Phase 4 Preview — Barron · USAAIO Round 2 Preparation, 2026
  2. DDPM linear noise schedule and ELBO decomposition verified with numpy 2.2.6 and torch 2.7.1+cpu, June 2026 — Real execution, verified
  3. Ho et al. 'Denoising Diffusion Probabilistic Models' (NeurIPS 2020) — arXiv:2006.11239
  4. Kingma & Welling 'Auto-Encoding Variational Bayes' (ICLR 2014) — arXiv:1312.6114
  5. Goodfellow et al. 'Generative Adversarial Networks' (NeurIPS 2014) — arXiv:1406.2661

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108