USAAIO Lesson 117, at the end of Phase 3. It is a structured gap analysis across all nine Phase 3 topic clusters - transformers, BERT and GPT, ViT, GNNs, NLP, computer vision, seq2seq, BLEU and ROUGE, and positional encoding - with a scored self-assessment and a way to adjust your study plan for the weak areas. It then previews Phase 4, covering the VAE and its ELBO, the GAN minimax objective, and the DDPM noise schedule, and closes on exam strategy: allocating your time, approaching a derivation, and mindset. All the math was verified by real Python execution with numpy 2.2.6 and torch 2.7.1+cpu. The lesson runs to 29 slides.
Subject: Machine Learning · 59 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 117 · End of Phase 3
Locate your weakest Phase 3 concepts, recalibrate your Phase 4 study plan, and get a concrete preview of VAE, GAN, and diffusion. The exam is closer than it feels.
Objectives
Warm-up
Discussion prompt
Before we open Lesson 117: Phase 3 Gap Analysis & Phase 4 Preview: without looking back, what was the main idea of NLP + CV Review & Exam Simulation, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
comprehensive NLP+CV exam-simulation deck covering tokenization/word2vec/BERT/LoRA/seq2seq/beam-search/BLEU/ROUGE and convolution/ResNet/ViT/U-Net/YOLO/IoU/NMS/CLIP. Emphasizes cross-modal connections through transformers, pattern recognition for exam pace, and gap-spotting for Phase 4.
Section
Part 1 of 4
Concept
Phase 3 (Lessons 87-116) covered transformers, NLP, and computer vision. Rate yourself 0-10 on each cluster before seeing the benchmarks.
| # | cluster | key lessons | your score |
|---|---|---|---|
| 1 | Transformer self-attention | 87-90 | /10 |
| 2 | BERT / GPT fine-tuning | 91, 94 | /10 |
| 3 | ViT patch embedding | 92 | /10 |
| 4 | GNN message passing | 95-97 | /10 |
| 5 | NER / sequence labeling | 99-101 | /10 |
| 6 | Object detection (IoU/NMS) | 105-107 | /10 |
| 7 | Seq2Seq + cross-attention | 88, 103 | /10 |
| 8 | BLEU / ROUGE metrics | 102-104 | /10 |
| 9 | Positional encoding variants | 89, 92 | /10 |
Clusters with a score of 7 or below go on your Phase 4 re-study list. Clusters at 8+ are consolidate-and-move-on.
Counterexample
Discussion prompt
Phase 3 (Lessons 87-116) covered transformers, NLP, and computer vision. Rate yourself 0-10 on each cluster before seeing the benchmarks.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Estimation
Predict first
Compute a weighted priority score: priority = (10 - score) * exam_weight. Here exam_weight is an estimate of how often the topic appears in USAAIO derivation problems.
Commit before you compute: what does Scoring example — running the audit come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Rank by gap * exam_weight to find where extra study hours yield the most exam points
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. A score-6 topic with high exam weight (like GNN or ObjDet) costs more than a score-7 on a low-weight topic.
Worked example
Compute a weighted priority score: priority = (10 - score) * exam_weight. Here exam_weight is an estimate of how often the topic appears in USAAIO derivation problems.
import numpy as np
scores = np.array([8, 7, 9, 6, 7, 6, 8, 5, 9])
wts = np.array([3, 2, 2, 1, 1, 2, 2, 1, 2])
gap = 10 - scores
priority = gap * wts
topics = ['self-attn','BERT/GPT','ViT','GNN','NER',
'ObjDet','Seq2Seq','BLEU','PosEnc']
order = np.argsort(-priority)
for i in order:
print(f'{priority[i]:3.0f} pts {topics[i]} (score {scores[i]})')Rank by gap * exam_weight to find where extra study hours yield the most exam points
Why: A score-6 topic with high exam weight (like GNN or ObjDet) costs more than a score-7 on a low-weight topic. Priority = gap x weight makes the trade-off explicit.
| priority pts | cluster | score /10 |
|---|---|---|
| 6 | BERT/GPT (wt=2) | 7 |
| 6 | ObjDet (wt=2) | 6 |
| 5 | BLEU (wt=1) | 5 |
| 4 | GNN (wt=1) | 6 |
| 4 | Seq2Seq (wt=2) | 8 |
Comparison
Comparison matrix
From Scoring example — running the audit: refill the score /10 column from what you know. The rest of the table is as it appeared.
| priority pts | cluster | score /10 |
|---|---|---|
| 6 | BERT/GPT (wt=2) | 7 |
| 6 | ObjDet (wt=2) | 6 |
| 5 | BLEU (wt=1) | 5 |
| 4 | GNN (wt=1) | 6 |
| 4 | Seq2Seq (wt=2) | 8 |
Concept
Phase 4 (Weeks 40-52) has 39 lessons. The first pass covers new content (VAE, GAN, Diffusion). Lessons marked 'review' are your reclaimed hours.
| phase | lessons | hours | strategy |
|---|---|---|---|
| new content | 118-130 | 26 h | attend fully; build each model |
| review slots | 131-135 | 10 h | drill top-3 gap clusters |
| mock exams | 136-152 | 34 h | timed, then targeted debrief |
| final sprint | 153-156 | 8 h | weak-area formula sheets only |
Rule: if a mock-exam debrief reveals a cluster re-entering the 'weak' list, pull one review slot from Week 49 and re-drill it before the next mock.
Trade off
Comparison matrix
From Phase 4 study plan re-allocation: every row here is a choice with a cost. Fill the lessons column, then say which row you would actually pick and what you give up for it.
| phase | lessons | hours | strategy |
|---|---|---|---|
| new content | 118-130 | 26 h | attend fully; build each model |
| review slots | 131-135 | 10 h | drill top-3 gap clusters |
| mock exams | 136-152 | 34 h | timed, then targeted debrief |
| final sprint | 153-156 | 8 h | weak-area formula sheets only |
Section
Part 2 of 4
Concept
A VAE (Kingma & Welling 2014) learns a continuous, structured latent space by maximizing the Evidence Lower BOund (ELBO). Unlike a plain AE, the latent code is regularized to approximate N(0,I).
\[ \mathcal{L}(\theta,\phi;x) = \underbrace{\mathbb{E}_{q_\phi(z|x)}[\log p_\theta(x|z)]}_{\text{reconstruction}} - \underbrace{D_{\mathrm{KL}}(q_\phi(z|x)\,\|\,p(z))}_{\text{regularization}} \]
The KL term forces the encoder distribution q(z|x) toward the prior p(z)=N(0,I). The reconstruction term rewards the decoder for recovering x from the sampled z.
Analogy
Discussion prompt
Explain VAE — the ELBO objective by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
A VAE (Kingma & Welling 2014) learns a continuous, structured latent space by maximizing the Evidence Lower BOund (ELBO). Unlike a plain AE, the latent code is regularized to approximate N(0,I).
Missing information
Discussion prompt
Verify the ELBO for a single 1D datapoint: x=2.0, encoder outputs mu_z=1.5, log_var=-0.693 (sigma^2=0.5), decoder outputs x_recon=1.8.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
The KL term dominates here because mu_z=1.5 is far from the prior mean 0. Training maximizes ELBO, so the encoder is pushed to reduce KL while the decoder reduces reconstruction error.
Worked example
Verify the ELBO for a single 1D datapoint: x=2.0, encoder outputs mu_z=1.5, log_var=-0.693 (sigma^2=0.5), decoder outputs x_recon=1.8.
import numpy as np
x_obs, mu_z, log_var, x_recon = 2.0, 1.5, np.log(0.5), 1.8
# Reconstruction: log p(x|z) ~ -0.5*(x-x_recon)^2 (unit decoder var)
recon = -0.5 * (x_obs - x_recon)**2
# KL(N(mu,sigma^2) || N(0,1))
sigma2 = np.exp(log_var)
kl = 0.5 * (mu_z**2 + sigma2 - 1 - log_var)
elbo = recon - kl
print(f'recon = {recon:.4f}')
print(f'KL = {kl:.4f}')
print(f'ELBO = {elbo:.4f}')recon = -0.0200, KL = 1.2216, ELBO = -1.2416
Why: The KL term dominates here because mu_z=1.5 is far from the prior mean 0. Training maximizes ELBO, so the encoder is pushed to reduce KL while the decoder reduces reconstruction error.
| term | formula | value |
|---|---|---|
| reconstruction | -0.5*(2.0-1.8)^2 | -0.0200 |
| KL | 0.5*(1.5^2 + 0.5 - 1 - log(0.5)) | 1.2216 |
| ELBO | recon - KL | -1.2416 |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
recon = -0.0200, KL = 1.2216, ELBO = -1.2416
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Verify the ELBO for a single 1D datapoint: x=2.0, encoder outputs mu_z=1.5, log_var=-0.693 (sigma^2=0.5), decoder outputs x_recon=1.8.
Concept
Sampling z ~ N(mu, sigma^2) is non-differentiable. The reparameterization trick writes z = mu + sigma * eps where eps ~ N(0,1), making the gradient flow through mu and sigma.
\[ z = \mu_\phi(x) + \sigma_\phi(x) \cdot \varepsilon, \quad \varepsilon \sim \mathcal{N}(0,I) \]
With eps=0.3, mu=1.5, sigma=0.7071: z = 1.5 + 0.7071*0.3 = 1.7121. The gradient of the loss w.r.t. mu flows through the addition, not the sampling.
Explain it
Discussion prompt
Explain Reparameterization trick to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Sampling z ~ N(mu, sigma^2) is non-differentiable. The reparameterization trick writes z = mu + sigma * eps where eps ~ N(0,1), making the gradient flow through mu and sigma.
Anomaly
Predict first
A student writes this, and it looks reasonable:
To generate a new sample, sample z from the AE's latent space and decode it.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: A plain AE's latent space has no enforced structure — latent codes cluster arbitrarily.
Use a VAE: the KL term regularizes the encoder to approximate N(0,I).
Why: A plain AE's latent space has no enforced structure — latent codes cluster arbitrarily. Sampling z from a uniform or Gaussian prior will almost always land in an 'empty' region and produce nonsense.
Trap
To generate a new sample, sample z from the AE's latent space and decode it.
Sample z ~ Uniform[-1, 1] and pass through AE decoder
Why: A plain AE's latent space has no enforced structure — latent codes cluster arbitrarily. Sampling z from a uniform or Gaussian prior will almost always land in an 'empty' region and produce nonsense.
Use a VAE: the KL term regularizes the encoder to approximate N(0,I).
Sample z ~ N(0,I) and pass through the VAE decoder
Why: Because the VAE ELBO penalizes any encoder distribution that deviates from N(0,I), the decoder learns to handle samples from the prior — so at inference, sampling z~N(0,I) produces coherent outputs.
Break the constraint
Discussion prompt
The rule this trap just fixed:
Because the VAE ELBO penalizes any encoder distribution that deviates from N(0,I), the decoder learns to handle samples from the prior — so at inference, sampling z~N(0,I) produces coherent outputs.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
A plain AE's latent space has no enforced structure — latent codes cluster arbitrarily. Sampling z from a uniform or Gaussian prior will almost always land in an 'empty' region and produce nonsense.
Concept
A GAN (Goodfellow 2014) trains a discriminator D to distinguish real from fake, and a generator G to fool D. They play a zero-sum game.
\[ \min_G \max_D \;V(D,G) = \mathbb{E}_{x\sim p_{\mathrm{data}}}[\log D(x)] + \mathbb{E}_{z\sim p_z}[\log(1-D(G(z)))] \]
D maximizes V (pushes D(real) toward 1, D(fake) toward 0). G minimizes V (pushes D(G(z)) toward 1). At Nash equilibrium D(x) = 0.5 everywhere and V = -2ln(2) = -1.3863.
Socratic
Discussion prompt
A GAN (Goodfellow 2014) trains a discriminator D to distinguish real from fake, and a generator G to fool D. They play a zero-sum game.
Suppose that were not true. What is the first thing in Lesson 117: Phase 3 Gap Analysis & Phase 4 Preview that would stop working?
Hint: Follow it one step downstream. The answer is whatever was quietly relying on it.
Answer:
D maximizes V (pushes D(real) toward 1, D(fake) toward 0). G minimizes V (pushes D(G(z)) toward 1). At Nash equilibrium D(x) = 0.5 everywhere and V = -2ln(2) = -1.3863.
Estimation
Predict first
Compute V(D,G) for a batch where D outputs 0.9, 0.85, 0.8 on real samples and 0.3, 0.2, 0.25 on generated samples.
Commit before you compute: what does GAN objective — real numbers come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: V(D,G) = -0.4528; at Nash equilibrium V = -1.3863
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. D is winning here (V > -1.3863 means G has not yet fooled D).
Worked example
Compute V(D,G) for a batch where D outputs 0.9, 0.85, 0.8 on real samples and 0.3, 0.2, 0.25 on generated samples.
import numpy as np
D_real = np.array([0.9, 0.85, 0.8])
D_fake = np.array([0.3, 0.2, 0.25])
V = np.mean(np.log(D_real) + np.log(1 - D_fake))
G_loss_sat = np.mean(np.log(1 - D_fake))
G_loss_nonsat = -np.mean(np.log(D_fake))
print(f'V(D,G) = {V:.4f}')
print(f'G saturating loss = {G_loss_sat:.4f}')
print(f'G non-saturating loss = {G_loss_nonsat:.4f}')
print(f'Nash V(D*,G*) = 2*log(0.5) = {2*np.log(0.5):.4f}')V(D,G) = -0.4528; at Nash equilibrium V = -1.3863
Why: D is winning here (V > -1.3863 means G has not yet fooled D). Training continues until V approaches -1.3863, at which point D can't do better than random guessing.
| quantity | formula | value |
|---|---|---|
| V(D,G) | E[logD(x)] + E[log(1-D(G(z)))] | -0.4528 |
| G saturating loss | E[log(1-D(G(z)))] | -0.2892 |
| G non-saturating loss | -E[log D(G(z))] | 1.3999 |
| Nash V(D,G) | 2*log(0.5) | -1.3863 |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
V(D,G) = -0.4528; at Nash equilibrium V = -1.3863
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Compute V(D,G) for a batch where D outputs 0.9, 0.85, 0.8 on real samples and 0.3, 0.2, 0.25 on generated samples.
Concept
DDPM (Ho et al. 2020) defines a forward process that gradually corrupts a clean sample x_0 by adding Gaussian noise over T=1000 steps with a linear beta schedule.
\[ q(x_t \mid x_0) = \mathcal{N}\!\left(\sqrt{\bar{\alpha}_t}\,x_0,\;(1-\bar{\alpha}_t)I\right), \quad \bar{\alpha}_t = \prod_{s=1}^{t}(1-\beta_s) \]
The closed form means you can jump to any noise level t in one step — no need to run t sequential steps during training. The network learns to reverse each small step.
Socratic
Discussion prompt
DDPM (Ho et al. 2020) defines a forward process that gradually corrupts a clean sample x_0 by adding Gaussian noise over T=1000 steps with a linear beta schedule.
Suppose that were not true. What is the first thing in Lesson 117: Phase 3 Gap Analysis & Phase 4 Preview that would stop working?
Hint: Follow it one step downstream. The answer is whatever was quietly relying on it.
Answer:
The closed form means you can jump to any noise level t in one step — no need to run t sequential steps during training. The network learns to reverse each small step.
Estimation
Predict first
Compute the signal fraction sqrt(alpha_bar_t) and noise std at key timesteps for T=1000, beta linearly spaced from 1e-4 to 0.02.
Commit before you compute: what does DDPM noise schedule — real values come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: By t=500 signal_fraction is 0.2803 and noise_std is 0.9599; by t=1000 the sample is nearly pure noise (signal=0.0064)
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The forward process destroys all structure by T=1000.
Worked example
Compute the signal fraction sqrt(alpha_bar_t) and noise std at key timesteps for T=1000, beta linearly spaced from 1e-4 to 0.02.
import numpy as np
T = 1000
betas = np.linspace(1e-4, 0.02, T)
alpha_bar = np.cumprod(1 - betas)
for t in [0, 49, 99, 249, 499, 749, 999]:
sb = np.sqrt(alpha_bar[t])
ns = np.sqrt(1 - alpha_bar[t])
print(f't={t+1:4d} signal={sb:.4f} noise={ns:.4f}')By t=500 signal_fraction is 0.2803 and noise_std is 0.9599; by t=1000 the sample is nearly pure noise (signal=0.0064)
Why: The forward process destroys all structure by T=1000. The reverse model (the neural network) learns to undo each small step. The closed form of q(x_t|x_0) means you can directly sample any t during training without simulating the Markov chain.
| t | sqrt(alpha_bar_t) — signal | noise std |
|---|---|---|
| 1 | 0.9999 | 0.0100 |
| 50 | 0.9854 | 0.1702 |
| 100 | 0.9471 | 0.3209 |
| 250 | 0.7239 | 0.6899 |
| 500 | 0.2803 | 0.9599 |
| 750 | 0.0579 | 0.9983 |
| 1000 | 0.0064 | 1.0000 |
Comparison
Comparison matrix
From DDPM noise schedule — real values: refill the noise std column from what you know. The rest of the table is as it appeared.
| t | sqrt(alpha_bar_t) — signal | noise std |
|---|---|---|
| 1 | 0.9999 | 0.0100 |
| 50 | 0.9854 | 0.1702 |
| 100 | 0.9471 | 0.3209 |
| 250 | 0.7239 | 0.6899 |
| 500 | 0.2803 | 0.9599 |
| 750 | 0.0579 | 0.9983 |
| 1000 | 0.0064 | 1.0000 |
Anomaly
Predict first
A student writes this, and it looks reasonable:
Train G by minimizing E[log(1 - D(G(z)))] — the original minimax objective.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Early in training, when G is weak and D(G(z)) is near 0, the gradient of log(1-D(G(z))) is nearly flat — the generator receives almost no signal to improve.
In practice, minimize -E[log D(G(z))] (non-saturating) — same optimum, stronger gradients early.
Why: Early in training, when G is weak and D(G(z)) is near 0, the gradient of log(1-D(G(z))) is nearly flat — the generator receives almost no signal to improve. Training stalls.
Trap
Train G by minimizing E[log(1 - D(G(z)))] — the original minimax objective.
Use saturating G loss = E[log(1-D(G(z)))]
Why: Early in training, when G is weak and D(G(z)) is near 0, the gradient of log(1-D(G(z))) is nearly flat — the generator receives almost no signal to improve. Training stalls.
In practice, minimize -E[log D(G(z))] (non-saturating) — same optimum, stronger gradients early.
Use non-saturating G loss = -E[log D(G(z))] = 1.3999 (vs saturating -0.2892)
Why: The non-saturating loss has a large gradient when D(G(z)) is near 0 (early training). Same Nash equilibrium, but the generator gets a useful learning signal from the first iteration.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
E[log(1 - D(G(z)))] — the original minimax objective.Section
Part 3 of 4
Concept
A typical USAAIO-style exam allocates roughly 180 minutes across three section types. Map your minutes to points-per-minute to decide where to spend effort.
| section | time (min) | points | pts/min |
|---|---|---|---|
| Multiple Choice (30 q) | 45 | 30 | 0.667 |
| Short Answer / Derivations (5 q) | 60 | 40 | 0.667 |
| Coding Problems (3 q) | 75 | 30 | 0.400 |
Multiple Choice and Short Answer are equal in pts/min. Coding costs more time per point — do it last and prioritize partial-credit comments if you're short on time.
Explain it
Discussion prompt
Explain Time allocation model to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
A typical USAAIO-style exam allocates roughly 180 minutes across three section types. Map your minutes to points-per-minute to decide where to spend effort.
Concept
For derivation questions: write the goal first (what you're deriving), set up notation (define all symbols), then proceed line by line with one algebraic step per line and a brief justification at each dimension-change.
Partial credit is awarded for correct structure even if the algebra has a sign error at step 3. Never skip steps — an examiner can't award credit for invisible work.
Analogy
Discussion prompt
Explain Derivation approach by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Partial credit is awarded for correct structure even if the algebra has a sign error at step 3. Never skip steps — an examiner can't award credit for invisible work.
Constraint
Discussion prompt
Run The gap-analysis + Phase 4 ramp-up recipe with this step confiscated:
Exam: time budget — 0.667 pts/min on MC and derivations, 0.40 on coding; coding is last
Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.
Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.
Answer:
gap * exam_weight to rank prioritiesPattern
gap * exam_weight to rank prioritiesSorting
Sort into buckets
These are the pieces of Lesson 117: Phase 3 Gap Analysis & Phase 4 Preview, out of order. Put each one back under the part of the lesson it belongs to.
Elimination
Eliminate the wrong options
In the VAE ELBO = E[log p(x|z)] - KL(q(z|x)||p(z)), the KL term for a Gaussian encoder with mu=0, sigma^2=1 equals:
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: B
Why: KL(N(mu,sigma^2)||N(0,1)) = 0.5(mu^2 + sigma^2 - 1 - log sigma^2). At mu=0, sigma^2=1: 0.5(0+1-1-log1) = 0.5*(0) = 0. Choices A and B both state this — B shows the algebra explicitly, confirming the formula.
Check
Compute mentally, then verify.
Check your understanding
In the VAE ELBO = E[log p(x|z)] - KL(q(z|x)||p(z)), the KL term for a Gaussian encoder with mu=0, sigma^2=1 equals:
Answer: B
Why: KL(N(mu,sigma^2)||N(0,1)) = 0.5(mu^2 + sigma^2 - 1 - log sigma^2). At mu=0, sigma^2=1: 0.5(0+1-1-log1) = 0.5*(0) = 0. Choices A and B both state this — B shows the algebra explicitly, confirming the formula.
Prediction
Predict first
In DDPM with T=1000 and a linear beta schedule (1e-4 to 0.02), the signal fraction sqrt(alpha_bar_t) at t=500 is approximately:
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: 0.28 (about 28% signal, 96% noise std)
Why: From real execution: sqrt(alpha_bar_500) = 0.2803, noise_std = 0.9599. The signal decays super-linearly (not linearly) because alpha_bar is a product — by t=500 less than 8% of the variance is signal.
Check
From the verified table: at t=500 (out of T=1000), what fraction of the original signal remains?
Check your understanding
In DDPM with T=1000 and a linear beta schedule (1e-4 to 0.02), the signal fraction sqrt(alpha_bar_t) at t=500 is approximately:
Answer: A
Why: From real execution: sqrt(alpha_bar_500) = 0.2803, noise_std = 0.9599. The signal decays super-linearly (not linearly) because alpha_bar is a product — by t=500 less than 8% of the variance is signal.
Elimination
Eliminate the wrong options
When D(G(z)) is near 0, the saturating generator loss E[log(1-D(G(z)))] has a gradient that is:
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: d/dx log(1-x) = -1/(1-x). Near x=0 this equals -1, which is small compared to the gradient of -log(x) = 1/x which blows up near x=0. The saturating loss provides weak gradients early — the non-saturating -E[log D(G(z))] has gradient 1/D(G(z)) -> large when D(G(z)) is small.
Check
Early training: D is strong, D(G(z)) ~ 0.05 for all generated samples.
Check your understanding
When D(G(z)) is near 0, the saturating generator loss E[log(1-D(G(z)))] has a gradient that is:
Answer: A
Why: d/dx log(1-x) = -1/(1-x). Near x=0 this equals -1, which is small compared to the gradient of -log(x) = 1/x which blows up near x=0. The saturating loss provides weak gradients early — the non-saturating -E[log D(G(z))] has gradient 1/D(G(z)) -> large when D(G(z)) is small.
Section
Project
Concept
Three deliverables: a scored Phase 3 gap table, a Phase 4 re-allocation plan (with hours per cluster), and a short derivation of the KL term for a Gaussian encoder.
| # | deliverable | tool / method |
|---|---|---|
| 1 | Phase 3 gap table (9 clusters, scored + prioritized) | numpy priority scoring |
| 2 | Phase 4 hour allocation (new content + review + mocks) | written plan, ~half page |
| 3 | Derive KL(N(mu,sigma^2) || N(0,1)) in closed form | pen and paper derivation |
Build rules: score honestly (not what you wish), show every algebra step in the KL derivation, and re-check your plan against the verified timing model (0.667 pts/min for MC/derivations).
Counterexample
Discussion prompt
Three deliverables: a scored Phase 3 gap table, a Phase 4 re-allocation plan (with hours per cluster), and a short derivation of the KL term for a Gaussian encoder.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Build rules: score honestly (not what you wish), show every algebra step in the KL derivation, and re-check your plan against the verified timing model (0.667 pts/min for MC/derivations).
Worked example
Your turn: fill in your own scores (0-10) for each cluster and run the priority computation. Which cluster tops your list?
Hint: exam weights for the nine clusters (self-attn=3, BERT=2, ViT=2, GNN=1, NER=1, ObjDet=2, Seq2Seq=2, BLEU=1, PosEnc=2) reflect typical USAAIO derivation frequency.
import numpy as np
# Replace with YOUR scores:
scores = np.array([8, 7, 9, 6, 7, 6, 8, 5, 9])
wts = np.array([3, 2, 2, 1, 1, 2, 2, 1, 2])
topics = ['self-attn','BERT/GPT','ViT','GNN','NER',
'ObjDet','Seq2Seq','BLEU','PosEnc']
priority = (10 - scores) * wts
for i in np.argsort(-priority):
print(f'{priority[i]:3.0f} {topics[i]:12s} score={scores[i]}')| priority pts | cluster | score |
|---|---|---|
| 6 | BERT/GPT | 7 |
| 6 | ObjDet | 6 |
| 5 | BLEU | 5 |
| 4 | GNN | 6 |
| 4 | Seq2Seq | 8 |
Trade off
Comparison matrix
From Milestone 1 — run the priority scorer: every row here is a choice with a cost. Fill the score column, then say which row you would actually pick and what you give up for it.
| priority pts | cluster | score |
|---|---|---|
| 6 | BERT/GPT | 7 |
| 6 | ObjDet | 6 |
| 5 | BLEU | 5 |
| 4 | GNN | 6 |
| 4 | Seq2Seq | 8 |
Worked example
Your turn: derive KL(N(mu, sigma^2) || N(0,1)) in closed form. Start from the definition of KL divergence between two Gaussians.
Hint: integrate log[p(z)/q(z|x)] under q(z|x). Use E_q[z^2] = mu^2 + sigma^2 and E_q[log sigma^2] = log sigma^2.
import numpy as np
# Verify closed-form KL for three (mu, sigma^2) pairs
test_cases = [(0.0, 1.0), (1.5, 0.5), (0.0, 2.0)]
print(f'{"mu":>6} {"sigma^2":>8} {"KL":>10} {"expected":>12}')
for mu, s2 in test_cases:
lv = np.log(s2)
kl = 0.5 * (mu**2 + s2 - 1 - lv)
print(f'{mu:>6.1f} {s2:>8.2f} {kl:>10.4f}')| mu | sigma^2 | KL(q||p) |
|---|---|---|
| 0.0 | 1.0 | 0.0000 (matches prior) |
| 1.5 | 0.5 | 1.2216 (far from prior) |
| 0.0 | 2.0 | 0.1534 (wider than prior) |
Comparison
Comparison matrix
From Milestone 2 — derive the Gaussian KL: refill the KL(q||p) column from what you know. The rest of the table is as it appeared.
| mu | sigma^2 | KL(q||p) |
|---|---|---|
| 0.0 | 1.0 | 0.0000 (matches prior) |
| 1.5 | 0.5 | 1.2216 (far from prior) |
| 0.0 | 2.0 | 0.1534 (wider than prior) |
Concept
With slides closed, explain: (1) which Phase 3 cluster is your personal top priority and why the gap*weight formula puts it there; (2) the ELBO in one sentence; (3) why GAN training uses the non-saturating G loss.
Stretch (homework, per syllabus): write your Phase 4 study priorities ordered by importance; re-read the USAAIO syllabus and check off each confident topic; download and read the DDPM paper (Ho et al. 2020) abstract and intro. Take one full day off before Phase 4 starts — consolidation happens during rest.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Phase 3 self-assessment — score the nine clusters · Phase 4 preview — VAE, GAN, Diffusion · Exam strategy — time, derivations, mindset · Your turn: Gap audit + Phase 4 plan. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
gap * exam_weight priority| topic | the one thing to remember |
|---|---|
| VAE ELBO | = reconstruction - KL; KL=0 when encoder = prior |
| reparameterization | z = mu + sigma * eps; makes sampling differentiable |
| GAN minimax | Nash equilibrium at D*=0.5, V=-1.3863 |
| non-saturating G loss | -E[log D(G(z))]; large gradient when D(G(z)) near 0 |
| DDPM q(x_t|x_0) | closed form: N(sqrt(alpha_bar_t)x_0, (1-alpha_bar_t)I) |
| exam timing | 0.667 pts/min MC+derivation; coding last |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.