USAAIO Lesson 14, from Week 5 on probability, fully worked. It builds surprise and entropy from a single axiom and computes the entropy of a running three-outcome weather distribution bit by bit. It then derives KL divergence as an extra coding cost and proves it asymmetric and non-negative, splits cross-entropy into H(P) + KL one algebra step at a time, derives mutual information from the joint table, and gets to the ELBO decomposition that underpins the VAE. Along the way it covers units, bits against nats, conditional entropy, and the closed form for the Gaussian KL. Every snippet runs standalone, and every number was produced by real execution. The lesson runs to 60 slides.
Subject: Machine Learning · 112 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 14 · Week 5 (Probability)
The math of uncertainty and the distance between distributions. We build surprise from one rule, turn it into entropy, derive KL divergence and prove it is asymmetric and ≥ 0, split cross-entropy into H(P) + KL, read mutual information off a joint table, and land on the ELBO that trains every VAE — no step skipped, every number run.
Objectives
−log p from one axiom and compute entropy H(X) = −Σ p log₂ p outcome by outcomeD_KL(P‖Q) = Σ P log₂(P/Q) as extra coding cost, and prove it is ≥ 0 and asymmetricH(P) + D_KL(P‖Q) step by step, and say exactly what minimizing it doesI(X;Y) from a joint table two ways and show it is 0 iff independentlog p(x) = ELBO + D_KL(q‖p), explain why it is a lower bound, and use the Gaussian-KL closed formWarm-up
Discussion prompt
Before we open Lesson 14: Information Theory: without looking back, what was the main idea of Positive (Semi)Definite Matrices, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
PSD vs PD via xᵀAx and eigenvalues, the Cholesky decomposition, why XᵀX and covariance are always PSD, and three ways to prove definiteness (definition, eigenvalues, Sylvester). Build a Cholesky-based PD checker from scratch.
Section
Part 1 of 6 — the one axiom
Concept
One distribution carries this whole lesson. A weather model says tomorrow is Sunny with probability 0.5, Cloudy 0.25, Rainy 0.25. Call it p.
| outcome | symbol | p(x) |
|---|---|---|
| Sunny | S | 0.5 |
| Cloudy | C | 0.25 |
| Rainy | R | 0.25 |
Every quantity today — entropy, KL, cross-entropy, mutual information — will be computed against this p, so the numbers compound instead of resetting.
Comparison
Comparison matrix
From The running example: tomorrow's weather: refill the symbol column from what you know. The rest of the table is as it appeared.
| outcome | symbol | p(x) |
|---|---|---|
| Sunny | S | 0.5 |
| Cloudy | C | 0.25 |
| Rainy | R | 0.25 |
Intuition
If a friend texts 'it's sunny in the desert', you learn almost nothing — you expected that. If they text 'it's snowing in the desert', you learn a lot. Rarer outcomes are more informative when they happen.
So information should be a decreasing function of probability: small p → big information, p = 1 → zero information (a sure thing tells you nothing).
We also want information to add for independent events: learning two independent facts should cost the sum of their individual costs. Those two wishes pin down the formula exactly.
Counterexample
Discussion prompt
We also want information to add for independent events: learning two independent facts should cost the sum of their individual costs. Those two wishes pin down the formula exactly.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Concept
The surprise (self-information) of an outcome with probability p is defined as:
\[ I(x) = -\log_2 p(x) = \log_2 \frac{1}{p(x)} \]
The log is the only function that turns 'independent probabilities multiply' into 'surprises add'. Base 2 gives units of bits; natural log would give nats.
Check the two wishes
Why: As p→1, −log p → 0 (a certain event is unsurprising). As p→0, −log p → ∞ (a near-impossible event is hugely informative). And −log(p·q) = −log p − log q, so independent surprises add. Both design goals are satisfied.
Analogy
Discussion prompt
Explain Surprise = −log p by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
The surprise (self-information) of an outcome with probability p is defined as:
Estimation
Predict first
Plug each p(x) into −log₂ p. Powers of two make these exact — no calculator needed:
Commit before you compute: what does Surprise of each weather outcome come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Rainy is twice as surprising as Sunny
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Sunny (half the time) costs 1 bit; each quarter-probability outcome costs 2 bits.
Worked example
Plug each p(x) into −log₂ p. Powers of two make these exact — no calculator needed:
\[ \begin{aligned} I(S) &= -\log_2 0.5 = -\log_2 2^{-1} = 1 \\ I(C) &= -\log_2 0.25 = -\log_2 2^{-2} = 2 \\ I(R) &= -\log_2 0.25 = 2 \end{aligned} \]
| outcome | p(x) | surprise −log₂p (bits) |
|---|---|---|
| Sunny | 0.5 | 1 |
| Cloudy | 0.25 | 2 |
| Rainy | 0.25 | 2 |
Rainy is twice as surprising as Sunny
Why: Sunny (half the time) costs 1 bit; each quarter-probability outcome costs 2 bits. Halving the probability adds exactly 1 bit of surprise — the signature of a log scale.
Trade off
Comparison matrix
From Surprise of each weather outcome: every row here is a choice with a cost. Fill the p(x) column, then say which row you would actually pick and what you give up for it.
| outcome | p(x) | surprise −log₂p (bits) |
|---|---|---|
| Sunny | 0.5 | 1 |
| Cloudy | 0.25 | 2 |
| Rainy | 0.25 | 2 |
Concept
Entropy is the surprise you expect on average — weight each outcome's surprise by how often it happens:
\[ H(X) = \mathbb{E}_{x \sim p}\big[-\log_2 p(x)\big] = -\sum_x p(x)\,\log_2 p(x) \]
It is the average number of bits needed to encode one sample under an optimal code. Uncertain distributions cost more bits; certain ones cost none.
Ranking
Put in order
Put the moves of Entropy of the weather p, term by term into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. p(S)·I(S). Sunny is common but cheap, contributing half a bit.
Worked example
Multiply each surprise by its probability, then sum. We already have the surprises [1, 2, 2]:
Sunny term: 0.5 × 1 = 0.5
Why: p(S)·I(S). Sunny is common but cheap, contributing half a bit.
Cloudy term: 0.25 × 2 = 0.5
Why: p(C)·I(C). Rarer but more surprising — also half a bit.
Rainy term: 0.25 × 2 = 0.5
Why: p(R)·I(R). Same as cloudy by symmetry.
Sum: H(p) = 0.5 + 0.5 + 0.5 = 1.5 bits
Why: On average it takes 1.5 bits to encode one day of this weather. Notice each outcome happens to contribute exactly 0.5 — a coincidence of these particular probabilities, not a rule.
\[ H(p) = -\big(0.5\log_2 0.5 + 0.25\log_2 0.25 + 0.25\log_2 0.25\big) = 1.5 \]
Notation
Annotate
From Entropy of the weather p, term by term — read this one piece at a time. What is each part doing?
On: \( H(p) = -\big(0.5\log_2 0.5 + 0.25\log_2 0.25 + 0.25\log_2 0.25\big) = 1.5 \)
Missing information
Discussion prompt
The vectorized formula is one line. −np.log2(p) gives the surprise vector, p * (...) weights it, np.sum averages:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Exactly the hand values. −log2 of a power of two is an integer.
Worked example
The vectorized formula is one line. −np.log2(p) gives the surprise vector, p * (...) weights it, np.sum averages:
import numpy as np
p = np.array([0.5, 0.25, 0.25]) # Sunny, Cloudy, Rainy
surprise = -np.log2(p) # bits of surprise per outcome
H = -np.sum(p * np.log2(p))
print(surprise) # [1. 2. 2.]
print(round(H, 4)) # 1.5surprise = [1., 2., 2.]
Why: Exactly the hand values. −log2 of a power of two is an integer.
H = 1.5 bits
Why: Matches the term-by-term sum. This four-line entropy is the atom every other quantity today is built from.
| expression | value (verified) |
|---|---|
| -np.log2(p) | [1. 2. 2.] |
| p * -np.log2(p) | [0.5 0.5 0.5] |
| H = -np.sum(p*np.log2(p)) | 1.5 |
Error analysis
Annotate
Walk the callouts on Entropy in code — same 1.5 bits. Each one is a place this is easy to get subtly wrong.
Concept
The most uncertain distribution over n outcomes is the uniform one — nothing is predictable — and its entropy is the ceiling log₂ n.
\[ H_{\max} = -\sum_{i=1}^{n} \tfrac{1}{n}\log_2 \tfrac{1}{n} = \log_2 n \]
For n = 3, that ceiling is log₂ 3 ≈ 1.585 bits. Our weather p scored 1.5, just under the max — the 0.5/0.25/0.25 skew makes it slightly more predictable than uniform.
A certain outcome sits at the other extreme: H = 0 (no surprise ever). Entropy always lands in [0, log₂ n].
Explain it
Discussion prompt
Explain Entropy is maximized by the uniform to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
The most uncertain distribution over n outcomes is the uniform one — nothing is predictable — and its entropy is the ceiling log₂ n.
Pattern
Predict first
The table runs: [0.5, 0.5] | 1.0000 | max — most uncertain · [0.6, 0.4] | 0.9710 | slightly predictable
In Coins: from fair to certain, given the rows so far: what is the next one — the row where distribution is [1.0, 0.0]?
Correct: [1.0, 0.0] | 0.0000 | certain — no surprise
| distribution | H (bits) | reading |
|---|---|---|
| [0.5, 0.5] | 1.0000 | max — most uncertain |
| [0.6, 0.4] | 0.9710 | slightly predictable |
| [1.0, 0.0] | 0.0000 | certain — no surprise |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. log2(2) = 1. Maximum uncertainty for two outcomes.
Worked example
Watch entropy fall as a two-outcome distribution goes from balanced to skewed to certain. The d > 0 mask handles the 0·log 0 = 0 convention:
import numpy as np
for dist in ([0.5, 0.5], [0.6, 0.4], [1.0, 0.0]):
d = np.array(dist)
nz = d > 0 # skip 0*log0 (defined as 0)
H = -np.sum(d[nz] * np.log2(d[nz]))
print(dist, round(float(H), 4))Fair coin [0.5,0.5] → 1.0 bit, the 2-outcome max
Why: log2(2) = 1. Maximum uncertainty for two outcomes.
Biased [0.6,0.4] → 0.9710, and certain [1,0] → 0
Why: The skew drops entropy below 1; a sure outcome hits the floor of 0. Skew always reduces entropy.
| distribution | H (bits) | reading |
|---|---|---|
| [0.5, 0.5] | 1.0000 | max — most uncertain |
| [0.6, 0.4] | 0.9710 | slightly predictable |
| [1.0, 0.0] | 0.0000 | certain — no surprise |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Biased [0.6,0.4] → 0.9710, and certain [1,0] → 0
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Watch entropy fall as a two-outcome distribution goes from balanced to skewed to certain. The d > 0 mask handles the 0·log 0 = 0 convention:
Concept
The only free choice in every formula today is the log base. Base 2 gives bits; the natural log gives nats. Nothing else changes — the two answers are proportional.
\[ H_{\text{bits}} = -\sum p\log_2 p, \qquad H_{\text{nats}} = -\sum p\ln p, \qquad H_{\text{nats}} = H_{\text{bits}}\cdot \ln 2 \]
Convention: classifiers use nats (framework losses use ln), while the coding/compression story is cleanest in bits. Report the base or your number is ambiguous.
Fill the middle
Fill in the blanks
From Same entropy, two units — one line has had its right-hand side removed. Put it back.
import numpy as np, math
p = np.array([0.5, 0.25, 0.25])
H_bits = **-np.sum(p * np.log2(p))**
H_nats = -np.sum(p * np.log(p)) # natural log
print(round(float(H_bits), 4)) # 1.5
print(round(float(H_nats), 4)) # 1.0397
print(round(float(H_bits * math.log(2)), 4)) # 1.0397 == nats
Why: H_bits is what everything below it consumes, so the wrong expression here fails later and somewhere else. Same distribution, same uncertainty — only the yardstick changed.
Worked example
Compute H(p) for the weather in both bases and confirm the ln 2 bridge between them:
import numpy as np, math
p = np.array([0.5, 0.25, 0.25])
H_bits = -np.sum(p * np.log2(p))
H_nats = -np.sum(p * np.log(p)) # natural log
print(round(float(H_bits), 4)) # 1.5
print(round(float(H_nats), 4)) # 1.0397
print(round(float(H_bits * math.log(2)), 4)) # 1.0397 == natsH = 1.5 bits = 1.0397 nats
Why: Same distribution, same uncertainty — only the yardstick changed. ln 2 ≈ 0.6931 converts bits to nats.
H_bits · ln2 = 1.0397 reproduces the nats value
Why: The two are one multiply apart. Divide nats by ln 2 to get back to bits.
| quantity | value (verified) |
|---|---|
| H(p) in bits (log₂) | 1.5000 |
| H(p) in nats (ln) | 1.0397 |
| 1.5 × ln 2 | 1.0397 |
Blank canvas
Draw it
Draw what Same entropy, two units just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Section
Part 2 of 6 — extra coding cost
Intuition
Entropy is the bit-cost of the optimal code for p. But suppose you built your code for a different distribution q and then the real data comes from p. Every message still gets through — but you overpay.
That overpayment — the extra bits per symbol from using q's code on p's data — is exactly what KL divergence measures. It is zero only when q = p (you had the right codebook all along).
Concept
The Kullback–Leibler divergence from P to Q weights the per-outcome log-ratio by P:
\[ D_{KL}(P \,\|\, Q) = \sum_x P(x)\,\log_2 \frac{P(x)}{Q(x)} = \mathbb{E}_{x\sim P}\!\left[\log_2 \frac{P(x)}{Q(x)}\right] \]
Read it as: under the true distribution P, how much bigger is −log Q (what you pay) than −log P (what you'd pay optimally), on average.
Two hard facts we'll prove: D_KL ≥ 0 always (zero iff P = Q), and it is not symmetric — D_KL(P‖Q) ≠ D_KL(Q‖P) in general. It is a divergence, not a distance.
Step zero
Discussion prompt
KL of weather vs uniform, term by term — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Ratios p/q = [1.5, 0.75, 0.75]
Answer:
Worked example
Let P = p = [0.5, 0.25, 0.25] and Q = q = [⅓, ⅓, ⅓] (uniform). Build D_KL(p‖q) one outcome at a time.
Ratios p/q = [1.5, 0.75, 0.75]
Why: 0.5÷(1/3)=1.5; 0.25÷(1/3)=0.75 twice. Where p exceeds q the ratio is >1 (positive log); where p is below q it is <1 (negative log).
Weighted logs: 0.5·log₂1.5, 0.25·log₂0.75, 0.25·log₂0.75
Why: Multiply each log-ratio by P(x). Numerically these are +0.292481, −0.103759, −0.103759.
\[ D_{KL}(p\|q) = 0.5\log_2 1.5 + 0.25\log_2 0.75 + 0.25\log_2 0.75 \]
Sum = 0.0850 bits
Why: 0.292481 − 0.103759 − 0.103759 = 0.084963. The negative terms don't sink it below zero — the P-weighting guarantees a non-negative total.
\[ D_{KL}(p\|q) \approx 0.0850 \text{ bits} \]
Faded example
Fill in the blanks
KL in code — the term vector, with the scaffolding fading: two lines are gone now — fill both.
import numpy as np
p = np.array([0.5, 0.25, 0.25])
q = np.array([1/3, 1/3, 1/3]) # uniform
terms = **p * np.log2(p / q)**
print(np.round(terms, 6)) # [ 0.292481 -0.103759 -0.103759]
print(round(float(np.sum(terms)), 4)) # 0.085
Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. Sunny (p>q) pushes the sum up; cloudy and rainy (p<q) pull it down — but not far enough to go negative.
Worked example
The same computation, so you can see the individual P·log(P/Q) contributions before they're summed:
import numpy as np
p = np.array([0.5, 0.25, 0.25])
q = np.array([1/3, 1/3, 1/3]) # uniform
terms = p * np.log2(p / q)
print(np.round(terms, 6)) # [ 0.292481 -0.103759 -0.103759]
print(round(float(np.sum(terms)), 4)) # 0.085One term is positive, two are negative
Why: Sunny (p>q) pushes the sum up; cloudy and rainy (p<q) pull it down — but not far enough to go negative.
D_KL(p‖q) = 0.085 bits
Why: Exactly the hand result. Coding p's weather with a uniform codebook wastes about 0.085 bits per day.
| outcome | p·log₂(p/q) |
|---|---|
| Sunny | +0.292481 |
| Cloudy | −0.103759 |
| Rainy | −0.103759 |
| D_KL(p‖q) = Σ | 0.084963 |
Discrimination
Sort into buckets
Sort these by p·log₂(p/q), from memory, without looking back at KL in code — the term vector. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Worked example
Now compute the reverse, D_KL(q‖p), weighting by q instead of p. If KL were a distance these would be equal:
import numpy as np
p = np.array([0.5, 0.25, 0.25])
q = np.array([1/3, 1/3, 1/3])
KL = lambda a, b: float(np.sum(a * np.log2(a / b)))
print(round(KL(p, q), 4), round(KL(q, p), 4)) # 0.085 0.0817
print(np.isclose(KL(p, q), KL(q, p))) # FalseD_KL(p‖q) = 0.0850 ≠ D_KL(q‖p) = 0.0817
Why: Same two distributions, different numbers — because the weighting distribution changed. The direction of KL is part of its meaning.
np.isclose(...) prints False
Why: Numerical proof of asymmetry. Never assume you can flip the arguments.
| direction | value (bits) | weighted by |
|---|---|---|
| D_KL(p‖q) | 0.0850 | p (true = weather) |
| D_KL(q‖p) | 0.0817 | q (true = uniform) |
| equal? | no — asymmetric |
Anomaly
Predict first
A student writes this, and it looks reasonable:
KL measures how far two distributions are, so D_KL(P‖Q) = D_KL(Q‖P) and it obeys the triangle inequality — just use whichever order is handy.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: We just measured 0.0850 ≠ 0.0817 on the SAME pair — the order changed the answer.
KL is a directional divergence: D_KL(P‖Q) averages the log-ratio under P, so the first argument (the 'truth' you weight by) is load-bearing.
Why: We just measured 0.0850 ≠ 0.0817 on the SAME pair — the order changed the answer. KL also fails the triangle inequality. Flipping arguments silently changes what you optimize.
Trap
KL measures how far two distributions are, so D_KL(P‖Q) = D_KL(Q‖P) and it obeys the triangle inequality — just use whichever order is handy.
Assume symmetry, pick either order
Why: We just measured 0.0850 ≠ 0.0817 on the SAME pair — the order changed the answer. KL also fails the triangle inequality. Flipping arguments silently changes what you optimize.
KL is a directional divergence: D_KL(P‖Q) averages the log-ratio under P, so the first argument (the 'truth' you weight by) is load-bearing.
Keep the order fixed to your problem
Why: Forward KL D_KL(P‖Q) is mean-seeking (covers all of P's support); reverse KL D_KL(Q‖P) is mode-seeking (locks onto one mode). VAEs minimize the reverse KL D_KL(q‖p) — choosing the wrong direction changes the whole behavior of the model.
Break the constraint
Discussion prompt
The rule this trap just fixed:
KL is a directional divergence: D_KL(P‖Q) averages the log-ratio under P, so the first argument (the 'truth' you weight by) is load-bearing.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
We just measured 0.0850 ≠ 0.0817 on the SAME pair — the order changed the answer. KL also fails the triangle inequality. Flipping arguments silently changes what you optimize.
Section
Part 3 of 6 — Gibbs' inequality
Concept
Jensen's inequality: for a concave function f (like log), the function of the average is at least the average of the function:
\[ f\big(\mathbb{E}[Z]\big) \;\ge\; \mathbb{E}\big[f(Z)\big] \qquad (f \text{ concave}) \]
The log curve bends downward, so a chord always lies below it — averaging inside the log beats averaging outside. Equality holds only when Z is constant. This single fact forces KL ≥ 0.
Ranking
Put in order
Put the moves of Prove D_KL(P‖Q) ≥ 0 into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Write −D_KL by inverting the ratio: −Σ P log(P/Q) = Σ P log(Q/P).
Worked example
Flip the sign inside the log
Why: Write −D_KL by inverting the ratio: −Σ P log(P/Q) = Σ P log(Q/P). We'll show this is ≤ 0.
\[ -D_{KL}(P\|Q) = \sum_x P(x)\,\log \frac{Q(x)}{P(x)} = \mathbb{E}_{P}\!\left[\log \frac{Q}{P}\right] \]
Apply Jensen with f = log (concave)
Why: Move the expectation inside the log: E[log Z] ≤ log E[Z]. Here Z = Q/P under P-weighting.
\[ \mathbb{E}_{P}\!\left[\log \frac{Q}{P}\right] \;\le\; \log \mathbb{E}_{P}\!\left[\frac{Q}{P}\right] \]
The inner expectation collapses to 1
Why: E_P[Q/P] = Σ P·(Q/P) = Σ Q = 1, since Q is a probability distribution. And log 1 = 0.
\[ \log \sum_x P(x)\cdot\frac{Q(x)}{P(x)} = \log \sum_x Q(x) = \log 1 = 0 \]
Chain it: −D_KL ≤ 0, so D_KL ≥ 0
Why: We showed −D_KL(P‖Q) ≤ 0. Multiply by −1 (flip the inequality): D_KL(P‖Q) ≥ 0. Equality holds exactly when Q/P is constant, i.e. Q = P. This is Gibbs' inequality.
\[ \boxed{\,D_{KL}(P\|Q) \ge 0,\quad \text{with } = 0 \iff P = Q\,} \]
Translation
\( \log \sum_x P(x)\cdot\frac{Q(x)}{P(x)} = \log \sum_x Q(x) = \log 1 = 0 \)
Draw it
Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.
Fill the middle
Fill in the blanks
From KL ≥ 0, tested on 1000 random pairs — one line has had its right-hand side removed. Put it back.
import numpy as np
rng = np.random.default_rng(0)
KL = lambda a, b: float(np.sum(a * np.log2(a / b)))
smallest = 1e9
for _ in range(1000):
a = rng.random(4); a /= a.sum() # random distribution
b = rng.random(4); b /= b.sum()
smallest = min(smallest, KL(a, b))
print(round(smallest, 4)) # > 0 : never went negative
p = np.array([0.5, 0.25, 0.25])
print(KL(p, p)) # 0.0 : zero only when a == b
Why: b is what everything below it consumes, so the wrong expression here fails later and somewhere else. Not one of a thousand random pairs produced a negative KL — the floor at 0 holds.
Worked example
Proof is one thing; let's stress-test it. Draw 1000 random distribution pairs, take the smallest KL seen, and check KL(p‖p) = 0:
import numpy as np
rng = np.random.default_rng(0)
KL = lambda a, b: float(np.sum(a * np.log2(a / b)))
smallest = 1e9
for _ in range(1000):
a = rng.random(4); a /= a.sum() # random distribution
b = rng.random(4); b /= b.sum()
smallest = min(smallest, KL(a, b))
print(round(smallest, 4)) # > 0 : never went negative
p = np.array([0.5, 0.25, 0.25])
print(KL(p, p)) # 0.0 : zero only when a == bsmallest KL over 1000 pairs = 0.0043 > 0
Why: Not one of a thousand random pairs produced a negative KL — the floor at 0 holds. The seed rng(0) makes this reproducible.
KL(p‖p) = 0.0 exactly
Why: A distribution has zero divergence from itself — the equality case of Gibbs.
| quantity | value (verified) |
|---|---|
| min KL over 1000 random pairs | 0.0043 |
| ever negative? | no |
| KL(p‖p) | 0.0 |
Sorting
Sort into buckets
These are the pieces of Lesson 14: Information Theory, out of order. Put each one back under the part of the lesson it belongs to.
Section
Part 4 of 6 — the classifier loss
Concept
Cross-entropy is the average number of bits to encode data from P using a code built for Q — the total cost, optimal part plus waste:
\[ H(P, Q) = -\sum_x P(x)\,\log_2 Q(x) = \mathbb{E}_{x\sim P}\big[-\log_2 Q(x)\big] \]
It looks almost like entropy, but the log is of Q (your model) while the weighting is by P (the truth). That mismatch is where the extra cost lives — and it factors cleanly.
Intuition
When a network classifies an image, the label is a one-hot P (all mass on the true class) and the softmax output is your Q. The training loss −log Q(true class) is cross-entropy H(P, Q) with a one-hot P.
With one-hot labels the entropy H(P) = 0, so the identity collapses to H(P,Q) = D_KL(P‖Q) exactly — the loss you minimize is the KL from the true label distribution to your prediction. That is why 'minimize cross-entropy' and 'match the target distribution' are the same instruction.
Frameworks compute it in nats (natural log) and average over the batch, but it is the same quantity we just defined — no new idea, only new packaging.
Step zero
Discussion prompt
Derive H(P,Q) = H(P) + D_KL(P‖Q) — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Start from the definition of cross-entropy
Answer:
Worked example
Start from the definition of cross-entropy
Why: Just the formula, nothing added.
\[ H(P,Q) = -\sum_x P(x)\log_2 Q(x) \]
Insert P(x)/P(x) inside the log's argument
Why: Multiply Q by P/P — a legal '×1' — so we can split it. This is the one clever move; everything after is algebra.
\[ = -\sum_x P(x)\log_2\!\Big(Q(x)\cdot\tfrac{P(x)}{P(x)}\Big) = -\sum_x P(x)\log_2\!\Big(P(x)\cdot\tfrac{Q(x)}{P(x)}\Big) \]
Split the log of a product into a sum of logs
Why: log(P · Q/P) = log P + log(Q/P), then distribute the −Σ P over the two pieces.
\[ = -\sum_x P(x)\log_2 P(x) \;-\; \sum_x P(x)\log_2 \frac{Q(x)}{P(x)} \]
Recognize the two pieces
Why: The first sum is H(P). The second is −Σ P log(Q/P) = +Σ P log(P/Q) = D_KL(P‖Q). Done.
\[ \boxed{\,H(P,Q) = H(P) + D_{KL}(P\|Q)\,} \]
Estimation
Predict first
Compute H(P,Q) directly and compare to H(P) + D_KL(P‖Q), reusing p and the uniform q:
Commit before you compute: what does Verify the identity on the weather data come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: np.isclose(...) → True
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Cross-entropy = irreducible entropy + avoidable KL penalty, confirmed numerically.
Worked example
Compute H(P,Q) directly and compare to H(P) + D_KL(P‖Q), reusing p and the uniform q:
import numpy as np
p = np.array([0.5, 0.25, 0.25])
q = np.array([1/3, 1/3, 1/3])
H = -np.sum(p * np.log2(p)) # entropy of p
KL = np.sum(p * np.log2(p / q)) # D_KL(p||q)
CE = -np.sum(p * np.log2(q)) # cross-entropy H(p,q)
print(round(float(CE), 4)) # 1.585
print(round(float(H + KL), 4)) # 1.585
print(np.isclose(CE, H + KL)) # TrueCE = 1.585 = H(p) 1.5 + KL 0.085
Why: The identity holds to floating point. Since q is uniform, −log2 q = log2 3 for every outcome, so H(p,q) = log2 3 ≈ 1.585 exactly.
np.isclose(...) → True
Why: Cross-entropy = irreducible entropy + avoidable KL penalty, confirmed numerically.
| term | value (bits) |
|---|---|
| H(p) (entropy of truth) | 1.5000 |
| D_KL(p‖q) (penalty) | 0.0850 |
| H(p,q) (cross-entropy) | 1.5850 |
Anomaly
Predict first
A student writes this, and it looks reasonable:
Training minimizes cross-entropy H(P,Q), so with enough training the loss should be drivable all the way to 0.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: H(P,Q) = H(P) + D_KL(P‖Q). H(P) is fixed by the data's own label noise — you can't optimize it away.
Because H(P) is a constant w.r.t. your model Q, minimizing cross-entropy is exactly minimizing D_KL(P‖Q).
Why: H(P,Q) = H(P) + D_KL(P‖Q). H(P) is fixed by the data's own label noise — you can't optimize it away. The loss floor is H(P), which is 1.5 bits here, not 0. Chasing 0 means overfitting the noise.
Trap
Training minimizes cross-entropy H(P,Q), so with enough training the loss should be drivable all the way to 0.
Expect loss → 0 on noisy labels
Why: H(P,Q) = H(P) + D_KL(P‖Q). H(P) is fixed by the data's own label noise — you can't optimize it away. The loss floor is H(P), which is 1.5 bits here, not 0. Chasing 0 means overfitting the noise.
Because H(P) is a constant w.r.t. your model Q, minimizing cross-entropy is exactly minimizing D_KL(P‖Q).
Training pushes Q → P; the floor is H(P)
Why: ∂H(P,Q)/∂Q only touches the KL term. So fitting a classifier by cross-entropy IS distribution matching — driving your predicted distribution Q toward the true label distribution P. The best achievable loss equals the label entropy H(P).
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
0.5, Cloudy 0.25, Rainy 0.25. Call it p.; The surprise (self-information) of an outcome with probability p is defined as:; The most uncertain distribution over n outcomes is the uniform one — nothing is predictable — and its entropy is the ceiling log₂ n.D_KL(P‖Q) = D_KL(Q‖P) and it obeys the triangle inequality — just use whichever order is handy.; Training minimizes cross-entropy H(P,Q), so with enough training the loss should be drivable all the way to 0.Section
Part 5 of 6 — shared bits
Concept
Mutual information needs two variables. Take a 2×2 joint P(X, Y) — say X = was it cloudy, Y = did it rain — with these joint probabilities:
| P(X,Y) | Y = 0 | Y = 1 | row sum p(x) |
|---|---|---|---|
| X = 0 | 0.30 | 0.20 | 0.50 |
| X = 1 | 0.10 | 0.40 | 0.50 |
| col sum p(y) | 0.40 | 0.60 | 1.00 |
The margins are the row/column sums: p(x) = [0.5, 0.5], p(y) = [0.4, 0.6]. If X and Y were independent, the joint would equal the outer product p(x)p(y) — it does not, so they share information.
Pattern
Step through it
Step through A joint distribution to work with one row at a time. What is driving the change, and what would the row after the last one be?
Picture it
Figure (svg): Two overlapping circles labeled H(X) and H(Y). The left-only crescent is H(X|Y), the right-only crescent is H(Y|X), the lens-shaped overlap in the middle is the mutual information I(X;Y), and the whole union is the joint entropy H(X,Y).
Discussion prompt
Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.
Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.
Answer:
Picture the uncertainty in X as one circle of area H(X) and the uncertainty in Y as another of area H(Y). When the variables are related, the circles overlap.
Intuition
Picture the uncertainty in X as one circle of area H(X) and the uncertainty in Y as another of area H(Y). When the variables are related, the circles overlap.
Figure (svg): Two overlapping circles labeled H(X) and H(Y). The left-only crescent is H(X|Y), the right-only crescent is H(Y|X), the lens-shaped overlap in the middle is the mutual information I(X;Y), and the whole union is the joint entropy H(X,Y).
The overlap is the mutual information — the uncertainty they share. The full union is the joint entropy H(X,Y), and each crescent is a conditional entropy (H(X|Y), H(Y|X)). Pull the circles apart until they just touch and the overlap — the MI — drops to zero: that is independence.
Concept
Mutual information I(X;Y) is how many bits knowing one variable saves you on the other. It is a KL divergence — the joint from the independence model:
\[ I(X;Y) = \sum_{x,y} P(x,y)\,\log_2 \frac{P(x,y)}{p(x)\,p(y)} = D_{KL}\big(P(x,y)\,\|\,p(x)p(y)\big) \]
Equivalently, in entropies — the shared area of two overlapping circles:
\[ I(X;Y) = H(X) + H(Y) - H(X,Y) \]
Because it is a KL, I(X;Y) ≥ 0, and it is 0 exactly when P(x,y) = p(x)p(y) — i.e. when X and Y are independent.
Faded example
Fill in the blanks
MI from the joint — the KL form, with the scaffolding fading: two lines are gone now — fill both.
import numpy as np
Pxy = np.array([[0.30, 0.20], # rows = X, cols = Y
[0.10, 0.40]])
px = Pxy.sum(axis=1) # marginal p(x)
py = Pxy.sum(axis=0) # marginal p(y)
MI = np.sum(Pxy * np.log2(Pxy / np.outer(px, py)))
print(px, py) # [0.5 0.5] [0.4 0.6]
print(round(float(MI), 4)) # 0.1245
Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. Row sums give p(x); column sums give p(y).
Worked example
Compute the margins from the table, then the joint-vs-independent KL directly:
import numpy as np
Pxy = np.array([[0.30, 0.20], # rows = X, cols = Y
[0.10, 0.40]])
px = Pxy.sum(axis=1) # marginal p(x)
py = Pxy.sum(axis=0) # marginal p(y)
MI = np.sum(Pxy * np.log2(Pxy / np.outer(px, py)))
print(px, py) # [0.5 0.5] [0.4 0.6]
print(round(float(MI), 4)) # 0.1245Margins come out [0.5, 0.5] and [0.4, 0.6]
Why: Row sums give p(x); column sums give p(y). np.outer(px,py) builds the independence model p(x)p(y).
I(X;Y) = 0.1245 bits
Why: The joint diverges from independence by 0.1245 bits — that's how much X and Y share.
| cell (x,y) | P(x,y) | p(x)p(y) | P·log₂(P / pxpy) |
|---|---|---|---|
| (0,0) | 0.30 | 0.20 | +0.175489 |
| (0,1) | 0.20 | 0.30 | −0.116993 |
| (1,0) | 0.10 | 0.20 | −0.100000 |
| (1,1) | 0.40 | 0.30 | +0.166015 |
| I(X;Y) = Σ | 0.124511 |
Worked example
The entropy form must give the identical number. Compute the three entropies and combine:
import numpy as np
Pxy = np.array([[0.30, 0.20], [0.10, 0.40]])
H = lambda d: -np.sum(d[d > 0] * np.log2(d[d > 0]))
Hx, Hy, Hxy = H(Pxy.sum(1)), H(Pxy.sum(0)), H(Pxy.ravel())
print(round(float(Hx), 4), round(float(Hy), 4), round(float(Hxy), 4))
print(round(float(Hx + Hy - Hxy), 4)) # 0.1245H(X)=1.0, H(Y)=0.9710, H(X,Y)=1.8464
Why: p(x) is uniform so H(X)=1; p(y)=[0.4,0.6] gives 0.9710; the flattened joint gives 1.8464.
H(X)+H(Y)−H(X,Y) = 0.1245 — matches the KL form
Why: Both routes to mutual information agree to 4 decimals. The joint entropy is LESS than H(X)+H(Y); the shortfall is exactly the shared information.
| quantity | value (bits) |
|---|---|
| H(X) | 1.0000 |
| H(Y) | 0.9710 |
| H(X,Y) | 1.8464 |
| H(X)+H(Y)−H(X,Y) | 0.1245 |
Pattern
Predict first
The table runs: H(Y) | 0.9710 · H(Y|X) | 0.8464
In The third route: I(X;Y) = H(Y) − H(Y|X), given the rows so far: what is the next one — the row where quantity is I(X;Y) = H(Y) − H(Y|X)?
Correct: I(X;Y) = H(Y) − H(Y|X) | 0.1245
| quantity | value (bits) |
|---|---|
| H(Y) | 0.9710 |
| H(Y|X) | 0.8464 |
| I(X;Y) = H(Y) − H(Y|X) | 0.1245 |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. Each row Pxy[i]/px[i] is the conditional p(Y|X=x); average their entropies weighted by p(x).
Worked example
The overlap picture says MI is also 'uncertainty in Y minus leftover uncertainty in Y once you know X'. The conditional entropy H(Y|X) is the right crescent. Build it and subtract:
import numpy as np
Pxy = np.array([[0.30, 0.20], [0.10, 0.40]])
H = lambda d: -np.sum(d[d > 0] * np.log2(d[d > 0]))
px = Pxy.sum(axis=1)
HYgivenX = sum(px[i] * H(Pxy[i] / px[i]) for i in range(2)) # Σ p(x) H(Y|X=x)
HY = H(Pxy.sum(axis=0))
print(round(float(HY), 4), round(float(HYgivenX), 4)) # 0.971 0.8464
print(round(float(HY - HYgivenX), 4)) # 0.1245H(Y) = 0.9710, H(Y|X) = 0.8464
Why: Each row Pxy[i]/px[i] is the conditional p(Y|X=x); average their entropies weighted by p(x). Knowing X shaves H(Y) from 0.971 down to 0.846.
I(X;Y) = H(Y) − H(Y|X) = 0.1245
Why: The reduction in Y's uncertainty from learning X — identical to both the KL form and H(X)+H(Y)−H(X,Y). Three routes, one number.
| quantity | value (bits) |
|---|---|
| H(Y) | 0.9710 |
| H(Y|X) | 0.8464 |
| I(X;Y) = H(Y) − H(Y|X) | 0.1245 |
Comparison
Comparison matrix
From The third route: I(X;Y) = H(Y) − H(Y|X): refill the value (bits) column from what you know. The rest of the table is as it appeared.
| quantity | value (bits) |
|---|---|
| H(Y) | 0.9710 |
| H(Y|X) | 0.8464 |
| I(X;Y) = H(Y) − H(Y|X) | 0.1245 |
Fill the middle
Fill in the blanks
From Independence forces I = 0 — one line has had its right-hand side removed. Put it back.
import numpy as np
# independent: joint = outer product of marginals -> I = 0
px = np.array([0.5, 0.5]); py = np.array([0.4, 0.6])
Pind = np.outer(px, py)
MI = np.sum(Pind * np.log2(Pind / np.outer(Pind.sum(1), Pind.sum(0))))
print(Pind)
print(round(float(MI), 6)) # 0.0
Why: Pind is what everything below it consumes, so the wrong expression here fails later and somewhere else. By construction the joint is the independence model, so the log-ratio inside the sum is log(1) = 0 in every cell.
Worked example
Build a joint that IS the outer product of the margins — genuine independence — and watch MI vanish:
import numpy as np
# independent: joint = outer product of marginals -> I = 0
px = np.array([0.5, 0.5]); py = np.array([0.4, 0.6])
Pind = np.outer(px, py)
MI = np.sum(Pind * np.log2(Pind / np.outer(Pind.sum(1), Pind.sum(0))))
print(Pind)
print(round(float(MI), 6)) # 0.0Pind exactly equals p(x)p(y)
Why: By construction the joint is the independence model, so the log-ratio inside the sum is log(1) = 0 in every cell.
I(X;Y) = 0.0
Why: Knowing one variable tells you nothing about the other. MI = 0 ⇔ independence — the cleanest test for dependence in all of statistics.
| joint used | independent? | I(X;Y) |
|---|---|---|
| [[0.3,0.2],[0.1,0.4]] | no | 0.1245 |
| outer(p(x), p(y)) | yes | 0.0 |
Blank canvas
Draw it
Draw what Independence forces I = 0 just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Section
Part 6 of 6 — the VAE objective
Intuition
Latent-variable models want the evidence log p(x) = log Σ_z p(x, z). But that sum runs over every possible latent z — for a real VAE, an integral over a continuous high-dimensional space that has no closed form.
So we introduce an easy, tunable approximation q(z|x) to the true-but-intractable posterior p(z|x), and instead of computing log p(x) we lower-bound it with a quantity we can compute and maximize. That bound is the ELBO.
Concept
The log-evidence splits exactly into the ELBO plus a KL gap to the true posterior — an identity, not an approximation:
\[ \log p(x) = \underbrace{\mathbb{E}_{q(z|x)}\!\left[\log \frac{p(x,z)}{q(z|x)}\right]}_{\text{ELBO}(q)} \;+\; D_{KL}\big(q(z|x)\,\|\,p(z|x)\big) \]
Since D_KL ≥ 0 (Part 3), the ELBO can never exceed log p(x) — it is a genuine lower bound. And the only slack is that KL: the tighter q matches the true posterior, the closer the bound.
Ranking
Put in order
Put the moves of Derive the split in three lines into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. log p(x) doesn't depend on z, so its expectation under any q(z|x) is itself.
Worked example
Write log p(x) as an expectation over q
Why: log p(x) doesn't depend on z, so its expectation under any q(z|x) is itself. Free move.
\[ \log p(x) = \mathbb{E}_{q(z|x)}\big[\log p(x)\big] \]
Expand p(x) via Bayes and multiply by q/q
Why: p(x) = p(x,z)/p(z|x). Insert q(z|x)/q(z|x) inside the log so both a p(x,z)/q and a q/p(z|x) piece appear.
\[ = \mathbb{E}_{q}\!\left[\log \frac{p(x,z)}{q(z|x)} + \log \frac{q(z|x)}{p(z|x)}\right] \]
Split the expectation into the two named pieces
Why: The first expectation is the ELBO; the second is E_q[log(q/p(z|x))] = D_KL(q‖p(z|x)). That's the boxed identity.
\[ = \text{ELBO}(q) + D_{KL}\big(q(z|x)\,\|\,p(z|x)\big) \]
Bound follows because KL ≥ 0
Why: Drop the non-negative KL term to get log p(x) ≥ ELBO(q). Maximizing the ELBO both raises the evidence bound AND shrinks the KL, pulling q toward the true posterior — the entire VAE training objective.
\[ \boxed{\,\log p(x) \ge \text{ELBO}(q)\,} \]
Notation
Annotate
From Derive the split in three lines — read this one piece at a time. What is each part doing?
On: \( \boxed{\,\log p(x) \ge \text{ELBO}(q)\,} \)
Missing information
Discussion prompt
A tiny discrete latent z (4 values) makes the identity concrete. Fix x, give a joint p(x,z), a candidate q, and check that ELBO + KL = log p(x) and ELBO ≤ log p(x):
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
The bound holds: a mismatched q gives an ELBO strictly less than the true log-evidence.
Worked example
A tiny discrete latent z (4 values) makes the identity concrete. Fix x, give a joint p(x,z), a candidate q, and check that ELBO + KL = log p(x) and ELBO ≤ log p(x):
import numpy as np
pxz = np.array([0.1, 0.2, 0.3, 0.05]) # joint p(x, z) over 4 z-values (fixed x)
px = pxz.sum() # marginal p(x)
post = pxz / px # true posterior p(z|x)
q = np.array([0.4, 0.3, 0.2, 0.1]) # approximate posterior
logpx = np.log(px)
elbo = np.sum(q * np.log(pxz / q)) # E_q[log p(x,z) - log q]
klqp = np.sum(q * np.log(q / post)) # D_KL(q || p(z|x))
print(round(float(logpx), 4)) # -0.4308
print(round(float(elbo), 4)) # -0.6644
print(round(float(elbo + klqp), 4)) # -0.4308 (== log p(x))
print(elbo <= logpx) # TrueELBO = −0.6644 sits below log p(x) = −0.4308
Why: The bound holds: a mismatched q gives an ELBO strictly less than the true log-evidence.
ELBO + D_KL(q‖p) = −0.4308 = log p(x)
Why: The gap between the ELBO and the evidence is EXACTLY the posterior KL — the identity is confirmed to 4 decimals. Improve q → shrink KL → tighten the bound.
| quantity | value (nats, verified) |
|---|---|
| log p(x) | −0.4308 |
| ELBO(q) | −0.6644 |
| D_KL(q‖p(z|x)) | 0.2336 |
| ELBO + KL | −0.4308 |
| ELBO ≤ log p(x)? | True |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
ELBO + D_KL(q‖p) = −0.4308 = log p(x)
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
A tiny discrete latent z (4 values) makes the identity concrete. Fix x, give a joint p(x,z), a candidate q, and check that ELBO + KL = log p(x) and ELBO ≤ log p(x):
Concept
Real VAEs use Gaussian q and prior, so the KL term has a formula — no sampling needed. For two 1-D Gaussians (in nats, natural log):
\[ D_{KL}\big(\mathcal{N}(\mu_1,\sigma_1^2)\,\|\,\mathcal{N}(\mu_2,\sigma_2^2)\big) = \log\frac{\sigma_2}{\sigma_1} + \frac{\sigma_1^2 + (\mu_1-\mu_2)^2}{2\sigma_2^2} - \frac12 \]
The VAE special case matches q = N(μ, σ²) to a standard-normal prior N(0,1), which collapses to ½(σ² + μ² − 1 − log σ²) — the KL regularizer in the VAE loss.
Explain it
Discussion prompt
Explain The Gaussian-KL closed form to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Real VAEs use Gaussian q and prior, so the KL term has a formula — no sampling needed. For two 1-D Gaussians (in nats, natural log):
Estimation
Predict first
Compute D_KL(N(0,1) ‖ N(1, var=2)) from the closed form, then confirm it by sampling log(p/q) under p:
Commit before you compute: what does Gaussian-KL: formula and Monte-Carlo agree come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Monte-Carlo = 0.3462, matches to 3 decimals
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Averaging log(p/q) over 2 million samples from p reproduces the analytic KL, validating the closed form.
Worked example
Compute D_KL(N(0,1) ‖ N(1, var=2)) from the closed form, then confirm it by sampling log(p/q) under p:
import numpy as np, math
def gauss_kl(m1, s1, m2, s2):
return math.log(s2/s1) + (s1**2 + (m1-m2)**2) / (2*s2**2) - 0.5
closed = gauss_kl(0, 1, 1, math.sqrt(2)) # N(0,1) || N(1, var=2)
rng = np.random.default_rng(0)
xs = rng.normal(0, 1, 2_000_000) # samples from p = N(0,1)
logp = -0.5*np.log(2*np.pi*1) - (xs-0)**2/(2*1)
logq = -0.5*np.log(2*np.pi*2) - (xs-1)**2/(2*2)
print(round(closed, 4)) # 0.3466 nats
print(round(float(np.mean(logp - logq)), 4)) # 0.3462 (MC)Closed form = 0.3466 nats
Why: Plug μ1=0, σ1=1, μ2=1, σ2²=2 into the formula. Because units are nats (natural log), not bits.
Monte-Carlo = 0.3462, matches to 3 decimals
Why: Averaging log(p/q) over 2 million samples from p reproduces the analytic KL, validating the closed form. The tiny gap is sampling noise.
| method | D_KL (nats, verified) |
|---|---|
| closed form | 0.3466 |
| Monte-Carlo (2e6 samples) | 0.3462 |
| agree? | yes (within noise) |
Error analysis
Annotate
Walk the callouts on Gaussian-KL: formula and Monte-Carlo agree. Each one is a place this is easy to get subtly wrong.
Concept
Step back: almost every quantity today was KL divergence in disguise. Recognizing that turns five formulas into one idea — the gap between what's true and what your model assumes.
| appears as | the KL underneath |
|---|---|
| compression waste | D_KL(P‖Q) = extra bits from the wrong codebook |
| cross-entropy loss | H(P,Q) − H(P) = D_KL(P‖Q) |
| mutual information | D_KL(P(x,y) ‖ p(x)p(y)) |
| ELBO slack | log p(x) − ELBO = D_KL(q ‖ p(z|x)) |
The two facts KL ≥ 0 and KL asymmetric therefore propagate everywhere: every loss has a floor, every bound is one-sided, and direction always matters.
Analogy
Discussion prompt
Explain One measure, wearing five hats by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Step back: almost every quantity today was KL divergence in disguise. Recognizing that turns five formulas into one idea — the gap between what's true and what your model assumes.
Ranking
Put in order
These are the steps of The information-theory toolkit, scrambled. Put them back in order before the next slide shows you.
−log₂ p — decreasing in p, adds for independent events; bits (log₂) or nats (ln)H(X) = −Σ p log₂ p — average surprise, in [0, log₂ n], max at uniformD_KL(P‖Q) = Σ P log₂(P/Q) — extra coding cost; ≥ 0 (Gibbs/Jensen), asymmetric, not a metricH(P,Q) = H(P) + D_KL(P‖Q) — the classifier loss; minimizing it minimizes KLI(X;Y) = H(X)+H(Y)−H(X,Y) = D_KL(P(x,y)‖p(x)p(y)) — 0 iff independentlog p(x) = ELBO + D_KL(q‖p) ≥ ELBO — maximize the bound; Gaussian-KL has a closed form (VAE)Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
−log₂ p — decreasing in p, adds for independent events; bits (log₂) or nats (ln)H(X) = −Σ p log₂ p — average surprise, in [0, log₂ n], max at uniformD_KL(P‖Q) = Σ P log₂(P/Q) — extra coding cost; ≥ 0 (Gibbs/Jensen), asymmetric, not a metricH(P,Q) = H(P) + D_KL(P‖Q) — the classifier loss; minimizing it minimizes KLI(X;Y) = H(X)+H(Y)−H(X,Y) = D_KL(P(x,y)‖p(x)p(y)) — 0 iff independentlog p(x) = ELBO + D_KL(q‖p) ≥ ELBO — maximize the bound; Gaussian-KL has a closed form (VAE)Edge cases
Discussion prompt
The information-theory toolkit works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
−log₂ p — decreasing in p, adds for independent events; bits (log₂) or nats (ln)H(X) = −Σ p log₂ p — average surprise, in [0, log₂ n], max at uniformD_KL(P‖Q) = Σ P log₂(P/Q) — extra coding cost; ≥ 0 (Gibbs/Jensen), asymmetric, not a metricH(P,Q) = H(P) + D_KL(P‖Q) — the classifier loss; minimizing it minimizes KLI(X;Y) = H(X)+H(Y)−H(X,Y) = D_KL(P(x,y)‖p(x)p(y)) — 0 iff independentlog p(x) = ELBO + D_KL(q‖p) ≥ ELBO — maximize the bound; Gaussian-KL has a closed form (VAE)Elimination
Eliminate the wrong options
Over 4 outcomes, which distribution has the HIGHEST entropy?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Entropy is maximized by the uniform distribution. With 4 equal outcomes, H = log₂4 = 2 bits — the ceiling for four outcomes, hence the most uncertain.
Check
Which distribution is the most uncertain?
Check your understanding
Over 4 outcomes, which distribution has the HIGHEST entropy?
Answer: A
Why: Entropy is maximized by the uniform distribution. With 4 equal outcomes, H = log₂4 = 2 bits — the ceiling for four outcomes, hence the most uncertain.
Check
What is actually true of D_KL(P‖Q)?
Check your understanding
Which statement about D_KL(P‖Q) is correct?
Answer: A
Why: By Gibbs/Jensen, KL is non-negative and zero only when P = Q, and it is directional — we measured 0.0850 vs 0.0817 on the same pair. It is a divergence, not a metric.
Prediction
Predict first
Minimizing the cross-entropy H(P,Q) over the model Q is equivalent to minimizing…
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: D_KL(P‖Q), because H(P) is constant in Q
Why: H(P,Q) = H(P) + D_KL(P‖Q). H(P) depends only on the fixed data, so optimizing Q can only shrink the KL term — training a classifier is distribution matching, and the loss floor is H(P), not 0.
Check
Decompose the loss before you answer.
Check your understanding
Minimizing the cross-entropy H(P,Q) over the model Q is equivalent to minimizing…
Answer: A
Why: H(P,Q) = H(P) + D_KL(P‖Q). H(P) depends only on the fixed data, so optimizing Q can only shrink the KL term — training a classifier is distribution matching, and the loss floor is H(P), not 0.
Elimination
Eliminate the wrong options
For two random variables, I(X;Y) = 0 exactly when…
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: I(X;Y) = D_KL(P(x,y) ‖ p(x)p(y)), and a KL is 0 iff its two arguments are equal. So I = 0 exactly when the joint equals the product of marginals — the definition of independence.
Check
Think about the KL form of I(X;Y).
Check your understanding
For two random variables, I(X;Y) = 0 exactly when…
Answer: A
Why: I(X;Y) = D_KL(P(x,y) ‖ p(x)p(y)), and a KL is 0 iff its two arguments are equal. So I = 0 exactly when the joint equals the product of marginals — the definition of independence.
Section
The project
Concept
Implement the three core quantities in NumPy on the weather p = [0.5, 0.25, 0.25] and uniform q, then verify the identity H(P,Q) = H(P) + D_KL(P‖Q) numerically. You derived every piece — now assemble it.
| # | requirement | tool |
|---|---|---|
| 1 | entropy(p) = −Σ p log₂ p | np.log2 |
| 2 | kl(p,q) = Σ p log₂(p/q), both directions | np.sum |
| 3 | Verify cross(p,q) == entropy(p) + kl(p,q) | np.isclose |
Build rules: type every line yourself, use log2 so answers are in bits, and confirm KL comes out asymmetric before you trust the code.
Counterexample
Discussion prompt
Build rules: type every line yourself, use log2 so answers are in bits, and confirm KL comes out asymmetric before you trust the code.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Worked example
Your turn: write entropy(p). Predict entropy([0.5, 0.25, 0.25]) out loud before you run it — you computed it by hand as 1.5.
Hint: -np.sum(p*np.log2(p)). A fair coin [0.5, 0.5] must give exactly 1.0 bit as a sanity check.
import numpy as np
def entropy(p):
p = np.asarray(p, dtype=float)
return -np.sum(p * np.log2(p))
print(round(float(entropy([0.5, 0.25, 0.25])), 4)) # 1.5
print(round(float(entropy([0.5, 0.5])), 4)) # 1.0| input | entropy (bits) |
|---|---|
| [0.5, 0.25, 0.25] | 1.5 |
| [0.5, 0.5] | 1.0 |
| [1.0, 0.0]* | 0.0 (with 0·log0 = 0) |
Worked example
Your turn: write kl(p, q) and compute both directions for p = [0.5, 0.25, 0.25], q = [⅓, ⅓, ⅓]. Predict: equal or not?
Hint: np.sum(p*np.log2(p/q)); swap the arguments for the reverse. They should come out 0.085 and 0.0817 — different.
import numpy as np
def kl(p, q):
p = np.asarray(p, dtype=float); q = np.asarray(q, dtype=float)
return np.sum(p * np.log2(p / q))
p = [0.5, 0.25, 0.25]; q = [1/3, 1/3, 1/3]
print(round(float(kl(p, q)), 4), round(float(kl(q, p)), 4)) # 0.085 0.0817| direction | value (bits) |
|---|---|
| kl(p, q) | 0.085 |
| kl(q, p) | 0.0817 |
| equal? | no — asymmetric |
Worked example
Your turn: write cross(p, q) directly and confirm it equals entropy(p) + kl(p, q) with np.isclose. Predict the value — since q is uniform it should be log₂ 3 ≈ 1.585.
Hint: cross = -np.sum(p*np.log2(q)); compare against entropy(p) + kl(p, q).
import numpy as np
def entropy(p): p = np.asarray(p, float); return -np.sum(p * np.log2(p))
def kl(p, q): p, q = np.asarray(p, float), np.asarray(q, float); return np.sum(p * np.log2(p / q))
def cross(p, q):p, q = np.asarray(p, float), np.asarray(q, float); return -np.sum(p * np.log2(q))
p = [0.5, 0.25, 0.25]; q = [1/3, 1/3, 1/3]
print(round(float(cross(p, q)), 4)) # 1.585
print(round(float(entropy(p) + kl(p, q)), 4)) # 1.585
print(np.isclose(cross(p, q), entropy(p) + kl(p, q))) # True| quantity | value (bits) |
|---|---|
| cross(p, q) | 1.585 |
| entropy(p) + kl(p, q) | 1.585 |
| identity holds (isclose) | True |
Trade off
Comparison matrix
From Milestone 3 — the identity: every row here is a choice with a cost. Fill the value (bits) column, then say which row you would actually pick and what you give up for it.
| quantity | value (bits) |
|---|---|
| cross(p, q) | 1.585 |
| entropy(p) + kl(p, q) | 1.585 |
| identity holds (isclose) | True |
Concept
import numpy as np
def entropy(p): p = np.asarray(p, float); return -np.sum(p * np.log2(p))
def kl(p, q): p, q = np.asarray(p, float), np.asarray(q, float); return np.sum(p * np.log2(p / q))
def cross(p, q): p, q = np.asarray(p, float), np.asarray(q, float); return -np.sum(p * np.log2(q))
p = np.array([0.5, 0.25, 0.25]) # weather: Sunny, Cloudy, Rainy
q = np.array([1/3, 1/3, 1/3]) # uniform guess
print('H(p) =', round(float(entropy(p)), 4))
print('KL(p||q) =', round(float(kl(p, q)), 4), ' KL(q||p) =', round(float(kl(q, p)), 4))
print('H(p,q) =', round(float(cross(p, q)), 4),
' = H(p)+KL =', round(float(entropy(p) + kl(p, q)), 4))| printed line | value |
|---|---|
| H(p) | 1.5 |
| KL(p||q) / KL(q||p) | 0.085 / 0.0817 |
| H(p,q) = H(p)+KL | 1.585 |
If H(p) reads 1.5, KL comes out asymmetric (0.085 ≠ 0.0817), and cross-entropy equals H(p)+KL = 1.585 — you've built the information-theory core of every classifier and VAE from the axioms up.
Comparison
Comparison matrix
From The full program: refill the value column from what you know. The rest of the table is as it appeared.
| printed line | value |
|---|---|
| H(p) | 1.5 |
| KL(p||q) / KL(q||p) | 0.085 / 0.0817 |
| H(p,q) = H(p)+KL | 1.585 |
Concept
Out loud, slides closed: explain (1) why surprise must be −log p, (2) why entropy is maximized by the uniform, (3) why minimizing cross-entropy equals minimizing KL, and (4) why the ELBO is a lower bound on log p(x).
Stretch (homework): re-derive D_KL ≥ 0 from Jensen without looking, then use the Gaussian-KL closed form to compute the VAE regularizer D_KL(N(μ, σ²) ‖ N(0,1)) = ½(σ² + μ² − 1 − ln σ²) — verified = 0.1366 nats at μ = 0.5, σ² = 0.8. This thread runs straight into the VAE (Week 44) and diffusion (Week 48).
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Surprise & entropy · KL divergence · Why KL ≥ 0 · Cross-entropy = H(P) + KL · Mutual information · The ELBO. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
−log₂ p from one axiom and compute entropy as average surprise — H(p) = 1.5 bits on the weather data, max at uniform log₂ nD_KL(P‖Q) = Σ P log₂(P/Q), prove ≥ 0 via Jensen, and show it is asymmetric (0.085 ≠ 0.0817)H(P) + D_KL(P‖Q) step by step, so minimizing the loss = minimizing KL, with floor H(P)0.1245 bits) and show I = 0 iff independentlog p(x) = ELBO + D_KL(q‖p) ≥ ELBO and use the Gaussian-KL closed form| idea | the one thing to remember |
|---|---|
| entropy | −Σ p log p; average surprise; max at uniform |
| KL | ≥ 0 (Jensen), asymmetric, NOT a metric |
| cross-entropy | H(P) + KL → minimizing CE = minimizing KL |
| mutual information | D_KL(joint ‖ product); 0 iff independent |
| ELBO | log p(x) = ELBO + KL(q‖p) ≥ ELBO |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.