Lesson 14: Information Theory

USAAIO Lesson 14, from Week 5 on probability, fully worked. It builds surprise and entropy from a single axiom and computes the entropy of a running three-outcome weather distribution bit by bit. It then derives KL divergence as an extra coding cost and proves it asymmetric and non-negative, splits cross-entropy into H(P) + KL one algebra step at a time, derives mutual information from the joint table, and gets to the ELBO decomposition that underpins the VAE. Along the way it covers units, bits against nats, conditional entropy, and the closed form for the Gaussian KL. Every snippet runs standalone, and every number was produced by real execution. The lesson runs to 60 slides.

Subject: Machine Learning · 112 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Information Theory

Title

USAAIO · Lesson 14 · Week 5 (Probability)

The math of uncertainty and the distance between distributions. We build surprise from one rule, turn it into entropy, derive KL divergence and prove it is asymmetric and ≥ 0, split cross-entropy into H(P) + KL, read mutual information off a joint table, and land on the ELBO that trains every VAE — no step skipped, every number run.

2. By the end of this lesson you can

Objectives

  1. Define surprise −log p from one axiom and compute entropy H(X) = −Σ p log₂ p outcome by outcome
  2. Derive KL divergence D_KL(P‖Q) = Σ P log₂(P/Q) as extra coding cost, and prove it is ≥ 0 and asymmetric
  3. Split cross-entropy as H(P) + D_KL(P‖Q) step by step, and say exactly what minimizing it does
  4. Compute mutual information I(X;Y) from a joint table two ways and show it is 0 iff independent
  5. State the ELBO decomposition log p(x) = ELBO + D_KL(q‖p), explain why it is a lower bound, and use the Gaussian-KL closed form
  6. Implement entropy, KL, and cross-entropy in NumPy and verify every identity numerically

3. What survived from Positive (Semi)Definite Matrices?

Warm-up

Discussion prompt

Before we open Lesson 14: Information Theory: without looking back, what was the main idea of Positive (Semi)Definite Matrices, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

PSD vs PD via xᵀAx and eigenvalues, the Cholesky decomposition, why XᵀX and covariance are always PSD, and three ways to prove definiteness (definition, eigenvalues, Sylvester). Build a Cholesky-based PD checker from scratch.

4. Surprise & entropy

Section

Part 1 of 6 — the one axiom

5. The running example: tomorrow's weather

Concept

One distribution carries this whole lesson. A weather model says tomorrow is Sunny with probability 0.5, Cloudy 0.25, Rainy 0.25. Call it p.

outcomesymbolp(x)
SunnyS0.5
CloudyC0.25
RainyR0.25

Every quantity today — entropy, KL, cross-entropy, mutual information — will be computed against this p, so the numbers compound instead of resetting.

6. Fill in: symbol for The running example: tomorrow's weather

Comparison

Comparison matrix

From The running example: tomorrow's weather: refill the symbol column from what you know. The rest of the table is as it appeared.

outcomesymbolp(x)
SunnyS0.5
CloudyC0.25
RainyR0.25

7. Rare events carry more information

Intuition

If a friend texts 'it's sunny in the desert', you learn almost nothing — you expected that. If they text 'it's snowing in the desert', you learn a lot. Rarer outcomes are more informative when they happen.

So information should be a decreasing function of probability: small p → big information, p = 1 → zero information (a sure thing tells you nothing).

We also want information to add for independent events: learning two independent facts should cost the sum of their individual costs. Those two wishes pin down the formula exactly.

8. Break it if you can: Rare events carry more information

Counterexample

Discussion prompt

We also want information to add for independent events: learning two independent facts should cost the sum of their individual costs. Those two wishes pin down the formula exactly.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

9. Surprise = −log p

Concept

The surprise (self-information) of an outcome with probability p is defined as:

\[ I(x) = -\log_2 p(x) = \log_2 \frac{1}{p(x)} \]

The log is the only function that turns 'independent probabilities multiply' into 'surprises add'. Base 2 gives units of bits; natural log would give nats.

Check the two wishes

Why: As p→1, −log p → 0 (a certain event is unsurprising). As p→0, −log p → ∞ (a near-impossible event is hugely informative). And −log(p·q) = −log p − log q, so independent surprises add. Both design goals are satisfied.

10. By analogy: Surprise = −log p

Analogy

Discussion prompt

Explain Surprise = −log p by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

The surprise (self-information) of an outcome with probability p is defined as:

11. Guess the shape of the answer: Surprise of each weather outcome

Estimation

Predict first

Plug each p(x) into −log₂ p. Powers of two make these exact — no calculator needed:

Commit before you compute: what does Surprise of each weather outcome come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Rainy is twice as surprising as Sunny

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Sunny (half the time) costs 1 bit; each quarter-probability outcome costs 2 bits.

12. Surprise of each weather outcome

Worked example

Plug each p(x) into −log₂ p. Powers of two make these exact — no calculator needed:

\[ \begin{aligned} I(S) &= -\log_2 0.5 = -\log_2 2^{-1} = 1 \\ I(C) &= -\log_2 0.25 = -\log_2 2^{-2} = 2 \\ I(R) &= -\log_2 0.25 = 2 \end{aligned} \]

outcomep(x)surprise −log₂p (bits)
Sunny0.51
Cloudy0.252
Rainy0.252

Rainy is twice as surprising as Sunny

Why: Sunny (half the time) costs 1 bit; each quarter-probability outcome costs 2 bits. Halving the probability adds exactly 1 bit of surprise — the signature of a log scale.

13. What each one costs: Surprise of each weather outcome

Trade off

Comparison matrix

From Surprise of each weather outcome: every row here is a choice with a cost. Fill the p(x) column, then say which row you would actually pick and what you give up for it.

outcomep(x)surprise −log₂p (bits)
Sunny0.51
Cloudy0.252
Rainy0.252

14. Entropy = average surprise

Concept

Entropy is the surprise you expect on average — weight each outcome's surprise by how often it happens:

\[ H(X) = \mathbb{E}_{x \sim p}\big[-\log_2 p(x)\big] = -\sum_x p(x)\,\log_2 p(x) \]

It is the average number of bits needed to encode one sample under an optimal code. Uncertain distributions cost more bits; certain ones cost none.

15. What has to happen first: Entropy of the weather p, term by term

Ranking

Put in order

Put the moves of Entropy of the weather p, term by term into the order they have to happen.

  1. Sunny term: 0.5 × 1 = 0.5
  2. Cloudy term: 0.25 × 2 = 0.5
  3. Rainy term: 0.25 × 2 = 0.5
  4. Sum: H(p) = 0.5 + 0.5 + 0.5 = 1.5 bits

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. p(S)·I(S). Sunny is common but cheap, contributing half a bit.

16. Entropy of the weather p, term by term

Worked example

Multiply each surprise by its probability, then sum. We already have the surprises [1, 2, 2]:

Sunny term: 0.5 × 1 = 0.5

Why: p(S)·I(S). Sunny is common but cheap, contributing half a bit.

Cloudy term: 0.25 × 2 = 0.5

Why: p(C)·I(C). Rarer but more surprising — also half a bit.

Rainy term: 0.25 × 2 = 0.5

Why: p(R)·I(R). Same as cloudy by symmetry.

Sum: H(p) = 0.5 + 0.5 + 0.5 = 1.5 bits

Why: On average it takes 1.5 bits to encode one day of this weather. Notice each outcome happens to contribute exactly 0.5 — a coincidence of these particular probabilities, not a rule.

\[ H(p) = -\big(0.5\log_2 0.5 + 0.25\log_2 0.25 + 0.25\log_2 0.25\big) = 1.5 \]

17. Decode the notation: Entropy of the weather p, term by term

Notation

Annotate

From Entropy of the weather p, term by term — read this one piece at a time. What is each part doing?

On: \( H(p) = -\big(0.5\log_2 0.5 + 0.25\log_2 0.25 + 0.25\log_2 0.25\big) = 1.5 \)

  • p(S)·I(S). Sunny is common but cheap, contributing half a bit.
  • p(C)·I(C). Rarer but more surprising — also half a bit.
  • On average it takes 1.5 bits to encode one day of this weather. Notice each outcome happens to contribute exactly 0.5 — a coincidence of these particular probabilities, not a rule.

18. What has to be given first: Entropy in code — same 1.5 bits

Missing information

Discussion prompt

The vectorized formula is one line. −np.log2(p) gives the surprise vector, p * (...) weights it, np.sum averages:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

Exactly the hand values. −log2 of a power of two is an integer.

19. Entropy in code — same 1.5 bits

Worked example

The vectorized formula is one line. −np.log2(p) gives the surprise vector, p * (...) weights it, np.sum averages:

import numpy as np
p = np.array([0.5, 0.25, 0.25])          # Sunny, Cloudy, Rainy
surprise = -np.log2(p)                    # bits of surprise per outcome
H = -np.sum(p * np.log2(p))
print(surprise)                          # [1. 2. 2.]
print(round(H, 4))                       # 1.5

surprise = [1., 2., 2.]

Why: Exactly the hand values. −log2 of a power of two is an integer.

H = 1.5 bits

Why: Matches the term-by-term sum. This four-line entropy is the atom every other quantity today is built from.

expressionvalue (verified)
-np.log2(p)[1. 2. 2.]
p * -np.log2(p)[0.5 0.5 0.5]
H = -np.sum(p*np.log2(p))1.5

20. Inspect it line by line: Entropy in code — same 1.5 bits

Error analysis

Annotate

Walk the callouts on Entropy in code — same 1.5 bits. Each one is a place this is easy to get subtly wrong.

  • Exactly the hand values. −log2 of a power of two is an integer.
  • Matches the term-by-term sum. This four-line entropy is the atom every other quantity today is built from.

21. Entropy is maximized by the uniform

Concept

The most uncertain distribution over n outcomes is the uniform one — nothing is predictable — and its entropy is the ceiling log₂ n.

\[ H_{\max} = -\sum_{i=1}^{n} \tfrac{1}{n}\log_2 \tfrac{1}{n} = \log_2 n \]

For n = 3, that ceiling is log₂ 3 ≈ 1.585 bits. Our weather p scored 1.5, just under the max — the 0.5/0.25/0.25 skew makes it slightly more predictable than uniform.

A certain outcome sits at the other extreme: H = 0 (no surprise ever). Entropy always lands in [0, log₂ n].

22. Teach it back: Entropy is maximized by the uniform

Explain it

Discussion prompt

Explain Entropy is maximized by the uniform to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

The most uncertain distribution over n outcomes is the uniform one — nothing is predictable — and its entropy is the ceiling log₂ n.

23. Predict the next row: Coins: from fair to certain

Pattern

Predict first

The table runs: [0.5, 0.5] | 1.0000 | max — most uncertain · [0.6, 0.4] | 0.9710 | slightly predictable

In Coins: from fair to certain, given the rows so far: what is the next one — the row where distribution is [1.0, 0.0]?

Correct: [1.0, 0.0] | 0.0000 | certain — no surprise

distributionH (bits)reading
[0.5, 0.5]1.0000max — most uncertain
[0.6, 0.4]0.9710slightly predictable
[1.0, 0.0]0.0000certain — no surprise

Why: The relationship between the columns, not the individual numbers, is what generates the next row. log2(2) = 1. Maximum uncertainty for two outcomes.

24. Coins: from fair to certain

Worked example

Watch entropy fall as a two-outcome distribution goes from balanced to skewed to certain. The d > 0 mask handles the 0·log 0 = 0 convention:

import numpy as np
for dist in ([0.5, 0.5], [0.6, 0.4], [1.0, 0.0]):
    d = np.array(dist)
    nz = d > 0                           # skip 0*log0 (defined as 0)
    H = -np.sum(d[nz] * np.log2(d[nz]))
    print(dist, round(float(H), 4))

Fair coin [0.5,0.5] → 1.0 bit, the 2-outcome max

Why: log2(2) = 1. Maximum uncertainty for two outcomes.

Biased [0.6,0.4] → 0.9710, and certain [1,0] → 0

Why: The skew drops entropy below 1; a sure outcome hits the floor of 0. Skew always reduces entropy.

distributionH (bits)reading
[0.5, 0.5]1.0000max — most uncertain
[0.6, 0.4]0.9710slightly predictable
[1.0, 0.0]0.0000certain — no surprise

25. Work backwards from the answer: Coins: from fair to certain

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Biased [0.6,0.4] → 0.9710, and certain [1,0] → 0

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Watch entropy fall as a two-outcome distribution goes from balanced to skewed to certain. The d > 0 mask handles the 0·log 0 = 0 convention:

26. Bits or nats — choosing the log base

Concept

The only free choice in every formula today is the log base. Base 2 gives bits; the natural log gives nats. Nothing else changes — the two answers are proportional.

\[ H_{\text{bits}} = -\sum p\log_2 p, \qquad H_{\text{nats}} = -\sum p\ln p, \qquad H_{\text{nats}} = H_{\text{bits}}\cdot \ln 2 \]

Convention: classifiers use nats (framework losses use ln), while the coding/compression story is cleanest in bits. Report the base or your number is ambiguous.

27. Restore the missing line: Same entropy, two units

Fill the middle

Fill in the blanks

From Same entropy, two units — one line has had its right-hand side removed. Put it back.

import numpy as np, math
p = np.array([0.5, 0.25, 0.25])
H_bits = **-np.sum(p * np.log2(p))**
H_nats = -np.sum(p * np.log(p)) # natural log
print(round(float(H_bits), 4)) # 1.5
print(round(float(H_nats), 4)) # 1.0397
print(round(float(H_bits * math.log(2)), 4)) # 1.0397 == nats

Why: H_bits is what everything below it consumes, so the wrong expression here fails later and somewhere else. Same distribution, same uncertainty — only the yardstick changed.

28. Same entropy, two units

Worked example

Compute H(p) for the weather in both bases and confirm the ln 2 bridge between them:

import numpy as np, math
p = np.array([0.5, 0.25, 0.25])
H_bits = -np.sum(p * np.log2(p))
H_nats = -np.sum(p * np.log(p))          # natural log
print(round(float(H_bits), 4))           # 1.5
print(round(float(H_nats), 4))           # 1.0397
print(round(float(H_bits * math.log(2)), 4))  # 1.0397 == nats

H = 1.5 bits = 1.0397 nats

Why: Same distribution, same uncertainty — only the yardstick changed. ln 2 ≈ 0.6931 converts bits to nats.

H_bits · ln2 = 1.0397 reproduces the nats value

Why: The two are one multiply apart. Divide nats by ln 2 to get back to bits.

quantityvalue (verified)
H(p) in bits (log₂)1.5000
H(p) in nats (ln)1.0397
1.5 × ln 21.0397

29. Draw the shape of it: Same entropy, two units

Blank canvas

Draw it

Draw what Same entropy, two units just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.

30. KL divergence

Section

Part 2 of 6 — extra coding cost

31. The wrong codebook costs extra

Intuition

Entropy is the bit-cost of the optimal code for p. But suppose you built your code for a different distribution q and then the real data comes from p. Every message still gets through — but you overpay.

That overpayment — the extra bits per symbol from using q's code on p's data — is exactly what KL divergence measures. It is zero only when q = p (you had the right codebook all along).

32. KL divergence, defined

Concept

The Kullback–Leibler divergence from P to Q weights the per-outcome log-ratio by P:

\[ D_{KL}(P \,\|\, Q) = \sum_x P(x)\,\log_2 \frac{P(x)}{Q(x)} = \mathbb{E}_{x\sim P}\!\left[\log_2 \frac{P(x)}{Q(x)}\right] \]

Read it as: under the true distribution P, how much bigger is −log Q (what you pay) than −log P (what you'd pay optimally), on average.

Two hard facts we'll prove: D_KL ≥ 0 always (zero iff P = Q), and it is not symmetric — D_KL(P‖Q) ≠ D_KL(Q‖P) in general. It is a divergence, not a distance.

33. Plan first: KL of weather vs uniform, term by term

Step zero

Discussion prompt

KL of weather vs uniform, term by term — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Ratios p/q = [1.5, 0.75, 0.75]

Answer:

  1. Ratios p/q = [1.5, 0.75, 0.75]
  2. Weighted logs: 0.5·log₂1.5, 0.25·log₂0.75, 0.25·log₂0.75
  3. Sum = 0.0850 bits

34. KL of weather vs uniform, term by term

Worked example

Let P = p = [0.5, 0.25, 0.25] and Q = q = [⅓, ⅓, ⅓] (uniform). Build D_KL(p‖q) one outcome at a time.

Ratios p/q = [1.5, 0.75, 0.75]

Why: 0.5÷(1/3)=1.5; 0.25÷(1/3)=0.75 twice. Where p exceeds q the ratio is >1 (positive log); where p is below q it is <1 (negative log).

Weighted logs: 0.5·log₂1.5, 0.25·log₂0.75, 0.25·log₂0.75

Why: Multiply each log-ratio by P(x). Numerically these are +0.292481, −0.103759, −0.103759.

\[ D_{KL}(p\|q) = 0.5\log_2 1.5 + 0.25\log_2 0.75 + 0.25\log_2 0.75 \]

Sum = 0.0850 bits

Why: 0.292481 − 0.103759 − 0.103759 = 0.084963. The negative terms don't sink it below zero — the P-weighting guarantees a non-negative total.

\[ D_{KL}(p\|q) \approx 0.0850 \text{ bits} \]

35. Finish it with less help: KL in code — the term vector

Faded example

Fill in the blanks

KL in code — the term vector, with the scaffolding fading: two lines are gone now — fill both.

import numpy as np
p = np.array([0.5, 0.25, 0.25])
q = np.array([1/3, 1/3, 1/3]) # uniform
terms = **p * np.log2(p / q)**
print(np.round(terms, 6)) # [ 0.292481 -0.103759 -0.103759]
print(round(float(np.sum(terms)), 4)) # 0.085

Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. Sunny (p>q) pushes the sum up; cloudy and rainy (p<q) pull it down — but not far enough to go negative.

36. KL in code — the term vector

Worked example

The same computation, so you can see the individual P·log(P/Q) contributions before they're summed:

import numpy as np
p = np.array([0.5, 0.25, 0.25])
q = np.array([1/3, 1/3, 1/3])            # uniform
terms = p * np.log2(p / q)
print(np.round(terms, 6))                # [ 0.292481 -0.103759 -0.103759]
print(round(float(np.sum(terms)), 4))    # 0.085

One term is positive, two are negative

Why: Sunny (p>q) pushes the sum up; cloudy and rainy (p<q) pull it down — but not far enough to go negative.

D_KL(p‖q) = 0.085 bits

Why: Exactly the hand result. Coding p's weather with a uniform codebook wastes about 0.085 bits per day.

outcomep·log₂(p/q)
Sunny+0.292481
Cloudy−0.103759
Rainy−0.103759
D_KL(p‖q) = Σ0.084963

37. Which is which, by p·log₂(p/q)

Discrimination

Sort into buckets

Sort these by p·log₂(p/q), from memory, without looking back at KL in code — the term vector. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

+0.292481
Sunny
−0.103759
Cloudy; Rainy
0.084963
D_KL(p‖q) = Σ
g1
p·log₂(p/q) is "+0.292481" for Sunny — that is what the table on "KL in code — the term vector" records, and it is the single property separating this group from the rest.
g2
p·log₂(p/q) is "−0.103759" for Cloudy, Rainy — that is what the table on "KL in code — the term vector" records, and it is the single property separating this group from the rest.
g3
p·log₂(p/q) is "0.084963" for D_KL(p‖q) = Σ — that is what the table on "KL in code — the term vector" records, and it is the single property separating this group from the rest.

38. Swap the arguments — KL is asymmetric

Worked example

Now compute the reverse, D_KL(q‖p), weighting by q instead of p. If KL were a distance these would be equal:

import numpy as np
p = np.array([0.5, 0.25, 0.25])
q = np.array([1/3, 1/3, 1/3])
KL = lambda a, b: float(np.sum(a * np.log2(a / b)))
print(round(KL(p, q), 4), round(KL(q, p), 4))   # 0.085 0.0817
print(np.isclose(KL(p, q), KL(q, p)))           # False

D_KL(p‖q) = 0.0850 ≠ D_KL(q‖p) = 0.0817

Why: Same two distributions, different numbers — because the weighting distribution changed. The direction of KL is part of its meaning.

np.isclose(...) prints False

Why: Numerical proof of asymmetry. Never assume you can flip the arguments.

directionvalue (bits)weighted by
D_KL(p‖q)0.0850p (true = weather)
D_KL(q‖p)0.0817q (true = uniform)
equal?no — asymmetric

39. Something is wrong here: treating KL as a distance

Anomaly

Predict first

A student writes this, and it looks reasonable:

KL measures how far two distributions are, so D_KL(P‖Q) = D_KL(Q‖P) and it obeys the triangle inequality — just use whichever order is handy.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: We just measured 0.0850 ≠ 0.0817 on the SAME pair — the order changed the answer.

KL is a directional divergence: D_KL(P‖Q) averages the log-ratio under P, so the first argument (the 'truth' you weight by) is load-bearing.

Why: We just measured 0.0850 ≠ 0.0817 on the SAME pair — the order changed the answer. KL also fails the triangle inequality. Flipping arguments silently changes what you optimize.

40. Trap: treating KL as a distance

Trap

The trap

KL measures how far two distributions are, so D_KL(P‖Q) = D_KL(Q‖P) and it obeys the triangle inequality — just use whichever order is handy.

Assume symmetry, pick either order

Why: We just measured 0.0850 ≠ 0.0817 on the SAME pair — the order changed the answer. KL also fails the triangle inequality. Flipping arguments silently changes what you optimize.

The fix

KL is a directional divergence: D_KL(P‖Q) averages the log-ratio under P, so the first argument (the 'truth' you weight by) is load-bearing.

Keep the order fixed to your problem

Why: Forward KL D_KL(P‖Q) is mean-seeking (covers all of P's support); reverse KL D_KL(Q‖P) is mode-seeking (locks onto one mode). VAEs minimize the reverse KL D_KL(q‖p) — choosing the wrong direction changes the whole behavior of the model.

41. Break it on purpose: treating KL as a distance

Break the constraint

Discussion prompt

The rule this trap just fixed:

KL is a directional divergence: D_KL(P‖Q) averages the log-ratio under P, so the first argument (the 'truth' you weight by) is load-bearing.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

We just measured 0.0850 ≠ 0.0817 on the SAME pair — the order changed the answer. KL also fails the triangle inequality. Flipping arguments silently changes what you optimize.

42. Why KL ≥ 0

Section

Part 3 of 6 — Gibbs' inequality

43. The tool: Jensen's inequality

Concept

Jensen's inequality: for a concave function f (like log), the function of the average is at least the average of the function:

\[ f\big(\mathbb{E}[Z]\big) \;\ge\; \mathbb{E}\big[f(Z)\big] \qquad (f \text{ concave}) \]

The log curve bends downward, so a chord always lies below it — averaging inside the log beats averaging outside. Equality holds only when Z is constant. This single fact forces KL ≥ 0.

44. What has to happen first: Prove D_KL(P‖Q) ≥ 0

Ranking

Put in order

Put the moves of Prove D_KL(P‖Q) ≥ 0 into the order they have to happen.

  1. Flip the sign inside the log
  2. Apply Jensen with f = log (concave)
  3. The inner expectation collapses to 1
  4. Chain it: −D_KL ≤ 0, so D_KL ≥ 0

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Write −D_KL by inverting the ratio: −Σ P log(P/Q) = Σ P log(Q/P).

45. Prove D_KL(P‖Q) ≥ 0

Worked example

Flip the sign inside the log

Why: Write −D_KL by inverting the ratio: −Σ P log(P/Q) = Σ P log(Q/P). We'll show this is ≤ 0.

\[ -D_{KL}(P\|Q) = \sum_x P(x)\,\log \frac{Q(x)}{P(x)} = \mathbb{E}_{P}\!\left[\log \frac{Q}{P}\right] \]

Apply Jensen with f = log (concave)

Why: Move the expectation inside the log: E[log Z] ≤ log E[Z]. Here Z = Q/P under P-weighting.

\[ \mathbb{E}_{P}\!\left[\log \frac{Q}{P}\right] \;\le\; \log \mathbb{E}_{P}\!\left[\frac{Q}{P}\right] \]

The inner expectation collapses to 1

Why: E_P[Q/P] = Σ P·(Q/P) = Σ Q = 1, since Q is a probability distribution. And log 1 = 0.

\[ \log \sum_x P(x)\cdot\frac{Q(x)}{P(x)} = \log \sum_x Q(x) = \log 1 = 0 \]

Chain it: −D_KL ≤ 0, so D_KL ≥ 0

Why: We showed −D_KL(P‖Q) ≤ 0. Multiply by −1 (flip the inequality): D_KL(P‖Q) ≥ 0. Equality holds exactly when Q/P is constant, i.e. Q = P. This is Gibbs' inequality.

\[ \boxed{\,D_{KL}(P\|Q) \ge 0,\quad \text{with } = 0 \iff P = Q\,} \]

46. Say it in words: Prove D_KL(P‖Q) ≥ 0

Translation

\( \log \sum_x P(x)\cdot\frac{Q(x)}{P(x)} = \log \sum_x Q(x) = \log 1 = 0 \)

Draw it

Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.

47. Restore the missing line: KL ≥ 0, tested on 1000 random pairs

Fill the middle

Fill in the blanks

From KL ≥ 0, tested on 1000 random pairs — one line has had its right-hand side removed. Put it back.

import numpy as np
rng = np.random.default_rng(0)
KL = lambda a, b: float(np.sum(a * np.log2(a / b)))
smallest = 1e9
for _ in range(1000):
a = rng.random(4); a /= a.sum() # random distribution
b = rng.random(4); b /= b.sum()
smallest = min(smallest, KL(a, b))
print(round(smallest, 4)) # > 0 : never went negative
p = np.array([0.5, 0.25, 0.25])
print(KL(p, p)) # 0.0 : zero only when a == b

Why: b is what everything below it consumes, so the wrong expression here fails later and somewhere else. Not one of a thousand random pairs produced a negative KL — the floor at 0 holds.

48. KL ≥ 0, tested on 1000 random pairs

Worked example

Proof is one thing; let's stress-test it. Draw 1000 random distribution pairs, take the smallest KL seen, and check KL(p‖p) = 0:

import numpy as np
rng = np.random.default_rng(0)
KL = lambda a, b: float(np.sum(a * np.log2(a / b)))
smallest = 1e9
for _ in range(1000):
    a = rng.random(4); a /= a.sum()       # random distribution
    b = rng.random(4); b /= b.sum()
    smallest = min(smallest, KL(a, b))
print(round(smallest, 4))                 # > 0 : never went negative
p = np.array([0.5, 0.25, 0.25])
print(KL(p, p))                           # 0.0 : zero only when a == b

smallest KL over 1000 pairs = 0.0043 > 0

Why: Not one of a thousand random pairs produced a negative KL — the floor at 0 holds. The seed rng(0) makes this reproducible.

KL(p‖p) = 0.0 exactly

Why: A distribution has zero divergence from itself — the equality case of Gibbs.

quantityvalue (verified)
min KL over 1000 random pairs0.0043
ever negative?no
KL(p‖p)0.0

49. Where does each piece belong: Lesson 14: Information Theory

Sorting

Sort into buckets

These are the pieces of Lesson 14: Information Theory, out of order. Put each one back under the part of the lesson it belongs to.

Surprise & entropy
The running example: tomorrow's weather; Rare events carry more information; Surprise = −log p
KL divergence
The wrong codebook costs extra; KL divergence, defined; KL of weather vs uniform, term by term
Why KL ≥ 0
The tool: Jensen's inequality; Prove D_KL(P‖Q) ≥ 0; KL ≥ 0, tested on 1000 random pairs
s1
Surprise & entropy is where Lesson 14: Information Theory puts The running example: tomorrow's weather, Rare events carry more information, Surprise = −log p. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
KL divergence is where Lesson 14: Information Theory puts The wrong codebook costs extra, KL divergence, defined, KL of weather vs uniform, term by term. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
Why KL ≥ 0 is where Lesson 14: Information Theory puts The tool: Jensen's inequality, Prove D_KL(P‖Q) ≥ 0, KL ≥ 0, tested on 1000 random pairs. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

50. Cross-entropy = H(P) + KL

Section

Part 4 of 6 — the classifier loss

51. Cross-entropy, defined

Concept

Cross-entropy is the average number of bits to encode data from P using a code built for Q — the total cost, optimal part plus waste:

\[ H(P, Q) = -\sum_x P(x)\,\log_2 Q(x) = \mathbb{E}_{x\sim P}\big[-\log_2 Q(x)\big] \]

It looks almost like entropy, but the log is of Q (your model) while the weighting is by P (the truth). That mismatch is where the extra cost lives — and it factors cleanly.

52. This IS the classifier loss you already use

Intuition

When a network classifies an image, the label is a one-hot P (all mass on the true class) and the softmax output is your Q. The training loss −log Q(true class) is cross-entropy H(P, Q) with a one-hot P.

With one-hot labels the entropy H(P) = 0, so the identity collapses to H(P,Q) = D_KL(P‖Q) exactly — the loss you minimize is the KL from the true label distribution to your prediction. That is why 'minimize cross-entropy' and 'match the target distribution' are the same instruction.

Frameworks compute it in nats (natural log) and average over the batch, but it is the same quantity we just defined — no new idea, only new packaging.

53. Plan first: Derive H(P,Q) = H(P) + D_KL(P‖Q)

Step zero

Discussion prompt

Derive H(P,Q) = H(P) + D_KL(P‖Q) — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Start from the definition of cross-entropy

Answer:

  1. Start from the definition of cross-entropy
  2. Insert P(x)/P(x) inside the log's argument
  3. Split the log of a product into a sum of logs
  4. Recognize the two pieces

54. Derive H(P,Q) = H(P) + D_KL(P‖Q)

Worked example

Start from the definition of cross-entropy

Why: Just the formula, nothing added.

\[ H(P,Q) = -\sum_x P(x)\log_2 Q(x) \]

Insert P(x)/P(x) inside the log's argument

Why: Multiply Q by P/P — a legal '×1' — so we can split it. This is the one clever move; everything after is algebra.

\[ = -\sum_x P(x)\log_2\!\Big(Q(x)\cdot\tfrac{P(x)}{P(x)}\Big) = -\sum_x P(x)\log_2\!\Big(P(x)\cdot\tfrac{Q(x)}{P(x)}\Big) \]

Split the log of a product into a sum of logs

Why: log(P · Q/P) = log P + log(Q/P), then distribute the −Σ P over the two pieces.

\[ = -\sum_x P(x)\log_2 P(x) \;-\; \sum_x P(x)\log_2 \frac{Q(x)}{P(x)} \]

Recognize the two pieces

Why: The first sum is H(P). The second is −Σ P log(Q/P) = +Σ P log(P/Q) = D_KL(P‖Q). Done.

\[ \boxed{\,H(P,Q) = H(P) + D_{KL}(P\|Q)\,} \]

55. Guess the shape of the answer: Verify the identity on the weather data

Estimation

Predict first

Compute H(P,Q) directly and compare to H(P) + D_KL(P‖Q), reusing p and the uniform q:

Commit before you compute: what does Verify the identity on the weather data come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: np.isclose(...) → True

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Cross-entropy = irreducible entropy + avoidable KL penalty, confirmed numerically.

56. Verify the identity on the weather data

Worked example

Compute H(P,Q) directly and compare to H(P) + D_KL(P‖Q), reusing p and the uniform q:

import numpy as np
p = np.array([0.5, 0.25, 0.25])
q = np.array([1/3, 1/3, 1/3])
H  = -np.sum(p * np.log2(p))              # entropy of p
KL =  np.sum(p * np.log2(p / q))          # D_KL(p||q)
CE = -np.sum(p * np.log2(q))              # cross-entropy H(p,q)
print(round(float(CE), 4))               # 1.585
print(round(float(H + KL), 4))           # 1.585
print(np.isclose(CE, H + KL))            # True

CE = 1.585 = H(p) 1.5 + KL 0.085

Why: The identity holds to floating point. Since q is uniform, −log2 q = log2 3 for every outcome, so H(p,q) = log2 3 ≈ 1.585 exactly.

np.isclose(...) → True

Why: Cross-entropy = irreducible entropy + avoidable KL penalty, confirmed numerically.

termvalue (bits)
H(p) (entropy of truth)1.5000
D_KL(p‖q) (penalty)0.0850
H(p,q) (cross-entropy)1.5850

57. Something is wrong here: what minimizing cross-entropy does

Anomaly

Predict first

A student writes this, and it looks reasonable:

Training minimizes cross-entropy H(P,Q), so with enough training the loss should be drivable all the way to 0.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: H(P,Q) = H(P) + D_KL(P‖Q). H(P) is fixed by the data's own label noise — you can't optimize it away.

Because H(P) is a constant w.r.t. your model Q, minimizing cross-entropy is exactly minimizing D_KL(P‖Q).

Why: H(P,Q) = H(P) + D_KL(P‖Q). H(P) is fixed by the data's own label noise — you can't optimize it away. The loss floor is H(P), which is 1.5 bits here, not 0. Chasing 0 means overfitting the noise.

58. Trap: what minimizing cross-entropy does

Trap

The trap

Training minimizes cross-entropy H(P,Q), so with enough training the loss should be drivable all the way to 0.

Expect loss → 0 on noisy labels

Why: H(P,Q) = H(P) + D_KL(P‖Q). H(P) is fixed by the data's own label noise — you can't optimize it away. The loss floor is H(P), which is 1.5 bits here, not 0. Chasing 0 means overfitting the noise.

The fix

Because H(P) is a constant w.r.t. your model Q, minimizing cross-entropy is exactly minimizing D_KL(P‖Q).

Training pushes Q → P; the floor is H(P)

Why: ∂H(P,Q)/∂Q only touches the KL term. So fitting a classifier by cross-entropy IS distribution matching — driving your predicted distribution Q toward the true label distribution P. The best achievable loss equals the label entropy H(P).

59. Which of these survive contact with Lesson 14: Information Theory?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
One distribution carries this whole lesson. A weather model says tomorrow is Sunny with probability 0.5, Cloudy 0.25, Rainy 0.25. Call it p.; The surprise (self-information) of an outcome with probability p is defined as:; The most uncertain distribution over n outcomes is the uniform one — nothing is predictable — and its entropy is the ceiling log₂ n.
Breaks
KL measures how far two distributions are, so D_KL(P‖Q) = D_KL(Q‖P) and it obeys the triangle inequality — just use whichever order is handy.; Training minimizes cross-entropy H(P,Q), so with enough training the loss should be drivable all the way to 0.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 14: Information Theory puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

60. Mutual information

Section

Part 5 of 6 — shared bits

61. A joint distribution to work with

Concept

Mutual information needs two variables. Take a 2×2 joint P(X, Y) — say X = was it cloudy, Y = did it rain — with these joint probabilities:

P(X,Y)Y = 0Y = 1row sum p(x)
X = 00.300.200.50
X = 10.100.400.50
col sum p(y)0.400.601.00

The margins are the row/column sums: p(x) = [0.5, 0.5], p(y) = [0.4, 0.6]. If X and Y were independent, the joint would equal the outer product p(x)p(y) — it does not, so they share information.

62. Watch it run: A joint distribution to work with

Pattern

Step through it

Step through A joint distribution to work with one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: P(X,Y) is X = 0
  2. Step 2: P(X,Y) is X = 1
  3. Step 3: P(X,Y) is col sum p(y)

63. Picture it first: Information as overlapping circles

Picture it

Figure (svg): Two overlapping circles labeled H(X) and H(Y). The left-only crescent is H(X|Y), the right-only crescent is H(Y|X), the lens-shaped overlap in the middle is the mutual information I(X;Y), and the whole union is the joint entropy H(X,Y).

The overlap is I(X;Y); the union is H(X,Y); each crescent is a conditional entropy.

Discussion prompt

Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.

Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.

Answer:

Picture the uncertainty in X as one circle of area H(X) and the uncertainty in Y as another of area H(Y). When the variables are related, the circles overlap.

64. Information as overlapping circles

Intuition

Picture the uncertainty in X as one circle of area H(X) and the uncertainty in Y as another of area H(Y). When the variables are related, the circles overlap.

Figure (svg): Two overlapping circles labeled H(X) and H(Y). The left-only crescent is H(X|Y), the right-only crescent is H(Y|X), the lens-shaped overlap in the middle is the mutual information I(X;Y), and the whole union is the joint entropy H(X,Y).

The overlap is I(X;Y); the union is H(X,Y); each crescent is a conditional entropy.

The overlap is the mutual information — the uncertainty they share. The full union is the joint entropy H(X,Y), and each crescent is a conditional entropy (H(X|Y), H(Y|X)). Pull the circles apart until they just touch and the overlap — the MI — drops to zero: that is independence.

65. Mutual information, two equivalent forms

Concept

Mutual information I(X;Y) is how many bits knowing one variable saves you on the other. It is a KL divergence — the joint from the independence model:

\[ I(X;Y) = \sum_{x,y} P(x,y)\,\log_2 \frac{P(x,y)}{p(x)\,p(y)} = D_{KL}\big(P(x,y)\,\|\,p(x)p(y)\big) \]

Equivalently, in entropies — the shared area of two overlapping circles:

\[ I(X;Y) = H(X) + H(Y) - H(X,Y) \]

Because it is a KL, I(X;Y) ≥ 0, and it is 0 exactly when P(x,y) = p(x)p(y) — i.e. when X and Y are independent.

66. Finish it with less help: MI from the joint — the KL form

Faded example

Fill in the blanks

MI from the joint — the KL form, with the scaffolding fading: two lines are gone now — fill both.

import numpy as np
Pxy = np.array([[0.30, 0.20], # rows = X, cols = Y
[0.10, 0.40]])
px = Pxy.sum(axis=1) # marginal p(x)
py = Pxy.sum(axis=0) # marginal p(y)
MI = np.sum(Pxy * np.log2(Pxy / np.outer(px, py)))
print(px, py) # [0.5 0.5] [0.4 0.6]
print(round(float(MI), 4)) # 0.1245

Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. Row sums give p(x); column sums give p(y).

67. MI from the joint — the KL form

Worked example

Compute the margins from the table, then the joint-vs-independent KL directly:

import numpy as np
Pxy = np.array([[0.30, 0.20],            # rows = X, cols = Y
                [0.10, 0.40]])
px = Pxy.sum(axis=1)                      # marginal p(x)
py = Pxy.sum(axis=0)                      # marginal p(y)
MI = np.sum(Pxy * np.log2(Pxy / np.outer(px, py)))
print(px, py)                             # [0.5 0.5] [0.4 0.6]
print(round(float(MI), 4))                # 0.1245

Margins come out [0.5, 0.5] and [0.4, 0.6]

Why: Row sums give p(x); column sums give p(y). np.outer(px,py) builds the independence model p(x)p(y).

I(X;Y) = 0.1245 bits

Why: The joint diverges from independence by 0.1245 bits — that's how much X and Y share.

cell (x,y)P(x,y)p(x)p(y)P·log₂(P / pxpy)
(0,0)0.300.20+0.175489
(0,1)0.200.30−0.116993
(1,0)0.100.20−0.100000
(1,1)0.400.30+0.166015
I(X;Y) = Σ0.124511

68. Cross-check with H(X)+H(Y)−H(X,Y)

Worked example

The entropy form must give the identical number. Compute the three entropies and combine:

import numpy as np
Pxy = np.array([[0.30, 0.20], [0.10, 0.40]])
H = lambda d: -np.sum(d[d > 0] * np.log2(d[d > 0]))
Hx, Hy, Hxy = H(Pxy.sum(1)), H(Pxy.sum(0)), H(Pxy.ravel())
print(round(float(Hx), 4), round(float(Hy), 4), round(float(Hxy), 4))
print(round(float(Hx + Hy - Hxy), 4))     # 0.1245

H(X)=1.0, H(Y)=0.9710, H(X,Y)=1.8464

Why: p(x) is uniform so H(X)=1; p(y)=[0.4,0.6] gives 0.9710; the flattened joint gives 1.8464.

H(X)+H(Y)−H(X,Y) = 0.1245 — matches the KL form

Why: Both routes to mutual information agree to 4 decimals. The joint entropy is LESS than H(X)+H(Y); the shortfall is exactly the shared information.

quantityvalue (bits)
H(X)1.0000
H(Y)0.9710
H(X,Y)1.8464
H(X)+H(Y)−H(X,Y)0.1245

69. Predict the next row: The third route: I(X;Y) = H(Y) − H(Y|X)

Pattern

Predict first

The table runs: H(Y) | 0.9710 · H(Y|X) | 0.8464

In The third route: I(X;Y) = H(Y) − H(Y|X), given the rows so far: what is the next one — the row where quantity is I(X;Y) = H(Y) − H(Y|X)?

Correct: I(X;Y) = H(Y) − H(Y|X) | 0.1245

quantityvalue (bits)
H(Y)0.9710
H(Y|X)0.8464
I(X;Y) = H(Y) − H(Y|X)0.1245

Why: The relationship between the columns, not the individual numbers, is what generates the next row. Each row Pxy[i]/px[i] is the conditional p(Y|X=x); average their entropies weighted by p(x).

70. The third route: I(X;Y) = H(Y) − H(Y|X)

Worked example

The overlap picture says MI is also 'uncertainty in Y minus leftover uncertainty in Y once you know X'. The conditional entropy H(Y|X) is the right crescent. Build it and subtract:

import numpy as np
Pxy = np.array([[0.30, 0.20], [0.10, 0.40]])
H = lambda d: -np.sum(d[d > 0] * np.log2(d[d > 0]))
px = Pxy.sum(axis=1)
HYgivenX = sum(px[i] * H(Pxy[i] / px[i]) for i in range(2))  # Σ p(x) H(Y|X=x)
HY = H(Pxy.sum(axis=0))
print(round(float(HY), 4), round(float(HYgivenX), 4))   # 0.971 0.8464
print(round(float(HY - HYgivenX), 4))                    # 0.1245

H(Y) = 0.9710, H(Y|X) = 0.8464

Why: Each row Pxy[i]/px[i] is the conditional p(Y|X=x); average their entropies weighted by p(x). Knowing X shaves H(Y) from 0.971 down to 0.846.

I(X;Y) = H(Y) − H(Y|X) = 0.1245

Why: The reduction in Y's uncertainty from learning X — identical to both the KL form and H(X)+H(Y)−H(X,Y). Three routes, one number.

quantityvalue (bits)
H(Y)0.9710
H(Y|X)0.8464
I(X;Y) = H(Y) − H(Y|X)0.1245

71. Fill in: value (bits) for The third route: I(X;Y) = H(Y) − H(Y|X)

Comparison

Comparison matrix

From The third route: I(X;Y) = H(Y) − H(Y|X): refill the value (bits) column from what you know. The rest of the table is as it appeared.

quantityvalue (bits)
H(Y)0.9710
H(Y|X)0.8464
I(X;Y) = H(Y) − H(Y|X)0.1245

72. Restore the missing line: Independence forces I = 0

Fill the middle

Fill in the blanks

From Independence forces I = 0 — one line has had its right-hand side removed. Put it back.

import numpy as np
# independent: joint = outer product of marginals -> I = 0
px = np.array([0.5, 0.5]); py = np.array([0.4, 0.6])
Pind = np.outer(px, py)
MI = np.sum(Pind * np.log2(Pind / np.outer(Pind.sum(1), Pind.sum(0))))
print(Pind)
print(round(float(MI), 6)) # 0.0

Why: Pind is what everything below it consumes, so the wrong expression here fails later and somewhere else. By construction the joint is the independence model, so the log-ratio inside the sum is log(1) = 0 in every cell.

73. Independence forces I = 0

Worked example

Build a joint that IS the outer product of the margins — genuine independence — and watch MI vanish:

import numpy as np
# independent: joint = outer product of marginals -> I = 0
px = np.array([0.5, 0.5]); py = np.array([0.4, 0.6])
Pind = np.outer(px, py)
MI = np.sum(Pind * np.log2(Pind / np.outer(Pind.sum(1), Pind.sum(0))))
print(Pind)
print(round(float(MI), 6))                # 0.0

Pind exactly equals p(x)p(y)

Why: By construction the joint is the independence model, so the log-ratio inside the sum is log(1) = 0 in every cell.

I(X;Y) = 0.0

Why: Knowing one variable tells you nothing about the other. MI = 0 ⇔ independence — the cleanest test for dependence in all of statistics.

joint usedindependent?I(X;Y)
[[0.3,0.2],[0.1,0.4]]no0.1245
outer(p(x), p(y))yes0.0

74. Draw the shape of it: Independence forces I = 0

Blank canvas

Draw it

Draw what Independence forces I = 0 just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.

75. The ELBO

Section

Part 6 of 6 — the VAE objective

76. Why we need a bound at all

Intuition

Latent-variable models want the evidence log p(x) = log Σ_z p(x, z). But that sum runs over every possible latent z — for a real VAE, an integral over a continuous high-dimensional space that has no closed form.

So we introduce an easy, tunable approximation q(z|x) to the true-but-intractable posterior p(z|x), and instead of computing log p(x) we lower-bound it with a quantity we can compute and maximize. That bound is the ELBO.

77. The exact decomposition

Concept

The log-evidence splits exactly into the ELBO plus a KL gap to the true posterior — an identity, not an approximation:

\[ \log p(x) = \underbrace{\mathbb{E}_{q(z|x)}\!\left[\log \frac{p(x,z)}{q(z|x)}\right]}_{\text{ELBO}(q)} \;+\; D_{KL}\big(q(z|x)\,\|\,p(z|x)\big) \]

Since D_KL ≥ 0 (Part 3), the ELBO can never exceed log p(x) — it is a genuine lower bound. And the only slack is that KL: the tighter q matches the true posterior, the closer the bound.

78. What has to happen first: Derive the split in three lines

Ranking

Put in order

Put the moves of Derive the split in three lines into the order they have to happen.

  1. Write log p(x) as an expectation over q
  2. Expand p(x) via Bayes and multiply by q/q
  3. Split the expectation into the two named pieces
  4. Bound follows because KL ≥ 0

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. log p(x) doesn't depend on z, so its expectation under any q(z|x) is itself.

79. Derive the split in three lines

Worked example

Write log p(x) as an expectation over q

Why: log p(x) doesn't depend on z, so its expectation under any q(z|x) is itself. Free move.

\[ \log p(x) = \mathbb{E}_{q(z|x)}\big[\log p(x)\big] \]

Expand p(x) via Bayes and multiply by q/q

Why: p(x) = p(x,z)/p(z|x). Insert q(z|x)/q(z|x) inside the log so both a p(x,z)/q and a q/p(z|x) piece appear.

\[ = \mathbb{E}_{q}\!\left[\log \frac{p(x,z)}{q(z|x)} + \log \frac{q(z|x)}{p(z|x)}\right] \]

Split the expectation into the two named pieces

Why: The first expectation is the ELBO; the second is E_q[log(q/p(z|x))] = D_KL(q‖p(z|x)). That's the boxed identity.

\[ = \text{ELBO}(q) + D_{KL}\big(q(z|x)\,\|\,p(z|x)\big) \]

Bound follows because KL ≥ 0

Why: Drop the non-negative KL term to get log p(x) ≥ ELBO(q). Maximizing the ELBO both raises the evidence bound AND shrinks the KL, pulling q toward the true posterior — the entire VAE training objective.

\[ \boxed{\,\log p(x) \ge \text{ELBO}(q)\,} \]

80. Decode the notation: Derive the split in three lines

Notation

Annotate

From Derive the split in three lines — read this one piece at a time. What is each part doing?

On: \( \boxed{\,\log p(x) \ge \text{ELBO}(q)\,} \)

  • log p(x) doesn't depend on z, so its expectation under any q(z|x) is itself. Free move.
  • p(x) = p(x,z)/p(z|x). Insert q(z|x)/q(z|x) inside the log so both a p(x,z)/q and a q/p(z|x) piece appear.
  • The first expectation is the ELBO; the second is E_q[log(q/p(z|x))] = D_KL(q‖p(z|x)). That's the boxed identity.

81. What has to be given first: The ELBO decomposition, numerically

Missing information

Discussion prompt

A tiny discrete latent z (4 values) makes the identity concrete. Fix x, give a joint p(x,z), a candidate q, and check that ELBO + KL = log p(x) and ELBO ≤ log p(x):

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

The bound holds: a mismatched q gives an ELBO strictly less than the true log-evidence.

82. The ELBO decomposition, numerically

Worked example

A tiny discrete latent z (4 values) makes the identity concrete. Fix x, give a joint p(x,z), a candidate q, and check that ELBO + KL = log p(x) and ELBO ≤ log p(x):

import numpy as np
pxz = np.array([0.1, 0.2, 0.3, 0.05])     # joint p(x, z) over 4 z-values (fixed x)
px  = pxz.sum()                            # marginal p(x)
post = pxz / px                            # true posterior p(z|x)
q = np.array([0.4, 0.3, 0.2, 0.1])         # approximate posterior
logpx = np.log(px)
elbo = np.sum(q * np.log(pxz / q))         # E_q[log p(x,z) - log q]
klqp = np.sum(q * np.log(q / post))        # D_KL(q || p(z|x))
print(round(float(logpx), 4))              # -0.4308
print(round(float(elbo), 4))               # -0.6644
print(round(float(elbo + klqp), 4))        # -0.4308  (== log p(x))
print(elbo <= logpx)                       # True

ELBO = −0.6644 sits below log p(x) = −0.4308

Why: The bound holds: a mismatched q gives an ELBO strictly less than the true log-evidence.

ELBO + D_KL(q‖p) = −0.4308 = log p(x)

Why: The gap between the ELBO and the evidence is EXACTLY the posterior KL — the identity is confirmed to 4 decimals. Improve q → shrink KL → tighten the bound.

quantityvalue (nats, verified)
log p(x)−0.4308
ELBO(q)−0.6644
D_KL(q‖p(z|x))0.2336
ELBO + KL−0.4308
ELBO ≤ log p(x)?True

83. Work backwards from the answer: The ELBO decomposition, numerically

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

ELBO + D_KL(q‖p) = −0.4308 = log p(x)

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

A tiny discrete latent z (4 values) makes the identity concrete. Fix x, give a joint p(x,z), a candidate q, and check that ELBO + KL = log p(x) and ELBO ≤ log p(x):

84. The Gaussian-KL closed form

Concept

Real VAEs use Gaussian q and prior, so the KL term has a formula — no sampling needed. For two 1-D Gaussians (in nats, natural log):

\[ D_{KL}\big(\mathcal{N}(\mu_1,\sigma_1^2)\,\|\,\mathcal{N}(\mu_2,\sigma_2^2)\big) = \log\frac{\sigma_2}{\sigma_1} + \frac{\sigma_1^2 + (\mu_1-\mu_2)^2}{2\sigma_2^2} - \frac12 \]

The VAE special case matches q = N(μ, σ²) to a standard-normal prior N(0,1), which collapses to ½(σ² + μ² − 1 − log σ²) — the KL regularizer in the VAE loss.

85. Teach it back: The Gaussian-KL closed form

Explain it

Discussion prompt

Explain The Gaussian-KL closed form to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Real VAEs use Gaussian q and prior, so the KL term has a formula — no sampling needed. For two 1-D Gaussians (in nats, natural log):

86. Guess the shape of the answer: Gaussian-KL: formula and Monte-Carlo agree

Estimation

Predict first

Compute D_KL(N(0,1) ‖ N(1, var=2)) from the closed form, then confirm it by sampling log(p/q) under p:

Commit before you compute: what does Gaussian-KL: formula and Monte-Carlo agree come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Monte-Carlo = 0.3462, matches to 3 decimals

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Averaging log(p/q) over 2 million samples from p reproduces the analytic KL, validating the closed form.

87. Gaussian-KL: formula and Monte-Carlo agree

Worked example

Compute D_KL(N(0,1) ‖ N(1, var=2)) from the closed form, then confirm it by sampling log(p/q) under p:

import numpy as np, math
def gauss_kl(m1, s1, m2, s2):
    return math.log(s2/s1) + (s1**2 + (m1-m2)**2) / (2*s2**2) - 0.5
closed = gauss_kl(0, 1, 1, math.sqrt(2))   # N(0,1) || N(1, var=2)
rng = np.random.default_rng(0)
xs = rng.normal(0, 1, 2_000_000)           # samples from p = N(0,1)
logp = -0.5*np.log(2*np.pi*1) - (xs-0)**2/(2*1)
logq = -0.5*np.log(2*np.pi*2) - (xs-1)**2/(2*2)
print(round(closed, 4))                    # 0.3466 nats
print(round(float(np.mean(logp - logq)), 4))  # 0.3462 (MC)

Closed form = 0.3466 nats

Why: Plug μ1=0, σ1=1, μ2=1, σ2²=2 into the formula. Because units are nats (natural log), not bits.

Monte-Carlo = 0.3462, matches to 3 decimals

Why: Averaging log(p/q) over 2 million samples from p reproduces the analytic KL, validating the closed form. The tiny gap is sampling noise.

methodD_KL (nats, verified)
closed form0.3466
Monte-Carlo (2e6 samples)0.3462
agree?yes (within noise)

88. Inspect it line by line: Gaussian-KL: formula and Monte-Carlo agree

Error analysis

Annotate

Walk the callouts on Gaussian-KL: formula and Monte-Carlo agree. Each one is a place this is easy to get subtly wrong.

  • Plug μ1=0, σ1=1, μ2=1, σ2²=2 into the formula. Because units are nats (natural log), not bits.
  • Averaging log(p/q) over 2 million samples from p reproduces the analytic KL, validating the closed form. The tiny gap is sampling noise.

89. One measure, wearing five hats

Concept

Step back: almost every quantity today was KL divergence in disguise. Recognizing that turns five formulas into one idea — the gap between what's true and what your model assumes.

appears asthe KL underneath
compression wasteD_KL(P‖Q) = extra bits from the wrong codebook
cross-entropy lossH(P,Q) − H(P) = D_KL(P‖Q)
mutual informationD_KL(P(x,y) ‖ p(x)p(y))
ELBO slacklog p(x) − ELBO = D_KL(q ‖ p(z|x))

The two facts KL ≥ 0 and KL asymmetric therefore propagate everywhere: every loss has a floor, every bound is one-sided, and direction always matters.

90. By analogy: One measure, wearing five hats

Analogy

Discussion prompt

Explain One measure, wearing five hats by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Step back: almost every quantity today was KL divergence in disguise. Recognizing that turns five formulas into one idea — the gap between what's true and what your model assumes.

91. Rebuild the recipe: The information-theory toolkit

Ranking

Put in order

These are the steps of The information-theory toolkit, scrambled. Put them back in order before the next slide shows you.

  1. Surprise −log₂ p — decreasing in p, adds for independent events; bits (log₂) or nats (ln)
  2. Entropy H(X) = −Σ p log₂ p — average surprise, in [0, log₂ n], max at uniform
  3. KL D_KL(P‖Q) = Σ P log₂(P/Q) — extra coding cost; ≥ 0 (Gibbs/Jensen), asymmetric, not a metric
  4. Cross-entropy H(P,Q) = H(P) + D_KL(P‖Q) — the classifier loss; minimizing it minimizes KL
  5. Mutual information I(X;Y) = H(X)+H(Y)−H(X,Y) = D_KL(P(x,y)‖p(x)p(y)) — 0 iff independent
  6. ELBO log p(x) = ELBO + D_KL(q‖p) ≥ ELBO — maximize the bound; Gaussian-KL has a closed form (VAE)

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

92. The information-theory toolkit

Pattern

  1. Surprise −log₂ p — decreasing in p, adds for independent events; bits (log₂) or nats (ln)
  2. Entropy H(X) = −Σ p log₂ p — average surprise, in [0, log₂ n], max at uniform
  3. KL D_KL(P‖Q) = Σ P log₂(P/Q) — extra coding cost; ≥ 0 (Gibbs/Jensen), asymmetric, not a metric
  4. Cross-entropy H(P,Q) = H(P) + D_KL(P‖Q) — the classifier loss; minimizing it minimizes KL
  5. Mutual information I(X;Y) = H(X)+H(Y)−H(X,Y) = D_KL(P(x,y)‖p(x)p(y)) — 0 iff independent
  6. ELBO log p(x) = ELBO + D_KL(q‖p) ≥ ELBO — maximize the bound; Gaussian-KL has a closed form (VAE)

93. Where does it stop working: The information-theory toolkit

Edge cases

Discussion prompt

The information-theory toolkit works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Surprise −log₂ p — decreasing in p, adds for independent events; bits (log₂) or nats (ln)
  2. Entropy H(X) = −Σ p log₂ p — average surprise, in [0, log₂ n], max at uniform
  3. KL D_KL(P‖Q) = Σ P log₂(P/Q) — extra coding cost; ≥ 0 (Gibbs/Jensen), asymmetric, not a metric
  4. Cross-entropy H(P,Q) = H(P) + D_KL(P‖Q) — the classifier loss; minimizing it minimizes KL
  5. Mutual information I(X;Y) = H(X)+H(Y)−H(X,Y) = D_KL(P(x,y)‖p(x)p(y)) — 0 iff independent
  6. ELBO log p(x) = ELBO + D_KL(q‖p) ≥ ELBO — maximize the bound; Gaussian-KL has a closed form (VAE)

94. Rule out three: Check yourself — maximum entropy

Elimination

Eliminate the wrong options

Over 4 outcomes, which distribution has the HIGHEST entropy?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. [0.25, 0.25, 0.25, 0.25]
  • B. [0.7, 0.1, 0.1, 0.1]
  • C. [1, 0, 0, 0]
  • D. [0.4, 0.4, 0.1, 0.1]

Survives elimination: A

Why: Entropy is maximized by the uniform distribution. With 4 equal outcomes, H = log₂4 = 2 bits — the ceiling for four outcomes, hence the most uncertain.

95. Check yourself — maximum entropy

Check

Which distribution is the most uncertain?

Check your understanding

Over 4 outcomes, which distribution has the HIGHEST entropy?

  • A. [0.25, 0.25, 0.25, 0.25] (correct)
  • B. [0.7, 0.1, 0.1, 0.1]
  • C. [1, 0, 0, 0]
  • D. [0.4, 0.4, 0.1, 0.1]

Answer: A

Why: Entropy is maximized by the uniform distribution. With 4 equal outcomes, H = log₂4 = 2 bits — the ceiling for four outcomes, hence the most uncertain.

Why B tempts people
Heavily skewed toward one outcome, so more predictable → lower entropy. Verified: H ≈ 1.3568 bits, below the uniform's 2.
Why C tempts people
A certain outcome has ZERO entropy — no surprise ever. That's the minimum, not the maximum.
Why D tempts people
Less skewed than B but still non-uniform, so its entropy (≈ 1.7219 bits) is below the uniform's 2 bits.

96. Check yourself — KL divergence

Check

What is actually true of D_KL(P‖Q)?

Check your understanding

Which statement about D_KL(P‖Q) is correct?

  • A. It is ≥ 0, equals 0 iff P = Q, and generally ≠ D_KL(Q‖P) (correct)
  • B. It is a symmetric distance metric obeying the triangle inequality
  • C. It can be negative whenever Q(x) > P(x) at some x
  • D. It equals H(P) − H(Q)

Answer: A

Why: By Gibbs/Jensen, KL is non-negative and zero only when P = Q, and it is directional — we measured 0.0850 vs 0.0817 on the same pair. It is a divergence, not a metric.

Why B tempts people
KL is NOT symmetric and violates the triangle inequality, so it fails two of the metric axioms. The asymmetry demo proved 0.0850 ≠ 0.0817.
Why C tempts people
Individual terms P·log(P/Q) can be negative (we saw two negative cloudy/rainy terms), but the P-weighted SUM is provably ≥ 0. The total never goes negative.
Why D tempts people
That is not the definition. KL is Σ P log(P/Q); it's cross-entropy MINUS entropy, H(P,Q) − H(P), not a difference of two entropies.

97. Answer it before you see the options: Check yourself — cross-entropy

Prediction

Predict first

Minimizing the cross-entropy H(P,Q) over the model Q is equivalent to minimizing…

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: D_KL(P‖Q), because H(P) is constant in Q

Why: H(P,Q) = H(P) + D_KL(P‖Q). H(P) depends only on the fixed data, so optimizing Q can only shrink the KL term — training a classifier is distribution matching, and the loss floor is H(P), not 0.

98. Check yourself — cross-entropy

Check

Decompose the loss before you answer.

Check your understanding

Minimizing the cross-entropy H(P,Q) over the model Q is equivalent to minimizing…

  • A. D_KL(P‖Q), because H(P) is constant in Q (correct)
  • B. the entropy H(P)
  • C. D_KL(Q‖P)
  • D. the mutual information I(P;Q)

Answer: A

Why: H(P,Q) = H(P) + D_KL(P‖Q). H(P) depends only on the fixed data, so optimizing Q can only shrink the KL term — training a classifier is distribution matching, and the loss floor is H(P), not 0.

Why B tempts people
H(P) is fixed by the data and is NOT a function of Q, so there is nothing about it to minimize — it's the irreducible floor of the loss.
Why C tempts people
Cross-entropy contains the FORWARD KL D_KL(P‖Q) — truth first — not the reverse. The direction matters (see the KL-as-distance trap).
Why D tempts people
Mutual information is between two variables of a joint distribution; it isn't what a single cross-entropy loss over labels minimizes.

99. Rule out three: Check yourself — mutual information

Elimination

Eliminate the wrong options

For two random variables, I(X;Y) = 0 exactly when…

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. X and Y are independent, i.e. P(x,y) = p(x)p(y)
  • B. X and Y are identical (Y = X)
  • C. H(X) = H(Y)
  • D. the joint entropy H(X,Y) is maximized

Survives elimination: A

Why: I(X;Y) = D_KL(P(x,y) ‖ p(x)p(y)), and a KL is 0 iff its two arguments are equal. So I = 0 exactly when the joint equals the product of marginals — the definition of independence.

100. Check yourself — mutual information

Check

Think about the KL form of I(X;Y).

Check your understanding

For two random variables, I(X;Y) = 0 exactly when…

  • A. X and Y are independent, i.e. P(x,y) = p(x)p(y) (correct)
  • B. X and Y are identical (Y = X)
  • C. H(X) = H(Y)
  • D. the joint entropy H(X,Y) is maximized

Answer: A

Why: I(X;Y) = D_KL(P(x,y) ‖ p(x)p(y)), and a KL is 0 iff its two arguments are equal. So I = 0 exactly when the joint equals the product of marginals — the definition of independence.

Why B tempts people
If Y = X they are maximally DEPENDENT: I(X;Y) = H(X) > 0, the opposite of 0. Knowing X tells you Y completely.
Why C tempts people
Equal marginal entropies say nothing about dependence — two variables can share information while having identical entropies, or differ in entropy while independent.
Why D tempts people
Maximizing H(X,Y) at fixed margins means the joint is as spread out as possible — that's the INDEPENDENT case, which does give I = 0, but the defining condition is P(x,y)=p(x)p(y), not 'H(X,Y) is maximized' as a standalone fact. The clean, always-correct characterization is independence.

101. Your turn: build the measures

Section

The project

102. Project: entropy, KL, cross-entropy from scratch

Concept

Implement the three core quantities in NumPy on the weather p = [0.5, 0.25, 0.25] and uniform q, then verify the identity H(P,Q) = H(P) + D_KL(P‖Q) numerically. You derived every piece — now assemble it.

#requirementtool
1entropy(p) = −Σ p log₂ pnp.log2
2kl(p,q) = Σ p log₂(p/q), both directionsnp.sum
3Verify cross(p,q) == entropy(p) + kl(p,q)np.isclose

Build rules: type every line yourself, use log2 so answers are in bits, and confirm KL comes out asymmetric before you trust the code.

103. Break it if you can: Project: entropy, KL, cross-entropy from scratch

Counterexample

Discussion prompt

Build rules: type every line yourself, use log2 so answers are in bits, and confirm KL comes out asymmetric before you trust the code.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

104. Milestone 1 — entropy

Worked example

Your turn: write entropy(p). Predict entropy([0.5, 0.25, 0.25]) out loud before you run it — you computed it by hand as 1.5.

Hint: -np.sum(p*np.log2(p)). A fair coin [0.5, 0.5] must give exactly 1.0 bit as a sanity check.

import numpy as np
def entropy(p):
    p = np.asarray(p, dtype=float)
    return -np.sum(p * np.log2(p))
print(round(float(entropy([0.5, 0.25, 0.25])), 4))   # 1.5
print(round(float(entropy([0.5, 0.5])), 4))          # 1.0
inputentropy (bits)
[0.5, 0.25, 0.25]1.5
[0.5, 0.5]1.0
[1.0, 0.0]*0.0 (with 0·log0 = 0)

105. Milestone 2 — KL both ways

Worked example

Your turn: write kl(p, q) and compute both directions for p = [0.5, 0.25, 0.25], q = [⅓, ⅓, ⅓]. Predict: equal or not?

Hint: np.sum(p*np.log2(p/q)); swap the arguments for the reverse. They should come out 0.085 and 0.0817 — different.

import numpy as np
def kl(p, q):
    p = np.asarray(p, dtype=float); q = np.asarray(q, dtype=float)
    return np.sum(p * np.log2(p / q))
p = [0.5, 0.25, 0.25]; q = [1/3, 1/3, 1/3]
print(round(float(kl(p, q)), 4), round(float(kl(q, p)), 4))   # 0.085 0.0817
directionvalue (bits)
kl(p, q)0.085
kl(q, p)0.0817
equal?no — asymmetric

106. Milestone 3 — the identity

Worked example

Your turn: write cross(p, q) directly and confirm it equals entropy(p) + kl(p, q) with np.isclose. Predict the value — since q is uniform it should be log₂ 3 ≈ 1.585.

Hint: cross = -np.sum(p*np.log2(q)); compare against entropy(p) + kl(p, q).

import numpy as np
def entropy(p): p = np.asarray(p, float); return -np.sum(p * np.log2(p))
def kl(p, q):   p, q = np.asarray(p, float), np.asarray(q, float); return np.sum(p * np.log2(p / q))
def cross(p, q):p, q = np.asarray(p, float), np.asarray(q, float); return -np.sum(p * np.log2(q))
p = [0.5, 0.25, 0.25]; q = [1/3, 1/3, 1/3]
print(round(float(cross(p, q)), 4))                 # 1.585
print(round(float(entropy(p) + kl(p, q)), 4))       # 1.585
print(np.isclose(cross(p, q), entropy(p) + kl(p, q)))  # True
quantityvalue (bits)
cross(p, q)1.585
entropy(p) + kl(p, q)1.585
identity holds (isclose)True

107. What each one costs: Milestone 3 — the identity

Trade off

Comparison matrix

From Milestone 3 — the identity: every row here is a choice with a cost. Fill the value (bits) column, then say which row you would actually pick and what you give up for it.

quantityvalue (bits)
cross(p, q)1.585
entropy(p) + kl(p, q)1.585
identity holds (isclose)True

108. The full program

Concept

import numpy as np

def entropy(p):  p = np.asarray(p, float); return -np.sum(p * np.log2(p))
def kl(p, q):    p, q = np.asarray(p, float), np.asarray(q, float); return np.sum(p * np.log2(p / q))
def cross(p, q): p, q = np.asarray(p, float), np.asarray(q, float); return -np.sum(p * np.log2(q))

p = np.array([0.5, 0.25, 0.25])          # weather: Sunny, Cloudy, Rainy
q = np.array([1/3, 1/3, 1/3])            # uniform guess
print('H(p)      =', round(float(entropy(p)), 4))
print('KL(p||q)  =', round(float(kl(p, q)), 4), ' KL(q||p) =', round(float(kl(q, p)), 4))
print('H(p,q)    =', round(float(cross(p, q)), 4),
      ' = H(p)+KL =', round(float(entropy(p) + kl(p, q)), 4))
printed linevalue
H(p)1.5
KL(p||q) / KL(q||p)0.085 / 0.0817
H(p,q) = H(p)+KL1.585

If H(p) reads 1.5, KL comes out asymmetric (0.085 ≠ 0.0817), and cross-entropy equals H(p)+KL = 1.585 — you've built the information-theory core of every classifier and VAE from the axioms up.

109. Fill in: value for The full program

Comparison

Comparison matrix

From The full program: refill the value column from what you know. The rest of the table is as it appeared.

printed linevalue
H(p)1.5
KL(p||q) / KL(q||p)0.085 / 0.0817
H(p,q) = H(p)+KL1.585

110. Show it off

Concept

Out loud, slides closed: explain (1) why surprise must be −log p, (2) why entropy is maximized by the uniform, (3) why minimizing cross-entropy equals minimizing KL, and (4) why the ELBO is a lower bound on log p(x).

Stretch (homework): re-derive D_KL ≥ 0 from Jensen without looking, then use the Gaussian-KL closed form to compute the VAE regularizer D_KL(N(μ, σ²) ‖ N(0,1)) = ½(σ² + μ² − 1 − ln σ²) — verified = 0.1366 nats at μ = 0.5, σ² = 0.8. This thread runs straight into the VAE (Week 44) and diffusion (Week 48).

111. Connect it up: Lesson 14: Information Theory

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Surprise & entropy · KL divergence · Why KL ≥ 0 · Cross-entropy = H(P) + KL · Mutual information · The ELBO. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

112. What you can do now

Recap

ideathe one thing to remember
entropy−Σ p log p; average surprise; max at uniform
KL≥ 0 (Jensen), asymmetric, NOT a metric
cross-entropyH(P) + KL → minimizing CE = minimizing KL
mutual informationD_KL(joint ‖ product); 0 iff independent
ELBOlog p(x) = ELBO + KL(q‖p) ≥ ELBO

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 14 (Week 5 — Information Theory) — Barron · USAAIO Round 2 Preparation, 2026
  2. Cover, Thomas — Elements of Information Theory, Ch. 2 (Entropy, Relative Entropy, Mutual Information) — Wiley, 2nd ed.
  3. Kingma, Welling — Auto-Encoding Variational Bayes (the ELBO / VAE)
  4. NumPy log2 / log
  5. Every entropy, KL, cross-entropy, MI, and Gaussian-KL value produced by real execution — numpy 2.2.6, verification run July 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108