USAAIO Lesson 29, from Week 10, fully worked. It recaps entropy from scratch, then derives mutual information I(X;Y) THREE equivalent ways, one algebra move at a time: as H(X)+H(Y)-H(X,Y), in the chain-rule form H(Y)-H(Y|X), and in the KL form D(p(x,y)||p(x)p(y)). All three are computed by hand on a single running 2x2 joint distribution and confirmed by real code. It then covers pointwise mutual information, proves the bounds I>=0 and I<=min(H), and works the trap in which correlation is zero yet the variables are dependent, on Y=X^2. It closes with mutual-information feature selection in sklearn, the InfoNCE and CLIP loss traced on a tiny batch, the link to compression, and a from-scratch mutual-information project. Every snippet runs standalone, and every number came from real execution. The lesson runs to 61 slides.
Subject: Machine Learning · 107 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 29 · Week 10 (Information Theory)
How much does knowing one variable tell you about another? We derive I(X;Y) three ways — entropy, chain rule, and KL — skipping no step, compute each by hand on one running joint, then ride it into feature selection and the InfoNCE loss behind CLIP.
Objectives
H(X) = −Σ p log₂ p and read it as bits of uncertaintyH(X)+H(Y)−H(X,Y), H(Y)−H(Y|X), and D_KL(p(x,y)‖p(x)p(y)) — and compute all three by hand to the same numberI ≥ 0, I ≤ min(H(X),H(Y)), and I = 0 ⟺ X ⊥ Y, then show a zero-correlation pair with large MIsklearn's scoresWarm-up
Discussion prompt
Before we open Lesson 29: Mutual Information & Contrastive Learning: without looking back, what was the main idea of Kernel Methods, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
feature maps and the kernel trick k(x,z)=φ(x)·φ(z), Mercer's theorem (a valid kernel has a PSD Gram matrix), the RBF kernel as an infinite-dimensional feature space, kernel composition rules, and kernel SVMs. Verify the polynomial kernel against its explicit feature map and solve XOR with an RBF SVM.
Section
Part 1 of 7 — the raw material
Intuition
Before mutual information there is entropy — how uncertain you are about a random variable, in bits. A fair coin (50/50) has 1 bit: you need exactly one yes/no answer to pin it down.
A rigged coin that lands heads 99% of the time is almost certain, so it carries less than a bit. A variable you already know carries 0 bits — nothing left to learn.
Mutual information will be entropy that two variables share — how much the uncertainty in one drops once you see the other. So we start by nailing entropy down cold.
Counterexample
Discussion prompt
Before mutual information there is entropy — how uncertain you are about a random variable, in bits. A fair coin (50/50) has 1 bit: you need exactly one yes/no answer to pin it down.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
A rigged coin that lands heads 99% of the time is almost certain, so it carries less than a bit. A variable you already know carries 0 bits — nothing left to learn.
Concept
The surprise of a single outcome with probability p is −log₂ p bits. A coin flip (p = 0.5) is −log₂ 0.5 = 1 bit of surprise; a 1-in-8 event is −log₂(1/8) = 3 bits — three times as surprising.
\[ \text{surprise}(x) = -\log_2 p(x) \]
Entropy will just be the average surprise over all outcomes. Rare things surprise more, certain things (p = 1) surprise not at all — −log₂ 1 = 0. Keep this: MI counts how much one variable's surprise drops once you see the other.
Concept
For a discrete variable X with outcome probabilities p(x), entropy is the expected surprise −log₂ p(x), averaged over outcomes:
\[ H(X) = -\sum_{x} p(x)\,\log_2 p(x) \]
Rare outcomes (small p) are surprising (large −log₂ p); certain outcomes (p = 1) carry zero surprise. Using log₂ makes the unit bits. We adopt the convention 0·log 0 = 0.
Analogy
Discussion prompt
Explain The entropy formula by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
For a discrete variable X with outcome probabilities p(x), entropy is the expected surprise −log₂ p(x), averaged over outcomes:
Estimation
Predict first
We mask out zero probabilities (log 0 is undefined but 0·log 0 = 0), then sum. Run this standalone:
Commit before you compute: what does Entropy in code — three sanity checks come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: The mask p[p > 0] is the one detail people forget
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. A distribution with a zero-probability outcome would otherwise feed log2(0) = −inf into the sum and poison it with NaN.
Worked example
We mask out zero probabilities (log 0 is undefined but 0·log 0 = 0), then sum. Run this standalone:
import numpy as np
def entropy(p):
p = np.asarray(p, dtype=float)
p = p[p > 0] # 0 log 0 = 0, so drop zeros
return -np.sum(p * np.log2(p)) # bits
print(entropy([0.5, 0.5])) # 1.0
print(entropy([0.4, 0.1, 0.1, 0.4])) # 1.7219...
print(entropy([1.0])) # 0.0 (no uncertainty)The mask p[p > 0] is the one detail people forget
Why: A distribution with a zero-probability outcome would otherwise feed log2(0) = −inf into the sum and poison it with NaN. Dropping zeros implements the 0·log 0 = 0 convention exactly.
| distribution | H (bits) | meaning |
|---|---|---|
| [0.5, 0.5] | 1.0 | fair coin — max uncertainty for 2 outcomes |
| [0.4, 0.1, 0.1, 0.4] | 1.7219 | our joint's 4 cells (used all lesson) |
| [1.0] | 0.0 | a sure thing — no uncertainty |
Comparison
Comparison matrix
From Entropy in code — three sanity checks: refill the H (bits) column from what you know. The rest of the table is as it appeared.
| distribution | H (bits) | meaning |
|---|---|---|
| [0.5, 0.5] | 1.0 | fair coin — max uncertainty for 2 outcomes |
| [0.4, 0.1, 0.1, 0.4] | 1.7219 | our joint's 4 cells (used all lesson) |
| [1.0] | 0.0 | a sure thing — no uncertainty |
Missing information
Discussion prompt
Sweep a coin's head-probability q from fair to certain and watch entropy fall. Uniform is maximally uncertain; a rigged coin is more predictable, so it carries fewer bits:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
A fair coin (q = 0.5) is the hardest to predict → 1 full bit. As q → 1 the outcome becomes certain and entropy → 0. Entropy is maximized by the uniform distribution — a fact we lean on for the MI bounds later.
Worked example
Sweep a coin's head-probability q from fair to certain and watch entropy fall. Uniform is maximally uncertain; a rigged coin is more predictable, so it carries fewer bits:
import numpy as np
def entropy(p):
p = np.asarray(p, dtype=float); p = p[p > 0]
return -np.sum(p * np.log2(p))
for q in [0.5, 0.75, 0.9, 0.99, 1.0]:
print(q, round(entropy([q, 1-q]), 4))H is 1.0 at q = 0.5 and slides to 0 at q = 1.0
Why: A fair coin (q = 0.5) is the hardest to predict → 1 full bit. As q → 1 the outcome becomes certain and entropy → 0. Entropy is maximized by the uniform distribution — a fact we lean on for the MI bounds later.
| q (P heads) | H([q, 1−q]) | reading |
|---|---|---|
| 0.50 | 1.0000 | fair — maximum uncertainty |
| 0.75 | 0.8113 | leaning heads |
| 0.90 | 0.4690 | fairly predictable |
| 0.99 | 0.0808 | almost certain |
| 1.00 | 0.0000 | a sure thing |
Trade off
Comparison matrix
From Entropy peaks at the fair coin: every row here is a choice with a cost. Fill the H([q, 1−q]) column, then say which row you would actually pick and what you give up for it.
| q (P heads) | H([q, 1−q]) | reading |
|---|---|---|
| 0.50 | 1.0000 | fair — maximum uncertainty |
| 0.75 | 0.8113 | leaning heads |
| 0.90 | 0.4690 | fairly predictable |
| 0.99 | 0.0808 | almost certain |
| 1.00 | 0.0000 | a sure thing |
Concept
One dataset carries the whole lesson. X is the weather (0 = dry, 1 = rain); Y is whether a commuter carries an umbrella (0 = no, 1 = yes). Their joint distribution p(x,y):
| p(x,y) | Y=0 (no umbrella) | Y=1 (umbrella) |
|---|---|---|
| X=0 (dry) | 0.40 | 0.10 |
| X=1 (rain) | 0.10 | 0.40 |
The mass sits on the diagonal: dry-and-no-umbrella and rain-and-umbrella are common; the off-diagonal mismatches are rare. So the two are related — but not perfectly. Quantifying exactly how related is the job of mutual information.
Explain it
Discussion prompt
Explain The running example: weather × umbrella to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
One dataset carries the whole lesson. X is the weather (0 = dry, 1 = rain); Y is whether a commuter carries an umbrella (0 = no, 1 = yes). Their joint distribution p(x,y):
Pattern
Predict first
The table runs: px = p(X) | [0.5, 0.5] | 0.4+0.1, 0.1+0.4 · py = p(Y) | [0.5, 0.5] | 0.4+0.1, 0.1+0.4
In Marginals: collapse the joint, given the rows so far: what is the next one — the row where quantity is Σ all cells?
Correct: Σ all cells | 1.0 | valid distribution
| quantity | value | how |
|---|---|---|
| px = p(X) | [0.5, 0.5] | 0.4+0.1, 0.1+0.4 |
| py = p(Y) | [0.5, 0.5] | 0.4+0.1, 0.1+0.4 |
| Σ all cells | 1.0 | valid distribution |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. Remember it by what SURVIVES: summing axis=1 keeps the row index x.
Worked example
The marginal p(x) sums each row (over all y); p(y) sums each column (over all x). Sum over the axis you want to erase:
import numpy as np
J = np.array([[0.4, 0.1],
[0.1, 0.4]]) # rows = X, cols = Y
px = J.sum(axis=1) # marginal of X: sum over Y
py = J.sum(axis=0) # marginal of Y: sum over X
print("px =", px) # [0.5 0.5]
print("py =", py) # [0.5 0.5]
print("sum J =", J.sum()) # 1.0axis=1 sums over columns (erases Y) → p(x); axis=0 erases X → p(y)
Why: Remember it by what SURVIVES: summing axis=1 keeps the row index x. Here both marginals come out [0.5, 0.5] — dry vs rain is 50/50, and so is umbrella vs not.
| quantity | value | how |
|---|---|---|
| px = p(X) | [0.5, 0.5] | 0.4+0.1, 0.1+0.4 |
| py = p(Y) | [0.5, 0.5] | 0.4+0.1, 0.1+0.4 |
| Σ all cells | 1.0 | valid distribution |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
axis=1 sums over columns (erases Y) → p(x); axis=0 erases X → p(y)
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
The marginal p(x) sums each row (over all y); p(y) sums each column (over all x). Sum over the axis you want to erase:
Fill the middle
Fill in the blanks
From Compute H(X), H(Y), H(X,Y) by hand — finish the line. Write what belongs on the right of the equals sign before you look.
H(X,Y) = -2\big(0.4\log_2 0.4\big) - 2\big(0.1\log_2 0.1\big)
Why: Producing the right-hand side unprompted is the difference between recognising this line and being able to use it. −0.4·log₂0.4 = 0.5288 and −0.1·log₂0.1 = 0.3322 (since log₂0.4 = −1.3219, log₂0.1 = −3.3219).
Worked example
Three entropies feed every MI formula. Marginals are [0.5, 0.5], so both give 1 bit:
\[ H(X) = -2\big(0.5\log_2 0.5\big) = -2(0.5)(-1) = 1 \text{ bit} \]
The joint entropy uses all four cells [0.4, 0.1, 0.1, 0.4]:
\[ H(X,Y) = -2\big(0.4\log_2 0.4\big) - 2\big(0.1\log_2 0.1\big) \]
= 2(0.5288) + 2(0.3322) = 1.0575 + 0.6644 = 1.7219 bits
Why: −0.4·log₂0.4 = 0.5288 and −0.1·log₂0.1 = 0.3322 (since log₂0.4 = −1.3219, log₂0.1 = −3.3219). Verified by real execution: H(X,Y) = 1.721928.
| entropy | value (bits) | from |
|---|---|---|
| H(X) | 1.0000 | marginal [0.5, 0.5] |
| H(Y) | 1.0000 | marginal [0.5, 0.5] |
| H(X,Y) | 1.7219 | all 4 cells [0.4,0.1,0.1,0.4] |
Concept
Notice H(X) + H(Y) = 2.0 but H(X,Y) = 1.7219. The joint is less uncertain than the two separately. If X and Y were independent they'd add to exactly 2.0; the shortfall is real structure.
That gap — 2.0 − 1.7219 = 0.2781 bits — is precisely the information they share. Hold that number: it is the mutual information, and the next part earns it three different ways.
Section
Part 2 of 7 — one quantity, three derivations
Picture it
Figure (svg): Two overlapping circles labeled H(X) and H(Y); the left-only crescent is H(X|Y), the right-only crescent is H(Y|X), and the lens-shaped overlap in the middle is labeled I(X;Y).
Discussion prompt
Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.
Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.
Answer:
Picture two overlapping circles: H(X) is the left circle, H(Y) the right. Where they overlap is information the two variables have in common — see one, and you learn something about the other.
Intuition
Picture two overlapping circles: H(X) is the left circle, H(Y) the right. Where they overlap is information the two variables have in common — see one, and you learn something about the other.
That overlap is mutual information I(X;Y). If the circles are disjoint (no overlap) the variables are independent and I = 0. If one circle sits entirely inside the other, knowing one fully determines the shared part.
Figure (svg): Two overlapping circles labeled H(X) and H(Y); the left-only crescent is H(X|Y), the right-only crescent is H(Y|X), and the lens-shaped overlap in the middle is labeled I(X;Y).
Concept
Mutual information is one number wearing three algebraic costumes. Each will be derived in full — but here they are up front so you know the destination:
\[ I(X;Y) = H(X) + H(Y) - H(X,Y) \]
\[ \;\;= H(Y) - H(Y\mid X) \]
\[ \;\;= D_{KL}\big(p(x,y)\,\|\,p(x)p(y)\big) \]
Form 1 is the set overlap. Form 2 is uncertainty reduced by observing X. Form 3 is the distance from independence. They are equal — provably, not coincidentally.
Concept
The ground-truth definition of MI is the expected pointwise log-ratio between the joint and the product of marginals:
\[ I(X;Y) = \sum_{x,y} p(x,y)\,\log_2 \frac{p(x,y)}{p(x)\,p(y)} \]
Everything else in this part is this expression, re-grouped. It is already Form 3 (a KL divergence). Watch it collapse into Forms 1 and 2 with nothing but log rules.
Ranking
Put in order
Put the moves of Derivation A → Form 1 (entropies) into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. The expected log-ratio of joint to product of marginals — our single source of truth.
Worked example
Start from the definition
Why: The expected log-ratio of joint to product of marginals — our single source of truth.
\[ I(X;Y) = \sum_{x,y} p(x,y)\,\log_2 \frac{p(x,y)}{p(x)p(y)} \]
Split the log of a quotient into a difference
Why: log(a/b) = log a − log b. The single log becomes log p(x,y) minus log[p(x)p(y)].
\[ = \sum_{x,y} p(x,y)\Big[\log_2 p(x,y) - \log_2 p(x)p(y)\Big] \]
Split log p(x)p(y) into log p(x) + log p(y)
Why: log of a product is a sum of logs. Now three log terms, each weighted by p(x,y).
\[ = \sum_{x,y} p(x,y)\log_2 p(x,y) - \sum_{x,y} p(x,y)\log_2 p(x) - \sum_{x,y} p(x,y)\log_2 p(y) \]
Notation
Annotate
From Derivation A → Form 1 (entropies) — read this one piece at a time. What is each part doing?
On: \( I(X;Y) = \sum_{x,y} p(x,y)\,\log_2 \frac{p(x,y)}{p(x)p(y)} \)
Step zero
Discussion prompt
Derivation A → Form 1 (finish) — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Term 1 is −H(X,Y)
Answer:
Worked example
Term 1 is −H(X,Y)
Why: Σ p(x,y) log₂ p(x,y) is exactly the negative joint entropy, by definition of H(X,Y).
\[ \sum_{x,y} p(x,y)\log_2 p(x,y) = -H(X,Y) \]
In term 2, sum out y first
Why: log₂ p(x) doesn't depend on y, so Σ_y p(x,y) collapses to the marginal p(x). The double sum becomes a single sum over x.
\[ \sum_{x,y} p(x,y)\log_2 p(x) = \sum_x p(x)\log_2 p(x) = -H(X) \]
Term 3 is −H(Y) by the same collapse over x
Why: Symmetric: summing out x turns p(x,y) into p(y).
\[ \sum_{x,y} p(x,y)\log_2 p(y) = -H(Y) \]
Assemble: −(−H(X,Y)) picks up the leading minus signs
Why: I = −H(X,Y) − (−H(X)) − (−H(Y)) = H(X) + H(Y) − H(X,Y). Form 1, derived with only log rules and marginalization.
\[ \boxed{\,I(X;Y) = H(X) + H(Y) - H(X,Y)\,} \]
Translation
\( \boxed{\,I(X;Y) = H(X) + H(Y) - H(X,Y)\,} \)
Draw it
Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.
Fill the middle
Fill in the blanks
From Form 1 numerically — it's the gap we spotted — one line has had its right-hand side removed. Put it back.
import numpy as np
def entropy(p):
p = np.asarray(p, dtype=float); p = p[p > 0]
return -np.sum(p * np.log2(p))
J = np.array([[0.4, 0.1], [0.1, 0.4]])
px, py = J.sum(1), J.sum(0)
HX = entropy(px) # 1.0
HY = entropy(py) # 1.0
HXY = entropy(J.ravel()) # 1.7219
mi = HX + HY - HXY
print(HX, HY, round(HXY, 4))
print("I(X;Y) =", round(mi, 4)) # 0.2781
Why: HY is what everything below it consumes, so the wrong expression here fails later and somewhere else. The two circles sum to 2.0 bits; the joint is only 1.7219 bits; the 0.2781-bit shortfall is the shared information.
Worked example
Plug our three entropies straight into Form 1. This is the 0.2781 we flagged in Part 1, now earned:
import numpy as np
def entropy(p):
p = np.asarray(p, dtype=float); p = p[p > 0]
return -np.sum(p * np.log2(p))
J = np.array([[0.4, 0.1], [0.1, 0.4]])
px, py = J.sum(1), J.sum(0)
HX = entropy(px) # 1.0
HY = entropy(py) # 1.0
HXY = entropy(J.ravel()) # 1.7219
mi = HX + HY - HXY
print(HX, HY, round(HXY, 4))
print("I(X;Y) =", round(mi, 4)) # 0.2781I = 1.0 + 1.0 − 1.7219 = 0.2781 bits
Why: The two circles sum to 2.0 bits; the joint is only 1.7219 bits; the 0.2781-bit shortfall is the shared information. Small but nonzero — weather and umbrella are related but not locked together.
| piece | value (bits) |
|---|---|
| H(X) + H(Y) | 2.0000 |
| H(X,Y) | 1.7219 |
| I(X;Y) = difference | 0.2781 |
Section
Part 3 of 7 — uncertainty reduced
Concept
H(Y|X) is the average uncertainty left in Y after you already know X — the leftover surprise once the weather is revealed. It is defined by the chain rule of entropy:
\[ H(Y\mid X) = H(X,Y) - H(X) \]
Read it as: the total joint uncertainty, minus the part X already accounts for. Whatever remains is the uncertainty in Y that X couldn't explain.
Ranking
Put in order
Put the moves of Derivation B → Form 2 into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. We just proved I = H(X) + H(Y) − H(X,Y).
Worked example
Start from Form 1
Why: We just proved I = H(X) + H(Y) − H(X,Y). Reuse it.
\[ I(X;Y) = H(X) + H(Y) - H(X,Y) \]
Regroup: pull H(Y) out front, keep H(X) − H(X,Y) together
Why: Pure rearrangement of the same three terms — associativity of addition, nothing added or dropped.
\[ = H(Y) + \big(H(X) - H(X,Y)\big) \]
Recognize H(X,Y) − H(X) = H(Y|X), so H(X) − H(X,Y) = −H(Y|X)
Why: The chain rule H(X,Y) = H(X) + H(Y|X) rearranges to H(Y|X) = H(X,Y) − H(X). Flip the sign to match our grouping.
\[ = H(Y) - \big(H(X,Y) - H(X)\big) = H(Y) - H(Y\mid X) \]
Form 2 established
Why: MI is how many bits of uncertainty about Y vanish once X is known. Same quantity, a genuinely different reading.
\[ \boxed{\,I(X;Y) = H(Y) - H(Y\mid X)\,} \]
Blank canvas
Draw it
Draw what Derivation B → Form 2 just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Worked example
Compute H(Y|X) = H(X,Y) − H(X), then subtract from H(Y). Must land on 0.2781 again:
import numpy as np
def entropy(p):
p = np.asarray(p, dtype=float); p = p[p > 0]
return -np.sum(p * np.log2(p))
J = np.array([[0.4, 0.1], [0.1, 0.4]])
px, py = J.sum(1), J.sum(0)
HY = entropy(py) # 1.0
HYgX = entropy(J.ravel()) - entropy(px) # H(X,Y) - H(X) = H(Y|X)
print("H(Y) =", round(HY, 4)) # 1.0
print("H(Y|X) =", round(HYgX, 4)) # 0.7219
print("I(X;Y) = H(Y) - H(Y|X) =", round(HY - HYgX, 4)) # 0.2781H(Y|X) = 1.7219 − 1.0 = 0.7219; I = 1.0 − 0.7219 = 0.2781
Why: Knowing the weather cuts umbrella-uncertainty from 1.0 bit down to 0.7219 bits — a drop of exactly 0.2781 bits. Same MI as Form 1, confirming the two derivations meet.
| quantity | value (bits) | reading |
|---|---|---|
| H(Y) | 1.0000 | umbrella uncertainty, knowing nothing |
| H(Y|X) | 0.7219 | umbrella uncertainty, knowing weather |
| I(X;Y) = drop | 0.2781 | bits weather tells you about umbrella |
Sorting
Sort into buckets
These are the pieces of Lesson 29: Mutual Information & Contrastive Learning, out of order. Put each one back under the part of the lesson it belongs to.
Concept
Form 1 is untouched if you swap X and Y — so I(X;Y) = I(Y;X). By the same algebra, MI also equals H(X) − H(X|Y). Weather tells you as much about umbrellas as umbrellas tell you about weather.
| direction | formula | value (bits) |
|---|---|---|
| Y from X | H(Y) − H(Y|X) | 1.0 − 0.7219 = 0.2781 |
| X from Y | H(X) − H(X|Y) | 1.0 − 0.7219 = 0.2781 |
Verified: H(X|Y) = H(X,Y) − H(Y) = 0.7219, so H(X) − H(X|Y) = 0.2781 — identical. Information is a shared quantity, not a one-way arrow.
Section
Part 4 of 7 — the KL view
Concept
From Lesson 14, the KL divergence measures how far one distribution q is from another p, in bits:
\[ D_{KL}(p \,\|\, q) = \sum_i p_i \log_2 \frac{p_i}{q_i} \;\ge\; 0 \]
It is 0 exactly when p = q, and strictly positive otherwise. MI's definition is a KL — between the true joint p(x,y) and the pretend-independent product p(x)p(y).
Ranking
Put in order
Put the moves of Derivation C → Form 3 (it's immediate) into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Set p = the true joint p(x,y) and q = the independent product p(x)p(y) in the KL definition.
Worked example
Write the KL of joint against product
Why: Set p = the true joint p(x,y) and q = the independent product p(x)p(y) in the KL definition.
\[ D_{KL}\big(p(x,y)\,\|\,p(x)p(y)\big) = \sum_{x,y} p(x,y)\,\log_2 \frac{p(x,y)}{p(x)p(y)} \]
That is the definition of I(X;Y) verbatim
Why: Compare to our source-of-truth expression from Part 2 — character for character the same sum. No manipulation needed.
\[ \boxed{\,I(X;Y) = D_{KL}\big(p(x,y)\,\|\,p(x)p(y)\big)\,} \]
So MI = 'how far the joint is from independent'
Why: This is the most useful mental model: MI literally measures the gap between reality and the world where X and Y ignore each other. Zero gap ⇒ independent.
Notation
Annotate
From Derivation C → Form 3 (it's immediate) — read this one piece at a time. What is each part doing?
On: \( \boxed{\,I(X;Y) = D_{KL}\big(p(x,y)\,\|\,p(x)p(y)\big)\,} \)
Estimation
Predict first
Build the product p(x)p(y) with an outer product, then sum the weighted log-ratio. Third road, same 0.2781:
Commit before you compute: what does Form 3 numerically come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: np.outer(px, py) builds the 'if-independent' joint [[0.25]×4]
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. With both marginals [0.5, 0.5], independence would put 0.25 in every cell.
Worked example
Build the product p(x)p(y) with an outer product, then sum the weighted log-ratio. Third road, same 0.2781:
import numpy as np
J = np.array([[0.4, 0.1], [0.1, 0.4]])
px, py = J.sum(1), J.sum(0)
prod = np.outer(px, py) # p(x)p(y): the independent joint
mi_kl = np.sum(J * np.log2(J / prod)) # KL(joint || product)
print("outer(px,py) =\n", prod)
print("I(X;Y) as KL =", round(mi_kl, 4)) # 0.2781np.outer(px, py) builds the 'if-independent' joint [[0.25]×4]
Why: With both marginals [0.5, 0.5], independence would put 0.25 in every cell. The real joint has 0.40/0.10 instead — that mismatch is what KL measures.
| cell | real p(x,y) | if independent | ratio |
|---|---|---|---|
| (dry, no) | 0.40 | 0.25 | 1.60 |
| (dry, yes) | 0.10 | 0.25 | 0.40 |
| (rain, no) | 0.10 | 0.25 | 0.40 |
| (rain, yes) | 0.40 | 0.25 | 1.60 |
Discrimination
Sort into buckets
Sort these by real p(x,y), from memory, without looking back at Form 3 numerically. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Concept
Inside the sum, each cell contributes p(x,y)·log₂[p(x,y)/(p(x)p(y))]. The bare log-ratio log₂[p(x,y)/(p(x)p(y))] is the pointwise MI (PMI) of one outcome pair — positive when the pair co-occurs more than chance, negative when less.
| pair | PMI (bits) | contribution | reading |
|---|---|---|---|
| (dry, no) | +0.6781 | +0.27123 | co-occur more than chance |
| (dry, yes) | −1.3219 | −0.13219 | rarer than chance |
| (rain, no) | −1.3219 | −0.13219 | rarer than chance |
| (rain, yes) | +0.6781 | +0.27123 | co-occur more than chance |
Sum the contributions: 0.27123 − 0.13219 − 0.13219 + 0.27123 = 0.2781. Individual PMIs can be negative, but their probability-weighted sum — the MI — is always ≥ 0. That guarantee is next.
Discrimination
Sort into buckets
Sort these by PMI (bits), from memory, without looking back at Pointwise mutual information. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Fill the middle
Fill in the blanks
From All three forms agree — the master check — one line has had its right-hand side removed. Put it back.
import numpy as np
def entropy(p):
p = np.asarray(p, dtype=float); p = p[p > 0]
return -np.sum(p * np.log2(p))
J = np.array([[0.4, 0.1], [0.1, 0.4]])
px, py = J.sum(1), J.sum(0)
form1 = entropy(px) + entropy(py) - entropy(J.ravel())
form2 = entropy(py) - (entropy(J.ravel()) - entropy(px))
form3 = np.sum(J * np.log2(J / np.outer(px, py)))
print(round(form1, 6), round(form2, 6), round(form3, 6))
Why: J is what everything below it consumes, so the wrong expression here fails later and somewhere else. Not a numerical coincidence: the three are algebraically the same object, and we proved each transformation.
Worked example
Three independent derivations, one dataset, one number. This is the payoff slide — verified by execution:
import numpy as np
def entropy(p):
p = np.asarray(p, dtype=float); p = p[p > 0]
return -np.sum(p * np.log2(p))
J = np.array([[0.4, 0.1], [0.1, 0.4]])
px, py = J.sum(1), J.sum(0)
form1 = entropy(px) + entropy(py) - entropy(J.ravel())
form2 = entropy(py) - (entropy(J.ravel()) - entropy(px))
form3 = np.sum(J * np.log2(J / np.outer(px, py)))
print(round(form1, 6), round(form2, 6), round(form3, 6))0.278072 0.278072 0.278072 — identical to 6 places
Why: Not a numerical coincidence: the three are algebraically the same object, and we proved each transformation. This triple-agreement is your always-available correctness check when you implement MI.
| form | expression | value (bits) |
|---|---|---|
| 1 — entropies | H(X)+H(Y)−H(X,Y) | 0.278072 |
| 2 — conditional | H(Y)−H(Y|X) | 0.278072 |
| 3 — KL | D(p(x,y) ‖ p(x)p(y)) | 0.278072 |
Section
Part 5 of 7 — I ≥ 0, I ≤ min(H), I = 0 ⟺ ⊥
Concept
Because MI is a KL divergence and KL ≥ 0 always (Gibbs' inequality, Lesson 14), mutual information is non-negative:
\[ I(X;Y) = D_{KL}\big(p(x,y)\,\|\,p(x)p(y)\big) \;\ge\; 0 \]
You can never lose information by observing a variable — at worst it's useless (I = 0). This is why individual PMIs may go negative but the average cannot.
Concept
KL is 0 iff its two arguments are equal. So I(X;Y) = 0 iff p(x,y) = p(x)p(y) for every cell — which is the definition of independence:
\[ I(X;Y) = 0 \;\Longleftrightarrow\; p(x,y) = p(x)p(y)\;\;\forall x,y \;\Longleftrightarrow\; X \perp Y \]
This is the strongest independence test there is: it certifies no dependence of any kind, linear or not. Contrast that with correlation, which only sees straight lines.
Worked example
Force independence by using the product p(x)p(y) as the joint. Its MI must be 0 — the KL of a distribution against itself:
import numpy as np
J = np.array([[0.4, 0.1], [0.1, 0.4]])
px, py = J.sum(1), J.sum(0)
Ind = np.outer(px, py) # force independence
ipx, ipy = Ind.sum(1), Ind.sum(0)
mi = np.sum(Ind * np.log2(Ind / np.outer(ipx, ipy)))
print("Ind =\n", Ind)
print("I(X;Y) for the product =", round(mi, 6)) # 0.0Every cell of Ind equals its own p(x)p(y), so each log-ratio is log(1) = 0
Why: Ind[i,k] = px[i]·py[k] by construction, and its marginals reproduce px, py, so Ind/outer(marginals) = 1 everywhere. log₂(1) = 0, sum = 0.
| joint used | MI (bits) | verdict |
|---|---|---|
| real J (diagonal-heavy) | 0.2781 | dependent |
| outer(px, py) | 0.000000 | independent — MI vanishes |
Comparison
Comparison matrix
From An independent joint scores exactly 0: refill the verdict column from what you know. The rest of the table is as it appeared.
| joint used | MI (bits) | verdict |
|---|---|---|
| real J (diagonal-heavy) | 0.2781 | dependent |
| outer(px, py) | 0.000000 | independent — MI vanishes |
Concept
Since H(Y|X) ≥ 0, Form 2 gives I = H(Y) − H(Y|X) ≤ H(Y). By symmetry I ≤ H(X) too. So MI is squeezed between 0 and the smaller of the two entropies:
\[ 0 \;\le\; I(X;Y) \;\le\; \min\big(H(X),\,H(Y)\big) \]
For our joint: I = 0.2781 ≤ min(1.0, 1.0) = 1.0 ✓. MI hits its ceiling H(Y) only when X fully determines Y (H(Y|X) = 0) — perfect dependence.
Anomaly
Predict first
A student writes this, and it looks reasonable:
The correlation between X and Y is ≈ 0, so they carry no information about each other — I(X;Y) = 0.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Correlation only measures LINEAR association.
Correlation ≈ 0 rules out only a linear relationship. To test real independence, compute the mutual information.
Why: Correlation only measures LINEAR association. It is blind to any symmetric curve, so a zero correlation says nothing about nonlinear dependence.
Trap
The correlation between X and Y is ≈ 0, so they carry no information about each other — I(X;Y) = 0.
Rule out dependence using the correlation coefficient
Why: Correlation only measures LINEAR association. It is blind to any symmetric curve, so a zero correlation says nothing about nonlinear dependence.
Conclude X and Y are independent
Why: False. Y = X² with X symmetric about 0 has correlation ≈ 0 yet Y is FULLY determined by X — maximal dependence hiding behind a zero correlation.
Correlation ≈ 0 rules out only a linear relationship. To test real independence, compute the mutual information.
Use I(X;Y) — it captures ANY dependence
Why: MI = 0 iff truly independent, full stop. It sees the parabola that correlation misses, so it is the honest independence test.
For Y = X²: correlation ≈ 0 but I = 0.9183 bits (large)
Why: The MI equals H(Y) here — X pins Y down completely. Next slide computes both numbers so you see the gap with your own eyes.
Break the constraint
Discussion prompt
The rule this trap just fixed:
MI = 0 iff truly independent, full stop. It sees the parabola that correlation misses, so it is the honest independence test.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
Correlation only measures LINEAR association. It is blind to any symmetric curve, so a zero correlation says nothing about nonlinear dependence.
Fill the middle
Fill in the blanks
From Y = X²: correlation ≈ 0, MI large — one line has had its right-hand side removed. Put it back.
import numpy as np
# X uniform on 1/3; Y = X^2 in ___. Fully determined, yet...
Jd = np.zeros((3, 2))
Jd[0, 1] = 1/3 # X=-1 -> Y=1
Jd[1, 0] = 1/3 # X= 0 -> Y=0
Jd[2, 1] = ___ # X= 1 -> Y=1
px, py = Jd.sum(1), Jd.sum(0)
m = Jd > 0
mi = np.sum(Jd[m] * np.log2(Jd[m] / np.outer(px, py)[m]))
xs = np.random.default_rng(0).uniform(-1, 1, 100000)
corr = np.corrcoef(xs, xs**2)[0, 1]
print("correlation(X, X^2) =", round(corr, 4)) # ~0
print("I(X; X^2) =", round(mi, 4), "bits") # 0.9183
Why: Jd[2, 1] is what everything below it consumes, so the wrong expression here fails later and somewhere else. MI equals H(Y) = H([1/3, 2/3]) = 0.9183 — the maximum, since X determines Y outright.
Worked example
Let X be uniform on {−1, 0, 1} and Y = X². Y is a deterministic function of X, so they're maximally dependent — but the parabola is symmetric, so correlation collapses:
import numpy as np
# X uniform on {-1, 0, 1}; Y = X^2 in {0, 1}. Fully determined, yet...
Jd = np.zeros((3, 2))
Jd[0, 1] = 1/3 # X=-1 -> Y=1
Jd[1, 0] = 1/3 # X= 0 -> Y=0
Jd[2, 1] = 1/3 # X= 1 -> Y=1
px, py = Jd.sum(1), Jd.sum(0)
m = Jd > 0
mi = np.sum(Jd[m] * np.log2(Jd[m] / np.outer(px, py)[m]))
xs = np.random.default_rng(0).uniform(-1, 1, 100000)
corr = np.corrcoef(xs, xs**2)[0, 1]
print("correlation(X, X^2) =", round(corr, 4)) # ~0
print("I(X; X^2) =", round(mi, 4), "bits") # 0.9183correlation ≈ 0 (about −0.006), yet I(X; X²) = 0.9183 bits
Why: MI equals H(Y) = H([1/3, 2/3]) = 0.9183 — the maximum, since X determines Y outright. Correlation is fooled by symmetry; MI is not. This is the exam's favorite gotcha.
| measure | value | sees nonlinear dependence? |
|---|---|---|
| correlation(X, X²) | ≈ 0 | no — blind to the parabola |
| I(X; X²) | 0.9183 bits | yes — flags full dependence |
| H(Y) (the ceiling) | 0.9183 bits | MI is maxed out |
Section
Part 6 of 7 — rank features by MI
Concept
A feature is useful when knowing it reduces uncertainty about the label — exactly I(feature; label). Rank features by their MI with y and keep the top ones; the rest are noise.
Unlike a correlation filter, an MI filter catches nonlinear predictors (the X²-style features a linear screen throws away). sklearn.feature_selection.mutual_info_classif estimates I for continuous features against a discrete label.
Pattern
Predict first
The table runs: 0 | 0.574 | y + noise | informative · 1 | 0.011 | pure noise | noise · 2 | 0.688 | 2y−1 + noise | informative · 3 | 0.004 | pure noise | noise
In MI finds the signal features, given the rows so far: what is the next one — the row where feature is 4?
Correct: 4 | 0.000 | pure noise | noise
| feature | MI score | built as | verdict |
|---|---|---|---|
| 0 | 0.574 | y + noise | informative |
| 1 | 0.011 | pure noise | noise |
| 2 | 0.688 | 2y−1 + noise | informative |
| 3 | 0.004 | pure noise | noise |
| 4 | 0.000 | pure noise | noise |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. The two features built from y light up; the three np.random noise columns land near zero.
Worked example
Build 5 features where only columns 0 and 2 carry signal about a binary label; the other three are pure noise. Let MI expose them:
import numpy as np
from sklearn.feature_selection import mutual_info_classif
rng = np.random.default_rng(0)
n = 500
y = rng.integers(0, 2, n)
f0 = y + 0.3*rng.normal(size=n) # informative
f2 = 2*y - 1 + 0.3*rng.normal(size=n) # informative
X = np.c_[f0, rng.normal(size=n), f2, rng.normal(size=(n, 2))]
scores = mutual_info_classif(X, y, random_state=0)
print(scores.round(3)) # [0.574 0.011 0.688 0.004 0.]Signal features 0 and 2 score ~0.57 and ~0.69; the rest score ~0
Why: The two features built from y light up; the three np.random noise columns land near zero. MI cleanly separates informative from useless — no linearity assumed. (nats here, since sklearn uses natural log — the ranking is what matters.)
| feature | MI score | built as | verdict |
|---|---|---|---|
| 0 | 0.574 | y + noise | informative |
| 1 | 0.011 | pure noise | noise |
| 2 | 0.688 | 2y−1 + noise | informative |
| 3 | 0.004 | pure noise | noise |
| 4 | 0.000 | pure noise | noise |
Pattern
Step through it
Step through MI finds the signal features one row at a time. What is driving the change, and what would the row after the last one be?
Ranking
Put in order
These are the steps of The mutual-information toolkit, scrambled. Put them back in order before the next slide shows you.
I(X;Y) = Σ p(x,y) log₂[p(x,y)/(p(x)p(y))] — the master expressionH(X)+H(Y)−H(X,Y) — the set-overlapH(Y)−H(Y|X) — uncertainty reduced by seeing XD_KL(p(x,y)‖p(x)p(y)) — distance from independence0 ≤ I ≤ min(H(X),H(Y)); I = 0 ⟺ X ⊥ Y (catches nonlinear, unlike correlation)I with the label; maximize a lower bound on I via InfoNCE (CLIP)Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
I(X;Y) = Σ p(x,y) log₂[p(x,y)/(p(x)p(y))] — the master expressionH(X)+H(Y)−H(X,Y) — the set-overlapH(Y)−H(Y|X) — uncertainty reduced by seeing XD_KL(p(x,y)‖p(x)p(y)) — distance from independence0 ≤ I ≤ min(H(X),H(Y)); I = 0 ⟺ X ⊥ Y (catches nonlinear, unlike correlation)I with the label; maximize a lower bound on I via InfoNCE (CLIP)Edge cases
Discussion prompt
The mutual-information toolkit works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
I(X;Y) = Σ p(x,y) log₂[p(x,y)/(p(x)p(y))] — the master expressionH(X)+H(Y)−H(X,Y) — the set-overlapH(Y)−H(Y|X) — uncertainty reduced by seeing XD_KL(p(x,y)‖p(x)p(y)) — distance from independence0 ≤ I ≤ min(H(X),H(Y)); I = 0 ⟺ X ⊥ Y (catches nonlinear, unlike correlation)I with the label; maximize a lower bound on I via InfoNCE (CLIP)Check
Interpret a zero. Remember MI is a KL divergence.
Check your understanding
I(X;Y) = 0 implies:
Answer: A
Why: I = 0 means the KL between the joint and the product of marginals is zero, so p(x,y) = p(x)p(y) everywhere — exact independence. MI captures all dependence, so zero MI is the strongest independence statement.
Elimination
Eliminate the wrong options
Y = X² with X symmetric about 0. Compared with correlation, mutual information will:
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Correlation measures only linear association, which cancels to ≈ 0 for the symmetric parabola. But Y is fully determined by X, so I(X;Y) = H(Y) is large (0.9183 bits in our demo). MI sees the nonlinear dependence correlation misses.
Check
Which measure sees nonlinear structure?
Check your understanding
Y = X² with X symmetric about 0. Compared with correlation, mutual information will:
Answer: A
Why: Correlation measures only linear association, which cancels to ≈ 0 for the symmetric parabola. But Y is fully determined by X, so I(X;Y) = H(Y) is large (0.9183 bits in our demo). MI sees the nonlinear dependence correlation misses.
Check
Relate MI to conditional entropy — Form 2.
Check your understanding
Which identity is correct?
Answer: A
Why: MI is the reduction in uncertainty about Y from knowing X: start with H(Y), subtract what remains, H(Y|X). We derived it from Form 1 by regrouping. Symmetrically it also equals H(X) − H(X|Y).
Elimination
Eliminate the wrong options
What is the tightest general upper bound on I(X;Y)?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: From I = H(Y) − H(Y|X) with H(Y|X) ≥ 0, we get I ≤ H(Y); by symmetry I ≤ H(X). So I ≤ min(H(X), H(Y)), attained when one variable fully determines the other (H(Y|X) = 0).
Check
How big can MI get? Use Form 2 and H(Y|X) ≥ 0.
Check your understanding
What is the tightest general upper bound on I(X;Y)?
Answer: A
Why: From I = H(Y) − H(Y|X) with H(Y|X) ≥ 0, we get I ≤ H(Y); by symmetry I ≤ H(X). So I ≤ min(H(X), H(Y)), attained when one variable fully determines the other (H(Y|X) = 0).
Section
Part 7 of 7 — contrastive learning
Intuition
CLIP and SimCLR don't have labels — they have pairs that belong together: an image and its caption, or two crops of the same photo. The trick is to make matched pairs' embeddings similar and mismatched pairs' embeddings different.
Doing so provably maximizes a lower bound on the mutual information between the two views. Contrastive learning is mutual-information maximization in disguise — which is why this lesson ends here.
Concept
For a true pair (x, y) among a batch of candidates {z_j} (the negatives plus the true y), InfoNCE is a softmax cross-entropy that says 'the true partner should win the similarity contest':
\[ \mathcal{L} = -\log \frac{\exp\!\big(\mathrm{sim}(x,y)/\tau\big)}{\sum_j \exp\!\big(\mathrm{sim}(x,z_j)/\tau\big)} \]
sim is cosine similarity; τ is a temperature that sharpens the softmax. Minimizing L pulls the true pair together and pushes the negatives apart — the exact cross-entropy machinery of Lesson 17, applied to representation learning.
Intuition
The temperature τ divides every similarity before the softmax. Small τ (like 0.1) multiplies the logits up (by 10), sharpening the softmax so it rewards the single best match harshly — the model is picky.
Large τ flattens the logits toward a uniform softmax, so the loss barely distinguishes the true partner from the negatives — the model gets lazy. CLIP tunes τ as a learnable parameter; too small overfits sharp boundaries, too large learns nothing.
In our trace τ = 0.1, so a perfect cosine of 1 becomes a logit of 10 — enormous, which is why aligned pairs drove the loss almost to zero.
Explain it
Discussion prompt
Explain What temperature τ does to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
In our trace τ = 0.1, so a perfect cosine of 1 becomes a logit of 10 — enormous, which is why aligned pairs drove the loss almost to zero.
Missing information
Discussion prompt
Three image embeddings, three text embeddings, matched by index. Build the similarity matrix S = (img·txtᵀ)/τ; the correct partner for row i is column i (the diagonal):
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Each row is a 3-way classification: which text matches this image? The target is the diagonal. Rows where the true partner is NOT the most similar (like row 2 below) pay a big loss — exactly the pairs the gradient will fix.
Worked example
Three image embeddings, three text embeddings, matched by index. Build the similarity matrix S = (img·txtᵀ)/τ; the correct partner for row i is column i (the diagonal):
import numpy as np
np.random.seed(0)
d = 4
img = np.random.randn(3, d)
txt = np.random.randn(3, d)
img = img / np.linalg.norm(img, axis=1, keepdims=True) # unit vectors
txt = txt / np.linalg.norm(txt, axis=1, keepdims=True)
tau = 0.1
S = img @ txt.T / tau # 3x3 logits
def softmax(r):
e = np.exp(r - r.max()); return e / e.sum()
loss = np.mean([-np.log(softmax(S[i])[i]) for i in range(3)])
print("S =\n", S.round(2))
print("InfoNCE loss =", round(loss, 4)) # 1.6539Row loss = −log(softmax(row)[correct index])
Why: Each row is a 3-way classification: which text matches this image? The target is the diagonal. Rows where the true partner is NOT the most similar (like row 2 below) pay a big loss — exactly the pairs the gradient will fix.
| row (image) | correct prob | row loss |
|---|---|---|
| 0 | 0.9992 | 0.0008 |
| 1 | 0.6837 | 0.3802 |
| 2 | 0.0102 | 4.5806 |
| mean | — | 1.6539 |
Pattern
Step through it
Step through InfoNCE traced on a tiny batch one row at a time. What is driving the change, and what would the row after the last one be?
Estimation
Predict first
Now make each image embedding equal its text partner (perfect alignment). The diagonal similarities dominate, softmax puts nearly all mass on the true index, and the loss drops toward zero:
Commit before you compute: what does When embeddings align, the loss collapses come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Diagonal = 1/τ = 10 (cosine of a vector with itself is 1); loss = 0.0188
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Aligned pairs push the correct softmax probability near 1, so −log(≈1) ≈ 0.
Worked example
Now make each image embedding equal its text partner (perfect alignment). The diagonal similarities dominate, softmax puts nearly all mass on the true index, and the loss drops toward zero:
import numpy as np
np.random.seed(1)
d = 4
emb = np.random.randn(3, d)
emb = emb / np.linalg.norm(emb, axis=1, keepdims=True)
tau = 0.1
S = emb @ emb.T / tau # img and txt already agree
def softmax(r):
e = np.exp(r - r.max()); return e / e.sum()
loss = np.mean([-np.log(softmax(S[i])[i]) for i in range(3)])
print("diagonal of S =", np.diag(S).round(1)) # [10. 10. 10.]
print("aligned loss =", round(loss, 4)) # 0.0188Diagonal = 1/τ = 10 (cosine of a vector with itself is 1); loss = 0.0188
Why: Aligned pairs push the correct softmax probability near 1, so −log(≈1) ≈ 0. Compare 0.0188 (aligned) with 1.6539 (random) and log(3) = 1.0986 (worst-case uniform guessing) — training drives representations from random toward aligned.
| batch state | mean loss | meaning |
|---|---|---|
| random embeddings | 1.6539 | partners not yet matched |
| worst case (uniform) | 1.0986 = log 3 | pure guessing among 3 |
| aligned embeddings | 0.0188 | true partner wins every row |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Diagonal = 1/τ = 10 (cosine of a vector with itself is 1); loss = 0.0188
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Now make each image embedding equal its text partner (perfect alignment). The diagonal similarities dominate, softmax puts nearly all mass on the true index, and the loss drops toward zero:
Concept
With a batch of N candidates, minimizing InfoNCE L maximizes a bound tying it to the mutual information between the two views x and y:
\[ I(X;Y) \;\ge\; \log_2 N - \mathcal{L} \]
Driving the loss L down pushes the lower bound on I(X;Y) up — so contrastive training literally increases the mutual information between an image and its caption. The bound tightens with bigger batches (log N), which is why CLIP trains on enormous batches of negatives.
Analogy
Discussion prompt
Explain Why InfoNCE is a lower bound on MI by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
With a batch of N candidates, minimizing InfoNCE L maximizes a bound tying it to the mutual information between the two views x and y:
Concept
Information theory ties straight to compression. Shannon's source coding theorem says entropy H(X) is the minimum average bits per symbol to encode X — you cannot compress below its entropy without losing information.
Minimum description length reframes learning as compression: the best model encodes the data in the fewest bits. Lower entropy and higher MI mean more exploitable structure — a model that finds high-MI features has literally found something compressible.
Section
Project — MI from scratch
Concept
Build MI from a joint distribution by two of the three formulas, confirm they match, then confirm an independent joint gives exactly 0. You've derived every piece — now assemble it yourself.
| # | requirement | tool |
|---|---|---|
| 1 | marginals + their entropies | J.sum(axis), entropy(p) |
| 2 | MI two ways, both equal | definition (Form 1) / KL (Form 3) |
| 3 | independent joint → MI = 0 | np.outer(px, py) |
Build rules: type every line yourself, mask out zero probabilities in the entropy (log 0 is undefined), and use log2 so the answer is in bits.
Counterexample
Discussion prompt
Build MI from a joint distribution by two of the three formulas, confirm they match, then confirm an independent joint gives exactly 0. You've derived every piece — now assemble it yourself.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Build rules: type every line yourself, mask out zero probabilities in the entropy (log 0 is undefined), and use log2 so the answer is in bits.
Worked example
Your turn: get the marginals and their entropies from the joint. Predict H(X) for a 50/50 marginal before you print.
Hint: px = J.sum(1), py = J.sum(0); entropy(p) = −Σ p log₂ p over p > 0.
import numpy as np
def entropy(p):
p = np.asarray(p, dtype=float); p = p[p > 0]
return -np.sum(p * np.log2(p))
J = np.array([[0.4, 0.1], [0.1, 0.4]])
px, py = J.sum(1), J.sum(0)
print(px, py) # [0.5 0.5] [0.5 0.5]
print(round(entropy(px), 4), round(entropy(py), 4)) # 1.0 1.0| quantity | value |
|---|---|
| px, py | [0.5, 0.5], [0.5, 0.5] |
| H(X) | 1.0 |
| H(Y) | 1.0 |
Worked example
Your turn: compute MI by the entropy definition (Form 1) and by the KL form (Form 3). Predict: do they match?
Hint: Form 1 is entropy(px)+entropy(py)−entropy(J.ravel()); Form 3 is np.sum(J*np.log2(J/np.outer(px,py))).
import numpy as np
def entropy(p):
p = np.asarray(p, dtype=float); p = p[p > 0]
return -np.sum(p * np.log2(p))
J = np.array([[0.4, 0.1], [0.1, 0.4]])
px, py = J.sum(1), J.sum(0)
mi_def = entropy(px) + entropy(py) - entropy(J.ravel())
mi_kl = np.sum(J * np.log2(J / np.outer(px, py)))
print(round(mi_def, 4), round(mi_kl, 4)) # 0.2781 0.2781| formula | value (bits) |
|---|---|
| Form 1 (entropies) | 0.2781 |
| Form 3 (KL) | 0.2781 |
Worked example
Your turn: build an independent joint with np.outer(px, py) and confirm its MI is 0 — the proof of the I = 0 ⟺ independent test.
Hint: for Ind = np.outer(px, py), the joint already equals the product, so every log-ratio is log(1) = 0.
import numpy as np
J = np.array([[0.4, 0.1], [0.1, 0.4]])
px, py = J.sum(1), J.sum(0)
Ind = np.outer(px, py) # independent joint
ipx, ipy = Ind.sum(1), Ind.sum(0)
mi = np.sum(Ind * np.log2(Ind / np.outer(ipx, ipy)))
print(round(mi, 6)) # 0.0| joint | MI (bits) |
|---|---|
| dependent J | 0.2781 |
| independent outer(px, py) | 0.0 |
Trade off
Comparison matrix
From Milestone 3 — independence gives 0: every row here is a choice with a cost. Fill the MI (bits) column, then say which row you would actually pick and what you give up for it.
| joint | MI (bits) |
|---|---|
| dependent J | 0.2781 |
| independent outer(px, py) | 0.0 |
Concept
import numpy as np
def entropy(p):
p = np.asarray(p, dtype=float); p = p[p > 0]
return -np.sum(p * np.log2(p))
def mutual_information(J):
px, py = J.sum(1), J.sum(0)
return np.sum(J * np.log2(J / np.outer(px, py))) # = H(X)+H(Y)-H(X,Y)
J = np.array([[0.4, 0.1], [0.1, 0.4]])
px, py = J.sum(1), J.sum(0)
print('MI(dependent) =', round(mutual_information(J), 4)) # 0.2781
print('MI(independent) =', round(mutual_information(np.outer(px, py)), 6)) # 0.0| output | value |
|---|---|
| MI(dependent) | 0.2781 |
| MI(independent) | 0.0 |
If MI agrees across the formulas and zeroes out for an independent joint — you can measure dependence the way representation learning does.
Comparison
Comparison matrix
From The full program: refill the value column from what you know. The rest of the table is as it appeared.
| output | value |
|---|---|
| MI(dependent) | 0.2781 |
| MI(independent) | 0.0 |
Concept
Out loud, slides closed: explain (1) the three roads from the definition to Forms 1, 2, and 3, (2) why zero MI is strictly stronger than zero correlation, and (3) what InfoNCE pulls together and pushes apart.
Stretch (homework): prove I(X;Y) ≥ 0 from D_KL ≥ 0, and use mutual_info_classif to rank the features of a 100-feature dataset. Mutual information returns in CLIP and contrastive pre-training (Week 47).
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Recap: entropy & one running joint · Mutual information: three faces · Form 2: conditional entropy · Form 3: distance from independence · Bounds & the independence test · Application: feature selection. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
H(X) = −Σ p log₂ p as bits of uncertainty, masking zerosH(X)+H(Y)−H(X,Y), H(Y)−H(Y|X), D_KL(p(x,y)‖p(x)p(y)) — and get 0.2781 bits three times on one jointI ≥ 0, I ≤ min(H(X),H(Y)), and I = 0 ⟺ independence, then show Y = X² has correlation ≈ 0 but MI 0.9183mutual_info_classif1.6539 → aligned 0.0188) and connect entropy to compression (MDL)| idea | the one thing to remember |
|---|---|
| MI = shared information | one number, three forms — all = 0.2781 here |
| Form 3 (KL) | distance of the joint from p(x)p(y) |
| MI = 0 | ⟺ independent — catches nonlinear, correlation can't |
| bounds | 0 ≤ I ≤ min(H(X), H(Y)) |
| feature selection | rank by I(feature; label) |
| InfoNCE | softmax over similarities = MI lower bound (CLIP) |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.