Lesson 29: Mutual Information & Contrastive Learning

USAAIO Lesson 29, from Week 10, fully worked. It recaps entropy from scratch, then derives mutual information I(X;Y) THREE equivalent ways, one algebra move at a time: as H(X)+H(Y)-H(X,Y), in the chain-rule form H(Y)-H(Y|X), and in the KL form D(p(x,y)||p(x)p(y)). All three are computed by hand on a single running 2x2 joint distribution and confirmed by real code. It then covers pointwise mutual information, proves the bounds I>=0 and I<=min(H), and works the trap in which correlation is zero yet the variables are dependent, on Y=X^2. It closes with mutual-information feature selection in sklearn, the InfoNCE and CLIP loss traced on a tiny batch, the link to compression, and a from-scratch mutual-information project. Every snippet runs standalone, and every number came from real execution. The lesson runs to 61 slides.

Subject: Machine Learning · 107 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Mutual Information & Contrastive Learning

Title

USAAIO · Lesson 29 · Week 10 (Information Theory)

How much does knowing one variable tell you about another? We derive I(X;Y) three ways — entropy, chain rule, and KL — skipping no step, compute each by hand on one running joint, then ride it into feature selection and the InfoNCE loss behind CLIP.

2. By the end of this lesson you can

Objectives

  1. Recall entropy H(X) = −Σ p log₂ p and read it as bits of uncertainty
  2. Derive mutual information three equivalent ways — H(X)+H(Y)−H(X,Y), H(Y)−H(Y|X), and D_KL(p(x,y)‖p(x)p(y)) — and compute all three by hand to the same number
  3. Prove I ≥ 0, I ≤ min(H(X),H(Y)), and I = 0 ⟺ X ⊥ Y, then show a zero-correlation pair with large MI
  4. Select features by mutual information with the label and read sklearn's scores
  5. Trace the InfoNCE contrastive loss on a real batch and connect entropy to compression (MDL)

3. What survived from Kernel Methods?

Warm-up

Discussion prompt

Before we open Lesson 29: Mutual Information & Contrastive Learning: without looking back, what was the main idea of Kernel Methods, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

feature maps and the kernel trick k(x,z)=φ(x)·φ(z), Mercer's theorem (a valid kernel has a PSD Gram matrix), the RBF kernel as an infinite-dimensional feature space, kernel composition rules, and kernel SVMs. Verify the polynomial kernel against its explicit feature map and solve XOR with an RBF SVM.

4. Recap: entropy & one running joint

Section

Part 1 of 7 — the raw material

5. Entropy is uncertainty measured in bits

Intuition

Before mutual information there is entropy — how uncertain you are about a random variable, in bits. A fair coin (50/50) has 1 bit: you need exactly one yes/no answer to pin it down.

A rigged coin that lands heads 99% of the time is almost certain, so it carries less than a bit. A variable you already know carries 0 bits — nothing left to learn.

Mutual information will be entropy that two variables share — how much the uncertainty in one drops once you see the other. So we start by nailing entropy down cold.

6. Break it if you can: Entropy is uncertainty measured in bits

Counterexample

Discussion prompt

Before mutual information there is entropy — how uncertain you are about a random variable, in bits. A fair coin (50/50) has 1 bit: you need exactly one yes/no answer to pin it down.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

A rigged coin that lands heads 99% of the time is almost certain, so it carries less than a bit. A variable you already know carries 0 bits — nothing left to learn.

7. Self-information: the surprise of one outcome

Concept

The surprise of a single outcome with probability p is −log₂ p bits. A coin flip (p = 0.5) is −log₂ 0.5 = 1 bit of surprise; a 1-in-8 event is −log₂(1/8) = 3 bits — three times as surprising.

\[ \text{surprise}(x) = -\log_2 p(x) \]

Entropy will just be the average surprise over all outcomes. Rare things surprise more, certain things (p = 1) surprise not at all — −log₂ 1 = 0. Keep this: MI counts how much one variable's surprise drops once you see the other.

8. The entropy formula

Concept

For a discrete variable X with outcome probabilities p(x), entropy is the expected surprise −log₂ p(x), averaged over outcomes:

\[ H(X) = -\sum_{x} p(x)\,\log_2 p(x) \]

Rare outcomes (small p) are surprising (large −log₂ p); certain outcomes (p = 1) carry zero surprise. Using log₂ makes the unit bits. We adopt the convention 0·log 0 = 0.

9. By analogy: The entropy formula

Analogy

Discussion prompt

Explain The entropy formula by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

For a discrete variable X with outcome probabilities p(x), entropy is the expected surprise −log₂ p(x), averaged over outcomes:

10. Guess the shape of the answer: Entropy in code — three sanity checks

Estimation

Predict first

We mask out zero probabilities (log 0 is undefined but 0·log 0 = 0), then sum. Run this standalone:

Commit before you compute: what does Entropy in code — three sanity checks come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: The mask p[p > 0] is the one detail people forget

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. A distribution with a zero-probability outcome would otherwise feed log2(0) = −inf into the sum and poison it with NaN.

11. Entropy in code — three sanity checks

Worked example

We mask out zero probabilities (log 0 is undefined but 0·log 0 = 0), then sum. Run this standalone:

import numpy as np

def entropy(p):
    p = np.asarray(p, dtype=float)
    p = p[p > 0]                      # 0 log 0 = 0, so drop zeros
    return -np.sum(p * np.log2(p))    # bits

print(entropy([0.5, 0.5]))            # 1.0
print(entropy([0.4, 0.1, 0.1, 0.4])) # 1.7219...
print(entropy([1.0]))                 # 0.0 (no uncertainty)

The mask p[p > 0] is the one detail people forget

Why: A distribution with a zero-probability outcome would otherwise feed log2(0) = −inf into the sum and poison it with NaN. Dropping zeros implements the 0·log 0 = 0 convention exactly.

distributionH (bits)meaning
[0.5, 0.5]1.0fair coin — max uncertainty for 2 outcomes
[0.4, 0.1, 0.1, 0.4]1.7219our joint's 4 cells (used all lesson)
[1.0]0.0a sure thing — no uncertainty

12. Fill in: H (bits) for Entropy in code — three sanity checks

Comparison

Comparison matrix

From Entropy in code — three sanity checks: refill the H (bits) column from what you know. The rest of the table is as it appeared.

distributionH (bits)meaning
[0.5, 0.5]1.0fair coin — max uncertainty for 2 outcomes
[0.4, 0.1, 0.1, 0.4]1.7219our joint's 4 cells (used all lesson)
[1.0]0.0a sure thing — no uncertainty

13. What has to be given first: Entropy peaks at the fair coin

Missing information

Discussion prompt

Sweep a coin's head-probability q from fair to certain and watch entropy fall. Uniform is maximally uncertain; a rigged coin is more predictable, so it carries fewer bits:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

A fair coin (q = 0.5) is the hardest to predict → 1 full bit. As q → 1 the outcome becomes certain and entropy → 0. Entropy is maximized by the uniform distribution — a fact we lean on for the MI bounds later.

14. Entropy peaks at the fair coin

Worked example

Sweep a coin's head-probability q from fair to certain and watch entropy fall. Uniform is maximally uncertain; a rigged coin is more predictable, so it carries fewer bits:

import numpy as np

def entropy(p):
    p = np.asarray(p, dtype=float); p = p[p > 0]
    return -np.sum(p * np.log2(p))

for q in [0.5, 0.75, 0.9, 0.99, 1.0]:
    print(q, round(entropy([q, 1-q]), 4))

H is 1.0 at q = 0.5 and slides to 0 at q = 1.0

Why: A fair coin (q = 0.5) is the hardest to predict → 1 full bit. As q → 1 the outcome becomes certain and entropy → 0. Entropy is maximized by the uniform distribution — a fact we lean on for the MI bounds later.

q (P heads)H([q, 1−q])reading
0.501.0000fair — maximum uncertainty
0.750.8113leaning heads
0.900.4690fairly predictable
0.990.0808almost certain
1.000.0000a sure thing

15. What each one costs: Entropy peaks at the fair coin

Trade off

Comparison matrix

From Entropy peaks at the fair coin: every row here is a choice with a cost. Fill the H([q, 1−q]) column, then say which row you would actually pick and what you give up for it.

q (P heads)H([q, 1−q])reading
0.501.0000fair — maximum uncertainty
0.750.8113leaning heads
0.900.4690fairly predictable
0.990.0808almost certain
1.000.0000a sure thing

16. The running example: weather × umbrella

Concept

One dataset carries the whole lesson. X is the weather (0 = dry, 1 = rain); Y is whether a commuter carries an umbrella (0 = no, 1 = yes). Their joint distribution p(x,y):

p(x,y)Y=0 (no umbrella)Y=1 (umbrella)
X=0 (dry)0.400.10
X=1 (rain)0.100.40

The mass sits on the diagonal: dry-and-no-umbrella and rain-and-umbrella are common; the off-diagonal mismatches are rare. So the two are related — but not perfectly. Quantifying exactly how related is the job of mutual information.

17. Teach it back: The running example: weather × umbrella

Explain it

Discussion prompt

Explain The running example: weather × umbrella to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

One dataset carries the whole lesson. X is the weather (0 = dry, 1 = rain); Y is whether a commuter carries an umbrella (0 = no, 1 = yes). Their joint distribution p(x,y):

18. Predict the next row: Marginals: collapse the joint

Pattern

Predict first

The table runs: px = p(X) | [0.5, 0.5] | 0.4+0.1, 0.1+0.4 · py = p(Y) | [0.5, 0.5] | 0.4+0.1, 0.1+0.4

In Marginals: collapse the joint, given the rows so far: what is the next one — the row where quantity is Σ all cells?

Correct: Σ all cells | 1.0 | valid distribution

quantityvaluehow
px = p(X)[0.5, 0.5]0.4+0.1, 0.1+0.4
py = p(Y)[0.5, 0.5]0.4+0.1, 0.1+0.4
Σ all cells1.0valid distribution

Why: The relationship between the columns, not the individual numbers, is what generates the next row. Remember it by what SURVIVES: summing axis=1 keeps the row index x.

19. Marginals: collapse the joint

Worked example

The marginal p(x) sums each row (over all y); p(y) sums each column (over all x). Sum over the axis you want to erase:

import numpy as np

J = np.array([[0.4, 0.1],
              [0.1, 0.4]])           # rows = X, cols = Y
px = J.sum(axis=1)                    # marginal of X: sum over Y
py = J.sum(axis=0)                    # marginal of Y: sum over X
print("px =", px)                     # [0.5 0.5]
print("py =", py)                     # [0.5 0.5]
print("sum J =", J.sum())             # 1.0

axis=1 sums over columns (erases Y) → p(x); axis=0 erases X → p(y)

Why: Remember it by what SURVIVES: summing axis=1 keeps the row index x. Here both marginals come out [0.5, 0.5] — dry vs rain is 50/50, and so is umbrella vs not.

quantityvaluehow
px = p(X)[0.5, 0.5]0.4+0.1, 0.1+0.4
py = p(Y)[0.5, 0.5]0.4+0.1, 0.1+0.4
Σ all cells1.0valid distribution

20. Work backwards from the answer: Marginals: collapse the joint

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

axis=1 sums over columns (erases Y) → p(x); axis=0 erases X → p(y)

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

The marginal p(x) sums each row (over all y); p(y) sums each column (over all x). Sum over the axis you want to erase:

21. Complete the line: Compute H(X), H(Y), H(X,Y) by hand

Fill the middle

Fill in the blanks

From Compute H(X), H(Y), H(X,Y) by hand — finish the line. Write what belongs on the right of the equals sign before you look.

H(X,Y) = -2\big(0.4\log_2 0.4\big) - 2\big(0.1\log_2 0.1\big)

Why: Producing the right-hand side unprompted is the difference between recognising this line and being able to use it. −0.4·log₂0.4 = 0.5288 and −0.1·log₂0.1 = 0.3322 (since log₂0.4 = −1.3219, log₂0.1 = −3.3219).

22. Compute H(X), H(Y), H(X,Y) by hand

Worked example

Three entropies feed every MI formula. Marginals are [0.5, 0.5], so both give 1 bit:

\[ H(X) = -2\big(0.5\log_2 0.5\big) = -2(0.5)(-1) = 1 \text{ bit} \]

The joint entropy uses all four cells [0.4, 0.1, 0.1, 0.4]:

\[ H(X,Y) = -2\big(0.4\log_2 0.4\big) - 2\big(0.1\log_2 0.1\big) \]

= 2(0.5288) + 2(0.3322) = 1.0575 + 0.6644 = 1.7219 bits

Why: −0.4·log₂0.4 = 0.5288 and −0.1·log₂0.1 = 0.3322 (since log₂0.4 = −1.3219, log₂0.1 = −3.3219). Verified by real execution: H(X,Y) = 1.721928.

entropyvalue (bits)from
H(X)1.0000marginal [0.5, 0.5]
H(Y)1.0000marginal [0.5, 0.5]
H(X,Y)1.7219all 4 cells [0.4,0.1,0.1,0.4]

23. Why H(X,Y) < H(X) + H(Y)

Concept

Notice H(X) + H(Y) = 2.0 but H(X,Y) = 1.7219. The joint is less uncertain than the two separately. If X and Y were independent they'd add to exactly 2.0; the shortfall is real structure.

That gap — 2.0 − 1.7219 = 0.2781 bits — is precisely the information they share. Hold that number: it is the mutual information, and the next part earns it three different ways.

24. Mutual information: three faces

Section

Part 2 of 7 — one quantity, three derivations

25. Picture it first: MI is shared uncertainty

Picture it

Figure (svg): Two overlapping circles labeled H(X) and H(Y); the left-only crescent is H(X|Y), the right-only crescent is H(Y|X), and the lens-shaped overlap in the middle is labeled I(X;Y).

H(X)+H(Y) double-counts the overlap; subtracting H(X,Y) leaves exactly one copy of it — the mutual information.

Discussion prompt

Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.

Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.

Answer:

Picture two overlapping circles: H(X) is the left circle, H(Y) the right. Where they overlap is information the two variables have in common — see one, and you learn something about the other.

26. MI is shared uncertainty

Intuition

Picture two overlapping circles: H(X) is the left circle, H(Y) the right. Where they overlap is information the two variables have in common — see one, and you learn something about the other.

That overlap is mutual information I(X;Y). If the circles are disjoint (no overlap) the variables are independent and I = 0. If one circle sits entirely inside the other, knowing one fully determines the shared part.

Figure (svg): Two overlapping circles labeled H(X) and H(Y); the left-only crescent is H(X|Y), the right-only crescent is H(Y|X), and the lens-shaped overlap in the middle is labeled I(X;Y).

H(X)+H(Y) double-counts the overlap; subtracting H(X,Y) leaves exactly one copy of it — the mutual information.

27. The three equivalent forms

Concept

Mutual information is one number wearing three algebraic costumes. Each will be derived in full — but here they are up front so you know the destination:

\[ I(X;Y) = H(X) + H(Y) - H(X,Y) \]

\[ \;\;= H(Y) - H(Y\mid X) \]

\[ \;\;= D_{KL}\big(p(x,y)\,\|\,p(x)p(y)\big) \]

Form 1 is the set overlap. Form 2 is uncertainty reduced by observing X. Form 3 is the distance from independence. They are equal — provably, not coincidentally.

28. The definition we build from

Concept

The ground-truth definition of MI is the expected pointwise log-ratio between the joint and the product of marginals:

\[ I(X;Y) = \sum_{x,y} p(x,y)\,\log_2 \frac{p(x,y)}{p(x)\,p(y)} \]

Everything else in this part is this expression, re-grouped. It is already Form 3 (a KL divergence). Watch it collapse into Forms 1 and 2 with nothing but log rules.

29. What has to happen first: Derivation A → Form 1 (entropies)

Ranking

Put in order

Put the moves of Derivation A → Form 1 (entropies) into the order they have to happen.

  1. Start from the definition
  2. Split the log of a quotient into a difference
  3. Split log p(x)p(y) into log p(x) + log p(y)

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. The expected log-ratio of joint to product of marginals — our single source of truth.

30. Derivation A → Form 1 (entropies)

Worked example

Start from the definition

Why: The expected log-ratio of joint to product of marginals — our single source of truth.

\[ I(X;Y) = \sum_{x,y} p(x,y)\,\log_2 \frac{p(x,y)}{p(x)p(y)} \]

Split the log of a quotient into a difference

Why: log(a/b) = log a − log b. The single log becomes log p(x,y) minus log[p(x)p(y)].

\[ = \sum_{x,y} p(x,y)\Big[\log_2 p(x,y) - \log_2 p(x)p(y)\Big] \]

Split log p(x)p(y) into log p(x) + log p(y)

Why: log of a product is a sum of logs. Now three log terms, each weighted by p(x,y).

\[ = \sum_{x,y} p(x,y)\log_2 p(x,y) - \sum_{x,y} p(x,y)\log_2 p(x) - \sum_{x,y} p(x,y)\log_2 p(y) \]

31. Decode the notation: Derivation A → Form 1 (entropies)

Notation

Annotate

From Derivation A → Form 1 (entropies) — read this one piece at a time. What is each part doing?

On: \( I(X;Y) = \sum_{x,y} p(x,y)\,\log_2 \frac{p(x,y)}{p(x)p(y)} \)

  • The expected log-ratio of joint to product of marginals — our single source of truth.
  • log(a/b) = log a − log b. The single log becomes log p(x,y) minus log[p(x)p(y)].
  • log of a product is a sum of logs. Now three log terms, each weighted by p(x,y).

32. Plan first: Derivation A → Form 1 (finish)

Step zero

Discussion prompt

Derivation A → Form 1 (finish) — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Term 1 is −H(X,Y)

Answer:

  1. Term 1 is −H(X,Y)
  2. In term 2, sum out y first
  3. Term 3 is −H(Y) by the same collapse over x
  4. Assemble: −(−H(X,Y)) picks up the leading minus signs

33. Derivation A → Form 1 (finish)

Worked example

Term 1 is −H(X,Y)

Why: Σ p(x,y) log₂ p(x,y) is exactly the negative joint entropy, by definition of H(X,Y).

\[ \sum_{x,y} p(x,y)\log_2 p(x,y) = -H(X,Y) \]

In term 2, sum out y first

Why: log₂ p(x) doesn't depend on y, so Σ_y p(x,y) collapses to the marginal p(x). The double sum becomes a single sum over x.

\[ \sum_{x,y} p(x,y)\log_2 p(x) = \sum_x p(x)\log_2 p(x) = -H(X) \]

Term 3 is −H(Y) by the same collapse over x

Why: Symmetric: summing out x turns p(x,y) into p(y).

\[ \sum_{x,y} p(x,y)\log_2 p(y) = -H(Y) \]

Assemble: −(−H(X,Y)) picks up the leading minus signs

Why: I = −H(X,Y) − (−H(X)) − (−H(Y)) = H(X) + H(Y) − H(X,Y). Form 1, derived with only log rules and marginalization.

\[ \boxed{\,I(X;Y) = H(X) + H(Y) - H(X,Y)\,} \]

34. Say it in words: Derivation A → Form 1 (finish)

Translation

\( \boxed{\,I(X;Y) = H(X) + H(Y) - H(X,Y)\,} \)

Draw it

Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.

35. Restore the missing line: Form 1 numerically — it's the gap we spotted

Fill the middle

Fill in the blanks

From Form 1 numerically — it's the gap we spotted — one line has had its right-hand side removed. Put it back.

import numpy as np

def entropy(p):
p = np.asarray(p, dtype=float); p = p[p > 0]
return -np.sum(p * np.log2(p))

J = np.array([[0.4, 0.1], [0.1, 0.4]])
px, py = J.sum(1), J.sum(0)
HX = entropy(px) # 1.0
HY = entropy(py) # 1.0
HXY = entropy(J.ravel()) # 1.7219
mi = HX + HY - HXY
print(HX, HY, round(HXY, 4))
print("I(X;Y) =", round(mi, 4)) # 0.2781

Why: HY is what everything below it consumes, so the wrong expression here fails later and somewhere else. The two circles sum to 2.0 bits; the joint is only 1.7219 bits; the 0.2781-bit shortfall is the shared information.

36. Form 1 numerically — it's the gap we spotted

Worked example

Plug our three entropies straight into Form 1. This is the 0.2781 we flagged in Part 1, now earned:

import numpy as np

def entropy(p):
    p = np.asarray(p, dtype=float); p = p[p > 0]
    return -np.sum(p * np.log2(p))

J = np.array([[0.4, 0.1], [0.1, 0.4]])
px, py = J.sum(1), J.sum(0)
HX  = entropy(px)                     # 1.0
HY  = entropy(py)                     # 1.0
HXY = entropy(J.ravel())             # 1.7219
mi  = HX + HY - HXY
print(HX, HY, round(HXY, 4))
print("I(X;Y) =", round(mi, 4))       # 0.2781

I = 1.0 + 1.0 − 1.7219 = 0.2781 bits

Why: The two circles sum to 2.0 bits; the joint is only 1.7219 bits; the 0.2781-bit shortfall is the shared information. Small but nonzero — weather and umbrella are related but not locked together.

piecevalue (bits)
H(X) + H(Y)2.0000
H(X,Y)1.7219
I(X;Y) = difference0.2781

37. Form 2: conditional entropy

Section

Part 3 of 7 — uncertainty reduced

38. Conditional entropy H(Y|X)

Concept

H(Y|X) is the average uncertainty left in Y after you already know X — the leftover surprise once the weather is revealed. It is defined by the chain rule of entropy:

\[ H(Y\mid X) = H(X,Y) - H(X) \]

Read it as: the total joint uncertainty, minus the part X already accounts for. Whatever remains is the uncertainty in Y that X couldn't explain.

39. What has to happen first: Derivation B → Form 2

Ranking

Put in order

Put the moves of Derivation B → Form 2 into the order they have to happen.

  1. Start from Form 1
  2. Regroup: pull H(Y) out front, keep H(X) − H(X,Y) together
  3. Recognize H(X,Y) − H(X) = H(Y|X), so H(X) − H(X,Y) = −H(Y|X)
  4. Form 2 established

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. We just proved I = H(X) + H(Y) − H(X,Y).

40. Derivation B → Form 2

Worked example

Start from Form 1

Why: We just proved I = H(X) + H(Y) − H(X,Y). Reuse it.

\[ I(X;Y) = H(X) + H(Y) - H(X,Y) \]

Regroup: pull H(Y) out front, keep H(X) − H(X,Y) together

Why: Pure rearrangement of the same three terms — associativity of addition, nothing added or dropped.

\[ = H(Y) + \big(H(X) - H(X,Y)\big) \]

Recognize H(X,Y) − H(X) = H(Y|X), so H(X) − H(X,Y) = −H(Y|X)

Why: The chain rule H(X,Y) = H(X) + H(Y|X) rearranges to H(Y|X) = H(X,Y) − H(X). Flip the sign to match our grouping.

\[ = H(Y) - \big(H(X,Y) - H(X)\big) = H(Y) - H(Y\mid X) \]

Form 2 established

Why: MI is how many bits of uncertainty about Y vanish once X is known. Same quantity, a genuinely different reading.

\[ \boxed{\,I(X;Y) = H(Y) - H(Y\mid X)\,} \]

41. Draw the shape of it: Derivation B → Form 2

Blank canvas

Draw it

Draw what Derivation B → Form 2 just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.

42. Form 2 numerically

Worked example

Compute H(Y|X) = H(X,Y) − H(X), then subtract from H(Y). Must land on 0.2781 again:

import numpy as np

def entropy(p):
    p = np.asarray(p, dtype=float); p = p[p > 0]
    return -np.sum(p * np.log2(p))

J = np.array([[0.4, 0.1], [0.1, 0.4]])
px, py = J.sum(1), J.sum(0)
HY   = entropy(py)                    # 1.0
HYgX = entropy(J.ravel()) - entropy(px)   # H(X,Y) - H(X) = H(Y|X)
print("H(Y)   =", round(HY, 4))       # 1.0
print("H(Y|X) =", round(HYgX, 4))     # 0.7219
print("I(X;Y) = H(Y) - H(Y|X) =", round(HY - HYgX, 4))  # 0.2781

H(Y|X) = 1.7219 − 1.0 = 0.7219; I = 1.0 − 0.7219 = 0.2781

Why: Knowing the weather cuts umbrella-uncertainty from 1.0 bit down to 0.7219 bits — a drop of exactly 0.2781 bits. Same MI as Form 1, confirming the two derivations meet.

quantityvalue (bits)reading
H(Y)1.0000umbrella uncertainty, knowing nothing
H(Y|X)0.7219umbrella uncertainty, knowing weather
I(X;Y) = drop0.2781bits weather tells you about umbrella

43. Where does each piece belong: Lesson 29: Mutual Information & Contrastive…

Sorting

Sort into buckets

These are the pieces of Lesson 29: Mutual Information & Contrastive Learning, out of order. Put each one back under the part of the lesson it belongs to.

Recap: entropy & one running joint
Entropy is uncertainty measured in bits; Self-information: the surprise of one outcome; The entropy formula
Mutual information: three faces
MI is shared uncertainty; The three equivalent forms; The definition we build from
Form 2: conditional entropy
Conditional entropy H(Y|X); Derivation B → Form 2; Form 2 numerically
s1
Recap: entropy & one running joint is where Lesson 29: Mutual Information & Contrastive Learning puts Entropy is uncertainty measured in bits, Self-information: the surprise of one outcome, The entropy formula. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
Mutual information: three faces is where Lesson 29: Mutual Information & Contrastive Learning puts MI is shared uncertainty, The three equivalent forms, The definition we build from. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
Form 2: conditional entropy is where Lesson 29: Mutual Information & Contrastive Learning puts Conditional entropy H(Y|X), Derivation B → Form 2, Form 2 numerically. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

44. MI is symmetric

Concept

Form 1 is untouched if you swap X and Y — so I(X;Y) = I(Y;X). By the same algebra, MI also equals H(X) − H(X|Y). Weather tells you as much about umbrellas as umbrellas tell you about weather.

directionformulavalue (bits)
Y from XH(Y) − H(Y|X)1.0 − 0.7219 = 0.2781
X from YH(X) − H(X|Y)1.0 − 0.7219 = 0.2781

Verified: H(X|Y) = H(X,Y) − H(Y) = 0.7219, so H(X) − H(X|Y) = 0.2781 — identical. Information is a shared quantity, not a one-way arrow.

45. Form 3: distance from independence

Section

Part 4 of 7 — the KL view

46. KL divergence, one line

Concept

From Lesson 14, the KL divergence measures how far one distribution q is from another p, in bits:

\[ D_{KL}(p \,\|\, q) = \sum_i p_i \log_2 \frac{p_i}{q_i} \;\ge\; 0 \]

It is 0 exactly when p = q, and strictly positive otherwise. MI's definition is a KL — between the true joint p(x,y) and the pretend-independent product p(x)p(y).

47. What has to happen first: Derivation C → Form 3 (it's immediate)

Ranking

Put in order

Put the moves of Derivation C → Form 3 (it's immediate) into the order they have to happen.

  1. Write the KL of joint against product
  2. That is the definition of I(X;Y) verbatim
  3. So MI = 'how far the joint is from independent'

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Set p = the true joint p(x,y) and q = the independent product p(x)p(y) in the KL definition.

48. Derivation C → Form 3 (it's immediate)

Worked example

Write the KL of joint against product

Why: Set p = the true joint p(x,y) and q = the independent product p(x)p(y) in the KL definition.

\[ D_{KL}\big(p(x,y)\,\|\,p(x)p(y)\big) = \sum_{x,y} p(x,y)\,\log_2 \frac{p(x,y)}{p(x)p(y)} \]

That is the definition of I(X;Y) verbatim

Why: Compare to our source-of-truth expression from Part 2 — character for character the same sum. No manipulation needed.

\[ \boxed{\,I(X;Y) = D_{KL}\big(p(x,y)\,\|\,p(x)p(y)\big)\,} \]

So MI = 'how far the joint is from independent'

Why: This is the most useful mental model: MI literally measures the gap between reality and the world where X and Y ignore each other. Zero gap ⇒ independent.

49. Decode the notation: Derivation C → Form 3 (it's immediate)

Notation

Annotate

From Derivation C → Form 3 (it's immediate) — read this one piece at a time. What is each part doing?

On: \( \boxed{\,I(X;Y) = D_{KL}\big(p(x,y)\,\|\,p(x)p(y)\big)\,} \)

  • Set p = the true joint p(x,y) and q = the independent product p(x)p(y) in the KL definition.
  • Compare to our source-of-truth expression from Part 2 — character for character the same sum. No manipulation needed.
  • This is the most useful mental model: MI literally measures the gap between reality and the world where X and Y ignore each other. Zero gap ⇒ independent.

50. Guess the shape of the answer: Form 3 numerically

Estimation

Predict first

Build the product p(x)p(y) with an outer product, then sum the weighted log-ratio. Third road, same 0.2781:

Commit before you compute: what does Form 3 numerically come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: np.outer(px, py) builds the 'if-independent' joint [[0.25]×4]

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. With both marginals [0.5, 0.5], independence would put 0.25 in every cell.

51. Form 3 numerically

Worked example

Build the product p(x)p(y) with an outer product, then sum the weighted log-ratio. Third road, same 0.2781:

import numpy as np

J = np.array([[0.4, 0.1], [0.1, 0.4]])
px, py = J.sum(1), J.sum(0)
prod = np.outer(px, py)               # p(x)p(y): the independent joint
mi_kl = np.sum(J * np.log2(J / prod)) # KL(joint || product)
print("outer(px,py) =\n", prod)
print("I(X;Y) as KL =", round(mi_kl, 4))   # 0.2781

np.outer(px, py) builds the 'if-independent' joint [[0.25]×4]

Why: With both marginals [0.5, 0.5], independence would put 0.25 in every cell. The real joint has 0.40/0.10 instead — that mismatch is what KL measures.

cellreal p(x,y)if independentratio
(dry, no)0.400.251.60
(dry, yes)0.100.250.40
(rain, no)0.100.250.40
(rain, yes)0.400.251.60

52. Which is which, by real p(x,y)

Discrimination

Sort into buckets

Sort these by real p(x,y), from memory, without looking back at Form 3 numerically. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

0.40
(dry, no); (rain, yes)
0.10
(dry, yes); (rain, no)
g1
real p(x,y) is "0.40" for (dry, no), (rain, yes) — that is what the table on "Form 3 numerically" records, and it is the single property separating this group from the rest.
g2
real p(x,y) is "0.10" for (dry, yes), (rain, no) — that is what the table on "Form 3 numerically" records, and it is the single property separating this group from the rest.

53. Pointwise mutual information

Concept

Inside the sum, each cell contributes p(x,y)·log₂[p(x,y)/(p(x)p(y))]. The bare log-ratio log₂[p(x,y)/(p(x)p(y))] is the pointwise MI (PMI) of one outcome pair — positive when the pair co-occurs more than chance, negative when less.

pairPMI (bits)contributionreading
(dry, no)+0.6781+0.27123co-occur more than chance
(dry, yes)−1.3219−0.13219rarer than chance
(rain, no)−1.3219−0.13219rarer than chance
(rain, yes)+0.6781+0.27123co-occur more than chance

Sum the contributions: 0.27123 − 0.13219 − 0.13219 + 0.27123 = 0.2781. Individual PMIs can be negative, but their probability-weighted sum — the MI — is always ≥ 0. That guarantee is next.

54. Which is which, by PMI (bits)

Discrimination

Sort into buckets

Sort these by PMI (bits), from memory, without looking back at Pointwise mutual information. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

+0.6781
(dry, no); (rain, yes)
−1.3219
(dry, yes); (rain, no)
g1
PMI (bits) is "+0.6781" for (dry, no), (rain, yes) — that is what the table on "Pointwise mutual information" records, and it is the single property separating this group from the rest.
g2
PMI (bits) is "−1.3219" for (dry, yes), (rain, no) — that is what the table on "Pointwise mutual information" records, and it is the single property separating this group from the rest.

55. Restore the missing line: All three forms agree — the master check

Fill the middle

Fill in the blanks

From All three forms agree — the master check — one line has had its right-hand side removed. Put it back.

import numpy as np

def entropy(p):
p = np.asarray(p, dtype=float); p = p[p > 0]
return -np.sum(p * np.log2(p))

J = np.array([[0.4, 0.1], [0.1, 0.4]])
px, py = J.sum(1), J.sum(0)
form1 = entropy(px) + entropy(py) - entropy(J.ravel())
form2 = entropy(py) - (entropy(J.ravel()) - entropy(px))
form3 = np.sum(J * np.log2(J / np.outer(px, py)))
print(round(form1, 6), round(form2, 6), round(form3, 6))

Why: J is what everything below it consumes, so the wrong expression here fails later and somewhere else. Not a numerical coincidence: the three are algebraically the same object, and we proved each transformation.

56. All three forms agree — the master check

Worked example

Three independent derivations, one dataset, one number. This is the payoff slide — verified by execution:

import numpy as np

def entropy(p):
    p = np.asarray(p, dtype=float); p = p[p > 0]
    return -np.sum(p * np.log2(p))

J = np.array([[0.4, 0.1], [0.1, 0.4]])
px, py = J.sum(1), J.sum(0)
form1 = entropy(px) + entropy(py) - entropy(J.ravel())
form2 = entropy(py) - (entropy(J.ravel()) - entropy(px))
form3 = np.sum(J * np.log2(J / np.outer(px, py)))
print(round(form1, 6), round(form2, 6), round(form3, 6))

0.278072 0.278072 0.278072 — identical to 6 places

Why: Not a numerical coincidence: the three are algebraically the same object, and we proved each transformation. This triple-agreement is your always-available correctness check when you implement MI.

formexpressionvalue (bits)
1 — entropiesH(X)+H(Y)−H(X,Y)0.278072
2 — conditionalH(Y)−H(Y|X)0.278072
3 — KLD(p(x,y) ‖ p(x)p(y))0.278072

57. Bounds & the independence test

Section

Part 5 of 7 — I ≥ 0, I ≤ min(H), I = 0 ⟺ ⊥

58. MI is never negative

Concept

Because MI is a KL divergence and KL ≥ 0 always (Gibbs' inequality, Lesson 14), mutual information is non-negative:

\[ I(X;Y) = D_{KL}\big(p(x,y)\,\|\,p(x)p(y)\big) \;\ge\; 0 \]

You can never lose information by observing a variable — at worst it's useless (I = 0). This is why individual PMIs may go negative but the average cannot.

59. I = 0 exactly when independent

Concept

KL is 0 iff its two arguments are equal. So I(X;Y) = 0 iff p(x,y) = p(x)p(y) for every cell — which is the definition of independence:

\[ I(X;Y) = 0 \;\Longleftrightarrow\; p(x,y) = p(x)p(y)\;\;\forall x,y \;\Longleftrightarrow\; X \perp Y \]

This is the strongest independence test there is: it certifies no dependence of any kind, linear or not. Contrast that with correlation, which only sees straight lines.

60. An independent joint scores exactly 0

Worked example

Force independence by using the product p(x)p(y) as the joint. Its MI must be 0 — the KL of a distribution against itself:

import numpy as np

J = np.array([[0.4, 0.1], [0.1, 0.4]])
px, py = J.sum(1), J.sum(0)
Ind = np.outer(px, py)                # force independence
ipx, ipy = Ind.sum(1), Ind.sum(0)
mi = np.sum(Ind * np.log2(Ind / np.outer(ipx, ipy)))
print("Ind =\n", Ind)
print("I(X;Y) for the product =", round(mi, 6))   # 0.0

Every cell of Ind equals its own p(x)p(y), so each log-ratio is log(1) = 0

Why: Ind[i,k] = px[i]·py[k] by construction, and its marginals reproduce px, py, so Ind/outer(marginals) = 1 everywhere. log₂(1) = 0, sum = 0.

joint usedMI (bits)verdict
real J (diagonal-heavy)0.2781dependent
outer(px, py)0.000000independent — MI vanishes

61. Fill in: verdict for An independent joint scores exactly 0

Comparison

Comparison matrix

From An independent joint scores exactly 0: refill the verdict column from what you know. The rest of the table is as it appeared.

joint usedMI (bits)verdict
real J (diagonal-heavy)0.2781dependent
outer(px, py)0.000000independent — MI vanishes

62. MI is capped by each entropy

Concept

Since H(Y|X) ≥ 0, Form 2 gives I = H(Y) − H(Y|X) ≤ H(Y). By symmetry I ≤ H(X) too. So MI is squeezed between 0 and the smaller of the two entropies:

\[ 0 \;\le\; I(X;Y) \;\le\; \min\big(H(X),\,H(Y)\big) \]

For our joint: I = 0.2781 ≤ min(1.0, 1.0) = 1.0 ✓. MI hits its ceiling H(Y) only when X fully determines Y (H(Y|X) = 0) — perfect dependence.

63. Something is wrong here: zero correlation ⇒ MI = 0

Anomaly

Predict first

A student writes this, and it looks reasonable:

The correlation between X and Y is ≈ 0, so they carry no information about each other — I(X;Y) = 0.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Correlation only measures LINEAR association.

Correlation ≈ 0 rules out only a linear relationship. To test real independence, compute the mutual information.

Why: Correlation only measures LINEAR association. It is blind to any symmetric curve, so a zero correlation says nothing about nonlinear dependence.

64. Trap: zero correlation ⇒ MI = 0

Trap

The trap

The correlation between X and Y is ≈ 0, so they carry no information about each other — I(X;Y) = 0.

Rule out dependence using the correlation coefficient

Why: Correlation only measures LINEAR association. It is blind to any symmetric curve, so a zero correlation says nothing about nonlinear dependence.

Conclude X and Y are independent

Why: False. Y = X² with X symmetric about 0 has correlation ≈ 0 yet Y is FULLY determined by X — maximal dependence hiding behind a zero correlation.

The fix

Correlation ≈ 0 rules out only a linear relationship. To test real independence, compute the mutual information.

Use I(X;Y) — it captures ANY dependence

Why: MI = 0 iff truly independent, full stop. It sees the parabola that correlation misses, so it is the honest independence test.

For Y = X²: correlation ≈ 0 but I = 0.9183 bits (large)

Why: The MI equals H(Y) here — X pins Y down completely. Next slide computes both numbers so you see the gap with your own eyes.

65. Break it on purpose: zero correlation ⇒ MI = 0

Break the constraint

Discussion prompt

The rule this trap just fixed:

MI = 0 iff truly independent, full stop. It sees the parabola that correlation misses, so it is the honest independence test.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

Correlation only measures LINEAR association. It is blind to any symmetric curve, so a zero correlation says nothing about nonlinear dependence.

66. Restore the missing line: Y = X²: correlation ≈ 0, MI large

Fill the middle

Fill in the blanks

From Y = X²: correlation ≈ 0, MI large — one line has had its right-hand side removed. Put it back.

import numpy as np

# X uniform on 1/3; Y = X^2 in ___. Fully determined, yet...
Jd = np.zeros((3, 2))
Jd[0, 1] = 1/3 # X=-1 -> Y=1
Jd[1, 0] = 1/3 # X= 0 -> Y=0
Jd[2, 1] = ___ # X= 1 -> Y=1
px, py = Jd.sum(1), Jd.sum(0)
m = Jd > 0
mi = np.sum(Jd[m] * np.log2(Jd[m] / np.outer(px, py)[m]))

xs = np.random.default_rng(0).uniform(-1, 1, 100000)
corr = np.corrcoef(xs, xs**2)[0, 1]
print("correlation(X, X^2) =", round(corr, 4)) # ~0
print("I(X; X^2) =", round(mi, 4), "bits") # 0.9183

Why: Jd[2, 1] is what everything below it consumes, so the wrong expression here fails later and somewhere else. MI equals H(Y) = H([1/3, 2/3]) = 0.9183 — the maximum, since X determines Y outright.

67. Y = X²: correlation ≈ 0, MI large

Worked example

Let X be uniform on {−1, 0, 1} and Y = X². Y is a deterministic function of X, so they're maximally dependent — but the parabola is symmetric, so correlation collapses:

import numpy as np

# X uniform on {-1, 0, 1}; Y = X^2 in {0, 1}. Fully determined, yet...
Jd = np.zeros((3, 2))
Jd[0, 1] = 1/3        # X=-1 -> Y=1
Jd[1, 0] = 1/3        # X= 0 -> Y=0
Jd[2, 1] = 1/3        # X= 1 -> Y=1
px, py = Jd.sum(1), Jd.sum(0)
m = Jd > 0
mi = np.sum(Jd[m] * np.log2(Jd[m] / np.outer(px, py)[m]))

xs = np.random.default_rng(0).uniform(-1, 1, 100000)
corr = np.corrcoef(xs, xs**2)[0, 1]
print("correlation(X, X^2) =", round(corr, 4))   # ~0
print("I(X; X^2)           =", round(mi, 4), "bits")  # 0.9183

correlation ≈ 0 (about −0.006), yet I(X; X²) = 0.9183 bits

Why: MI equals H(Y) = H([1/3, 2/3]) = 0.9183 — the maximum, since X determines Y outright. Correlation is fooled by symmetry; MI is not. This is the exam's favorite gotcha.

measurevaluesees nonlinear dependence?
correlation(X, X²)≈ 0no — blind to the parabola
I(X; X²)0.9183 bitsyes — flags full dependence
H(Y) (the ceiling)0.9183 bitsMI is maxed out

68. Application: feature selection

Section

Part 6 of 7 — rank features by MI

69. Rank features by MI with the label

Concept

A feature is useful when knowing it reduces uncertainty about the label — exactly I(feature; label). Rank features by their MI with y and keep the top ones; the rest are noise.

Unlike a correlation filter, an MI filter catches nonlinear predictors (the X²-style features a linear screen throws away). sklearn.feature_selection.mutual_info_classif estimates I for continuous features against a discrete label.

70. Predict the next row: MI finds the signal features

Pattern

Predict first

The table runs: 0 | 0.574 | y + noise | informative · 1 | 0.011 | pure noise | noise · 2 | 0.688 | 2y−1 + noise | informative · 3 | 0.004 | pure noise | noise

In MI finds the signal features, given the rows so far: what is the next one — the row where feature is 4?

Correct: 4 | 0.000 | pure noise | noise

featureMI scorebuilt asverdict
00.574y + noiseinformative
10.011pure noisenoise
20.6882y−1 + noiseinformative
30.004pure noisenoise
40.000pure noisenoise

Why: The relationship between the columns, not the individual numbers, is what generates the next row. The two features built from y light up; the three np.random noise columns land near zero.

71. MI finds the signal features

Worked example

Build 5 features where only columns 0 and 2 carry signal about a binary label; the other three are pure noise. Let MI expose them:

import numpy as np
from sklearn.feature_selection import mutual_info_classif

rng = np.random.default_rng(0)
n = 500
y  = rng.integers(0, 2, n)
f0 = y + 0.3*rng.normal(size=n)         # informative
f2 = 2*y - 1 + 0.3*rng.normal(size=n)   # informative
X  = np.c_[f0, rng.normal(size=n), f2, rng.normal(size=(n, 2))]
scores = mutual_info_classif(X, y, random_state=0)
print(scores.round(3))                  # [0.574 0.011 0.688 0.004 0.]

Signal features 0 and 2 score ~0.57 and ~0.69; the rest score ~0

Why: The two features built from y light up; the three np.random noise columns land near zero. MI cleanly separates informative from useless — no linearity assumed. (nats here, since sklearn uses natural log — the ranking is what matters.)

featureMI scorebuilt asverdict
00.574y + noiseinformative
10.011pure noisenoise
20.6882y−1 + noiseinformative
30.004pure noisenoise
40.000pure noisenoise

72. Watch it run: MI finds the signal features

Pattern

Step through it

Step through MI finds the signal features one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: feature is 0
  2. Step 2: feature is 1
  3. Step 3: feature is 2
  4. Step 4: feature is 3
  5. Step 5: feature is 4

73. Rebuild the recipe: The mutual-information toolkit

Ranking

Put in order

These are the steps of The mutual-information toolkit, scrambled. Put them back in order before the next slide shows you.

  1. Definition: I(X;Y) = Σ p(x,y) log₂[p(x,y)/(p(x)p(y))] — the master expression
  2. Form 1 (entropies): H(X)+H(Y)−H(X,Y) — the set-overlap
  3. Form 2 (conditional): H(Y)−H(Y|X) — uncertainty reduced by seeing X
  4. Form 3 (KL): D_KL(p(x,y)‖p(x)p(y)) — distance from independence
  5. Bounds: 0 ≤ I ≤ min(H(X),H(Y)); I = 0 ⟺ X ⊥ Y (catches nonlinear, unlike correlation)
  6. Uses: rank features by I with the label; maximize a lower bound on I via InfoNCE (CLIP)

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

74. The mutual-information toolkit

Pattern

  1. Definition: I(X;Y) = Σ p(x,y) log₂[p(x,y)/(p(x)p(y))] — the master expression
  2. Form 1 (entropies): H(X)+H(Y)−H(X,Y) — the set-overlap
  3. Form 2 (conditional): H(Y)−H(Y|X) — uncertainty reduced by seeing X
  4. Form 3 (KL): D_KL(p(x,y)‖p(x)p(y)) — distance from independence
  5. Bounds: 0 ≤ I ≤ min(H(X),H(Y)); I = 0 ⟺ X ⊥ Y (catches nonlinear, unlike correlation)
  6. Uses: rank features by I with the label; maximize a lower bound on I via InfoNCE (CLIP)

75. Where does it stop working: The mutual-information toolkit

Edge cases

Discussion prompt

The mutual-information toolkit works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Definition: I(X;Y) = Σ p(x,y) log₂[p(x,y)/(p(x)p(y))] — the master expression
  2. Form 1 (entropies): H(X)+H(Y)−H(X,Y) — the set-overlap
  3. Form 2 (conditional): H(Y)−H(Y|X) — uncertainty reduced by seeing X
  4. Form 3 (KL): D_KL(p(x,y)‖p(x)p(y)) — distance from independence
  5. Bounds: 0 ≤ I ≤ min(H(X),H(Y)); I = 0 ⟺ X ⊥ Y (catches nonlinear, unlike correlation)
  6. Uses: rank features by I with the label; maximize a lower bound on I via InfoNCE (CLIP)

76. Check yourself — what MI = 0 means

Check

Interpret a zero. Remember MI is a KL divergence.

Check your understanding

I(X;Y) = 0 implies:

  • A. X and Y are independent (correct)
  • B. X and Y are uncorrelated but possibly dependent
  • C. X perfectly predicts Y
  • D. X and Y have equal entropy

Answer: A

Why: I = 0 means the KL between the joint and the product of marginals is zero, so p(x,y) = p(x)p(y) everywhere — exact independence. MI captures all dependence, so zero MI is the strongest independence statement.

Why B tempts people
That is what zero CORRELATION allows. Zero MI is strictly stronger — it rules out nonlinear dependence too, so 'possibly dependent' is exactly what it forbids.
Why C tempts people
Perfect prediction means MAXIMAL MI (I = H(Y), the ceiling), the opposite of zero. Confusing 'no shared information' with 'all shared information'.
Why D tempts people
Equal entropy is unrelated to MI; two independent variables can have wildly different entropies, and two dependent ones can have equal entropy.

77. Rule out three: Check yourself — MI vs correlation

Elimination

Eliminate the wrong options

Y = X² with X symmetric about 0. Compared with correlation, mutual information will:

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. be large (it detects the dependence) while correlation is ≈ 0
  • B. also be ≈ 0, like the correlation
  • C. be negative, since the parabola dips
  • D. equal the correlation exactly

Survives elimination: A

Why: Correlation measures only linear association, which cancels to ≈ 0 for the symmetric parabola. But Y is fully determined by X, so I(X;Y) = H(Y) is large (0.9183 bits in our demo). MI sees the nonlinear dependence correlation misses.

78. Check yourself — MI vs correlation

Check

Which measure sees nonlinear structure?

Check your understanding

Y = X² with X symmetric about 0. Compared with correlation, mutual information will:

  • A. be large (it detects the dependence) while correlation is ≈ 0 (correct)
  • B. also be ≈ 0, like the correlation
  • C. be negative, since the parabola dips
  • D. equal the correlation exactly

Answer: A

Why: Correlation measures only linear association, which cancels to ≈ 0 for the symmetric parabola. But Y is fully determined by X, so I(X;Y) = H(Y) is large (0.9183 bits in our demo). MI sees the nonlinear dependence correlation misses.

Why B tempts people
MI is NOT zero here — X fully determines Y, so they are maximally dependent. Only correlation is fooled by the symmetry; assuming MI behaves like correlation is the whole trap.
Why C tempts people
MI is a KL divergence, so it is always ≥ 0 — it can never be negative, regardless of the shape of the relationship.
Why D tempts people
They differ precisely in this case — that is the entire point. Correlation ≈ 0, MI large; if they always matched, MI would add nothing.

79. Check yourself — the conditional form

Check

Relate MI to conditional entropy — Form 2.

Check your understanding

Which identity is correct?

  • A. I(X;Y) = H(Y) − H(Y|X) (correct)
  • B. I(X;Y) = H(Y) + H(Y|X)
  • C. I(X;Y) = H(X,Y) − H(X)
  • D. I(X;Y) = H(Y|X)

Answer: A

Why: MI is the reduction in uncertainty about Y from knowing X: start with H(Y), subtract what remains, H(Y|X). We derived it from Form 1 by regrouping. Symmetrically it also equals H(X) − H(X|Y).

Why B tempts people
Adding the conditional entropy gives more than H(Y), which cannot be the information GAINED — MI SUBTRACTS the leftover uncertainty, it does not add it.
Why C tempts people
H(X,Y) − H(X) equals H(Y|X), the conditional entropy itself (the chain rule) — that is the leftover uncertainty, not the mutual information.
Why D tempts people
H(Y|X) is the uncertainty REMAINING after knowing X — the opposite of the information gained. MI is H(Y) minus this, not this.

80. Rule out three: Check yourself — the upper bound

Elimination

Eliminate the wrong options

What is the tightest general upper bound on I(X;Y)?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. min(H(X), H(Y))
  • B. H(X) + H(Y)
  • C. H(X,Y)
  • D. 1 bit, always

Survives elimination: A

Why: From I = H(Y) − H(Y|X) with H(Y|X) ≥ 0, we get I ≤ H(Y); by symmetry I ≤ H(X). So I ≤ min(H(X), H(Y)), attained when one variable fully determines the other (H(Y|X) = 0).

81. Check yourself — the upper bound

Check

How big can MI get? Use Form 2 and H(Y|X) ≥ 0.

Check your understanding

What is the tightest general upper bound on I(X;Y)?

  • A. min(H(X), H(Y)) (correct)
  • B. H(X) + H(Y)
  • C. H(X,Y)
  • D. 1 bit, always

Answer: A

Why: From I = H(Y) − H(Y|X) with H(Y|X) ≥ 0, we get I ≤ H(Y); by symmetry I ≤ H(X). So I ≤ min(H(X), H(Y)), attained when one variable fully determines the other (H(Y|X) = 0).

Why B tempts people
H(X) + H(Y) equals H(X,Y) + I, which is generally far larger than I — it is an upper bound, but a very loose one, not the tightest.
Why C tempts people
H(X,Y) = H(X) + H(Y) − I ≥ I in typical cases but is not the standard tight bound; the sharp ceiling is min(H(X), H(Y)).
Why D tempts people
1 bit is only the ceiling when both variables are single fair bits. For higher-entropy variables MI can exceed 1 bit, so a fixed 1-bit cap is wrong.

82. Application: InfoNCE & CLIP

Section

Part 7 of 7 — contrastive learning

83. Learn representations by contrast

Intuition

CLIP and SimCLR don't have labels — they have pairs that belong together: an image and its caption, or two crops of the same photo. The trick is to make matched pairs' embeddings similar and mismatched pairs' embeddings different.

Doing so provably maximizes a lower bound on the mutual information between the two views. Contrastive learning is mutual-information maximization in disguise — which is why this lesson ends here.

84. The InfoNCE loss

Concept

For a true pair (x, y) among a batch of candidates {z_j} (the negatives plus the true y), InfoNCE is a softmax cross-entropy that says 'the true partner should win the similarity contest':

\[ \mathcal{L} = -\log \frac{\exp\!\big(\mathrm{sim}(x,y)/\tau\big)}{\sum_j \exp\!\big(\mathrm{sim}(x,z_j)/\tau\big)} \]

sim is cosine similarity; τ is a temperature that sharpens the softmax. Minimizing L pulls the true pair together and pushes the negatives apart — the exact cross-entropy machinery of Lesson 17, applied to representation learning.

85. What temperature τ does

Intuition

The temperature τ divides every similarity before the softmax. Small τ (like 0.1) multiplies the logits up (by 10), sharpening the softmax so it rewards the single best match harshly — the model is picky.

Large τ flattens the logits toward a uniform softmax, so the loss barely distinguishes the true partner from the negatives — the model gets lazy. CLIP tunes τ as a learnable parameter; too small overfits sharp boundaries, too large learns nothing.

In our trace τ = 0.1, so a perfect cosine of 1 becomes a logit of 10 — enormous, which is why aligned pairs drove the loss almost to zero.

86. Teach it back: What temperature τ does

Explain it

Discussion prompt

Explain What temperature τ does to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

In our trace τ = 0.1, so a perfect cosine of 1 becomes a logit of 10 — enormous, which is why aligned pairs drove the loss almost to zero.

87. What has to be given first: InfoNCE traced on a tiny batch

Missing information

Discussion prompt

Three image embeddings, three text embeddings, matched by index. Build the similarity matrix S = (img·txtᵀ)/τ; the correct partner for row i is column i (the diagonal):

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

Each row is a 3-way classification: which text matches this image? The target is the diagonal. Rows where the true partner is NOT the most similar (like row 2 below) pay a big loss — exactly the pairs the gradient will fix.

88. InfoNCE traced on a tiny batch

Worked example

Three image embeddings, three text embeddings, matched by index. Build the similarity matrix S = (img·txtᵀ)/τ; the correct partner for row i is column i (the diagonal):

import numpy as np

np.random.seed(0)
d = 4
img = np.random.randn(3, d)
txt = np.random.randn(3, d)
img = img / np.linalg.norm(img, axis=1, keepdims=True)   # unit vectors
txt = txt / np.linalg.norm(txt, axis=1, keepdims=True)
tau = 0.1
S = img @ txt.T / tau                                     # 3x3 logits

def softmax(r):
    e = np.exp(r - r.max()); return e / e.sum()

loss = np.mean([-np.log(softmax(S[i])[i]) for i in range(3)])
print("S =\n", S.round(2))
print("InfoNCE loss =", round(loss, 4))                  # 1.6539

Row loss = −log(softmax(row)[correct index])

Why: Each row is a 3-way classification: which text matches this image? The target is the diagonal. Rows where the true partner is NOT the most similar (like row 2 below) pay a big loss — exactly the pairs the gradient will fix.

row (image)correct probrow loss
00.99920.0008
10.68370.3802
20.01024.5806
mean—1.6539

89. Watch it run: InfoNCE traced on a tiny batch

Pattern

Step through it

Step through InfoNCE traced on a tiny batch one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: row (image) is 0
  2. Step 2: row (image) is 1
  3. Step 3: row (image) is 2
  4. Step 4: row (image) is mean

90. Guess the shape of the answer: When embeddings align, the loss collapses

Estimation

Predict first

Now make each image embedding equal its text partner (perfect alignment). The diagonal similarities dominate, softmax puts nearly all mass on the true index, and the loss drops toward zero:

Commit before you compute: what does When embeddings align, the loss collapses come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Diagonal = 1/τ = 10 (cosine of a vector with itself is 1); loss = 0.0188

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Aligned pairs push the correct softmax probability near 1, so −log(≈1) ≈ 0.

91. When embeddings align, the loss collapses

Worked example

Now make each image embedding equal its text partner (perfect alignment). The diagonal similarities dominate, softmax puts nearly all mass on the true index, and the loss drops toward zero:

import numpy as np

np.random.seed(1)
d = 4
emb = np.random.randn(3, d)
emb = emb / np.linalg.norm(emb, axis=1, keepdims=True)
tau = 0.1
S = emb @ emb.T / tau               # img and txt already agree

def softmax(r):
    e = np.exp(r - r.max()); return e / e.sum()

loss = np.mean([-np.log(softmax(S[i])[i]) for i in range(3)])
print("diagonal of S =", np.diag(S).round(1))   # [10. 10. 10.]
print("aligned loss  =", round(loss, 4))         # 0.0188

Diagonal = 1/τ = 10 (cosine of a vector with itself is 1); loss = 0.0188

Why: Aligned pairs push the correct softmax probability near 1, so −log(≈1) ≈ 0. Compare 0.0188 (aligned) with 1.6539 (random) and log(3) = 1.0986 (worst-case uniform guessing) — training drives representations from random toward aligned.

batch statemean lossmeaning
random embeddings1.6539partners not yet matched
worst case (uniform)1.0986 = log 3pure guessing among 3
aligned embeddings0.0188true partner wins every row

92. Work backwards from the answer: When embeddings align, the loss collapses

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Diagonal = 1/τ = 10 (cosine of a vector with itself is 1); loss = 0.0188

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Now make each image embedding equal its text partner (perfect alignment). The diagonal similarities dominate, softmax puts nearly all mass on the true index, and the loss drops toward zero:

93. Why InfoNCE is a lower bound on MI

Concept

With a batch of N candidates, minimizing InfoNCE L maximizes a bound tying it to the mutual information between the two views x and y:

\[ I(X;Y) \;\ge\; \log_2 N - \mathcal{L} \]

Driving the loss L down pushes the lower bound on I(X;Y) up — so contrastive training literally increases the mutual information between an image and its caption. The bound tightens with bigger batches (log N), which is why CLIP trains on enormous batches of negatives.

94. By analogy: Why InfoNCE is a lower bound on MI

Analogy

Discussion prompt

Explain Why InfoNCE is a lower bound on MI by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

With a batch of N candidates, minimizing InfoNCE L maximizes a bound tying it to the mutual information between the two views x and y:

95. The compression link

Concept

Information theory ties straight to compression. Shannon's source coding theorem says entropy H(X) is the minimum average bits per symbol to encode X — you cannot compress below its entropy without losing information.

Minimum description length reframes learning as compression: the best model encodes the data in the fewest bits. Lower entropy and higher MI mean more exploitable structure — a model that finds high-MI features has literally found something compressible.

96. Your turn: compute MI

Section

Project — MI from scratch

97. Project: mutual information from a joint

Concept

Build MI from a joint distribution by two of the three formulas, confirm they match, then confirm an independent joint gives exactly 0. You've derived every piece — now assemble it yourself.

#requirementtool
1marginals + their entropiesJ.sum(axis), entropy(p)
2MI two ways, both equaldefinition (Form 1) / KL (Form 3)
3independent joint → MI = 0np.outer(px, py)

Build rules: type every line yourself, mask out zero probabilities in the entropy (log 0 is undefined), and use log2 so the answer is in bits.

98. Break it if you can: Project: mutual information from a joint

Counterexample

Discussion prompt

Build MI from a joint distribution by two of the three formulas, confirm they match, then confirm an independent joint gives exactly 0. You've derived every piece — now assemble it yourself.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Build rules: type every line yourself, mask out zero probabilities in the entropy (log 0 is undefined), and use log2 so the answer is in bits.

99. Milestone 1 — marginals & entropies

Worked example

Your turn: get the marginals and their entropies from the joint. Predict H(X) for a 50/50 marginal before you print.

Hint: px = J.sum(1), py = J.sum(0); entropy(p) = −Σ p log₂ p over p > 0.

import numpy as np

def entropy(p):
    p = np.asarray(p, dtype=float); p = p[p > 0]
    return -np.sum(p * np.log2(p))

J = np.array([[0.4, 0.1], [0.1, 0.4]])
px, py = J.sum(1), J.sum(0)
print(px, py)                                  # [0.5 0.5] [0.5 0.5]
print(round(entropy(px), 4), round(entropy(py), 4))   # 1.0 1.0
quantityvalue
px, py[0.5, 0.5], [0.5, 0.5]
H(X)1.0
H(Y)1.0

100. Milestone 2 — MI two ways

Worked example

Your turn: compute MI by the entropy definition (Form 1) and by the KL form (Form 3). Predict: do they match?

Hint: Form 1 is entropy(px)+entropy(py)−entropy(J.ravel()); Form 3 is np.sum(J*np.log2(J/np.outer(px,py))).

import numpy as np

def entropy(p):
    p = np.asarray(p, dtype=float); p = p[p > 0]
    return -np.sum(p * np.log2(p))

J = np.array([[0.4, 0.1], [0.1, 0.4]])
px, py = J.sum(1), J.sum(0)
mi_def = entropy(px) + entropy(py) - entropy(J.ravel())
mi_kl  = np.sum(J * np.log2(J / np.outer(px, py)))
print(round(mi_def, 4), round(mi_kl, 4))       # 0.2781 0.2781
formulavalue (bits)
Form 1 (entropies)0.2781
Form 3 (KL)0.2781

101. Milestone 3 — independence gives 0

Worked example

Your turn: build an independent joint with np.outer(px, py) and confirm its MI is 0 — the proof of the I = 0 ⟺ independent test.

Hint: for Ind = np.outer(px, py), the joint already equals the product, so every log-ratio is log(1) = 0.

import numpy as np

J = np.array([[0.4, 0.1], [0.1, 0.4]])
px, py = J.sum(1), J.sum(0)
Ind = np.outer(px, py)                         # independent joint
ipx, ipy = Ind.sum(1), Ind.sum(0)
mi = np.sum(Ind * np.log2(Ind / np.outer(ipx, ipy)))
print(round(mi, 6))                            # 0.0
jointMI (bits)
dependent J0.2781
independent outer(px, py)0.0

102. What each one costs: Milestone 3 — independence gives 0

Trade off

Comparison matrix

From Milestone 3 — independence gives 0: every row here is a choice with a cost. Fill the MI (bits) column, then say which row you would actually pick and what you give up for it.

jointMI (bits)
dependent J0.2781
independent outer(px, py)0.0

103. The full program

Concept

import numpy as np

def entropy(p):
    p = np.asarray(p, dtype=float); p = p[p > 0]
    return -np.sum(p * np.log2(p))

def mutual_information(J):
    px, py = J.sum(1), J.sum(0)
    return np.sum(J * np.log2(J / np.outer(px, py)))   # = H(X)+H(Y)-H(X,Y)

J = np.array([[0.4, 0.1], [0.1, 0.4]])
px, py = J.sum(1), J.sum(0)
print('MI(dependent)   =', round(mutual_information(J), 4))            # 0.2781
print('MI(independent) =', round(mutual_information(np.outer(px, py)), 6))  # 0.0
outputvalue
MI(dependent)0.2781
MI(independent)0.0

If MI agrees across the formulas and zeroes out for an independent joint — you can measure dependence the way representation learning does.

104. Fill in: value for The full program

Comparison

Comparison matrix

From The full program: refill the value column from what you know. The rest of the table is as it appeared.

outputvalue
MI(dependent)0.2781
MI(independent)0.0

105. Show it off

Concept

Out loud, slides closed: explain (1) the three roads from the definition to Forms 1, 2, and 3, (2) why zero MI is strictly stronger than zero correlation, and (3) what InfoNCE pulls together and pushes apart.

Stretch (homework): prove I(X;Y) ≥ 0 from D_KL ≥ 0, and use mutual_info_classif to rank the features of a 100-feature dataset. Mutual information returns in CLIP and contrastive pre-training (Week 47).

106. Connect it up: Lesson 29: Mutual Information & Contrastive Learning

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Recap: entropy & one running joint · Mutual information: three faces · Form 2: conditional entropy · Form 3: distance from independence · Bounds & the independence test · Application: feature selection. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

107. What you can do now

Recap

ideathe one thing to remember
MI = shared informationone number, three forms — all = 0.2781 here
Form 3 (KL)distance of the joint from p(x)p(y)
MI = 0⟺ independent — catches nonlinear, correlation can't
bounds0 ≤ I ≤ min(H(X), H(Y))
feature selectionrank by I(feature; label)
InfoNCEsoftmax over similarities = MI lower bound (CLIP)

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 29 (Week 10 — Mutual Information & Contrastive Learning) — Barron · USAAIO Round 2 Preparation, 2026
  2. Cover, Thomas — Elements of Information Theory, Ch. 2 (Entropy, Relative Entropy, and Mutual Information) — Wiley, 2nd ed.
  3. van den Oord, Li, Vinyals — Representation Learning with Contrastive Predictive Coding (InfoNCE)
  4. scikit-learn mutual_info_classif
  5. Every entropy, MI value, MI score, and InfoNCE loss produced by real execution — numpy 2.2.6 + scikit-learn 1.9.0, each block run standalone, July 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108