Chapter 20 of Trappe & Washington: Shannon's measure of uncertainty and what it settles about secrecy. Covers the four requirements that force the entropy formula, joint and conditional entropy with the chain rule and the three standard inequalities, Huffman codes and the H <= L < H+1 compression bound, perfect secrecy defined as H(P|C) = H(P) with the one-time pad proof and the general two-condition theorem, the entropy of English measured by Shannon's prediction experiment, redundancy, and unicity distance — plus the boundary the chapter draws, that RSA has H(P|C) = 0 and is secure anyway.
Subject: Cryptography · 68 slides · diagram-first lesson
Open the interactive version of this deck
Title
Cryptography · Chapter 20
Shannon's measure of uncertainty, and what it means to say a ciphertext reveals nothing
Objectives
This course has said 'reveals nothing' many times. This chapter makes the phrase into an equation, and then checks which systems satisfy it.
Figure (svg): Four experiments ordered by how uncertain their outcome is, with the entropy of each.
Warm-up
Roll a standard six-sided die. Let A be the event that the number of dots is odd, and B the event that it is at least 3.
Discussion prompt
Compare how much you learn from being told A, from being told B, and from being told A ∩ B.
Hint: Count the possibilities each one leaves open.
Answer:
A leaves three possibilities — 1, 3, 5. B leaves four — 3, 4, 5, 6. A ∩ B leaves two — 3 and 5.
So A ∩ B tells you more, and the reason is that it leaves you less uncertain. Information and uncertainty are two views of the same quantity: information gained is uncertainty removed.
That is the whole plan of the chapter. Define a number that measures uncertainty, and then 'this ciphertext reveals nothing about the plaintext' becomes 'the uncertainty about the plaintext is unchanged by seeing the ciphertext' — a statement about two numbers being equal.
And the payoff is that vague security claims become checkable. Chapter 4 asserted that the one-time pad is unbreakable; here that claim is proved, and the proof shows exactly which two properties are doing the work — properties that can then be looked for in other systems.
Section
Section 20.1 · pp. 370-372
Concept
An experiment X with outcomes in a finite set 𝒳. Each outcome gets a probability, written p_X(x) = p(X = x), and they sum to 1. The outcome is called a random variable.
Two experiments can be bundled into a joint random variable Z = (X, Y) with outcomes in 𝒳 × 𝒴, and the joint probability is written p_{X,Y}(x, y).
The individual distribution is recovered by summing out, p_X(x) = Σ_y p_{X,Y}(x, y) — the marginal.
Independent — X and Y are independent when p_{X,Y}(x, y) = p_X(x)p_Y(y) for every x and y. Equivalently, when p_Y(y | x) = p_Y(y): knowing X does not change the odds for Y.
That last phrasing is the one that matters here. Independence is exactly the statement that one variable carries no information about the other, and perfect secrecy will turn out to be independence of plaintext and ciphertext.
Worked example
Draw a card from a standard deck. Let X be the suit and Y the value.
𝒳 has four elements and 𝒴 has thirteen
Why: So the joint variable Z = (X, Y) has 52 outcomes, one per card.
All cards are equally likely, so p(X = x, Y = y) = 1/52
Why: Each specific card, once.
And p(X = x) = 1/4 while p(Y = y) = 1/13
Why: Thirteen cards per suit, four cards per value.
\[ \tfrac{1}{52} = \tfrac{1}{4} \cdot \tfrac{1}{13} \]
Verify: so suit and value are independent
Why: Which matches intuition: learning that the card is a heart tells you nothing about whether it is a seven. The factorisation of the joint probability is the formal content of that sentence.
Figure (svg): Two joint distributions side by side: suit and value of a card, which factor, and two die events, which do not.
Worked example
Roll a die. Let X = 1 if the number of dots is odd and 0 otherwise. Let Y = 1 if the number is at least 2 and 0 otherwise.
p(Z = (1, 0)) is the probability that the roll is odd and less than 2
Why: That is the single outcome 1, so the probability is 1/6.
But p(X = 1) = 1/2 and p(Y = 0) = 1/6
Why: Three odd faces out of six; one face below 2.
\[ \tfrac{1}{6} \ne \tfrac{1}{2} \cdot \tfrac{1}{6} = \tfrac{1}{12} \]
This is the case that matters for cryptanalysis. A cipher is broken to the extent that ciphertext and plaintext fail to be independent, and the rest of the chapter is about measuring that failure.
Verify: so X and Y are dependent
Why: And the dependence is real: if you learn Y = 0 then the roll is 1, so X = 1 with certainty. Knowing Y has told you something about X, which is precisely what the failed factorisation records.
Figure (svg): Two joint distributions side by side: suit and value of a card, which factor, and two die events, which do not.
Concept
When p_X(x) > 0, the conditional probability of Y = y given X = x is
\[ p_Y(y \mid x) = \frac{p_{X,Y}(x, y)}{p_X(x)} \]
Read it as a restriction. Confine attention to the cases where X = x; those have total probability p_X(x), and the fraction of that total coming from Y = y is the conditional probability.
Bayes's theorem reverses the conditioning:
\[ p_X(x \mid y) = \frac{p_X(x) \, p_Y(y \mid x)}{p_Y(y)} \]
The proof is to write both sides out from the definition — there is nothing to it. Its importance is directional: an attacker usually knows how plaintexts turn into ciphertexts, p(c | m), and wants p(m | c). Bayes is the bridge, and every ciphertext-only attack in Chapter 2 was an informal use of it.
Prediction
Predict first
Are X and Y independent?
Correct: No — knowing the first flip changes the odds for the total
The general trap: independence of the underlying experiments does not survive taking functions of them. Y = f(X, second flip) inherits information from X by construction.
And this is why joint entropy needs its own definition. H(X, Y) is not generally H(X) + H(Y); the gap is exactly the information the two share, and it is zero only when they are independent.
Why: If X is heads, then Y is 1 or 2 — it cannot be 0. So p(Y = 0 | X = heads) = 0 while p(Y = 0) = 1/4. The flips are independent of each other; X and Y are not, because Y is built from X.
Section
Section 20.2 · pp. 372-378
Concept
Rather than invent a formula, Shannon wrote down what any measure of uncertainty ought to satisfy, and derived the formula from those.
The fourth is the substantive one. It says uncertainty is additive across a two-stage experiment, discounted by how often you reach the second stage. Everything else follows from it.
Figure (svg): Shannon's four requirements on any measure of uncertainty, from which the entropy formula follows uniquely.
Worked example
The book's illustration is worth doing, because the requirement is the least obvious of the four.
Record only whether a die roll is even or odd
Why: Two equally likely outcomes, so the uncertainty is H(½, ½) = 1 bit.
Now split 'even' into the suboutcomes 2 and {4, 6}
Why: Three outcomes altogether: 2 with probability ⅙, {4, 6} with probability ⅓, and odd with probability ½.
Within the even branch, the split is ⅓ versus ⅔
Why: Because 2 is one of the three even faces and {4, 6} is the other two.
\[ H\bigl(\tfrac16, \tfrac13, \tfrac12\bigr) = H\bigl(\tfrac12, \tfrac12\bigr) + \tfrac12 \, H\bigl(\tfrac23, \tfrac13\bigr) \]
Verify: the added term is weighted by ½ because you only face that choice half the time
Why: Numerically: 1.459 = 1 + 0.5 × 0.918. And this additivity is what forces a logarithm — it is the one function turning the multiplication of probabilities into the addition of uncertainties.
Figure (svg): Requirement four illustrated: splitting the even outcomes of a die into two suboutcomes and accounting for the added uncertainty.
Concept
Theorem. Any H satisfying the four requirements must have the form −λ Σ p_k log₂ p_k for some constant λ ≥ 0.
Taking λ = 1 and base-2 logarithms defines the entropy:
\[ H(X) = -\sum_{x \in \mathcal X} p(x) \log_2 p(x) \]
Since log₂ p(x) ≤ 0, entropy is never negative — there is no such thing as negative uncertainty. And the convention 0 log₂ 0 = 0 handles impossible outcomes, justified by the limit of x log₂ x as x → 0.
The units are bits, because the logarithm is base 2. Equivalently, H(X) is the expected value of −log₂ p(X): the average surprise, where an outcome of probability p carries surprise −log₂ p.
The theorem is what makes entropy non-arbitrary. It is not a convenient formula that behaves well; it is the only formula that behaves well, which is why the same quantity turns up in compression, in thermodynamics and in cryptography.
Figure (svg): Shannon's four requirements on any measure of uncertainty, from which the entropy formula follows uniquely.
Worked example
The three calculations everything else is compared against.
A fair coin: two outcomes at ½ each
Why: H = −(½ log₂ ½ + ½ log₂ ½) = 1 bit. The flip delivers exactly one bit of information.
A biased coin with probability p of heads
Why: H(X) = −p log₂ p − (1−p) log₂ (1−p), maximised at p = ½ and falling to 0 at either end. A coin that always lands heads carries no information.
An n-sided fair die: n outcomes at 1/n each
Why: H = −n · (1/n) log₂ (1/n) = log₂ n. So a six-sided die is 2.58 bits and a ten-sided die 3.32.
\[ H\bigl(\tfrac1n, \ldots, \tfrac1n\bigr) = \log_2 n \]
Verify: both intuitions from the opening are captured by one number
Why: More outcomes raises log₂ n; a flatter distribution over fixed outcomes raises H toward that maximum. Entropy is the single quantity that orders every experiment by uncertainty.
Figure (svg): The binary entropy function, peaking at one bit when the two outcomes are equally likely.
Worked example
A second reading of the same number, and the one that makes it concrete.
Flip two coins and let X be the number of heads
Why: Outcomes 0, 1, 2 with probabilities ¼, ½, ¼.
H(X) = −(¼ log₂ ¼ + ½ log₂ ½ + ¼ log₂ ¼) = 3/2
Why: One and a half bits.
Now count yes-no questions. Ask “is there exactly one head?”
Why: Half the time the answer is yes and you are done.
The other half of the time, ask “are there two heads?”
Why: That settles it. So the average is ½ · 1 + ½ · 2 = 3/2 questions.
Verify: the average number of questions equals the entropy
Why: Which is the bridge to Section 20.3: an optimal question strategy is an optimal binary code, and the entropy is the number of bits per symbol it needs. Entropy measures storage as directly as it measures uncertainty.
Figure (svg): Four experiments ordered by how uncertain their outcome is, with the entropy of each.
Notation
Every term in the sum is doing a specific job.
Annotate
On: \( H(X) = -\sum_x p(x) \log_2 p(x) \)
That last note is the chapter's most important caveat, and the source of most misuse of the word 'entropy' outside information theory.
Concept
Two variables need two more definitions, and the second is where cryptography enters.
\[ H(X, Y) = -\sum_x \sum_y p_{X,Y}(x,y) \log_2 p_{X,Y}(x,y) \]
The joint entropy is just the entropy of the bundled variable Z = (X, Y) — nothing new.
\[ H(Y \mid X) = \sum_x p_X(x) \, H(Y \mid X = x) \]
The conditional entropy is the uncertainty remaining in Y once X is known. It is a weighted average over the possible values of X, weighted by how likely each is.
The weighting is essential and easy to get wrong. The book warns that the unweighted sum −Σ p_Y(y|x) log₂ p_Y(y|x) is not conditional entropy and does not behave like one: for independent X and Y it would make the uncertainty in Y increase when X is revealed, which is nonsense.
Figure (svg): The chain rule for entropies shown as a decomposition of the joint uncertainty into two parts.
Worked example
The identity that ties the three quantities together.
Start from H(X, Y) = −ΣΣ p(x,y) log₂ p(x,y)
Why: The definition.
Substitute p(x,y) = p_X(x) p_Y(y|x)
Why: The definition of conditional probability, rearranged.
The logarithm of a product splits into a sum, so the whole expression splits into two double sums
Why: One involving log p_X(x), the other log p_Y(y|x).
In the first, sum over y first: Σ_y p(x,y) = p_X(x)
Why: That double sum collapses to −Σ_x p_X(x) log₂ p_X(x) = H(X). The second is H(Y|X) by definition.
\[ H(X, Y) = H(X) + H(Y \mid X) \]
Verify: the uncertainty in both equals the uncertainty in X plus what remains in Y once X is known
Why: It is a bookkeeping identity, exact rather than approximate, and it is the engine of the perfect-secrecy proof in Section 20.4.
Figure (svg): The chain rule for entropies shown as a decomposition of the joint uncertainty into two parts.
Concept
Three results, all of which have plain-English readings.
H(X) ≤ log₂ |𝒳| — Uncertainty is greatest when all outcomes are equally likely, with equality exactly in that case. The uniform distribution is the most uncertain one available.
H(X, Y) ≤ H(X) + H(Y) — The information in a pair is at most the sum of the information in each, because the two may overlap. Equality exactly when they are independent.
H(Y | X) ≤ H(Y) — Conditioning reduces entropy. Learning X can only reduce your uncertainty about Y — it can never make you more uncertain.
The third follows from the first two and the chain rule in one line: H(X) + H(Y|X) = H(X, Y) ≤ H(X) + H(Y), and cancel H(X).
Note carefully what the third does not say. On average, learning X reduces uncertainty about Y. A particular value of X can raise it — you can receive a message that leaves you more confused than before. It is the weighted average that cannot increase, which is exactly why the weighting in the definition was not optional.
Figure (svg): The chain rule for entropies shown as a decomposition of the joint uncertainty into two parts.
Socratic
The book explicitly warns against defining H(Y|X) as the unweighted sum over x and y.
Discussion prompt
What goes wrong with the unweighted version, and what does the weighting represent?
Hint: Try it on independent X and Y.
Answer:
Take X and Y independent, so p_Y(y|x) = p_Y(y) for every x. The unweighted sum then adds one copy of H(Y) for each value of x, giving |𝒳| · H(Y).
Which says that learning an irrelevant fact makes you more uncertain — by a factor equal to the number of things it could have been. Plainly wrong, and wrong in a way that would break every later result.
The weights p_X(x) are the fix, and they are not a normalisation trick: they say that the uncertainty remaining when X = x should count in proportion to how often X = x actually happens.
With them, independence gives Σ_x p_X(x) H(Y) = H(Y), which is the correct answer — an irrelevant fact leaves your uncertainty exactly where it was.
The general moral is worth keeping: a conditional quantity in probability is almost always an expectation over the conditioning variable, and definitions that forget the expectation tend to fail on the independent case first. It is a cheap and reliable sanity check.
Definition probe
Some hold always, some only under conditions.
Sort into buckets
Sort each statement.
Faded example
Four blanks.
Fill in the blanks
Entropy is H(X) = −Σ p(x) log₂ p(x), measured in bits when the logarithm is base 2. Conditional entropy H(Y|X) is a weighted average of H(Y | X = x) over the values of X. The chain rule says H(X, Y) = H(X) + H(Y|X).
Why: The weighting in the third blank is the detail the book flags explicitly, and the one that makes conditioning-reduces-entropy come out right.
Section
Section 20.3 · pp. 378-381
Concept
Shannon's original motivation was compression, and the connection to entropy is direct: the average bits needed per symbol is essentially the entropy.
The example. Four letters a, b, c, d with frequencies 0.5, 0.3, 0.1, 0.1. Fixed-length coding gives every letter two bits, for an average of 2.
But frequencies differ, so give the common letters short codes: a = 1, b = 01, c = 001, d = 000. The average is 1(0.5) + 2(0.3) + 3(0.1) + 3(0.1) = 1.7 bits.
Morse code was the same idea by hand. Morse asked printers which letters were used most and gave them the shortest symbols — e is a single dot, t a single dash, while x is −··− and z is −−··.
Huffman's algorithm makes this optimal rather than clever. List the outputs with their probabilities, assign 0 and 1 to the two smallest, merge them into a single output with the combined probability, and repeat until one remains. Read the codes backwards through the merges.
Figure (svg): The Huffman tree for four symbols with probabilities one half, three tenths, one tenth and one tenth.
Worked example
The four-letter example, step by step.
Start with a: 0.5, b: 0.3, c: 0.1, d: 0.1
Why: The two smallest are c and d, tied — the tie is broken arbitrarily.
Assign 1 to c and 0 to d, and merge into a node of probability 0.2
Why: The list is now a: 0.5, b: 0.3, {c,d}: 0.2.
The two smallest are now b and {c,d}
Why: Assign 1 to b and 0 to {c,d}, merging into a node of probability 0.5.
Two nodes remain, a: 0.5 and the merged node: 0.5
Why: Assign 1 to a and 0 to the other. Done.
Read backwards: a got only a 1, so a = 1
Why: b got a 1 at its merge and a 0 above it, so b = 01. c got 1, then 0, then 0 — c = 001. And d = 000.
Verify: average length 1.7, against an entropy of 1.685
Why: Within 0.015 bits of the theoretical floor, and no code can do better than the entropy. The fixed-length code needed 2 bits, so this is a 15% saving on a four-letter alphabet.
Figure (svg): The Huffman tree for four symbols with probabilities one half, three tenths, one tenth and one tenth.
Concept
A useful feature of Huffman encoding is that a message can be read one letter at a time, left to right, with no lookahead.
The string 011000 can only be bad. After two bits — 01 — the first letter is already known to be b, because no other codeword starts 01.
Reverse the codewords and it breaks. With b = 10 and c = 100, the message 101000 could begin bb or ba, and nothing is decidable until the end.
Worse, a careless assignment destroys uniqueness altogether. If a were coded 0 rather than 1, then aaa and d would both be 000 and the code would not be decodable at all.
Huffman's construction avoids both faults automatically, because every symbol ends up at a leaf of the tree: no codeword is a prefix of another. That is a structural consequence of merging from the bottom, not something the algorithm checks for.
Figure (svg): Why the prefix-free property matters: one code can be decoded left to right, the reversed one cannot.
Concept
Theorem. If L is the average number of bits per output for the Huffman code of X, then H(X) ≤ L < H(X) + 1.
The lower bound says entropy is a floor: no code, Huffman or otherwise, can average fewer bits per symbol than the entropy. Compression has a hard limit, and the limit is an information-theoretic quantity.
The upper bound says Huffman is close to it — within one bit per symbol, always. In the example, H = 1.685 and L = 1.7.
The gap comes from integrality. Codewords are whole numbers of bits, while the ideal length −log₂ p(x) generally is not. Coding blocks of several symbols at once amortises the rounding, which is how arithmetic coding beats Huffman.
And this closes the loop with the question-counting reading. Entropy is the average number of yes-no questions, an optimal code is an optimal question strategy, and Huffman builds one greedily from the bottom up.
Figure (svg): The Huffman tree for four symbols with probabilities one half, three tenths, one tenth and one tenth.
Anomaly
In the example, c and d both had probability 0.1, and the choice of which got 0 was arbitrary.
Predict first
Does the choice matter?
Correct: No — the average length is the same either way, though the codes differ
Ties higher up the tree can change the code lengths, giving genuinely different codes with the same average — one might use lengths 1, 2, 3, 3 and another 2, 2, 2, 2 for a suitable distribution. Both are optimal.
Practically the arbitrariness matters for interoperability, which is why deployed formats either transmit the code table or fix a canonical tie-breaking rule so that both ends build the identical tree — DEFLATE does the latter.
Why: Swapping the labels swaps c and d's codewords, both of which have length 3. The average length is unchanged, and both codes are prefix-free. Huffman codes are not unique; their optimal average length is.
Estimation
English text stored one byte per character, with an entropy of roughly 1 bit per letter.
Predict first
Roughly what compression ratio should a good compressor achieve?
Correct: About 8 to 1
The chapter's own estimate is more conservative: comparing 4.7 bits (random letters) with about 1 bit gives roughly 4:1 of redundancy within the alphabet, before counting the wasted bits of an 8-bit byte.
And note the cryptographic reading of the same number. The redundancy that compression removes is exactly what a cryptanalyst uses to recognise a correct decryption. Compressing before encrypting removes the tell — good practice, and the reason unicity distance rises when you do it.
Why: Eight bits per character stored against about one bit of actual information gives a ceiling near 8:1. Real general-purpose compressors reach 3:1 or 4:1 on English prose, and specialised models get closer to the bound — the gap measures how much structure the model fails to capture.
Real world
The quantity is not confined to textbooks.
Discussion prompt
Name four places a working engineer computes an entropy.
Hint: Compression, passwords, randomness testing, and machine learning.
Answer:
Compression. Huffman coding is inside DEFLATE, so it runs in gzip, PNG and ZIP; JPEG and MP3 use it for their final entropy-coding stage. Arithmetic and range coders beat it by fractions of a bit and are used where that matters.
Password and key strength. 'This password has 40 bits of entropy' is this exact quantity computed over the generating distribution — and it is routinely misapplied, because entropy is a property of the process, not of the string it produced. A password generated by flipping coins has entropy; the same characters typed by a person do not.
Randomness testing. Entropy estimators are the first-line health check on a hardware RNG, and NIST SP 800-90B specifies min-entropy estimates for exactly this. Chapter 5's failures are what they exist to catch.
Machine learning. Cross-entropy is the standard loss for classification, and decision trees split on information gain — H(Y) − H(Y|X), which is this chapter's third inequality used as an objective function.
The through-line: whenever the question is 'how much does this tell me', the answer is a difference of entropies, and the field it appears in changes only the names.
Section
Section 20.4 · pp. 381-385
Concept
A cipher system has plaintexts 𝒫, ciphertexts 𝒞 and keys 𝒦. Each plaintext has some probability of occurring; the key is chosen independently of the plaintext.
The question: if Eve intercepts a ciphertext, how much has her uncertainty about the plaintext dropped? That is H(P) − H(P | C).
Perfect secrecy — A cryptosystem has perfect secrecy when H(P | C) = H(P). Seeing the ciphertext leaves the uncertainty about the plaintext exactly where it was.
Equivalently, P and C are independent — which is the sharpest way to say it. The ciphertext is not merely hard to invert; it is statistically unrelated to the message.
Note what the definition does not mention: computation, time, or the attacker's resources. This is a statement about distributions, and it holds against an adversary of unlimited power.
Figure (svg): A cipher whose ciphertext reduces the uncertainty in the plaintext, contrasted with one where it does not.
Worked example
The book's example makes the leak numerical.
Plaintexts a, b, c with probabilities .5, .3, .2; keys k₁, k₂ each with probability .5
Why: And the encryption table: k₁ sends a, b, c to U, V, W; k₂ sends them to U, W, V.
p(U) = .5(.5) + .5(.5) = .50, and p(V) = p(W) = .25
Why: U arises from plaintext a under either key, which is already suspicious.
Seeing U determines the plaintext outright — it must be a
Why: Both keys send a to U and nothing else to U.
Seeing V gives p(b | V) = (.3)(.5)/.25 = .6 and p(c | V) = .4
Why: The prior odds were .3 and .2; the ciphertext has revised them.
H(P) = −(.5 log₂ .5 + .3 log₂ .3 + .2 log₂ .2) = 1.485 bits
Why: The uncertainty before the interception.
Verify: H(P | C) = 0.485 bits, so exactly one bit has leaked
Why: Compute .5 · H(1) + .25 · H(.6, .4) + .25 · H(.6, .4) = 0 + .25(.971) + .25(.971) = .485. The cipher gave away a full bit of a 1.485-bit message — and note the source: only two keys for three plaintexts.
Figure (svg): A cipher whose ciphertext reduces the uncertainty in the plaintext, contrasted with one where it does not.
Concept
Theorem. The one-time pad has perfect secrecy.
The setup: an alphabet of Z letters, plaintexts and ciphertexts of length L, and Z^L keys — one shift per position — each chosen with probability Z^{−L}.
Step one: every ciphertext is equally likely. For each plaintext x and ciphertext c there is exactly one key sending x to c, so summing over the (x, k) pairs that produce c gives
\[ p_C(c) = \frac{1}{Z^L} \sum_{x \in \mathcal P} p_P(x) = \frac{1}{Z^L} \]
Step two: therefore H(K) = H(C) = log₂(Z^L), since keys and ciphertexts are both uniform over Z^L possibilities.
Figure (svg): The one-time pad proof: computing the joint entropy of plaintext, key and ciphertext two different ways.
Worked example
The argument computes H(P, K, C) twice and equates the answers.
Knowing P and K determines C, so the triple carries no more information than the pair (P, K)
Why: Hence H(P, K, C) = H(P, K).
And P and K are independent, so H(P, K) = H(P) + H(K)
Why: The joint-entropy inequality, at equality because of independence.
Knowing P and C determines K, since the pad is c − p
Why: Hence H(P, K, C) = H(P, C) as well.
And by the chain rule H(P, C) = H(P | C) + H(C)
Why: Conditioning on C first this time.
\[ H(P) + H(K) = H(P \mid C) + H(C) \]
Verify: since H(K) = H(C), the two cancel and H(P | C) = H(P)
Why: Perfect secrecy. And notice exactly which facts were used: every key equally likely, and each plaintext-ciphertext pair reachable by exactly one key. Those two properties are the entire content of the theorem.
Figure (svg): The one-time pad proof: computing the joint entropy of plaintext, key and ciphertext two different ways.
Concept
The proof used nothing about shifts, so it generalises immediately.
Theorem. A cryptosystem has perfect secrecy if (1) every key has probability 1/#𝒦, and (2) for each plaintext x and ciphertext c there is exactly one key k with e_k(x) = c.
Condition (2) forces #𝒞 = #𝒦, and with #𝒫 = #𝒞 = #𝒦 the converse holds too: perfect secrecy implies both conditions.
Which delivers the bad news of Chapter 4 as a theorem. Perfect secrecy requires at least as many keys as messages, so the key must be as long as everything you will ever send. That is not a weakness of the one-time pad's design; it is a lower bound on any system with the property.
And it explains why the leaky example leaked. Three plaintexts, two keys — condition (2) could not possibly hold, and the missing key is exactly the bit that escaped.
Figure (svg): A cipher whose ciphertext reduces the uncertainty in the plaintext, contrasted with one where it does not.
Anomaly
RSA is believed secure. Consider its conditional entropy.
Predict first
What is H(P | C) for RSA?
Correct: Zero — the ciphertext completely determines the plaintext
And this is not a criticism of RSA. Entropy does not know about computation. It says the information is present; it says nothing about the billions of years needed to extract it.
So information-theoretic security and computational security are different properties, and almost every deployed cipher has the second and not the first. AES has H(P|C) = 0 too, once the ciphertext is longer than the key.
The relevant measure for RSA is computational complexity, which is where Chapter 9 lives. Information theory sets what is possible with unlimited power; complexity theory sets what is possible with realistic power, and modern cryptography is built almost entirely on the second.
Which is why the one-time pad, alone in having the stronger property, is nevertheless almost never used. The property costs a key as long as the message, and the weaker property is enough.
Why: Given n, e and c, the plaintext is determined: there is exactly one m with mᵉ ≡ c. So no uncertainty remains and H(P|C) = 0. Every bit of the message is present in what Eve holds.
Socratic
H(P|C) = 0 for RSA even though nobody can recover P from C.
Discussion prompt
Is entropy the wrong tool, or is it answering a different question?
Hint: What does the quantity actually quantify?
Answer:
It is answering a different question. Entropy quantifies how much information is present in a random variable, and information presence has nothing to do with extraction cost.
The analogy: a locked safe with the combination written inside. The information is fully determined by the safe's contents; the difficulty is opening it. Entropy describes the contents and is silent about the lock.
This is a strength when you want an unconditional guarantee. A one-time pad's security does not depend on any assumption about the adversary's machine, which no computational argument can promise — and which is why the property is worth having when the key-management cost is affordable.
And it is a weakness when you want a practical one, because almost everything useful is only computationally secure. The right response is not to abandon entropy but to use it for what it settles: it decides what unlimited power could do, and complexity decides the rest.
A concrete instance of the split appears in Chapter 17 and 19. Shamir's shares and Alice's ignorance of which square root Bob used are information-theoretic; Bob's inability to lie and RSA's security are computational. Naming which kind a guarantee is tells you exactly what future advances could break.
Definition probe
Security claims made across the course.
Sort into buckets
Sort each by which kind of security it is.
Section
Section 20.5 · pp. 385-392
Concept
Four successive estimates, each using more context than the last.
Each drop is conditioning reducing entropy, measured. Seeing q makes u almost certain; seeing studyin makes g almost certain. The letters are not independent, so context keeps paying.
The entropy of English is defined as the limit of H(Lᴺ)/N as N → ∞ — the average information per letter in a very long text, or equivalently the uncertainty in the next letter given everything before it.
Figure (svg): Successive estimates of the information per letter of English, falling as more context is taken into account.
Concept
The limit cannot be computed directly — tabulating 100-gram frequencies is impossible. Shannon's method sidesteps the problem entirely.
Imagine an optimal predictor. Given the text so far, it guesses the next letter. If right, record 1. If wrong, it guesses again; if right this time, record 2. And so on.
The sequence of numbers determines the text, because the predictor is deterministic: knowing it guessed correctly on its second attempt tells you which letter that was. So text and number sequence have the same entropy — and the numbers can be tallied.
Shannon proposed using a person as the predictor, on the grounds that a fluent English speaker is close to optimal. On itissunnytoday the guess counts run 2, 1, 1, 1, 4, 3, 2, 1, 4, 1, 1, 1, 1, 1 — mostly ones, which is the redundancy showing up directly.
\[ \sum_{i=1}^{26} i(q_i - q_{i+1}) \log_2 i \;\le\; H_{\text{English}} \;\le\; -\sum_{i=1}^{26} q_i \log_2 q_i \]
Figure (svg): Shannon's prediction experiment: a text turned into a sequence of guess-counts that can be turned back into the text.
Worked example
The numbers from Table 20.1, with spaces included and punctuation ignored, so 27 symbols.
Of 102 guesses, 79 were right first time, 8 took two tries, 3 took three, and so on
Why: So q₁ = 79/102, q₂ = 8/102, q₃ = 3/102, q₄ = q₅ = 2/102, q₆ = 3/102, and a handful of singletons.
The upper bound is the entropy of the q distribution
Why: −(79/102 log₂ 79/102 + ⋯) ≈ 1.42 bits.
The lower bound is Shannon's other formula
Why: 1·(79/102 − 8/102) log₂ 1 + 2·(8/102 − 3/102) log₂ 2 + ⋯ ≈ 0.72 bits. The first term vanishes since log₂ 1 = 0.
So the entropy of English is somewhere near 1 bit per letter, perhaps slightly above
Why: The bounds are approximate — there is experimental error, and the limit as N → ∞ is being estimated from one finite sample.
Verify: Huffman-coding the guess numbers gives 171 bits for 102 letters, or 1.68 per letter
Why: Against 510 bits for a naive five-bits-per-letter encoding. And 1.68 is within one bit of the estimate, exactly as the H ≤ L < H + 1 theorem requires.
Figure (svg): Shannon's prediction experiment: a text turned into a sequence of guess-counts that can be turned back into the text.
Concept
Comparing what a letter could carry with what it does carry gives one number that will do a great deal of work.
\[ R = 1 - \frac{H_{\text{English}}}{\log_2 26} \;\approx\; 0.75 \]
English is about 75% redundant. A long message in standard written English is roughly four times longer than its optimally compressed form.
Redundancy is what makes cryptanalysis possible. Every attack in Chapter 2 — frequency analysis, digram counts, guessing the — worked because English is far from random. A cryptanalyst recognises the right key because only the right key produces text with English's statistics.
Which suggests a countermeasure the book returns to: compress before encrypting. Compression removes redundancy, so the compressed plaintext looks nearly random, and there is no longer a statistical signal that says 'this decryption is correct'.
Figure (svg): English redundancy shown as the gap between the alphabet's capacity and the information actually carried.
Concept
Given a ciphertext, how many keys decrypt it to something meaningful? For a short message, several. For a long one, presumably just the right one.
Unicity distance n₀ — The length of ciphertext at which a unique meaningful plaintext is expected.
\[ n_0 = \frac{\log_2 |\mathcal K|}{R \log_2 |\mathcal L|} \]
The reading is a balance of two quantities. The numerator is how many bits of key uncertainty must be eliminated. The denominator is how many bits of redundancy each ciphertext letter supplies to eliminate them. Divide, and you get the number of letters needed.
So more keys raise it and more redundant plaintext lowers it — which is why compressing before encrypting raises the unicity distance, and why a cipher's key space is only half of what determines its resistance to ciphertext-only attack.
Figure (svg): Unicity distance for three ciphers, showing how many ciphertext letters are needed before the decryption is unique.
Worked example
With R = 0.75 and log₂ 26 = 4.70, so the denominator is about 3.5.
Substitution cipher: 26! keys, so log₂ 26! ≈ 88.4 bits
Why: n₀ = 88.4 / 3.5 ≈ 25.1 letters.
So about 25 letters of ciphertext usually pin down a unique plaintext
Why: Which matches practice — a two-line substitution cryptogram has one solution, and a five-letter one does not.
Affine cipher: 312 keys, so log₂ 312 ≈ 8.3 bits
Why: n₀ ≈ 2.35 letters. A very rough figure, and clearly a few more than two are needed in practice, but the message is right: almost nothing suffices.
One-time pad on a message of length N: 26^N keys
Why: n₀ = N log₂ 26 / 3.5 ≈ 1.33N.
Verify: for the one-time pad, n₀ exceeds the message length
Why: You would need more ciphertext than exists before a unique decryption became likely — which never happens. That is unicity distance's way of stating perfect secrecy: every plaintext of the right length remains possible, forever.
Figure (svg): Unicity distance for three ciphers, showing how many ciphertext letters are needed before the decryption is unique.
Prediction
Predict first
What happens to the unicity distance?
Correct: It rises sharply, and in the limit becomes infinite
Which is the practical lesson. Compressing before encrypting genuinely strengthens a system against ciphertext-only attack, at no cost in key length. It is one of the rare free improvements in the subject.
With one large caveat, from Chapter 14: compression before encryption leaks length, and when an attacker can inject chosen data into the same compressed stream, the length becomes an oracle. The CRIME and BREACH attacks on TLS did exactly that.
So the advice is conditional, and the condition is about the threat model: compression helps against a passive analyst and can be fatal against an active one who controls part of the plaintext. Both facts follow from the same property — that compressed length depends on content.
Why: R is in the denominator, so as R approaches 0 the unicity distance grows without bound. With no redundancy there is nothing for the cryptanalyst to recognise: every key produces a plausible-looking output, and none can be ruled out.
Discrimination
Five properties of a cryptosystem.
Sort into buckets
Sort each.
Matching
Five definitions from the chapter.
Match the pairs
Why: The third is mutual information, and it is the quantity a cryptanalyst is trying to make large and a designer trying to make zero — perfect secrecy is exactly the statement that it is zero. The fourth and fifth are the applied consequences: redundancy is what a cryptanalyst exploits, and unicity distance is how much ciphertext they need before they can.
Trade off
The chapter's central distinction. Fill the blanks.
Comparison matrix
| Information-theoretic | Computational | |
|---|---|---|
| H(P | C) | equals H(P) | zero |
| Assumes about the adversary | nothing | bounded resources, unproved hardness |
| Key length required | at least the message length | fixed and short |
| Survives a quantum computer | yes | depends on the problem |
| Used in practice | rarely — diplomatic links, some key distribution | almost everywhere |
The stronger guarantee is available and almost nobody buys it, because the price is fixed by a theorem rather than by engineering. That is a genuinely unusual situation in the subject.
Edge cases
n₀ = log₂|𝒦| / (R log₂|ℒ|) gave 25.1 for a substitution cipher and 2.35 for an affine cipher.
Discussion prompt
What is the formula assuming, and where does it mislead?
Hint: It is called a rough estimate for a reason.
Answer:
It assumes plaintexts are typical English with a fixed redundancy, and it treats all keys as equally likely and all wrong decryptions as equally likely to look meaningless. Real texts vary, and real wrong decryptions vary too.
It is an average, not a guarantee. It says a unique plaintext is expected at that length, not that no shorter ciphertext is ever uniquely decipherable and not that every longer one is.
The book flags the affine case explicitly: 2.35 letters is clearly too few in practice, and a handful more are needed. The estimate captures the order of magnitude, which is what makes it useful, and no more.
And at 25 letters for a substitution cipher there is another effect: several letters will not have appeared at all, so several keys decrypt to the same plaintext. The plaintext becomes unique before the key does — and the formula does not distinguish them.
Where it is most reliable is at the extremes, which is where it is used: 1.33N for a one-time pad is not a rough estimate but a statement of principle, and it is the honest way to say that the requirement is never met.
Ranking
Five distributions.
Put in order
Why: 0, then about 0.47, then 1, then about 1 to 1.5 for English in context, then 4.7 for a uniform letter. The interesting comparison is the last two: the same alphabet carries 4.7 bits when random and about 1 in real text, and that gap of nearly 4 bits per letter is the redundancy every classical cryptanalytic technique lives on.
Constraint
A system needs user secrets with at least 60 bits of entropy.
Discussion prompt
Specify the policy, and say what entropy is a property of.
Hint: Entropy belongs to the generating process, not the string.
Answer:
Entropy is a property of the process, not the output. 'correct horse battery staple' has no entropy as a string; it has entropy because of how it was chosen. Any policy that scores the characters typed is measuring the wrong thing.
So generate, do not judge. Four words drawn uniformly from a 7776-word list give 4 log₂ 7776 ≈ 51.7 bits; five words give 64.6. That is the calculation, and it depends only on the list size and the draw count.
Random characters: log₂ 95 ≈ 6.55 bits per printable ASCII character, so ten characters give 65 bits — if they are genuinely uniform. Human-chosen ten-character passwords carry perhaps 20.
Composition rules make it worse. Requiring an uppercase letter and a digit shrinks the space and concentrates users on predictable patterns, so the measured entropy falls even as the rule looks stricter. NIST's current guidance drops these rules for exactly this reason.
And the honest measure for guessing resistance is min-entropy, −log₂(max p), not Shannon entropy: an attacker tries the most likely candidate first and does not care about the average. A distribution with high Shannon entropy and one very likely outcome is weak, and only min-entropy sees that.
Cost model
Four symbols, and each one is a design lever.
Annotate
On: \( n_0 = \frac{\log_2 |\mathcal K|}{R \log_2 |\mathcal L|} \)
The formula's practical message is that compression is worth more than key length against this class of attack, which is not what most people's intuition suggests.
Missing information
A common claim in security documentation.
Discussion prompt
What must be true for the number to mean anything?
Hint: Entropy is computed over a distribution — whose?
Answer:
Which distribution. Entropy is defined over a generating process. If the claim is about a key produced by a flawed RNG, the true distribution is not the assumed one and the number is fiction — Chapter 5's Debian and Dual_EC failures were exactly this.
Shannon entropy or min-entropy. For guessing resistance, min-entropy is the correct measure, and it can be far lower. A distribution that is 50% one value and uniform over 2¹²⁸ others has huge Shannon entropy and one bit of min-entropy.
Whether the bits are independent. 128 bits from a source with internal correlations may have far less usable entropy, which is why extractors and conditioning functions exist.
Whether it survived processing. Hashing a 20-bit input to 256 bits produces a 256-bit string with 20 bits of entropy. The length of the output says nothing — Chapter 11's low-entropy commitment trap.
And what it is protecting against. 128 bits is ample against brute force and irrelevant against a side channel, a bad protocol, or someone reading the key off a disk.
Two truths and a lie
Two of these overstate it.
Eliminate the wrong options
Which statement is correct?
Survives elimination: a
Why: The precise statement is about the distribution of the message content given the ciphertext, and it is unconditional in the adversary's power. Everything outside that — key handling, message length, traffic patterns, endpoint security — is untouched, which is why a perfectly secret cipher can still sit inside a thoroughly broken system.
Commit first
A one-time pad system with genuinely random pads, used exactly once, distributed by courier.
Predict first
What leaks?
Correct: Message length and timing, which the definition does not cover
Padding to a fixed length and sending constant-rate cover traffic are the countermeasures, and both are expensive — which is why most systems accept the leak.
The general point is the one the trap slide makes: a theorem protects exactly what it quantifies over. H(P|C) = H(P) quantifies over message content, and length is not content.
And note that the same gap appears in modern systems. Encrypted messengers leak message sizes and timing; encrypted web traffic leaks page fingerprints. The cryptography is not what fails.
Why: Content is protected unconditionally. But the ciphertext has the same length as the plaintext, so length is public, and when messages are sent is public too. Traffic analysis has produced actionable intelligence from nothing else many times over.
Error analysis
From a security review.
Annotate
Every error here is a confusion between a representation and the process that produced it, which is the single most common way information-theoretic language gets misused in practice.
Explain it to yourself
The unicity distance rises when redundancy falls.
Discussion prompt
Explain the mechanism, in terms of what a cryptanalyst is actually doing.
Hint: How does an analyst know a trial key is right?
Answer:
The analyst tries a key and looks at the output. With redundant plaintext, a wrong key produces something that is obviously not English, and the right key produces something that obviously is. The recognition is the attack.
Redundancy is what makes that recognition possible. English uses about 1 bit per letter out of a possible 4.7, so the vast majority of letter strings are not English, and a wrong decryption lands in that majority.
Compress the plaintext and the output looks nearly random, so a wrong key produces something statistically indistinguishable from what the right key produces. There is no signal to recognise, and the analyst has no way to rank candidate keys.
In the limit of perfect compression, every key yields a plausible plaintext, which is the one-time pad's situation reached from the other direction — and it is why n₀ goes to infinity as R goes to zero.
The caveat is Chapter 14's. Compressed length depends on content, so an attacker who can inject text into the same stream learns about the rest of it from the size — CRIME and BREACH. Against a passive analyst compression is a free win; against an active one it can be the whole vulnerability.
Figure (svg): English redundancy shown as the gap between the alphabet's capacity and the information actually carried.
Explain it
Someone asks what entropy is, and does not want a logarithm.
Discussion prompt
Explain it, and give them a way to check their understanding.
Hint: Yes-no questions.
Answer:
Entropy is the average number of yes-no questions needed to determine the outcome, when you ask as cleverly as possible.
A fair coin: one question, so one bit. A coin that always lands heads: no questions, so zero. A four-sided die: two questions — is it less than 3, then is it odd — so two bits.
Unequal probabilities let you do better than counting outcomes, because you ask about the likely cases first. Two coins with 0, 1, 2 heads takes 1.5 questions on average, not 2: ask 'exactly one head?' and half the time you are finished.
The check: ask them to work out the questions for a four-outcome distribution of 1/2, 1/4, 1/8, 1/8. The answer is 1.75, and the strategy is to ask about the most likely outcome first each time — which is Huffman's algorithm, arrived at by intuition.
And the payoff to mention: this is why compression and uncertainty are the same subject. The best code is the best question strategy written down, and the number of bits it needs is the amount of information there was to begin with.
Figure (svg): Four experiments ordered by how uncertain their outcome is, with the entropy of each.
Pattern
One quantity, and every result in the chapter is a statement about differences of it.
And the limitation runs through all of it: entropy counts information present and is silent about the cost of extracting it. RSA has H(P|C) = 0 and is secure; that single fact marks the boundary between this chapter and the rest of the subject.
Figure (svg): A cipher whose ciphertext reduces the uncertainty in the plaintext, contrasted with one where it does not.
Trap
The trap. Perfect secrecy is H(P|C) = H(P). RSA and AES both have H(P|C) = 0 — the ciphertext determines the plaintext completely. So by the chapter's own measure they leak everything, and a system that leaks everything is broken.
Both halves of the premise are true. The conclusion is wrong, and seeing why is the point of the chapter.
Entropy measures information present, not information obtainable. The plaintext is determined by the ciphertext in the same sense that the factorisation of a 2048-bit number is determined by the number: fully, and uselessly.
Almost every deployed cipher has H(P|C) = 0, and this is unavoidable. The general theorem says perfect secrecy needs as many keys as messages, so any cipher with a short key fails it by construction — the property is not something better engineering could supply.
The relevant measure is computational, and that is a different subject with different tools: reductions, hardness assumptions, and concrete security bounds. Chapter 9's analysis of RSA is that subject, and it is where security is actually argued.
The right use of information theory is where it is decisive: it proves the one-time pad unconditionally secure, it sets the compression floor, it fixes the minimum key length for perfect secrecy, and it predicts unicity distance. None of those needs a hardness assumption, which is exactly why they are worth having.
The habit worth taking: when you meet a security claim, ask which kind it is. Information-theoretic claims are unconditional and expensive; computational claims are cheap and rest on unproved assumptions. Confusing them produces both false alarms and false confidence.
Check
Work it out before clicking.
Check your understanding
For a random variable with n possible outcomes, when is H(X) largest?
Answer: B
Why: H(X) ≤ log₂|𝒳| with equality exactly for the uniform distribution — the first of the three inequalities. Uniformity is maximum uncertainty, which is why a good key is drawn uniformly and why the binary entropy curve peaks at p = ½.
Check
Consider what the proof actually used.
Check your understanding
Which two properties make the one-time pad perfectly secret?
Answer: B
Why: Those are exactly the two conditions of the general theorem, and the proof uses nothing else — not XOR, not shifts, not anything specific to the pad. Any cipher with both properties is perfectly secret.
Check
Consider what raises and lowers it.
Check your understanding
Which change most increases a cipher's unicity distance?
Answer: B
Why: R sits in the denominator, so driving redundancy toward zero sends n₀ toward infinity. The key space is inside a logarithm in the numerator, so doubling it adds a single bit and moves n₀ by a fraction of a letter.
Connect it up
The chapter is six quantities and the relations between them.
Draw it
Write out H(X), H(X,Y), H(Y|X), mutual information, R and n₀ — the definition of each in symbols, its plain-English reading, and one worked number from this chapter. Then draw the three inequalities and mark which are identities and which carry hypotheses. Finally, write the two conditions of the general perfect-secrecy theorem and, beside them, why RSA satisfies neither and is secure anyway.
That last line is the one to be able to say cleanly. It is the boundary between information-theoretic and computational security, and everything after this chapter lives on the far side of it.
Exit ticket
One question, about the chapter's central caveat.
Predict first
RSA has H(P|C) = 0. What does that tell you?
Correct: The ciphertext contains all the information needed to recover the plaintext, and entropy says nothing about the cost of doing so
Why: Given n, e and c there is exactly one m with mᵉ ≡ c, so no uncertainty remains and the conditional entropy is zero. Entropy measures information present, not work required — and RSA's security is entirely a claim about the work, which is a claim information theory is not equipped to make.
Recap
Shannon's framework, and what it settles.
Chapter 21 next. Elliptic curves: the group law on a cubic curve, why the discrete logarithm problem is harder there than in a finite field, and how that hardness buys equivalent security at a fraction of the key length.
Figure (svg): Unicity distance for three ciphers, showing how many ciphertext letters are needed before the decryption is unique.
Want this taught 1-on-1? Alexander tutors Cryptography — $55/session, free consultation.