Chapter 4 of Trappe & Washington: binary and ASCII encoding, the one-time pad and its XOR arithmetic, the total collapse that follows reusing a pad, perfect secrecy defined through conditional probability and proved for the pad, the counting bound that makes perfect secrecy impossible for any short-key system, and the ciphertext indistinguishability game that becomes the book's working definition of security.
Subject: Cryptography · 60 slides · diagram-first lesson
Open the interactive version of this deck
Title
Cryptography · Chapter 4
The only unbreakable cipher in this book, what it costs, and what breaks it the instant you economise
Objectives
Every cipher in Chapter 2 fell. This one does not, and the reason is not that nobody has been clever enough — it is a theorem. The chapter is short, and its real subject is what the word secure should mean.
Figure (svg): The provable security thread through the book: perfect secrecy, then a reduction to Diffie-Hellman, then the random oracle model.
Warm-up
Chapter 2 broke seven ciphers with statistics; Chapter 1 argued that key length bounds only brute force.
Discussion prompt
Write down a definition of 'this cipher cannot be broken' that does not mention how much computing power the attacker has.
Hint: Think about what Eve believes about the message before and after she sees the ciphertext.
Answer:
The definition the chapter arrives at: seeing the ciphertext does not change Eve's beliefs about the plaintext at all. Whatever she thought was likely before, she thinks is exactly as likely after.
Notice what this rules out. It is not 'Eve cannot compute the message in reasonable time' — that is a claim about her resources. It is 'the ciphertext contains no information about the message', which is a claim about the ciphertext.
A cipher meeting it is secure against an adversary with unlimited computing power. Only one cipher in this book qualifies.
Section
Section 4.1 · pp. 88-89
Concept
The one-time pad operates on bits, so the first job is turning a message into bits. ASCII assigns each character a number, and the number is written in binary.
\[ \texttt{A} = 65 = 1000001, \quad \texttt{Z} = 90 = 1011010, \quad \text{space} = 32 = 0100000 \]
The letters A to Z are 65 through 90, and space is 32. That is seven bits each, and the choice matters more than it looks — the encoding is not neutral, and Section 4.3 will pick it apart.
The message need not be text at all. A digitised video or audio signal is already a bit string, and the pad treats it identically.
Figure (svg): The ASCII codes for A, H, Z and space in binary, showing that letters begin with 1 and space begins with 0.
Worked example
Turn HA into bits, ready for the pad.
H is the 8th letter, so its ASCII code is 64 + 8 = 72
Why: A = 65 is the anchor to memorise; everything else is an offset from it.
72 in binary: 64 + 8 = 1001000
Why: Seven bits, because ASCII's printable range fits in seven.
A = 65 = 1000001
Why: The leading 1 and the trailing 1, with five zeros between.
\[ \texttt{HA} \;\longmapsto\; 1001000\,1000001 \]
Verify: both blocks begin with 1, as every letter A-Z does
Why: 65 through 90 all lie between 64 and 127, so bit seven is always set. Space, at 32, is the exception — and that single structural fact is what breaks a reused pad two sections from now.
Figure (svg): The ASCII codes for A, H, Z and space in binary, showing that letters begin with 1 and space begins with 0.
Definition probe
Sort by what can be concluded from bit seven of a 7-bit ASCII block.
Sort into buckets
Classify each statement about a 7-bit block b.
Translation
Fluency here is worth ten minutes, because everything from Chapter 5 to Chapter 8 is stated in bits.
Match the pairs
Why: The capitals occupy 65 through 90, so every one of them has bit seven set; space at 32 does not. Reading these four off without a table is the skill Section 4.3's attack actually requires — the whole break turns on spotting which XOR blocks could only have come from a space.
Explain it to yourself
ASCII was designed in 1963 for teleprinters, with no thought of cryptanalysis.
Discussion prompt
Explain how a decision about character encoding ends up being a cryptographic weakness, and what a cryptographer would have chosen instead.
Hint: Look at where the letters sit relative to space, and what that does to the top bit.
Answer:
The weakness is structure in the plaintext alphabet. Capitals sit in 65-90 and space at 32, so one bit of every character is almost perfectly predictable. XOR two such characters and that bit announces whether exactly one of them was a space.
An encoding that scattered the characters uniformly across all 128 values would leak nothing from any single bit position, and the Section 4.3 attack would have to start from statistics rather than from certainty.
But the deeper point is that no encoding fixes it. Language has redundancy — Chapter 20 measures it at about 1.5 bits per letter against a possible 4.7 — and redundancy is what every ciphertext-only attack consumes. Changing the encoding relocates the redundancy; it does not remove it.
Which is why compression before encryption is genuinely useful, and why it is not sufficient: it reduces redundancy without eliminating the structure a determined attacker can find.
Section
Section 4.2 · pp. 89-91
Concept
Three requirements, and every one of them is load-bearing.
\[ c_i = m_i \oplus k_i, \qquad m_i = c_i \oplus k_i \]
Encryption XORs the key onto the message bit by bit. Decryption XORs the same key onto the ciphertext. There is no second algorithm, because XOR is its own inverse: m ⊕ k ⊕ k = m.
Figure (svg): The one-time pad encryption of 00101001 with key 10101100, XORed bit by bit to give 10000101.
Worked example
The book's own example. Message 00101001, key 10101100.
Column 1: 0 ⊕ 1 = 1
Why: The rules are 0⊕0 = 0, 0⊕1 = 1, 1⊕1 = 0 — addition mod 2, with no carry.
Column 2: 0 ⊕ 0 = 0. Column 3: 1 ⊕ 1 = 0. Column 4: 0 ⊕ 0 = 0
Why: No column affects any other. There is no diffusion here at all, and that is deliberate.
Continue through all eight columns
Why: Giving 10000101.
\[ 00101001 \oplus 10101100 = 10000101 \]
Verify: decrypt: 10000101 ⊕ 10101100 = 00101001
Why: The original message returns. XOR is an involution, so one routine serves both directions — which is also true of Enigma, and for the same structural reason.
Figure (svg): Decryption by XORing the same key onto the ciphertext, returning the original plaintext.
Picture it
Suppose the ciphertext is FIOWPSLQNTISJQL, from the letter-shift variant of the pad. The plaintext could be wewillwinthewar. It could equally be theduckwantsout.
Figure (svg): One fifteen-letter ciphertext with two completely different plaintexts, each requiring a different key that is equally likely.
Every 15-letter message is reachable from that ciphertext under some key, and every key is equally likely. So the ciphertext leaks nothing but the length — and Section 4.4 turns that observation into a proof.
Prediction
Eve somehow learns the first 20 bits of the message.
Predict first
What does this tell her about the rest?
Correct: The corresponding 20 bits of the key, and nothing more
The same reasoning covers chosen-plaintext and chosen-ciphertext attacks: each would reveal the part of the key used during the attack, and that part is useless unless it is reused.
Unless it is reused. Hold that clause — it is the whole of the next section.
Why: She recovers those 20 key bits exactly, by XORing the known plaintext with the ciphertext. But the key is random, so those bits are statistically independent of every other bit of the key — knowing them tells her literally nothing about bit 21. This is precisely what fails for every other cipher in the book, where the key is short and structured and a fragment constrains the whole.
Socratic
The theorem needs a truly random key. The book notes this is very difficult in practice, and mentions coin-flipping (too slow) and Geiger counters (needs care).
Discussion prompt
Why is 'looks random' not good enough, and what goes wrong with a Geiger counter used carelessly?
Hint: The proof uses the assumption that every key has probability exactly 1/N.
Answer:
The proof depends on uniformity, not on appearance. It computes P(C = c | M = m) = 1/N using the fact that every key is equally likely. A generator whose outputs merely pass statistical tests can still have a key distribution far from uniform, and the moment it does, the equality in the proof fails.
The Geiger counter problem is bias. Counting clicks and recording the parity gives 0 and 1 with unequal probability unless the counting interval is chosen carefully, and the radioactive source's decay rate itself drifts. A generator producing 0 with probability 0.52 is not a one-time pad, and Section 4.5 quantifies exactly how much such a bias costs.
The deeper point: randomness is a property of the process, not of the output. No amount of inspecting a bit string tells you whether it was generated uniformly, which is why this is an engineering problem that has repeatedly gone wrong in deployed systems.
And it is why Chapter 5 exists: pseudorandom generators are what people actually use, and they trade the proof for practicality.
Trade off
It is provably unbreakable and almost nobody uses it. Fill the blanks and the reason becomes obvious.
Comparison matrix
| Property | One-time pad | A block cipher such as AES |
|---|---|---|
| Key length | as long as the message | 128 or 256 bits, whatever the message length |
| Key reuse | never, not once | safe across many messages |
| Security guarantee | unconditional — a theorem | computational — no known attack |
| Resists unlimited computing power? | yes | no |
| Key distribution problem | gigabytes of key, delivered in advance | one short key, or a key agreement protocol |
The pad moves the entire problem into key distribution: to send a gigabyte securely you must already have securely sent a gigabyte. That circularity is why it is reserved for diplomatic and nuclear-command links where couriers are affordable.
Invariant
The claim is that each ciphertext bit is independent of its message bit. Step through what a random key bit does.
Step through it
Which step fails if the key bit is 1 with probability 0.6 instead of 0.5?
Both middle rows. With a biased key, P(c = 1) becomes 0.6 when m = 0 and 0.4 when m = 1, so the ciphertext now leans towards the truth — and the leaning accumulates over the message.
Matching
The pad has three hypotheses, and every real-world failure of a stream cipher violates one of them by name.
Match the pairs
Why: Three hypotheses, and between them they account for every documented break of a pad-shaped cipher. Notice that the first and last are the same violation reached by different routes — one by manufacturing error, one by a field that was sized without asking how often it would wrap. The second and third are not accidents at all: they are deliberate engineering trades, made knowingly, and they are what Chapter 5 studies.
Step zero
You hold C₁ and C₂, encrypted under the same keystream, and you know the plaintexts are capitals and spaces in 7-bit ASCII.
Discussion prompt
What is the very first thing you compute, and what do you look at in the result?
Hint: The key has to go before anything else can start.
Answer:
Compute C₁ ⊕ C₂. This is step zero, and it is the only step that involves the ciphertexts at all. Everything afterwards is a language problem.
Then split it into 7-bit blocks and read the leading bit of each. A leading 1 means exactly one of the two characters was a space, which pins that position in both messages down to 'space here, some letter there'.
And look for all-zero blocks, which mean the two messages agree in that position — often the start of a stereotyped opening shared by both.
The reason this ordering matters: attacking C₁ directly is hopeless, because it genuinely is a one-time pad ciphertext. The reuse is the only opening, so the first move must be the one that exploits it.
Figure (svg): Two ciphertexts made with the same pad, XORed together so the key cancels and the two plaintexts remain.
Section
Section 4.3 · pp. 91-94
Concept
Alice sends messages to Bob, Carla and Dante. She is lazy and uses the same key for all three. Eve does one XOR.
\[ C_1 \oplus C_2 = (M_1 \oplus K) \oplus (M_2 \oplus K) = M_1 \oplus M_2 \]
The key has vanished — not been weakened, vanished. Eve now holds M₁ ⊕ M₂, M₁ ⊕ M₃ and M₂ ⊕ M₃, and her problem no longer involves the key at all. It is a pure language problem of the kind Chapter 2 solved seven times.
This is why the word one-time is in the name. It is not advice; it is a hypothesis of the theorem, and dropping it does not degrade the guarantee gracefully.
Figure (svg): Two ciphertexts made with the same pad, XORed together so the key cancels and the two plaintexts remain.
Worked example
The book's method, using capitals and spaces in 7-bit ASCII. The lever is the leading bit.
Every letter A-Z has leading bit 1; space has leading bit 0
Why: So in M₁ ⊕ M₂, a block with leading bit 1 arises from a letter XORed with a space — the only combination whose leading bits differ.
Block three of M₁ ⊕ M₂ is 1100100, and block three of M₁ ⊕ M₃ is 1100001
Why: Both have leading bit 1, so position three is a space in one message and a letter in the other.
1100100 must be space ⊕ 1000100, so the letter is 1000100 = 68 = D... and the book's reading gives H
Why: Deduce which side the space is on by using both equations at once: since positions three of M₁ ⊕ M₂ and M₁ ⊕ M₃ are both letter-like, M₁ holds the space and M₂ and M₃ hold letters.
A block of 0000000 means the two messages agree in that position
Why: The first block of M₁ ⊕ M₂ is all zeros, so M₁ and M₂ start with the same letter.
Ambiguous positions need language statistics: 'space A A' and 'A space space' can give the same XORs
Why: Resolve them from the surrounding letters, exactly as in Chapter 2's substitution attack.
Verify: every recovered letter must make the other two messages read as English too
Why: Three messages constrain each other, so a wrong guess in one propagates into visible nonsense in the others. That mutual checking is what makes the attack reliable rather than speculative.
Figure (svg): The ASCII codes for A, H, Z and space in binary, showing that letters begin with 1 and space begins with 0.
Anomaly
Reusing a pad twice sounds like it should halve the security, or leak a little. It does not.
Predict first
What is the right way to describe the security of a twice-used pad?
Correct: The proof no longer applies at all; it is a running-key cipher and falls to language statistics
This is the general shape of provable security and it is worth over-learning: a proof buys you exactly the statement proved, under exactly its hypotheses, and nothing in the neighbourhood.
The practical corollary is that the hypotheses are the specification. 'Use a fresh random key of full length, once' is not guidance around the theorem — it is the theorem.
Why: Perfect secrecy is proved under the hypothesis that each key encrypts one message. Use the key twice and the hypothesis is false, so nothing is being claimed any more — and what is left, M₁ ⊕ M₂, is a running-key cipher that classical cryptanalysis handles. Security proofs are not quantities that erode; they are implications whose hypotheses either hold or do not.
Figure (svg): Two ciphertexts made with the same pad, XORed together so the key cancels and the two plaintexts remain.
Real world
Two-time pad is not a textbook curiosity. It has broken real systems, repeatedly.
Discussion prompt
Name a historical case and a modern one, and say what forced the reuse in each.
Hint: One is a Soviet cipher programme; one is a wireless standard from Chapter 14.
Answer:
Venona. Soviet diplomatic traffic in the 1940s used one-time pads, but wartime production pressure led to duplicate pad pages being issued. US cryptanalysts found the duplicates and read thousands of messages over decades. The mathematics was perfect; the manufacturing was not.
WEP, the subject of Section 14.3. Its 24-bit initialisation vector is far too short, so on a busy network the same IV — and therefore the same RC4 keystream — recurs within hours. Same key, two messages, keystream cancels.
And the general pattern: in both cases the reuse was forced by an engineering constraint, not chosen. Pads were expensive to produce; the IV field was sized to fit an existing header.
So the lesson is not 'do not reuse keys', which everyone already agrees with. It is that a design must make reuse structurally impossible, because wherever reuse is merely discouraged it eventually happens.
Elimination
A system XORs plaintext with a keystream generated from a shared key. Two messages have gone out under the same keystream.
Eliminate the wrong options
Which change prevents the problem recurring?
Survives elimination: c
Why: The only real fix is to guarantee that no keystream is ever produced twice, and that means a per-message nonce that is structurally incapable of repeating — a counter, not a random 24-bit field. This is exactly what CTR mode does in Chapter 6 and what WEP failed to do in Chapter 14. Note that (a) and (b) both make the attack harder while leaving it possible, which is the characteristic shape of a fix that is not a fix.
Ranking
Each is a way of producing the bits that get XORed onto the message.
Put in order
Why: The first is a reused pad and falls to language statistics with no search at all. The second has only about a million keys, so an attacker enumerates them; Proposition 4.4 already ruled out perfect secrecy, and here the gap is exploitable by hand. The third also has no perfect secrecy — the counting bound does not care how strong the generator is — but the search is infeasible, so it has computational security. Only the fourth satisfies the theorem's hypotheses. Note the jump between the third and fourth is a change in the kind of guarantee, not its size.
Figure (svg): Proposition 4.4 as a picture: perfect secrecy requires at least as many keys as messages.
Estimation
A generator emits 1 with probability 0.55 instead of 0.5. Eve sees a ciphertext bit and wants to guess the plaintext bit.
Predict first
Roughly how often is she right?
Correct: 55%
This is why the definition demands uniform keys rather than unpredictable-looking ones: a small, systematic departure from uniformity is not a small departure from security.
It also shows how the indistinguishability game degrades gracefully where perfect secrecy does not. Eve wins 0.55 rather than 0.5 — a quantifiable advantage — which is exactly the sort of statement the game-based definition is designed to make.
Why: The ciphertext bit is c = m ⊕ k. If Eve guesses that k was the more likely value (1), she is right 55% of the time, and her guess for m is right exactly when her guess for k is. So she wins 55% against the 50% baseline. That 5% edge looks small, but it applies independently to every bit, and over a long message the accumulated advantage lets her reconstruct the plaintext with near certainty.
Notation
Four symbols, and every one of them is doing work. It is the first formal security definition in the book.
Annotate
On: \( P(M = m \mid C = c) = P(M = m) \)
Compare with the computational version in later chapters: the equality becomes 'differs by a negligible amount', and 'every adversary' becomes 'every efficient adversary'.
Section
Section 4.4 · pp. 94-97
Concept
Before defining secrecy we need one tool. The conditional probability of B given A restricts attention to the cases where A happened.
\[ P(B \mid A) = \frac{P(A \cap B)}{P(A)} \]
The book's example: A is 'it rains Saturday', B is 'it rains Sunday'. These are not independent, since P(B | A) > P(B) — knowing about Saturday shifts your belief about Sunday.
Independent events — A and B are independent exactly when P(B | A) = P(B): learning A changes nothing about B. This is the form the definition of secrecy will take.
The cryptographic translation is direct. A is 'the ciphertext was c'. B is 'the message was m'. Secrecy should mean these are independent.
Figure (svg): A prior distribution over messages next to the posterior after seeing the ciphertext, identical bar for bar.
Concept
A cryptosystem has perfect secrecy when, for every plaintext m and every ciphertext c:
\[ P(M = m \mid C = c) = P(M = m) \]
In words: knowledge of the ciphertext never changes the probability that a given plaintext occurs. Eve's beliefs after eavesdropping are exactly her beliefs before.
Read the definition carefully for what it does not say. It does not say Eve cannot guess the message — she can, and if the message is one of two she will be right half the time. It says eavesdropping gave her no advantage in doing so. Her prior may already be very good; the ciphertext simply adds nothing to it.
Note also that plaintexts do not have equal probabilities. attack at noon is more probable than two plus two equals seven, and perfect secrecy preserves that skew rather than removing it.
Figure (svg): A prior distribution over messages next to the posterior after seeing the ciphertext, identical bar for bar.
Worked example
Assume N keys, each with probability 1/N. Two steps, and the first is the one that does the work.
Fix any plaintext m and any ciphertext c. If c is the ciphertext, the key must have been k = m ⊕ c
Why: XOR is invertible, so exactly one key maps this m to this c — not several, and not none.
So P(C = c | M = m) equals the probability that the key is m ⊕ c, which is 1/N
Why: Independent of both m and c. Every message reaches every ciphertext with the same probability.
Sum over messages: P(C = c) = Σ_m P(M = m)·P(C = c | M = m) = (1/N)·Σ_m P(M = m) = 1/N
Why: Using Σ P(M = m) = 1, since the message is certainly something.
\[ P(M = m \mid C = c) = \frac{P(C = c \mid M = m) \, P(M = m)}{P(C = c)} = \frac{(1/N) \, P(M = m)}{1/N} = P(M = m) \]
Verify: the 1/N cancels top and bottom for every m and c
Why: Which is the definition, so the pad has perfect secrecy. The whole proof rests on one fact: every message-to-ciphertext pair is achieved by exactly one key, and all keys are equally likely.
Figure (svg): A prior distribution over messages next to the posterior after seeing the ciphertext, identical bar for bar.
Fill the middle
One line of the proof is doing all the labour. Complete it.
Fill in the blanks
P(C = c \mid M = m) = P\bigl(K = m ⊕ c\bigr) = 1/N
Why: Because XOR is invertible, the key that carries m to c is uniquely determined as m ⊕ c — exactly one key, always. Since every key has probability 1/N, this probability is 1/N no matter which m and c you picked. That uniformity across all pairs is precisely what makes the ciphertext uninformative, and it is where the uniform-key hypothesis enters the proof.
Counterexample
The proof needs every key to have probability 1/N. Suppose instead one key has probability 1/2 and the rest share the remainder.
Discussion prompt
Show that perfect secrecy fails, and say what Eve does with the flaw.
Hint: Work out P(C = c | M = m) for the c that the popular key produces.
Answer:
The proof breaks at step two. P(C = c | M = m) is now the probability of the key m ⊕ c, which is 1/2 for one particular c and much less for the others. The quantity is no longer independent of c, and the cancellation at the end does not happen.
Concretely: for each m there is a ciphertext that is far more likely than the rest. So on seeing a ciphertext, Eve computes which m would have produced it under the popular key, and that m is now the odds-on favourite. Her posterior differs from her prior, which is exactly the failure of the definition.
What Eve does: guess that the popular key was used. She is right half the time, and each time she is right she reads the message completely.
This is Section 4.5's subject quantified: the security of a one-time pad is exactly the quality of its random number generator, and biases translate directly into an advantage for Eve rather than into a vague weakening.
Edge cases
Proposition 4.4 in the book gives a hard lower bound, and it is a pure counting argument with no cryptography in it.
Discussion prompt
Argue that a system with fewer keys than messages cannot have perfect secrecy, whatever its algorithm.
Hint: Fix a ciphertext c and count how many messages could have produced it.
Answer:
The argument. Fix a ciphertext c. Each key k decrypts c to exactly one message, so the set of messages reachable from c has at most (number of keys) elements.
If there are fewer keys than messages, some message m is not in that set — no key decrypts c to m. So P(M = m | C = c) = 0.
But P(M = m) > 0, since m was a possible message. So P(M = m | C = c) ≠ P(M = m), and perfect secrecy fails by the definition.
Why this bound is brutal: it is unconditional. No cleverness in the cipher can evade a counting argument, so any system with a short key — every practical system — is definitively without perfect secrecy, and must aim at computational security instead.
It also explains the pad's key length requirement. To have as many keys as messages of length n bits, the key needs n bits. The pad is not wasteful; it is exactly at the bound.
Figure (svg): Proposition 4.4 as a picture: perfect secrecy requires at least as many keys as messages.
Two truths and a lie
Two of these overstate the guarantee. One is exactly right.
Eliminate the wrong options
Which statement is correct?
Survives elimination: a
Why: Perfect secrecy is exactly the statement that the posterior equals the prior — the ciphertext contributes zero information. It does not make guessing hard, and it does not conceal length. Both of those exclusions have caused real failures, and knowing precisely what a guarantee covers is more useful than knowing that it is strong.
Comparison
This chapter is the hinge between two kinds of guarantee. Fill the blanks.
Comparison matrix
| Perfect secrecy | Computational security | |
|---|---|---|
| Adversary assumed to be | unlimited | bounded — running in feasible time |
| Guarantee holds | forever | until a better algorithm or faster machine appears |
| Key length required | at least the message length | fixed and short — 128 or 256 bits |
| Achieved by | the one-time pad | AES, RSA, everything deployed |
| Rests on | a counting theorem | an unproved hardness assumption |
The last row is the honest one. Everything practical in this book rests on a belief that some problem is hard, and none of those beliefs has ever been proved.
Explain it
A colleague reads that the one-time pad is unbreakable and proposes using it for the company's backups.
Discussion prompt
Explain, without dismissing the idea, why it does not help — and name the one situation where it would.
Hint: Work out how much key a nightly backup would need, and how it gets to the other end.
Answer:
The arithmetic first. A 2 TB nightly backup needs 2 TB of true random key, generated fresh each night, delivered to the restore site before it can be used, and destroyed afterwards. Delivering 2 TB securely is the problem the encryption was supposed to solve.
So the pad relocates the problem without shrinking it. Any channel able to carry the key securely could have carried the backup securely.
And key management gets harder, not easier: the key must be stored until the restore, never reused, and destroyed reliably — three operational requirements that AES with a 256-bit key does not have at all.
Where it does make sense: a low-volume, extremely high-value link where a courier is affordable and the traffic is small — the Moscow-Washington hotline, diplomatic cables, nuclear command authorisation. Small messages, enormous stakes, and a physical channel that already exists.
The reframing worth offering: the pad is not a stronger AES; it is a different trade. It converts a computational assumption into a logistics problem, and for almost everyone the logistics problem is the harder one.
Error analysis
A design document proposes a stream cipher and calls it a one-time pad.
Annotate
The scheme is fine; the claim is wrong. Rewrite it as 'computationally secure assuming the PRNG is secure' and everything in it becomes defensible.
Discrimination
Classify each scheme by the kind of guarantee it can honestly claim.
Sort into buckets
Sort each one.
Real world
Ciphertext indistinguishability is not a textbook nicety — it is what modern schemes are actually certified against.
Discussion prompt
Name three concrete design features of deployed cryptography that exist purely to win this game.
Hint: Think about what has to be different every time a message is encrypted.
Answer:
Randomised padding in RSA (OAEP). Textbook RSA loses the game in one step because it is deterministic. OAEP mixes random bits into the plaintext before exponentiation, so encrypting the same message twice gives different ciphertexts.
Initialisation vectors and nonces in block cipher modes. CBC's IV and CTR's nonce exist so that identical plaintexts produce different ciphertexts. ECB has neither, which is why the ECB-encrypted image still shows its picture — a failure of exactly this definition.
Ephemeral keys and forward secrecy in TLS. A fresh Diffie-Hellman exchange per session means recorded traffic cannot be linked or retrospectively decrypted after a long-term key compromise.
The common thread is that determinism is the enemy. Any scheme where the same input reliably gives the same output leaks equality, and equality is usually enough.
And that is Chapter 2's lesson restated formally: the substitution cipher hides the letters and not the pattern, and the pattern is most of the message.
Section
Section 4.5 · pp. 97-100
Concept
Perfect secrecy is stated with probabilities over all messages. There is an equivalent formulation that is easier to test and has become the standard way modern cryptography defines security: a game.
Guessing blindly wins half the time. The system has ciphertext indistinguishability if no strategy beats 1/2 by any significant margin.
Note what the definition does: it makes Alice, not Bob, choose the messages. She may pick the two most distinguishable messages she can think of, and the system must still defeat her. That is what makes the definition strong.
Figure (svg): The ciphertext indistinguishability game: Alice supplies two messages, Bob encrypts one at random, Alice guesses which.
Worked example
The book's example, and it takes one observation.
Alice submits m₀ = CAT and m₁ = DOG
Why: She is allowed to choose, so she picks a pair no single shift can confuse.
Bob returns the ciphertext PNG
Why: One of the two, encrypted under a shift Alice does not know.
Check DOG against PNG: D→P is +12, O→N is −1, G→G is 0
Why: Three different shifts, so no single key does this. PNG is not a shift of DOG.
Therefore it must be CAT
Why: Alice answers with certainty.
Verify: Alice wins with probability 1, not 1/2
Why: So the shift cipher has no indistinguishability whatever. And note she never found the key — the definition is violated without any key recovery, which is the point of defining security this way.
Figure (svg): The shift cipher losing the indistinguishability game: PNG cannot be a shift of DOG, so the answer is forced.
Prediction
RSA's encryption function is public: anyone can compute c = mᵉ mod n.
Predict first
Alice submits m₀ and m₁ and receives c. Can she win?
Correct: Yes — she encrypts both messages herself and compares with c
The general principle: any deterministic public-key encryption fails indistinguishability, because the adversary can always re-encrypt candidates and compare.
It is Chapter 2's ECB lesson at a different scale — determinism leaks equality — and it is why randomised padding is not an optional extra but part of the scheme.
Why: Textbook RSA is deterministic and its encryption key is public, so Alice simply computes m₀ᵉ mod n and m₁ᵉ mod n and sees which equals c. She wins with certainty, having done two public operations and broken nothing. This is why real RSA never encrypts a bare message: OAEP padding mixes in random bits so that the same plaintext gives a different ciphertext every time.
Figure (svg): The shift cipher losing the indistinguishability game: PNG cannot be a shift of DOG, so the answer is forced.
Worked example
Show the pad achieves exactly 1/2, using the previous section's result.
Bob picks b at random, so P(m₀) = P(m₁) = 1/2 from Alice's point of view
Why: This is her prior, before she sees anything.
Perfect secrecy gives P(M = m₀ | C = c) = P(M = m₀) = 1/2
Why: The ciphertext does not move the probability, by the theorem just proved.
Likewise P(M = m₁ | C = c) = 1/2
Why: The two possibilities remain exactly balanced.
So whatever c Alice receives, both candidates are equally likely and her best strategy is a coin flip
Why: There is no strategy at all, not merely no known strategy.
Verify: her success probability is exactly 1/2, the blind-guessing rate
Why: So the pad has ciphertext indistinguishability. Note the direction of the argument: perfect secrecy implies indistinguishability, which is why the game is a usable substitute for the probabilistic definition.
Figure (svg): The ciphertext indistinguishability game: Alice supplies two messages, Bob encrypts one at random, Alice guesses which.
Hypothesis
Perfect secrecy is a precise statement about probabilities. The CI game looks less rigorous.
Predict first
Why has modern cryptography adopted game-based definitions instead?
Correct: They can be relaxed to computational security by bounding the adversary's resources, which the probabilistic definition cannot
The game also composes. Chapter 12's random oracle model, Chapter 10's reduction to Computational Diffie-Hellman, and modern chosen-ciphertext definitions are all the same game with the adversary given more powers.
So this short section is where the book's definition of security shifts from information-theoretic to computational, and every later security claim is stated in the second style.
Why: Proposition 4.4 shows perfect secrecy is unattainable for any practical system — a counting argument no algorithm can dodge. The game survives the relaxation: keep the same setup, but require only that no adversary running in feasible time beats 1/2 by more than a negligible margin. That is the definition every modern scheme is proved against, and it has no counterpart in the unconditional formulation.
Constraint
Pads are unwieldy, so people use a short seed and a pseudorandom generator: a 100-bit seed expanding to a 1-million-bit key.
Discussion prompt
What exactly is lost, and what is the practical risk if the seed were only 20 bits?
Hint: Count the keys, then apply Proposition 4.4, then imagine enumerating.
Answer:
What is lost is the proof, immediately. A 100-bit seed gives 2¹⁰⁰ possible keys, while there are 2^1000000 possible messages. Proposition 4.4 says perfect secrecy is impossible — not unlikely, impossible — regardless of how good the generator is.
What is gained is practicality: 100 bits to distribute instead of a megabyte, which is the difference between a usable system and a courier network.
With a 20-bit seed the loss becomes concrete. An attacker generates all 2²⁰ ≈ a million keystreams, and given a ciphertext and a guessed plaintext, checks whether any seed connects them. A million trials is nothing.
So the seed length becomes the security parameter, and the system's guarantee changes kind: from 'no adversary, however powerful' to 'no adversary who cannot search 2¹⁰⁰ seeds'. That is a computational claim, and Chapter 5 is about building generators worthy of it.
The honest summary of the trade: the one-time pad's proof is real and its key distribution is impossible; every practical stream cipher keeps the shape and gives up the theorem.
Figure (svg): Proposition 4.4 as a picture: perfect secrecy requires at least as many keys as messages.
Intuition
Section 20.4 proves this chapter's result again using entropy, and the restatement is worth previewing because it makes the key-length requirement feel inevitable rather than unlucky.
Entropy H(X) — The average number of bits needed to describe an outcome of X — a measure of how much you do not know. H is zero for a certainty and maximal for a uniform distribution.
\[ \text{perfect secrecy} \;\Longleftrightarrow\; H(M \mid C) = H(M) \;\Longrightarrow\; H(K) \geq H(M) \]
The first equivalence says the ciphertext leaves your uncertainty about the message untouched — the same statement as P(M | C) = P(M), written with entropy instead of probability.
The second is Proposition 4.4 again: the key must carry at least as much uncertainty as the message. A 128-bit key has at most 128 bits of entropy, so it cannot conceal a megabyte. Counting keys and measuring entropy give the same bound because they are the same fact.
Figure (svg): Uncertainty about the message before and after seeing the ciphertext, unchanged, with the key's entropy shown as at least as large.
Commit first
Two systems protect the same 1 MB file for the next fifty years. System A uses a genuine one-time pad. System B uses AES-256.
Predict first
Which is more likely to still be unread in 2076?
Correct: System A — its guarantee is a theorem and cannot expire
Which points at the real answer: the honest bet depends on whether the key material was truly random, delivered securely, never reused and reliably destroyed. Venona lost on the third of those, not on the mathematics.
So the practical ranking is often the reverse of the theoretical one — a well-implemented computational scheme beats a badly operated unconditional one, every time.
Why: The pad's guarantee is unconditional: no advance in algorithms, hardware or computational model touches it, and Chapter 25's quantum algorithms are irrelevant to it. AES-256's guarantee is an assumption that has held for two decades, which is good evidence and not a proof. Over a fifty-year horizon the difference matters — but only if the pad's operational requirements were genuinely met, and that is where such systems actually fail.
Hypothesis
The pad is unusable, so Chapter 5 generates the keystream from a short seed.
Predict first
Which of the pad's three hypotheses does this necessarily break?
Correct: Both the first and the second, and the second is the one Proposition 4.4 punishes
So Chapter 5's whole subject is the remaining question: given that the proof is gone, how good can the generator be? Section 5.1 asks what pseudorandom should mean, 5.2 gives a construction that is fast and linear — and breaks — and 5.3 gives RC4.
Keeping the hypothesis-by-hypothesis accounting is the most useful thing to carry forward: every stream cipher failure in this book is one of these three lines.
Why: The keystream is as long as the message, but the KEY is the seed, and it is short — so the number of distinct keys collapses and Proposition 4.4 forbids perfect secrecy outright. The output is also pseudorandom rather than random, which breaks the first hypothesis. The third can be preserved, and it is the one careful designs do preserve, using a per-message nonce so no keystream ever recurs.
Blank canvas
Take the constraint seriously rather than dismissing the pad. Some organisations really do run one.
Draw it
Sketch a system for a diplomatic link carrying at most 10 kB per day between two fixed sites, using genuine one-time pads. Address: how the key material is generated and how you would test the generator; how it is delivered and how much a courier run carries; how both ends stay synchronised on which pad page is next; how used material is destroyed and audited; and what happens when a page is lost or a message is corrupted in transit. Then mark which of your answers would fail if the traffic rose to 10 GB per day.
The last mark is the point of the exercise. Nothing in the design is mathematically hard; the whole system is logistics, and it collapses on volume rather than on cryptanalysis.
Pattern
This chapter contains the book's first real proof of security, and the shape of it recurs in Chapters 10, 12 and 19.
The habit to take away: when a system claims a proof, ask what the hypotheses are and whether the deployment satisfies them. Venona and WEP both used mathematics that was correct and hypotheses that were false.
Figure (svg): The provable security thread through the book: perfect secrecy, then a reduction to Diffie-Hellman, then the random oracle model.
Trap
The trap. Security degrades gradually. A one-time pad used twice is roughly half as strong; used ten times it is a tenth as strong. So reuse is a matter of degree, and reusing a pad on two very short messages is a small risk worth taking under pressure.
This reasoning is exactly what produced Venona: wartime production pressure, a judgement that a little duplication was an acceptable trade.
Why it fails. Perfect secrecy is an implication with hypotheses, and one of them is that the key encrypts one message. Use it twice and the hypothesis is false, so the theorem says nothing at all — there is no residual half-guarantee to fall back on.
What is left is not a weakened one-time pad. It is a different cipher: Eve computes C₁ ⊕ C₂ = M₁ ⊕ M₂, a running-key cipher with no key in it, broken by the language statistics of Chapter 2.
And more reuse makes it easier, not proportionally harder. Three messages give three pairwise XORs that constrain each other, and the constraints are what make the attack reliable — this is why the book's example uses three messages rather than two.
The correct mental model: security proofs are cliffs, not slopes. Hypotheses hold and you have the guarantee, or they do not and you have nothing. Design so the hypotheses cannot be violated, rather than trusting that they will not be.
Check
Work it out before you click.
Check your understanding
A one-time pad ciphertext is 10000101 and Eve learns the plaintext was 00101001. What does she now know?
Answer: B
Why: XORing plaintext with ciphertext recovers the key: 00101001 ⊕ 10000101 = 10101100. But the key is used once and is random, so those bits are independent of everything else — knowing them helps with no other message. The pad is not unbreakable in the sense of hiding the key from a known plaintext; it is unbreakable in the sense that the recovered key is worthless.
Check
Apply Proposition 4.4 directly.
Check your understanding
A system encrypts 1000-bit messages using 128-bit keys. What can be said about its secrecy?
Answer: B
Why: There are 2¹²⁸ keys and 2¹⁰⁰⁰ messages. Fix a ciphertext: at most 2¹²⁸ messages decrypt to it, so the vast majority of messages are excluded, giving them posterior probability 0 against a positive prior. Perfect secrecy fails by counting alone — no property of the cipher can change it. This is why every practical system, AES included, targets computational security instead.
Check
Read the scheme carefully before answering.
Check your understanding
A scheme encrypts by XORing the message with SHA-256 of a fixed shared secret. Does it have ciphertext indistinguishability?
Answer: C
Why: The secret is fixed, so the hash output is fixed, so every message is XORed with the same keystream. Alice submits m₀ and m₁, receives c, and computes c ⊕ m₀: if the result is the same keystream she saw on any previous message she knows the answer — and even in isolation, two encryptions under this scheme XOR to M₁ ⊕ M₂. Determinism is the failure, exactly as with textbook RSA.
Connect it up
This chapter is one theorem and one impossibility result. Both are short, and writing them out is the fastest way to own them.
Draw it
On one page: (1) state the three hypotheses of the one-time pad and mark, beside each, the step of the proof that consumes it; (2) write the definition of perfect secrecy as a formula and say in plain words what it promises and what it does not; (3) give the counting argument for Proposition 4.4 in three lines; (4) describe the CI game and show why the shift cipher and textbook RSA both lose it; (5) finish with the trade a pseudorandom generator makes — what is gained, what is lost, and what the security parameter becomes.
Point five is what Chapter 5 spends its length on, so it is worth arriving there with the trade already clear.
Exit ticket
One question about the shape of the guarantee, not its content.
Predict first
Why is the one-time pad, despite being provably unbreakable, almost never used?
Correct: The key must be as long as the message, truly random, and never reused — so distributing keys is as hard as sending the message securely
Why: The pad does not solve the security problem; it relocates it. To send a gigabyte securely you must first have delivered a gigabyte of random key securely, so the key distribution problem is exactly as large as the problem you started with. XOR is in fact the fastest operation available, so speed is not the issue, and no cipher is stronger — nothing else in the book has an unconditional proof. The pad is used exactly where couriers are affordable: diplomatic links and nuclear command channels.
Recap
One cipher, one theorem, and a definition of security the rest of the book uses.
Chapter 5 next. Since a true pad is unusable, generate a keystream from a short seed. That is a stream cipher, and it keeps the pad's shape while giving up its proof — so the whole question becomes how good the generator is.
Figure (svg): The provable security thread through the book: perfect secrecy, then a reduction to Diffie-Hellman, then the random oracle model.
Want this taught 1-on-1? Alexander tutors Cryptography — $55/session, free consultation.