Chapter 5 of Trappe & Washington: pseudorandom bit generation from linear congruential generators, one-way functions and the Blum-Blum-Shub quadratic residue generator; linear feedback shift registers, their periods, and the known-plaintext attack that recovers a length-m recurrence from 2m bits by linear algebra over GF(2); and RC4's key scheduling and generation algorithms together with the biases that ended its deployment.
Subject: Cryptography · 60 slides · diagram-first lesson
Open the interactive version of this deck
Title
Cryptography · Chapter 5
Keystreams from short seeds: linear congruential, shift registers, and RC4 — and how the first two fall
Objectives
Chapter 4 ended with a trade: give up the one-time pad's proof to get a key short enough to distribute. This chapter is that trade carried out three times, and two of the three results are broken.
Figure (svg): Four keystream generators compared on speed against how well their unpredictability is understood.
Warm-up
Chapter 4 proved the one-time pad secure under three hypotheses, and a stream cipher breaks the second of them by design.
Discussion prompt
Given that perfect secrecy is now definitively unavailable, what property should we be asking a keystream generator for instead?
Hint: Section 4.5 gave a definition that survives being weakened.
Answer:
Unpredictability. Not 'looks random', not 'passes statistical tests' — the requirement is that an adversary who has seen some output bits cannot predict the next one significantly better than by guessing.
This is the indistinguishability game of Section 4.5 applied to the keystream: no efficient strategy should beat 1/2 by a noticeable margin at predicting the next bit.
It is a strictly weaker demand than perfect secrecy — Proposition 4.4 already killed that — but it is achievable, testable, and it is the right bar. The three generators in this chapter can be sorted by how convincingly they clear it, and two of them do not clear it at all.
Section
Section 5.1 · pp. 105-107
Concept
Every cryptographic algorithm in this book needs random bits — as a key for DES or AES, as a one-time pad, as padding, as a nonce.
Natural randomness works in principle. Thermal noise in a semiconductor resistor is a good source. But it is impractical at scale for two reasons the book names: sampling a physical process is inherently slow, and it is hard to guarantee an adversary is not observing the same process.
So we want randomness generated in software, from a short seed. That is a pseudorandom generator, and the whole question is what 'random enough' means.
Seed — The short secret input from which the generator produces its long output. Everything the generator produces is a deterministic function of the seed, so the seed's entropy is a hard ceiling on the output's.
Figure (svg): A one-way function generator: a seed plus a counter feeds a one-way function, and the least significant bit of each output is the keystream.
Concept
The generator behind the standard C library's rand(), and behind most language runtimes' default random numbers.
\[ x_n \equiv a \, x_{n-1} + b \pmod m \]
The seed is x₀; a, b and m are parameters. It is fast, it has a long period if the parameters are chosen well, and its output passes many statistical tests.
The book's verdict is unambiguous: suitable for experimental purposes, highly discouraged for cryptographic purposes. The reason is that it is predictable — an eavesdropper can use knowledge of some outputs to predict future ones with high probability, even without knowing a, b and m.
And the result generalises: it has been shown that any polynomial congruential generator is cryptographically insecure. The failure is structural, not a matter of bad parameters.
Figure (svg): A linear congruential generator's outputs plotted in pairs, falling on a small number of parallel lines rather than filling the square.
Worked example
Suppose Eve sees the outputs 7, 38, 61 from a generator with unknown a, b and modulus m = 64.
Write the two equations: 38 ≡ 7a + b and 61 ≡ 38a + b (mod 64)
Why: Two unknowns, two equations. This is the affine known-plaintext attack from Chapter 2, one modulus larger.
Subtract: 23 ≡ 31a (mod 64)
Why: The constant b drops out — exactly as in Section 2.2.
31 is odd, so gcd(31, 64) = 1 and 31 is invertible mod 64. Its inverse is 31, since 31 · 31 = 961 = 15 · 64 + 1
Why: The extended Euclidean algorithm of Section 3.2 again.
a ≡ 31 · 23 = 713 = 11 · 64 + 9, so a = 9. Then b ≡ 38 − 7 · 9 = −25 ≡ 39 (mod 64)
Why: Two outputs determined the whole generator.
Verify: check the third output: 9 · 38 + 39 = 381 = 5 · 64 + 61 ≡ 61 ✓
Why: The parameters reproduce the observed sequence, so Eve can now generate every future bit. Three observed values, two equations, one gcd — and the generator is finished.
Figure (svg): Two observed outputs giving two linear congruences whose solution is the generator's parameters.
Notation
Four symbols, and the reason the generator is forbidden is visible in all four.
Annotate
On: \( x_n \equiv a \, x_{n-1} + b \pmod m \)
Compare with the RC4 round two sections on: there, the index j depends on the array's contents and the output is an entry selected by other entries — no equation to write down.
Counterexample
A program generates a key by seeding a generator with the current time in seconds.
Discussion prompt
How large is the effective key space, and does using a strong generator rescue it?
Hint: Count the plausible seeds rather than the possible outputs.
Answer:
Count the seeds. If the attacker knows the key was generated within a given day, there are 86 400 possible seeds. Within a month, about 2.6 million. That is the entire key space, whatever the generator does with it.
A strong generator does not help at all. The output is a deterministic function of the seed, so the number of distinct possible keys equals the number of distinct possible seeds. SHA-512 applied to a million seeds yields a million keys, not 2⁵¹².
This is Proposition 4.4 again, and it is the most useful thing that proposition does: the count of keys is bounded by the entropy of the seed, and no downstream processing can manufacture entropy that was never there.
It has happened. The Debian OpenSSL bug of 2008 removed entropy from the seeding of the key generator, leaving about 32 768 possible keys per architecture. Every SSH and TLS key generated on affected systems for two years was enumerable, and the generator itself was never at fault.
The rule: audit the seed, not the algorithm. The algorithm is the part that is usually right.
Two truths and a lie
Two of these confuse a property of the output with a property of the process.
Eliminate the wrong options
Which statement states the requirement correctly?
Survives elimination: a
Why: Unpredictability is a statement about what an adversary can compute, not about how the output looks. That is why it is stated as the game of Section 4.5 — quantified over strategies rather than over surface features. Option (c) is worth dwelling on, because the instinct behind it is exactly backwards: a generator that avoids coincidences is more distinguishable, not less.
Concept
The fix is to put something non-invertible between the state and the output. Take a one-way function f and a random seed s.
\[ x_j = f(s + j), \qquad b_j = \text{least significant bit of } x_j \]
The sequence b₁, b₂, … is then pseudorandom, and its unpredictability is inherited: predicting the next bit means learning something about f's input from its output, which is what one-way means it cannot do.
The book names two popular choices for f: DES (Chapter 7) and SHA, the Secure Hash Algorithm (Chapter 11). The cryptographic generator in OpenSSL, which secures a large fraction of internet traffic, is built on SHA in exactly this way.
This is the practical answer, and it is what almost every real system uses. Its security is conditional — on f being one-way — which is the trade Chapter 4 said we would be making.
Figure (svg): A one-way function generator: a seed plus a counter feeds a one-way function, and the least significant bit of each output is the keystream.
Analogy
Real systems layer several of these. Confusing their roles is a common design error.
Match the pairs
Why: The layered answer real systems use: collect genuine but slow entropy into a pool, then stretch it with a one-way function into as many bits as needed. Each layer does the job the other cannot — physics supplies unpredictability, the hash supplies throughput. The failure mode is treating the third row as if it belonged in the fourth, which is the Debian bug and a hundred smaller ones.
Missing information
A specification reads: "Session keys are generated using SHA-256."
Discussion prompt
List what a reviewer still cannot determine, and rank the gaps by how badly each could fail.
Hint: SHA-256 is a function. A generator is a function plus an input plus a state discipline.
Answer:
Where the seed comes from, and how much entropy it has. This is the gap that matters most: the whole key space is bounded by it, and it is the one that has caused real catastrophes.
Whether the counter or state can repeat. If the generator restarts after a reboot with the same seed, it re-emits the same keystream — the two-time pad of Section 4.3, arriving through an operational path rather than a cryptographic one. Virtual machine snapshots have caused exactly this.
Whether the state is protected. If an attacker reads the internal state, can they compute past outputs as well as future ones? Backtracking resistance is a design property and it is usually not free.
How often it is reseeded. A generator that never mixes in fresh entropy stays compromised forever after a single state leak.
Ranking by damage: seed entropy first, state repetition second, backtracking third, reseeding fourth. Naming the hash function — the one thing the spec does say — is close to the least important decision in the list.
Concept
Also called the quadratic residue generator, and the strongest of the three in argument. Choose two large primes p and q both congruent to 3 mod 4, set n = pq, and pick x coprime to n.
\[ x_0 \equiv x^2 \pmod n, \qquad x_j \equiv x_{j-1}^2 \pmod n, \qquad b_j = \text{lsb of } x_j \]
Each step squares the previous value mod n, and the output bit is the parity — whether the number is odd or even, which is free to check.
Predicting the next bit can be shown to be as hard as distinguishing quadratic residues mod n, which is as hard as factoring n. So the generator is very likely unpredictable, on the same assumption RSA rests on.
Its problem is speed: one modular squaring of a 2048-bit number per output bit. Extracting the k least significant bits per step helps, and remains secure as long as k ≤ log₂log₂ n — but even so it is orders of magnitude slower than a hash-based generator.
Figure (svg): The Blum-Blum-Shub iteration: repeated squaring mod n, with only the parity of each value emitted.
Estimation
One modular squaring of a 2048-bit number per output bit, at very roughly 10 microseconds per squaring.
Predict first
Roughly how long to generate a 256-bit AES key?
Correct: A few milliseconds
The book's own remedy is to extract the k least significant bits per squaring rather than one, which stays secure while k ≤ log₂log₂ n — about 11 bits for a 2048-bit modulus, so an order-of-magnitude speed-up.
Note the shape of the answer: this is Chapter 1's hybrid pattern again. Use the expensive, well-justified primitive on something short, and a fast one on the bulk.
Why: 256 bits at about 10 microseconds each is roughly 2.6 milliseconds — fine for a single key, and that is the honest verdict: BBS is perfectly usable for generating occasional key material. Where it fails is bulk keystream: encrypting a one-gigabyte file needs 8 × 10⁹ squarings, about a day. So the slowness is not a reason to reject the design, it is a reason to use it for the small job and something else for the large one.
Real world
Chapter 4 proved that the pad's security is exactly the quality of its randomness. That is not only a statement about pads.
Discussion prompt
Name two deployed failures where every primitive was sound and the random number generator was the entire flaw.
Hint: One is a Linux distribution's patch; one is a games console.
Answer:
Debian OpenSSL, 2006-2008. A patch silencing a memory-analysis warning removed most of the entropy from the seeding path, leaving roughly 32 768 possible keys per architecture. SSH host keys, TLS certificates and DSA signing keys generated over two years were all enumerable. RSA and SHA were untouched and irrelevant.
Sony PlayStation 3, 2010. ECDSA signatures require a fresh random nonce per signature; the implementation used a constant. Two signatures with the same nonce let anyone solve directly for the private signing key, which was then published. The curve, the hash and the signature scheme were all fine.
The common shape: the randomness requirement was a hypothesis of a security proof, stated in the specification, and violated by an implementation decision that looked unrelated to security.
Which is why this chapter comes where it does. Randomness is not a preliminary to cryptography; it is a component with its own failure modes, and it is the component most often got wrong.
Comparison
Fill the blanks. The pattern in the last column is the chapter's argument.
Comparison matrix
| Generator | Speed | Security basis |
|---|---|---|
| Linear congruential | very fast | none — provably predictable |
| One-way function (SHA, DES) | fast | f is one-way — an assumption |
| Blum-Blum-Shub | very slow — a modular squaring per bit | factoring is hard — a well-studied assumption |
| LFSR | fastest of all | none — linear algebra recovers it |
| RC4 | fast | no proof, and known biases |
Nothing in the table is both fast and provably unpredictable, and that is not an accident of the state of the art — it is the shape of the field.
Definition probe
Some generators fail because of bad parameters. Others fail whatever parameters you choose.
Sort into buckets
Sort each weakness.
Section
Section 5.2 · pp. 107-113
Concept
The book is explicit about the setting: this is a method for when speed matters more than security — cable television, where there is a lot of data and rarely an economic case for an expensive attack. Its real use is as one building block inside more complex systems.
Fix a length m and coefficients c₀ … c_{m−1}, all mod 2, and a recurrence:
\[ x_{n+m} \equiv c_0 x_n + c_1 x_{n+1} + \cdots + c_{m-1} x_{n+m-1} \pmod 2 \]
Specify the initial values x₁ … x_m and every later bit follows. In hardware this is m flip-flops in a row: each tick shifts everything right, and the tapped positions XOR into the left end. A few dozen transistors, one bit per clock cycle.
Figure (svg): A three-stage linear feedback shift register implementing the recurrence x n plus 3 equals x n plus 1 plus x n.
Worked example
Take the recurrence x_{n+5} ≡ x_n + x_{n+2} with seed x₁…x₅ = 0, 1, 0, 0, 0.
x₆ = x₁ + x₃ = 0 + 0 = 0
Why: Only two taps, so each new bit costs one XOR.
x₇ = x₂ + x₄ = 1 + 0 = 1, and x₈ = x₃ + x₅ = 0 + 0 = 0
Why: The window slides one place each time.
x₉ = x₄ + x₆ = 0, x₁₀ = x₅ + x₇ = 0 + 1 = 1, x₁₁ = x₆ + x₈ = 0
Why: Building up 01000010010…
\[ 01000010010110011111000110111010100001001011001111 \]
Verify: the sequence repeats after 31 terms, and 2⁵ − 1 = 31
Why: A length-5 register has 32 possible states, and the all-zero state is a dead end, so 31 is the maximum possible period — this recurrence achieves it. Ten bits of information (5 seed, 5 coefficients) produced 31 bits of keystream.
Figure (svg): The first sixteen bits of the period-31 sequence, with the five seed bits marked and the recurrence producing the rest.
Invariant
The recurrence and the hardware are the same object. Step the register x₄₊ₙ ≡ xₙ + xₙ₊₁ and watch the state move.
Step through it
Why is the all-zero state fatal rather than merely unlucky?
Because the feedback is a sum of cells: all zeros in gives zero out, forever. Every LFSR must be seeded non-zero, and that is why the period is 2^m − 1 rather than 2^m.
Translation
The same LFSR gets written three ways in the literature. Being able to move between them saves a lot of confusion.
Match the pairs
Why: The coefficients, the taps and the polynomial x⁴ + x + 1 over GF(2) are three notations for one object, and the period is a property of that polynomial — maximal exactly when it is primitive, which is the finite-field material of Section 3.11 in use. The attacker's linear system solves for the coefficient vector, so all four rows describe the thing that 2m known bits reveal.
Commit first
Two stream ciphers are offered. Cipher A is an LFSR of length 128 with a maximal period. Cipher B is RC4 with a 128-bit key.
Predict first
Which needs more known plaintext to break?
Correct: B — there is no known attack recovering its key from known plaintext at that size
The Berlekamp-Massey algorithm does the job in O(n²) and finds the shortest LFSR generating any sequence, so the attack is not merely possible but routine.
This is why the linear complexity of a keystream — the length of the shortest LFSR producing it — is a standard measure, and why it is checked before a period ever is.
Why: Cipher A falls to 256 known bits — 32 bytes, less than one packet header — because its output is a linear function of its state, and the period plays no part in the attack. RC4 has biases and distinguishers but no known key-recovery attack from known plaintext at 128 bits. The astronomically larger period buys cipher A nothing at all, which is the whole point of this section.
Worked example
Exactly the one-time pad's arithmetic, with the pad replaced by LFSR output.
Plaintext 1011001110001111, keystream 0100001001011001
Why: The keystream is the first sixteen bits of the sequence just generated.
XOR column by column: 1⊕0 = 1, 0⊕1 = 1, 1⊕0 = 1, 1⊕0 = 1, …
Why: No carries, no interaction between columns.
\[ 1011001110001111 \oplus 0100001001011001 = 1111000111010110 \]
Decryption adds the same keystream to the ciphertext
Why: Which works because XOR is its own inverse — the receiver needs only the seed and the coefficients.
Verify: 1111000111010110 ⊕ 0100001001011001 = 1011001110001111
Why: The plaintext returns. The advantage over Vigenère is period: a short key gave Vigenère a period of 5 or 6, and the coincidence count found it. Here 62 bits of key give a period over two billion.
Figure (svg): A sixteen-bit plaintext XORed with sixteen bits of LFSR keystream to give the ciphertext.
Estimation
The book states that the recurrence x_{n+31} ≡ x_n + x_{n+3}, with any non-zero seed, has a period of 2³¹ − 1.
Predict first
So how many bits of key material do 31 seed bits plus 31 coefficients produce?
Correct: About 2.1 billion
A maximal-length LFSR of length m has period 2^m − 1, achieved when the associated polynomial is primitive over GF(2) — which is the finite-field material of Section 3.11 doing real work.
Hold the number, and then watch the next four slides take the cipher apart using twelve known bits.
Why: 2³¹ − 1 = 2 147 483 647, so 62 bits of specification produce more than two billion bits before repeating. Against Chapter 2's Vigenère cipher, where a six-letter key gave a period of six and coincidence counting found it in an afternoon, this is an enormous improvement — and it is exactly the improvement that turns out not to be the thing that matters.
Concept
The cipher succumbs easily to a known-plaintext attack, and 'easily' is not an exaggeration.
First, the plaintext and ciphertext drop out. XOR them together and you have the keystream directly, so the attack reduces to: given a segment of the sequence, find the recurrence.
And finding the recurrence is a linear system. Each consecutive window of m+1 known bits gives one equation in the unknowns c₀ … c_{m−1}. Take m of them and solve over GF(2).
\[ \begin{pmatrix} x_1 & x_2 & \cdots & x_m \\ x_2 & x_3 & \cdots & x_{m+1} \\ \vdots & & & \vdots \\ x_m & x_{m+1} & \cdots & x_{2m-1} \end{pmatrix} \begin{pmatrix} c_0 \\ c_1 \\ \vdots \\ c_{m-1} \end{pmatrix} = \begin{pmatrix} x_{m+1} \\ x_{m+2} \\ \vdots \\ x_{2m} \end{pmatrix} \]
So 2m known bits determine an LFSR of length m completely. The two-billion-bit period is irrelevant: the attacker never waits for it.
Figure (svg): Three attempted recurrence lengths as matrix equations over GF(2): length two gives a wrong answer, length three has no solution, length four works.
Worked example
The book's example. Known segment 011010111100, from a sequence of period 15. The length m is unknown, so try lengths in turn.
Length 2: from x₃ = c₀x₁ + c₁x₂ and x₄ = c₀x₂ + c₁x₃, get 1 ≡ c₁ and 0 ≡ c₀ + c₁, so c₀ = c₁ = 1
Why: A solution exists, so length 2 looks plausible — until it is tested.
Test it: the recurrence predicts x₆ = x₄ + x₅ = 0 + 1 = 1, but the actual x₆ is 0. Reject
Why: Solving is not enough; the candidate must regenerate the bits already known.
Length 3: the matrix has rows (0,1,1), (1,1,0), (1,0,1) and right-hand side (0,1,0)
Why: Every column of the matrix sums to 0 mod 2, while the right-hand side sums to 1. So no solution exists at all — the determinant is 0 and the system is inconsistent.
Length 4: rows (0,1,1,0), (1,1,0,1), (1,0,1,0), (0,1,0,1), right-hand side (1,0,1,1). This solves to c = (1, 1, 0, 0)
Why: Giving the recurrence x_{n+4} ≡ x_n + x_{n+1}.
Verify: regenerate from x₁…x₄ = 0,1,1,0: x₅ = 0+1 = 1, x₆ = 1+1 = 0, x₇ = 1+0 = 1, x₈ = 0+1 = 1, x₉ = 1, x₁₀ = 1, x₁₁ = 0, x₁₂ = 0
Why: That is 011010111100 — every known bit reproduced. So the recurrence is right, and every future bit of keystream is now computable. Twelve bits of known plaintext, and the cipher is finished.
Figure (svg): Three attempted recurrence lengths as matrix equations over GF(2): length two gives a wrong answer, length three has no solution, length four works.
Cost model
Put a number on it so it can be compared with the period.
Annotate
On: \( \text{data} = 2m \text{ bits}, \qquad \text{work} = O(m^3) \text{ by elimination, or } O(m^2) \text{ by Berlekamp-Massey} \)
Whenever a specification advertises a parameter, find the attack's cost formula and check whether that parameter appears in it.
Socratic
Trying length 5 on the same sequence gives a matrix whose determinant is 0 mod 2, and so does length 6.
Discussion prompt
Explain why over-guessing the length must produce a singular matrix — and how that turns into a length-finding algorithm.
Hint: Look at what the true recurrence says about the rows of the matrix.
Answer:
Because the rows become linearly dependent. With the true recurrence x_{n+4} ≡ x_n + x_{n+1}, the fifth row of a 5 × 5 matrix is the sum of the first and second rows — entry by entry, since x₅ ≡ x₁ + x₂, x₆ ≡ x₂ + x₃, and so on all the way across.
A matrix with one row a linear combination of others is singular, exactly as over the reals. So once the assumed length exceeds the true one, the determinant is 0 for every larger size.
This gives the algorithm. Compute the determinant for increasing lengths. It is non-zero up to the true length and zero thereafter, so the last non-singular size is m. Then solve that system for the coefficients.
It also explains the length-3 case: there the determinant was 0 and the system was inconsistent, which correctly says 'no length-3 recurrence fits'. Zero determinant means either no solution or many, and testing the candidate against the known bits settles which.
The efficient version of all this is the Berlekamp-Massey algorithm, which finds the shortest LFSR generating a sequence in O(n²) — and it is why LFSRs are never used alone.
Discrimination
Linear complexity is the length of the shortest LFSR generating a sequence. Sort by whether the sequence has low linear complexity.
Sort into buckets
Sort each sequence.
Trap
The trap. The LFSR x_{n+31} ≡ x_n + x_{n+3} has a period of 2 147 483 647. Vigenère fell because its period was 5 and Kasiski found it by counting coincidences. A period of two billion cannot be found that way, so the keystream is secure.
The reasoning has real pedigree: it is Chapter 2's lesson applied faithfully, and it gets exactly the wrong answer.
Why it fails. The attacker does not look for the period. She solves for the recurrence, and that needs only 2m = 62 known bits — about eight bytes of known plaintext, which a file header supplies for free.
Period measures how long until the keystream repeats. It says nothing about how much of the keystream is determined by a short prefix, and for a linear recurrence every bit is determined by 2m of them.
This is Chapter 1's argument in a new costume. There, a large key space did not imply security because Eve used structure instead of search. Here, a large period does not imply security because Eve uses linear algebra instead of waiting.
The general form worth keeping: a parameter being large only helps against attacks that scale with that parameter. Always ask which attack the number bounds, and whether it is the best one.
The practical consequence is that LFSRs are never used alone. They appear inside constructions that break the linearity — combined nonlinearly, irregularly clocked, or filtered — which is exactly how A5/1 in GSM and E0 in Bluetooth are built.
Fill the middle
Complete the rule that makes the attack quantitative.
Fill in the blanks
\text2m m \textat most 2^m − 1 ___ \text___ ___
Why: Each equation needs a window of m+1 consecutive bits, and m equations are required, so the windows span 2m bits. The period can be as large as 2^m − 1 when the recurrence is primitive. Setting these side by side is the whole point: the attack cost is linear in m while the period is exponential in it, so making the register longer does almost nothing against an attacker and a great deal against a coincidence count.
Explain it
A colleague who does not know matrix methods asks why a two-billion-bit period does not protect the cipher.
Discussion prompt
Give the explanation in plain terms, and say what the right question to ask about any keystream is.
Hint: The distinction is between how long until it repeats and how much of it is determined by the start.
Answer:
The period answers the wrong question. It says how long before the keystream comes round again. The attacker is not waiting for that; she is asking how much of the keystream is fixed once you know a little of it.
For a shift register, the answer is: all of it. Every bit is a fixed sum of a few earlier bits, so once you know which earlier bits, everything follows. And she can work out which by seeing about twice the register length — a few dozen bits — because the rule is a system of simple equations she can solve.
The analogy that lands: an arithmetic sequence 3, 7, 11, 15, … runs forever without repeating, and two terms tell you every future term. Length of run and predictability are unrelated.
The right question about any keystream: how many bits does an attacker need before she can compute the next one? For a shift register of length m the answer is 2m. For a good generator there is no known finite answer at all.
Faded example
Fill in the four steps that turn known plaintext into the whole keystream.
Fill in the blanks
First XOR the known plaintext with the ciphertext to recover a segment of the keystream. Then guess a register length m and build an m × m matrix from sliding windows of the known bits. Solve over GF(2) for the coefficients. Finally test the recurrence against the known bits — if it fails, increase m and repeat.
Why: The first step is what makes the attack a keystream problem rather than a cipher problem — the plaintext and ciphertext are never needed again. The last step is the one people omit, and it matters: at length 2 the system solved and gave a wrong answer, so solvability alone proves nothing and the candidate must regenerate every bit already known.
Real world
The book positions LFSRs as a building block. They were shipped in hundreds of millions of devices.
Discussion prompt
Name two deployed systems built from LFSRs and say what each did to try to defeat the linear-algebra attack.
Hint: One secures GSM voice calls; one is in every Bluetooth radio.
Answer:
A5/1, the GSM voice encryption cipher. Three LFSRs of lengths 19, 22 and 23, irregularly clocked: at each step a majority vote among three bits decides which registers advance. The irregular clocking is what breaks linearity, because the recurrence no longer holds at fixed offsets.
E0, Bluetooth's stream cipher. Four LFSRs whose outputs are combined through a small nonlinear finite-state machine with memory, so the output is not a linear function of the register contents.
Both have been broken, though not by the plain attack in this section. A5/1 fell to time-memory trade-offs and to the short 64-bit key; E0 to correlation attacks that exploit the residual linearity leaking through the combiner.
The pattern is worth naming: adding nonlinearity to a linear core buys difficulty, not security. Correlation attacks work by finding the linear component that survives, and there usually is one.
Which is why modern designs — ChaCha20, AES-CTR — do not start from a linear core at all.
Ranking
All four are stream ciphers of some kind. Rank by the amount of known plaintext an attacker needs.
Put in order
Why: Vigenère needs about 6 known letters — one per key position — and then the key is fully recovered. The length-31 LFSR needs 62 bits, roughly 8 bytes. RC4 has no known attack that recovers the key from known plaintext at 128 bits, so the requirement is effectively unbounded for that purpose. And a one-time pad is never broken by known plaintext at all: the recovered key bits are independent of every other bit. Note the first two are both tiny, and that the LFSR's two-billion period bought it a factor of ten in this ranking.
Elimination
You must keep the speed of a shift register but want to resist the known-plaintext attack.
Eliminate the wrong options
Which change addresses the real weakness?
Survives elimination: c
Why: The weakness is linearity itself, so the only real fix is to destroy it. Nonlinear combining and irregular clocking are exactly what A5/1 and E0 do, and they genuinely raise the cost — the plain matrix attack no longer applies. Note that (a) and (b) both make a number bigger without touching the attack, which is the same mistake as the trap slide, and (d) misreads what the attack assumes.
Section
Section 5.3 · pp. 113-114
Concept
Developed by Rivest, and for two decades the most widely deployed stream cipher in the world — SSL/TLS, WEP, WPA-TKIP — because it is fast and remarkably simple.
The algorithm was originally a trade secret. It was leaked to the internet in 1994 and has been extensively analysed ever since, which is a small case study in Kerckhoffs's principle: secrecy delayed the analysis and did not prevent it.
The key is a binary string between 40 and 256 bits. Everything else is a permutation of the 256 byte values, held in an array S, and two loops that shuffle it.
Figure (svg): The RC4 key scheduling algorithm swapping entries of the 256-byte array S during its first five steps.
Worked example
The book's trace, with key 10100000… of length 40. S starts as S[i] = i, and j starts at 0.
i = 0: j := (j + S[0] + key[0]) mod 256 = 0 + 0 + 1 = 1, then swap S[0] and S[1]
Why: Now S[0] = 1 and S[1] = 0. The key's first bit chose the swap.
i = 1: j := (1 + S[1] + key[1]) = 1 + 0 + 0 = 1, so S[1] swaps with itself
Why: A no-op step. The array is unchanged here.
i = 2: j := (1 + S[2] + key[2]) = 1 + 2 + 1 = 4, so swap S[2] and S[4]
Why: Giving S[2] = 4 and S[4] = 2.
i = 3: j := (4 + S[3] + key[3]) = 4 + 3 + 0 = 7, so swap S[3] and S[7]
Why: Giving S[3] = 7 and S[7] = 3.
i = 4: j := (7 + S[4] + key[4]) = 7 + 2 + 0 = 9 — note S[4] is now 2, not 4
Why: So swap S[4] and S[9], giving S[4] = 9 and S[9] = 2. The array's own evolving contents feed back into j, which is where the nonlinearity comes from.
Verify: continue to i = 255; S remains a permutation throughout, because every step is a swap
Why: A swap can never duplicate or lose a value, so S is a permutation of 0…255 at every moment. That invariant is what makes the generator's output well-defined, and it is worth noticing because it is also the source of the biases.
Figure (svg): The RC4 key scheduling algorithm swapping entries of the 256-byte array S during its first five steps.
Sorting
RC4 is two algorithms, and its strengths and weaknesses divide between them fairly cleanly.
Sort into buckets
Sort each property by where it originates.
Anomaly
RC4-drop[n] discards the first n keystream bytes — often 768 or 3072 — before encrypting anything.
Predict first
What does discarding output actually fix?
Correct: The strongest biases are concentrated in the early bytes, so discarding them removes the easiest distinguishers
The tell that it is a mitigation: it removes the biases that have been found, and says nothing about biases not yet found. Later work did find distinguishers extending far beyond the dropped prefix.
Compare with the LFSR situation. There the flaw was structural and no amount of running forward helps; here the flaw is also structural, and dropping bytes reduces its visibility rather than its existence.
Why: The array S carries visible traces of the key scheduling for the first few hundred rounds, and the known biases — the second byte, the first two bytes jointly, the S[0] distribution — all live there. Running the generator forward mixes S further, so the measurable structure fades. It is a mitigation, not a repair: the book says plainly that even RC4-drop is not recommended where high security is required.
Figure (svg): The observed probability that RC4's second output byte is zero, at twice the value a random keystream would give.
Explain it to yourself
The linear-algebra attack demolished a shift register with a two-billion period. RC4 has 256 bytes of state and no such attack is known.
Discussion prompt
Explain precisely what in RC4's design prevents the matrix method, and why that is not the same as being secure.
Hint: Write down what an equation for RC4's next output byte would have to look like.
Answer:
There is no linear equation to write. In an LFSR the next bit is a fixed sum of earlier bits, so the unknowns are coefficients and the relation is linear. In RC4, j is updated by adding S[i] — a value from the array, used as data — and the output is S[S[i] + S[j]], an array entry selected by other array entries. Indices depend on contents, so the relation is not linear in anything.
Concretely: the swap makes the state a permutation that changes as it is read, and array indexing by data-dependent indices is not expressible as a matrix over GF(2).
But immunity to one attack is not security. RC4 has no proof of anything. What it has is thirty years of analysis, and that analysis found the Mantin-Shamir biases, the Fluhrer-Mantin-Shamir key-recovery attack on related keys, and eventually plaintext-recovery attacks good enough to remove it from TLS in 2015.
The general lesson: defeating a specific attack is a design achievement and not a security argument. The honest claim for RC4 was always 'no better attack is known', and eventually a better attack was known.
Trade off
Chapter 6 is about the other shape. Fill the blanks — this is the decision an engineer actually makes.
Comparison matrix
| Property | Stream cipher | Block cipher |
|---|---|---|
| Output length | exactly the input length | rounded up to a whole block, so padding is needed |
| Effect of one flipped ciphertext bit | flips exactly that plaintext bit | destroys a whole block, and often the next one |
| Can an attacker flip a chosen plaintext bit? | yes, trivially — which is why integrity protection is mandatory | not without wrecking the block |
| Needs a mode of operation | no | yes — Chapter 6 |
| Consequence of reusing key material | total — the keystream cancels, as in Section 4.3 | leaks equality of blocks in ECB; other modes vary |
Row two cuts both ways and is worth remembering: a stream cipher's bit-exact error behaviour is a virtue on a noisy channel and a serious hazard without a MAC, because an attacker can flip any plaintext bit she likes.
Break the constraint
A designer proposes: keep RC4, but XOR each output byte with the byte before it, to smooth out the biases.
Discussion prompt
Does this fix the problem? Argue it through rather than guessing.
Hint: Ask what an attacker who knows the transformation can compute.
Answer:
It does not. The transformation is public under Kerckhoffs's principle, and it is invertible: given the modified stream, the attacker recovers the original by a running XOR. Any distinguisher on the original applies to the modified stream after one cheap step.
Worse, it can create new structure. XORing adjacent bytes of a biased stream produces its own biases, and there is no reason to expect them to be smaller — differencing a sequence with a known distribution gives a sequence with a computable distribution, not a flat one.
The general principle: an invertible public post-processing step cannot remove a distinguisher. If the output was distinguishable before, it is distinguishable after, because the distinguisher composes with the inverse of the transformation.
What would actually help: a non-invertible step — hashing the output, which costs more than using a hash-based generator in the first place — or a different cipher. Which is what happened: TLS moved to AES-GCM and ChaCha20-Poly1305.
Concept
With S shuffled, the output loop is four lines and produces one byte per round.
i := 0
j := 0
while generating output:
i := (i + 1) mod 256
j := (j + S[i]) mod 256
swap(S[i], S[j])
output S[(S[i] + S[j]) mod 256]Each output byte is XORed with the corresponding plaintext byte, exactly as in every other cipher in this chapter. The array keeps changing as it goes, so the generator has 256 bytes of state and an enormous period.
Notice what makes RC4 different from an LFSR: j depends on the contents of S, and the output is an entry of S selected by other entries of S. Nothing here is a linear function of the state, so the matrix attack has nothing to attack.
Figure (svg): One round of the RC4 generator: advance i, derive j from S, swap, then output the entry indexed by the sum.
Prediction
Mantin and Shamir showed that RC4's second output byte is 0 with probability about 2/256, where a random keystream would give 1/256.
Predict first
Why does a bias this small matter?
Correct: It loses the indistinguishability game, and it lets a repeated plaintext be recovered from many sessions
Mantin and Shamir also found P(first two bytes both 0) = 3/256² rather than 1/256², and there are biases in the KSA's output too — S[0] = 1 about 37% more often than it should be, and S[0] = 255 about 26% less often.
Hence RC4-drop[n], which discards the first n bytes before using the keystream. The book notes plainly that even this version is not recommended where high security is required.
Why: Section 4.5 requires that no strategy distinguish the keystream from random by a noticeable margin — and 'the second byte is 0 twice as often' is precisely such a strategy, so RC4 fails the definition outright. It is also practically exploitable: if the same plaintext byte is encrypted in many sessions, as an HTTP cookie is, the bias accumulates across sessions and the byte can be recovered. That attack is what finally removed RC4 from TLS.
Figure (svg): The observed probability that RC4's second output byte is zero, at twice the value a random keystream would give.
Error analysis
From a design document for a wireless product, written in 2003.
Annotate
This is WEP, and Section 14.3 dissects it properly. Every primitive in it was reasonable; the assembly was not.
Edge cases
RC4 accepts keys from 40 to 256 bits, and the book advises against small sizes.
Discussion prompt
Work out where the boundary sits, and say why a legal option can be an indefensible one.
Hint: Compare 2⁴⁰ with Chapter 1's brute-force arithmetic.
Answer:
2⁴⁰ ≈ 1.1 × 10¹². At Chapter 1's generous 10⁹ trials per second that is about eighteen minutes on one machine, and seconds on a cluster. Forty-bit keys were an export-control artefact, not an engineering judgement.
2⁵⁶ ≈ 7.2 × 10¹⁶ was DES's key size and fell to purpose-built hardware in 1998. 2⁸⁰ is around the edge of feasibility for a well-funded adversary today.
2¹²⁸ is comfortably beyond exhaustion by the calculation in Section 1.1.3, and is the modern floor.
Why a legal option can be indefensible: a specification that permits a range will be deployed at the cheap end, because the cheap end is faster and satisfies the letter of the spec. Offering 40 bits guarantees that some products ship with 40 bits.
The design lesson generalises well past RC4: do not make insecure configurations available. Modern protocol design removes weak options rather than documenting them, which is why TLS 1.3 deleted whole families of ciphersuites instead of deprecating them.
Picture it
Two axes, five generators. The empty corner — fast and provably unpredictable — is empty for a reason, and the reason is that nobody has found a construction that occupies it.
Figure (svg): Four keystream generators compared on speed against how well their unpredictability is understood.
Practical cryptography lives in the middle: fast enough to use, resting on an assumption well-studied enough to bet on. That is not a compromise anyone is embarrassed about; it is what the field is.
Constraint
You are encrypting live video from a fleet of drones: 40 Mbit/s per link, an 8-bit microcontroller on the drone, a lossy radio channel that drops and corrupts packets, and a two-year deployment.
Discussion prompt
Which of this chapter's options survive the constraints, and which constraint eliminates which option?
Hint: Take the constraints one at a time and cross options off.
Answer:
The throughput eliminates Blum-Blum-Shub immediately. A modular squaring per bit at 40 Mbit/s is not close to feasible on any processor, let alone an 8-bit one.
The lossy channel favours a stream cipher. A flipped bit corrupts one bit rather than a whole block, and there is no padding to be desynchronised. But it also demands that each packet be independently decryptable, so the keystream must be re-derivable from a per-packet nonce.
The 8-bit processor argues for RC4 on paper — it is byte-oriented and was designed for exactly this class of hardware. The two-year deployment argues against it: RC4's biases were already fatal in TLS, and a system shipping today should not be built on it.
The LFSR is out whatever the hardware. Known plaintext is abundant in video — headers, black frames, static backgrounds — and 2m bits is nothing.
The answer is ChaCha20 with a per-packet nonce: designed for constrained processors, no known distinguisher, and a nonce discipline that makes keystream reuse structurally impossible. Add a MAC, because a stream cipher without integrity lets an attacker flip any plaintext bit she chooses — and in a video feed that is a real capability.
Notice that the constraints did most of the work, and the security argument only had to settle the last two candidates.
Scale up
Grow an LFSR and track both numbers side by side. This is the trap slide made quantitative.
Step through it
Multiplying the period by 10²⁹ multiplied the attack cost by what?
By four. The period is exponential in m and the attack is linear in m, so every bit added to the register buys the defender a factor of two and costs the attacker two bits of data. Growing m is not a defence.
Pattern
Three generators, three breaks, and the same three questions each time.
Notice that none of the three is about the period, and none is about the key length — the two numbers that a specification sheet actually reports. The parameters that get advertised are not the parameters that decide the outcome.
The positive version: a good stream cipher is nonlinear in its state, structurally incapable of repeating a keystream, and has no known distinguisher. ChaCha20 was designed against exactly this list.
Figure (svg): Four keystream generators compared on speed against how well their unpredictability is understood.
Check
Work it out before you click.
Check your understanding
An LFSR of length 40 has period 2⁴⁰ − 1 ≈ 10¹². How much known plaintext does an attacker need to recover the recurrence?
Answer: C
Why: 2m bits, so 80 — about ten bytes. Each equation needs a window of 41 consecutive bits and 40 equations are required, spanning 80 bits in total. The period is irrelevant to the attack: the cost is linear in the register length while the period is exponential in it, which is exactly why a large period is no defence.
Check
Think about which property the application actually needs.
Check your understanding
You need to generate a 256-bit AES key. Which source is appropriate?
Answer: C
Why: A key must be unpredictable, and only the one-way-function construction offers that. It is the method the book recommends and the one OpenSSL uses. The seed must come from real entropy — a time-seeded generator has only as many possible outputs as there were plausible seconds, which is a few million.
Check
Apply Section 4.5's definition.
Check your understanding
A keystream generator's first output byte is 0 with probability 3/256 instead of 1/256. What follows?
Answer: B
Why: A measurable bias is a distinguisher, so the generator loses the game of Section 4.5 by definition. Practically, if the same plaintext byte is encrypted in many sessions — a cookie, a fixed header — the attacker collects ciphertexts and the biased keystream lets the byte be recovered statistically. This is the reasoning that removed RC4 from TLS.
Connect it up
This chapter is one decision made three ways. Write it out once.
Draw it
Draw three columns: linear congruential, LFSR, RC4. For each, record (1) the state and how big it is, (2) how the next output is computed from the state, (3) the best known attack and how much data it needs, (4) whether the failure is structural or a parameter choice. Then add a fourth column for what a modern design such as ChaCha20 does differently, and underneath write the three questions from the pattern slide that decide all four.
Row (2) is where the whole answer lives: linear in the state means broken, and everything else in the table follows from it.
Exit ticket
One question, and it is the one the whole chapter argues.
Predict first
What property must a keystream generator have, that period and key length do not guarantee?
Correct: Unpredictability — no efficient strategy guesses the next bit better than chance, given all previous bits
Why: The LFSR with a two-billion period, a large state and blistering speed is recovered from twelve bits, because its output is a linear function of its state. Period, speed and state size are all properties the specification advertises and none of them is the property that matters. Unpredictability is the requirement, it is the indistinguishability game of Section 4.5 applied bit by bit, and it is what a one-way function or a hardness assumption is for.
Recap
Three generators and the same lesson three times.
Chapter 6 next. Stream ciphers process a bit at a time and inherit the pad's fragility about reuse. Block ciphers collect a whole block and transform it at once — which removes some problems, introduces the question of what to do with the second block, and makes modes of operation a security topic rather than a formatting one.
Figure (svg): Four keystream generators compared on speed against how well their unpredictability is understood.
Want this taught 1-on-1? Alexander tutors Cryptography — $55/session, free consultation.