Chapter 9 of Trappe & Washington: RSA set up and worked through with the book's own numbers, its four-line correctness proof from Euler's theorem, and the equivalence of factoring, computing phi(n) and finding d. Covers the Fermat, Miller-Rabin and Solovay-Strassen compositeness tests; Fermat factorization, Pollard's p-1 method and the x^2 = y^2 principle behind the quadratic sieve; the low-exponent and Wiener attacks that break RSA without factoring anything; the RSA-129 challenge; and treaty verification, which is a digital signature three chapters early.
Subject: Cryptography · 61 slides · diagram-first lesson
Open the interactive version of this deck
Title
Cryptography · Chapter 9
Public key cryptography that works: multiply two primes, publish the product, and keep the factors
Objectives
Every cipher so far has assumed Alice and Bob already share a key. This chapter drops that assumption, and Chapter 3's number theory stops being preparation and becomes the security itself.
Figure (svg): Bob publishes n and e, Alice encrypts with them, and only Bob's private d decrypts.
Warm-up
Alice wants to send Bob a message. They have never met, there is no courier, and Eve sees everything on the channel.
Discussion prompt
Why is every cipher from Chapters 2 to 8 useless here, and what would have to be true for a solution to exist?
Hint: Ask what Alice would have to send first.
Answer:
Every symmetric cipher needs the key to arrive first, and the only channel available is the one Eve is watching. Alice sends the key, Eve reads it, and every subsequent message is open. There is no arrangement of symmetric primitives that escapes this.
So a solution requires something new: a way for Bob to publish enough for Alice to encrypt, while withholding enough that nobody else can decrypt. The encryption and decryption keys must be different, and the second must not be computable from the first.
Diffie and Hellman described this possibility in 1976 without a practical implementation. Rivest, Shamir and Adleman supplied one in 1977, based on the difficulty of factoring.
A footnote worth knowing: British agency CESG documents released in 1997 showed James Ellis had the concept in 1970 and Clifford Cocks had written down a version of RSA in 1973 — with e equal to n. Secrecy rules kept it unpublished for twenty-four years.
Figure (svg): Multiplication is easy and factoring is hard, with the private key sitting on the hard side.
Section
Section 9.1 · pp. 171-177
Concept
Bob does all the setup. Alice needs to know nothing secret.
\[ c \equiv m^{e} \pmod n, \qquad m \equiv c^{d} \pmod n \]
Alice writes her message as a number m < n — breaking it into blocks if it is larger — and sends c. Note that Alice, who could be an enemy of Bob, never needs p or q.
Figure (svg): Bob publishes n and e, Alice encrypts with them, and only Bob's private d decrypts.
Worked example
The book's own numbers, with letters numbered a = 01 through z = 26 rather than from 0 — otherwise a leading a would vanish as a leading 00.
Bob picks p = 885320963 and q = 238855417, so n = 211463707796206571
Why: Nine-digit primes give an eighteen-digit modulus. Real RSA uses primes of about 300 digits.
He chooses e = 9007 and publishes (n, e)
Why: gcd(9007, (p−1)(q−1)) = 1, which is what makes d exist.
Alice encodes cat as c=03, a=01, t=20, giving m = 30120
Why: The leading zero of 03 is dropped, which is why the numbering starts at 1.
She computes c ≡ 30120⁹⁰⁰⁷ ≡ 113535859035722866 (mod n)
Why: By square-and-multiply from Section 3.5 — about 14 squarings, not 9006 multiplications.
Bob knows p and q, so he knows (p−1)(q−1) = 211463706672030192 and finds d = 116402471153538991
Why: One run of the extended Euclidean algorithm.
Verify: c^d ≡ 113535859035722866^116402471153538991 ≡ 30120 (mod n)
Why: The message returns. Every one of these numbers was recomputed before it was written down — the point of a worked example is that it is checkable, and this one checks.
Figure (svg): The book's RSA example: two primes, their product, the exponents, and the message travelling as ciphertext.
Worked example
Four lines, and the only ingredient is Euler's theorem from Section 3.6.
φ(n) = φ(pq) = (p−1)(q−1), since p and q are distinct primes
Why: Section 3.6's multiplicativity of φ.
de ≡ 1 (mod φ(n)) means de = 1 + kφ(n) for some integer k
Why: This is what the extended Euclidean algorithm arranged.
So c^d ≡ (m^e)^d = m^(de) = m^(1 + kφ(n)) = m · (m^φ(n))^k
Why: Just index laws.
Euler's theorem says m^φ(n) ≡ 1 (mod n) whenever gcd(m, n) = 1
Why: Which is overwhelmingly likely — p and q are enormous, so m almost certainly has neither as a factor.
\[ c^{d} \equiv m \cdot 1^{k} \equiv m \pmod n \]
Verify: and when gcd(m, n) ≠ 1, Bob still recovers the message
Why: The book leaves this to Exercise 37; the Chinese Remainder Theorem handles the case where p or q divides m. So decryption is correct for every message, not merely almost all of them.
Figure (svg): The correctness proof of RSA in four lines, resting on Euler's theorem from Section 3.6.
Notation
Four symbols and two congruences carry the whole system. Each has a job.
Annotate
On: \( n = pq, \quad \gcd(e, \varphi(n)) = 1, \quad de \equiv 1 \!\pmod{\varphi(n)}, \quad c \equiv m^{e} \!\pmod{n} \)
Every line of this appeared in Chapter 3. RSA is number theory with a security claim attached.
Socratic
Eve knows c ≡ mᵉ (mod n) and she knows e. Taking an e-th root sounds elementary.
Discussion prompt
Why does that not work, and what does the book's small example show?
Hint: Compare taking a cube root of an integer with taking a cube root modulo something.
Answer:
Over the integers, roots work because of order. The cube root of 3 is 1.4422…, and you find it by narrowing an interval — bigger inputs give bigger outputs, so bisection converges.
Modular arithmetic has no order. The book's example: if m³ ≡ 3 (mod 85), computing 1.4422… on a calculator and reducing mod 85 gives nothing, because the map x ↦ x³ mod 85 jumps around with no monotonicity to exploit.
A case-by-case search would find m = 7, and for a 600-digit modulus a case-by-case search is Chapter 1's impossible calculation.
The general point is important. Every technique for inverting a function over the reals — bisection, Newton's method, logarithms — depends on continuity or order, and modular arithmetic has neither. That is why the problems in this chapter are hard, and it is why they are attacked with number theory rather than with analysis.
It also explains why knowing d makes it easy: d converts the root extraction into another exponentiation, which is cheap. The trapdoor is not extra computing power, it is a shortcut.
Prediction
Eve has n, e and c. The obvious route to d is through φ(n).
Predict first
How hard is it to compute φ(n) from n alone?
Correct: As hard as factoring n
\[ p + q = n - \varphi(n) + 1, \qquad pq = n \;\Longrightarrow\; p, q \text{ are the roots of } x^2 - (p+q)x + n \]
The book also shows the other direction — if Eve finds d she can probably factor n — so all three of 'factor n', 'compute φ(n)' and 'find d' stand or fall together. That is a clean statement of exactly what RSA's security rests on.
Why: For n = pq, knowing φ(n) = (p−1)(q−1) = n − p − q + 1 gives p + q, and knowing both the sum and product of p and q means solving a quadratic — so φ(n) yields the factorization immediately. Conversely the factorization gives φ(n). The two problems are equivalent, so there is no back route to d that avoids factoring.
Fill the middle
Complete the two conditions Bob must satisfy.
Fill in the blanks
\text(p−1)(q−1) e \text(p−1)(q−1) \gcd\bigl(e, ___\bigr) = 1, \quad \text___ d \equiv e^___ \!\pmod___}
Why: Both blanks are φ(n) = (p−1)(q−1), and both are the modulus for the exponent arithmetic rather than n. This is Chapter 3's rule — bases mod n, exponents mod φ(n) — and it is the single most common error in implementing RSA by hand. Note that d is computed with one run of the extended Euclidean algorithm, which is logarithmic time; the expensive part of key generation is finding p and q, not finding d.
Explain it
A colleague says it cannot possibly work: if the encryption rule is public, anyone can reverse it.
Discussion prompt
Answer the objection precisely, and say where the colleague's intuition comes from.
Hint: Their intuition is correct about functions in general and wrong about this one.
Answer:
The intuition is right for most functions. If you can compute f, you can usually invert it — try inputs, use the structure, use order and continuity. Nearly every function anyone meets is like this.
Modular exponentiation is not. Computing mᵉ mod n is a dozen squarings. Recovering m needs an e-th root mod n, and modular arithmetic has no order to bisect on and no continuity to follow — Section 9.1's example is that knowing m³ ≡ 3 (mod 85) tells you nothing useful about m, because 1.4422… is not an answer here.
And there is an asymmetry of knowledge, not of ability. Bob is not cleverer or better equipped; he simply knows two numbers whose product everyone else can see. That knowledge turns the root extraction into another exponentiation.
Where the colleague should stay sceptical: this is a computational claim, not a theorem. It holds because nobody has found a fast factoring algorithm, and Chapter 25 shows a quantum computer would find one. The right response to the objection is not 'it is impossible' but 'nobody has managed it in fifty years, and here is exactly what would change that'.
Definition probe
RSA involves six quantities and mixing up which are public is the fastest route to a broken implementation.
Sort into buckets
Sort each quantity.
Section
Section 9.2 · pp. 177-183
Concept
The book states two theorems of the same shape, and both are more damaging than they first look.
Coppersmith: if n has m digits and you know the first m/4 or the last m/4 digits of p, you can efficiently factor n. With 300-digit primes, knowing half of p finishes it.
And this bites a plausible implementation. Suppose p is found by taking a random 150-digit N and testing N·10¹⁵⁰ + k for k = 1, 3, 5, … until a prime appears — which happens for k under 10 000. An attacker who knows that method knows 147 of the last 150 digits, since they are all zero except the last three or four. Trying the theorem for each k under 10 000 factors n.
The second theorem: if you have at least the last m/4 digits of d, you can find d in time linear in e log₂ e. So for a small e — and e = 65537 is small — knowing a quarter of d recovers all of it.
The lesson is that RSA's secrets are brittle. A symmetric key leaks gracefully: knowing half of an AES key still leaves 2⁶⁴ to search. Knowing half of an RSA prime leaves nothing.
Figure (svg): The safe and unsafe regions for RSA's exponents: e too small and d too small are both attackable.
Concept
A small encryption exponent speeds up encryption, and e = 3 is the smallest useful choice. It is also dangerous.
If the same message m is sent to three recipients with e = 3 and different moduli n₁, n₂, n₃, then the Chinese Remainder Theorem of Section 3.4 combines the three ciphertexts into m³ mod n₁n₂n₃. But m³ is smaller than that product, so the congruence is an equation — and an ordinary integer cube root recovers m.
No factoring, no key recovery. Three ciphertexts and a cube root.
The book's recommended choice is e = 65537 = 2¹⁶ + 1. It is prime, so it is very likely coprime to (p−1)(q−1); and being one more than a power of two, x^65537 costs sixteen squarings and one multiplication.
The general fix is randomised padding — OAEP — which makes the message different every time, so identical plaintexts never produce a solvable system. Chapter 4's indistinguishability requirement, arriving with teeth.
Figure (svg): The low exponent attack: three ciphertexts of the same message under e = 3 combine by the CRT into a plain integer cube.
Concept
Bob wants fast decryption, so he picks a small d and derives e from it. Wiener showed that Eve can then find d easily.
Start from ed − kφ(n) = 1 and divide by dφ(n):
\[ \left| \frac{e}{\varphi(n)} - \frac{k}{d} \right| = \frac{1}{d \, \varphi(n)} \]
Since φ(n) = n − p − q + 1 is close to n, the fraction e/n is very nearly k/d. And Section 3.12's theorem says the convergents of a continued fraction are the best rational approximations for their denominator size — so if d is small, k/d must be one of the convergents of e/n.
There are only about log n convergents. Eve computes them, tests each as a candidate, and recovers d in seconds. The attack works whenever d is below roughly n^(1/4).
So do not choose a small d. The right way to speed up decryption is the Chinese Remainder Theorem — compute mod p and mod q separately for a fourfold speed-up, as Section 3.4 showed — which costs nothing in security.
Figure (svg): Wiener's attack: the convergents of the continued fraction of e over n, one of which is k over d.
Two truths and a lie
Two of these will get a system broken. One is the standard advice.
Eliminate the wrong options
Which statement is correct?
Survives elimination: a
Why: e = 65537 is prime, so gcd(e, φ(n)) = 1 for almost every key, and it is 2¹⁶ + 1, so exponentiation is sixteen squarings and one multiply — nearly as fast as e = 3 without the attack. A full-length d resists Wiener, and the CRT gives roughly a fourfold decryption speed-up for free. Note that both wrong options trade security for speed, and in both cases the speed was available a safer way.
Figure (svg): The safe and unsafe regions for RSA's exponents: e too small and d too small are both attackable.
Explain it to yourself
The condition is stated as part of the setup, alongside choosing the primes.
Discussion prompt
Explain what breaks without it, and connect it to a failure you have already seen.
Hint: d is defined as an inverse. When do inverses exist?
Answer:
Without it, d does not exist. d is defined by de ≡ 1 (mod φ(n)), which is exactly asking for e's inverse mod φ(n) — and Section 3.3 says that inverse exists precisely when gcd(e, φ(n)) = 1.
So the map m ↦ mᵉ is not invertible mod n. Several messages share a ciphertext and no receiver could tell which was sent.
You have seen this exact failure before. Chapter 2's affine cipher with α = 13 sent both input and alter to ERRER, because gcd(13, 26) ≠ 1. It is the same condition, the same consequence, and the same one-line reason.
Which is why e = 65537 is a good default beyond its speed: it is prime, so gcd(e, φ(n)) = 1 unless 65537 happens to divide φ(n) — overwhelmingly unlikely, and cheap to check.
The recurring shape: encryption must be a bijection, and over Z/n the condition for that is always a gcd.
Commit first
You are auditing an RSA deployment and have time to examine one thing thoroughly.
Predict first
Which is most likely to be broken?
Correct: The padding scheme and the randomness behind it
Bleichenbacher's attack is the RSA analogue of Chapter 6's padding oracle: the server reports whether a decrypted value has valid padding, and about a million adaptive queries recover the plaintext. It was published in 1998 and has been rediscovered in deployed systems repeatedly since — ROBOT in 2017 found it still live in major TLS stacks.
A close second would be key generation entropy: large-scale surveys of internet-facing RSA keys have found thousands sharing a prime factor, because devices generated keys at first boot with almost no entropy. A shared factor means one gcd between two public keys breaks both.
Why: Padding is where deployed RSA actually fails, and it fails in two directions: no padding at all loses indistinguishability and enables low-exponent attacks, while badly implemented PKCS#1 v1.5 padding produces the Bleichenbacher oracle, which recovers plaintext from an error-message difference. Modulus size is almost always adequate and is the one thing every audit checklist already contains.
Section
Section 9.3 · pp. 183-188
Concept
Bob needs 300-digit primes. Trial division is hopeless: there are about 3 × 10¹⁴⁷ primes below 10¹⁵⁰, more than the particles in the universe, and at 10¹⁰ primes per second the calculation takes 3 × 10¹³⁰ years.
The book's aside is worth quoting: you could sit on the beach for twenty years, buy a computer a thousand times faster, and cut the runtime to 3 × 10¹²⁷ years — a very large saving.
But primality testing and factorization are not the same problem, and the gap is enormous. It is far easier to prove a number composite than to factor it, and there are many large integers known to be composite that have never been factored.
Basic Principle. If x² ≡ y² (mod n) but x ≢ ±y (mod n), then n is composite and gcd(x − y, n) is a non-trivial factor. Example: 12² ≡ 2² (mod 35) and 12 ≢ ±2, so 35 is composite and gcd(10, 35) = 5.
Figure (svg): Primality testing is far cheaper than factoring: a 300-digit number is proved composite in milliseconds and factored in never.
Concept
Fermat's theorem says that if p is prime and p ∤ a, then a^(p−1) ≡ 1 (mod p). Run it backwards as a test.
\[ a^{n-1} \not\equiv 1 \pmod n \;\Longrightarrow\; n \text{ is composite} \]
One modular exponentiation, which is cheap by square-and-multiply. If the answer is not 1, n is definitely composite — and the test gives no factor, which is the crucial point.
The book is precise about the direction: these are called primality tests but they are really compositeness tests. Failing proves composite; passing proves nothing on its own.
And there are composites that pass for every a coprime to n — the Carmichael numbers, of which 561 is the smallest. So the Fermat test alone is not enough, which is why the next two tests exist.
Figure (svg): The Fermat test as a one-way conclusion: failing proves composite, passing proves nothing.
Concept
Miller-Rabin strengthens the Fermat test by looking at the square roots taken along the way.
Write n − 1 = 2^k · m with m odd. Compute b₀ = a^m (mod n), then square repeatedly. If n is prime, the sequence must reach 1, and the step before it must be −1 — because mod a prime the only square roots of 1 are ±1.
\[ a^{m}, \; a^{2m}, \; a^{4m}, \; \ldots, \; a^{2^{k}m} = a^{n-1} \]
If the sequence reaches 1 without passing through −1, we have found a square root of 1 that is neither 1 nor −1 — and the Basic Principle then says n is composite, with a factor available from a gcd.
The book gives the error rate: the probability that Miller-Rabin fails to recognise a composite, for a randomly chosen a, is under 1/4. Repeat with several a and the failure probability falls geometrically. And it finds a 300-digit prime in under 400 tests on average, because the primes are dense enough.
Figure (svg): The Miller-Rabin sequence of repeated squarings, with the check that 1 must be reached via minus one.
Worked example
Two small numbers that show what each test does and does not catch.
561 = 3 · 11 · 17 is composite, but a⁵⁶⁰ ≡ 1 (mod 561) for every a coprime to 561
Why: It is the smallest Carmichael number, and the Fermat test never detects it.
Miller-Rabin does detect it: write 560 = 2⁴ · 35 and square from a³⁵
Why: For most a the sequence reaches 1 without passing through −1, which proves compositeness and yields a factor by gcd.
For 299: the book notes Miller-Rabin with a = 2 reports it is not prime
Why: And 299 = 13 · 23.
Importantly, the test reports composite without necessarily naming the factors
Why: It is a compositeness test that sometimes hands you a factor as a side effect, not a factoring algorithm.
Verify: this asymmetry is what makes RSA key generation practical
Why: Bob generates a random 300-digit odd number and runs Miller-Rabin. Under 400 attempts on average finds a prime, at one exponentiation each — milliseconds. Meanwhile Eve, given the product, has no comparable shortcut. Key generation is cheap and key breaking is not.
Figure (svg): Primality testing is far cheaper than factoring: a 300-digit number is proved composite in milliseconds and factored in never.
Concept
A second test of similar character, using the Jacobi symbol from Section 3.10.
\[ \left(\frac{a}{n}\right) \equiv a^{(n-1)/2} \pmod n \]
For a prime this is Euler's criterion and always holds. For a composite it fails for at least half the a. So compute both sides — the Jacobi symbol by quadratic reciprocity, which needs no factorization, and the power by square-and-multiply — and compare.
The book's examples: the test shows 15 is not prime, and it correctly reports that 341 is composite where a naive Fermat test with a = 2 would not.
Both tests are compositeness tests, and both are probabilistic in the same direction: they never call a prime composite, and they may occasionally call a composite prime. Miller-Rabin is generally preferred because its error bound of 1/4 is better than Solovay-Strassen's 1/2, and it is cheaper.
Figure (svg): The Solovay-Strassen test comparing the Jacobi symbol with a modular power, both computable without factoring.
Translation
Five procedures from this chapter, and it matters which of them hands you a factor.
Match the pairs
Why: The top two prove a number composite without factoring it, which is the asymmetry that makes RSA key generation cheap — Bob can find 300-digit primes in milliseconds. The bottom three are factoring methods, and the middle two only work on moduli with a specific weakness, which is why key generation must actively test for both. Only the quadratic sieve and its successors work on a general modulus, and they are the reason key sizes have had to grow.
Faded example
Fill in the reasoning that makes the method work.
Fill in the blanks
If p − 1 has only small prime factors, then p − 1 probably divides B!. So b ≡ a^(B!) ≡ 1 (mod p) by Fermat's theorem, which means p divides b − 1 — and since p also divides n, the factor appears in gcd(b − 1, n).
Why: The whole method is Fermat's theorem pointed at factoring rather than at primality: a^(p−1) ≡ 1, so any exponent that is a multiple of p−1 also gives 1 mod p, and B! is a multiple of p−1 whenever p−1 is smooth. The gcd then separates p from the other factors, for which the congruence almost certainly fails. It is why RSA key generation must reject a prime whose predecessor is smooth.
Section
Section 9.4 · pp. 188-192
Concept
Express n as a difference of two squares and the factorization falls out.
\[ n = x^2 - y^2 = (x + y)(x - y) \]
Compute n + 1², n + 2², n + 3², … until one is a perfect square. The book's example: 295927 + 3² = 295936 = 544², so 295927 = 547 · 541.
It works well when n is a product of two primes that are close together, and it takes |p − q|/2 steps. Two random 300-digit primes will differ by something around 300 digits, so it is hopeless against a properly generated modulus.
But it is the reason the book advises choosing p and q of slightly different lengths — just to be safe. A generator that happened to produce nearby primes would hand the modulus over for nothing.
Figure (svg): Fermat factorization on 295927, finding a perfect square after three steps.
Concept
Pollard's method from 1974. If some prime factor p of n has p − 1 with only small prime factors, n falls.
Why it works: if p − 1 has only small prime factors, it likely divides B!, say B! = (p−1)k. Then Fermat gives b ≡ a^(B!) ≡ (a^(p−1))^k ≡ 1 (mod p), so p divides b − 1 and appears in the gcd. For another prime factor q it is unlikely that b ≡ 1 (mod q), unless q − 1 is also smooth.
The consequence for key generation is a real requirement: the book says if p − 1 has only small prime factors then p should be rejected and replaced. A prime is not enough; it must be a prime whose predecessor is not smooth.
Section 21.3's elliptic curve method generalises this, and its advantage is that it does not depend on p − 1 being smooth — it gets to try many groups instead of one.
Figure (svg): The p minus 1 method: a smooth p minus 1 divides B factorial, so Fermat's theorem forces a gcd to reveal p.
Concept
Every modern factoring method rests on the Basic Principle from Section 9.3. Find x and y with x² ≡ y² (mod n) but x ≢ ±y, and gcd(x − y, n) is a factor.
The problem is finding such a pair. The method is to collect many relations — numbers whose squares mod n factor completely over a fixed set of small primes, the factor base — and then find a subset whose product has every prime to an even power. That product is a perfect square on both sides, which is exactly the pair wanted.
Finding the right subset is a linear algebra problem over GF(2): one column per prime in the factor base, one row per relation, and a dependency in the matrix gives the square.
The quadratic sieve is an efficient way of generating relations; the number field sieve is a further improvement and is the fastest known general method. Both are sub-exponential — much better than trial division, and still far from polynomial, which is the gap RSA lives in.
Figure (svg): The x squared congruent to y squared principle: a non-trivial square root of one gives a factor by a single gcd.
Section
Section 9.5 · pp. 192-194
Concept
In 1977 Rivest, Shamir and Adleman published a 129-digit modulus, the exponent e = 9007, and a ciphertext, and offered $100 to anyone who could read the message before April 1, 1982.
Their estimate, using the factoring methods of 1977, was that it would take 4 × 10¹⁶ years.
It was factored in 1994, by Atkins, Graff, Lenstra and Leyland. The effort used 524 339 small primes below 16 333 610 as the factor base, allowed up to two large primes per relation, and ran on 1600 computers belonging to 600 people in their spare time, over about eight months.
The resulting matrix had 524 339 columns and 569 466 rows. It was sparse, so it could be stored; Gaussian elimination reduced it in under twelve hours, and another 45 hours of computation produced 205 dependencies. The first three gave the trivial factorization; the fourth gave p and q.
The plaintext was the magic words are squeamish ossifrage — a squeamish ossifrage being an overly sensitive hawk. The phrase was chosen so nobody could claim the prize by guessing the message and checking that it encrypted correctly.
Figure (svg): The RSA-129 challenge: estimated at 4 times 10 to the 16 years in 1977 and factored in eight months in 1994.
Anomaly
The 1977 estimate was not incompetent — it was a correct calculation using the best methods then known.
Predict first
What went wrong with the forecast?
Correct: Factoring algorithms improved dramatically, and algorithmic progress is much harder to forecast than hardware progress
Compare with DES in Chapter 7. There, Diffie and Hellman's 1977 forecast was accurate to about a factor of a hundred over twenty-one years, because the attack was exhaustive search and its cost tracks hardware.
The general lesson: parameters defended by exhaustive search can be forecast; parameters defended by 'no good algorithm is known' cannot. RSA key sizes have had to grow repeatedly — 512 bits, then 1024, now 2048 or 3072 — and each increase was a response to algorithmic progress, not to faster machines.
Which is also why Chapter 25's quantum computing section matters so much: Shor's algorithm is not a faster machine, it is a different cost function.
Why: Hardware improved by perhaps four orders of magnitude over those seventeen years; the estimate was out by more than twenty. The rest came from the quadratic sieve and its successors, which changed the exponent in the cost function rather than the constant. Nothing about RSA itself was broken — a 129-digit modulus was factored, and the algorithm is used today with 617-digit ones.
Figure (svg): The RSA-129 challenge: estimated at 4 times 10 to the 16 years in 1977 and factored in eight months in 1994.
Edge cases
Not every large number is hard to factor. RSA moduli are chosen to be hard.
Discussion prompt
Which large numbers are easy to factor, and what does that tell you about generating a modulus?
Hint: Every factoring method in this chapter has a precondition. List them.
Answer:
Numbers with a small factor fall to trial division in moments, so both primes must be large.
Numbers whose factors are close together fall to Fermat factorization in |p − q|/2 steps — hence the advice to make p and q of slightly different lengths.
Numbers where some p − 1 is smooth fall to Pollard's p−1 method, so key generation must test for and reject them.
Numbers with structured digits fall to Coppersmith if a quarter of the digits are predictable — so no fixed patterns in the generator.
And numbers sharing a factor with another public key fall to a single gcd. This is not hypothetical: surveys of internet-facing keys have found thousands of hosts whose moduli share a prime, because they generated keys at first boot with almost no entropy.
So the hardness is a property of the generation process, not of the size. A 4096-bit modulus can be trivially factorable, and knowing all five preconditions is what a key generator is for.
Cost model
RSA's security parameter is the cost of the best factoring algorithm, and its shape explains every key-size recommendation you have seen.
Annotate
On: \( L_n\left[\tfrac{1}{3}, c\right] = \exp\!\left(c \, (\ln n)^{1/3} (\ln\ln n)^{2/3}\right) \)
Three families, three cost shapes: exponential for AES, square-root for elliptic curves, sub-exponential for RSA. Every key-size table in the world is these three formulas set equal to each other.
Section
Section 9.6 · pp. 194-195
Concept
Countries A and B have signed a nuclear test ban treaty. A wants seismic sensors inside B, and two requirements look contradictory:
The solution reverses RSA. A chooses n = pq and the exponents, gives B the pair (n, e), and keeps p, q and d secret. The tamper-proof sensor, buried in B's territory, collects data x and computes y ≡ x^d (mod n) — using the private exponent.
Both x and y are sent to B, which checks y^e ≡ x. If it holds, B knows y corresponds to exactly the data x — which it can read in full — and forwards the pair to A. A checks the same relation.
A can trust x because producing a valid y for a chosen x would require decrypting an RSA message, which is believed hard. B could pick y first and set x ≡ y^e, but then x would be meaningless data and A would notice.
Figure (svg): Treaty verification by running RSA backwards: the sensor signs with the private key and both countries verify with the public one.
Socratic
The sensor applies the private exponent and everyone verifies with the public one. Nothing is hidden — B reads all the data.
Discussion prompt
This is not encryption. What is it, and what property is it providing?
Hint: Ask what the pair (x, y) proves and to whom.
Answer:
It is a digital signature, and Chapter 13 is entirely about it. The pair (x, y) proves that whoever produced it holds d, and anyone with (n, e) can verify that.
The property is authentication and integrity, not confidentiality. The data travels in the clear on purpose — that is B's requirement — and what RSA supplies is a proof that the data was not altered.
And it gives non-repudiation, the objective Chapter 1 said symmetric cryptography cannot provide. Only A's sensor holds d, so A cannot later claim the reading was fabricated, and B cannot fabricate one.
The structural insight is that RSA's two exponents are symmetric in use. Encrypt with the public and only the holder of the private can read; apply the private and anyone can verify who applied it. The same equation, read in two directions, gives two different services.
A caution the book returns to in Chapter 13: signing a raw message this way is not safe in general — the message must be hashed first, or an attacker can construct valid signatures by multiplying existing ones. Here it works because the sensor signs only its own readings.
Section
Section 9.7 · pp. 195-197
Concept
Diffie and Hellman described the concept in 1976 with no public implementation. Here is the general shape any realisation must satisfy.
There is a set M of possible messages and a set K of keys — in RSA a key is a triple (e, d, n) with ed ≡ 1 mod φ(n). For each key there is an encryption function E_k and a decryption function D_k, both mapping M to M, and:
The book names three other realisations covered later: ElGamal on discrete logarithms (Chapter 10), NTRU on lattices (Chapter 23), and McEliece on error-correcting codes (Chapter 24). Knapsack-based systems also exist; some versions have been broken and they are generally suspected to be weaker.
Figure (svg): Four public key systems, each resting on a different problem, with two of them surviving Shor's algorithm.
Counterexample
Two devices independently generate RSA moduli n₁ = p·q₁ and n₂ = p·q₂, having drawn the same p because both had almost no entropy at first boot.
Discussion prompt
How quickly can an attacker who has collected both public keys break them, and what does this say about where RSA fails?
Hint: The attacker does not need to factor anything.
Answer:
One gcd. gcd(n₁, n₂) = p, computed by the Euclidean algorithm in microseconds even for 2048-bit numbers. Then q₁ = n₁/p and q₂ = n₂/p, and both private keys follow.
And the attacker can do this at scale. Given a collection of N public keys, a batch-gcd computation finds every shared factor across the whole set in near-linear time. There is no need to guess which pairs to try.
This has happened, on a large scale. Surveys scanning the public internet in 2012 found tens of thousands of TLS and SSH hosts whose keys shared a factor — embedded devices generating keys on first boot before any entropy had accumulated.
What it says about RSA: the modulus size was irrelevant, the algorithm was correct, the exponents were fine, and the keys were worthless. The failure was Chapter 5's — the randomness feeding the generator — and it is the single most common way RSA is broken in the wild.
It is also an argument for the layered generator design of Section 5.1: collect real entropy slowly into a pool, and refuse to produce keys until the pool has enough.
Trade off
Four public key systems, and the choice between them is made on these rows. Fill the blanks.
Comparison matrix
| RSA | Elliptic curve (Ch. 21) | |
|---|---|---|
| Key size for ~128-bit security | 3072 bits | 256 bits |
| Best known attack | number field sieve — sub-exponential | generic — square root of the group size |
| Signature size | as large as the modulus | about twice the curve size |
| Verification speed | very fast with e = 65537 | slower than RSA verification |
| Survives Shor's algorithm? | no | no — both fall |
The last row is the one people get wrong: elliptic curves are not post-quantum. Shor's algorithm solves discrete logarithms as readily as factoring, so Chapters 23 and 24 — lattices and codes — are where the post-quantum candidates actually are.
Scale up
Every increase below was forced by algorithmic progress, not by faster hardware.
Step through it
Why has the recommended size always stayed roughly triple the record?
Because the cost curve is sub-exponential and shallow: the next few hundred bits are not far beyond the last. A margin that would be enormous for a symmetric key is a modest one here, which is the practical consequence of that cube root in the exponent.
Comparison
RSA's security rests on three problems that are all the same problem. Fill the blanks.
Comparison matrix
| Problem | Given | Equivalent to |
|---|---|---|
| Factor n | n | the other two |
| Compute φ(n) | n | factoring n |
| Find d | n and e | probably factoring n |
| Take an e-th root mod n | n, e and c | the RSA problem — possibly easier than factoring |
The last row is the honest caveat: nobody has proved that breaking RSA requires factoring. The RSA problem might be easier, and that gap has never been closed.
Discrimination
Several attacks in this chapter recover the plaintext or the key without ever factoring n.
Sort into buckets
Sort each attack.
Estimation
RSA-129 fell in 1994. RSA-768 (232 digits) fell in 2009. The number field sieve is sub-exponential.
Predict first
What modulus size is recommended for new systems today?
Correct: 2048 bits or more
The disparity is worth internalising. Symmetric security scales with the key length directly; RSA security scales with roughly the cube root of the modulus length, so getting more security costs disproportionately more.
It is the main practical argument for elliptic curves in Chapter 21: a 256-bit elliptic curve key gives about the same security as a 3072-bit RSA key, because the best attack on the curve group is far weaker.
Why: 2048 bits is the current minimum and 3072 is recommended for long-lived data. 512-bit moduli have been factorable for decades and 1024 is deprecated. Note how much larger these are than a symmetric key: 3072 bits of RSA gives roughly the security of a 128-bit AES key, because the number field sieve is far better than exhaustive search while nothing comparable exists against AES.
Figure (svg): The RSA-129 challenge: estimated at 4 times 10 to the 16 years in 1977 and factored in eight months in 1994.
Real world
RSA is nearly fifty years old and still ubiquitous, though its role has narrowed.
Discussion prompt
Name where RSA is still standard, where it has been displaced, and why in each case.
Hint: Chapter 1's hybrid pattern says what it was always for.
Answer:
Still standard: certificates and code signing. Almost every TLS certificate chain and every signed software package uses RSA signatures. The operations are infrequent, the verification cost is tiny with e = 65537, and the format is universally supported.
Displaced: key exchange. TLS 1.3 removed RSA key transport entirely in favour of ephemeral Diffie-Hellman — mainly for forward secrecy, since with RSA key transport a later compromise of the server's private key decrypts all recorded traffic.
Displaced: new signature deployments, increasingly, by elliptic curves. A 256-bit ECDSA or Ed25519 key matches a 3072-bit RSA key for security, with far smaller signatures and faster signing.
Never used: bulk encryption. Chapter 1's rule of thumb — public key methods should not encrypt large quantities of data — has held since 1977. RSA moves keys and signs hashes.
And on the horizon: Chapter 25's Shor's algorithm retires RSA entirely if a large quantum computer is built, which is why the post-quantum standards are lattice-based rather than RSA with a longer modulus.
Error analysis
From a code review of a licensing system.
Annotate
Four decisions, four documented attacks, and the modulus size — the only thing usually discussed — never comes into it.
Missing information
A vendor states: “We use RSA-4096.”
Discussion prompt
List what remains undetermined, and rank the gaps by how likely each is to be the real problem.
Hint: This chapter's attacks are mostly not about the modulus.
Answer:
Is there randomised padding? Without OAEP, RSA is deterministic — losing indistinguishability — and open to low-exponent and chosen-ciphertext attacks. This is the likeliest real problem.
How were p and q generated? Coppersmith punishes any structure in the primes, and the p−1 method punishes a smooth p−1. A 4096-bit modulus from a bad generator is worth nothing; large-scale surveys have found shared factors across thousands of deployed keys because of weak entropy at boot.
Which exponents? e = 3 and a small d are both attacks, and neither is visible from the key size.
Is RSA being used for bulk encryption? If so, either it is unusably slow or it is being applied block by block, which reproduces ECB's failure on top of RSA.
Is the signature scheme hashing first? Signing a raw message allows forgeries by multiplying signatures, as Chapter 13 shows.
And the ranking: padding, then key generation, then exponents. The modulus size — the only thing the claim states — is the parameter least likely to be the weakness, which is this course's most-repeated point.
Constraint
Write the requirements for a key generator that resists everything in this chapter.
Discussion prompt
List the checks p and q must pass and the exponent rules, with the attack each one blocks.
Hint: Walk the chapter's attacks in order and write the defence for each.
Answer:
p and q from a strong random source, with no fixed digits and no derived structure — against Coppersmith, and against the weak-entropy failures that have produced shared factors in the wild.
p and q of slightly different lengths, so |p − q| is large — against Fermat factorization, which needs |p − q|/2 steps.
Reject p if p − 1 is smooth, and likewise q — against Pollard's p−1 method. The book states this explicitly as a key-generation requirement.
Both primes verified by several Miller-Rabin witnesses, since one witness leaves a failure probability under 1/4 and several make it negligible.
e = 65537 — prime, so coprime to φ(n) for almost every key, and cheap to exponentiate — against the low exponent attack.
d full length, computed from e rather than chosen — against Wiener. Use the CRT for decryption speed, and verify each signature after producing it, against the Bellcore fault attack that Section 3.4 flagged.
Six requirements, and only one of them is about the size of anything.
Ranking
Rank by realistic likelihood of breaking a deployed RSA system.
Put in order
Why: Implementation errors break real systems constantly and are the subject of most of Section 9.2. Algorithmic progress is steady and has forced key sizes up repeatedly — RSA-129 in 1994, RSA-768 in 2009 — so it is a certainty on a long horizon. Shor's algorithm would end RSA outright but requires hardware that does not yet exist. And exhaustive search of the key space is not a thing anyone attempts, because factoring the modulus is enormously cheaper than trying every d.
Pattern
RSA is the first of four such systems in this book, and all four have the same five parts.
And one warning that generalises. RSA's security is not proved equivalent to factoring: taking e-th roots mod n might be easier, and nobody has closed that gap in fifty years. Chapter 10 shows ElGamal's security does reduce to a named problem, which is a stronger position to be in.
Figure (svg): Multiplication is easy and factoring is hard, with the private key sitting on the hard side.
Trap
The trap. RSA-129 was factored in 1994, and RSA-768 in 2009. The lesson is that the modulus must be large. Use 4096 bits and RSA is safe.
This is the reasoning behind most RSA configuration guidance, and it addresses one attack out of the six in this chapter.
Why it fails. Look at what actually breaks deployments. The low exponent attack recovers the message from three ciphertexts with e = 3 and never touches the modulus. Wiener's attack recovers a small d from a continued fraction and never touches the modulus. Coppersmith needs half the digits of p, which a structured prime generator hands over regardless of size. Deterministic encryption without padding loses indistinguishability at any modulus size.
A 4096-bit modulus with e = 3, no padding and structured primes is broken four different ways, none of which the size affects.
And even for factoring, size is a moving target rather than a fix. The number field sieve is sub-exponential, and each algorithmic improvement shortens every deployed key at once. Doubling the modulus does not double the security — RSA security scales roughly with the cube root of the modulus length, which is why 3072 bits is needed to match a 128-bit AES key.
The rule, for the sixth time in this course: a parameter defends against the attack that consumes it, and nothing else. Ask what the best attack is before asking how big the number should be.
Check
Work it out before you click.
Check your understanding
For RSA with p = 11, q = 13 and e = 7, what is d?
Answer: B
Why: φ(n) = (11−1)(13−1) = 120, and d must satisfy 7d ≡ 1 (mod 120). Checking: 7 · 103 = 721 = 6 · 120 + 1, so d = 103. Note the modulus is φ(n) = 120 and not n = 143 — exponents reduce mod φ(n), which is Chapter 3's rule and the commonest slip here. And always multiply back: one multiplication settles it.
Check
Read the parameters carefully.
Check your understanding
A system uses a 2048-bit modulus, e = 3, and no padding. The same message is sent to three different recipients. What is the fastest attack?
Answer: B
Why: The CRT gives m³ mod n₁n₂n₃, and since m < each nᵢ, m³ is smaller than the product — so the congruence is an equation over the integers and an ordinary cube root recovers m. This takes milliseconds and never touches any modulus. Randomised padding would prevent it by making the three plaintexts different.
Check
The distinction is the reason RSA key generation is practical.
Check your understanding
Miller-Rabin reports that a 300-digit number n is composite. What do you now know?
Answer: B
Why: A Miller-Rabin failure is a proof of compositeness: the witness exhibits a violation of a property every prime has. The test is probabilistic in only one direction — it may pass a composite, but it never fails a prime. And it names no factors, which is exactly why generating primes is cheap while factoring their product is not.
Connect it up
The algorithm is six lines. Its attacks are six more. Both fit on a page.
Draw it
Write the four setup steps and the two operations, then the four-line correctness proof with the theorem it uses named. Beside it, list the six attacks from this chapter — Coppersmith on partial p, partial d, low exponent, Wiener, Fermat factorization, p−1 — and against each write the key-generation or parameter rule that blocks it. Finish with the three equivalent hard problems, and the one honest caveat about whether breaking RSA requires factoring.
The caveat is the sentence most often left out: nobody has proved that taking e-th roots mod n is as hard as factoring, and fifty years have not closed the gap.
Exit ticket
One question, about what makes RSA a cryptosystem rather than a hard problem.
Predict first
What exactly is RSA's trapdoor?
Correct: Knowledge of p and q, which makes φ(n) and therefore d computable — while computing φ(n) from n alone is as hard as factoring
Why: A trapdoor is extra information that collapses a hard problem for its holder. Bob knows the factorization, so φ(n) = (p−1)(q−1) is arithmetic and d follows from one run of the extended Euclidean algorithm. Eve has n alone, and computing φ(n) from n is provably equivalent to factoring. Note that e is public — the security rests entirely on the asymmetry of information about the factors, which is what distinguishes a cryptosystem from a merely difficult problem.
Recap
The first system in this course that needs no pre-shared secret.
Chapter 10 next. A different hard problem: given g and gˣ mod p, find x. It gives Diffie-Hellman key exchange, the ElGamal cryptosystem, and — unlike RSA — a security proof that reduces to a named computational problem.
Figure (svg): Multiplication is easy and factoring is hard, with the private key sitting on the hard side.
Want this taught 1-on-1? Alexander tutors Cryptography — $55/session, free consultation.