Chapter 10 of Trappe & Washington: the discrete logarithm problem and the free parity bit; the Pohlig-Hellman algorithm worked through the book's own example mod 41; baby step giant step as a time-memory trade; index calculus and why its absence in elliptic curve groups is worth a factor of twelve in key size; bit commitment with its hiding and binding requirements; Diffie-Hellman key exchange together with the intruder-in-the-middle attack it does not resist; and the ElGamal cryptosystem, whose security reduces to the Computational Diffie-Hellman problem.
Subject: Cryptography · 60 slides · diagram-first lesson
Open the interactive version of this deck
Title
Cryptography · Chapter 10
A second hard problem: Diffie-Hellman key exchange, ElGamal, and a security proof that reduces to something named
Objectives
RSA bet everything on factoring. This chapter makes a different bet, and gets three things RSA does not have: a key exchange with no encryption in it, a cryptosystem that is randomised by construction, and a security reduction to a stated problem.
Figure (svg): Exponentiation is cheap in one direction and the discrete logarithm has no known feasible route back.
Warm-up
RSA already solves the key-distribution problem. Chapter 10 builds two more systems on a different hard problem.
Discussion prompt
Why bother? Give two reasons that are not 'variety for its own sake'.
Hint: One is about what happens if factoring falls; one is about what RSA cannot do at all.
Answer:
Independence of assumptions. If someone finds a fast factoring algorithm tomorrow, every RSA key in the world dies at once. Discrete logarithms are a separate bet, so a system built on them survives. Diversification is not decoration — it is the only hedge available when nothing is proved.
Key agreement without encryption. Diffie-Hellman lets two parties compute a shared secret with no message ever encrypted and no key ever transmitted. RSA can move a key only by encrypting it, which means the server's long-term private key can decrypt every recorded session later. Diffie-Hellman with fresh exponents gives forward secrecy, which RSA key transport structurally cannot.
And a third, which the chapter delivers: ElGamal's security reduces to a named problem — the Computational Diffie-Hellman problem — whereas nobody has ever proved that breaking RSA requires factoring. A reduction to a well-studied problem is a stronger security statement.
Section
Section 10.1 · pp. 211-212
Concept
Fix a prime p and a primitive root α mod p. Section 3.7 showed that the powers of α run through every non-zero residue, so for any β there is exactly one x with:
\[ \beta \equiv \alpha^{x} \pmod p, \qquad 0 \le x < p - 1 \]
That x is the discrete logarithm of β to the base α, written L_α(β). Computing β from x is a handful of squarings by Section 3.5. Computing x from β is the problem this chapter is about, and no feasible general method is known.
The name is exact: over the reals, log is the inverse of exponentiation and is easy because exponentiation is monotone. In Z/p there is no order to exploit, so the same inverse becomes intractable — the same reason RSA's e-th roots are hard.
Why a primitive root — If α is not a primitive root its powers cover only a subgroup, so some β have no logarithm at all and the problem is ill-posed. Section 3.7's φ(p−1) primitive roots are the legal bases.
Figure (svg): Exponentiation is cheap in one direction and the discrete logarithm has no known feasible route back.
Notation
Four objects, and confusing which is secret is the fastest way to misread every protocol in this chapter.
Annotate
On: \( \beta \equiv \alpha^{x} \pmod p \)
Compare RSA: there the secret was a factorization, here it is an exponent. Both are one number that makes an infeasible computation trivial.
Definition probe
The problem appears in several disguises across the rest of the book.
Sort into buckets
Sort each task.
Section
Section 10.2 · pp. 212-218
Concept
Before any algorithm, notice that one bit of x comes out of a single exponentiation.
Since α is a primitive root, p − 1 is the smallest exponent giving 1, so α^((p−1)/2) is a square root of 1 that is not 1 — meaning it is −1. Raise β ≡ αˣ to the (p−1)/2 power:
\[ \beta^{(p-1)/2} \equiv \alpha^{x(p-1)/2} \equiv (-1)^{x} \pmod p \]
So if β^((p−1)/2) ≡ +1 then x is even, and if it is −1 then x is odd. One exponentiation, one bit.
The book's example: to solve 2ˣ ≡ 9 (mod 11), compute 9⁵ ≡ 1 (mod 11), so x is even — and indeed x = 6.
This is not a curiosity. It says the discrete log leaks its low bit for free, which is why Section 10.3's bit commitment uses the second bit rather than the first, and why any protocol treating x as uniformly hidden must be examined carefully.
Figure (svg): The parity test: raising beta to the p minus one over two power gives plus or minus one according to the parity of x.
Concept
Pohlig and Hellman extended the parity trick to every prime power dividing p − 1.
\[ p - 1 = \prod_i q_i^{r_i} \]
For each prime power q^r, the algorithm computes x mod q^r digit by digit in base q — each digit costing a search over q values. The results are then combined by the Chinese Remainder Theorem of Section 3.4 into x mod (p−1).
The cost is dominated by the largest prime factor, because that is where the per-digit search over q values happens. If every q is small the whole computation is quick; if one q is large the algorithm is infeasible.
Which gives the design rule: if a discrete logarithm is to be hard, p − 1 must have a large prime factor. Choosing a prime p at random is not enough — p − 1 has to be checked.
Figure (svg): Pohlig-Hellman on p = 41: solve the discrete log modulo each prime power of p minus 1, then recombine by the CRT.
Worked example
The book's example. p = 41, α = 7 is a primitive root, and p − 1 = 40 = 2³ · 5.
Work modulo 8 first, one bit at a time, using the parity idea repeatedly
Why: Each step raises to a power (p−1)/2^k and reads off the next binary digit of x mod 8.
The three digits give x ≡ x₀ + 2x₁ + 4x₂ ≡ 1 + 4 ≡ 5 (mod 8)
Why: So x is 5 modulo the power-of-two part.
Now modulo 5: compute β^((p−1)/5) ≡ 12⁸ ≡ 18 and α^((p−1)/5) ≡ 7⁸ ≡ 37 (mod 41)
Why: Then find which power of 37 equals 18.
37⁰ ≡ 1, 37¹ ≡ 37, 37² ≡ 16, 37³ ≡ 18 — so x ≡ 3 (mod 5)
Why: Four trials, because q = 5 is small. This search is what becomes infeasible for a large q.
Combine by the CRT: x ≡ 5 (mod 8) and x ≡ 3 (mod 5) give x ≡ 13 (mod 40)
Why: Section 3.4, doing the gluing.
Verify: 7¹³ ≡ 12 (mod 41) ✓
Why: The answer checks. And note what made it easy: 40 = 2³ · 5 has no prime factor above 5, so every search was over at most five values. A prime p with p − 1 = 2q for a large prime q would have made the second stage hopeless.
Figure (svg): Pohlig-Hellman on p = 41: solve the discrete log modulo each prime power of p minus 1, then recombine by the CRT.
Concept
A general method that works whatever the factorization of p − 1, trading memory for time.
Let m = ⌈√(p−1)⌉ and write the unknown x as x = im + j with 0 ≤ i, j < m. Then:
\[ \beta \equiv \alpha^{im + j} \;\Longrightarrow\; \beta \alpha^{-im} \equiv \alpha^{j} \]
The cost is about √(p−1) time and √(p−1) storage. This is Chapter 6's meet-in-the-middle attack in a new setting: split the unknown into two halves, tabulate one, and search the other.
It also sets the security floor. Any group of order N can be attacked in √N steps by a generic method, so a 256-bit group gives about 128 bits of security — which is exactly the elliptic-curve sizing in Chapter 21.
Figure (svg): Baby step giant step: a table of small powers, then large strides, meeting somewhere in the middle.
Concept
The most powerful method for discrete logs mod p, and the reason discrete-log key sizes look like RSA's rather than like elliptic curves'.
Note the shape: collect relations, solve a linear system. That is exactly the quadratic sieve's structure from Section 9.4, and the two algorithms have essentially the same sub-exponential running time.
The consequence is practical. Because index calculus exists, discrete logs mod p need a modulus of a few thousand bits, just like RSA. Elliptic curve groups have no index calculus — the small primes have no analogue — so only the generic √N attacks apply, and 256 bits suffices.
Figure (svg): Index calculus: collect relations that factor over a small factor base, then solve a linear system for the logarithms.
Ranking
Each has a precondition, and knowing which is which decides how parameters must be chosen.
Put in order
Why: Pohlig-Hellman is the least general — it needs p − 1 to have only small prime factors, which correct parameter choice prevents. Index calculus needs a factor base, so it works in Z/p but not in an elliptic curve group. Baby step giant step is fully generic, working in any group at √N cost, which makes it the security floor. Trial multiplication is also fully generic and costs N, so it is the most general and the worst. Ordering by generality is the useful way to read attacks: the general ones set the parameter sizes and the special ones set the parameter conditions.
Figure (svg): Baby step giant step: a table of small powers, then large strides, meeting somewhere in the middle.
Cost model
The same protocol in two different groups needs wildly different key sizes, and this is why.
Annotate
On: \( \text{generic: } O\bigl(\sqrt{N}\bigr) \qquad \text{index calculus: } L_p\!\left[\tfrac{1}{3}, c\right] \)
Three attacks, three conditions on parameters: large enough order, no small-factor decomposition, and no factor base.
Prediction
The group has order about 2⁸⁰ and you have unlimited memory.
Predict first
Roughly how many group operations?
Correct: 2⁴⁰
The storage is also 2⁴⁰ entries, which is the real constraint — though Pollard's rho method achieves the same √N time with negligible memory, so the storage requirement is not a defence.
This is the single most useful number in parameter selection: security level ≈ half the group's bit length, for any group with no better attack. A 256-bit elliptic curve gives 128 bits, exactly.
Why: About √N = 2⁴⁰, which is roughly a trillion — feasible for a determined attacker, and the reason an 80-bit group is no longer considered adequate. The square root is a hard floor: it applies in every group, cannot be beaten by any method that ignores the group's structure, and it is why a group must have roughly twice as many bits as the security level you want.
Figure (svg): Baby step giant step: a table of small powers, then large strides, meeting somewhere in the middle.
Translation
Four ways to attack a discrete logarithm, four different cost shapes.
Match the pairs
Why: Read the third row carefully, because it is the one that catches people: Pohlig-Hellman's cost depends on the largest prime factor of the group order, not on the order itself. A 2048-bit group whose order factors into small primes offers the security of its largest factor and nothing more. And the fourth row applies only where a factor base exists, which is the whole difference between finite fields and elliptic curves.
Section
Section 10.3 · pp. 218-219
Concept
The book's story: Alice claims she can predict football results and wants to sell the method. Bob will not pay without proof; Alice will not predict this weekend's games because Bob would simply bet and not pay. Showing last week's predictions proves nothing.
The requirement is a bit b sent so that:
Physically: Alice locks the bit in a box and sends it; later she sends the key. The mathematical version must work without the two being in a room together.
The construction. Alice and Bob agree a large prime p ≡ 3 (mod 4) and a primitive root α. Alice picks a random x < p−1 whose second bit is b, and sends β ≡ αˣ. Later she sends x, and Bob reads b from x mod 4.
Figure (svg): Bit commitment: Alice sends a value that hides her bit and that she cannot later change.
Worked example
Two properties, two different reasons, and they come from different places.
Hiding: Bob cannot compute discrete logs, so he cannot find x from β
Why: And Section 10.2 showed he cannot even get x mod 4 without solving the problem — the first bit is free from the parity test, which is precisely why the construction uses the second.
Binding: Bob checks that β ≡ αˣ when Alice reveals x
Why: So Alice must supply an x consistent with what she already sent.
And β ≡ αˣ has a unique solution with x < p − 1, because α is a primitive root
Why: There is no second x she could substitute. She is bound by mathematics, not by good behaviour.
Verify: apply it to the football problem
Why: Alice commits to b = 1 for a home win and b = 0 for a loss, one commitment per game, before kickoff. Bob learns nothing in advance, so he cannot bet on it. After the games Alice reveals, and she cannot have changed her predictions. Both parties' objections are met at once.
Figure (svg): Bit commitment: Alice sends a value that hides her bit and that she cannot later change.
Socratic
The book gives a second construction using any one-way function, and it is the one used in practice.
Discussion prompt
Describe it, and say why the random padding on both sides of the bit is necessary.
Hint: Without padding, how many possible inputs are there?
Answer:
The construction: Alice takes a random 100-bit string, then the bit b, then another random 100-bit string, applies a one-way function to the 201-bit result, and sends the output. Later she reveals the whole string; Bob applies the function and compares.
Why padding is essential: without it there are exactly two possible inputs, 0 and 1, so Bob computes both hashes and reads the bit immediately. Hiding fails completely. The 100 random bits make the input space too large to search.
Why on both sides: padding only before the bit leaves the bit in a known position at a known end, and some functions leak more about their input's tail than its head. Padding both sides is cheap insurance against structure in the function.
Binding comes from collision resistance: to change b, Alice would need a second 201-bit string with the same hash, and Chapter 11 is about why that is infeasible.
And this is the version deployed everywhere — in coin-flipping protocols, in sealed-bid auctions, in the commit-reveal schemes of blockchain applications — because a hash is far cheaper than a modular exponentiation.
Real world
The football story is a toy. The primitive is deployed widely.
Discussion prompt
Name three real uses and say what each one needs hiding and binding for.
Hint: Auctions, randomness, and anything where two parties must act simultaneously without trusting each other.
Answer:
Sealed-bid auctions. Every bidder commits to a bid, then all bids are opened together. Hiding stops anyone bidding one unit above a rival; binding stops anyone changing their bid after seeing the others.
Fair coin flipping over a network. Alice commits to a bit, Bob announces his, then Alice opens hers and the XOR is the result. Neither can bias the outcome — which is exactly Chapter 18's subject, arriving eight chapters early.
Blockchain commit-reveal schemes. Voting and randomness on a public ledger use commit-then-reveal precisely because everything on the chain is visible immediately, so a plain submission would be front-run.
And a fourth worth knowing: zero-knowledge protocols. Chapter 19's Feige-Fiat-Shamir scheme has the prover commit before the verifier challenges, and the commitment is what stops the prover tailoring an answer to the challenge.
The common shape is simultaneity between mutually distrusting parties over a channel that cannot deliver simultaneously. Commitment manufactures it from a hard problem.
Faded example
Four lines, and a system.
Fill in the blanks
Bob's public key is (p, α, β) with β ≡ α^b. Alice picks a fresh random k, sends r ≡ α^k and t ≡ β^k m. Bob recovers the message by computing t · r^(−b), which works because the α^(bk) terms cancel.
Why: The structure is a one-time pad in the multiplicative group: β^k is a uniformly random mask, and r ≡ α^k is the information Bob needs to reconstruct that mask using his private b. The freshness of k is what makes the mask uniform — reuse it and the masks cancel between two ciphertexts, which is Section 4.3's failure in a new group.
Section
Section 10.4 · pp. 219-221
Concept
The problem is establishing a key for AES between two widely separated parties. RSA is one answer; this is another, and its security is tied directly to the discrete logarithm problem.
\[ K \equiv \alpha^{xy} \pmod p \]
They now share K without ever transmitting it. They need not use all of K — the book suggests taking the middle 56 bits for a DES key, and in practice a hash of K is used.
Figure (svg): Diffie-Hellman key exchange: Alice and Bob each send one public value and both compute the same shared secret.
Worked example
Eve has p, α, αˣ and αʸ. Track exactly what she would need.
If Eve can compute discrete logs, she is finished
Why: Take the discrete log of αˣ to get x, then raise αʸ to the power x to get α^(xy) ≡ K.
But she does not need x or y — she only needs K
Why: This is the crucial observation, and it defines a separate problem.
Computational Diffie-Hellman problem — Given p, a primitive root α, and the values αˣ and αʸ mod p, compute α^(xy) mod p. Solving the discrete log problem solves this — but nobody has proved the converse.
So breaking Diffie-Hellman requires solving CDH, which is at most as hard as the discrete log problem
Why: It might be strictly easier. Fifty years have not settled it, and the honest statement of Diffie-Hellman's security is 'CDH is believed hard', not 'discrete logs are hard'.
Verify: compare with RSA's situation
Why: There, breaking RSA requires taking e-th roots mod n, which might be easier than factoring, and nobody has closed that gap either. The pattern is general: a system's security rests on the problem an attacker actually has to solve, which is usually a weakening of the famous one.
Figure (svg): Three related problems: the discrete logarithm, computational Diffie-Hellman, and decision Diffie-Hellman.
Concept
Alice completes the protocol and holds a shared key. The protocol says nothing about whose key it is.
Eve intercepts. She runs one Diffie-Hellman exchange with Alice using her own exponent, and a second with Bob. Alice ends up sharing a key with Eve; Bob ends up sharing a different key with Eve; both believe they are talking to each other. Eve decrypts, reads, re-encrypts and forwards.
Neither end sees anything wrong, because nothing in the protocol binds αˣ to Alice's identity. This is the intruder-in-the-middle attack, and it is the opening subject of Chapter 15.
It is the padlock-substitution problem from Chapter 1, in its exact form. Public key cryptography does not authenticate public keys — it makes authenticating them urgent.
The fix is always the same: sign the exchanged values, or authenticate them against something pre-shared. TLS signs the Diffie-Hellman values with a certificate-backed key, which is why certificates exist.
Figure (svg): The intruder-in-the-middle attack on unauthenticated Diffie-Hellman: Eve runs two exchanges and relays between them.
Counterexample
Alice and Bob run the protocol exactly as specified over an open network.
Discussion prompt
Write out Eve's attack step by step, and say precisely which security property fails and which holds.
Hint: Eve does not need to solve any hard problem.
Answer:
The attack. Eve intercepts αˣ from Alice and sends Bob αᵉ instead. She intercepts αʸ from Bob and sends Alice αᵉ. Alice computes K₁ = α^(xe); Eve computes the same from αˣ and e. Bob computes K₂ = α^(ye); Eve computes that too. She now decrypts with one key and re-encrypts with the other.
What fails: authentication. Neither party has any evidence about who sent the value they received.
What holds: the discrete logarithm problem. Eve solved nothing. She was simply in a position to substitute, and the protocol has no mechanism to detect substitution.
Why this matters as a general lesson: confidentiality against a passive eavesdropper and security against an active one are different properties, and Diffie-Hellman as stated provides only the first. Chapter 1's four attack models were making exactly this distinction.
And it is not hypothetical. Any attacker on the network path — a hostile Wi-Fi access point, a compromised router, a state-level tap — is in position. It is the reason TLS spends most of its handshake on authentication rather than on key agreement.
Real world
TLS 1.3 removed RSA key transport entirely, keeping only ephemeral Diffie-Hellman for key agreement.
Discussion prompt
What property does that buy, and what attack does it defeat?
Hint: Think about an adversary who records traffic now and compromises the server later.
Answer:
With RSA key transport, the client encrypts a session key under the server's long-term public key. Anyone who records the session and later obtains that private key — by compromise, by subpoena, by the server being decommissioned carelessly — can decrypt the recorded traffic.
With ephemeral Diffie-Hellman, both sides generate fresh exponents per session and destroy them afterwards. The session key was never transmitted in any form, and the long-term key was used only to sign the exchange, not to protect it.
So a later key compromise reveals nothing about past sessions. This is forward secrecy, and it is a property RSA key transport cannot have, because the thing that protects the session key is exactly the long-term secret.
The long-term key is still needed — to sign the Diffie-Hellman values and defeat the intruder-in-the-middle attack. Its role changes from confidentiality to authentication, which is precisely the split this chapter's two sections make.
Practical scale: this is why bulk recording of encrypted traffic became far less valuable in the 2010s. The recording is still possible; decrypting it later is not.
Section
Section 10.5 · pp. 221-223
Concept
Published by ElGamal in 1985. Bob chooses a large prime p, a primitive root α, and a secret b, and computes β ≡ α^b. His public key is the triple (p, α, β).
\[ t \, r^{-b} \equiv \beta^{k} m (\alpha^{k})^{-b} \equiv \alpha^{bk} m \, \alpha^{-bk} \equiv m \pmod p \]
Note the ciphertext is a pair, so it is twice as long as the message. That is the price of randomisation, and it is a price RSA pays too once OAEP padding is added.
Figure (svg): ElGamal encryption: a fresh random k per message produces a two-part ciphertext (r, t).
Worked example
The argument is short and worth following, because it is stronger than anything available for textbook RSA.
k is chosen at random, so β^k is a random non-zero element mod p
Why: β is fixed; raising it to a random exponent lands anywhere in the group.
So t ≡ β^k m is the message multiplied by a random element
Why: Multiplication by a uniform random group element is exactly a one-time pad, in the multiplicative group rather than in bits.
Therefore t alone is uniformly random and gives Eve no information about m
Why: Provided m ≠ 0, which must be avoided — as with any masking scheme, zero is a fixed point.
And r ≡ α^k does not help her: recovering k from r is a discrete logarithm problem
Why: Though if Eve does find k she computes tβ^(−k) = m immediately, so k is as sensitive as b.
Verify: compare with textbook RSA
Why: There, encryption is deterministic and public, so Alice's own ciphertext can be recomputed by anyone testing a guess — the indistinguishability game of Section 4.5 is lost outright. ElGamal is randomised by construction and wins that game without needing a padding scheme bolted on.
Figure (svg): ElGamal encryption: a fresh random k per message produces a two-part ciphertext (r, t).
Anomaly
Alice encrypts m₁ and m₂ to Bob and, through carelessness, uses the same k for both.
Predict first
What does Eve learn?
Correct: The ratio m₁/m₂, and therefore m₂ if she ever learns m₁
The book states the requirement explicitly: a different random k must be used for each message. It is a hypothesis of the security argument, not advice.
This is the two-time pad of Chapter 4 arriving for the third time — after Venona and after WEP. The pattern is now unmistakable: any scheme that masks a message with a value derived from a nonce collapses when the nonce repeats.
And it is far worse in the signature version. Reusing k in the ElGamal or DSA signature scheme of Chapter 13 lets an attacker solve directly for the private key, not merely for a message ratio. That is exactly how the PlayStation 3's signing key was recovered in 2010.
Why: With the same k, both ciphertexts share the same r, which is immediately visible. And t₁/t₂ ≡ (β^k m₁)/(β^k m₂) ≡ m₁/m₂ — the random mask cancels exactly as the keystream did in Section 4.3. If Eve ever learns one plaintext, she gets the other for free; and even without that, the ratio of two messages is often enough.
Comparison
Two public key systems, two hard problems. Fill the blanks.
Comparison matrix
| RSA | ElGamal | |
|---|---|---|
| Hard problem | factoring n | discrete logarithms mod p |
| Ciphertext length | same as the modulus | twice the modulus — a pair (r, t) |
| Randomised? | no, unless padding is added | yes, by construction |
| Security reduces to | nothing named — the RSA problem is its own assumption | the Computational Diffie-Hellman problem |
| Encryption cost | one exponentiation, cheap with e = 65537 | two exponentiations |
Row four is the one that matters to a theorist and row two to an engineer, which is a fair summary of why both systems survive.
Concept
The chapter has quietly introduced three distinct problems, and confusing them makes security claims meaningless.
ElGamal's security reduces to CDH, which is a genuine theorem and more than RSA has. But the reduction is to CDH, not to DL — so the honest claim is 'as hard as CDH', and CDH might be strictly easier than the discrete log.
Figure (svg): Three related problems: the discrete logarithm, computational Diffie-Hellman, and decision Diffie-Hellman.
Two truths and a lie
Two of these overstate what has been proved.
Eliminate the wrong options
Which statement is correct?
Survives elimination: a
Why: The precise statement is that breaking Diffie-Hellman means solving CDH, and that solving DL would solve CDH — so DL hardness is sufficient for security and has never been shown necessary. Option (c) is the more dangerous error in practice, because it confuses a passive-eavesdropper guarantee with security against an active attacker, and every real network has active attackers in it.
Socratic
The protocol produces α^(xy) mod p, a number of a few thousand bits. AES needs 256.
Discussion prompt
Why not simply take 256 bits of it, and what is done instead?
Hint: Ask whether the bits of a random group element are uniformly distributed.
Answer:
α^(xy) is not a uniform bit string. It is a uniform element of a group, which is a different thing: the group has p−1 elements and the bit string has 2^k values, so the map from one to the other is biased. The high bits are skewed by wherever p falls relative to a power of two.
And it has structure Eve can test. The parity leak of Section 10.2 means the quadratic residuosity of α^(xy) is computable from the public values, so at least one bit is not secret at all.
In a subgroup the problem is sharper still: if the shared value lies in a subgroup of order q, it takes only q of the p−1 possible values, and the pattern of which ones is public.
The fix is a key derivation function — hash the shared value, usually with a salt and a context string, using HKDF or similar. A hash of a high-entropy but non-uniform input is a uniform output, which is exactly what Chapter 11's functions are for.
The general principle: a shared secret and a key are different objects. One has entropy, the other must have entropy uniformly distributed across its bits, and turning the first into the second is a step, not an afterthought.
Cost model
Understanding where the time goes explains most of the engineering decisions around it.
Annotate
On: \( \text{2 exponentiations per party, each } \approx 1.5\log_2(x) \text{ modular multiplications} \)
The last note is the one with teeth: an optimisation that looks purely about performance changed the attacker's cost model entirely.
Explain it
A colleague asks why TLS bothers with Diffie-Hellman when RSA could just deliver the session key.
Discussion prompt
Explain the difference in terms of what an adversary who records traffic today can do in five years.
Hint: Ask what has to be stolen, and when.
Answer:
With RSA delivery: the session key travels, encrypted under the server's long-term key. An adversary records the encrypted session now and stores it. Five years later she obtains the server's private key — by a breach, a court order, or the hardware being sold on — and decrypts everything she recorded.
With ephemeral Diffie-Hellman: the session key never travels at all. Both ends compute it from exponents they generate for that session and discard afterwards. The long-term key only signs, so stealing it lets the adversary impersonate the server from then on and reveals nothing about the past.
The one-sentence version: with RSA delivery, one theft decrypts everything ever recorded; with ephemeral Diffie-Hellman, one theft decrypts nothing that already happened.
Why it changed practice: bulk recording of encrypted traffic only pays off if it can be decrypted later. Forward secrecy removes the payoff, which is why TLS 1.3 made it mandatory rather than optional.
The caveat to state honestly: it protects past sessions, not future ones, and only if the ephemeral exponents really are discarded. An implementation that caches them for performance throws the property away silently.
Constraint
You are specifying key agreement for a product with a fifteen-year field life, running on modest embedded hardware.
Discussion prompt
Specify the group, the exponent length and the surrounding protocol, with a reason for each.
Hint: Every constraint in the question points to the same family of groups.
Answer:
Use an elliptic curve group — Curve25519 or P-256. The embedded hardware makes 3072-bit modular exponentiation expensive, and the curve gives the same 128-bit security at 256 bits, several times faster and with far smaller messages.
Ephemeral exponents per session, discarded afterwards, for forward secrecy over a fifteen-year horizon during which any long-term key will very likely leak.
Exponents of 256 bits, matching the group — longer buys nothing, since the generic attack costs √N regardless.
Authenticate the exchange, by signing the ephemeral values with a device key provisioned at manufacture. Without this the intruder-in-the-middle attack applies and everything else is irrelevant.
Derive the key with a KDF, not by truncating the shared point — the x-coordinate is not a uniform bit string.
And provide algorithm agility: a version field and a negotiation mechanism, so the curve can be replaced within the field life. Chapter 7's DES lesson — build in the ability to change the parameter, because fifteen years is long enough that you will need to.
Elimination
A product uses an unauthenticated Diffie-Hellman exchange over the public internet.
Eliminate the wrong options
Which change addresses the actual weakness?
Survives elimination: c
Why: The deployment's weakness is that nothing binds the exchanged values to an identity, so an attacker on the path completes two exchanges and relays. Only authentication fixes that, and it has to come from outside the protocol — a signature verified against something the peer already trusts. Note that all four options are good practice and three of them defend against attacks that require Eve to do mathematics; the one that matters defends against her doing none.
Fill the middle
Fill the two blanks and the whole protocol is on the page.
Fill in the blanks
\text(α^y)^x \alpha^(α^x)^y; \; \text___ \alpha^___; \; \text___ ___; \; \text___ ___; \; \text___ \alpha^___
Why: Both sides raise the value they received to their own secret exponent, and commutativity of exponentiation does the rest: (α^y)^x = α^(yx) = α^(xy) = (α^x)^y. Nothing secret crosses the channel, and the shared key is never transmitted in any form — which is what gives the protocol forward secrecy when the exponents are fresh per session.
Explain it to yourself
Choosing a large random prime p is not sufficient for the discrete log problem to be hard.
Discussion prompt
Explain what goes wrong when p − 1 is smooth, and describe the standard fix.
Hint: Pohlig-Hellman plus the Chinese Remainder Theorem.
Answer:
Pohlig-Hellman decomposes the problem. It solves for x modulo each prime power dividing p − 1, at a cost driven by the size of the largest prime factor, then glues the answers with the CRT. If every prime factor of p − 1 is small, every sub-problem is a short search and x falls out.
So the effective difficulty is set by the largest prime factor of p − 1, not by the size of p. A 2048-bit p whose predecessor factors into primes under a million offers no security at all.
The standard fix is a safe prime: choose p = 2q + 1 with q also prime. Then p − 1 = 2q, the largest prime factor is q ≈ p/2, and Pohlig-Hellman gains nothing beyond one bit — the parity, which was free anyway.
A common alternative is to work in a prime-order subgroup: pick a large prime q dividing p − 1 and use a generator of the order-q subgroup rather than a primitive root. Then all exponent arithmetic happens mod q and Pohlig-Hellman has nothing to decompose.
The subgroup version has a bonus the book flags: with β chosen as a power of α^t, the discrete log is automatically 0 mod t, so the algorithm's partial information is information about nothing. That idea is used in the Digital Signature Algorithm of Chapter 13.
Estimation
Index calculus attacks discrete logs mod p sub-exponentially, at essentially the same cost as the number field sieve attacks factoring.
Predict first
What size of p is recommended for 128-bit security?
Correct: 3072 bits
The comparison is worth memorising because it comes up constantly: 128-bit symmetric ≈ 3072-bit RSA ≈ 3072-bit finite-field Diffie-Hellman ≈ 256-bit elliptic curve.
It also explains a real-world attack. Logjam (2015) exploited servers using 512-bit Diffie-Hellman groups for export compatibility, and — worse — the fact that many servers shared the same 1024-bit group. Index calculus front-loads most of its work per group rather than per key, so precomputing on one widely-shared group breaks every connection using it.
Why: About 3072 bits, the same as RSA — and for the same reason, since index calculus and the number field sieve have the same sub-exponential shape. This is why elliptic curves are so attractive: their groups admit no index calculus, so only the generic √N attack applies and 256 bits suffices for the same security.
Figure (svg): Three related problems: the discrete logarithm, computational Diffie-Hellman, and decision Diffie-Hellman.
Error analysis
From a VPN product's configuration documentation.
Annotate
The fourth point is the fatal one: without authentication the other three do not matter, because Eve never needs to attack the mathematics.
Discrimination
Some of this chapter's problems are mathematical and some are protocol-level.
Sort into buckets
Sort each one.
Matching
This chapter has built three different things from one hard problem.
Match the pairs
Why: One hard problem, four services — and the fourth is the one Diffie-Hellman lacks and must borrow. The general lesson is that a cryptographic primitive provides exactly one property, and a real protocol is an assembly of several. Almost every failure in Chapter 14 is an assembly that left one of these out.
Edge cases
The scheme's two properties come from two different sources, so they fail in different circumstances.
Discussion prompt
Under what conditions does hiding fail, and under what conditions does binding fail?
Hint: Hiding rests on a computational assumption; binding rests on uniqueness.
Answer:
Hiding fails if Bob can compute discrete logs — a computational assumption, so it is only as good as the parameters. Too small a p, or a smooth p − 1, and Bob simply solves for x and reads b.
Hiding also fails if the bit is in the wrong place. The parity of x is free from Section 10.2's test, so committing to b as the first bit of x would be no commitment at all. The book uses the second bit precisely for this reason.
Binding does not rest on a computational assumption at all — it holds because β ≡ αˣ has a unique solution below p−1. Alice could have unlimited computing power and still not find a second x. This is an unconditional guarantee, like the one-time pad's.
The asymmetry is worth naming: this scheme is computationally hiding and perfectly binding. The hash-based version is the opposite — computationally binding, because a collision would break it, and statistically hiding because of the random padding.
And you cannot have both perfectly. If commitments are perfectly hiding then several openings must exist, so a computationally unbounded Alice could find one. It is a genuine trade, and knowing which side a scheme falls on tells you which party you are trusting.
Commit first
You are auditing a Diffie-Hellman deployment and have time for one thorough check.
Predict first
What do you examine?
Correct: Whether the exchange is authenticated
The runner-up is the random exponents, for Chapter 5's reason: a predictable x is as fatal as a published one, and it has broken deployed systems repeatedly.
Then p − 1's factorization, then p's size — in that order, because a smooth p − 1 destroys security at any size while a merely small p is a graded rather than total failure.
Why: Without authentication, the intruder-in-the-middle attack breaks the deployment completely and requires no mathematics, no precomputation and no special parameters — only a position on the network path. Every other item on the list matters, and every one is irrelevant if Eve is simply substituting values. Check that the mathematics is being used before checking whether it is being used well.
Scale up
Three attacks, three cost curves. This table is how every key-size recommendation is produced.
Step through it
Why is the fourth row twelve times smaller than the third for the same security?
Because only the generic column applies to it. Index calculus needs a factor base, and an elliptic curve group has no notion of a small element to build one from. That single absence is worth a factor of twelve in key size and is the entire commercial argument for Chapter 21.
Missing information
The phrase appears in security questionnaires constantly.
Discussion prompt
List what remains undetermined, in the order you would ask about it.
Hint: This chapter has given you exactly five questions.
Answer:
Is it authenticated, and by what? Without this the protocol provides nothing against an active attacker, and everything else is moot.
Are the exponents ephemeral? Reused exponents throw away forward secrecy, which is the main reason to choose Diffie-Hellman over RSA key transport in the first place.
Which group, and is it shared? A 1024-bit group is within reach, and a group shared across many deployments amortises the attacker's precomputation — the Logjam lesson.
Does the group order have a large prime factor? Otherwise Pohlig-Hellman decomposes the problem regardless of size.
Is the shared value passed through a key derivation function? α^(xy) is a group element with exploitable structure, not a uniform key.
Five questions, and the phrase answers none of them. As with 'we use AES' and 'we use RSA-4096', naming the primitive is the smallest part of specifying the system.
Picture it
Before the recap, fix the hierarchy. Solving a problem higher in this list solves every problem below it, and none of the reverse implications has ever been proved.
Figure (svg): Three related problems: the discrete logarithm, computational Diffie-Hellman, and decision Diffie-Hellman.
So a system reducing to the middle row has a weaker security guarantee than one reducing to the top — and ElGamal reduces to the middle row. Saying that precisely is the difference between a security claim and a slogan.
Pattern
Chapter 9 gave the shape of a public key system. Chapter 10 fills it in differently and adds three things.
The chapter also supplies the sharpest single fact about key sizes in the book: whether index calculus applies to your group decides whether you need 3072 bits or 256. Everything about elliptic curves in Chapter 21 follows from that one absence.
Figure (svg): Three related problems: the discrete logarithm, computational Diffie-Hellman, and decision Diffie-Hellman.
Trap
The trap. Alice and Bob run Diffie-Hellman over an open network and end up with a shared secret that was never transmitted. Eve saw every message and cannot compute the key. So the channel is now secure and they can start sending AES-encrypted traffic.
This is the natural reading of the protocol, and every step of it is true.
Why it fails. Alice has a shared key with somebody. The protocol contains nothing that identifies who. Eve runs one exchange with Alice and another with Bob, holds both keys, and relays traffic between them — decrypting, reading, and re-encrypting as she goes.
She solved no hard problem. She needed only to be on the path, which is exactly where a hostile access point, a compromised router or a network tap already is.
The precise statement is that Diffie-Hellman is secure against a passive eavesdropper and provides nothing against an active attacker. Chapter 1's attack models were drawing this line, and it is the difference between the two halves of that list.
The fix is authentication, and it must come from outside the protocol — a signature over the exchanged values, verified against a certificate or a pre-shared key. That is why TLS spends most of its handshake on identity rather than on key agreement, and why Chapter 15 exists as a separate chapter.
The habit to build: after establishing what a primitive guarantees, ask against which adversary. 'Secure' with no adversary named is not a claim.
Check
Work it out before you click.
Check your understanding
You know β ≡ αˣ (mod p) with α a primitive root, and you compute β^((p−1)/2) ≡ −1. What do you know about x?
Answer: B
Why: β^((p−1)/2) ≡ α^(x(p−1)/2) ≡ (−1)ˣ, since α^((p−1)/2) ≡ −1 for a primitive root. A result of −1 therefore means x is odd. This is one bit of the secret exponent recovered by a single exponentiation, and it is why Section 10.3 commits to the second bit of x rather than the first.
Check
Apply the design rule from Section 10.2.
Check your understanding
Which prime p is a poor choice for Diffie-Hellman, regardless of its size?
Answer: B
Why: Pohlig-Hellman solves the discrete log modulo each prime power of p − 1 and glues the answers by the CRT, at a cost set by the largest prime factor. If every factor is small the problem decomposes into short searches and falls, whatever the size of p. This is why safe primes p = 2q + 1 are used.
Check
Read the encryption steps carefully.
Check your understanding
Alice encrypts two messages to Bob with ElGamal, using the same random k for both. What has she leaked?
Answer: C
Why: t₁/t₂ ≡ (β^k m₁)/(β^k m₂) ≡ m₁/m₂, and the reused r makes the reuse obvious in the first place. This is Section 4.3's two-time pad in the multiplicative group: the mask is the same, so dividing cancels it. Knowing one plaintext then yields the other.
Connect it up
Three constructions and three attacks, and they fit together tightly.
Draw it
Write the discrete log problem in one line, then the three attacks — Pohlig-Hellman, baby step giant step, index calculus — with the precondition and cost of each. Beside them write the parameter rule each one forces. Then write out Diffie-Hellman in four lines and mark the attack it does not resist, and ElGamal in four lines with the requirement on k. Finish with the three problems DL, CDH and DDH and which one ElGamal's security actually reduces to.
The last item is the one worth being precise about: 'as hard as the discrete log problem' is what people say, and 'as hard as CDH' is what has been proved.
Exit ticket
One question, about the property that changed how the internet encrypts traffic.
Predict first
What does ephemeral Diffie-Hellman give that RSA key transport cannot?
Correct: Forward secrecy — the session key is never transmitted, so a later compromise of the long-term key does not decrypt recorded sessions
Why: With RSA key transport the session key is encrypted under the server's long-term public key, so anyone who records the traffic and later obtains that private key can decrypt everything. With ephemeral Diffie-Hellman the session key is computed independently at both ends from exponents that are destroyed afterwards, and the long-term key only signs. This is why TLS 1.3 removed RSA key transport entirely. Note that option 3 is precisely what Diffie-Hellman does not provide — authentication must be added from outside.
Recap
A second hard problem, and three constructions built on it.
Chapter 11 next. A different kind of primitive entirely: a hash function compresses any input to a fixed-length digest, and it is the tool behind integrity, signatures, password storage and commitments. It is not encryption at all — nothing about it is reversible, and that is the point.
Figure (svg): Three related problems: the discrete logarithm, computational Diffie-Hellman, and decision Diffie-Hellman.
Want this taught 1-on-1? Alexander tutors Cryptography — $55/session, free consultation.