Chapter 21: Elliptic Curves

Chapter 21 of Trappe & Washington: the chord-and-tangent group law with worked chord and tangent computations, the point at infinity and negation, curves modulo p and Hasse's theorem, the elliptic curve discrete logarithm problem and why index calculus has no analogue, Koblitz encoding of messages as points, Lenstra's factorisation method and its unification with the p-1 method and trial division on singular curves, curves in characteristic 2 over GF(2^n), and the translation table that rebuilds ElGamal, Diffie-Hellman and ElGamal signatures on a curve — with every numeric example from the book verified.

Subject: Cryptography · 79 slides · diagram-first lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Elliptic Curves

Title

Cryptography · Chapter 21

A group law drawn with a straightedge, and the cryptosystems it rebuilds at a fraction of the key size

2. What you will be able to do

Objectives

Miller and Koblitz proposed elliptic curves for cryptography in the mid-1980s, and Lenstra used them to factor integers. Both applications come from the same construction: a way to add two points on a cubic curve and get a third.

Figure (svg): An elliptic curve over the reals with a chord through two points, its third intersection, and the reflection that defines the sum.

The whole group law in one picture: a line meets a cubic three times, so any two points name a third.

3. Why would anyone want a different group?

Warm-up

Diffie-Hellman, ElGamal and DSA are all built on one thing: a finite group in which exponentiation is easy and taking logarithms is hard. The integers mod p supply one.

Discussion prompt

What would be gained by finding a different group with the same property?

Hint: Think about what limits the key size in Chapter 10.

Answer:

The security of mod-p discrete logs is limited by index calculus, from Section 10.2. That attack is subexponential, so the prime has to be large — 2048 bits or more — to stay ahead of it.

If a group had no index calculus attack, the best known algorithm would be generic — baby-step giant-step or Pollard rho, both taking about √n operations. Then n only needs to be about 2²⁵⁶ for 128-bit security, and the numbers being manipulated are 256 bits rather than 3072.

That is a factor of ten or more in key size, and a corresponding saving in bandwidth, storage and hardware. Blake and coauthors estimate that a 4096-bit conventional system is matched by a 313-bit elliptic curve system.

So the question is where to find such a group. The answer is the points on a cubic curve, with an addition law that has been studied since the nineteenth century for entirely unrelated reasons.

And a warning worth stating at the outset: 'no known attack' is not 'no attack'. The whole advantage rests on the absence of an algorithm, not on a proof that none exists.

4. The Addition Law

Section

Section 21.1 · pp. 393-401

5. What an elliptic curve is

Concept

An elliptic curve E is the graph of an equation

\[ E: \; y^2 = x^3 + a x^2 + b x + c \]

with a, b, c drawn from whatever field is appropriate — the rationals, the reals, or the integers mod a prime p. Together with a point at infinity, written ∞.

Over the reals the graph has two shapes, depending on the cubic. Three real roots gives two components, as with y² = x(x+1)(x−1); one real root gives a single connected curve, as with y² = x³ + 73.

Technical point: the cubic must have no repeated roots. Curves like y² = (x−1)²(x+2) are excluded, and Section 21.3.1 shows what happens if you use one anyway.

Historical point: elliptic curves are not ellipses. The name comes from elliptic integrals such as ∫dx/√(x³+bx+c), which arise when computing the arc length of an ellipse.

Figure (svg): An elliptic curve over the reals with a chord through two points, its third intersection, and the reflection that defines the sum.

The whole group law in one picture: a line meets a cubic three times, so any two points name a third.

6. Elliptic Curve Addition Law: chord and tangent

Concept

The reason elliptic curves matter is that any two points produce a third.

  1. Draw the line L through P₁ and P₂ — or the tangent at P₁ if the two points coincide
  2. A line meets a cubic in three points, so L meets E in a third point Q
  3. Reflect Q through the x-axis (change y to −y) to get P₃
  4. Define P₁ + P₂ = P₃

This is not addition of points in the plane. The coordinates of P₃ have nothing to do with the coordinates of P₁ plus those of P₂; the name is borrowed because the operation turns out to satisfy the same axioms.

An equivalent statement worth remembering: P + Q + R = ∞ exactly when P, Q and R are collinear. Every fact about the group law can be read off from that one sentence.

Figure (svg): An elliptic curve over the reals with a chord through two points, its third intersection, and the reflection that defines the sum.

The whole group law in one picture: a line meets a cubic three times, so any two points name a third.

7. Adding two points on y² = x³ + 73

Worked example

The chord case, with numbers.

Take P₁ = (2, 9) and P₂ = (3, 10)

Why: Both on the curve: 2³ + 73 = 81 = 9², and 3³ + 73 = 100 = 10².

The line through them is y = x + 7

Why: Slope (10−9)/(3−2) = 1, through (2, 9).

Substitute into the curve: (x+7)² = x³ + 73, so x³ − x² − 14x + 24 = 0

Why: A cubic whose three roots are the x-coordinates of the three intersections.

Two roots are already known — x = 2 and x = 3 — and the roots sum to 1

Why: Minus the coefficient of x². So 2 + 3 + x = 1 and the third root is x = −4.

From the line, y = −4 + 7 = 3, so Q = (−4, 3)

Why: Check: (−4)³ + 73 = 9 = 3². ✓

Verify: reflect to get (2, 9) + (3, 10) = (−4, −3)

Why: Notice the trick that made this easy: knowing two roots of a cubic and the sum of all three gives the third without any factoring. Every addition formula in the chapter is that observation written out.

Figure (svg): An elliptic curve over the reals with a chord through two points, its third intersection, and the reflection that defines the sum.

The whole group law in one picture: a line meets a cubic three times, so any two points name a third.

8. Doubling a point: the tangent case

Worked example

Now add P₃ = (−4, −3) to itself.

Differentiate the curve implicitly: 2y dy = 3x² dx

Why: So dy/dx = 3x²/(2y), which at (−4, −3) is 48/(−6) = −8.

The tangent line is y = −8(x + 4) − 3

Why: Slope −8 through P₃.

Substituting gives x³ − 64x² + ⋯ = 0, so the three roots sum to 64

Why: Minus the coefficient of x², as before.

The line is tangent, so x = −4 is a double root

Why: That is the algebraic meaning of tangency, and it is why the same counting argument still applies.

Hence (−4) + (−4) + x = 64, giving x = 72, and y = −8(76) − 3 = −611

Why: From the line equation.

Verify: reflect: 2P₃ = (72, 611)

Why: Check: 72³ + 73 = 373321 = 611². ✓ And note how fast the coordinates grow — two additions took a point with single-digit coordinates to six digits, which is why curves over finite fields are the ones used in practice.

Figure (svg): Doubling a point: the tangent line at P meets the curve once more, and the reflection of that point is 2P.

Doubling is the same construction with the chord degenerating to a tangent — one formula, two cases.

9. Infinity, negation and subtraction

Concept

Three conventions complete the group, and each is forced rather than chosen.

Lines through ∞ are vertical. So the line through P = (x, y) and ∞ meets E again at (x, −y); reflecting gives back P. Hence P + ∞ = P, and ∞ is the identity.

The line through (x, y) and (x, −y) is vertical, so its third intersection is ∞, and reflecting ∞ gives ∞ — which is what it means to say ∞ sits at both the top and the bottom of the y-axis. Hence

\[ (x, y) + (x, -y) = \infty \qquad \text{so} \qquad -(x, y) = (x, -y) \]

And subtraction is defined the obvious way: P − Q means P + (−Q), where −Q is Q reflected in the x-axis. Negation costs one sign flip, which will matter in Section 21.5 — on a curve, subtraction is as cheap as addition, while mod p division is far more expensive than multiplication.

Figure (svg): The point at infinity sitting at the top and bottom of the y-axis, making vertical lines meet the curve three times.

The identity element is not on the graph, and adding it is what makes the construction a group rather than a curiosity.

10. The formulas

Concept

For computation the geometry can be dropped. Let E be y² = x³ + bx + c with P₁ = (x₁, y₁) and P₂ = (x₂, y₂). Then P₁ + P₂ = (x₃, y₃) where

\[ x_3 = m^2 - x_1 - x_2, \qquad y_3 = m(x_1 - x_3) - y_1 \]

\[ m = \begin{cases} (y_2 - y_1)/(x_2 - x_1) & \text{if } P_1 \ne P_2 \\ (3x_1^2 + b)/(2y_1) & \text{if } P_1 = P_2 \end{cases} \]

If the slope is infinite, P₃ = ∞ — the vertical-line case. And ∞ + P = P for every P.

Note the shape of x₃ = m² − x₁ − x₂. It is the sum-of-roots identity from the worked examples, rearranged: the three roots sum to m², two of them are known, so the third is what remains.

Figure (svg): Doubling a point: the tangent line at P meets the curve once more, and the reflection of that point is 2P.

Doubling is the same construction with the chord degenerating to a tangent — one formula, two cases.

11. The points form a group

Concept

Two facts, one easy and one not, make the construction useful.

Commutative: P + Q = Q + P. Obvious, since the line through two points does not depend on their order.

Associative: (P + Q) + R = P + (Q + R). This is not obvious — it can be proved by a long computation with the formulas, or elegantly with projective geometry — and it is what makes the set of points an abelian group, with ∞ as the identity.

So multiples are well defined: kP means P added to itself k times, and the grouping does not matter. Negative multiples work too, with (−3)P = 3(−P).

And that is the whole cryptographic content of the chapter. Once you have an abelian group in which the operation is cheap and inverting a multiple is hard, every discrete-log protocol in the course can be rebuilt inside it without a new idea.

Figure (svg): The point at infinity sitting at the top and bottom of the y-axis, making vertical lines meet the curve three times.

The identity element is not on the graph, and adding it is what makes the construction a group rather than a curiosity.

12. Computing 100P efficiently

Worked example

The additive analogue of successive squaring from Chapter 3.

2P = P + P, then 4P = 2P + 2P, 8P = 4P + 4P, and so on

Why: Each doubling costs one addition, so 64P is reached in six.

Write 100 in binary: 1100100₂ = 64 + 32 + 4

Why: The set bits name which doubled values to combine.

100P = 64P + 32P + 4P

Why: Two more additions on top of the six doublings — eight in total.

Compare the naive route: 99 additions

Why: And even 4P computed as ((P+P)+P)+P takes three additions where 2P + 2P takes two.

Verify: the cost is logarithmic in k, not linear

Why: About 2 log₂ k operations. Which is exactly the modular exponentiation cost from Chapter 3, and it is why the translation table in Section 21.5 works: the expensive operation on each side has the same complexity profile.

13. Reading the addition formulas

Notation

Two lines of algebra encode the whole geometric construction.

Annotate

On: \( x_3 = m^2 - x_1 - x_2, \qquad y_3 = m(x_1 - x_3) - y_1 \)

  • Substituting the line into the cubic gives x³ − m²x² + ⋯ = 0, whose roots sum to m². Two roots are x₁ and x₂, so the third is m² − x₁ − x₂.
  • Follow the line from (x₁, y₁) to the third point and then negate: the line gives y₁ + m(x₃ − x₁), and the reflection flips the sign, giving m(x₁ − x₃) − y₁.
  • A secant slope when the points differ, and the implicit derivative when they coincide. The tangent case is the limit of the chord case, so it is one formula with a removable singularity rather than two rules.
  • A quotient a/b is a·b⁻¹ with b⁻¹ found by the extended Euclidean algorithm, which requires gcd(b, p) = 1. When the modulus is composite, that requirement can fail — and Section 21.3 turns the failure into a factoring algorithm.
  • One inversion plus a few multiplications per addition. The inversion dominates, which is why implementations use projective coordinates to defer it.

The last two notes are the chapter in miniature: the same failed inversion is a bug when factoring is not the goal and the whole method when it is.

14. Why reflect the third point?

Socratic

The construction takes the third intersection and flips its sign before calling it the sum.

Discussion prompt

Why not simply define P + Q to be the third intersection?

Hint: Check the group axioms.

Answer:

Without the reflection there is no identity element. You would need a point O with the property that the line through P and O meets the curve again at P, and no such point exists.

And associativity fails. The unreflected operation is commutative but does not group properly, so it is not the operation of a group at all.

With the reflection everything works, because the rule becomes 'three collinear points sum to ∞'. That statement is symmetric in all three points, which is exactly the symmetry a group law needs.

A useful way to see it: the reflection turns the geometric relation P + Q + R = ∞ into the algebraic relation P + Q = −R. The curve knows about triples of collinear points; the group law is what you get by choosing ∞ as a base point and rewriting.

And a different choice of base point gives a different, isomorphic, group law — which is why the theory is really about the curve, not about the coordinates.

Figure (svg): The point at infinity sitting at the top and bottom of the y-axis, making vertical lines meet the curve three times.

The identity element is not on the graph, and adding it is what makes the construction a group rather than a curiosity.

15. What if the two points share an x-coordinate?

Anomaly

P = (x, y) and Q = (x, −y) with y ≠ 0.

Predict first

What is P + Q?

  • The point (x, 0)
  • ∞ — the line is vertical, so its third intersection is the point at infinity
  • Undefined, since the slope formula divides by zero
  • P itself

Correct: ∞ — the line is vertical, so its third intersection is the point at infinity

In an implementation this is a special case that must be tested for, and forgetting it is a classic source of bugs — the code divides by zero, or worse, computes a modular inverse of 0 and produces nonsense.

There is a second special case: P = Q with y = 0. Then the tangent is vertical, so 2P = ∞ and P is a point of order 2. On y² = x³ + 4x + 4 mod 5 the point (2, 0) is exactly this.

And in the factoring algorithm these cases are not errors but results. A denominator that is zero modulo one prime factor and nonzero modulo another is what the gcd detects.

Why: The slope formula has x₂ − x₁ = 0 in the denominator, which is exactly the signal that the line is vertical. Vertical lines pass through ∞, so the third intersection is ∞ and reflecting gives ∞ back. So Q = −P.

16. Elliptic Curves Mod p

Section

Section 21.2 · pp. 401-409

17. Curves over a finite field

Concept

The same definitions with the coefficients and coordinates taken mod p. Elliptic curves mod p are finite sets of points, and these are the ones useful in cryptography.

Listing them is a matter of substitution. For each x in 0, 1, …, p−1, compute x³ + bx + c and ask whether it is a square mod p. If it is a nonzero square there are two values of y; if it is zero there is one; if it is a nonsquare there are none.

Arithmetic works exactly as before, with one change: a quotient a/b means a·b⁻¹ where b⁻¹b ≡ 1 (mod p), which requires gcd(b, p) = 1. Over a prime modulus that is automatic for b ≢ 0.

Over a composite modulus n it is not automatic, and the situations where it fails are precisely the key to factoring — Section 21.3. For now, when working mod a composite, pretend it is prime: if something goes wrong you usually learn the factorisation.

Figure (svg): The points of an elliptic curve modulo five, found by substituting each x and testing whether the result is a square.

A curve mod p is a finite set of points, and that finiteness is what makes it useful for cryptography.

18. Listing the points mod 5

Worked example

Take E: y² ≡ x³ + 2x − 1 (mod 5). Substitute each x in turn.

xx³ + 2x − 1reducedy
0−142, 3
122none — 2 is not a square mod 5
21111, 4
3322none
47111, 4

The squares mod 5 are 0, 1 and 4

Why: Since 1² = 1, 2² = 4, 3² = 4, 4² = 1. So 2 and 3 are nonsquares, and two of the five x-values give nothing.

Verify: the curve has seven points: (0,2), (0,3), (2,1), (2,4), (4,1), (4,4) and ∞

Why: Three of the five x-values gave two points each, which matches the heuristic that x³ + bx + c is a square about half the time. Hence roughly p points, plus ∞ — about p + 1 in total.

Figure (svg): The points of an elliptic curve modulo five, found by substituting each x and testing whether the result is a square.

A curve mod p is a finite set of points, and that finiteness is what makes it useful for cryptography.

19. Adding points mod 5

Worked example

On E: y² ≡ x³ + 4x + 4 (mod 5), whose points are (0,2), (0,3), (1,2), (1,3), (2,0), (4,2), (4,3) and ∞.

Compute (1, 2) + (4, 3). The slope is (3 − 2)/(4 − 1) = 1/3

Why: A fraction, which mod 5 means 1 · 3⁻¹.

3⁻¹ ≡ 2 (mod 5), since 3 · 2 = 6 ≡ 1

Why: So m ≡ 2.

x₃ ≡ m² − x₁ − x₂ ≡ 4 − 1 − 4 ≡ −1 ≡ 4

Why: Reducing mod 5 throughout.

y₃ ≡ m(x₁ − x₃) − y₁ ≡ 2(1 − 4) − 2 ≡ −8 ≡ 2

Why: So the sum is (4, 2).

Verify: (4, 2) is on the list of points, as it must be

Why: Closure is guaranteed by the group law, but checking it on a small example is the fastest way to catch a sign error in the formulas.

Figure (svg): The points of an elliptic curve modulo five, found by substituting each x and testing whether the result is a square.

A curve mod p is a finite set of points, and that finiteness is what makes it useful for cryptography.

20. Hasse's theorem

Concept

How many points does a curve mod p have? The heuristic says about p + 1. Hasse made that precise in the 1930s.

\[ \bigl| N - p - 1 \bigr| < 2\sqrt p \]

So the count is pinned to an interval of width about 4√p around p + 1 — and conversely, every N in that interval is achieved by some curve mod p.

Two consequences matter later. For cryptography, N is essentially p, so choosing a 256-bit prime gives a group of about 2²⁵⁶ elements. For factoring, N varies across curves within Hasse's interval, and that variation is exactly what Lenstra's method exploits.

Counting points is not trivial for large p. Listing them is hopeless beyond about 10²⁰, and the practical algorithms are due to Schoof, Atkin and Elkies — polynomial time, and essential, because a curve cannot be used until its point count is known.

Figure (svg): Hasse's interval around p plus one, narrowing in relative terms as the prime grows.

Hasse's theorem: the number of points is essentially p, and the uncertainty is only of order √p.

21. A larger example, mod 2773

Worked example

Take E: y² ≡ x³ + 4x + 4 (mod 2773) and P = (1, 3). Compute 2P.

The tangent slope is (3x² + 4)/(2y) = 7/6 at (1, 3)

Why: From the doubling formula.

Invert 6 mod 2773 by the extended Euclidean algorithm: 2311 · 6 ≡ 1

Why: So 1/6 becomes 2311.

m ≡ 7 · 2311 ≡ 16177 ≡ 2312 (mod 2773)

Why: Since 16177 − 5 · 2773 = 2312.

x₃ ≡ 2312² − 1 − 1 ≡ 1771

Why: And y₃ ≡ 2312(1 − 1771) − 3 ≡ 705.

Verify: 2P = (1771, 705)

Why: Note what the calculation needed: gcd(6, 2773) = 1, so the inverse existed. In the next section the same curve, one step further, will produce a denominator whose gcd with 2773 is not 1 — and 2773 is not prime.

Figure (svg): Lenstra's method: multiples of P reach infinity at different times modulo each prime factor, and the gcd catches the gap.

Every factoring method in this course is a way of making two primes behave differently; here it is when the multiples run out.

22. Discrete logarithms on elliptic curves

Concept

The classical problem: given g and h with h ≡ gᵏ (mod p), find k. The elliptic version substitutes the group.

Elliptic curve discrete logarithm problem — Given points A and B on E with B = kA, find the integer k.

It does not look like a logarithm, because the group is written additively — kA is a multiple, not a power. But it is the same problem in a different notation, and the name is kept.

There is no good general attack. Two classical methods partly transfer and one does not:

  1. Pohlig-Hellman transfers. If n, the smallest integer with nA = ∞, has only small prime factors, k can be found modulo each prime power and recombined by CRT. Defeated by choosing E and A so that n has a large prime factor.
  2. Baby-step giant-step transfers, but needs about √n memory, which is impractical at cryptographic sizes.
  3. Index calculus does not transfer at all — and that is the whole story.

Figure (svg): Why index calculus does not transfer: subtracting a small point can produce a large one, so there is no notion of progress.

The security of every deployed curve rests on the absence of this one algorithm, not on a proof.

23. Why index calculus does not transfer

Concept

Index calculus, from Section 10.2, works by expressing elements in terms of a factor base of small primes. The elliptic version fails for a precise reason.

There is no good analogue of 'small'. You might try points with small coordinates as the factor base, but the analogy breaks at the crucial step.

When factoring an integer, dividing off a prime makes the quotient smaller, so repeated division visibly makes progress and terminates.

On a curve, subtracting a point with small coordinates can produce a point with large coordinates. The doubling example in Section 21.1 went from (−4, −3) to (72, 611) in one step. So there is no way to tell whether a decomposition is getting closer to finishing.

Hence the best known attacks are generic, taking about √n operations — and a 256-bit curve gives 128 bits of security, where a mod-p group would need about 3072 bits for the same. That single fact is the entire practical case for elliptic curves.

It is worth being precise about the status of this claim. No one has proved that index calculus cannot be adapted; the claim is that forty years of trying has not produced one. That is good evidence and it is not a theorem.

Figure (svg): Why index calculus does not transfer: subtracting a small point can produce a large one, so there is no notion of progress.

The security of every deployed curve rests on the absence of this one algorithm, not on a proof.

24. Which attacks carry over to curves?

Definition probe

Four techniques from earlier chapters.

Sort into buckets

Sort each by whether it works against elliptic curve discrete logs.

Transfers
Pohlig-Hellman; Baby-step giant-step; Pollard rho
Does not transfer
Index calculus
works
All three are generic group algorithms — they use only the group operation and never look inside the elements. So they work in any group, and their cost is about √n, which is why n must be around 2²⁵⁶.
no
Index calculus is not generic: it exploits the fact that integers factor into primes, and there is no analogue of a small prime on a curve. Its absence is the reason curve keys can be so much shorter.

25. Koblitz Encoding: messages as points

Concept

To encrypt with a curve, a message must first become a point — and unlike the mod-p case, that is not simply a matter of reading the message as a number.

There is no known deterministic polynomial-time algorithm for writing down points on an arbitrary curve mod p. But probabilistic methods are fast, and Koblitz's is the standard one.

The idea: embed m in the x-coordinate, and adjust a few spare bits until x³ + bx + c happens to be a square.

  1. Fix K so that a failure rate of 2⁻ᴷ is acceptable, and require (m+1)K < p
  2. For j = 0, 1, …, K−1, set x = mK + j and test whether x³ + bx + c is a square mod p
  3. If it is, take the square root y and set P_m = (x, y). If not, increment j
  4. Recover the message as m = ⌊x/K⌋

Each try succeeds about half the time, so the chance of failing all K times is about 2⁻ᴷ — and K = 30 makes it about one in a billion.

26. Encoding a message with p = 179

Worked example

On y² = x³ + 2x + 7 (mod 179), with a failure rate of 2⁻¹⁰ acceptable, so K = 10.

The constraint (m+1)K < 179 gives 0 ≤ m ≤ 16

Why: So this curve encodes a very small message — the point of the example is the mechanism, not the capacity.

Take m = 5. The candidate x-values are 50 through 59

Why: x = mK + j = 50 + j for j = 0, …, 9.

At x = 51: 51³ + 2·51 + 7 ≡ 121 (mod 179)

Why: And 121 = 11², so a square root exists.

So P_m = (51, 11)

Why: j = 1 worked, which is typical — about half of all j succeed.

Verify: recovery is m = ⌊51/10⌋ = 5

Why: The division discards exactly the j that was used to search, and nothing else. Note that when p ≡ 3 (mod 4) the square root is a single exponentiation by (p+1)/4 — the formula from Chapter 19.

Figure (svg): The points of an elliptic curve modulo five, found by substituting each x and testing whether the result is a square.

A curve mod p is a finite set of points, and that finiteness is what makes it useful for cryptography.

27. What is the chance the encoding fails?

Prediction

Predict first

Roughly how often does it fail to find a point?

  • About half the time
  • About 1 in a thousand
  • About 1 in a billion
  • Never

Correct: About 1 in a billion

The cost is small and the benefit is a hard guarantee, which is the usual shape of a probabilistic construction: pay a few bits, drive the failure rate below anything that will occur in practice.

Modern systems mostly avoid the problem entirely. ECIES encrypts the message with a symmetric cipher and uses the curve only for key agreement, so no message ever has to become a point. Koblitz encoding matters when the plaintext itself must live in the group — as in textbook elliptic ElGamal.

Why: Each j succeeds with probability about 1/2, independently, so all thirty fail with probability about 2⁻³⁰ — around one in a billion. The failure rate is tunable by choosing K, at a cost of log₂ K bits of message capacity.

28. Factoring with Elliptic Curves

Section

Section 21.3 · pp. 409-417

29. Lenstra Elliptic Curve Factorization

Concept

The failed inversion that would be a bug elsewhere is the algorithm here.

Choosing a curve is done backwards. Pick a point P and a coefficient b first, then choose c so that P lies on y² = x³ + bx + c. Far more efficient than picking the curve and hunting for a point.

Then compute a large multiple of P modulo n, typically B!P by successive doubling. Every addition needs a modular inverse, and every inverse needs a gcd.

When a gcd comes out strictly between 1 and n, that gcd is a factor and the algorithm stops.

Why it works. By CRT, a curve mod n = pq behaves like a pair of curves, one mod p and one mod q. The multiples of P reach ∞ on the two curves at different times, because the two point counts are unrelated. A denominator that is 0 mod p and nonzero mod q is exactly that difference, and the gcd isolates it.

Figure (svg): Lenstra's method: multiples of P reach infinity at different times modulo each prime factor, and the gcd catches the gap.

Every factoring method in this course is a way of making two primes behave differently; here it is when the multiples run out.

30. Factoring 2773

Worked example

Continuing the earlier example: E: y² ≡ x³ + 4x + 4 (mod 2773) with P = (1, 3), chosen by fixing P and b and solving 3² ≡ 1 + 4 + c for c = 4.

2P = (1771, 705), computed earlier

Why: That step needed 6⁻¹ mod 2773, and gcd(6, 2773) = 1, so it went through.

Now compute 3P = 2P + P. The slope is (705 − 3)/(1771 − 1) = 702/1770

Why: A chord, not a tangent, so the first slope formula applies.

Try to invert 1770 mod 2773 — and gcd(1770, 2773) = 59

Why: The extended Euclidean algorithm reports the gcd on its way to the inverse, so the failure is detected for free.

Verify: 2773 = 59 × 47

Why: What happened: 3P = ∞ on E mod 59 while 4P = ∞ on E mod 47. The slope was infinite mod 59 and finite mod 47, so the denominator was 0 mod 59 and nonzero mod 47 — and the gcd separated them.

Figure (svg): Lenstra's method: multiples of P reach infinity at different times modulo each prime factor, and the gcd catches the gap.

Every factoring method in this course is a way of making two primes behave differently; here it is when the multiples run out.

31. Factoring 455839 with 8!P

Worked example

A larger run, showing the method as it is actually used.

E: y² ≡ x³ + 5x − 5 (mod 455839), P = (1, 1)

Why: Again chosen by fixing the point and solving for c.

Compute 2!P, then 3!P = 3(2!P), then 4!P = 4(3!P), and so on

Why: Multiplying by factorials builds up a highly composite multiplier cheaply.

Everything is fine through 7!P, but 8!P requires inverting 599

Why: And gcd(599, 455839) = 599.

So 455839 = 599 × 761

Why: Recovered from a single failed inversion.

Verify: the reason: #E(mod 599) = 640 = 2⁷ · 5, and 8! is a multiple of 640

Why: So 8!P = ∞ on the curve mod 599. But #E(mod 761) = 777 = 3 · 7 · 37, which does not divide 8!, so 8!P is an ordinary point there. Reaching ∞ means dividing by 0, and that is the failure the gcd caught.

Figure (svg): Factoring 455839: successive factorial multiples of P until an inversion fails.

The smooth side reaches ∞ first, and the failed inversion is the algorithm's output rather than an error.

32. Smoothness, and why this beats p − 1

Concept

The general principle. On E mod p, the smallest m with mP = ∞ divides the point count N, by Lagrange's theorem, so NP = ∞. If N is a product of small primes, then B! is a multiple of N for a modest B, and B!P = ∞.

B-smooth — An integer all of whose prime factors are at most B. Smoothness has driven the x² ≡ y² method, the p − 1 method, and the index calculus attack; here it appears once more.

Now the comparison with Section 9.4's p − 1 method. That method needs p − 1 to be smooth, and p − 1 is determined by p. If it has a large prime factor, the method fails and there is nothing to adjust.

The elliptic curve method needs N = #E(mod p) to be smooth — and N changes with the curve. Hasse's theorem says N lies near p in an interval wide enough to contain many integers, and enough of them are smooth that a random curve has a fair chance.

So a failure is not the end but a retry. Run fourteen curves in parallel for a fifty-digit number, more for larger ones, and one of them is likely to have a smooth N. That is the whole advantage: not a better test, but an unlimited supply of independent attempts.

Figure (svg): Why the elliptic curve method beats p minus one: a fresh curve gives a fresh point count, so a bad draw can be retried.

The advantage is not a better test but an unlimited supply of retries, which is what turns a lucky method into a reliable one.

33. Where is ECM actually the right tool?

Socratic

Lenstra's method is not what breaks RSA moduli.

Discussion prompt

What is it good at, and what beats it elsewhere?

Hint: Its running time depends on the size of the factor, not the size of n.

Answer:

Its cost depends mainly on the size of the smallest prime factor, not on the size of n. So it excels at pulling a 10- or 20-digit factor out of a very large number — which no other general method does efficiently.

It is the method of choice for numbers of medium size, around 40 to 50 digits, and for finding small factors before handing a number to something heavier.

For large numbers with two large factors — an RSA modulus — the quadratic sieve and the number field sieve are far superior, because their cost depends on the size of n and not on the factors.

So a real factoring pipeline runs them in sequence: trial division, then Pollard rho, then ECM to strip medium factors, then the number field sieve on what remains. Each is best in a range and useless outside it.

And the practical relevance to RSA is indirect but real. ECM is why an RSA modulus must not have any small-ish factor, and why key generation must produce two primes of equal size — a 2048-bit modulus with a 60-digit factor would fall to ECM in an afternoon.

34. Singular curves: two old methods in disguise

Concept

The construction assumed the cubic has no repeated roots. What if it does? The answer is a genuine surprise.

The discriminant 4b³ + 27c² is zero exactly when there is a multiple root — the cubic analogue of b² − 4ac for quadratics. Working mod a composite n, the gcd of n and the discriminant might land strictly between 1 and n, which is already a factor, so you stop.

A double root, as in y² = x³ − 3x + 2 = (x−1)²(x+2): to each point associate the number (y + √3(x−1))/(y − √3(x−1)). Adding points corresponds to multiplying these numbers, so factoring with this curve is essentially the p − 1 method.

A triple root, y² = x³: associate x/y to each point. Then mP has associated number m, and adding points corresponds to adding integers. Factoring here amounts to computing gcd(2, n), gcd(3, n), … — trial division.

So p − 1 and trial division are both special cases of Lenstra's algorithm, recovered by letting the curve degenerate. That is not a coincidence but a statement about what the group of a singular curve becomes: the multiplicative group in one case and the additive group in the other.

Figure (svg): Singular curves in disguise: a double root reduces to the p minus one method and a triple root to trial division.

The two oldest factoring methods turn out to be Lenstra's algorithm run on degenerate curves — a genuinely surprising unification.

35. The double-root curve mod 143

Worked example

Concretely, on y² = x³ − 3x + 2 (mod 143), with 143 = 11 · 13.

3 is a square mod 143: 82² ≡ 3

Why: Convenient, so √3 can be replaced by 82 throughout. If it were not, a different curve would be chosen.

Take P = (−1, 2) and compute multiples

Why: 2P = (2, 141), 3P = (112, 101), 4P = (10, 20) — and computing 5P finds the factor 11.

The number attached to P is (2 + 82(−1−1))/(2 − 82(−1−1)) ≡ 80 (mod 143)

Why: Using the ratio of the two tangent lines at the singular point.

The numbers for P, 2P, 3P, 4P are 80, 108, 60, 81

Why: And the powers 80¹, 80², 80³, 80⁴ mod 143 are 80, 108, 60, 81 — identical.

Verify: 80⁵ ≡ 45, and 45 ≡ 1 (mod 11) but not mod 13

Why: Which is exactly the statement that 5P = ∞ mod 11 and not mod 13. Point addition has become multiplication, ∞ has become 1, and the elliptic curve method has become the p − 1 method.

Figure (svg): Singular curves in disguise: a double root reduces to the p minus one method and a triple root to trial division.

The two oldest factoring methods turn out to be Lenstra's algorithm run on degenerate curves — a genuinely surprising unification.

36. The discriminant shares a factor with n

Anomaly

You choose a curve mod n and compute gcd(4b³ + 27c², n), finding it equals 59 with n = 2773.

Predict first

What do you do?

  • Discard the curve and pick another
  • Stop — that gcd is a nontrivial factor of n, which was the goal
  • Use the curve anyway, since it is only singular mod one factor
  • Increase B and continue

Correct: Stop — that gcd is a nontrivial factor of n, which was the goal

This is the general pattern of the method: every arithmetic operation that could fail is a chance to find a factor, and 'failure' is the success condition.

It is worth checking the discriminant before starting rather than discovering the degeneracy mid-run, and implementations do — it costs one gcd and can save the whole computation.

Why: A gcd strictly between 1 and n is a factor, however it was obtained. The curve is singular mod 59 and non-singular mod 47 — and detecting that difference is precisely what the whole algorithm is trying to do, so arriving there early is a win.

37. Elliptic Curves in Characteristic 2

Section

Section 21.4 · pp. 417-420

38. Why the equation must change

Concept

Many applications use curves over GF(2ⁿ), because binary arithmetic suits hardware. Of the fifteen curves NIST recommended in 1999, ten are over binary fields.

But the equation y² = x³ + bx + c fails mod 2. Differentiating gives 2y y′ = 0, since 2 = 0, so every tangent line is vertical and 2P = ∞ for every point. More precisely, the curve is singular — the partial derivatives vanish simultaneously.

So the general Weierstrass form is needed:

\[ E: \; y^2 + a_1 x y + a_3 y = x^3 + a_2 x^2 + a_4 x + a_6 \]

Over any field where 2 and 3 are invertible, a change of variables reduces this to y² = x³ + bx + c. In characteristic 2 or 3 it cannot, which is exactly why the longer form exists.

The addition law keeps its structure — three collinear points still sum to ∞, and lines through ∞ are still vertical — but finding −P is no longer just flipping the sign of y.

Figure (svg): An elliptic curve over the field of two elements, where the usual equation degenerates and negation changes.

Binary fields suit hardware, and the price is a longer equation and a negation rule that must be relearned.

39. Adding points on a curve mod 2

Worked example

Take E: y² + y ≡ x³ + x (mod 2). Its points are (0,0), (0,1), (1,0), (1,1) and ∞.

Compute (0, 0) + (1, 1). The line through them is y = x

Why: Both points satisfy it.

Substituting gives x² + x ≡ x³ + x, that is x²(x + 1) ≡ 0

Why: Roots x = 0, 0, 1 — so x = 0 is a double root and the line is tangent at (0, 0).

The third intersection has x = 0 and lies on y = x, so it is (0, 0) again

Why: Hence (0,0) + (0,0) + (1,1) = ∞.

Now find −(0, 0): the vertical line x = 0 meets E where y² + y = 0, that is y = 0 or 1

Why: So the other point is (0, 1), and (0,0) + (0,1) = ∞.

Verify: therefore (0, 0) + (1, 1) = (0, 1)

Why: The answer is not obtained by flipping a sign — over GF(2) there are no signs to flip. Negation must be computed from the vertical line each time, and that is the practical difference the longer equation forces.

Figure (svg): An elliptic curve over the field of two elements, where the usual equation degenerates and negation changes.

Binary fields suit hardware, and the price is a longer equation and a negation rule that must be relearned.

40. GF(4) and the fields actually used

Concept

Curves mod 2 are far too small, so finite fields GF(2ⁿ) are used instead. The smallest interesting one is GF(4).

GF(4) = {0, 1, ω, ω²} with x + x = 0 for all x and 1 + ω = ω². From these, ω³ = ω · ω² = ω(1 + ω) = ω + ω² = ω + 1 + ω = 1, so ω² is the inverse of ω and every nonzero element is invertible.

Curves over a finite field are treated exactly like curves over the integers — the same equation, the same collinearity rule, the same procedure for listing points.

For cryptographic use, n is at least 150, giving a field of about 2¹⁵⁰ elements and a curve with about that many points.

A note on the current landscape: binary curves were attractive when hardware multipliers were scarce, and they have fallen out of favour. Modern practice prefers prime-field curves such as P-256 and Curve25519, partly because binary-field discrete logs have seen real algorithmic progress in related settings — a reminder that 'no known attack' can change.

41. Adding points over GF(4)

Worked example

Take E: y² + xy = x³ + ω over GF(4).

List the points by substituting each x

Why: x = 0 gives y² = ω, so y = ω². x = 1 gives y² + y = 1 + ω = ω², which has no solution. x = ω gives y² + ωy = ω², solved by y = 1 and y = ω². x = ω² gives no solution.

So E has the points (0, ω²), (ω, 1), (ω, ω²) and ∞

Why: Four points — a small group, but enough to demonstrate the law.

Compute (0, ω²) + (ω, ω²). The line through them is y = ω²

Why: Both share that y-coordinate.

Substituting: ω⁴ + ω²x = x³ + ω, which becomes x³ + ω²x = 0, with roots x = 0, ω, ω

Why: So the third intersection is (ω, ω²), and (0,ω²) + (ω,ω²) + (ω,ω²) = ∞.

Verify: so the answer is −(ω, ω²), found by the vertical line x = ω, which gives (ω, 1)

Why: Hence (0, ω²) + (ω, ω²) = (ω, 1). The procedure is identical to the real case; only the arithmetic of the field has changed.

Figure (svg): An elliptic curve over the field of two elements, where the usual equation degenerates and negation changes.

Binary fields suit hardware, and the price is a longer equation and a negation rule that must be relearned.

42. Elliptic Curve Cryptosystems

Section

Section 21.5 · pp. 420-424

43. The translation table

Concept

Any discrete-log system can be converted mechanically. The dictionary is the content of the section.

mod pelliptic curve
nonzero numbers mod ppoints on E
multiplication mod paddition of points
1 (multiplicative identity)∞ (additive identity)
division mod psubtraction of points
exponentiation gᵏinteger multiple kP
p − 1N, the number of points
Fermat: aᵖ⁻¹ ≡ 1NP = ∞ (Lagrange)
solve gᵏ ≡ h for ksolve kP = Q for k

Three notes the book attaches. Addition and subtraction of points cost the same, whereas multiplication mod p is far cheaper than division. Both mod-p operations are individually simpler than the curve operations. And the curve discrete log is believed harder than the mod-p one at the same size.

Figure (svg): The dictionary translating a discrete-log cryptosystem modulo p into its elliptic curve counterpart.

Not an analogy but a substitution: swap the group and every protocol built on it comes along unchanged.

44. Not every group works: a cautionary case

Concept

The table might suggest any group will do. The book's fourth note shows otherwise, and it is worth pausing on.

Take the integers mod m under addition. The analogues are: addition mod m, identity 0, subtraction, the multiple ka = a + ⋯ + a, the count m, the relation ma ≡ 0, and the discrete log problem 'solve ka ≡ b (mod m) for k'.

And that problem is easy. The extended Euclidean algorithm solves it in a few steps.

So the difficulty of a discrete logarithm depends entirely on the binary operation, not on the size of the group. Multiplication mod p is hard to invert in this sense; addition mod m is not; point addition on a curve appears to be the hardest of the three.

Which is why the search for new groups is a real research programme rather than a formality. A group is a candidate only if inverting its multiples resists everything known, and most groups fail that test immediately.

Figure (svg): The dictionary translating a discrete-log cryptosystem modulo p into its elliptic curve counterpart.

Not an analogy but a substitution: swap the group and every protocol built on it comes along unchanged.

45. Elliptic Curve ElGamal

Concept

Recall the classical version: Bob publishes p, α and β ≡ αˢ; Alice sends y₁ ≡ αᵏ and y₂ ≡ xβᵏ; Bob recovers x ≡ y₂y₁⁻ˢ.

Now read down the translation table. Bob chooses a curve E mod p, a point α on it, and a secret integer s, and publishes β = sα.

  1. Alice represents her message as a point x on E — Koblitz encoding, from Section 21.2
  2. She picks a random k and computes y₁ = kα and y₂ = x + kβ
  3. She sends the pair (y₁, y₂)
  4. Bob decrypts by computing x = y₂ − sy₁

Why it works: sy₁ = s(kα) = k(sα) = kβ, so subtracting it from y₂ = x + kβ leaves x. Every step is the classical protocol with multiplication replaced by addition.

The book notes a more workable variant due to Menezes and Vanstone, which avoids encoding the message as a point at all — the ancestor of the ECIES construction used today.

46. An ElGamal exchange with p = 8831

Worked example

Concrete numbers, all checkable.

Take p = 8831, G = (4, 11) and b = 3, forcing c = 45

Why: Since 11² = 121 and 4³ + 3·4 = 76, we need c = 45. Choosing the point first and solving for c is the standard trick.

Bob's secret is s_B = 3, and he publishes s_B G = (413, 1808)

Why: One point multiplication.

Alice has the message point P_m = (5, 1743) and picks k = 8

Why: Check: 5³ + 3·5 + 45 = 185, and 1743² ≡ 185 (mod 8831). ✓

She sends kG = (5415, 6321) and P_m + k(s_B G) = (6626, 3576)

Why: Two point multiplications and one addition.

Bob computes s_B(kG) = 3(5415, 6321) = (673, 146)

Why: The shared masking point, reached from the other side.

Verify: (6626, 3576) − (673, 146) = (6626, 3576) + (673, −146) = (5, 1743)

Why: The original message point. Subtraction is addition of the reflection, exactly as Section 21.1 defined it.

Figure (svg): The dictionary translating a discrete-log cryptosystem modulo p into its elliptic curve counterpart.

Not an analogy but a substitution: swap the group and every protocol built on it comes along unchanged.

47. Elliptic Curve Diffie-Hellman

Concept

The shortest translation in the chapter, because the original protocol has so few moving parts.

Alice and Bob agree on a curve and a public base point G. Alice picks N_A at random, Bob picks N_B, and they publish N_A G and N_B G while keeping the multipliers private.

Alice computes N_A(N_B G); Bob computes N_B(N_A G). These are equal, because integer multiples of a point commute — the same associativity that made 100P well defined.

\[ N_A(N_B G) = (N_A N_B) G = N_B(N_A G) \]

An eavesdropper sees G, N_A G and N_B G and must solve the elliptic curve discrete logarithm problem to recover either multiplier. That is the whole security argument, and it is Chapter 10's argument with the group swapped.

Figure (svg): Elliptic curve Diffie-Hellman: both parties reach the same point from opposite directions.

The Diffie-Hellman idea unchanged, with the group swapped; every step of Chapter 10's protocol carries over verbatim.

48. A key exchange with p = 7211

Worked example

The numbers from the book.

p = 7211, b = 1 and G = (3, 5), which forces c = 7206

Why: Since 25 = 27 + 3 + c gives c = −5 ≡ 7206 (mod 7211).

Alice picks N_A = 12 and publishes N_A G = (1794, 6375)

Why: Four doublings and a couple of additions.

Bob picks N_B = 23 and publishes N_B G = (3861, 1242)

Why: Likewise.

Alice computes 12 · (3861, 1242) = (1472, 2098)

Why: Using Bob's published point.

Bob computes 23 · (1794, 6375) = (1472, 2098)

Why: Using Alice's.

Verify: the same point, so the same shared secret

Why: In practice only the x-coordinate is kept and it is run through a key derivation function — the raw point has structure and is not a uniform bit string, which is Chapter 10's lesson repeated.

Figure (svg): Elliptic curve Diffie-Hellman: both parties reach the same point from opposite directions.

The Diffie-Hellman idea unchanged, with the group swapped; every step of Chapter 10's protocol carries over verbatim.

49. Elliptic Curve ElGamal Digital Signatures

Concept

The signature version needs a little more care, because it mixes integers with points.

Alice fixes a curve E mod p and a point A, computes the number of points N, and requires 0 ≤ m < N. Her private integer is a, and B = aA is public along with p, E, N, A.

  1. She picks a random k with 1 ≤ k < N and gcd(k, N) = 1, and computes R = kA = (x, y)
  2. She computes s ≡ k⁻¹(m − ax) (mod N)
  3. She sends the signed message (m, R, s) — where R is a point and m, s are integers

Bob verifies by computing V₁ = xB + sR and V₂ = mA, and accepting if V₁ = V₂.

Note the x doing double duty. The x-coordinate of R is used as an integer in the equation for s. That choice is arbitrary — any way of assigning integers to points would work — and the classical scheme makes the same arbitrary choice with r.

Figure (svg): The elliptic curve ElGamal signature and the cancellation that makes verification work.

The verification is an identity, and the subtle step is that N times any point is ∞ — Lagrange doing the work Fermat did before.

50. Why the verification works

Worked example

A short computation, with one subtlety worth spelling out.

V₁ = xB + sR, and B = aA while R = kA

Why: So V₁ = xaA + s(kA).

Substitute s = k⁻¹(m − ax)

Why: V₁ = xaA + k⁻¹(m − ax)(kA).

The k⁻¹ and k cancel, leaving V₁ = xaA + (m − ax)A = mA = V₂

Why: Which is the verification condition.

The subtlety: k⁻¹k is not 1 but 1 + tN for some integer t

Why: It is an inverse modulo N, not an exact inverse.

\[ k^{-1}k A = (1 + tN)A = A + t(NA) = A + t\infty = A \]

Verify: the extra multiple of N vanishes because NA = ∞

Why: By Lagrange's theorem, N times any point is the identity. That is the exact analogue of Fermat's aᵖ⁻¹ ≡ 1 in the classical scheme, and it is the row of the translation table that makes the whole signature scheme go through.

Figure (svg): The elliptic curve ElGamal signature and the cancellation that makes verification work.

The verification is an identity, and the subtle step is that N times any point is ∞ — Lagrange doing the work Fermat did before.

51. Reading the signature equation

Notation

One congruence carries the whole scheme.

Annotate

On: \( s \equiv k^{-1}(m - a x) \pmod N \)

  • The only place the long-term secret enters. It is multiplied by x and hidden behind the subtraction and the k⁻¹.
  • Must be fresh, uniform, and coprime to N. Two signatures with the same k give two equations in a and k, and the private key falls out — the failure of Chapter 13 and Chapter 19, now for the third time.
  • An arbitrary but standard way of turning the point R into a number. Any injective-enough map would serve.
  • Not p. The exponent arithmetic lives modulo the group order, which is why N must be computed before the curve can be used — and why Schoof's algorithm matters in practice.
  • Otherwise k⁻¹ mod N does not exist and the signature cannot be formed. Choosing N prime, as standardised curves do, makes this automatic for any k in range.

ECDSA differs from this in details of how R is reduced to an integer, but the shape and the nonce requirement are identical — which is why the PlayStation 3 key recovery applies verbatim.

52. Key sizes: what the advantage actually is

Concept

The practical claim from the chapter's opening: certain conventional systems with a 4096-bit key can be replaced by 313-bit elliptic curve systems.

The reason is entirely about attacks. Index calculus makes mod-p discrete logs subexponential, so the modulus must grow much faster than the security level. Generic attacks on curves take about √N steps, so N need only be about 2²ˢ for s bits of security.

Hence the rough correspondence: 128-bit security needs a 3072-bit RSA modulus or a 256-bit curve; 256-bit security needs 15360 bits of RSA or a 512-bit curve. The gap widens as the security level rises.

The savings compound. Shorter keys mean shorter signatures and certificates, less bandwidth per handshake, smaller hardware multipliers, and less power — which is why constrained devices moved to curves first.

And the caveat is worth repeating. The advantage exists because index calculus has no elliptic analogue, and that is an empirical fact about the state of the art. It is the one assumption on which the entire size advantage rests.

Figure (svg): Comparable security levels for RSA and elliptic curve key sizes, showing the growth gap.

The entire practical case for elliptic curves in one chart, and the reason is an attack that does not exist.

53. Where elliptic curves actually run

Real world

This is the most widely deployed public-key mathematics in the world.

Discussion prompt

Name four places curve arithmetic is running right now.

Hint: Web traffic, messaging, cryptocurrency, and device identity.

Answer:

TLS. Essentially every HTTPS handshake uses elliptic curve Diffie-Hellman for key agreement — X25519 or P-256 — because it is fast and the keys are small. The RSA key exchange it replaced is removed entirely in TLS 1.3.

Signatures. Ed25519 and ECDSA sign software updates, SSH sessions, certificates and packages. An Ed25519 public key is 32 bytes against 256 for a 2048-bit RSA key.

Cryptocurrencies. Bitcoin and Ethereum addresses are derived from secp256k1 public keys, and every transaction is an ECDSA signature. The nonce-reuse failure has drained real wallets.

Device and platform identity. Secure elements, TPMs, passkeys and FIDO2 authenticators use curves because the key material fits in constrained storage and the arithmetic fits in a small coprocessor.

And the curve choice is itself a story. NIST's P-curves have unexplained seed constants, and Curve25519 was designed with rigid, publicly justified parameters partly in response — Chapter 5's Dual_EC affair made that suspicion mainstream.

54. Find the flaws in this ECDSA implementation

Error analysis

From a code review.

Annotate

  • Predictable and repeatable — two signatures in the same second share a nonce, and one reused nonce yields the private key by a subtraction. RFC 6979's deterministic nonce derivation exists for exactly this.
  • The fatal one. Invalid-curve attacks feed a point from a different, weaker curve; the victim's scalar multiplication then leaks the private key modulo a small order, and a few queries recover it entirely.
  • Necessary and nowhere near sufficient. A full check verifies the point satisfies the curve equation, has coordinates in range, and lies in the correct subgroup.
  • Not constant time, so it leaks how many leading characters matched — Chapter 14's timing channel. And comparing encodings rather than values invites malleability, where a re-encoded signature verifies but compares unequal.

The second is the one that ends the system. Every scalar multiplication on attacker-supplied input must be preceded by validation, and libraries that get this right do it before anything else.

55. Curves versus integers mod p

Trade off

The same protocols in two groups. Fill the blanks.

Comparison matrix

Integers mod pElliptic curve
Best known attackindex calculus, subexponentialgeneric, about √N steps
Size for 128-bit security3072 bits256 bits
Cost of the group operationone multiplicationseveral multiplications and an inversion
Inverse elementexpensive — extended Euclidfree — flip the sign of y
Setup requiredchoose a safe primechoose a curve and count its points

The curve operation is individually more expensive and the numbers are twelve times smaller, so curves win overall — and the point-counting requirement is why standardised curves exist rather than everyone generating their own.

56. Match each classical object to its curve counterpart

Matching

Reading down the translation table.

Match the pairs

  • m1. The identity 1
  • m2. Exponentiation gᵏ
  • m3. Fermat's aᵖ⁻¹ ≡ 1
  • m4. p − 1
  • m5. Division mod p
  • n1. The point at infinity
  • n2. The multiple kP
  • n3. NP = ∞, by Lagrange's theorem
  • n4. N, the number of points on E
  • n5. Subtraction of points

Why: The third pairing is the one that does real work in the signature proof: Fermat's theorem is the special case of Lagrange's for the multiplicative group mod p, and on a curve the general version applies directly. Everything else is notation.

57. Which of these make a curve unsuitable?

Discrimination

Five properties a candidate curve might have.

Sort into buckets

Sort each.

Unsuitable
The point count N has only small prime factors; The cubic has a repeated root; N equals p exactly; The curve is defined over GF(2ⁿ) with n = 8
Fine
N is prime and about 2²⁵⁶
bad
A smooth N falls to Pohlig-Hellman. A repeated root makes the curve singular, and its group collapses to the multiplicative or additive group where discrete logs are easy. N = p is an anomalous curve, attackable in polynomial time by the Smart-Satoh-Araki attack. And a 256-element field gives a group far too small to search-proof.
ok
A large prime order near 2²⁵⁶ is exactly what standardised curves aim for: it defeats Pohlig-Hellman outright and puts generic attacks at about 2¹²⁸ steps.

58. What does the size advantage depend on?

Edge cases

A 256-bit curve is claimed to match a 3072-bit RSA modulus.

Discussion prompt

What would have to change for that claim to fail?

Hint: The advantage rests on an absence.

Answer:

An index calculus algorithm for curves would end it. The claim is not that none exists but that none has been found in forty years of effort. If one appeared, curve sizes would have to grow the way prime sizes did.

A better generic algorithm would weaken both sides equally, since √N is a proved lower bound for generic attacks — so that particular worry is bounded.

A structural weakness in a specific curve family would be narrower, affecting some curves and not others. This has happened: supersingular curves fall to the MOV attack, which maps the problem into a finite field where index calculus does apply, and anomalous curves with N = p fall in polynomial time.

And a quantum computer ends both, since Shor's algorithm solves discrete logs and factoring alike — and curves fall faster, because their smaller keys need fewer qubits. The size advantage inverts into a disadvantage, which is why post-quantum migration is urgent for curves specifically.

So the honest statement is layered: curves are secure against everything currently known, more efficiently than the alternatives, and their advantage is contingent on an algorithmic absence that has held up well but is not a theorem.

59. Order these by how completely they break a curve system

Ranking

Five things that can go wrong.

Put in order

  1. The point count N has a large prime factor but is not prime
  2. The shared secret is used raw instead of through a KDF
  3. A nonce is reused across two signatures
  4. Incoming points are not validated as being on the curve
  5. The curve is anomalous, with N = p

Why: A large prime factor keeps Pohlig-Hellman at bay, so a non-prime N with a large factor is a minor cofactor concern. A raw shared secret has bias and structure, which is real but usually not immediately exploitable. Nonce reuse hands over one private key. Missing point validation lets an attacker choose the curve and extract the key over a few queries. And an anomalous curve is broken in polynomial time for everyone, permanently — the arithmetic itself is the vulnerability.

60. Choose parameters for a new deployment

Constraint

A team needs authenticated key exchange on constrained hardware and asks which curve to use.

Discussion prompt

Give the recommendation and the reasoning.

Hint: The first decision is whether to choose a curve at all.

Answer:

Do not generate a curve. Point counting, subgroup checks and twist security are all easy to get wrong, and a bad curve is invisible from the equation. Use a standardised one.

X25519 for key agreement and Ed25519 for signatures, if the ecosystem permits. Their parameters have public rigid justifications, the arithmetic is designed to be implementable in constant time, and every input is a valid point — which removes the invalid-curve attack class entirely.

P-256 where standards or certification require it, with a library that validates points and runs in constant time. It is not weaker in any known way; it is harder to implement safely.

Derive keys with a KDF, never use the raw shared point. And derive signature nonces deterministically per RFC 6979 or Ed25519's built-in scheme, so a weak RNG cannot cause nonce reuse.

And plan for migration. A quantum computer breaks all of this, so new long-lived deployments should be built to carry a hybrid — a curve exchange combined with a post-quantum one, which is what TLS is already doing.

61. Reading the security estimate

Cost model

One comparison governs every key-size decision.

Annotate

On: \( \text{generic attack} \approx \sqrt N \quad \text{vs} \quad \text{index calculus} \approx \exp\bigl(c (\log p)^{1/3}(\log\log p)^{2/3}\bigr) \)

  • Pollard rho on a curve of order N. Fully generic, proved optimal for generic algorithms, so N ≈ 2²⁵⁶ gives 2¹²⁸ work.
  • The number field sieve's shape, subexponential in log p. It grows much more slowly than the key, so the key must grow much faster than the security level.
  • Doubling the security level doubles the curve size and roughly quintuples the RSA size. At 128 bits the ratio is 12:1; at 256 bits it is 30:1.
  • Hasse's theorem ties N to p, so choosing a 256-bit prime gives a 256-bit group — no separate tuning is needed once the curve is fixed.
  • Shor's algorithm, which is polynomial for both. Against a quantum adversary this whole comparison is void.

Being able to say where each estimate comes from is what lets you evaluate a key-size recommendation instead of copying one.

62. What elliptic curves give you

Two truths and a lie

Two of these claim more than is true.

Eliminate the wrong options

Which statement is correct?

  • a. Comparable security to mod-p systems at much smaller key sizes, because no index calculus attack is known for curves
  • b. Provably harder discrete logarithms than the mod-p case
  • c. Security against quantum computers, since curves are not based on factoring

Survives elimination: a

Why: The correct claim is empirical and carefully hedged: comparable security, smaller keys, because of an attack that has not been found. Both of the others upgrade an absence of evidence into a guarantee, which is the standard way this subject's claims get overstated.

63. Where would a curve deployment fail first?

Commit first

A team deploys P-256 ECDH and ECDSA using a well-regarded library, with a 256-bit curve and a good RNG.

Predict first

What is the realistic failure?

  • Someone solves the elliptic curve discrete log problem
  • An implementation flaw — a timing leak, a missing point validation, or a nonce failure
  • The curve turns out to be anomalous
  • Hasse's theorem is wrong

Correct: An implementation flaw — a timing leak, a missing point validation, or a nonce failure

Scalar multiplication is the danger zone. A naive double-and-add branches on the bits of the secret, so its timing and power profile leak the key directly — Chapter 14's subject, and the reason Montgomery ladders and constant-time libraries exist.

Point validation is the second. An unvalidated point can come from a different curve with a weak group, and the victim's own scalar multiplication then computes the leak.

And nonce failure is the third, for the third time in this course. The Sony PS3 key, several Bitcoin wallets, and Android's SecureRandom bug all reduce to the same subtraction.

Which is the chapter's real conclusion. The choice of group is a solved problem; the difficulty has moved entirely into implementing the group operation without leaking.

Why: The mathematics of P-256 has held for twenty-five years. Real elliptic curve failures are almost entirely implementation failures: timing and cache leaks in scalar multiplication, missing validation of received points, and nonce generation. Every publicly documented break of a deployed curve system is in this category.

64. Explain the group law to someone who knows no algebra

Explain it

A colleague asks what 'adding points on a curve' means.

Discussion prompt

Explain it, and say why anyone would do it.

Hint: Start with the picture, not the formulas.

Answer:

Start with the picture. Draw a curve that looks like a sideways loop. Pick two points on it and draw the straight line through them. That line hits the curve in exactly one more place. Flip that third point across the horizontal axis, and call the result the 'sum' of the first two.

It is not ordinary addition — the coordinates have nothing to do with each other. It is a rule for combining two points into a third, and the name is borrowed because the rule obeys the same laws that ordinary addition does.

Now do it over and over. Starting from a point P and adding it to itself a thousand times gives some other point. That is easy to compute — about twenty steps, by doubling.

And here is the useful part: going backwards is hard. Given P and the final point, working out that it took a thousand steps appears to require trying essentially all the possibilities.

So the number of steps is a secret and the final point can be published. Two people can each publish a point, each apply their own step count to the other's, and land on the same place — which nobody watching can reach. That is the key exchange, and it is why anyone bothers.

Figure (svg): An elliptic curve over the reals with a chord through two points, its third intersection, and the reflection that defines the sum.

The whole group law in one picture: a line meets a cubic three times, so any two points name a third.

65. Why does a failed inversion factor n?

Explain it to yourself

In Lenstra's method, computing 3P mod 2773 failed and produced the factor 59.

Discussion prompt

Explain the mechanism, using the Chinese remainder theorem.

Hint: One curve mod n behaves like two curves.

Answer:

By CRT, arithmetic mod 2773 is arithmetic mod 59 and mod 47 carried out in parallel. So the curve mod 2773 behaves like a pair of curves, and the point P is really a pair of points.

The two components have different orders. On E mod 59, 3P = ∞; on E mod 47, it takes 4P. The multiples run out at different times because the two point counts are unrelated.

Reaching ∞ means a vertical line, which means a slope denominator of zero. So at the third step the denominator was 0 mod 59 and nonzero mod 47 — that is, divisible by 59 and not by 47.

A number divisible by one factor and not the other has a nontrivial gcd with n, and the extended Euclidean algorithm reports it while trying to invert. So the inversion failure is the factorisation.

And the general principle, which the book states plainly: you cannot separate p and q while they behave identically. Every factoring method is a way of making them behave differently — the p − 1 method uses the orders of the multiplicative groups, and this one uses the orders of curve groups, which have the advantage of being re-drawable.

Figure (svg): Lenstra's method: multiples of P reach infinity at different times modulo each prime factor, and the gcd catches the gap.

Every factoring method in this course is a way of making two primes behave differently; here it is when the multiples run out.

66. How fast do the coordinates grow?

Prediction

Predict first

What happens to the size of the coordinates as you keep doubling?

  • They stay about the same
  • They roughly double in digit count with each doubling
  • They grow slowly, by one digit at a time
  • They eventually return to small values

Correct: They roughly double in digit count with each doubling

Which is why rational points are hopeless for computation — and, incidentally, why the theory of heights is a central tool in the number theory of elliptic curves.

Mod p the problem vanishes, because every coordinate is reduced below p. That is a second reason, beyond finiteness, that cryptography works over finite fields: the numbers never grow.

Why: The height of the coordinates grows roughly quadratically, so the number of digits doubles with each doubling of the point. Going from (−4, −3) to (72, 611) is one step; a few more and the numerators run to hundreds of digits.

67. Why choose the point before the curve?

Socratic

Lenstra's method picks P and b first, then solves for c.

Discussion prompt

Why not pick the curve first and then look for a point on it?

Hint: Compare the cost of the two searches.

Answer:

Finding a point on a given curve is a search. You try x-values and test whether x³ + bx + c is a square, succeeding about half the time — and mod a composite n you cannot even test for squareness reliably, because you do not know the factorisation.

Solving for c is one subtraction. Given P = (x₀, y₀) and b, set c = y₀² − x₀³ − bx₀. The point is on the curve by construction, with no search and no square roots.

And the resulting curve is random enough, which is all the method needs — the point count varies across the b values just as it would across independently chosen curves.

The same trick appears in the cryptosystem examples. The book fixes G = (4, 11) and b = 3, then takes c = 45. It is the standard way to produce a curve-with-basepoint in one step.

A general habit worth noticing: when a construction requires an object satisfying a constraint, check whether the constraint can be solved for one of the parameters instead of searched for. It converts a probabilistic loop into arithmetic surprisingly often.

68. A point of order two

Anomaly

On y² ≡ x³ + 4x + 4 (mod 5), one of the points is (2, 0).

Predict first

What is 2·(2, 0)?

  • (2, 0) again
  • ∞, because the tangent at a point with y = 0 is vertical
  • (4, 0)
  • Undefined

Correct: ∞, because the tangent at a point with y = 0 is vertical

Points of order 2 are exactly the points with y = 0, that is, the roots of the cubic. A curve over a field where the cubic splits completely has three of them, plus ∞, giving a subgroup of order 4.

This matters for parameter choice. Such points make the group order even, so it cannot be prime — which is why standardised curves publish a cofactor, and why implementations must either clear it or use a curve designed so that it does not matter.

And it is one more special case an implementation must handle, alongside P + (−P) and P + ∞. Formula-based code that forgets any of them produces a wrong answer rather than an error, which is worse.

Why: The doubling slope is (3x² + b)/(2y), and y = 0 makes the denominator zero — the signal for a vertical line. So the tangent is vertical, its third intersection is ∞, and 2P = ∞. The point has order 2.

69. Sort by what has to be checked before a curve is usable

Definition probe

Some properties are read off the equation; others require computation.

Sort into buckets

Sort each check.

Cheap — arithmetic on the coefficients
The cubic has no repeated root; The base point actually lies on the curve
Requires counting points
The group order has a large prime factor; The order is not equal to p
cheap
The discriminant 4b³ + 27c² is one expression, and substituting the point into the equation is one evaluation. Both are instant regardless of the size of p.
hard
Both need N, and computing N for a cryptographic-size prime requires Schoof's algorithm or its Atkin-Elkies refinements. This is the real reason standardised curves exist: the validation is expensive and easy to skip.

70. Complete the addition law

Faded example

Four blanks.

Fill in the blanks

To add P and Q, draw the line through them — or the tangent at P if they coincide — take the third intersection with the curve, and reflect it through the x-axis. The identity element is ∞, and the negative of (x, y) is (x, −y).

Why: Every one of these follows from the single rule that three collinear points sum to ∞ — the tangent case because tangency is a double root, the identity because vertical lines pass through infinity, and negation because a vertical line joins (x, y) to (x, −y).

71. How big is a curve group in practice?

Estimation

P-256 uses a prime p of 256 bits.

Predict first

Roughly how many points does the curve have?

  • About 2¹²⁸
  • About 2²⁵⁶
  • About 2⁵¹²
  • It depends entirely on b and c

Correct: About 2²⁵⁶

And the security follows immediately: generic attacks take about √N ≈ 2¹²⁸ steps, which is the 128-bit security level P-256 is named for.

Note how little freedom there is. Once the field is chosen, the group size is determined to within a fraction of a percent — so a designer picks the prime for the security level and then searches within Hasse's interval for a curve whose order is prime.

Why: Hasse's theorem pins N to within 2√p of p + 1, so N is about p — around 2²⁵⁶. The coefficients shift N only within that narrow interval, so the group size is fixed by the prime rather than by the curve.

72. Swap the group, keep the protocol

Pattern

The chapter contains one idea applied twice, and it is worth stating in the abstract.

  1. Find a group where the operation is cheap and inverting a multiple is hard. The points of a curve, under chord-and-tangent addition, are such a group.
  2. Then every discrete-log protocol transfers mechanically — Diffie-Hellman, ElGamal, ElGamal signatures — by reading down a translation table. No new cryptographic idea is required.
  3. The advantage is measured against attacks, not against the group. Curves win because index calculus has no analogue there, so the best known attack is generic at √N.
  4. And the same structure, run over a composite modulus, factors integers — because a curve mod pq is two curves that run out of multiples at different times.

The unification at the end of Section 21.3 is the most striking thing in the chapter. Let the curve degenerate and Lenstra's algorithm becomes the p − 1 method, or trial division. Two methods that look nothing alike are the same algorithm on singular curves.

And the difficulty has moved. Choosing the group is now a solved problem with standardised answers; the risk lives entirely in implementing scalar multiplication without leaking, and in validating what arrives from the network.

Figure (svg): The dictionary translating a discrete-log cryptosystem modulo p into its elliptic curve counterpart.

Not an analogy but a substitution: swap the group and every protocol built on it comes along unchanged.

73. Smaller keys mean elliptic curves are stronger

Trap

The trap

The trap. A 256-bit curve gives the security of a 3072-bit RSA modulus. So curve arithmetic is intrinsically harder to attack, more is packed into each bit, and a curve system is stronger than an RSA system of comparable size.

The size comparison is accurate. The conclusion drawn from it is not.

The fix

The comparison is about attacks, not about strength. Curves need fewer bits because index calculus does not apply to them, so the best known attack is generic. It is a statement about which algorithms exist, and it would evaporate the day an elliptic index calculus was published.

Nothing is proved for either problem. Neither factoring nor discrete logarithms — in any group — has a proof of hardness. Both rest on the failure of sustained attempts, which is evidence rather than certainty.

Against quantum computers the ordering reverses. Shor's algorithm handles both, and a 256-bit curve needs fewer qubits than a 3072-bit modulus. The efficiency that makes curves attractive classically makes them fall sooner quantumly.

And in practice the mathematics is not what fails. Every documented break of a deployed curve system has been an implementation defect — a timing leak, an unvalidated point, a repeated nonce. RSA's failures are the same shape. Neither system is broken by its number theory.

The accurate claim is narrow and still worth having: at present, curves deliver a given level of resistance to known classical attacks using far fewer bits. Efficiency, on current knowledge — not strength, and not permanence.

74. Check: the group law

Check

Work it out before clicking.

Check your understanding

On an elliptic curve, what is (x, y) + (x, −y)?

  • A. (x, 0)
  • B. ∞, the identity element (correct)
  • C. (2x, 0)
  • D. Undefined

Answer: B

Why: The line through the two points is vertical, so its third intersection with the curve is ∞, and reflecting ∞ gives ∞. Hence the two points are negatives of each other, and −(x, y) = (x, −y) — negation on a curve costs one sign flip, unlike inversion mod p.

Why A tempts people
That would be the midpoint in ordinary plane arithmetic. The curve's addition has nothing to do with coordinate averages.
Why C tempts people
Same confusion — adding points is not adding coordinates.
Why D tempts people
The slope formula divides by zero, but that is the signal for the vertical case, not a breakdown. The answer is well defined and it is ∞.

75. Check: why Lenstra beats p − 1

Check

Consider what each method requires.

Check your understanding

What is the advantage of the elliptic curve factoring method over the p − 1 method?

  • A. It is faster per operation
  • B. A new curve gives a fresh point count, so a non-smooth draw can be retried indefinitely (correct)
  • C. It does not require any gcd computations
  • D. It works on numbers with more than two prime factors

Answer: B

Why: The p − 1 method needs p − 1 to be smooth, and p − 1 is fixed by p — if it has a large prime factor there is nothing to change. The curve method needs #E(mod p) to be smooth, and that number varies from curve to curve within Hasse's interval, so each new curve is an independent attempt.

Why A tempts people
It is slower per operation — a curve addition costs several multiplications and an inversion, against one multiplication mod p.
Why C tempts people
Gcd computations are the heart of it; the method detects factors precisely when a gcd comes out nontrivial.
Why D tempts people
Both methods find one factor at a time and both handle more than two factors. That is not the distinction.

76. Check: what the key-size advantage rests on

Check

Consider why 256 bits suffices.

Check your understanding

Why can elliptic curve keys be so much shorter than RSA keys for the same security?

  • A. Curve arithmetic is more complex, so each operation hides more
  • B. No index calculus attack is known for curves, so the best known attack is generic at about √N (correct)
  • C. Elliptic curve discrete logs are proved to be exponentially hard
  • D. Curves use a larger alphabet of operations

Answer: B

Why: Index calculus makes mod-p discrete logs and factoring subexponential, so those moduli must grow rapidly with the security level. Nothing similar is known for curves, so attacks are generic and N ≈ 2²ˢ suffices for s bits — 256 bits for 128-bit security.

Why A tempts people
Complexity of the operation is not security; the group law is completely public and easy to compute.
Why C tempts people
Nothing of the kind is proved. The hardness is an empirical claim resting on decades of unsuccessful attack.
Why D tempts people
There is one operation, point addition. The size of the operation set is not a security parameter.

77. Write the translation table from memory

Connect it up

The table is the chapter's reusable content.

Draw it

Write the eight rows of the mod-p to elliptic-curve dictionary. Then use it to derive elliptic Diffie-Hellman, elliptic ElGamal and the elliptic ElGamal signature from their classical versions, showing at each step which row you used. Beside the signature, write the two-line verification and mark where NA = ∞ is needed. Finish with three lines: what makes a curve unsuitable, why the key-size advantage exists, and what breaks curve systems in practice.

Deriving the protocols from the table rather than memorising them is the point — it is what lets you read a new curve-based scheme and know immediately what it is doing.

78. Exit ticket

Exit ticket

One question, about where the advantage comes from.

Predict first

What single fact makes elliptic curve cryptography practical?

  • Curve arithmetic is faster than modular arithmetic
  • No index calculus analogue exists for curves, so the best known attack is generic and keys can be much shorter
  • Elliptic curve discrete logarithms are proved hard
  • Curves resist quantum computers

Correct: No index calculus analogue exists for curves, so the best known attack is generic and keys can be much shorter

Why: Curve operations are individually slower than modular multiplication, nothing is proved hard, and Shor's algorithm breaks curves faster than RSA. The entire advantage is that index calculus has no elliptic analogue, so attacks are generic at about √N and a 256-bit group gives 128-bit security.

79. What to carry into Chapter 22

Recap

A group built from geometry, and everything that follows from having it.

Chapter 22 next. Pairing-based cryptography: bilinear maps on elliptic curves, which turn the group into something that supports operations no ordinary group can — identity-based encryption, three-party key agreement in one round, and short signatures.

Figure (svg): Comparable security levels for RSA and elliptic curve key sizes, showing the growth gap.

The entire practical case for elliptic curves in one chart, and the reason is an attack that does not exist.

Sources

  1. Introduction to Cryptography with Coding Theory, 3rd edition — Wade Trappe and Lawrence C. Washington — Pearson, 2020 (ISBN 978-0-13-485906-4)
  2. Chapter 21 — Elliptic Curves (sections 21.1-21.7) — Trappe & Washington, 3rd edition, pp. 393-424

Want this taught 1-on-1? Alexander tutors Cryptography — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108