Chapter 25 of Trappe & Washington, the final chapter: the three-polarizer experiment and the mathematics of photon polarization and qubits; quantum key distribution with the 25% error rate that betrays an eavesdropper; what a quantum computer does with a superposition and why measurement is the bottleneck; the discrete and quantum Fourier transforms as period finders; and Shor's algorithm worked end to end on n = 21 — the collapse to a residue class, the peaks at 85, 171, 341 and 427, the continued-fraction step giving r = 6, and the classical gcd finish. Closes with what a quantum computer would and would not break across the whole course.
Subject: Cryptography · 60 slides · diagram-first lesson
Open the interactive version of this deck
Title
Cryptography · Chapter 25
A machine that would break most of this course, and a channel that no machine can break
Objectives
The last chapter, and it points in two directions. Shor's algorithm would undo nearly every public-key system in the book. Quantum key distribution offers a guarantee that no algorithm can touch.
Figure (svg): The three-polarizer experiment: two crossed filters block all light, and inserting a third between them lets light through.
Warm-up
Take three polarising filters — sunglass lenses will do. Set A horizontal and C vertical, and shine a light through both. No light arrives: the crossed filters block everything. Now insert B at 45° between them.
Discussion prompt
Light now reaches the wall. How can adding a filter increase the light that gets through?
Hint: What does a filter do to a photon that passes it?
Answer:
Classically it is impossible. A filter can only remove light, so adding one between two that already block everything should change nothing.
**The resolution is that a filter does not merely select — it changes the photon.** A photon emerging from filter A is horizontally polarised, whatever it was before.
So B does not filter the horizontal photons; it re-prepares them. A horizontal photon meeting a 45° filter passes with probability ½, and if it passes it is now polarised at 45° — no longer horizontal.
And a 45° photon passes the vertical filter with probability ½ too. So the light reaching the wall is ½ × ½ × ½ = one eighth of the original intensity, where without B it was zero.
That is the entire chapter in one experiment. Measurement forces a state into the measured value, and this is not ignorance being resolved — the photon genuinely had no definite vertical polarisation until B was asked.
Section
Section 25.1 · pp. 509-513
Concept
Light is an electromagnetic wave: an electric field travelling orthogonally to a magnetic field. Polarization is the direction the electric field lies in, and there is no constraint on that direction.
A photon's polarization is a unit vector in a two-dimensional complex vector space — though real numbers suffice for everything here. The dot product is (a,b)·(c,d) = a c̄ + b d̄, so the squared length is |a|² + |b|².
Choose a basis, written |↑⟩ and |→⟩ in the physicists' ket notation. An arbitrary polarization is
\[ a| \!\uparrow \rangle + b| \!\rightarrow \rangle, \qquad |a|^2 + |b|^2 = 1 \]
A different orthogonal basis would do equally well, for instance the 45° rotation |↖⟩ and |↗⟩. Nothing distinguishes one basis as the true one, and that fact is what quantum key distribution exploits.
Figure (svg): A qubit as a unit vector in a two-dimensional space, with measurement probabilities given by the squared coefficients.
Concept
A Polaroid filter measures the polarity of a photon. There are exactly two outcomes: aligned with the filter, or perpendicular to it.
If a|↑⟩ + b|→⟩ meets a vertical filter, the probability of coming out vertically polarised is |a|², and the probability of being measured horizontal — and so not passing — is |b|².
And a measurement forces the photon into a definite state. After being measured as |→⟩, the photon is |→⟩ from then on. Measure it again with a horizontal filter and it always passes; with a vertical one it never does.
The 45° case is the important one. Since
\[ | \!\uparrow \rangle = \tfrac{1}{\sqrt2}| \!\nwarrow \rangle + \tfrac{1}{\sqrt2}| \!\nearrow \rangle \]
a vertical photon passes a 45° filter with probability (1/√2)² = ½, and fails with probability ½. A state perfectly known in one basis is completely unknown in the other.
Figure (svg): The two measurement bases, each a rotation of the other, with a state definite in one being maximally uncertain in the other.
Worked example
Now the experiment resolves, step by step.
The source emits randomly polarised light, so half the photons pass filter A
Why: And every one that passes is now in the state |→⟩ — measurement has re-prepared it.
With only A and C, the |→⟩ photons meet a vertical filter
Why: The probability of passing is |b|² with b = 0, so none get through. The wall is dark, as observed.
Insert B at 45°. It measures the |→⟩ photons in the {|↖⟩, |↗⟩} basis
Why: Each passes with probability ½ — a 4:1 reduction from the original beam, counting A's loss too.
And a photon leaving B is now |↗⟩, not |→⟩
Why: This is the crucial step: B did not select horizontal photons, it created diagonal ones.
Verify: the |↗⟩ photons pass the vertical filter with probability ½, giving 1/8 of the original intensity
Why: Zero without B, one eighth with it. There is no way to tell this story with filters that merely select, which is why the experiment is a genuine demonstration and not a puzzle about hidden properties.
Figure (svg): The three-polarizer experiment: two crossed filters block all light, and inserting a third between them lets light through.
Socratic
A natural objection: perhaps each photon always had a definite polarisation in every direction, and the filters simply reveal it.
Discussion prompt
Does that explanation work for this experiment?
Hint: Try to assign every photon definite answers for all three filters in advance.
Answer:
For this experiment alone, a determined sceptic can construct an explanation. The three-filter demonstration shows that measurement changes things, which a sufficiently contrived classical model could mimic.
But the mathematics we need does not depend on settling that. The rules — amplitudes, |a|² probabilities, collapse on measurement — predict the experiment exactly, and they are what the protocols are built on.
The decisive experiments are Bell's, from 1964, and they do rule out hidden properties: correlations between entangled particles violate an inequality that any local hidden-variable theory must satisfy. Those violations have been measured repeatedly, and the 2022 Nobel Prize was awarded for them.
So the honest position for this chapter is: the mathematics is not a convenient description of underlying classical facts. There are no underlying classical facts, and the security argument in Section 25.2 depends on that being true rather than merely useful.
Which is worth pausing on, because it is unlike every other security argument in the course. Every previous guarantee rested on a computational assumption. This one rests on physics being what it appears to be.
Section
Section 25.2 · pp. 513-516
Concept
Take a two-dimensional complex vector space and a pair of orthogonal unit vectors, called |0⟩ and |1⟩ — for instance either of the polarization bases from the last section.
Qubit — A unit vector in that space: a|0⟩ + b|1⟩ with |a|² + |b|² = 1. For the present discussion, a polarised photon.
Measuring in the {|0⟩, |1⟩} basis gives 0 with probability |a|² and 1 with probability |b|², and afterwards the qubit is whichever was observed.
Note what a qubit is not. It is not a bit that is secretly 0 or 1 and merely unknown. The superposition is a real state with observable consequences — the three-filter experiment is precisely such a consequence.
Figure (svg): A qubit as a unit vector in a two-dimensional space, with measurement probabilities given by the squared coefficients.
Concept
Alice and Bob need two channels: a quantum channel carrying photons isolated from the environment, and an ordinary classical channel. Eve may listen to the classical channel and may measure and resend photons on the quantum one.
Two bases are used. B₁ = {|↑⟩, |→⟩} and B₂ = {|↖⟩, |↗⟩}. In either basis, 0 is the first element and 1 the second.
About half the bits survive, and they form a shared secret that can key a conventional cipher.
Figure (svg): The quantum key distribution protocol: Alice and Bob choose bases independently and keep only the bits where they agree.
Worked example
The book's example, checked position by position.
Alice's bits are 0, 1, 1, 1, 0, 0, 1, 0, with bases B₁, B₂, B₁, B₁, B₂, B₂, B₁, B₂
Why: So she sends |↑⟩, |↗⟩, |→⟩, |→⟩, |↖⟩, |↖⟩, |→⟩, |↖⟩.
Bob measures with bases B₂, B₂, B₂, B₁, B₂, B₁, B₁, B₂
Why: Chosen independently and at random.
Comparing the two lists, the bases agree at positions 2, 4, 5, 7 and 8
Why: Five out of eight — close to the expected half.
At those positions Bob measured |↗⟩, |→⟩, |↖⟩, |→⟩, |↖⟩
Why: Where the bases agreed, his measurement is guaranteed to match what Alice sent.
Verify: both hold the string 1, 1, 0, 1, 0
Why: Which becomes the key. In practice they would run far more photons and take, say, the first 128 surviving bits as an AES key. Notice that neither party chose the key: it emerged from two independent random choices, which is a property no classical exchange has.
Figure (svg): The quantum key distribution protocol: Alice and Bob choose bases independently and keep only the bits where they agree.
Worked example
The security argument, and it is a probability calculation rather than a hardness assumption.
Eve must measure the photons to learn anything, and measurement forces each into a definite state
Why: She then resends what she observed. She cannot copy a photon and pass the original along.
Suppose Alice sends |→⟩ and Bob happens to use the matching basis B₁
Why: Without Eve, Bob is certain to measure |→⟩.
If Eve also used B₁ — half the time — the photon passes through unchanged and Bob is correct
Why: No damage done, and Eve learned the bit.
If Eve used B₂, she measures |↖⟩ or |↗⟩ with equal probability, and resends that
Why: Bob then measures that diagonal state in B₁ and gets the right answer only half the time.
\[ \tfrac12 \cdot 0 + \tfrac12 \cdot \tfrac12 = \tfrac14 \]
Verify: so Bob is wrong 25% of the time on bits where he and Alice used the same basis
Why: Alice and Bob sacrifice a sample of their agreed bits and compare them over the classical channel. Errors well above the channel's own noise floor mean Eve. Implementations have run over more than 100 km of ordinary fibre.
Figure (svg): Why eavesdropping is detectable: Eve's measurement collapses the photon and introduces a 25% error rate.
Prediction
Predict first
Why can she not?
Correct: The no-cloning theorem — an unknown quantum state cannot be copied
The no-cloning theorem is what makes the whole protocol work. Every classical eavesdropper copies bits silently; a quantum eavesdropper cannot copy and so must destroy.
It also explains an odd feature of quantum information: it cannot be backed up, forwarded, or amplified. A conventional repeater on a fibre link would break QKD, which is why range is limited and quantum repeaters are an active research problem.
And note the shape of the guarantee. It is not that copying is expensive — it is that no physical process performs it. That is a stronger statement than anything computational hardness can offer.
Why: Copying an unknown quantum state is impossible in principle: any operation that duplicated arbitrary states would have to be both linear and non-linear at once. So Eve must measure, and measuring is exactly what disturbs the state and reveals her.
Definition probe
The protocol uses a quantum channel and a classical one.
Sort into buckets
Sort each item by where it is sent.
Socratic
It offers security from physics rather than from computational assumptions — an apparently unbeatable proposition.
Discussion prompt
What are the practical objections?
Hint: Consider what it needs, what it provides, and what it does not.
Answer:
It needs an authenticated classical channel to begin with. Without authentication Eve conducts a man-in-the-middle attack, measuring and re-sending on both sides. So a shared secret is already needed — which is much of what QKD was meant to establish.
It needs dedicated hardware and dedicated fibre. Photons cannot be amplified or routed through ordinary switches, so range is limited to a few hundred kilometres without trusted relay nodes — and a trusted relay is exactly the assumption QKD was supposed to remove.
It delivers only a key. Everything above it — authentication, signatures, integrity — still needs conventional cryptography, so it replaces one component rather than the stack.
And the theoretical guarantee is about the idealised protocol, not the device. Real implementations have been attacked through detector blinding, timing side channels and imperfect sources — Chapter 14 applies to quantum hardware exactly as it does to smart cards.
The consensus reflected in the standards bodies is that post-quantum algorithms are the practical answer and QKD is a specialised tool for a few high-value links. Both the UK's NCSC and the American NSA have said as much. The physics is genuinely unassailable; the engineering is what limits it.
Section
Section 25.3 · pp. 516-524
Concept
A classical computer takes a binary input and gives a binary output, handling several inputs one at a time.
A quantum computer takes qubits and returns qubits, and the inputs may be linear combinations of basic states. It operates on all the basic states in that combination simultaneously — in effect, a massively parallel machine.
So given a function f, it can accept a superposition of every input and return
\[ \frac1C \sum_x |x, f(x)\rangle \]
Which looks like a complete list of all the values of f. And here is the problem: measuring forces the state into one result. You get |x₀, f(x₀)⟩ for a random x₀ and everything else is destroyed.
You get one look. So the skill in programming a quantum computer is arranging for the outputs you want to appear with much higher probability than the others — and that is exactly what Shor's algorithm does.
Figure (svg): A quantum computer evaluating a function on every input at once, and the measurement that destroys all but one answer.
Concept
Shor's algorithm does not attack n directly. It reduces factoring to period-finding.
Recall from Section 9.4.1: if we can find nontrivial a and r with aʳ ≡ 1 (mod n), we have a good chance of factoring n.
Now consider the sequence 1, a, a², a³, … (mod n). If aʳ ≡ 1 then the sequence repeats every r terms, since a^{j+r} ≡ a^j · a^r ≡ a^j.
So the period of that sequence is exactly the r we want. Measuring the period — or a multiple of it — gives an r with aʳ ≡ 1 (mod n).
The whole quantum content is period-finding. Everything before it is preparation and everything after it is the classical method from Chapter 9. That division is worth holding onto: the quantum computer supplies one number.
Figure (svg): Shor's algorithm as a pipeline from superposition to a candidate period.
Concept
Classically, Fourier transforms find the period of a periodic sequence, and they work here too. For a sequence a₀, …, a_{2ᵐ−1}:
\[ F(x) = \frac{1}{\sqrt{2^m}} \sum_{c=0}^{2^m-1} e^{2\pi i c x / 2^m} a_c \]
Take 1, 3, 7, 2, 1, 3, 7, 2 — length 8, period 4, so frequency 2. The transform gives F(1) = F(3) = F(5) = F(7) = 0, with nonzero values only at 0, 2, 4 and 6.
The reason is cancellation. With ζ = e^{2πi/8}, the sum for √8·F(1) is 1 + 3ζ + 7ζ² + 2ζ³ + ζ⁴ + 3ζ⁵ + 7ζ⁶ + 2ζ⁷, and since ζ⁴ = −1 the terms cancel in pairs.
In general, when the period divides the length, all nonzero values occur at multiples of the frequency. The transform reads the period off directly.
Figure (svg): Discrete Fourier transforms of two periodic sequences, with nonzero values only at multiples of the frequency.
Worked example
The realistic case, and the source of all the difficulty.
Take 1, 0, 0, 1, 0, 0, 1, 0 — length 8, almost period 3
Why: The sequence was cut off before completing its last period.
The absolute value of its transform has peaks at 0, 3 and 5, continuing to 8, 11, 13, 16, …
Why: Not exact multiples of anything, but spaced at an average distance of 8/3.
Dividing the length by that average spacing gives 8/(8/3) = 3
Why: The period, recovered approximately.
A longer example: 1,0,0,0,0,1,0,0,0,0,1,0,0,0,0,1 with 16 terms and period about 5
Why: The peaks are spaced around 3 apart, so the frequency is about 3 and the period about 5.
The peak at 0 is unsurprising: F(0) is just the sum of the sequence divided by the square root of its length.
Verify: peaks occur at approximate multiples of the frequency, and the period is approximately the number of peaks
Why: The mismatch introduces noise — values near zero away from the peaks rather than exactly zero. Everything in Shor's algorithm has to survive that blur, which is why the final step needs continued fractions rather than simple division.
Figure (svg): Discrete Fourier transforms of two periodic sequences, with nonzero values only at multiples of the frequency.
Concept
Choose m so that n² ≤ 2ᵐ < 2n². Start with m qubits all in state |0⟩, then rotate each in turn into an equal superposition:
\[ \frac{1}{\sqrt{2^m}}\bigl(|0\rangle + |1\rangle + |2\rangle + \cdots + |2^m-1\rangle\bigr) \]
Every possible input is now superimposed in a single state, written in decimal for brevity.
Choose a random a with 1 < a < n, assuming gcd(a, n) = 1 — otherwise a factor has already been found. The computer evaluates f(x) = aˣ (mod n) across the whole superposition:
\[ \frac{1}{\sqrt{2^m}}\bigl(|0, a^0\rangle + |1, a^1\rangle + \cdots + |2^m-1, a^{2^m-1}\rangle\bigr) \]
So far this is no better than a classical computer, because measuring the whole system gives one random pair and obliterates the rest.
Figure (svg): Shor's algorithm as a pipeline from superposition to a candidate period.
Worked example
The book's example, which it notes is about the largest that quantum computers running Shor's algorithm have handled.
n = 21, and 21² = 441 ≤ 512 < 882 = 2·21², so m = 9 and 2ᵐ = 512
Why: Nine qubits in the first register.
Choose a = 11. Then 11ˣ mod 21 runs 1, 11, 16, 8, 4, 2, 1, 11, 16, …
Why: Period 6, though we do not yet know that.
The state is the superposition of |x, 11ˣ⟩ over all 512 values of x
Why: A full list of the function's values, inaccessible all at once.
Now measure only the second register, and suppose the result is 2
Why: The system collapses to the terms with 11ˣ ≡ 2 — and nothing else is learned.
Verify: the remaining state is |5⟩ + |11⟩ + |17⟩ + ⋯ + |509⟩, with 85 terms
Why: All ≡ 5 (mod 6). The partial measurement did not destroy the structure — it isolated one residue class, which is a perfectly periodic sequence with the period we want. That is the key manoeuvre: measure just enough to collapse, not enough to lose.
Figure (svg): The n = 21 example: measuring the second register collapses to one residue class, spaced six apart.
Concept
The state |5⟩ + |11⟩ + ⋯ + |509⟩ contains the period, and reading it out is the difficulty.
Measuring now just gives some x with 11ˣ ≡ 2 (mod 21), which is useless on its own.
Two measurements would suffice. With x and y satisfying 11ˣ ≡ 11ʸ we would get 11^{x−y} ≡ 1 (mod 21), and Chapter 9's method would take over.
But the first measurement puts the system into that state, so a second measurement returns the same answer. There is no second sample.
So the period must be extracted from the single state that survives — and the quantum Fourier transform is exactly the tool for that, because it measures frequencies rather than values.
Figure (svg): The n = 21 example: measuring the second register collapses to one residue class, spaced six apart.
Worked example
The transform is defined on basic states and extended by linearity.
\[ \mathrm{QFT}(|x\rangle) = \frac{1}{\sqrt{2^m}} \sum_{c=0}^{2^m-1} e^{2\pi i c x/2^m}|c\rangle \]
Applying it to the collapsed state gives a combination Σ g(c)|c⟩
Why: Where g is the discrete Fourier transform of the indicator sequence 0,0,0,0,0,1,0,0,0,0,0,1,… — a 1 at every x ≡ 5 (mod 6).
The period is 6 and the length is 512, so the frequency is about 512/6 ≈ 85.33
Why: Not an integer, so the peaks will blur.
The peaks appear at c = 0, 85, 171, 256, 341, 427
Why: Approximate multiples of 85.33, exactly as the noisy examples predicted.
Around the peak at 341 the values run 0.305, 0.439, 0.773, 3.111, 1.567, 0.631, 0.398, 0.291
Why: A sharp spike with a small skirt. At c = 0 and 256 the value is 3.757 with neighbours around 0.015.
Verify: the probability of measuring one of 85, 171, 341, 427 is about 0.456
Why: Since probability is proportional to |g(c)|², and 3.111² divided by the total is about 0.114 for each of the four. Nearly a coin flip that the measurement lands on something useful — and if it does not, the algorithm simply restarts.
Figure (svg): The Fourier transform of the collapsed state, with sharp peaks at approximate multiples of 512/6.
Worked example
Suppose the measurement gives c = 427.
We expect 427 ≈ j·f₀ for some j, where r·f₀ ≈ 2ᵐ = 512
Why: The peaks are at approximate multiples of the fundamental frequency.
Dividing: 427/512 ≈ j/r, and 427/512 ≈ 0.834
Why: A rational approximation problem — recover a fraction with small denominator.
The continued fraction of 427/512 is [0; 1, 5, 42, 2], with convergents 0, 1, 5/6, 211/253, 427/512
Why: Chapter 3's algorithm, applied to a measured number.
Since r must be less than n = 21, take the last denominator below 21: r = 6
Why: The convergent 5/6 — and 0.834 ≈ 5/6 = 0.833, an excellent fit.
Verify: check 11⁶ ≡ 1 (mod 21) ✓
Why: Shor showed there is a high chance that |c/2ᵐ − j/r| < 1/2n², and continued fractions find the unique j/r with r < n satisfying that. If the check fails, restart with a new a — the algorithm is probabilistic and cheap to repeat.
Figure (svg): Continued fractions turning a measured frequency into the period.
Worked example
With r = 6 in hand, Chapter 9's method takes over and no quantum mechanics is involved.
Write r = 2ᵏm with m odd: 6 = 2 · 3
Why: So m = 3.
Compute b₀ ≡ a^m ≡ 11³ ≡ 1331 ≡ 8 (mod 21)
Why: Since 1331 = 63 · 21 + 8.
Square successively: b₁ ≡ 8² ≡ 64 ≡ 1 (mod 21)
Why: We have reached 1, so b₀ = 8 is the last value that was not 1.
Compute gcd(b₀ − 1, n) = gcd(7, 21) = 7
Why: A nontrivial factor.
Verify: 21 = 7 × 3
Why: Done. If the gcd had come out trivial, the procedure restarts with a new a — and a few attempts almost always suffice. Note how little of the work was quantum: one number, r, came from the quantum computer, and everything else is Chapter 9.
Figure (svg): The classical finish: from the period to a factor of 21.
Concept
Shor's algorithm factors in probabilistic polynomial time, and the same machinery solves discrete logarithms.
What survives: symmetric ciphers and hashes lose half their effective key length to Grover's algorithm, so AES-256 still offers 128 bits. Lattice systems from Chapter 23 and McEliece from Chapter 24 have no known quantum attack. And the one-time pad is untouched, since its security is information-theoretic.
The book's caveat, still fair: quantum computers are not yet a reality, current versions handle only a few qubits, and n = 21 is about the limit for Shor's algorithm as implemented.
But the threat is present tense for long-lived secrets. Traffic recorded today can be decrypted whenever the machine arrives, which is why the migration described in Chapter 23 is already under way.
Figure (svg): What a quantum computer would and would not break across the course.
Notation
One formula does the work that no classical step in the algorithm could.
Annotate
On: \( \mathrm{QFT}(|x\rangle) = \frac{1}{\sqrt{2^m}}\sum_{c=0}^{2^m-1} e^{2\pi i cx/2^m}|c\rangle \)
This is the pattern of nearly every quantum algorithm: arrange interference so that wrong answers cancel and right ones reinforce, then look once.
Prediction
Predict first
What would measuring the whole state give?
Correct: One random pair |x₀, a^{x₀}⟩ — everything else destroyed, and no period
Measuring the second register alone collapses far less. It fixes the function's value but leaves a superposition over all the x's that produce it — and those x's form an arithmetic progression with the period as its common difference.
So the partial measurement is constructive. It throws away the parts of the state that were not periodic and keeps one clean progression, which is precisely the input the Fourier transform needs.
And note that the measured value itself is discarded. We never care that the result was 2; only that it isolated a residue class. The information used is the structure that survives, not the number obtained.
This is the characteristic move of quantum algorithm design: measure the minimum needed to shape the state, and never more.
Why: A full measurement collapses to a single basic state, giving one x and one value of aˣ. That is exactly what a classical computer would produce with one evaluation, and the parallelism is wasted.
Anomaly
The peaks were at 0, 85, 171, 256, 341 and 427, and the measurement gives 256.
Predict first
What happens?
Correct: 256/512 = 1/2 gives r = 2, which fails the check 11² ≡ 1 (mod 21), so the run is discarded
c = 0 is similarly useless, giving 0/512 and no information — and both c = 0 and c = 256 are among the largest peaks, at 3.757.
Which is why the success probability is quoted as about 0.456 rather than the full weight of all six peaks: only four of them yield a usable j/r.
The verification step is essential and cheap. Checking aʳ ≡ 1 (mod n) is one modular exponentiation, so bad runs are detected immediately and discarded.
And this is normal for the algorithm. It is probabilistic in several places — the choice of a, the collapse, the measurement of the transform, and whether Chapter 9's method yields a nontrivial gcd. A handful of repetitions is expected, and each is cheap.
Why: 256/512 is exactly 1/2, whose continued fraction gives r = 2. But 11² = 121 ≡ 16 (mod 21), not 1, so the candidate fails verification and the algorithm restarts with a new a or a new run.
Definition probe
The algorithm mixes quantum and classical work.
Sort into buckets
Sort each step.
Faded example
Five blanks.
Fill in the blanks
Choose m with n² ≤ 2ᵐ < 2n², superpose all inputs, and compute f(x) = aˣ mod n across the superposition. Measure the second register, collapsing to one residue class. Apply the quantum Fourier transform to that state and measure, obtaining c. Use continued fractions on c/2ᵐ to recover r, check that aʳ ≡ 1 (mod n), and finish with the method of Section 9.4.1.
Why: Note how much of it is classical. Only the superposition, the parallel evaluation and the transform need a quantum computer; the rest is Chapters 3 and 9.
Trade off
This chapter contains a threat and a defence. Fill the blanks.
Comparison matrix
| Shor's algorithm | Quantum key distribution | |
|---|---|---|
| What it needs | a large fault-tolerant quantum computer | a quantum channel and dedicated hardware |
| Status | n = 21 is about the record | deployed over 100+ km of fibre |
| What it delivers | breaks RSA, discrete logs, elliptic curves | a shared key, with eavesdropping detectable |
| Security rests on | n/a — it is an attack | the laws of physics, not on hardness |
| Main limitation | the machine does not exist yet | needs an authenticated classical channel to start |
The asymmetry is worth noticing: the attack needs technology nobody has, and the defence needs technology that exists but does not scale. Neither is the practical answer today, which is why post-quantum algorithms are.
Discrimination
Six systems from the course.
Sort into buckets
Sort each.
Edge cases
The book says the first full-scale machine is probably many years off, and that n = 21 is about the record for Shor's algorithm.
Discussion prompt
What actually stands in the way?
Hint: Count the qubits needed and think about errors.
Answer:
**Factoring RSA-2048 needs several thousand logical qubits**, and logical qubits are built from many physical ones because physical qubits are noisy. Current estimates run to millions of physical qubits.
Error correction is the central obstacle. Quantum states decohere, and quantum error correction — a direct descendant of Chapter 24's ideas, adapted to a setting where you cannot copy or freely measure — costs a large overhead.
Machines with hundreds of physical qubits exist and are not close. The gap between a demonstration factoring 21 and one factoring a 2048-bit modulus is many orders of magnitude, not a few years of refinement.
But the timeline is not the whole risk. Store-now-decrypt-later means traffic recorded today is exposed whenever the machine appears, so anything needing decades of confidentiality must migrate before it exists.
Which is why the standards bodies acted in advance. NIST's post-quantum algorithms were finalised in 2024 and hybrid deployment is already running in TLS — built, as Chapter 23 described, on lattices, with McEliece from Chapter 24 as the conservative alternative.
Ranking
Five quantum tasks.
Put in order
Why: QKD needs single photons and detectors — deployed commercially, and not a computer at all. Factoring 21 has been demonstrated. A 512-bit number would need thousands of logical qubits. RSA-2048 needs several thousand more. And Grover against AES-256 needs about 2¹²⁸ sequential quantum operations, which is not merely a hardware problem but a running-time one — it will not happen.
Constraint
A national archive encrypts records that must remain confidential for seventy-five years, and asks about quantum risk.
Discussion prompt
What is the advice?
Hint: Start from the confidentiality horizon, not from the technology.
Answer:
Seventy-five years is far beyond any plausible estimate for a working quantum computer, so the risk must be treated as certain rather than speculative. This is one of the clearest cases for early migration.
Migrate key establishment first, because recorded ciphertext is the exposure. Signatures matter less: forging one requires a machine at the time of forgery, which cannot be done retroactively.
Use hybrid key exchange — a classical scheme combined with ML-KEM — so that a break of either alone is not enough. SIKE's classical collapse in 2022 is the argument for not trusting a new scheme alone.
Consider Classic McEliece for the archive specifically. Its assumption dates to 1978 and has never been broken, and its enormous key is affordable when keys are static and long-lived — which is exactly an archive's situation, and unlike a TLS handshake.
And use AES-256 rather than AES-128 for the bulk encryption, since Grover halves the effective key length. That is a one-line change that removes the symmetric side of the problem entirely.
Cost model
One number decides whether QKD works.
Annotate
On: \( \Pr[\text{Bob wrong}] = \tfrac12 \cdot 0 + \tfrac12 \cdot \tfrac12 = \tfrac14 \)
A security guarantee that is a probability calculation from physics rather than an assumption about computational difficulty — the only one of its kind in this course.
Missing information
A product is advertised as quantum-safe.
Discussion prompt
What needs checking?
Hint: Which algorithms, which layer, and what assumption.
Answer:
Which algorithms, and for what. Key exchange and signatures are separate problems with separate urgencies, and a product may have migrated one and not the other.
Whether it is hybrid. A post-quantum scheme alone is a bet on a young assumption. Combining it with a classical one costs little and protects against a classical break of the new scheme — which has already happened once, to SIKE.
Whether the symmetric layer was raised. Grover halves effective key lengths, so AES-128 becomes 64-bit-equivalent against a quantum attacker. Post-quantum key exchange over AES-128 is a half-migration.
Whether it means QKD or post-quantum algorithms. These are entirely different technologies with different requirements, and marketing frequently blurs them. QKD needs dedicated fibre; post-quantum algorithms are a software change.
And whether long-stored data was re-encrypted. Migrating new traffic does nothing for ciphertext an adversary recorded years ago, which is the actual threat for anything with a long confidentiality horizon.
Two truths and a lie
Two of these claim too much.
Eliminate the wrong options
Which statement is correct?
Survives elimination: a
Why: The guarantee is detection, not prevention, and it is conditional on the classical channel being authenticated. Both conditions are real: the physics is unassailable and the protocol around it has ordinary requirements.
Commit first
An organisation encrypts traffic with TLS using elliptic curve key exchange and AES-256.
Predict first
What should they actually worry about?
Correct: Traffic recorded now being decrypted in fifteen or twenty years
AES-256 is fine. Grover reduces it to 128-bit equivalent security, which is ample, and the algorithm's sequential structure makes even that impractical.
So the migration priority is exactly one thing: key establishment, for traffic whose confidentiality must outlast the arrival of a quantum computer.
And the answer is already deployed. Hybrid X25519 + ML-KEM runs on a large share of Chrome's TLS connections, which means most readers have already used post-quantum cryptography without noticing.
Why: No machine capable of breaking a 256-bit curve exists or is close. But recorded ciphertext keeps indefinitely, and the elliptic curve key exchange protecting it today would fall to a machine built at any point in the future. That is the store-now-decrypt-later risk, and it is the only quantum risk that is live right now.
Explain it
A colleague understands that RSA rests on factoring being hard, and asks how a quantum computer changes that.
Discussion prompt
Explain it without quantum mechanics.
Hint: Start with the classical reduction.
Answer:
Start with a classical fact they may not know: factoring n is easy if you can find the period of the sequence 1, a, a², a³, … mod n. That reduction is ordinary number theory and has nothing to do with quantum mechanics.
So the whole problem becomes: find the period of a sequence you can compute but not tabulate. Classically that means computing terms one at a time, and there are far too many.
A quantum computer computes all the terms at once, in a single superposed state. That alone is useless, because looking at the state gives one random term and destroys the rest.
The trick is the Fourier transform. It converts a periodic pattern into sharp peaks at the frequency, and it can be applied to the whole state in one operation. Then one measurement lands on a peak about half the time, and a peak tells you the frequency, which gives the period.
Then it is ordinary arithmetic again: the period gives a number r with aʳ ≡ 1 mod n, and a gcd finishes the factorisation. So the quantum machine contributes exactly one number, and everything else is the mathematics they already know.
Figure (svg): Shor's algorithm as a pipeline from superposition to a candidate period.
Explain it to yourself
The collapsed state is a sum over x ≡ 5 (mod 6), and the transform produces peaks at multiples of about 512/6.
Discussion prompt
Explain the mechanism, in terms of the exponentials.
Hint: What happens when you sum roots of unity along an arithmetic progression?
Answer:
Each term contributes e^{2πicx/512}, a point on the unit circle. Summing over an arithmetic progression of x means stepping around the circle by a fixed angle each time.
If the step lands the points evenly around the circle, they cancel and the sum is near zero. That happens for most values of c.
But if the step is a whole number of turns, every term points the same way and they add — giving a large peak. That happens when c is close to a multiple of 512/6.
Since 512/6 is not an integer, the alignment is never perfect, so the peaks are large rather than infinite and have a small skirt around them. That is the blur, and it is why continued fractions are needed afterwards.
So the transform is a resonance test. It asks, for each frequency c, whether the state's periodicity is in step with it — and only frequencies matching the hidden period pass. The period was in the state all along; the transform makes it the thing most likely to be measured.
Figure (svg): The Fourier transform of the collapsed state, with sharp peaks at approximate multiples of 512/6.
Concept
The protocol needs both, and confusing their roles is the commonest misunderstanding of it.
The quantum channel carries polarised photons isolated from the environment. Only Alice sends on it, only once per photon, and nothing sent on it can be copied.
The classical channel carries ordinary messages that Eve may read in full. Bob's basis choices and Alice's replies go here, and neither reveals a bit.
Announcing bases is safe because a basis says nothing about the value measured in it. Announcing a bit would of course be fatal, which is why the sacrificed error-checking sample is discarded afterwards and never reused.
But the classical channel must be authenticated. Otherwise Eve impersonates Bob to Alice and Alice to Bob, running two protocols and relaying between them — the man-in-the-middle attack of Chapter 10, unchanged by any amount of quantum mechanics.
Which is the honest statement of what QKD provides: given a way to authenticate, it produces a shared key whose secrecy is guaranteed by physics. It does not remove the need for that authentication, and a small shared secret is usually what supplies it.
Figure (svg): The quantum key distribution protocol: Alice and Bob choose bases independently and keep only the bits where they agree.
Concept
Unlike the quantum computer, this technology exists and is deployed.
The book notes implementations over more than 100 km of conventional fibre optic cable, and that range has since been extended considerably, including via satellite links.
The range limit comes from no-cloning, which is the same property that provides the security. Photons cannot be amplified by a repeater without being measured, so loss accumulates and cannot be undone.
Trusted-node relays extend the range by decrypting and re-encrypting at each hop — which works, and reintroduces exactly the trusted intermediary the technology was meant to avoid. Genuine quantum repeaters remain a research problem.
Deployments exist in banking and government networks, and China's Micius satellite demonstrated intercontinental key distribution. But the volume is tiny compared with conventional cryptography, and the reasons are engineering rather than physics.
Figure (svg): Why eavesdropping is detectable: Eve's measurement collapses the photon and introduces a 25% error rate.
Socratic
Bob broadcasts which basis he used for every photon, and Eve hears all of it.
Discussion prompt
Why is that safe, and what would happen if the announcement came earlier?
Hint: Consider what Eve could do with the basis before the photon arrives.
Answer:
A basis is not a bit. Knowing Bob measured photon 4 in B₁ says nothing about whether he got 0 or 1 — both are equally likely from Eve's position.
And by the time it is announced, the photons are gone. Eve cannot go back and re-measure them, because measurement is destructive and she had only one chance.
If the bases were announced in advance, the protocol would collapse completely. Eve would measure every photon in the correct basis, learn every bit, and resend faithfully — introducing no errors at all.
So the ordering is the security, exactly as in Chapter 19's commit-challenge-respond: the choice must be made before the adversary learns anything that would let them exploit it.
Which is worth noticing as the last instance of a pattern. Nearly every protocol in this book depends on someone committing before someone else speaks, and the reason is always the same — that is how you manufacture the effect of simultaneity over a sequential channel.
Anomaly
After comparing bases, Alice and Bob hold what should be the same string. They compare a sample and find 3% disagreement, well below the 25% Eve would cause.
Predict first
What should they conclude?
Correct: Channel noise — but they must still correct the errors and reduce Eve's possible knowledge
Information reconciliation comes first: Alice and Bob exchange parity information over the classical channel to locate and fix the mismatches, at the cost of leaking a little to Eve.
Then privacy amplification. They compress the reconciled string with a hash so that Eve's partial knowledge — from noise-disguised eavesdropping plus the reconciliation leakage — is reduced to a negligible amount. A shorter key that Eve knows nothing about beats a longer one she partly knows.
Both steps are classical, and they are what turns a raw sifted string into a usable key. The textbook protocol stops before them, which is why real QKD systems are considerably more involved than the five steps described.
And this is the last appearance of a familiar shape: the cryptographic core is small and clean, and most of the engineering is in the machinery around it.
Why: Real fibre and detectors introduce errors, so a small discrepancy is expected. But the errors must be reconciled — otherwise the keys differ and nothing encrypted with them decrypts — and some of the disagreement could in principle be a partial eavesdropper rather than noise.
Estimation
Shor's algorithm has factored 21.
Predict first
Roughly how many physical qubits would RSA-2048 need on current error-correction estimates?
Correct: About 20 million
The gap is many orders of magnitude, not a few years of refinement. That is why the book's caution — a full-scale machine is probably many years off — still reads accurately.
But the estimates have fallen repeatedly as error-correction schemes improve, which is the reason nobody treats the timeline as safe. A factor-of-ten improvement in overhead moves the date substantially.
And note where quantum error correction came from. It is a direct descendant of Chapter 24, adapted to a setting where you cannot copy a state or measure it freely — which makes it a much harder problem than the classical case.
Why: Several thousand logical qubits are needed, and each logical qubit currently costs a thousand or more physical ones because of error correction. Published estimates cluster in the millions to tens of millions. Machines today have hundreds to low thousands of noisy physical qubits.
Matching
Shor's algorithm reuses a great deal of the course.
Match the pairs
Why: Only the superposition and the Fourier transform are new. The reduction to period-finding, the rational approximation, and the machinery that would make a quantum computer work at all are all things the course has already covered — which is a fair summary of how the last chapter relates to the rest.
Definition probe
The course has produced two kinds of security claim.
Sort into buckets
Sort each.
Error analysis
From a vendor datasheet.
Annotate
The first is the instructive one: a chain of security is not stronger than its authentication, and adding an unconditionally secure component beneath a quantum-breakable one buys nothing at all.
Pattern
The chapter's two halves rest on the same two facts about measurement.
And the two applications point in opposite directions. Shor's algorithm exploits the parallelism to break the hardness assumptions this course was built on. Quantum key distribution exploits the fragility of measurement to offer a guarantee that no computation can defeat.
Both are the same physics, and the difference is only which property you build on — which is the last instance of a pattern that has run through the whole book: the structure an attacker exploits and the structure a designer needs are usually the same structure.
Figure (svg): A quantum computer evaluating a function on every input at once, and the measurement that destroys all but one answer.
Trap
The trap. Shor's algorithm factors in polynomial time and solves discrete logarithms. RSA, Diffie-Hellman, ElGamal, DSA and elliptic curves all fall. That is nearly every public-key system in this book, so a quantum computer would end cryptography as we know it.
The list is accurate and the conclusion is much too broad.
Symmetric cryptography survives. Grover's algorithm gives only a square-root speedup, so AES-256 retains 128 bits of security and hash functions degrade similarly. Doubling the key length is the entire response, and it is a configuration change.
Several public-key families survive too. Lattice problems from Chapter 23 and the McEliece system from Chapter 24 have no known quantum attack — which is why NIST's standards are built on them and why deployment has already begun.
And the one-time pad is untouchable. Chapter 20 proved H(P|C) = H(P), an information-theoretic guarantee that holds against an adversary of unlimited power, quantum or otherwise. Its problems were always key distribution, never the mathematics.
The machine also does not exist. The book's n = 21 remains near the record, and breaking RSA-2048 needs error-corrected qubits by the million. The threat is real for long-lived secrets and is not present-tense for a session key.
The accurate claim is narrow: a large fault-tolerant quantum computer would break the public-key systems that rest on factoring or discrete logarithms, leave symmetric cryptography needing longer keys, and not touch lattice, code-based or information-theoretic security. Which is a serious problem with a known solution already being deployed — not the end of the subject.
Check
Work it out before clicking.
Check your understanding
In the three-polarizer experiment, why does inserting filter B at 45° let light through?
Answer: B
Why: A photon leaving filter A is horizontal, and a horizontal photon never passes a vertical filter. But measurement at 45° changes the state: half the photons pass B and emerge diagonal, and a diagonal photon passes the vertical filter half the time. The final intensity is 1/8, up from 0.
Check
Consider what Eve's measurement does.
Check your understanding
If Eve measures and resends every photon, what error rate does Bob see on bits where his basis matched Alice's?
Answer: C
Why: Eve picks the right basis half the time, causing no error. The other half she collapses the photon into the wrong basis, and Bob is then wrong half the time. So the total is ½ · 0 + ½ · ½ = ¼ — bits that should have agreed perfectly disagree a quarter of the time.
Check
Consider what the quantum part actually delivers.
Check your understanding
What does the quantum portion of Shor's algorithm produce?
Answer: B
Why: The quantum computer produces one measured value c near a multiple of 2ᵐ/r. Continued fractions turn c/2ᵐ into j/r and give r, and then the entirely classical method of Section 9.4.1 turns aʳ ≡ 1 (mod n) into a factor.
Connect it up
One threat, one defence, both from the same physics.
Draw it
Write the QKD protocol in five steps and, beside it, the calculation ½·0 + ½·½ = ¼ with a sentence on what each half means. Then write Shor's algorithm in six steps, marking which need a quantum computer and which are Chapters 3 and 9. Finish with two lists: what a quantum computer breaks, and what survives — and one line saying why the survivors survive.
The last line is the one worth being able to say: Shor exploits hidden periodic structure, and systems whose hardness has no such structure — lattices, codes, symmetric ciphers — are untouched by it.
Exit ticket
One question, on where the security in this chapter comes from.
Predict first
What makes quantum key distribution different from every other scheme in this course?
Correct: Its security rests on the laws of physics rather than on a computational hardness assumption
Why: Every other system's security rests on a problem being hard — factoring, discrete logs, lattice reduction — and any of those could fall to a better algorithm. QKD's guarantee is that measurement disturbs a quantum state, which no algorithm can change. Its limitations are practical: dedicated hardware, limited range, and the need for an authenticated classical channel.
Recap
The last chapter, pointing both ways.
And that closes the book. It began with the shift cipher and frequency analysis, and ends with a machine that would undo most of what came between — and with the systems already being deployed to replace them. The pattern held throughout: a construction's structure is what makes it possible and what an attacker works with, and the security of a system is only ever as precise as the claim you can state about it.
Figure (svg): What a quantum computer would and would not break across the course.
Want this taught 1-on-1? Alexander tutors Cryptography — $55/session, free consultation.