Chapter 24: Error Correcting Codes

Chapter 24 of Trappe & Washington, the book's longest: repetition, parity, two-dimensional parity, Hamming, ISBN and Hadamard codes with their code rates; the general theory of (n, M, d) codes, Hamming distance and the detect-s / correct-t thresholds; the Singleton, sphere-packing and Gilbert-Varshamov bounds with MDS and perfect codes; linear codes with generator and parity check matrices, syndrome decoding and duals; Golay, cyclic, BCH and Reed-Solomon codes; and the McEliece cryptosystem, whose trapdoor is an efficient decoder for a disguised Goppa code — with every numeric example verified.

Subject: Cryptography · 91 slides · diagram-first lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Error Correcting Codes

Title

Cryptography · Chapter 24

Adding redundancy so a noisy channel can be undone — and turning a decoder into a trapdoor

2. What you will be able to do

Objectives

This chapter is not about hiding information but about preserving it. The last section shows why it belongs in a cryptography book: an efficient decoder that only you possess is a public-key trapdoor.

Figure (svg): The coding pipeline: a message is encoded to a codeword, corrupted by a noisy channel, and decoded back.

Every code in the chapter is a different answer to one question: what redundancy is worth its bandwidth?

3. How do you talk in a noisy room?

Warm-up

Two options: raise your voice, or repeat yourself. The second is the one that concerns us.

Discussion prompt

What does repeating yourself actually accomplish, stated precisely?

Hint: Think about what the listener can do that they could not before.

Answer:

It adds redundancy, so that the message can be reconstructed even when part of what arrives is wrong. The information was already there; the repetition makes it recoverable.

And it separates two abilities. With two repetitions the listener can tell that something went wrong — the two copies disagree — but cannot tell which is right. With three, a majority vote decides.

**Those are error detection and error correction**, and they are genuinely different requirements. Detection is cheap; correction costs roughly twice as much.

The whole chapter is about buying correction more cheaply than repetition does. Repeating a 4-bit message three times costs 12 bits and corrects one error. The Hamming code does the same job in 7.

And the cryptographic twist comes at the end. If correcting errors requires a decoder that only you have, then deliberately-introduced errors become an encryption you alone can undo.

4. Introduction

Section

Section 24.1 · pp. 461-472

5. Repetition Code

Concept

The alphabet is {A, B, C, D} and the channel corrupts each symbol with probability p = 0.1. Sending C alone gives a 90% chance of success — too risky.

So send CCC and take the majority. The message is recovered if all three arrive correctly, or if exactly one is wrong:

\[ (0.9)^3 + 3(0.9)^2(0.1) = 0.972 \]

Detection and correction differ here. If CBC arrives, that could be one error from CCC or two from BBB — so with up to two errors we detect but cannot correct. With at most one error we correct.

And two repetitions would not do. From CB you cannot tell whether BB or CC was sent, so two copies detect one error and correct none.

Note that codes can use any alphabet. Replacing A, B, C, D by 00, 01, 10, 11 gives the binary codewords 000000, 010101, 101010, 111111 — and most codes in practice use binary strings or integers mod a prime.

Figure (svg): The three-fold repetition code correcting a single error by majority vote.

The crudest possible code, and it already shows the two ideas: detection needs distance, correction needs more of it.

6. Parity Check

Concept

Send seven bits, then an eighth chosen so the total number of 1s is even.

  1. 0110010 becomes 01100101
  2. 1100110 becomes 11001100

A single error is discovered immediately, because the received word has an odd number of 1s.

But it cannot be located. An error in any of the eight positions produces the same symptom, so the only sensible response is to ask for a retransmission.

The code rate is 7/8 — very efficient, and it buys detection only. This is the cheapest possible non-trivial code, and it is still used everywhere: in serial protocols, in memory parity, and as a component of larger schemes.

7. Two-dimensional parity locates the error

Worked example

The same idea in two directions gives correction rather than mere detection.

Arrange 20 data bits into a 4 × 5 array

Why: The bits 10011011001100101011, read row by row.

Add a parity bit to each row and each column

Why: Giving a 5 × 6 array. The bottom-right corner is the parity of the column-parity bits — and it comes out the same either way.

Suppose the bit in row 3, column 4 is corrupted in transmission

Why: The receiver rebuilds the array and recomputes both sets of parities.

Row 3's parity is odd and column 4's parity is odd; every other check passes

Why: Two failing checks, one intersection.

Verify: the error is at row 3, column 4 — located exactly, so it can be flipped

Why: Two errors are detected but not located: errors in positions 2 and 3 of row 2 produce the same column-parity failures as errors in positions 2 and 3 of row 5. The pattern of failures no longer has a unique intersection.

Figure (svg): A two-dimensional parity array with row and column checks locating a single bit error at their intersection.

Two one-dimensional detectors crossed to make a two-dimensional locator — the same idea as a syndrome, in the simplest case.

8. Hamming Code: the [7,4] example

Concept

The first code in the chapter that corrects an error efficiently. Messages are blocks of four bits, multiplied mod 2 on the right by

\[ G = [I_4 \mid P] = \begin{pmatrix} 1&0&0&0&1&1&0 \\ 0&1&0&0&1&0&1 \\ 0&0&1&0&0&1&1 \\ 0&0&0&1&1&1&1 \end{pmatrix} \]

The first four columns are the identity, so the first four output bits are the message itself. The remaining three provide the redundancy.

Decoding uses H = [Pᵀ | I₃]. Multiply the received vector by Hᵀ. A zero result means no correction is needed; a nonzero result is the transpose of a column of H, and that column's position is the bit in error.

Figure (svg): The Hamming [7,4] generator and parity check matrices, with a syndrome pointing at the corrupted bit.

The syndrome is not a hint about the error — it is the error's address, written in binary.

9. Encoding and correcting with the [7,4] code

Worked example

One message, one error, verified end to end.

The message 1100 encodes as (1,1,0,0)·G = 1100011

Why: The first four bits are the message; the last three are 0, 1, 1.

Suppose 1100001 is received — the sixth bit has flipped

Why: Two bits differ from what was sent, in the sense that the receiver does not know which.

Compute (1,1,0,0,0,0,1)·Hᵀ = (0, 1, 0)

Why: Three inner products mod 2.

(0, 1, 0) is the transpose of the 6th column of H

Why: Every column of H is distinct, so the match is unique.

Verify: flip bit 6 to recover 1100011, and read off the message 1100

Why: Had the syndrome been (0,0,0) there would have been no error to correct. The procedure looks mysterious here and is explained completely in Section 24.5 — the columns of H are the binary numerals for the positions.

Figure (svg): The Hamming [7,4] generator and parity check matrices, with a syndrome pointing at the corrupted bit.

The syndrome is not a hint about the error — it is the error's address, written in binary.

10. Code rate

Concept

The Hamming code is a large improvement on repetition, and the comparison is worth making precise.

Hamming [7,4]: four information bits are sent as seven, detecting up to two errors and correcting one.

Repetition: to achieve the same, each of the four bits must be sent three times — twelve bits.

\[ R_{\text{Hamming}} = \tfrac47 \approx 0.57 \qquad \text{versus} \qquad R_{\text{repetition}} = \tfrac{4}{12} = \tfrac13 \]

Higher is generally better, as long as too much correcting capability is not lost. Sending a message unencoded has rate 1 and is useless in a noisy channel.

So every code is a point on a trade-off curve, and Section 24.3 gives the bounds that say which points are even possible.

Figure (svg): Code rates of the chapter's examples, from the wasteful repetition code to the efficient Hamming code.

Higher is better, as long as too much correcting power is not lost — and that trade is the whole engineering problem.

11. ISBN Code

Concept

A ten-digit codeword assigned to each book — the first edition of this book was 0-13-061814-4. The first digit gives the language, the next two the publisher, the next six an identifier, and the tenth is a check digit chosen so that

\[ \sum_{j=1}^{10} j \, a_j \equiv 0 \pmod{11} \]

The modulus is 11, not 10. The first nine digits come from 0–9, but a₁₀ may be 10, written X.

The receiver computes S = Σ j·xⱼ mod 11. If S ≡ 0, no error is detected; otherwise an error has occurred — though it cannot be corrected.

Books published from 2007 use a 13-digit ISBN with a slightly different sum, working mod 10.

Figure (svg): The ISBN weighted checksum detecting a single error and a transposition.

Two of the commonest human errors caught by one weighted sum — and it works because the modulus is prime.

12. Why ISBN catches a transposition

Worked example

Two error types are guaranteed to be caught, and each needs a line of algebra.

A single error at position k: xₖ = aₖ + e with e ≠ 0

Why: Every other digit is correct.

Then S = Σ j·aⱼ + ke ≡ ke (mod 11)

Why: The correct part sums to 0 by construction.

Since 11 is prime and 1 ≤ k ≤ 10, ke ≢ 0 whenever e ≢ 0

Why: So the error is detected. This is exactly where the prime modulus is needed — mod 10, k = 2 and e = 5 would give ke = 10 ≡ 0.

A transposition swaps aₖ and aₗ, so xₗ = aₖ and xₖ = aₗ

Why: One of the commonest errors when copying numbers by hand.

\[ S \equiv (k - l)(a_l - a_k) \pmod{11} \]

Verify: both factors are nonzero when the digits differ, so the product is nonzero mod 11

Why: The check is chosen deliberately for the errors humans actually make: mistyping one digit and swapping two adjacent ones. A code is only as good as its match to the error model.

Figure (svg): The ISBN weighted checksum detecting a single error and a transposition.

Two of the commonest human errors caught by one weighted sum — and it works because the modulus is prime.

13. Hadamard Code

Concept

Used by the Mariner spacecraft in 1969 to send pictures back to Earth. There are 64 codewords: the 32 rows of a 32 × 32 matrix H, and the 32 rows of −H.

H is built from binary expansions. Number rows and columns 0 to 31; writing i = a₄a₃a₂a₁a₀ and j = b₄b₃b₂b₁b₀ in binary,

\[ h_{ij} = (-1)^{a_0 b_0 + a_1 b_1 + \cdots + a_4 b_4} \]

For example i = 31 and j = 3 give exponent 2, so h₃₁,₃ = 1.

The rows are mutually orthogonal: the dot product of two distinct rows is 0, and of a row with itself is 32.

Each pixel's darkness was a 6-bit number, mapped to one of the 64 codewords and transmitted as 32 values of ±1.

Figure (svg): The Hadamard code's decoding by dot products, with the correct row standing out from the rest.

Correction by correlation rather than by algebra, and it is why a code with a dreadful rate flew to Mars.

14. Decoding by correlation

Worked example

The Hadamard decoder does not do algebra — it measures which row the received signal resembles.

Take the dot product of the received vector with each of the 32 rows of H

Why: Thirty-two numbers.

With no errors, 31 of them are 0 and one is ±32

Why: Orthogonality. +32 identifies a row of H; −32 the corresponding row of −H.

With one error, all are ±2 except one, which is ±30

Why: Each flipped entry moves every dot product by 2, so the gap between the right answer and the rest shrinks slowly.

With seven errors, all lie between −14 and 14 except one, between 16 and 30 in absolute value

Why: Still unambiguous, so 7 errors are corrected.

Verify: with eight errors the ranges overlap and correction fails

Why: Detection still works up to 15 errors, since 16 changes are needed to reach another codeword — which is the minimum distance. This is a (32, 64, 16) code with rate 6/32: a dreadful rate bought a very high correction capability, and for a weak signal from deep space that was the right trade. The engineering name for this calculation is coding gain.

Figure (svg): The Hadamard code's decoding by dot products, with the correct row standing out from the rest.

Correction by correlation rather than by algebra, and it is why a code with a dreadful rate flew to Mars.

15. Why not just send a stronger signal?

Prediction

Predict first

Why not transmit a stronger signal with a more efficient code?

  • A stronger signal was technically impossible
  • It was an engineering trade — the weaker signal with heavy coding used less energy overall
  • Codes with higher rates did not exist in 1969
  • The receiver could not handle higher rates

Correct: It was an engineering trade — the weaker signal with heavy coding used less energy overall

Coding gain is the practical currency. It measures how much transmitter power a code saves for a given error rate, and it is what turns a mathematical choice into an engineering one.

The same calculation still runs today, in mobile networks, satellite links and storage. Modern codes — turbo codes and LDPC — approach Shannon's limit closely enough that the remaining gain is fractions of a decibel, which is why they replaced everything in this chapter for high-rate links.

And note the connection to Chapter 20. Shannon's channel coding theorem says reliable communication is possible at any rate below the channel capacity. This chapter builds the codes; that theorem says how good they could possibly get.

Why: The book puts it directly: the other option was to increase signal strength and use a higher-rate code, which would have transmitted faster and potentially used less energy. In this case the weaker signal more than offset the loss in speed. That calculation is called coding gain, and it decides which code a real system uses.

16. Sort by what each code can do

Definition probe

Five codes from Section 24.1.

Sort into buckets

Sort each by its capability.

Detects errors only
Single parity check; ISBN
Corrects errors
Hamming [7,4]; Two-dimensional parity; Hadamard (32, 64, 16)
detect
Both produce a single yes/no answer with no location information. A parity failure could come from any of the eight bits; an ISBN checksum failure could come from any digit. Detection is cheap and the response is retransmission.
correct
Each of these localises the error. Two-dimensional parity intersects a failing row with a failing column; the Hamming syndrome names the position directly; the Hadamard decoder identifies the nearest codeword by correlation.

17. Error Correcting Codes

Section

Section 24.2 · pp. 472-478

18. Codes, alphabets and block codes

Concept

The framework. A sender encodes a message into codewords made of symbols from an alphabet; these cross a noisy channel; the receiver decodes by correcting the errors and recovering the message.

Binary codes use {0, 1}; ternary codes use three symbols; a q-ary code uses q of them.

Code of length n — A nonempty subset of Aⁿ, where A is the alphabet. Its elements are codewords or code vectors.

These are block codes — every codeword has the same length. Codes with variable-length codewords exist and are mentioned at the end of the chapter.

A random subset of Aⁿ would be hopeless to decode, so useful codes carry extra structure. The commonest requirement is that A be a finite field and the code be a subspace, which is what Section 24.4 studies.

Figure (svg): The coding pipeline: a message is encoded to a codeword, corrupted by a noisy channel, and decoded back.

Every code in the chapter is a different answer to one question: what redundancy is worth its bandwidth?

19. Hamming distance

Concept

To decode we need a measure of how close two vectors are.

Hamming distance d(u, v) — The number of places where u and v differ. So d(10101010, 10111000) = 2, and d(fourth, eighth) = 4.

Its importance is that it counts errors: d(u, v) is the minimum number of symbol changes needed to turn u into v.

It is a metric, satisfying three properties: it is non-negative and zero only when u = v; it is symmetric; and it obeys the triangle inequality d(u, v) ≤ d(u, w) + d(w, v).

The triangle inequality has a one-line proof: if u and v differ in a place, then w must differ from at least one of them there. So the places where u and v differ are counted among those where u differs from w plus those where v differs from w.

Figure (svg): Hamming spheres around codewords, showing why minimum distance controls detection and correction.

One number, d(C), controls everything a code can do — and the two thresholds differ by a factor of two.

20. Detection, correction, and the two thresholds

Worked example

The minimum distance d(C) — the smallest distance between distinct codewords — controls everything.

A code detects up to s errors if d(C) ≥ s + 1

Why: If a codeword is changed in s or fewer places, the result cannot be a different codeword, because that would need at least s + 1 changes. So the receiver knows something is wrong.

A code corrects up to t errors if d(C) ≥ 2t + 1

Why: Meaning nearest-neighbour decoding gives the right answer whenever at most t errors occurred.

Suppose c was sent and r received with d(c, r) ≤ t, and some other codeword c₁ had d(c₁, r) ≤ t

Why: Then by the triangle inequality 2t + 1 ≤ d(C) ≤ d(c, c₁) ≤ d(c, r) + d(c₁, r) ≤ 2t.

Verify: a contradiction, so every other codeword is at distance at least t + 1 and nearest-neighbour decoding succeeds

Why: Geometrically: spheres of radius t around the codewords do not overlap, which is exactly what d ≥ 2t+1 says. And note the factor of two — correcting costs about twice the distance that detecting does.

Figure (svg): Hamming spheres around codewords, showing why minimum distance controls detection and correction.

One number, d(C), controls everything a code can do — and the two thresholds differ by a factor of two.

21. Notation and code rate

Concept

A code of length n with M codewords and minimum distance d is called an (n, M, d) code.

Linear codes get square brackets: an [n, k, d] binary linear code is an (n, 2ᵏ, d) code. The two notations cause less confusion than one might expect.

  1. The binary repetition code {000, 111} is a (3, 2, 3) code
  2. The Hamming [7,4] code is a (7, 16, 3) code
  3. The Hadamard code is a (32, 64, 16) code, correcting 7 errors since 16 ≥ 2·7 + 1

\[ R = \frac{\log_q M}{n} \]

The code rate R is the ratio of input data symbols to transmitted symbols — the fraction of bandwidth carrying actual data. For the Hadamard code, log₂(64)/32 = 6/32.

Figure (svg): Code rates of the chapter's examples, from the wasteful repetition code to the efficient Hamming code.

Higher is better, as long as too much correcting power is not lost — and that trade is the whole engineering problem.

22. Equivalent codes

Concept

Two operations produce a code that is essentially the same:

  1. Positional permutation — permute the order of the entries in every codeword
  2. Symbol permutation — fix a position and apply a permutation of the alphabet to that position in every codeword

Two codes are equivalent if one can be obtained from the other by a series of these. Every code equivalent to an (n, M, d) code is also an (n, M, d) code.

But the converse fails: for some n, M, d there are several inequivalent codes with the same parameters. So the parameters do not determine the code.

This matters for the McEliece cryptosystem in Section 24.10, whose entire security rests on disguising a code by scrambling and permuting it — producing an equivalent code that nobody can recognise.

23. Reading the two thresholds

Notation

Two inequalities govern what a code can do.

Annotate

On: \( \text{detect } s \iff d(C) \ge s+1, \qquad \text{correct } t \iff d(C) \ge 2t+1 \)

  • s errors must not be able to reach another codeword. If the nearest codeword is s+1 away, s changes cannot arrive at one.
  • Spheres of radius t around distinct codewords must not overlap, or a received vector would have two equally close candidates. Non-overlap requires the centres to be more than 2t apart.
  • Correcting costs roughly twice the distance detecting does — which is the formal version of the warm-up's observation that two repetitions detect and three correct.
  • The minimum over all pairs of distinct codewords. One bad pair caps the whole code, which is why constructions aim to make every pair far apart.
  • Any efficient algorithm. The definitions say nearest-neighbour decoding gives the right answer, not that you can find the nearest neighbour quickly — and for large codes you cannot.

That last note motivates the rest of the chapter: linear, cyclic and BCH codes exist because their structure makes the nearest neighbour findable.

24. Complete the framework

Faded example

Four blanks.

Fill in the blanks

An (n, M, d) code has length n, M codewords, and minimum distance d. It can detect up to d − 1 errors and correct up to (d−1)/2. Its code rate is log_q M divided by n. Decoding by choosing the nearest codeword in Hamming distance is called nearest neighbour decoding.

Why: Detection tolerates d−1 errors and correction only ⌊(d−1)/2⌋ — the same factor of two, seen from the other side.

25. Bounds on General Codes

Section

Section 24.3 · pp. 478-486

26. The Singleton bound and MDS codes

Concept

We want d large so that many errors can be corrected, and M large so the rate is close to 1. These pull against each other, and the bounds say by how much.

\[ M \le q^{\, n - d + 1} \]

The proof is a projection argument. Delete the first d−1 entries of every codeword. Two distinct codewords differ in at least d places, so after deleting d−1 they still differ somewhere — the truncations are distinct. There are at most q^{n−d+1} truncations, so at most that many codewords.

A corollary: the code rate is at most 1 − (d−1)/n. So a large relative minimum distance forces a small rate.

MDS code — One meeting the Singleton bound with equality — maximum distance separable, the largest possible d for its n and M. Reed-Solomon codes are the important example.

Figure (svg): The three bounds on codes, with the Hamming [7,4] code measured against each.

Three inequalities that between them fence in every possible code, and a way to judge how good a real one is.

27. Hamming spheres and the sphere-packing bound

Concept

A Hamming sphere B(c, t) is the set of vectors within distance t of the codeword c. Counting its elements is a combinatorial exercise:

\[ |B(c, r)| = \binom{n}{0} + \binom{n}{1}(q-1) + \cdots + \binom{n}{r}(q-1)^r \]

Because there are C(n, m) ways to choose which m positions differ, and q−1 choices of a different symbol at each.

The Hamming bound, also called the sphere-packing bound: if d ≥ 2t+1 then the spheres of radius t around the codewords do not overlap, so their total size is at most qⁿ:

\[ M \le \frac{q^n}{\sum_{j=0}^{t} \binom{n}{j}(q-1)^j} \]

Perfect code — One with d = 2t+1 meeting the Hamming bound with equality — the spheres of radius t tile the whole space with nothing left over. Hamming codes and the Golay code G₂₃ are perfect.

Figure (svg): The three bounds on codes, with the Hamming [7,4] code measured against each.

Three inequalities that between them fence in every possible code, and a way to judge how good a real one is.

28. Measuring the examples against the bounds

Worked example

Three codes, three verdicts.

The repetition code (3, 2, 3): Singleton gives 2 ≤ 2³⁻³⁺¹ = 2 — an MDS code

Why: And the Hamming bound gives 2 ≤ 2³/(1 + 3) = 2, so it is also perfect.

The Hamming [7,4] code, a (7, 16, 3) code: Singleton gives 16 ≤ 2⁵ = 32

Why: Not MDS — there is room for a better code at these parameters, in principle.

But the Hamming bound gives 16 ≤ 2⁷/(1 + 7) = 16

Why: Equality, so it is perfect. The 16 spheres of radius 1 contain 16 × 8 = 128 = 2⁷ vectors: the whole space, exactly once.

The Gilbert-Varshamov bound guarantees a (7, M, 3) code with M ≥ 128/29 ≈ 4.4

Why: So a code with at least 5 codewords must exist.

Verify: the Hamming code has 16, more than three times the guarantee

Why: Which shows how weak the lower bound is — and the book notes that codes with efficient decoding that also exceed Gilbert-Varshamov are relatively rare, so this one is doing very well.

Figure (svg): The three bounds on codes, with the Hamming [7,4] code measured against each.

Three inequalities that between them fence in every possible code, and a way to judge how good a real one is.

29. The Gilbert-Varshamov bound

Concept

The upper bounds say what cannot exist. This lower bound says what must.

\[ A_q(n, d) \ge \frac{q^n}{\sum_{j=0}^{d-1} \binom{n}{j}(q-1)^j} \]

The proof is a greedy construction. Pick any vector c₁ and delete every vector within distance d−1 of it. Pick c₂ from what remains; it is automatically at distance ≥ d from c₁. Repeat.

Each choice deletes at most the size of a sphere of radius d−1, so the process continues until M times that quantity exceeds qⁿ — which gives the bound.

The asymptotic version says that for any x with 0 < x < 1 − 1/q there are codes with d/n → x whose rate approaches at least H_q(x), an explicit entropy-like expression. Chapter 20's function, in a different role.

And it can be beaten. Tsfasman, Vladut and Zink showed in 1982 that Goppa codes from algebraic geometry exceed the asymptotic bound for certain x and q — one of the celebrated results of the subject.

30. What does a perfect code guarantee?

Prediction

Predict first

Does perfect mean best?

  • Yes — no better code exists at those parameters
  • No — it means the spheres tile the space exactly, and other codes can be better for particular purposes
  • Yes, but only for binary codes
  • It means the code corrects all errors

Correct: No — it means the spheres tile the space exactly, and other codes can be better for particular purposes

Every vector in the space is within distance 1 of exactly one codeword, which sounds ideal and has a hidden cost: there is no room to detect a double error. Any two-error pattern lands inside some sphere and is silently miscorrected.

Which is why extended Hamming codes exist — add one overall parity bit, giving d = 4, and the code corrects one error while detecting two. It is no longer perfect and it is more useful.

The complete list of perfect codes is known, and it is short: Hamming codes, the binary Golay code G₂₃, a ternary [11, 6, 5] Golay code, the trivial codes, and binary repetition codes of odd length. A rare instance of a classification being finished.

Why: The book issues this caveat explicitly. Perfect means the radius-t spheres cover the space with nothing left over — an elegant property, not an optimality claim. Reed-Solomon codes are not perfect and have better error-correcting capability in the situations they are designed for.

31. Linear Codes

Section

Section 24.4 · pp. 486-497

32. Linear Code

Concept

When you speak on a mobile phone, your voice is coded, transmitted, and decoded before your friend hears it. The decoding delay is critical — a few seconds would make conversation impossible. So codes need structure that makes decoding fast, and the standard structure is linearity.

Linear code — A k-dimensional subspace of Fⁿ, where F is a finite field. Called an [n, k] code, or [n, k, d] if the minimum distance is d.

For binary codes the definition is simpler: a set of 2ᵏ binary n-tuples such that the sum of any two codewords is a codeword.

Many codes already met are linear. The repetition code {000, 111} is a one-dimensional subspace; the parity check code is a [8, 7] code; the Hamming [7,4] code is spanned by the four rows of G.

The ISBN code is not linear. It lives in ℤ₁₁¹⁰ but is not closed under linear combinations, because X is not allowed among the first nine entries.

Figure (svg): A linear code as the row space of a generator matrix, with the parity check matrix annihilating it.

Linearity is not a mathematical nicety — it turns decoding from a search into a matrix multiplication.

33. Weight, and why linearity helps immediately

Concept

The Hamming weight wt(u) is the number of nonzero places in u — that is, d(u, 0).

Proposition. For a linear code, d(C) equals the smallest weight among nonzero codewords.

\[ d(C) = \min\{\, \mathrm{wt}(u) \;\mid\; 0 \ne u \in C \,\} \]

The proof is one observation: d(v, w) = wt(v − w), and v − w is itself a codeword by linearity. So distances between pairs are weights of single codewords.

The computational saving is large. For an arbitrary code, finding d(C) means comparing all M(M−1)/2 pairs. For a linear code it means checking M weights — a quadratic problem becomes linear.

And this is the first of several such savings. Linearity turns nearly every question about the code from a search over pairs into a computation on a matrix.

Figure (svg): A linear code as the row space of a generator matrix, with the parity check matrix annihilating it.

Linearity is not a mathematical nicety — it turns decoding from a search into a matrix multiplication.

34. Generator and parity check matrices

Concept

To build an [n, k] code, take a k × n generating matrix G of rank k; the codewords are the vectors vG.

Systematic form is G = [I_k | P]. Then the first k bits of a codeword are the message — the information symbols — and the remaining n−k are the check symbols.

A parity check matrix H for C is one with the property that v ∈ C if and only if vHᵀ = 0.

\[ G = [I_k \mid P] \;\Longrightarrow\; H = [-P^T \mid I_{n-k}] \]

Any k × n matrix of rank k generates a code, and row and column operations bring it to systematic form — so nothing is lost by assuming that shape.

Figure (svg): A linear code as the row space of a generator matrix, with the parity check matrix annihilating it.

Linearity is not a mathematical nicety — it turns decoding from a search into a matrix multiplication.

35. Why H is a parity check matrix

Worked example

Two steps: show H annihilates the code, then show nothing else is annihilated.

Take the i-th row of G: it is (0,…,1,…,0, p_{i,1},…,p_{i,n−k}) with the 1 in position i

Why: A codeword.

The j-th column of Hᵀ is (−p_{1,j},…,−p_{n−k,j}, 0,…,1,…,0) with the 1 in position n−k+j

Why: From the shape of H = [−Pᵀ | I].

Their dot product is 1·(−p_{i,j}) + p_{i,j}·1 = 0

Why: So Hᵀ annihilates every row of G, hence every codeword by linearity.

Hᵀ contains I_{n−k} as a submatrix, so it has rank n−k, so its left null space has dimension k

Why: Standard linear algebra: the left null space of an m × n matrix of rank r has dimension n − r.

Verify: C is contained in a space of dimension k and has dimension k, so they are equal

Why: Hence vHᵀ = 0 characterises the code exactly. If a received vector gives a nonzero result, there is definitely an error; if zero, it is a codeword — and since errors are more likely to be few than to be enough to reach another codeword, the best guess is that none occurred.

Figure (svg): A linear code as the row space of a generator matrix, with the parity check matrix annihilating it.

Linearity is not a mathematical nicety — it turns decoding from a search into a matrix multiplication.

36. Syndrome decoding

Concept

The parity check matrix does more than detect. It makes decoding tractable.

Write the received vector as y = c + e, where c is the codeword sent and e the error pattern. Then

\[ y H^T = c H^T + e H^T = 0 + e H^T = e H^T \]

Syndrome — The vector s = yHᵀ. It depends only on the error pattern, not on which codeword was sent.

That is the whole gain. One table indexed by syndrome serves every codeword at once — rather than a table indexed by received vector, which would have qⁿ entries.

The book's construction makes this concrete: list the codewords in the first row, then repeatedly pick a remaining vector of smallest weight and add it to the first row. Each row shares a syndrome, and its leading vector is the most likely error pattern.

Even so, syndrome decoding is too slow for large codes — the table has q^{n−k} rows. That limitation is what drives the rest of the chapter toward codes with algebraic decoders.

Figure (svg): Syndrome decoding: the syndrome depends only on the error, not on the codeword.

The single most useful consequence of linearity, and the reason cellular calls decode fast enough to hold a conversation.

37. Dual codes

Concept

Fⁿ has a dot product defined as usual — with one surprise: over ℤ₂, (0,1,0,1,1,1)·(0,1,0,1,1,1) = 0, so a nonzero vector can be orthogonal to itself. The dot product does not measure length here, but it is still useful.

Dual code C⊥ — The set of vectors orthogonal to every codeword of C.

Proposition. If C is an [n, k] code with generating matrix G = [I_k | P], then C⊥ is an [n, n−k] code with generating matrix H = [−Pᵀ | I_{n−k}] — and G is a parity check matrix for C⊥.

So G and H swap roles, which is a genuinely useful duality: results about a code translate into results about its dual, and some codes are easier to analyse from one side than the other.

A code with C = C⊥ is self-dual. The Golay code G₂₄ is the important example, and self-duality forces n = 2k — the code is exactly half the space.

38. Why does linearity make decoding tractable?

Socratic

For an arbitrary code, decoding means finding the nearest codeword among M candidates.

Discussion prompt

List everything linearity buys, and say what it still does not.

Hint: Storage, computation, and the limits.

Answer:

Storage. A linear code is specified by a k × n matrix rather than a list of qᵏ codewords. For the [1024, 524] code in Section 24.10 that is the difference between a matrix and a list with 2⁵²⁴ entries.

Minimum distance. Computing d(C) becomes a search over weights rather than over pairs — M operations rather than M².

Detection. One matrix multiplication, vHᵀ, decides membership in the code.

Decoding. The syndrome depends only on the error, so a single table of q^{n−k} entries suffices instead of qⁿ.

What it does not buy is efficiency at scale. q^{n−k} is still enormous for a real code, so syndrome decoding alone is impractical. That is why the chapter goes on to cyclic, BCH and Reed-Solomon codes, whose extra algebraic structure gives decoders that never build a table at all.

Figure (svg): Syndrome decoding: the syndrome depends only on the error, not on the codeword.

The single most useful consequence of linearity, and the reason cellular calls decode fast enough to hold a conversation.

39. Hamming Codes

Section

Section 24.5 · pp. 497-500

40. The general Hamming Code construction

Concept

A family of single-error-correcting codes that encode and decode easily. They were originally used to control errors in long-distance telephone calls.

  1. Length n = 2ᵐ − 1
  2. Dimension k = 2ᵐ − m − 1
  3. Minimum distance d = 3

The construction is through the parity check matrix. Build an m × n matrix whose columns are all the nonzero binary m-tuples — there are exactly 2ᵐ − 1 of them, which is why n has that form.

Then move columns so the matrix ends with I_m, giving systematic form. The order of the other columns is irrelevant, which is a way of saying that all Hamming codes of a given length are equivalent.

Figure (svg): The general Hamming code construction: a parity check matrix whose columns are every nonzero binary m-tuple.

The construction explains the decoder: the syndrome is a column of H, and columns are in bijection with bit positions.

41. Decoding the [15, 11] Hamming code

Worked example

The algorithm is three lines, and the example shows why it works.

Compute the syndrome s = yHᵀ. If s = 0, there are no errors — return y

Why: One matrix multiplication.

Otherwise find the column of H equal to sᵀ, at position j

Why: Every nonzero m-tuple appears as a column exactly once, so the match exists and is unique.

Flip the j-th bit of y

Why: And the result is the transmitted codeword, provided at most one error occurred.

Example: y = 000010000011001 gives s = (1, 1, 1, 1)

Why: Which is the transpose of the 11th column of H.

Verify: flip bit 11 to get 000010000010001; the first 11 bits are the message 00001000000

Why: And now the mystery of Section 24.1 dissolves: the syndrome is a column of H, columns correspond one-to-one with positions, so the syndrome is the error's address. The construction and the decoder are the same fact seen twice.

Figure (svg): The general Hamming code construction: a parity check matrix whose columns are every nonzero binary m-tuple.

The construction explains the decoder: the syndrome is a column of H, and columns are in bijection with bit positions.

42. Golay Codes

Section

Section 24.6 · pp. 500-508

43. Golay Code

Concept

Two of the most famous binary codes. The [24, 12, 8] extended Golay code G₂₄ was used by Voyager I and Voyager II during 1979–1981 to protect the colour pictures of Jupiter and Saturn on their way back to Earth.

The [23, 12, 7] code G₂₃ is closely related, obtained by deleting a coordinate — and it is one of the very few perfect codes.

G₂₄'s generating matrix is 12 × 24, of the form [I₁₂ | B]. The last 11 columns of the first row come from the squares mod 11 — namely 0, 1, 3, 4, 5, 9 — giving the vector (1,1,0,1,1,1,0,0,0,1,0). The remaining rows are cyclic shifts of it, with one exceptional final row.

G₂₄ is self-dual, meaning it equals its own dual code — a rare and highly structured property, and part of why the Golay codes appear all over mathematics, from sporadic simple groups to the Leech lattice.

As an [24, 12, 8] code it corrects 3 errors and detects 4, at rate 1/2. For a spacecraft at Jupiter, that was the right balance between the Hadamard code's extravagance and a rate that would have taken too long to transmit.

44. Cyclic Codes

Section

Section 24.7 · pp. 508-518

45. Cyclic Code

Concept

A code is cyclic if every cyclic shift of a codeword is a codeword:

\[ (c_1, \ldots, c_n) \in C \;\Longrightarrow\; (c_n, c_1, \ldots, c_{n-1}) \in C \]

This looks like an odd thing to require. Why should it matter that the shifts of a codeword are codewords?

The point is structure. Cyclic codes have a great deal of it, which makes them easier to study — and in the case of BCH codes, that structure yields an efficient decoding algorithm. Structure is the currency of this whole chapter.

The algebraic reformulation is where the power comes from. Identify the vector (a₀, …, a_{n−1}) with the polynomial a₀ + a₁X + ⋯ + a_{n−1}X^{n−1}, working modulo Xⁿ − 1. Then a cyclic shift is just multiplication by X — because X·X^{n−1} = Xⁿ ≡ 1.

Figure (svg): A cyclic code generated by a polynomial, with every cyclic shift of a codeword also a codeword.

Rewriting vectors as polynomials turns a shift into multiplication by X — and that is where the fast decoders come from.

46. A cyclic [7, 3, 4] code

Worked example

The book's example, both ways.

Take the generator matrix whose rows are 1011100, 0101110, 0010111

Why: Each row is the previous one shifted right by one.

Its eight codewords are 0000000 and the seven cyclic shifts of 1011100

Why: So the code is genuinely cyclic, and the minimum weight is 4 — a [7, 3, 4] code.

Algebraically: let g(X) = 1 + X² + X³ + X⁴ in ℤ₂[X]/(X⁷ − 1)

Why: Then g(X)·1 gives the first row, g(X)·X the second, g(X)·X² the third.

Every codeword is g(X)f(X) for some f of degree ≤ 2

Why: And taking f of higher degree gives nothing new, since g(X)(X³ + X² + 1) = X⁷ − 1 ≡ 0.

Verify: the code is exactly the set of multiples of g(X) modulo X⁷ − 1

Why: g is called the generator polynomial, and it must divide Xⁿ − 1. So constructing cyclic codes of length n amounts to factoring Xⁿ − 1 — a finite, purely algebraic search that replaces any hunting for good matrices.

Figure (svg): A cyclic code generated by a polynomial, with every cyclic shift of a codeword also a codeword.

Rewriting vectors as polynomials turns a shift into multiplication by X — and that is where the fast decoders come from.

47. BCH Codes

Section

Section 24.8 · pp. 518-528

48. BCH Code

Concept

A class of cyclic codes discovered around 1959 by Bose and Ray-Chaudhuri, and independently by Hocquenghem. They matter because good decoding algorithms exist that correct multiple errors — and they are used in satellites.

The construction needs a primitive nth root of unity. For a finite field F with q = pᵐ elements and n not divisible by p, there is a larger field F′ containing F and an element α with αⁿ = 1 but αᵏ ≠ 1 for 1 ≤ k < n.

The calculations happen in F′, but the code lives over F. For example with F = ℤ₂ and n = 3, take F′ = GF(4), where ω is a primitive cube root of unity.

A primitive nth root of unity exists in a field with q′ elements exactly when n divides q′ − 1 — which is the condition that decides how large the auxiliary field must be.

A BCH code is then a cyclic code whose generator polynomial has a run of consecutive powers of α as roots, and the next slide explains why that forces the distance up.

49. The BCH bound

Concept

The theorem that makes the construction work, and it is the reason the roots are chosen consecutively.

Theorem. Let C be a cyclic code over F with generator polynomial g, and let α be a primitive nth root of unity. If

\[ g(\alpha^{\ell}) = g(\alpha^{\ell+1}) = \cdots = g(\alpha^{\ell+\delta}) = 0 \]

then d ≥ δ + 2.

The proof is a Vandermonde determinant argument. Suppose a codeword had weight w with 1 ≤ w < δ + 2. Its nonzero coefficients satisfy a homogeneous linear system whose matrix is Vandermonde in distinct powers of α — hence has nonzero determinant, forcing all the coefficients to be zero. Contradiction.

So minimum distance is designed rather than discovered. Want d ≥ 5? Arrange for three consecutive powers of α to be roots. That is an extraordinary amount of control, and it is what the algebra was for.

Figure (svg): A cyclic code generated by a polynomial, with every cyclic shift of a codeword also a codeword.

Rewriting vectors as polynomials turns a shift into multiplication by X — and that is where the fast decoders come from.

50. Reed-Solomon Codes

Section

Section 24.9 · pp. 528-532

51. Reed-Solomon Code

Concept

Constructed in 1960, an example of BCH codes, and used in spacecraft communications and on compact discs.

Let F have q elements and n = q − 1, so F contains a primitive nth root of unity α. Choose d with 1 ≤ d < n and set

\[ g(X) = (X - \alpha)(X - \alpha^2) \cdots (X - \alpha^{d-1}) \]

The BCH bound gives d(C) ≥ d. And since g has degree d−1 it has at most d nonzero coefficients, so the codeword formed from them has weight at most d — hence the minimum weight is exactly d.

The dimension is n − deg(g) = n + 1 − d, so this is a cyclic [n, n+1−d, d] code with q^{n−d+1} codewords.

Which meets the Singleton bound exactly. Reed-Solomon codes are MDS: the best possible distance for their length and size.

Figure (svg): A Reed-Solomon code over the integers mod 7, generated by a product of root factors.

Roots chosen consecutively force the distance up by the BCH bound, and the count then meets Singleton exactly.

52. A Reed-Solomon code over ℤ₇

Worked example

Small enough to write down completely.

F = ℤ₇, so q = 7 and n = 6

Why: A primitive sixth root of unity mod 7 is the same as a primitive root, and α = 3 works.

Choose d = 4, so g(X) = (X − 3)(X − 3²)(X − 3³)

Why: And 3² ≡ 2, 3³ ≡ 6 (mod 7), so the roots are 3, 2 and 6.

Multiplying out gives g(X) = X³ + 3X² + X + 6

Why: Coefficients mod 7.

The generating matrix has rows 613100, 061310, 006131

Why: The coefficients of g, shifted — the cyclic structure made visible.

Verify: 7³ = 343 codewords, minimum weight 4, and Singleton gives 7^(6−4+1) = 343

Why: Equality, confirming MDS. Note how the design worked: three consecutive roots forced d ≥ 4 by the BCH bound, and the degree of g forced d ≤ 4. The two squeezed together give exactly 4.

Figure (svg): A Reed-Solomon code over the integers mod 7, generated by a product of root factors.

Roots chosen consecutively force the distance up by the BCH bound, and the count then meets Singleton exactly.

53. Burst errors, CDs and spacecraft

Concept

Real errors are often not randomly distributed but come in bursts — a scratch on a CD corrupts many adjacent bits, and a burst of solar energy does the same to a spacecraft signal.

Reed-Solomon codes suit this exactly, because they work at the level of symbols rather than bits.

Take F = GF(2⁸), so each symbol is a byte and n = 255. With d = 33, a codeword is 255 bytes — 222 information bytes and 33 check bytes — transmitted as 2040 bits.

A burst of 121 consecutive corrupted bits touches at most 16 bytes, since 121 = 15·8 + 1. And 16 < d/2, so the errors are corrected.

The same 121 errors scattered randomly through the 2040 bits would corrupt far more bytes, and decoding would fail. So the choice of code depends on the type of errors expected — which is the practical version of this chapter's whole message.

Figure (svg): Why Reed-Solomon suits burst errors: many corrupted bits fall inside few corrupted symbols.

A symbol-level code turns a scratch on a CD into a handful of bad bytes instead of a hundred bad bits.

54. The McEliece Cryptosystem

Section

Section 24.10 · pp. 532-536

55. McEliece Cryptosystem

Concept

The section that puts coding theory in a cryptography book, and the idea is simple.

Suppose you have a binary string of length 1024 with 50 errors. There are C(1024, 50) ≈ 3 × 10⁸⁵ possible error locations, so exhaustive search is hopeless. But if you have an efficient decoding algorithm that nobody else has, only you can correct them.

  1. Bob picks G, the generating matrix of an (n, k) linear code with minimum distance d and a fast decoder
  2. He picks S, an invertible k × k matrix mod 2, and P, an n × n permutation matrix
  3. His public key is G₁ = SGP. He keeps S, G and P secret
  4. Alice encrypts x as y ≡ xG₁ + e (mod 2), where e is random of weight t

The public key is a generating matrix for a code that looks random, and finding the nearest codeword in a random linear code is hard. Bob's advantage is that he knows the scrambling.

Figure (svg): The McEliece cryptosystem: a scrambled generator matrix as public key, with deliberate errors only the owner can correct.

The trapdoor is an efficient decoder for a code that has been disguised as a random one.

56. How Bob decrypts

Worked example

Four steps, each undoing one layer of the disguise.

Compute y₁ ≡ yP⁻¹

Why: Since P is a permutation matrix, e₁ = eP⁻¹ is still a binary string of weight t — permuting positions does not change the weight. So y₁ ≡ xSG + e₁.

Apply the code's own error decoder to y₁

Why: This is the step nobody else can perform, and it yields the nearest codeword x₁.

Find x₀ with x₀G ≡ x₁

Why: In systematic form that is just the first k bits of x₁.

Compute x ≡ x₀S⁻¹

Why: Undoing the scrambling matrix, and out comes the message.

Verify: the trapdoor is the decoder, not the matrix

Why: S and P hide which code G₁ generates; without knowing that, an attacker faces general nearest-codeword decoding, which is NP-hard. The book notes that S provides only a little security on its own — once a decoding algorithm for the code generated by GP is known, a chosen-plaintext attack recovers S exactly as with the Hill cipher.

Figure (svg): The McEliece cryptosystem: a scrambled generator matrix as public key, with deliberate errors only the owner can correct.

The trapdoor is an efficient decoder for a code that has been disguised as a random one.

57. A McEliece encryption with the [7,4] code

Worked example

The book's toy example, using the Hamming [7,4] code so that every step is checkable.

G is the Hamming [7,4] generator; Bob picks an invertible S and a permutation P, and publishes G₁ = SGP

Why: A 4 × 7 matrix that looks nothing like a Hamming generator.

Alice's message is x = (1, 0, 1, 1), and she picks e = (0, 1, 0, 0, 0, 0, 0) of weight 1

Why: Weight 1 because the Hamming code corrects one error — t must not exceed the code's capability.

She sends y = xG₁ + e = (0, 0, 0, 1, 1, 0, 0)

Why: The ciphertext.

Bob computes y₁ = yP⁻¹ = (0, 0, 1, 0, 0, 0, 1)

Why: Undoing the permutation, so the error is back in a position the Hamming decoder understands.

The syndrome via H locates and flips the bad bit, giving x₁ = (0, 0, 1, 0, 0, 1, 1)

Why: Ordinary Hamming decoding.

Verify: x₀ = (0, 0, 1, 0) is the first four bits, and x = x₀S⁻¹ = (1, 0, 1, 1)

Why: The original message. A real deployment uses a [1024, 524] Goppa code with t = 50, where the same procedure runs against C(1024, 50) ≈ 3 × 10⁸⁵ possible error patterns.

Figure (svg): The McEliece cryptosystem: a scrambled generator matrix as public key, with deliberate errors only the owner can correct.

The trapdoor is an efficient decoder for a code that has been disguised as a random one.

58. Parameters, and the key-size problem

Concept

McEliece suggested a [1024, 512, 101] Goppa code. Goppa codes have parameters n = 2ᵐ, d = 2t+1, k = n − mt, so m = 10 and t = 50 give a [1024, 524, 101] code correcting up to 50 errors.

For given m and t there are many inequivalent Goppa codes, which is essential: the attacker cannot simply try the one code with those parameters.

The system seems reasonably secure — and it has stood since 1978, which is longer than RSA has been under attack in its modern parameter ranges.

The disadvantage is the key size. G₁ is 1024 × 524 bits, about 67 kilobytes, against a few hundred bytes for RSA. That single fact is why McEliece was never widely deployed despite never being broken.

And it is why the system has come back. Its security does not rest on factoring or discrete logs, so no quantum attack is known — and Classic McEliece, its direct descendant, is a NIST post-quantum finalist. The huge key that killed it in 1978 is now an acceptable price.

Figure (svg): The trade McEliece makes: very fast operations and strong security, against an unusually large public key.

One of the oldest public-key systems, never broken, and almost never used — the key size decided it.

59. What the chapter leaves out

Concept

Coding theory is a vast subject explored by both mathematicians and engineers, and this chapter touches a handful of its ideas.

Convolutional codes. Everything here is a block code, mapping fixed-length blocks to fixed-length codewords. When data arrive continuously it is better to map a stream of data to a stream of coded symbols — avoiding the delay of waiting for a whole block. The book's analogy is exact: block codes are to convolutional codes as block ciphers are to stream ciphers.

Efficient decoding. Syndrome decoding beats searching, and is still too slow for large codes. The original BCH and Reed-Solomon decoders could not handle more than a few errors; Berlekamp and Massey supplied the efficient approach, and research continues.

Erasures rather than errors. On the internet, data travel in packets that are sometimes dropped entirely rather than corrupted. TCP handles this by retransmission; erasure codes handle it by redundancy, and knowing which symbols are missing makes the problem easier than not knowing.

And connections elsewhere in mathematics — dense sphere packings in high dimensions among them, which is how the Golay code leads to the Leech lattice and from there to the sporadic simple groups.

One practical warning the book gives: when doing cryptography, error control must be combined properly, or the receiver may not be able to decrypt at all. A single flipped bit in a block cipher's ciphertext destroys an entire block.

60. Why does error correction belong in a cryptography book?

Socratic

Chapters 1 through 23 were about confidentiality and authenticity. This one is about noise.

Discussion prompt

Give three reasons the subjects belong together.

Hint: One is the McEliece system; the others are practical.

Answer:

A decoder can be a trapdoor. McEliece's system is built entirely on the gap between decoding a code you know and decoding one you do not — the same shape as any public-key system, with a different hard problem underneath.

Encryption amplifies errors. A good cipher has the avalanche property, so one flipped ciphertext bit corrupts an entire block of plaintext. Error correction must be applied outside the encryption, or a single channel error destroys far more than one bit.

The order matters, and it is a real design decision. Encrypt-then-code protects the ciphertext in transit; code-then-encrypt leaves the redundancy encrypted and useless for correction. The first is correct and is what real systems do.

And both fields are about structured redundancy. Coding adds redundancy so errors can be undone; cryptanalysis exploits redundancy so ciphertexts can be broken — Chapter 20's unicity distance is exactly this. The same quantity, wanted in one setting and unwanted in the other.

There is a fourth connection worth knowing: Shannon founded both subjects, in two 1948–49 papers. Information theory underlies the channel coding theorem and the definition of perfect secrecy alike.

Figure (svg): The McEliece cryptosystem: a scrambled generator matrix as public key, with deliberate errors only the owner can correct.

The trapdoor is an efficient decoder for a code that has been disguised as a random one.

61. Encrypting after coding

Anomaly

A system applies a Reed-Solomon code to the plaintext, then encrypts the result with AES before transmission.

Predict first

What goes wrong?

  • Nothing — the redundancy is still there
  • A channel error corrupts a whole AES block on decryption, producing far more symbol errors than the code can correct
  • The code becomes linear over the wrong field
  • AES cannot encrypt coded data

Correct: A channel error corrupts a whole AES block on decryption, producing far more symbol errors than the code can correct

So the correct order is encrypt-then-code. The code protects the ciphertext, errors are corrected before decryption, and the cipher only ever sees a clean block.

This also matches the authentication ordering from Chapter 12. Encrypt-then-MAC is preferred for a related reason: the outermost layer should be the one that checks integrity, so nothing is processed before it has been validated.

And the book flags the general point: error control must be combined carefully with cryptography, or the receiver may not be able to decrypt at all. Getting the layering wrong turns a recoverable channel error into total message loss.

Why: The receiver decrypts first, and one flipped ciphertext bit avalanches into about half the bits of a 16-byte block. The Reed-Solomon decoder then sees 16 corrupted bytes from a single channel error — likely beyond its correcting power, and certainly a catastrophic amplification.

62. Sort by where the structure comes from

Definition probe

Each code family gains its decoding efficiency from a different kind of structure.

Sort into buckets

Sort each.

Algebraic — polynomials or matrices over a field
Hamming codes; Reed-Solomon codes
Combinatorial or geometric
The repetition code; Hadamard codes
alg
Hamming codes are defined by a parity check matrix whose columns enumerate the nonzero m-tuples, and the syndrome is a matrix product. Reed-Solomon codes are multiples of a polynomial with prescribed roots, and their decoders are polynomial algorithms.
comb
Repetition needs only a majority vote. The Hadamard code decodes by correlation against orthogonal rows — a geometric argument about inner products, with no field arithmetic at all.

63. Match each code to where it was used

Matching

Five codes with real deployments.

Match the pairs

  • m1. Hamming codes
  • m2. Hadamard code
  • m3. Golay G₂₄
  • m4. Reed-Solomon
  • m5. Goppa codes
  • r1. Long-distance telephone error control
  • r2. Mariner's pictures from Mars, 1969
  • r3. Voyager at Jupiter and Saturn, 1979-81
  • r4. Compact discs and spacecraft burst errors
  • r5. The McEliece public key cryptosystem

Why: The progression through the spacecraft is instructive: Mariner used a rate-6/32 code because its signal was very weak, Voyager a rate-1/2 code a decade later with better transmitters. As the channel improves, the rate can rise — which is coding gain being spent differently.

64. Rate against correction

Trade off

Every code sits somewhere on this trade. Fill the blanks.

Comparison matrix

Hamming [7,4]Hadamard (32,64,16)
Code rate4/7 ≈ 0.576/32 ≈ 0.19
Errors corrected17
Errors detected215
Perfect?yes — meets the Hamming boundno
Suited toa fairly clean channela very weak or noisy signal

The Singleton corollary makes the trade formal: R ≤ 1 − (d−1)/n, so a large relative distance forces a small rate. No cleverness escapes it.

65. Which claims are true?

Discrimination

Five statements about codes.

Sort into buckets

Sort each.

True
For a linear code, minimum distance equals minimum nonzero weight; Reed-Solomon codes are MDS; A code with d = 5 corrects 2 errors
False
A perfect code is the best possible code at its parameters; The syndrome depends on which codeword was sent
true
Linearity makes distances into weights since d(v,w) = wt(v−w). Reed-Solomon meets Singleton with equality by construction. And d ≥ 2t+1 with d = 5 gives t = 2.
false
Perfect means the spheres tile the space exactly — elegant, not optimal, and the book warns about this explicitly. And the syndrome yHᵀ = (c+e)Hᵀ = eHᵀ depends only on the error, which is the entire point of syndrome decoding.

66. How far can a code be pushed?

Edge cases

The bounds fence in what is possible.

Discussion prompt

What do they leave open, and where does the real difficulty lie?

Hint: Existence versus construction versus decoding.

Answer:

The bounds are about existence, not construction. Gilbert-Varshamov proves a good code exists by a greedy argument that gives no way to write one down, and no way to decode it if you could.

Three separate problems hide here. Does a code with these parameters exist? Can it be constructed explicitly? Does it have an efficient decoder? A yes to the first says nothing about the others.

The gap is where the whole subject lives. Random linear codes achieve excellent parameters and are useless because decoding them is NP-hard — which is precisely what McEliece's security rests on.

Beating Gilbert-Varshamov was a landmark. Tsfasman, Vladut and Zink did it in 1982 with algebraic-geometry codes, and the book notes that codes which both exceed the bound and decode efficiently remain relatively rare.

And modern practice has moved past all of it. Turbo codes and LDPC codes approach Shannon's channel capacity within a fraction of a decibel, using iterative probabilistic decoding rather than algebra. The bounds in this chapter still hold; the codes that meet them best are not the ones in this chapter.

67. Order these by error-correcting capability

Ranking

Five codes.

Put in order

  1. Single parity check [8,7]
  2. Hamming [7,4]
  3. Golay [24,12,8]
  4. Hadamard (32,64,16)
  5. Reed-Solomon over GF(2⁸) with d = 33

Why: Parity corrects nothing. Hamming corrects 1 (d = 3). Golay corrects 3 (d = 8 gives t = 3). Hadamard corrects 7 (d = 16). Reed-Solomon with d = 33 corrects 16 symbol errors — and each symbol is a byte, so in the burst case that is up to 121 bit errors. The ordering is by correction capability, and the code rates run almost exactly the other way.

68. How much redundancy does a CD carry?

Estimation

A Reed-Solomon code over GF(2⁸) with n = 255 and d = 33.

Predict first

What fraction of the transmitted bytes are check bytes?

  • About 3%
  • About 13%
  • About 50%
  • About 90%

Correct: About 13%

Real CDs use a cross-interleaved scheme with two Reed-Solomon codes, which does better still and is why a disc keeps playing through a scratch several millimetres wide.

Interleaving is the other half of the trick. Consecutive bytes of a codeword are written far apart on the disc, so a physical scratch is spread thinly across many codewords rather than destroying one — converting a burst into scattered single-symbol errors, which is the case the code handles best.

Which is a general design pattern: if your code handles one error model well, add a transformation that converts the errors you actually get into the errors it likes.

Why: There are 33 check bytes out of 255, or about 13% — a rate of 222/255 ≈ 0.87. That buys correction of 16 symbol errors, which in a burst covers up to 121 consecutive corrupted bits. An excellent return for 13% overhead, and it is why symbol-level codes are used where bursts dominate.

69. Find the flaws in this error-control design

Error analysis

From a satellite link specification.

Annotate

  • It detects an odd number of bit errors in 512 bits and nothing else. Two errors — likely in a burst — cancel and pass silently. This is a detection scheme for a nearly-clean channel, deployed on a noisy one.
  • The text says coding happens after encryption, which is correct — but decoding must then happen before decryption at the receiver. If the order is reversed anywhere, one channel error avalanches into a whole corrupted AES block.
  • Backwards. Burst errors are exactly what symbol-level codes such as Reed-Solomon handle well, because many corrupted bits fall inside few corrupted symbols. Choosing a bit-level code for bursts discards the main defence.
  • The round-trip delay to geostationary orbit is about half a second, and a deep-space link is far worse. Forward error correction exists precisely so that retransmission is not needed — Voyager could not have asked Jupiter to repeat itself.

Every fault is a mismatch between the error model and the code chosen for it, which is this chapter's central practical lesson.

70. Choose a code for a deep-space probe

Constraint

A probe near Saturn sends images. The signal is very weak, errors come in bursts from solar activity, and retransmission is impossible.

Discussion prompt

What would you specify, and why?

Hint: Consider the error model, the rate, and the impossibility of asking again.

Answer:

Forward error correction only. Round-trip time to Saturn is over two hours, so any scheme requiring retransmission is useless. Everything must be correctable at the receiver from what arrives.

A symbol-level outer code — Reed-Solomon over GF(2⁸). Bursts corrupt runs of adjacent bits, and byte-level symbols turn a long burst into a handful of bad symbols.

Plus interleaving between the code and the channel, so that a burst is spread across many codewords rather than concentrated in one. This is the single most effective addition, and it costs only buffer memory.

And a low-rate inner code for the weak signal, exactly as Voyager did: a convolutional code inner, Reed-Solomon outer. The inner code cleans up random bit errors, and the outer one mops up the bursts the inner decoder gets wrong.

This concatenated design is what actually flew, and it is the historically correct answer: Voyager used a convolutional inner code with a Golay or Reed-Solomon outer code, and the combination outperformed either alone.

71. Reading the sphere-packing bound

Cost model

One inequality limits every code.

Annotate

On: \( M \le \frac{q^n}{\sum_{j=0}^{t} \binom{n}{j}(q-1)^j} \)

  • The total number of vectors in the space — everything a receiver could possibly see.
  • The size of one Hamming sphere of radius t, counting all the vectors that must decode to a single codeword.
  • The spheres cannot overlap or decoding would be ambiguous, so their total size is at most the whole space. Nothing deeper than counting.
  • A perfect code: the spheres tile the space with nothing left over. Rare, and the complete list is known.
  • Whether such a code can be constructed, or decoded efficiently. Every bound in the chapter is about existence, and existence is the easy part.

A packing argument, and the same one that appears in sphere packing, lattices and combinatorial design — which is why coding theory keeps meeting the rest of mathematics.

72. What does 'corrects t errors' not tell you?

Missing information

A code is specified as correcting up to 3 errors.

Discussion prompt

What still needs to be known before it can be used?

Hint: Errors of what kind, detected how, at what cost.

Answer:

What an 'error' is. Bit errors, symbol errors and erasures are different quantities. A Reed-Solomon code correcting 16 symbol errors handles far more bit errors when they cluster.

What happens beyond t. Some codes fail silently, miscorrecting to the wrong codeword; others detect that they cannot decode. The difference matters enormously for a system that must know when to give up.

Whether an efficient decoder exists. The definition only says nearest-neighbour decoding gives the right answer, not that the nearest neighbour is findable. For a random linear code it is not.

The code rate. Correcting 3 errors at rate 0.9 and at rate 0.1 are entirely different propositions, and the rate is what the bandwidth budget cares about.

And the error model of the actual channel. A code tuned for random errors can perform badly on bursts and vice versa — which is why the CD design pairs a code with an interleaver rather than just picking a stronger code.

73. What McEliece's security rests on

Two truths and a lie

Two of these are wrong.

Eliminate the wrong options

Which statement is correct?

  • a. Decoding a general linear code is hard, and S and P disguise which code G₁ generates
  • b. The security comes from the difficulty of inverting the matrix S
  • c. The system is secure because the error vector cannot be guessed

Survives elimination: a

Why: The trapdoor is possession of an efficient decoder for a specific code, and the disguise is what prevents anyone else from recognising which code it is. Both halves are needed: a public code with a known decoder would be useless, and a hidden code with no decoder would be undecryptable by its own owner.

74. Why did McEliece never get deployed?

Commit first

The system dates from 1978, has never been broken, and is fast.

Predict first

What kept it out of use?

  • It was broken shortly after publication
  • The public key is around 67 kilobytes, against a few hundred bytes for RSA
  • It is too slow to encrypt
  • It requires a trusted authority

Correct: The public key is around 67 kilobytes, against a few hundred bytes for RSA

Encryption and decryption are actually faster than RSA's, since they are matrix operations over GF(2) rather than modular exponentiation. Speed was never the problem.

And the situation has changed. Classic McEliece is a NIST post-quantum finalist precisely because its assumption is old, well-studied and unbroken — and modern networks can afford a large key when the alternative is a broken one.

Its keys are still the largest among the candidates, around a megabyte for the highest security levels. The trade is now openly discussed as key size against confidence, and for long-lived static keys the confidence can be worth it.

Why: G₁ is a 1024 × 524 binary matrix, so the public key runs to tens of kilobytes. In the 1980s and 1990s that was prohibitive for certificates, smart cards and network protocols, and RSA's compact keys won on that ground alone.

75. Explain error correction to someone who has not seen it

Explain it

A colleague asks how a scratched CD still plays.

Discussion prompt

Explain it, and give them a way to see why it must cost something.

Hint: Start with spelling.

Answer:

Start with words. If someone writes 'recieve', you know what they meant, because English words are spread out — no other word is one letter away. That spacing is redundancy, and it is what lets you correct.

A code does this deliberately. It uses only certain patterns as valid, chosen so that any two differ in many places. A corrupted pattern is close to exactly one valid one, and you snap to it.

The cost is that most patterns are wasted. If every possible pattern were valid there would be no spacing and no correction — which is exactly what makes the rate less than one.

On a CD the spacing is deliberate in two ways. The code adds check bytes, and consecutive bytes of a codeword are written far apart on the disc — so a scratch damages one byte here and one there rather than wiping out a whole codeword.

And the check they can do themselves: ask why a code correcting more errors must waste more bandwidth. Answer: more spacing between valid patterns means fewer valid patterns fit, so each transmission carries less. That is the Singleton bound in words.

Figure (svg): Hamming spheres around codewords, showing why minimum distance controls detection and correction.

One number, d(C), controls everything a code can do — and the two thresholds differ by a factor of two.

76. Why is the syndrome the error's address?

Explain it to yourself

In a Hamming code, the syndrome of a single-error vector is a column of H, and that column's position is the error's position.

Discussion prompt

Explain why, from the construction of H.

Hint: What is e Hᵀ when e has a single 1?

Answer:

The syndrome depends only on the error: yHᵀ = (c + e)Hᵀ = eHᵀ, since codewords are annihilated.

If e has a single 1 in position j, then eHᵀ is exactly the j-th column of H. Multiplying a matrix by a vector with one 1 selects a column.

And in a Hamming code every nonzero m-tuple appears as a column exactly once, by construction. So the map from position to syndrome is a bijection onto the nonzero syndromes.

Therefore the syndrome names the position uniquely, and there is nothing to look up beyond finding which column it is — which, if the columns are arranged as binary numerals, is just reading the syndrome as a number.

Which explains why n = 2ᵐ − 1. The columns must be all the nonzero m-tuples, and there are exactly 2ᵐ − 1 of those. The code's length is forced by the requirement that every single-error syndrome be distinct — the construction and the decoder are one idea.

Figure (svg): The general Hamming code construction: a parity check matrix whose columns are every nonzero binary m-tuple.

The construction explains the decoder: the syndrome is a column of H, and columns are in bijection with bit positions.

77. Why does the ISBN use modulus 11?

Prediction

Predict first

What would break if the modulus were 10?

  • Nothing — the check would work the same
  • Some single errors would go undetected, since ke can be ≡ 0 mod 10 with k and e both nonzero
  • The check digit could not be computed
  • Transpositions would still be caught but single errors would not

Correct: Some single errors would go undetected, since ke can be ≡ 0 mod 10 with k and e both nonzero

The same primality argument covers transpositions, where the checksum shifts by (k−l)(aₗ−aₖ). Both factors are nonzero when the positions and digits differ, so the product cannot vanish.

The cost is the digit X. The check value can be 10, which is not a decimal digit, so a special symbol is needed — a small price for detecting every single error and every transposition.

The 13-digit ISBN works mod 10 and gives this up. It uses weights 1 and 3 alternately, which catches all single errors but misses transpositions of digits differing by 5. The trade was made for compatibility with barcode standards.

Why: Mod 10, take k = 2 and e = 5: then ke = 10 ≡ 0, so an error of 5 in the second position is invisible. Mod 11 the modulus is prime, so a product of two nonzero residues is never zero, and every single error changes the checksum.

78. Why does the minimum distance control everything?

Socratic

Detection, correction and the bounds are all expressed in terms of one number.

Discussion prompt

Why is d(C) the right quantity, rather than, say, the average distance?

Hint: Consider what an adversarial channel would do.

Answer:

Because the worst case is what fails. A code is only as good as its closest pair of codewords: those two are the ones a small number of errors can confuse, however far apart everything else is.

Average distance would be misleading. A code could have a huge average distance and one pair at distance 2, and that single pair caps detection at one error regardless of the average.

And it makes the geometry clean. Spheres of radius t around every codeword are disjoint exactly when d ≥ 2t+1, and that one condition gives both thresholds and the sphere-packing bound.

For linear codes it is even cheaper to compute, since d(C) is the minimum nonzero weight — a search over M codewords rather than M² pairs.

The general habit is worth naming: when a guarantee must hold for every input, the right parameter is a minimum or a maximum, never an average. The same reasoning made min-entropy the right measure for password guessing in Chapter 20.

79. Sort by what each matrix does

Definition probe

Linear codes involve several matrices with distinct jobs.

Sort into buckets

Sort each.

Defines or checks the code
G = [I_k | P]; H = [−Pᵀ | I_{n−k}]
Disguises the code
S in the McEliece system; P, the permutation matrix in McEliece
code
G's rows span the code and turn a message into a codeword; H's rows span the dual and test membership by vHᵀ = 0. Between them they define the code completely.
hide
Neither S nor P changes what the code can do — SGP generates an equivalent code with the same parameters. Their only job is to make the published matrix unrecognisable, so that nobody else can apply the fast decoder.

80. A perfect code miscorrects

Anomaly

A Hamming [7,4] codeword is sent and two bits are corrupted.

Predict first

What does the decoder do?

  • It reports that decoding failed
  • It silently corrects to the wrong codeword, since every vector lies in some sphere of radius 1
  • It corrects both errors
  • It returns the received word unchanged

Correct: It silently corrects to the wrong codeword, since every vector lies in some sphere of radius 1

So perfection has a hidden cost: there is no room left for a decoding failure signal, which for many applications is the more useful outcome.

The fix is the extended Hamming code. Add one overall parity bit, making d = 4. Now single errors are corrected, double errors produce a syndrome pattern that matches no single column, and the decoder can say 'I cannot decode this'.

Which is why extended codes are what actually get deployed in memory error correction — a SECDED code, single error correcting and double error detecting, is the industry standard for exactly this reason.

Why: Perfection means the radius-1 spheres tile the whole space with nothing left over — so every received vector, including a double-error one, lies in exactly one sphere and gets 'corrected' to its centre. The decoder cannot tell that anything is wrong.

81. Complete the linear code machinery

Faded example

Four blanks.

Fill in the blanks

A linear [n, k] code is a subspace of Fⁿ. Its generator matrix in systematic form is [I_k | P], and its parity check matrix is [−Pᵀ | I]. A vector v is a codeword exactly when vHᵀ = 0. For a linear code, the minimum distance equals the smallest weight of a nonzero codeword.

Why: The last one is where the computational saving lives: distances between pairs become weights of single codewords, because v − w is itself a codeword.

82. How big is the McEliece key?

Estimation

The public key G₁ is a 524 × 1024 binary matrix.

Predict first

Roughly how large is it?

  • About 1 kilobyte
  • About 67 kilobytes
  • About 5 megabytes
  • About a gigabyte

Correct: About 67 kilobytes

Systematic form helps a little. If G₁ is put in the form [I | Q] only the 524 × 500 part need be transmitted, roughly halving the key — but it also leaks structure, so the trade is not free.

Modern Classic McEliece parameters are larger still, up to about a megabyte at the highest security level. The proposal is explicit that this is the price of an assumption nobody has broken in nearly fifty years.

And where it fits is niche but real: static long-term keys that are distributed once and used many times, rather than keys exchanged in every TLS handshake.

Why: 524 × 1024 = 536 576 bits, or about 67 kilobytes. Against a few hundred bytes for an RSA public key, that is two to three orders of magnitude larger — and in the 1980s it was decisive.

83. Reading the code rate

Notation

One ratio decides whether a code is affordable.

Annotate

On: \( R = \frac{\log_q M}{n} \)

  • How many symbols of real information the code carries. For a linear [n, k] code it is simply k.
  • How many symbols are actually transmitted. The difference n − k is the redundancy paid for.
  • R ≤ 1 − (d−1)/n, so demanding a large relative distance directly caps the rate. There is no way round it.
  • Correction. Mariner's 6/32 corrected 7 errors out of 32 symbols, which is what let a very weak signal get through at all.
  • Decoding cost, burst tolerance, or whether failures are silent. Two codes with the same rate can behave completely differently on a real channel.

Rate is the first number quoted about a code and never the only one that matters — the error model decides the rest.

84. Structure buys decodability

Pattern

Every step of the chapter adds structure, and every addition buys a faster decoder.

  1. An arbitrary code requires comparing a received word with every codeword — hopeless beyond toy sizes.
  2. A linear code is a subspace, so it is stored as a matrix, its distance is a minimum weight, and its syndrome depends only on the error. Decoding becomes a table lookup of q^{n−k} rows.
  3. A cyclic code is a set of polynomial multiples of g(X) mod Xⁿ − 1, so shifting is multiplication by X and constructing codes means factoring Xⁿ − 1.
  4. A BCH code has consecutive powers of α as roots, so the BCH bound designs the minimum distance rather than discovering it — and the algebra yields multi-error decoders.
  5. A Reed-Solomon code takes this to the limit, meeting the Singleton bound exactly and working at symbol level, which is what makes it right for bursts.

And the same progression is what makes McEliece possible. A random linear code has excellent parameters and no efficient decoder; a Goppa code has both. Hiding one inside the appearance of the other is the trapdoor.

The recurring exchange is structure for speed, paid for in security margin. It appeared with RSA's multiplicativity, with pairing curves, and with ring-LWE — and here it appears in reverse: the structure is the secret, and disguising it is the whole system.

Figure (svg): A linear code as the row space of a generator matrix, with the parity check matrix annihilating it.

Linearity is not a mathematical nicety — it turns decoding from a search into a matrix multiplication.

85. More redundancy always means better protection

Trap

The trap

The trap. Error correction works by adding redundancy, and more redundancy means more errors can be corrected. So a code with a lower rate is always safer, and when in doubt, add more check symbols.

The first clause is true and the conclusion does not follow.

The fix

Redundancy must match the error model. A bit-level code and a symbol-level code with identical rates perform completely differently on bursts. Reed-Solomon corrects 121 burst bit-errors that a bit-level code of the same rate cannot touch.

Rate costs are real and sometimes decisive. Mariner's rate of 6/32 was correct for a very weak signal and would be absurd on a fibre link. The book's coding gain calculation is precisely the question of whether the redundancy is worth its bandwidth in this channel.

A perfect code has no margin for detection. Every received vector decodes to something, so a double error is silently miscorrected. Adding one parity bit to make an extended Hamming code sacrifices perfection and gains the ability to say 'I cannot decode this' — which is often more valuable than another correction.

And placement matters as much as quantity. Coding inside encryption is nearly useless; interleaving with the same code and the same rate can be transformative. The redundancy has to be arranged against the errors that actually occur.

The accurate statement is narrow: redundancy chosen to match the channel's error model, placed outside the encryption, and interleaved against bursts, buys correction proportional to the distance it creates. Every clause is doing work, and 'add more check bits' does none of them.

86. Check: the two thresholds

Check

Work it out before clicking.

Check your understanding

A code has minimum distance 7. How many errors can it correct?

  • A. 7
  • B. 6
  • C. 3 (correct)
  • D. 4

Answer: C

Why: Correction needs d ≥ 2t + 1, so 7 ≥ 2t + 1 gives t ≤ 3. It can also detect up to 6 errors, since detection needs only d ≥ s + 1. That gap between 6 and 3 is the factor of two that runs through the whole chapter.

Why A tempts people
That is not either threshold — d itself is the distance, not a count of correctable errors.
Why B tempts people
That is the number of detectable errors, d − 1. Detecting is much cheaper than correcting.
Why D tempts people
Four errors could carry a received word past the halfway point between two codewords, making the nearest neighbour the wrong one.

87. Check: what the syndrome sees

Check

Consider a linear code with parity check matrix H.

Check your understanding

A codeword c is sent and y = c + e is received. What is yHᵀ?

  • A. cHᵀ
  • B. eHᵀ — it depends only on the error, not on which codeword was sent (correct)
  • C. Zero, always
  • D. y itself

Answer: B

Why: Codewords satisfy cHᵀ = 0 by definition, so yHᵀ = cHᵀ + eHᵀ = eHᵀ. This is the property that makes syndrome decoding possible: one table indexed by syndrome covers every codeword at once, rather than a table with qⁿ entries.

Why A tempts people
That term is zero — it is what being a codeword means.
Why C tempts people
Zero only when e is itself a codeword, which includes e = 0. A nonzero syndrome is exactly how an error is detected.
Why D tempts people
The syndrome has length n − k, not n, and it is a linear image of y rather than y itself.

88. Check: the McEliece trapdoor

Check

Consider what Bob knows that Eve does not.

Check your understanding

What is the trapdoor in the McEliece cryptosystem?

  • A. The matrix S
  • B. Knowing which code G₁ generates, and hence having an efficient decoder for it (correct)
  • C. The weight of the error vector
  • D. The permutation P

Answer: B

Why: S and P disguise the code so that G₁ looks like the generator of a random linear code, and decoding a random linear code is NP-hard. Bob's advantage is that he knows the disguise and can strip it, revealing a Goppa code with a fast decoder.

Why A tempts people
S alone provides little security — the book notes that once a decoder for the code generated by GP is known, a chosen-plaintext attack recovers S as with a Hill cipher.
Why C tempts people
The weight t is public; Alice must know it to build a valid ciphertext.
Why D tempts people
P is part of the disguise but not the trapdoor by itself. Undoing the permutation is useless without an efficient decoder for the underlying code.

89. Build the code family table

Connect it up

Eleven codes, one framework.

Draw it

Make a table with a row for each code in this chapter: repetition, parity, two-dimensional parity, Hamming, ISBN, Hadamard, Golay, cyclic, BCH, Reed-Solomon, Goppa. For each give (n, M, d) or [n, k, d] where known, the code rate, errors detected and corrected, and the structure its decoder exploits. Then mark which are linear, which are cyclic, which are perfect and which are MDS. Finish by writing the McEliece scheme in four lines and marking which row of the table supplies its trapdoor.

The classification columns are the point: linear, cyclic, perfect and MDS are four independent properties, and knowing which a code has tells you immediately what tools apply to it.

90. Exit ticket

Exit ticket

One question, on why this chapter is in this book.

Predict first

What makes an error-correcting code into a public-key cryptosystem?

  • Encrypting the codewords
  • An efficient decoder that only the key holder possesses, applied to deliberately introduced errors
  • Using a code with a very large minimum distance
  • Publishing the parity check matrix

Correct: An efficient decoder that only the key holder possesses, applied to deliberately introduced errors

Why: McEliece publishes a scrambled generator matrix, so anyone can encode and add errors, but only the owner knows which code it really is and therefore how to decode. Decoding a general linear code is NP-hard; decoding a Goppa code you recognise is fast. That gap is the trapdoor.

91. What to carry into Chapter 25

Recap

Redundancy, deliberately added and systematically exploited.

Chapter 25 next, and it is the last. Quantum techniques: how a quantum computer would break factoring and discrete logs — undoing most of this course — and how quantum key distribution offers a form of security that no computation can defeat.

Figure (svg): The trade McEliece makes: very fast operations and strong security, against an unusually large public key.

One of the oldest public-key systems, never broken, and almost never used — the key size decided it.

Sources

  1. Introduction to Cryptography with Coding Theory, 3rd edition — Wade Trappe and Lawrence C. Washington — Pearson, 2020 (ISBN 978-0-13-485906-4)
  2. Chapter 24 — Error Correcting Codes (sections 24.1-24.13) — Trappe & Washington, 3rd edition, pp. 461-508

Want this taught 1-on-1? Alexander tutors Cryptography — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108