Chapter 8 of Trappe & Washington: AES as ten to fourteen rounds of four layers on a 4x4 byte state, with SubBytes, ShiftRows, MixColumns and AddRoundKey each given its own slide and its own worked example. Covers the GF(2^8) arithmetic the cipher runs on, the inversion of every layer including InvMixColumns, the reason the final round omits MixColumns, and Section 8.4's published design rationale — the S-box built algebraically to avoid DES's trapdoor suspicions, and a round count set at the attack frontier plus four.
Subject: Cryptography · 60 slides · diagram-first lesson
Open the interactive version of this deck
Title
Cryptography · Chapter 8
Rijndael: four layers, one algebraic S-box, and a design process built to answer DES's unanswered questions
Objectives
Chapter 7 ended with five questions for reading a block cipher. AES answers every one of them differently from DES, and its designers wrote down why — which is itself the chapter's biggest departure from what came before.
Figure (svg): DES and AES compared on structure, S-box design, block size and design process.
Warm-up
You have Chapter 7's five questions and twenty years of hindsight. NIST opens a competition in 1997.
Discussion prompt
Name three changes you would require of a replacement, and say what each fixes.
Hint: One parameter, one structural choice, and one thing about the process itself.
Answer:
A longer key, and a choice of lengths. 56 bits was outdated within months of publication. AES offers 128, 192 and 256, all beyond exhaustive search by an enormous margin, and the choice is written into the standard so it need not be renegotiated later.
A larger block. 64 bits means the birthday bound bites at 2³² blocks — 32 GB — which is the Sweet32 attack that eventually retired triple DES. AES uses 128.
An open process with published rationale. DES's S-box changes turned out to be a strengthening, and nobody outside could tell for fifteen years. AES ran as a five-year public competition in which candidates were expected to explain every design decision.
All three are in the standard, and the third is the one with the longest reach: Section 8.4 exists precisely because the designers were required to write it.
Section
Section 8.1 · pp. 160-161
Concept
Rijndael, designed by Daemen and Rijmen, selected in 2000 and standardised as AES. Keys of 128, 192 or 256 bits; the book restricts to 128 for simplicity.
Ten rounds for a 128-bit key, twelve for 192, fourteen for 256. Each round takes 128 bits in and gives 128 bits out, using a round key derived from the original. There is also a round 0 whose key is the original key itself.
Four layers make up a round, and each has one job:
Figure (svg): One AES round as four layers in sequence: SubBytes, ShiftRows, MixColumns, AddRoundKey.
Notation
The whole cipher is one line, once the abbreviations are unpacked.
Annotate
On: \( \text{ARK}; \; \bigl[\text{SB},\text{SR},\text{MC},\text{ARK}\bigr]^{9}; \; \text{SB},\text{SR},\text{ARK} \)
Compare with DES's line: IP, sixteen Feistel rounds, swap, IP⁻¹. No permutation here exists purely for implementation convenience.
Prediction
AES-128 uses ten rounds, AES-192 twelve, AES-256 fourteen. The block size is 128 bits in all three.
Predict first
Why should the round count depend on the key length rather than the block length?
Correct: More key material needs more rounds for every key bit to have influenced every output bit, and a longer key also gives related-key attacks more to work with
This is a genuinely counter-intuitive result: the best known related-key attacks are on AES-256, not AES-128, so the larger key is not uniformly stronger.
For ordinary use the distinction does not matter — none of these attacks is practical, and related-key attacks require an adversary who can force encryptions under related keys. But it is a good reminder that 'bigger key' and 'more secure' are not synonyms, which is Chapter 1's argument once more.
Why: The key schedule stretches the key across the round keys, and a 256-bit key is spread over more round-key material than a 128-bit one. More rounds ensure every key bit has been mixed thoroughly, and they add margin against related-key attacks, which have historically been more effective against AES-256 than AES-128 precisely because its key schedule is comparatively simpler relative to its length.
Figure (svg): The number of AES rounds against the number of rounds any known attack can beat brute force on.
Definition probe
Every major decision in Rijndael is a response to something specific about its predecessor.
Sort into buckets
Sort each AES design choice by the DES problem it addresses.
Section
Section 8.2 · pp. 161-166
Concept
Group the 128 input bits into 16 bytes and arrange them in a 4 × 4 matrix, filled column by column — a₀,₀, a₁,₀, a₂,₀, a₃,₀, then a₀,₁, and so on.
Each byte is an element of the finite field GF(2⁸) from Section 3.11. For AES's purposes three facts about it matter:
The book notes that other choices of irreducible polynomial would presumably give equally good algorithms — the particular one is a convention, not a security parameter.
Figure (svg): The 128-bit block arranged as a four by four matrix of bytes, filled column by column.
Worked example
The bytes go into the matrix down the columns, which trips people up because matrices are usually read across rows.
Take the 16 bytes in order: a₀,₀ a₁,₀ a₂,₀ a₃,₀ a₀,₁ a₁,₁ …
Why: The first four bytes fill the first column, top to bottom.
So byte 0 is at row 0 column 0, byte 1 at row 1 column 0, byte 4 at row 0 column 1
Why: Byte index i goes to row i mod 4, column ⌊i/4⌋.
A 16-byte input therefore fills four columns of four
Why: And a 128-bit key fills a state-shaped matrix in exactly the same order.
Note what this means for the layers: ShiftRows acts across the rows, so it mixes bytes that arrived far apart in the input
Why: Bytes 0, 5, 10 and 15 end up in the first column after ShiftRows — one from each original column.
Verify: count: 16 bytes in, 16 positions filled, none twice
Why: Column-major order is a convention, but it is the convention the layers were designed around: ShiftRows only spreads bytes usefully because the input filled the columns rather than the rows.
Figure (svg): The 128-bit block arranged as a four by four matrix of bytes, filled column by column.
Concept
Every byte of the state is replaced using a single 16 × 16 lookup table. Write the byte as abcdefgh; the top four bits give the row and the bottom four the column, both numbered 0 to 15.
The book's example: the input byte 10001011 means row 8 and column 11, whose entry is 61 — and 61 in binary is 00111101. That is the output.
Unlike DES's eight tabulated 6-to-4 boxes, AES has one box taking 8 bits to 8, and it is a permutation. It has to be: AES is not a Feistel cipher, so every layer must be individually invertible.
And it is not tabulated arbitrarily. It is built from the map x ↦ x⁻¹ in GF(2⁸), followed by an affine transformation over GF(2). Section 8.4 gives the reason.
Figure (svg): The SubBytes lookup: the input byte's top four bits pick the row and its bottom four pick the column of the S-box table.
Worked example
Follow the book's example all the way through.
Input byte 10001011
Why: One of the sixteen bytes of the state.
Split it: the first four bits 1000 give the row, the last four 1011 give the column
Why: 1000 is 8 and 1011 is 11, so we want row 8 and column 11 — the ninth row and the twelfth column, since both are numbered from 0.
The table entry there is 61
Why: The book's table is in decimal.
Convert to binary: 61 = 32 + 16 + 8 + 4 + 1 = 00111101
Why: Eight bits out for eight bits in.
\[ \texttt{10001011} \;\longmapsto\; \texttt{00111101} \]
Verify: the map is a bijection on all 256 byte values
Why: It must be, because AES inverts each layer separately to decrypt. This is the structural difference from DES, whose S-boxes throw away two bits per box and rely on the Feistel construction to restore invertibility.
Figure (svg): The SubBytes lookup: the input byte's top four bits pick the row and its bottom four pick the column of the S-box table.
Socratic
DES's S-boxes were tables with no published rationale. AES's is an algebraic formula.
Discussion prompt
What does defining the S-box by a formula buy, and why inversion specifically?
Hint: One answer is about trust and one is about mathematics.
Answer:
The trust answer, which the book states directly: the S-box was constructed in an explicit and simple algebraic way so as to avoid any suspicions of trapdoors — the mysteries about the DES S-boxes that haunted that cipher for fifteen years. A formula anyone can recompute cannot hide a hand-chosen weakness.
The mathematical answer: x ↦ x⁻¹ in GF(2⁸) is about as non-linear as a byte permutation can be. Its differential and linear properties are provably near-optimal, so it resists both Section 7.3's differential cryptanalysis and linear cryptanalysis, and also interpolation attacks.
Why an affine step is added afterwards. Pure inversion has algebraic structure that is too clean — it has few terms when written as a polynomial, and fixed points. Composing with an affine map over GF(2) breaks that structure without harming the differential properties.
And 0 is handled by convention: 0 has no inverse, so it is mapped to itself before the affine step, keeping the map a permutation.
The general lesson worth taking: a design decision that can be recomputed from a stated principle is auditable; a table is not. That is a security property of the process rather than the algorithm, and it is the thing AES most clearly improved on.
Anomaly
The map x ↦ x⁻¹ in GF(2⁸) sends 1 to 1, since 1 is its own inverse. It also has to do something with 0, which has no inverse.
Predict first
Why does AES compose the inversion with an affine transformation afterwards?
Correct: To remove fixed points and break the map's simple algebraic description, without harming its differential properties
Fixed points matter because a byte passing through unchanged is a small piece of structure an attacker can look for, exactly as Enigma's lack of fixed points was in Chapter 2.
The algebraic simplicity matters because of interpolation attacks, which reconstruct a cipher as a polynomial. Section 8.4 lists these explicitly among what the S-box design resists.
The pattern is worth naming: take a component with provably good statistics, then compose with something cheap to spoil its algebraic tidiness. That is a standard move in cipher design.
Why: Inversion is already non-linear and already a permutation once 0 is mapped to itself by convention. What it is not is structurally messy: as a polynomial over GF(2⁸) it has very few terms, and it has fixed points such as 0 and 1. The affine step scrambles that description — removing fixed points and raising the algebraic complexity — while leaving the differential and linear properties, which come from the inversion, untouched.
Concept
The four rows of the state are rotated cyclically to the left by offsets of 0, 1, 2 and 3.
\[ \begin{pmatrix} b_{00} & b_{01} & b_{02} & b_{03} \\ b_{11} & b_{12} & b_{13} & b_{10} \\ b_{22} & b_{23} & b_{20} & b_{21} \\ b_{33} & b_{30} & b_{31} & b_{32} \end{pmatrix} \]
Row 0 does not move. Row 1 moves one place left, row 2 two places, row 3 three places, all wrapping round.
The point is that MixColumns works down columns. Without ShiftRows, the four bytes of a column would stay together forever and AES would decompose into four independent 32-bit ciphers. ShiftRows is what breaks each column apart so that MixColumns mixes different bytes each round.
Section 8.4 adds that ShiftRows was included specifically to resist truncated differentials and the Square attack — the latter named after Square, Rijndael's own predecessor.
Figure (svg): ShiftRows rotating row zero by nothing, row one by one, row two by two and row three by three positions to the left.
Worked example
Take a state whose entries are labelled by position and follow each row.
Row 0: no shift. b₀₀ b₀₁ b₀₂ b₀₃ stays as it is
Why: The offset is 0, so this row's bytes keep their columns.
Row 1: shift left by 1. b₁₀ b₁₁ b₁₂ b₁₃ becomes b₁₁ b₁₂ b₁₃ b₁₀
Why: b₁₀ wraps round to the end.
Row 2: shift left by 2, giving b₂₂ b₂₃ b₂₀ b₂₁
Why: Two bytes wrap.
Row 3: shift left by 3, giving b₃₃ b₃₀ b₃₁ b₃₂
Why: Equivalently a shift right by one.
Verify: look down the first column: it now holds b₀₀, b₁₁, b₂₂, b₃₃
Why: Four bytes that came from four different columns. Before ShiftRows they were b₀₀, b₁₀, b₂₀, b₃₀ — all from one. That is the entire purpose: MixColumns is about to mix this column, and it now has material from across the whole state.
Figure (svg): ShiftRows rotating row zero by nothing, row one by one, row two by two and row three by three positions to the left.
Prediction
SubBytes acts on single bytes. MixColumns acts within a column. ShiftRows moves bytes along rows.
Predict first
Which of these operations moves information between columns?
Correct: ShiftRows only
This is why the pair ShiftRows-then-MixColumns is the diffusion engine: one moves bytes between columns, the other mixes within a column, and alternating them reaches every byte in two rounds.
It also explains why Section 8.4 lists ShiftRows' purpose as resisting truncated differentials and the Square attack — both of which exploit exactly the kind of column-local structure that ShiftRows destroys.
Why: SubBytes changes a byte's value but never its position. MixColumns mixes the four bytes of a column with each other and nothing else. AddRoundKey XORs position-for-position. So ShiftRows is the sole layer moving data across columns — and without it AES would be four independent 32-bit ciphers running side by side, each with a quarter of the state.
Figure (svg): ShiftRows rotating row zero by nothing, row one by one, row two by two and row three by three positions to the left.
Concept
Treat each column of the state as a vector of four elements of GF(2⁸) and multiply it by a fixed matrix.
\[ \begin{pmatrix} 02 & 03 & 01 & 01 \\ 01 & 02 & 03 & 01 \\ 01 & 01 & 02 & 03 \\ 03 & 01 & 01 & 02 \end{pmatrix} \begin{pmatrix} c_{0j} \\ c_{1j} \\ c_{2j} \\ c_{3j} \end{pmatrix} = \begin{pmatrix} d_{0j} \\ d_{1j} \\ d_{2j} \\ d_{3j} \end{pmatrix} \]
The arithmetic is field arithmetic: addition is XOR, and multiplication by 02 and 03 is a shift-and-XOR — both extremely cheap in hardware, which is why those coefficients were chosen.
The diffusion guarantee is precise and it is the reason the matrix looks like this: change one input byte and all four output bytes change. Change two input bytes and at least three output bytes change. That is a property of the matrix, provable rather than hoped for.
Figure (svg): MixColumns multiplying each column of the state by a fixed four by four matrix over GF(2 to the 8).
Worked example
The matrix is not arbitrary; every entry is chosen against a constraint.
Multiplication by 01 is the identity — free
Why: No work at all.
Multiplication by 02 is a left shift, with a conditional XOR of 00011011 if the top bit was set
Why: That conditional XOR is the reduction modulo X⁸ + X⁴ + X³ + X + 1. One shift, one compare, one XOR.
Multiplication by 03 is multiplication by 02 then XOR with the original, since 03 = 02 ⊕ 01
Why: So it costs one extra XOR on top of the 02 case.
The matrix is circulant — each row is the previous one rotated — so one routine handles all four rows
Why: This is what lets a hardware implementation use one small circuit four times.
Verify: check the diffusion claim on the matrix's structure
Why: Every column of the matrix contains at least three non-zero entries other than 01, which forces a single changed input byte into all four outputs. The design constraint was: cheapest possible coefficients that still achieve maximal diffusion — and 01, 02, 03 is the answer.
Figure (svg): A byte as an element of GF(2 to the 8), with addition being XOR and the field defined by an irreducible polynomial of degree eight.
Invariant
Flip a single bit of one byte of the state and follow how far its influence has reached after each layer.
Step through it
How many rounds does full diffusion take?
Two, as Section 8.4 states — every one of the 128 output bits depends on every one of the 128 input bits after two rounds. DES needs about five, because half its block sits idle each round.
Concept
AddRoundKey is the simplest layer: XOR the 128-bit round key into the state. It is the only place the key enters, and it is its own inverse.
The key schedule produces the round keys — eleven of them for AES-128, since there are ten rounds plus round 0. Its design choices are stated in Section 8.4:
That last point is easy to underrate. Without round constants, a key with a symmetric structure could produce identical round keys, and identical rounds mean the cipher has a structure an attacker can exploit — a slide attack.
Figure (svg): The key schedule expanding one 128-bit key into eleven round keys, using the S-box and a round constant at each step.
Estimation
Ten rounds, each with SubBytes on 16 bytes, ShiftRows, MixColumns on four columns, and AddRoundKey.
Predict first
Roughly how many byte-level operations is that, before hardware acceleration?
Correct: About 1000
Compare DES: sixteen rounds of bit-level permutations and 6-to-4 table lookups, which is far more awkward in software because bit permutations are cheap in hardware and expensive in registers.
AES was explicitly designed to be fast in both — byte-oriented operations, a small table, and coefficients of 01, 02 and 03. That dual target was one of the competition's stated criteria.
Why: Per round: 16 S-box lookups, a free byte permutation, four column multiplications of about 16 byte-operations each, and 16 XORs — on the order of 100 byte-operations. Ten rounds gives roughly 1000 for a 16-byte block, or about 60 operations per byte. That is why AES was fast enough in software to be adopted universally, and why AES-NI, which does a whole round in one instruction, was such a large further step.
Translation
Because AES is not a Feistel cipher, decryption must invert each layer individually.
Match the pairs
Why: Three of the four inverses cost something extra. InvSubBytes needs a second 256-byte table; InvMixColumns' coefficients 09, 0B, 0D and 0E take several shift-and-XOR steps each where encryption's 02 and 03 take one or two. So AES decryption is meaningfully slower than encryption — which is a concrete reason to prefer Chapter 6's CTR, CFB and OFB modes, none of which ever calls the decryption function.
Fill the middle
Complete the round. Each blank is a layer's single purpose.
Fill in the blanks
\textconfusion diffusion. \quad \textthe key enters ___. \quad \text___ ___.
Why: Shannon's two ingredients plus the key. Every block cipher is an arrangement of exactly these three functions, and being able to name which layer does which is what lets you read a new cipher's specification quickly. Note that AddRoundKey on its own provides no security at all — it is a XOR with a value that becomes known if any of the others fails.
Discrimination
Chapter 7's lesson: find the non-linear component, because that is the cipher.
Sort into buckets
Sort AES's layers.
Ranking
A useful exercise, because the answer is very lopsided and knowing that shapes how you read any cipher.
Put in order
Why: AddRoundKey alone is a one-time pad with a key that repeats every block — no security at all beyond the key's secrecy, and none against known plaintext. ShiftRows only permutes bytes and is worth little on its own, though its unique role in moving data between columns makes it indispensable in combination. MixColumns is a linear map, so it is invertible and keyless. SubBytes is the only non-linear layer, and without it every other layer composes into one affine map that a linear system solves. The ordering is really 'everything else, then the S-box'.
Sorting
Retrieval across Chapters 5 to 8. Every component of every cipher so far falls into one of these.
Sort into buckets
Sort each component.
Three buckets, and reading a new cipher's specification is largely the exercise of filling them in.
Section
Section 8.3 · pp. 166-168
Concept
Because AES is not a Feistel cipher, decryption is genuinely a different algorithm — each layer must be inverted individually.
\[ \text{InvMixColumns} = \begin{pmatrix} 0E & 0B & 0D & 09 \\ 09 & 0E & 0B & 0D \\ 0D & 09 & 0E & 0B \\ 0B & 0D & 09 & 0E \end{pmatrix} \]
Notice the cost asymmetry. Encryption multiplies by 01, 02 and 03 — a shift and an XOR. Decryption multiplies by 09, 0B, 0D and 0E, which take several shifts each. AES decryption is meaningfully slower than encryption, which is a real reason to prefer the modes of Chapter 6 that never call the decryption function.
Figure (svg): AES decryption rewritten to have the same layer order as encryption, by swapping the commuting steps and transforming the round keys.
Worked example
The naive inverse runs the inverse layers in reverse order, and it does not have encryption's shape. Two observations fix that.
Encryption is: ARK; then nine times SB, SR, MC, ARK; then SB, SR, ARK
Why: MixColumns is missing from the last round.
So the naive decryption is: ARK, ISR, ISB; then nine times ARK, IMC, ISR, ISB; then ARK
Why: Each layer inverted, in reverse order. Correct, but structurally different from encryption.
First observation: SB then SR is the same as SR then SB
Why: Because SubBytes acts on one byte at a time and ShiftRows only moves bytes around. Substituting then moving equals moving then substituting. So ISR and ISB can be swapped too.
Second observation: ARK then IMC can be exchanged for IMC then a transformed round key
Why: Because IMC is linear, IMC(x ⊕ k) = IMC(x) ⊕ IMC(k). Apply IMC to the round key in advance and the two steps commute.
Verify: after both rearrangements, decryption reads ARK; nine times ISB, ISR, IMC, ARK; then ISB, ISR, ARK
Why: Exactly encryption's shape with inverted layers and modified round keys — which is why the last round omits MixColumns. With MC in the last round the two sequences would not line up, and an implementation would need two structurally different code paths instead of one parameterised one.
Figure (svg): AES decryption rewritten to have the same layer order as encryption, by swapping the commuting steps and transforming the round keys.
Socratic
The omission looks like an inconsistency, and it is the one thing about AES's structure that surprises people.
Discussion prompt
Give the reason, and say what would be lost by keeping it.
Hint: Think about the shape of the decryption algorithm, and about whether the last MixColumns would add security.
Answer:
The reason is structural symmetry. With MixColumns present in the last round, the rearrangement in the previous slide does not produce encryption's shape, and decryption needs a genuinely different sequence of operations rather than the same sequence with different tables.
And it costs nothing in security. MixColumns is a public, invertible linear map with no key in it. An attacker who has the ciphertext can simply apply InvMixColumns herself, so a final MixColumns adds no confusion, no key dependence, and no work she cannot undo for free.
Which is exactly the same argument as DES's initial permutation — a public invertible step at the boundary is decoration. The difference is that AES's designers removed the decoration and DES's designers kept it for hardware-loading reasons.
The practical payoff: one implementation, one code structure, two sets of tables. On constrained hardware that halves the footprint.
The general principle worth extracting: a public invertible transformation at the very start or end of a cipher contributes nothing. Anything that is not keyed and not the last non-linear step can be peeled off by the attacker.
Error analysis
From an embedded firmware design document.
Annotate
Three of the four decisions trade a large amount of security for a small amount of code, and the fourth trades all of it for none.
Concept
The inverse of MixColumns is multiplication by the inverse matrix, and its entries are where AES's asymmetry lives.
\[ \begin{pmatrix} 0E & 0B & 0D & 09 \\ 09 & 0E & 0B & 0D \\ 0D & 09 & 0E & 0B \\ 0B & 0D & 09 & 0E \end{pmatrix} \]
Compare with encryption's 01, 02, 03. Multiplication by 02 is one shift and a conditional XOR; by 03 it is that plus one more XOR. Multiplication by 0E, 0B, 0D and 09 takes three or four such steps each, because each is a sum of several powers of two.
So InvMixColumns is roughly three times the work of MixColumns, and AES decryption is correspondingly slower than encryption. The asymmetry was a deliberate choice: encryption is the more frequent operation in most protocols, and the designers put the cheap coefficients there.
It also has a practical consequence you have already met. Chapter 6's CTR, CFB and OFB modes never invoke the block cipher's decryption function at all — they only ever generate keystream — so a system using them needs neither InvSubBytes nor InvMixColumns, halving the code and the table space on constrained hardware.
Figure (svg): The encryption and decryption MixColumns matrices side by side, with the cost of each coefficient marked.
Section
Section 8.4 · pp. 168-169
Concept
This section is the chapter's real subject. DES's designers withheld their rationale and it cost fifteen years of suspicion; AES's published theirs, and here it is.
Figure (svg): The number of AES rounds against the number of rounds any known attack can beat brute force on.
Two truths and a lie
Two of these misread the margin argument in ways that lead to bad decisions.
Eliminate the wrong options
Which statement is correct?
Survives elimination: a
Why: The design argument is explicit and quantified: attacks work up to six rounds, none works at seven, ten rounds are specified. That structure — frontier plus margin — is what a well-designed cipher's round count should look like, and it is what lets the community track a cipher's health over decades. DES's sixteen rounds against a frontier around fifteen was a much thinner margin, and it is one of the few places where the older design looks worse on its own terms.
Real world
AES is not merely standardised; it is implemented in silicon on essentially every modern processor.
Discussion prompt
What changed when AES moved into the instruction set, and what security problem did that solve as a side effect?
Hint: Think about how a software implementation looks up S-box entries.
Answer:
Speed: the AES-NI instructions perform a whole round in a single instruction, taking AES from a few cycles per byte to well under one. That is what made encrypting all traffic and all disks practical rather than a considered choice.
And a security problem solved as a side effect: cache-timing attacks. A software AES implements SubBytes as a table lookup, and which table entry is fetched depends on the data — so the pattern of cache hits and misses leaks information about the state. An attacker sharing the machine, even in a different VM, could recover keys this way.
Hardware AES has no table. The S-box is computed in logic in constant time, so there is no data-dependent memory access to observe.
The general point is important beyond AES: an algorithm's security properties are stated about the mathematics, and an implementation can violate them without changing a single output. Timing, cache behaviour, power draw and electromagnetic emission are all channels the specification never mentions.
This is why constant-time implementation is a discipline of its own, and why Chapter 14's theme — sound primitives, failed assembly — keeps recurring.
Trade off
The standard offers three key lengths. Fill the blanks — the honest comparison is less one-sided than it looks.
Comparison matrix
| AES-128 | AES-256 | |
|---|---|---|
| Rounds | 10 | 14 |
| Relative speed | fastest | about 40% slower |
| Brute-force resistance | 2¹²⁸ — beyond any adversary | 2²⁵⁶ — also beyond any adversary |
| Best known related-key attack | none of note | better than on AES-128 |
| Resists a large quantum computer? | Grover halves it to 2⁶⁴ — inadequate | Grover halves it to 2¹²⁸ — still ample |
Rows three and four together are the surprise: on the attack that matters both are equally unbreakable, and on related-key attacks the longer key is weaker. Row five is the real modern argument for AES-256 — it is the one place the extra length buys something identifiable.
Explain it to yourself
Section 8.4 states that treating all bits uniformly diffuses the input bits faster, and that two rounds suffice for full diffusion.
Discussion prompt
Explain the mechanism, and work out roughly why DES needs more.
Hint: Count how much of the block is transformed in one round of each.
Answer:
In a Feistel round, half the block is copied across untouched. L_i is exactly R_{i−1} — no transformation at all. So in one DES round, only 32 of the 64 bits have been through the round function.
In an AES round, all 16 bytes go through all four layers. Nothing is idle. SubBytes touches every byte, ShiftRows moves every row, MixColumns mixes every column.
Counting the spread: in AES, one changed byte becomes four after one MixColumns, and ShiftRows then scatters those four into four different columns, so the second MixColumns fills all sixteen. Two rounds, full diffusion.
In DES, a changed bit enters one or two S-boxes, whose four-bit outputs P scatters — but only into the half of the block that is active, and the other half must wait a round to be touched at all. That roughly doubles the number of rounds needed, and about five are required for a bit to reach everywhere.
What Feistel buys in exchange is that encryption and decryption are the same code, which mattered a great deal in 1975 hardware. AES pays for two code paths and gets faster diffusion, fewer rounds, and a cipher that is easier to analyse layer by layer.
Explain it
A colleague argues that a cipher's security is a mathematical property, so how it was chosen is irrelevant.
Discussion prompt
Make the case that the process is part of the security, using DES and AES as the two data points.
Hint: Ask what evidence exists that a cipher is strong, and where that evidence comes from.
Answer:
The only evidence a cipher is strong is failed attacks by competent people. There is no proof of security for any practical cipher — Chapter 4 established that computational security rests on assumptions. So the evidence is entirely social: who tried, how hard, and for how long.
A closed process produces no such evidence. DES's S-boxes were excellent and nobody could tell. Fifteen years of expert suspicion was aimed at a decision that had strengthened the cipher, and the same authority defended the 56-bit key, which had weakened it.
AES's competition generated the evidence deliberately: fifteen submissions, five years, and every candidate publicly attacked by every other candidate's team. The rationale requirement meant each design decision had to be defended in public.
And it produced something DES never had — a documented margin. Section 8.4 states that attacks reach six rounds and ten are specified, which is a claim anyone can check and track over time.
So the process is not separate from the security; it is the mechanism that produces the evidence for it. Your colleague is right that security is a mathematical property, and wrong that we have any way to establish it other than adversarial public review.
Constraint
You are specifying encryption for medical records that must stay confidential for fifty years, on servers with AES-NI, with a compliance requirement to document every choice.
Discussion prompt
Choose a key length, a mode and an implementation strategy, and give a defensible reason for each.
Hint: The fifty-year horizon changes one of these decisions and not the others.
Answer:
AES-256, and the reason is the horizon rather than classical security. AES-128 is beyond any classical adversary, but Grover's algorithm halves the effective key length against a large quantum computer, taking 128 to 64 — inadequate — and 256 to 128, which remains ample. Over fifty years that possibility has to be priced in.
An authenticated mode: AES-GCM. Chapter 6's conclusion — no mode there provides integrity, and assembling encryption plus a MAC by hand is where padding oracles and ordering mistakes come from. GCM is CTR plus a tag, and it is hardware-accelerated alongside AES-NI.
Hardware AES via AES-NI, both for speed and to avoid the cache-timing channel that table-based software implementations expose on shared servers.
A documented rekeying and algorithm-agility plan. The DES lesson: build in the ability to increase parameters, because over fifty years you will need to. Store an algorithm identifier with every record.
And a nonce discipline enforced structurally, not by convention — GCM's security collapses entirely on nonce reuse, so the counter must be one that cannot reset on a reboot or a VM restore.
Note that four of the five decisions are about the surroundings and one is about the cipher. That ratio is about right, and it is the course's recurring point.
Missing information
The phrase appears on product pages constantly. Chapters 4 through 8 supply the questions it does not answer.
Discussion prompt
List what a reviewer still cannot determine, and say which single omission is most likely to be the real vulnerability.
Hint: Almost none of the questions are about the cipher.
Answer:
Which mode? ECB leaks the pattern of equal blocks; CBC needs a random unpredictable IV; CTR needs a nonce that cannot repeat. The phrase names a primitive and says nothing about how it is used.
Is there integrity protection? No mode in Chapter 6 provides it. Without a MAC, an attacker can flip chosen plaintext bits in the stream modes and mount padding-oracle attacks against CBC.
Where does the key come from? Chapter 5: a key from a weak or time-seeded generator has a fraction of its nominal entropy, and this has broken real systems while the cipher stood.
Is the implementation constant-time? A table-based software AES leaks through cache timing, which recovers keys across VM boundaries. The mathematics is untouched and the key is gone.
How are keys stored and rotated? A key in a config file next to the ciphertext is not a key.
The likeliest real vulnerability is the missing integrity protection, because it is both the commonest omission and the one with a direct, cheap exploit path. The cipher named in the phrase is the one component nobody has broken.
Scale up
Track the best public attack on reduced-round AES-128 against the ten rounds the standard specifies.
Step through it
What would an attack on eight-round AES mean?
That the margin had shrunk from four rounds to two — a significant result, widely reported, and not a break. Being able to make that distinction is what lets you read cryptography news without either panicking or ignoring it.
Picture it
Everything in Section 8.2 on one picture. If you can reproduce this from memory, you can reconstruct the cipher.
Figure (svg): One AES round as four layers in sequence: SubBytes, ShiftRows, MixColumns, AddRoundKey.
Confusion, diffusion, diffusion, key — and ten of these, with the third layer dropped from the last one.
Counterexample
Suppose an implementation omits SubBytes, keeping ShiftRows, MixColumns and AddRoundKey for all ten rounds.
Discussion prompt
Show that the result is breakable with a single known plaintext-ciphertext pair.
Hint: Ask what kind of function each remaining layer is, and what their composition must therefore be.
Answer:
Every remaining layer is linear over GF(2). ShiftRows permutes bytes, MixColumns is matrix multiplication over GF(2⁸) — which is linear over GF(2) as well — and AddRoundKey is XOR.
So the whole ten-round cipher is an affine map: c = A·m ⊕ B·k, where A and B are fixed public matrices computable from the specification, and k is the key.
One known pair gives 128 linear equations in the 128 unknown key bits. Gaussian elimination over GF(2) on a 128 × 128 matrix takes microseconds.
And the round count is irrelevant — a composition of a hundred linear maps is still one linear map. Adding rounds to a linear cipher adds nothing whatever.
This is the fourth time this argument has appeared: the LFSR in Chapter 5, the Hill cipher in Chapter 6, DES-without-S-boxes in Chapter 7, and now AES. The non-linear component is not a component; it is the cipher.
Elimination
Four proposed modifications to a deployed AES implementation.
Eliminate the wrong options
Which one is a real security problem rather than a cosmetic or performance change?
Survives elimination: c
Why: Identical round keys make every round identical, which is precisely what the round constants in the key schedule exist to prevent. A cipher built from identical rounds is vulnerable to slide attacks, where an attacker shifts one encryption against another by a round and matches them up, defeating the round count entirely. Section 8.4 names this: the round constants eliminate symmetries by making each round different. Note that (a) and (b) are changes with no security consequence and (d) has a side-channel one — only (c) breaks the design argument.
Edge cases
Section 8.4 notes that the number of rounds could easily be increased if needed.
Discussion prompt
Suppose an attack reached nine rounds tomorrow. What would actually have to happen, and what does that tell you about designing for change?
Hint: Think about everything that has AES-128's round count baked into it.
Answer:
The algorithm change is trivial. The round function is unchanged; the key schedule already generates round keys by a repeatable rule; adding four more rounds is a loop bound and four more round constants.
The deployment change is enormous. AES-128's ten rounds are in hardware instruction sets, in smartcards, in HSMs, in protocol specifications and in stored ciphertexts. A 12-round AES-128 would be a different algorithm needing a new identifier, and it would not interoperate with anything.
Which is why the practical response would be to migrate to AES-256 — already standardised, already in silicon, already carrying fourteen rounds. The margin is spent by moving to an existing option, not by editing one.
And that is the design lesson: the way to build in room for change is to standardise several parameter sets up front and give every ciphertext an algorithm identifier, so a migration is a configuration change rather than a redesign.
DES had neither. It had one key length and no negotiation mechanism, which is why the response to its weakness was the awkward bridge of triple encryption. AES learned that lesson explicitly, and TLS's ciphersuite negotiation is the same idea at protocol level.
Faded example
The numbers that get quoted. Fill each blank.
Fill in the blanks
AES operates on a block of 128 bits, arranged as a 4 × 4 matrix of bytes filled down the columns. With a 128-bit key it runs 10 rounds, each consisting of SubBytes, ShiftRows, MixColumns and AddRoundKey — except the last, which omits MixColumns. Each byte is an element of the field GF(2⁸).
Why: The block size is fixed at 128 bits for all three key lengths — only the key length and the round count vary. The 4 × 4 byte matrix is what makes the layer descriptions so compact: one acts on entries, one on rows, one on columns, and one XORs the whole thing. And GF(2⁸) is Section 3.11's field, which is why that section exists at all.
Explain it to yourself
DES's S-boxes take six bits to four and are not invertible. AES's S-box takes eight to eight and is.
Discussion prompt
Explain why AES cannot afford a non-invertible layer, and what it gives up as a result.
Hint: Compare how each cipher recovers the plaintext.
Answer:
A Feistel cipher never inverts its round function. It recomputes it. The right half survives the round untouched, so on the way back the same f(R, K) can be calculated again and XORed away. f can destroy information freely.
AES transforms the entire state every round, so there is nothing preserved to recompute from. The only way back is to invert each layer, and a layer that lost information cannot be inverted.
What AES gives up is design freedom. Its S-box must be a permutation on 256 values, its MixColumns matrix must be invertible over GF(2⁸), and ShiftRows must be a bijection on positions. DES's designers could choose any tables they liked.
What it gains is that every bit is transformed every round, so full diffusion arrives in two rounds instead of five — and the cipher is far easier to analyse, because each layer can be reasoned about independently.
The trade in one line: Feistel buys freedom in the round function at the cost of half the block sitting idle; a substitution-permutation network buys speed of diffusion at the cost of every component having to be invertible.
Commit first
Consider a well-implemented system: AES-128 in GCM mode, keys from the OS entropy pool, hardware acceleration, nonces from a persisted counter.
Predict first
Over the next twenty years, which component is most likely to be the cause of a real compromise?
Correct: The nonce discipline — a counter reset by a restore, a rollback or a clone
This is Chapter 4's lesson arriving in a modern setting: the one-time pad's hypotheses were violated by wartime pad duplication, and GCM's are violated by a restored snapshot. In both cases the mathematics is untouched.
It is also why nonce-misuse-resistant modes such as AES-GCM-SIV exist — they degrade to leaking only message equality rather than the whole key, which is the difference between an incident and a catastrophe.
Why: AES has stood since 1998 with the best attacks stuck at six rounds out of ten, and 128 bits is beyond exhaustion by a margin the whole universe cannot close. GCM's tag is sound when used correctly. The nonce is the fragile part: GCM's security collapses entirely on nonce reuse — not degrades, collapses, revealing the authentication key — and counters are reset by VM restores, database rollbacks, image cloning and failovers, none of which look like security events to the people who cause them.
Real world
AES is probably the most widely executed algorithm in the world after basic arithmetic.
Discussion prompt
Name four places it is running on the device you are reading this on, and note which mode each uses.
Hint: Storage, network, memory and authentication all use it, for different reasons and in different modes.
Answer:
Disk encryption — BitLocker, FileVault, LUKS — in XTS mode, which is built on CTR-like seekability because a filesystem must read sector 200 000 without reading the ones before it.
Every HTTPS connection — AES-GCM, which is CTR mode plus an authentication tag, negotiated in the TLS handshake and accelerated by AES-NI.
Wi-Fi — WPA2 and WPA3 use AES-CCMP, counter mode with a CBC-MAC, replacing the RC4-based TKIP that inherited WEP's problems from Chapter 14.
Password managers, secure enclaves and encrypted messaging, generally in GCM or a similar authenticated mode.
Notice the pattern: every one of them is an authenticated mode, and every one is built on CTR or something like it. Chapter 6's five modes are not equally used — the industry converged on the seekable, parallel, padding-free one and then bolted authentication onto it.
So the practical shape of modern symmetric cryptography is: AES as the primitive, counter mode as the construction, a MAC for integrity, and a nonce discipline as the thing that must not fail.
Pattern
Chapter 7 ended with five questions for reading any block cipher. Here they are with both answers side by side.
| Question | DES | AES |
|---|---|---|
| What is non-linear? | 8 S-boxes, 6 bits to 4, tabulated | 1 S-box, 8 to 8, from x inverse in GF(2^8) |
| How does the key enter? | XOR into the expanded half-block | XOR into the whole state, once per round |
| What provides diffusion? | the permutation P | ShiftRows and MixColumns together |
| How many rounds, and why? | 16, against a frontier around 15 | 10, against a frontier of 6 — four rounds of margin |
| What is bookkeeping? | IP, IP inverse, the parity bits | nothing — even the final MixColumns was removed |
The last row is the sharpest difference. DES carried a public invertible permutation at each end for hardware-loading convenience; AES removed the one public invertible step that was not doing work. Neither affects security, and the contrast says a great deal about how carefully each was specified.
Figure (svg): DES and AES compared on structure, S-box design, block size and design process.
Trap
The trap. AES-256 has twice the key bits, so it is twice as secure — or, more carefully, 2¹²⁸ times as hard to brute force. Any system that can afford the 40% slowdown should use it, and any system using AES-128 is cutting a corner.
This is the reasoning behind a great many 'military-grade 256-bit encryption' claims, and behind a great many procurement requirements.
Why it fails. Both key spaces are already beyond exhaustion by margins that dwarf the difference. Chapter 1's arithmetic: 2¹²⁸ at a trillion keys per second takes about 10¹⁹ years. Multiplying an already-impossible number by another impossible number changes nothing an adversary could ever consume.
And on some measures AES-256 is weaker. The best known related-key attacks are against AES-256, not AES-128, because its key schedule is comparatively simple relative to the amount of key it must spread. Neither attack is practical, but the direction is the opposite of what the slogan implies.
There is one real argument for AES-256, and it is not key length as such. Grover's algorithm gives a quantum square-root speed-up on exhaustive search, taking 2¹²⁸ down to 2⁶⁴ — genuinely inadequate — and 2²⁵⁶ down to 2¹²⁸, which is fine. For data with a multi-decade confidentiality requirement that is a defensible reason.
The honest summary: choose AES-256 for a long confidentiality horizon, not because 256 is twice 128. And notice that in every real system the mode, the nonce discipline, the key generation and the integrity protection are all far more likely to fail than either cipher — which is Chapter 6's lesson, and Chapter 14's.
Check
Use the book's row-and-column convention for AES.
Check your understanding
The byte 11010110 arrives at SubBytes. Which row and column of the S-box table are used?
Answer: A
Why: The top four bits give the row and the bottom four the column. 1101 is 13 and 0110 is 6, so it is row 13, column 6 — with both numbered from 0. This is simpler than DES's convention, where the outer two bits gave the row and the inner four the column.
Check
Think about what each layer uniquely provides.
Check your understanding
AES is implemented without ShiftRows, keeping all other layers. What is the consequence?
Answer: B
Why: MixColumns mixes within a column and SubBytes acts within a byte, so with ShiftRows removed no operation ever moves data between columns. Each of the four columns evolves entirely independently, giving four separate 32-bit ciphers — each with a fraction of the state and vastly less security than one 128-bit cipher.
Check
Recall the argument from Section 8.3.
Check your understanding
Why does AES's final round omit MixColumns?
Answer: C
Why: Two halves to the answer, and both matter. Structurally, the omission is what lets the inverse cipher be rearranged into encryption's layer order, so one implementation serves both directions. And a final MixColumns would contribute nothing: it is a keyless invertible public map, so the attacker simply applies InvMixColumns to the ciphertext herself — exactly the argument that makes DES's initial permutation decorative.
Connect it up
AES is short enough to hold entirely in your head, which is worth doing once.
Draw it
Write the round as four named layers with one line each on what it does and why it is there. Draw the 4 × 4 state and mark which layer acts on entries, which on rows, and which on columns. Note the three key lengths with their round counts, and the two facts about GF(2⁸) that AES needs. Finish with the two reasons the last round omits MixColumns, and the one sentence from Section 8.4 about why the S-box is algebraic.
That last sentence is the chapter's real content: the S-box was made explicit and algebraic to avoid the suspicions that haunted DES. It is a decision about trust, made in the design of the mathematics.
Exit ticket
One question, about the difference that outlasted the cipher.
Predict first
What is the most consequential difference between how DES and AES were chosen?
Correct: AES was selected by open public competition with published design rationale, so its strength is supported by evidence anyone can check
Why: No practical cipher has a proof of security, so the only evidence for one is that competent people attacked it publicly and failed. DES's closed process produced a genuinely excellent cipher that nobody outside could verify — fifteen years of suspicion aimed at a decision that had helped, while a decision that had hurt was defended by the same authority. AES's competition was designed to generate the evidence, and Section 8.4's published rationale is what lets the margin be tracked decades later.
Recap
The cipher securing most of the world's traffic, and the design argument behind every choice in it.
Chapter 9 next. Everything so far has been symmetric, and every chapter has assumed Alice and Bob already share a key. RSA is where that assumption is finally dropped — and where Chapter 3's number theory stops being preparation and becomes the security itself.
Figure (svg): DES and AES compared on structure, S-box design, block size and design process.
Want this taught 1-on-1? Alexander tutors Cryptography — $55/session, free consultation.