Chapter 11 of Trappe & Washington: the three defining properties of a cryptographic hash and why collision resistance implies preimage resistance; the toy XOR hash broken three ways; the discrete log hash and its provable-but-unusable trade; the Merkle-Damgard construction together with the length extension attack that follows from outputting the internal state; SHA-2's compression function; and the sponge construction behind SHA-3, whose hidden capacity removes length extension structurally.
Subject: Cryptography · 60 slides · diagram-first lesson
Open the interactive version of this deck
Title
Cryptography · Chapter 11
A fingerprint of any length of data — not encryption, not reversible, and the tool behind most of the rest of the book
Objectives
Every primitive so far has been reversible by design. This one is not, and its usefulness comes from that. Hash functions carry integrity in Chapter 12, signatures in Chapter 13, protocols in Chapter 15, and they have already appeared in Chapters 5, 7 and 10.
Figure (svg): The three defining properties of a cryptographic hash function and what each one forbids.
Warm-up
It takes any input and returns a fixed-size digest. Nothing can be recovered from the digest.
Discussion prompt
If nothing is recoverable, what is the point? Name three uses you have already met in this course.
Hint: Chapters 5, 7 and 10 each used one without naming it as a chapter subject.
Answer:
Chapter 5: the practical pseudorandom generator was b_j = the last bit of f(s + j), with f a one-way function — and the book named SHA as one of the two standard choices.
Chapter 7: password storage. The system stores h(salt ‖ password) and never the password, so a stolen file does not reveal credentials.
Chapter 10: bit commitment. Alice hashes random padding plus her bit; the hash hides it and collision resistance binds her to it.
And the general answer: a hash lets you compare or commit to data without holding it, and lets you detect change without storing the original. Those are different jobs from encryption, and none of them needs reversibility.
Which is why the chapter is placed here: Chapter 12's message authentication, Chapter 13's signatures and Chapter 16's blockchains are all built on this one primitive.
Figure (svg): A hash function taking inputs of any length and producing a fixed-length digest.
Section
Section 11.1 · pp. 226-230
Concept
A cryptographic hash function h takes a message of arbitrary length and produces a message digest of fixed length — 256 bits, say. Three properties are required.
Read the second property carefully. It is a stronger demand than 'the original message cannot be recovered', because it forbids finding any preimage — and h is many-to-one, so preimages are plentiful.
Figure (svg): The three defining properties of a cryptographic hash function and what each one forbids.
Concept
The set of possible messages is vastly larger than the set of possible digests. A 256-bit hash has 2²⁵⁶ outputs and accepts inputs of every length, so by the pigeonhole principle there are infinitely many pairs colliding.
So the requirement cannot be that collisions do not exist. It is that nobody can find one.
The book states the practical form: if Bob produces a message m and its digest h(m), Alice wants to be reasonably certain Bob does not know another m′ with h(m′) = h(m) — even if both are allowed to be random strings of symbols.
That last clause matters. Eve does not have to produce two meaningful documents; a collision between two strings of noise is enough to break the property, and Chapter 12 shows how a noise collision is turned into a meaningful one.
Figure (svg): Many inputs mapping into a much smaller set of digests, so collisions are unavoidable.
Worked example
The book gives a short argument, and it is worth following because it explains why the two properties are listed separately at all.
Suppose H is not preimage resistant — so there is an efficient way to find preimages
Why: This is the assumption we want to contradict.
Take a random x and compute y = H(x)
Why: A perfectly ordinary computation.
Use the preimage-finding ability to obtain x′ with H(x′) = y = H(x)
Why: By assumption this is efficient.
Because H is many-to-one, x′ is very likely different from x
Why: There are far more preimages than the one we started with, so landing back on x would be a coincidence.
Verify: x and x′ collide, contradicting collision resistance
Why: So collision resistance is the stronger property. They are listed separately because they are used in different circumstances: password storage needs preimage resistance, and signatures need collision resistance.
Figure (svg): The three defining properties of a cryptographic hash function and what each one forbids.
Definition probe
Different applications need different guarantees, and getting this mapping right decides which hash is acceptable where.
Sort into buckets
Sort each use by the property it depends on most.
Concept
A hash can be built from any hard problem. Take a large prime p and two primitive roots α and β, split the message into two halves x₀ and x₁, and set:
\[ h(m) = \alpha^{x_0} \beta^{x_1} \pmod p \]
Finding a collision turns out to be equivalent to computing the discrete logarithm of β to the base α — so this function is collision resistant exactly as far as Chapter 10's problem is hard. That is an unusually clean security argument: a provable reduction, not an assumption about a hand-built table.
And it is unusable. It employs modular exponentiation, so its cost is about that of RSA or ElGamal. Fast enough for a key exchange; hopeless for the massive inputs hash functions are actually applied to — a video file, a disk image, every packet on a link.
This is the chapter's first trade, and it is the same one Chapter 5 made with Blum-Blum-Shub: the construction with the best security argument is far too slow, so practice uses bit-level operations with no proof and a great deal of scrutiny.
Figure (svg): The discrete log hash: provable collision resistance at the cost of a modular exponentiation per block.
Two truths and a lie
Two of these describe something else entirely.
Eliminate the wrong options
Which statement describes a cryptographic hash function?
Survives elimination: a
Why: The two wrong options are the two things people most often mistake a hash for, and both errors have practical consequences: treating a hash as encryption leads to expecting the data back, and treating it as compression leads to expecting the digest to grow with the input. The defining features are that it is unkeyed, public, fixed-output and irreversible by design.
Estimation
SHA-256 produces 256-bit digests.
Predict first
Roughly how many distinct digests exist?
Correct: About 10⁷⁷
This is the number that makes a digest usable as an identifier. Git names every object by its hash precisely because an accidental collision is beyond consideration — the design assumed no adversary, which is why SHA-1's break required Git to add a collision detector.
It also explains why digests are compared for equality rather than searched: a digest is an address in a space nobody can enumerate.
Why: 2²⁵⁶ ≈ 1.2 × 10⁷⁷, which is roughly the number of atoms in the observable universe. The set is finite and the set of inputs is not, so collisions exist in unlimited supply — but the space is large enough that stumbling on one by chance is impossible, and finding one deliberately costs 2¹²⁸ by the birthday bound.
Section
Section 11.2 · pp. 230-231
Concept
The book gives a hash with the right shape and none of the security, explicitly warning it should never be used in any system. It is the right thing to study, because breaking it shows what the properties are for.
Break the message into n-bit blocks m₁, …, m_l, padding the last with zeros. Stack them as rows of an array. The i-th bit of the digest is the XOR down the i-th column:
\[ h_i = m_{1i} \oplus m_{2i} \oplus \cdots \oplus m_{li} \]
This does take arbitrary input to an n-bit digest, and it is extremely fast. It is not cryptographically secure, and the reason is worth working out rather than being told.
Figure (svg): The toy XOR hash: stack the message blocks as rows and XOR down each column.
Worked example
Each break corresponds to one property failing.
Collisions: swap any two blocks
Why: XOR is commutative, so reordering the rows changes nothing about the column sums. Two completely different messages, same digest, found in zero work.
Collisions again: append two identical blocks
Why: m ‖ b ‖ b has the same digest as m, since b ⊕ b = 0. Unlimited collisions on demand.
Preimages: choose all but one block freely, then solve for the last
Why: Set the final block to the XOR of the target digest with everything already chosen. One pass, and the digest is whatever you wanted.
Verify: all three attacks cost less than computing the hash itself
Why: The function is fast, deterministic and compressing — and completely insecure. So those three properties are not what makes a hash cryptographic; the difficulty of inversion and collision is, and it has to be engineered deliberately.
Figure (svg): The toy XOR hash: stack the message blocks as rows and XOR down each column.
Socratic
It compresses, it is fast, and it accepts any input. Compare it with what a real hash does.
Discussion prompt
Name the two structural properties it lacks, and connect each to a cipher-design idea from Chapters 7 and 8.
Hint: One is about linearity; one is about position.
Answer:
It is linear. XOR is addition over GF(2), so h(m₁ ⊕ m₂) = h(m₁) ⊕ h(m₂) and the whole function is a matrix. This is the fourth appearance of the same failure — the LFSR, the Hill cipher, DES without S-boxes, AES without SubBytes. A cryptographic primitive needs a non-linear component.
It is position-blind. Every block enters the sum the same way, so reordering is invisible. Real hashes make each block's contribution depend on where it sits — through chaining, through rotations, through round constants. It is the same reason AES has round constants: to make step i different from step j.
And a third, following from the first two: no diffusion. Flipping one input bit flips exactly one output bit. A real hash aims for the avalanche property, where one input bit changes about half the output bits.
So the design targets are the same as a block cipher's — non-linearity, position dependence, diffusion — which is why Section 11.4's SHA-256 looks so much like a cipher's round function, with rotations, XORs and non-linear mixing of a state.
Counterexample
The three properties are separate requirements, and a function can satisfy some and not others.
Discussion prompt
Construct a function that is preimage resistant but not collision resistant, and say why such a function is dangerous in practice.
Hint: Take a good hash and make it ignore one bit of its input.
Answer:
The construction: define g(m) = SHA-256(m with its last bit set to 0). Given a digest y, finding any m with g(m) = y still requires inverting SHA-256, so g is preimage resistant.
But collisions are free. m and m-with-the-last-bit-flipped always collide, at zero cost. One property intact, the other destroyed completely.
Why this is dangerous in practice: a function like this passes every casual test. Digests look random, no input can be recovered, and an implementer checking 'is it one-way?' gets a yes. The failure appears only when an attacker is allowed to choose both messages — which is exactly the signature setting.
And it is not hypothetical. MD5 today is precisely this shape: preimages remain infeasible while collisions take seconds. Anyone reasoning 'MD5 cannot be reversed, so it is fine here' is making this error.
The rule: name the property the application needs before choosing the function, and check that property specifically. 'Is it secure?' has no answer.
Real world
Hash functions are the most-invoked cryptographic primitive after none at all.
Discussion prompt
Name four places a hash ran on your behalf in the last hour, and what property each use needed.
Hint: Version control, package management, the network, and logging in.
Answer:
Git. Every commit, tree and blob is named by its digest, and a commit's hash covers its parent's — so the whole history is a hash chain. Needs collision resistance, which is why Git moved to a collision-detecting SHA-1 and is migrating to SHA-256.
Package managers. npm, apt and pip verify downloads against published digests. Needs collision resistance, because an attacker who controls a mirror supplies both the file and, potentially, influences the published digest.
TLS. Every record carries a MAC — HMAC or a polynomial MAC — and the certificate chain is verified by checking signatures over hashes. Needs collision resistance for the certificates and unforgeability for the records.
Logging in. The password was hashed with a salt and a slow function; the session token is very likely a MAC. Needs preimage resistance and deliberate slowness.
And a fifth: the deduplication in your backup system, and the content-addressed storage behind most CDNs. Both silently rely on collisions being impossible to find, which is a security assumption dressed as an engineering convenience.
Section
Section 11.3 · pp. 231-233
Concept
Invented independently by Ralph Merkle in 1979 and Ivan Damgård in 1989, and used by almost every hash function until recently.
The ingredient is a compression function f taking two bitstrings H and M and producing H′ = f(H, M) of the same length as H. For SHA-256, M is 512 bits and H is 256.
Pad the message so its length is a multiple of 512 and split it into blocks M₀ ‖ M₁ ‖ … ‖ M_{n−1}. Set an initial value IV, then feed the blocks in one at a time. The final output is the hash.
\[ H_0 = IV, \qquad H_{i+1} = f(H_i, M_i), \qquad h(M) = H_n \]
The book calls the construction very natural: blocks are read one at a time and stirred into the mix with everything before them. It also streams — the whole message never needs to be in memory — which is why it dominated for thirty years.
Figure (svg): The Merkle-Damgard construction: message blocks fed one at a time into a compression function chained from an initial value.
Concept
Alice wants Bob to know her message M is untampered. They share a secret K, so she sends M together with H(K‖M). Bob recomputes and compares. Eve does not know K, so she cannot produce her own tag — apparently.
But she does not need K. Because the construction is iterative, the value H(K‖M) is the internal state after processing K‖M. Eve simply continues the computation:
\[ H(K \| M \| M'') = f\bigl(H(K \| M), \, M''\bigr) \]
She appends her own blocks M″, computes the new tag from the old one, and sends M‖M″ with a tag Bob will accept as authentic. She never learned K and never broke the compression function.
Using M‖K instead thwarts this one — but the book immediately notes it opens a different problem: if Eve can find collisions for H, she finds a good M₁ and a bad M₂ with H(M₁) = H(M₂), gets Alice to authenticate M₁, and reuses the tag on M₂.
Figure (svg): The length extension attack: because the hash is the internal state, Eve continues the computation from it.
Worked example
Trace the attack so its cost is clear — it is not an approximation or a search.
Eve intercepts M and t = H(K‖M). She does not know K, and she does not know |K|
Why: The length matters only for the padding, and she can guess it — there are few plausible values.
She sets her hash function's internal state to t
Why: Legitimate, because in Merkle-Damgård the output is the state. No implementation trick is needed; many libraries expose exactly this.
She constructs the padding that K‖M would have received, and appends her chosen suffix M″
Why: The forged message is M ‖ padding ‖ M″, which is a slightly odd-looking but valid message.
She runs the compression function over M″ from state t, obtaining t′
Why: One or two invocations of f — microseconds.
Verify: Bob computes H(K ‖ M ‖ padding ‖ M″) and gets exactly t′
Why: The forgery verifies. Total cost: two hash computations and a guess at |K|. This is not a theoretical weakness; it broke real API authentication schemes that used H(secret ‖ message) as a signature, and it is the reason HMAC exists.
Figure (svg): The length extension attack: because the hash is the internal state, Eve continues the computation from it.
Anomaly
Length extension is not a flaw in SHA-256's compression function. Every Merkle-Damgård hash has it, whatever f is.
Predict first
What is the underlying structural mistake?
Correct: The construction outputs its entire internal state, so an attacker can resume the computation
HMAC's answer: HMAC(K, M) = H((K ⊕ opad) ‖ H((K ⊕ ipad) ‖ M)). The outer hash means the value an attacker sees is not the state of any computation she can continue.
SHA-3's answer: keep part of the state — the capacity — permanently hidden and never output it. Section 11.5's sponge.
And the general lesson: a construction can be insecure with a perfect primitive inside it. This is Chapter 6's modes lesson, in a new setting.
Why: Merkle-Damgård returns H_n, which is exactly the state after the last block. Anyone holding a digest holds a fully resumable computation. No strengthening of f helps, because f is never attacked — it is used, correctly, by the attacker. The fix is structural: either do not output the whole state, which is what the sponge does, or wrap the hash so the attacker never sees a resumable value, which is what HMAC does.
Figure (svg): The length extension attack: because the hash is the internal state, Eve continues the computation from it.
Invariant
Four blocks, one compression function, and the state carried between them.
Step through it
At which step does an attacker who holds only the digest have everything she needs to continue?
The last one. H₄ is both the answer and a valid starting state, so appending another block is a single call to f. Nothing about f had to be weak.
Explain it
A developer has written an API that authenticates requests with SHA256(secret + body) and asks why you flagged it.
Discussion prompt
Explain the attack in terms they will act on, and give the fix in one line.
Hint: Avoid the word 'state' until you have said what the attacker does.
Answer:
Start with what the attacker does: she takes one request she has seen — a legitimate one, with its signature — and produces a longer request ending in whatever she chooses, together with a signature your server will accept. She never learns the secret.
Then say why: SHA-256 works through the message a block at a time, keeping a running value, and the signature you publish is that running value. So she loads it and keeps going from where you stopped.
Then the concrete consequence: if the body is action=read&user=alice, she can produce action=read&user=alice<padding>&action=delete with a valid signature. Whether that is exploitable depends on how the body is parsed, and many parsers take the last value of a repeated key.
The fix, in one line: use HMAC-SHA256 instead of SHA256, which is one function call in every language's standard library.
And the general point worth leaving them with: a hash is not a MAC, and you cannot make it one by putting the key in the input. HMAC exists because that specific instinct is wrong.
Sorting
Retrieval across the course so far. These are the four symmetric tools now available.
Sort into buckets
Sort each requirement by the primitive that meets it.
Five requirements, four tools, and the commonest production error is using row two's tool for row three's job.
Section
Section 11.4 · pp. 233-237
Concept
Unlike block ciphers, where many good options exist, only a few hash functions are used in practice. The book names the SHA family, the MD family, and RIPEMD-160.
SHA-256 processes 512-bit blocks with a 256-bit state, running 64 rounds of bit-level operations: rotations, shifts, XORs, modular additions and non-linear functions of three words. No modular exponentiation anywhere — that is the whole reason it is fast enough to use.
Figure (svg): The SHA family over time: SHA-0 and SHA-1 broken, SHA-2 in wide use, SHA-3 standardised alongside it.
Worked example
The book warns the detail is technical and gives it to convey the flavour. Here is the shape, which is what transfers.
Expand the 512-bit block into 64 words of 32 bits, by a recurrence mixing earlier words with rotations and shifts
Why: The message schedule. Its job is to make each round see a different function of the block — the same role as AES's key schedule.
Initialise eight 32-bit working variables from the current chaining value
Why: The 256-bit state, viewed as eight words.
Run 64 rounds. Each round mixes the working variables using rotations, XORs, additions mod 2³², and two non-linear choice functions
Why: Rotation and XOR give diffusion; the choice functions and the additions give non-linearity. Compare DES's S-boxes and AES's SubBytes — same requirement, cheaper components.
Each round also adds one word of the expanded message and one round constant
Why: Round constants again, for the same reason as AES: to make every round different and defeat symmetry attacks.
Verify: add the working variables back into the chaining value at the end
Why: This final feed-forward is what makes the compression function one-way. Without it the round function would be invertible and an attacker could run the whole thing backwards from the digest. It is the Davies-Meyer construction, and it is the single most important line in the design.
Figure (svg): The SHA-256 compression function: message expansion, sixty-four rounds on eight working words, then a feed-forward addition.
Estimation
SHA-256 produces a 256-bit digest, and no attack better than generic is known.
Predict first
How many hash computations does finding a collision take?
Correct: 2¹²⁸
This is why digests are twice as long as the security level wanted: SHA-256 for 128-bit security, SHA-512 for 256-bit. Compare AES, where the key is the security level directly.
It is also why SHA-1's 160 bits gave only 80 bits of collision resistance, which is what the 2017 attack finally reached — and why MD5's 128 bits, giving 64, fell far earlier.
Why: About 2¹²⁸, by the birthday bound of Section 12.1: with 2^(n/2) random digests a collision becomes likely, because the number of pairs grows as the square of the number of samples. So an n-bit hash gives n/2 bits of collision resistance and n bits of preimage resistance — the two numbers are always a factor of two apart in the exponent, and quoting the wrong one is the commonest error about hash security.
Figure (svg): The work needed to find a preimage against the work needed to find a collision, for a hash of n bits.
Prediction
The 2017 SHA-1 collision was produced by Google and CWI Amsterdam.
Predict first
Roughly what did it take?
Correct: About 6500 CPU-years and 100 GPU-years, run in parallel
Note the two numbers: the nominal birthday bound was 2⁸⁰ and the actual attack cost about 2⁶³. That gap of seventeen bits is what years of cryptanalysis bought, and it is the usual pattern — a hash weakens gradually before it falls.
Which is why standards bodies deprecate on the trend rather than on the break. SHA-1 was deprecated for certificates in 2011, six years before the collision, on exactly this reasoning.
Why: Around 2⁶³ hash computations — nine quintillion — which was well below SHA-1's nominal 2⁸⁰ collision resistance thanks to cryptanalytic shortcuts, and affordable for a large organisation with cloud capacity. The point of the demonstration was that it was expensive but ordinary, which is exactly the threshold at which a hash must be retired.
Figure (svg): The work needed to find a preimage against the work needed to find a collision, for a hash of n bits.
Cost model
Two numbers, and every question about whether a hash is adequate reduces to them.
Annotate
On: \( \text{preimage: } 2^{n} \qquad \text{collision: } 2^{n/2} \)
The last note is the one that matters: the formulas tell you when a hash is definitely too small, never that it is definitely large enough.
Section
Section 11.5 · pp. 237-242
Concept
In 2006 NIST announced a competition for a new hash function to serve alongside SHA-2 — not to replace it. The requirement was to be at least as secure as SHA-2 with the same four output sizes.
Fifty-one entries were submitted. Keccak won in 2012 and was standardised as FIPS-202, becoming SHA-3. It was designed by Bertoni, Daemen, Van Assche and Peeters — Daemen also being a co-designer of AES.
The competition's purpose is worth stating plainly. SHA-1 had just fallen and SHA-2 uses the same construction, so a structural break in Merkle-Damgård would have taken both at once. NIST wanted a standard resting on different foundations, held in reserve.
Keccak duly uses a completely different construction: sponge functions.
Figure (svg): The SHA family over time: SHA-0 and SHA-1 broken, SHA-2 in wide use, SHA-3 standardised alongside it.
Concept
The state is b bits, split as b = r + c, where r is the rate and c the capacity. A function f maps b bits to b bits — and unlike a compression function, f is one-to-one, which Merkle-Damgård could not permit because its input is larger than its output.
The capacity is never touched by the message and never output. That single fact is the whole security argument: an attacker who holds a digest holds at most r bits of a b-bit state, so she cannot resume the computation.
Length extension therefore does not apply, structurally rather than by patching. And the squeeze phase gives arbitrary-length output for free, which is why SHA-3 also provides the SHAKE extendable-output functions.
Figure (svg): The sponge construction: message blocks absorbed into the rate portion of the state, then the hash squeezed out.
Comparison
Two constructions for the same job. Fill the blanks.
Comparison matrix
| Merkle-Damgård | Sponge | |
|---|---|---|
| Inner function | compression: shrinks its input | a permutation — one-to-one, same size in and out |
| What the output is | the entire internal state | part of the state; the capacity stays hidden |
| Length extension | applies, structurally | does not apply |
| Variable-length output | no | yes — keep squeezing |
| Used by | MD5, SHA-1, SHA-2 | SHA-3 / Keccak |
Row two is the whole difference, and rows three and four are both consequences of it. A construction that hides part of its state gets a security property and a feature at the same time.
Socratic
SHA-2 is unbroken and widely deployed. SHA-3 is standardised alongside it rather than replacing it.
Discussion prompt
What is the reasoning, and where have you seen this argument before?
Hint: SHA-1 and SHA-2 share something that SHA-3 does not share with either.
Answer:
SHA-1 and SHA-2 use the same construction. When SHA-1's collision resistance fell, the immediate worry was whether the break was about SHA-1's compression function or about Merkle-Damgård itself. If the latter, SHA-2 would follow.
It turned out to be the compression function, and SHA-2 stands. But NIST could not know that in 2006, and by 2012 it had a standardised alternative resting on entirely different foundations.
This is Chapter 1's diversification argument, in a new place. There it was RSA, ElGamal, NTRU and McEliece resting on four independent hard problems, so that a break of one leaves the others. Here it is two hash constructions.
**And it is why the competition asked for something different rather than something better.** A faster Merkle-Damgård hash would have added nothing to the hedge, whatever its margin.
The general principle: when nothing is proved, the only real protection is independence. Having two standards that would fall to the same insight is having one standard.
Trade off
Fill the blanks. Whether speed is a virtue depends entirely on the job.
Comparison matrix
| Use | Is speed good? | Why |
|---|---|---|
| Verifying a 4 GB download | yes | the defender hashes once and wants it instant |
| Authenticating API requests | yes | the key is high-entropy, so guessing is hopeless regardless of speed |
| Storing passwords | no — speed is the attacker's asset | the defender hashes once per login; the attacker hashes billions of times |
| Deriving a key from a Diffie-Hellman value | yes | the input already has full entropy |
| Deriving a key from a passphrase | no | the input has low entropy, so the attacker can enumerate it |
The rule underneath every row: slow down the hash exactly when the input is guessable. High-entropy inputs need speed; low-entropy inputs need cost.
Analogy
Almost nothing in this chapter is new — the same four failures keep reappearing in new settings.
Match the pairs
Why: Four weaknesses, four earlier chapters, and the same four ideas. Linearity is always solvable; determinism always leaks equality; a correct primitive in the wrong construction is always insecure; and a small input space is always enumerable regardless of what function is applied to it. Recognising which of the four you are looking at is most of what reading a new design consists of.
Edge cases
MD5 collisions take seconds, and MD5 is still shipped in a great deal of software.
Discussion prompt
Name a use where MD5 remains defensible and one where it emphatically is not, and give the principle that separates them.
Hint: Ask whether anyone hostile chooses the input.
Answer:
Defensible: a checksum against accidental corruption. A disk error or a truncated download will not produce a colliding file, because nothing is choosing the corruption. The property needed is that random change is detected, and MD5 does that perfectly.
Not defensible: signing a certificate, a binary, or any document. The attacker chooses both messages, so collision resistance is the binding requirement and MD5 has none. The 2008 rogue CA attack did exactly this.
Also not defensible: deduplication in a system that accepts untrusted uploads. Two colliding files silently merge, so an attacker uploads a benign file that displaces a hostile one, or vice versa.
The separating principle: is there an adversary who chooses the input? Where the answer is no, a broken-for-collisions hash is fine; where it is yes, it is worthless.
And a caution about drift: a system built with no adversary in mind acquires one when it becomes popular. Git chose SHA-1 as a content identifier, not a security mechanism, and then had to add collision detection when it turned out to be both.
Commit first
You are specifying integrity protection for records that must remain verifiable for fifty years.
Predict first
What do you choose?
Correct: SHA3-512, or SHA-512 — and record which was used, with a plan to re-hash
The re-hashing plan matters as much as the algorithm. An archive whose digests cannot be upgraded without re-verifying every record against the old ones has no migration path at all.
A common design is to store digests under several functions from the start, so a break in one leaves the others standing — Chapter 1's diversification, at the record level.
Why: Over fifty years the specific function will very likely weaken, so two things matter more than the choice itself: a large margin now, and the ability to change later. A 512-bit digest gives 256-bit collision resistance, comfortable even against Grover; recording the algorithm identifier alongside each digest makes migration a data operation rather than a redesign — which is Chapter 7's DES lesson applied to hashes.
Notation
Every hash is quoted with a handful of numbers, and each one means something different.
Annotate
On: \( \text{SHA-256: } n = 256, \quad \text{block} = 512, \quad \text{rounds} = 64 \)
The last note is the one that causes real bugs: H(K‖M) looks like a keyed hash and is not one.
Discrimination
MD5 and SHA-1 are 'broken', and the word covers two very different situations.
Sort into buckets
Sort each statement.
Error analysis
From an API's authentication documentation.
Annotate
Every one of these has appeared in a real API. The first is the one that has appeared most often.
Real world
A collision is two random-looking strings with the same digest. That sounds harmless.
Discussion prompt
How did the 2008 rogue CA certificate and the 2017 SHA-1 collision turn a collision into a real attack?
Hint: The attacker controls parts of both documents, and documents have places where arbitrary bytes are ignored.
Answer:
Documents have slack. A PDF has comment fields and unused object streams; a certificate has extension fields; an executable has padding. An attacker constructs two documents that differ only in slack bytes and collide.
Then she gets the innocent one signed. In 2008 a team obtained a legitimate MD5-signed certificate from a commercial CA while holding a colliding certificate for an intermediate CA. The signature transferred, giving them the power to issue certificates for any domain.
The 2017 SHA-1 collision was two PDFs that displayed completely different content and shared a digest — deliberately chosen to be visibly different, so the consequence was undeniable.
And Chapter 12 makes it worse: the birthday attack lets an attacker prepare 2^(n/2) variants of a good document and 2^(n/2) of a bad one, and match them up. Slack bytes make generating variants free.
Which is why collision resistance, not preimage resistance, is the property signatures need — and why a hash whose collisions are findable cannot be used for signing, even though nothing about it can be reversed.
Ranking
Rank by the actual cost of finding a collision today, not by the nominal digest length.
Put in order
Why: MD5 collisions take seconds on a laptop, far below its nominal 2⁶⁴. SHA-1's first collision cost a large but affordable computation in 2017, below its nominal 2⁸⁰. SHA-256 has no known attack better than the generic 2¹²⁸, and SHA3-512 gives 2²⁵⁶. Note that the first two are ordered by practical cost rather than by digest length, and that the gap between nominal and practical is exactly what 'a hash is broken' means.
Figure (svg): The work needed to find a preimage against the work needed to find a collision, for a hash of n bits.
Explain it to yourself
Every other primitive in this course has a key. A hash function does not.
Discussion prompt
Explain why, and say what has to change when you do want a keyed digest.
Hint: Ask what the key would be protecting.
Answer:
A hash makes no secrecy claim, so there is nothing for a key to protect. Its properties — preimage and collision resistance — are about the difficulty of computation, not about who knows what. Anyone can compute h(m), and that is intended.
This is what makes it publicly verifiable. Anyone can check that a file matches a published digest, or that a signature covers a given document, without holding any secret. A keyed function could not serve that role.
When you do want a keyed digest — to prove a message came from someone who knows K — you need a message authentication code, and you cannot get one by feeding the key in as data. H(K‖M) falls to length extension; H(M‖K) falls to collisions.
HMAC is the construction that works: H((K ⊕ opad) ‖ H((K ⊕ ipad) ‖ M)). Two nested hashes and two constants, and it has a proof of security given reasonable assumptions on H.
The general lesson, and it recurs: you cannot obtain a new security property by using a primitive in a new way. Composition needs a construction with its own argument, and Chapter 12's MACs are exactly that.
Commit first
A system uses SHA-256 correctly for password storage, file integrity and signature hashing.
Predict first
What is most likely to go wrong?
Correct: A construction error — H(K‖M) used as a MAC, or unsalted password hashing
This is the same conclusion as Chapters 6 and 10 reached about modes and about authentication: the primitive is the part that has been analysed for decades, and the assembly is the part written last week.
The practical checklist for a hash: is it a MAC? then HMAC. Is it a password? then Argon2 with a salt. Is it a signature? then a collision-resistant hash, so not MD5 or SHA-1.
Why: SHA-256 itself has no known weakness and is not close to one. What breaks is the surrounding construction: using a raw hash as a MAC exposes length extension; hashing passwords without a salt permits precomputed tables; hashing passwords with a fast function permits billions of guesses per second. Every one of these is a correct primitive used in an incorrect construction.
Faded example
Three lines, and the length extension attack is visible in them.
Fill in the blanks
Set H₀ = IV. For each block, H_H_n = f(H_i, M_i). The hash is internal state, which is also the continue after the last block — and that is precisely why an attacker can ___ the computation.
Why: Writing the recurrence out makes the attack self-evident: the value returned to the world is the same object the algorithm would use to keep going. The sponge's fix is visible in the same notation — it returns only part of the state, so what an attacker holds is not enough to continue.
Edge cases
Shorter digests are cheaper to store and transmit. There is a floor.
Discussion prompt
What is the shortest useful digest, and what decides it?
Hint: The birthday bound, and what the digest is being used for.
Answer:
For collision resistance, the digest must be twice the security level. 128-bit security needs 256 bits, because 2^(n/2) is the generic collision cost. There is no design that beats this — it is a property of the output size, not of the function.
For preimage resistance alone, n bits gives n bits. So a use needing only preimage resistance — a password hash, a key derivation — can safely use a shorter digest, and this is why 128-bit outputs still appear in such roles.
Below about 64 bits, nothing is safe: 2³² operations is seconds, so even accidental collisions become likely in a large enough dataset. Git's original 160-bit SHA-1 hashes were chosen against accidental collision and turned out to need adversarial resistance too.
And truncation is legitimate if done to a standardised length: SHA-512/256 is SHA-512 truncated to 256 bits with a different IV, and it is faster than SHA-256 on 64-bit hardware while giving the same 128-bit collision resistance.
The rule: decide which property the application needs, double it if the answer is collision resistance, and never go below the doubled number to save bytes.
Figure (svg): The work needed to find a preimage against the work needed to find a collision, for a hash of n bits.
Matching
Four constructions built on hash functions, each answering a specific failure.
Match the pairs
Why: Four constructions, four specific failures — and none of them is a weakness in the hash function itself. This is the chapter's practical summary: SHA-256 is sound, and everything that goes wrong goes wrong in how it is wrapped. Note also that the second row is Chapter 6's ECB lesson again, and the third is Chapter 6's modes lesson again.
Constraint
One system, three uses: verifying downloaded files, authenticating API requests, and storing user passwords.
Discussion prompt
Choose a construction for each and say what would go wrong with the obvious alternative.
Hint: Each job needs a different property, and only one of them wants speed.
Answer:
File verification: SHA-256, plain. Needs collision resistance, because an attacker may supply both the file and the published digest through a compromised mirror. MD5 would be wrong — collisions are trivial, and a malicious file matching a published MD5 is producible.
API authentication: HMAC-SHA256. Needs a keyed construction. Plain SHA256(secret ‖ body) falls to length extension, and SHA256(body ‖ secret) falls to collisions. HMAC is the construction with a proof.
Password storage: Argon2id with a per-user salt. Needs preimage resistance and deliberate slowness. SHA-256 would be a serious error — not because it is weak, but because it is fast, and speed is the attacker's asset here. Chapter 7's lesson.
The pattern to notice: the same primitive family serves all three, and the right answer differs every time. Speed is a virtue in the first, irrelevant in the second, and a vulnerability in the third.
And what is common: in all three, the failure mode is choosing the construction by habit rather than by asking which property the job needs.
Missing information
The phrase appears in privacy policies and security questionnaires constantly.
Discussion prompt
List what remains undetermined, and note which omission is most likely to matter.
Hint: Ask what the hashing is for, and what an attacker gets to choose.
Answer:
What is being hashed, and is the input space small? Hashing an email address or a phone number is not anonymisation: there are only so many phone numbers, and an attacker enumerates them all in minutes. This is the most common real misuse of the phrase.
Is it salted, and per-record? Without a salt, identical inputs give identical digests, so the digest set reveals the multiset of values — Chapter 6's ECB failure once more.
Is it slow, if it is a password? SHA-256 is fast by design, which is exactly wrong for password storage.
Is it keyed, if it is authenticating? A raw hash is not a MAC, and using it as one exposes length extension.
Which property is being relied on? Collision resistance and preimage resistance are different guarantees, and a function can have one without the other.
The likeliest problem is the first, because it is the one people do not realise is a problem: hashing does not anonymise a value drawn from a small set, and no choice of function fixes that.
Pattern
The chapter has shown a toy, a provable-but-slow construction, and two real families. The same four questions distinguish them.
And the practical summary of the whole chapter: the functions are sound and the constructions around them are where things break. Length extension, unsalted passwords, fast password hashing and raw-hash MACs are four failures with a correct primitive in every one.
Figure (svg): The sponge construction: message blocks absorbed into the rate portion of the state, then the hash squeezed out.
Trap
The trap. Hashing is one-way and irreversible. So replacing each user's email address with SHA-256 of that address anonymises the dataset — the digests cannot be reversed, so the identities are gone.
This reasoning appears in privacy policies, in advertising identifier schemes, and in data-sharing agreements, and it is stated in exactly these terms.
Why it fails. Preimage resistance says you cannot find an input for a digest by searching the whole input space. But an email address is not drawn from the whole input space — it is drawn from a list of a few billion real addresses, and hashing all of them takes minutes on one machine.
The attacker never inverts anything. She computes the forward direction over a dictionary and matches, exactly as in Chapter 7's password attack. The hash was not broken; it was applied to a small domain.
And determinism makes it worse. The same address always gives the same digest, so hashed datasets from different sources join on the digest — which is usually the very linkage the anonymisation was meant to prevent. Chapter 6's ECB lesson, in a privacy setting.
What actually helps: a secret salt or key held by one party only, which turns the digest into a MAC and prevents dictionary attacks — at the cost of no longer being publicly verifiable. Or genuine anonymisation techniques that add noise, which is a different field.
The rule to carry: a one-way function protects an input only if the input is unpredictable. Hashing a value from a small or enumerable set protects nothing at all.
Check
Work it out before you click.
Check your understanding
A certificate authority signs the hash of each certificate it issues. Which hash property must hold?
Answer: B
Why: An attacker prepares two certificates with the same digest — one innocuous, one granting authority — gets the CA to sign the innocuous one, and transfers the signature to the other. She chose both messages, so collision resistance is the binding requirement. This is exactly the 2008 rogue CA attack against MD5-signed certificates.
Check
Read the construction carefully.
Check your understanding
A server authenticates messages with SHA256(secret ‖ message). Eve sees one valid message and its tag. What can she do?
Answer: C
Why: The tag is the internal state of SHA-256 after processing secret ‖ message, so Eve loads it as her starting state, appends her own blocks and continues. She needs a guess at the secret's length for the padding, and nothing else. This is why HMAC exists, and it has broken real API authentication schemes.
Check
Apply the relationship between digest length and security.
Check your understanding
A hash function produces 160-bit digests. What is its collision resistance?
Answer: B
Why: About 2^(n/2) = 2⁸⁰ by the birthday bound, because the number of pairs among k samples grows as k², so a collision becomes likely once k reaches the square root of the output space. This is SHA-1's situation, and 2⁸⁰ is what the 2017 collision finally reached — nominally, and with a great deal of optimisation.
Connect it up
Three properties, two constructions, and a list of things that break around them.
Draw it
Write the three properties with the precise statement of each, and the one-line argument that collision resistance implies preimage resistance. Draw the Merkle-Damgård chain and mark where length extension enters. Draw the sponge and mark the capacity, noting what it prevents. Write the two security levels for an n-bit digest and why they differ by a factor of two in the exponent. Finish with four constructions — HMAC, salting, slow hashing, and the sponge's capacity — and the specific failure each one fixes.
That last list is the chapter in practice: the primitive is fine, and the four failures are all in how it was wrapped.
Exit ticket
One question, about what a hash actually promises.
Predict first
Why does hashing a value not make it anonymous?
Correct: Because a value from a small or predictable set can be found by hashing every candidate and matching — preimage resistance protects only unpredictable inputs
Why: The forward direction is cheap by design, so an attacker who can enumerate the plausible inputs simply hashes them all. Email addresses, phone numbers, national identifiers and dates of birth are all enumerable, so hashing them protects nothing. Nothing is reversed and no property of the function is violated — the function is being asked to do something it never promised, which is to protect a predictable input.
Recap
The first primitive in this course with no key and no inverse.
Chapter 12 next. The birthday attack that produces that factor of two, multicollisions, the random oracle model, message authentication codes, password protocols — and blockchains, which are hash chains with an economic argument attached.
Figure (svg): The work needed to find a preimage against the work needed to find a collision, for a hash of n bits.
Want this taught 1-on-1? Alexander tutors Cryptography — $55/session, free consultation.