CS 161, Lesson 22, in 51 slides. It explains what a cryptographic hash is and why it is avalanche-prone, deterministic, and unkeyed, in section 7.1, then gives the three security properties in section 7.2: preimage resistance, second-preimage resistance, and collision resistance. It covers using hashes for integrity and the trusted-channel limit, also in section 7.2, then real algorithms, including SHA-2's length-extension flaw and the birthday bound, in section 7.3, and closes with the lowest-hash verification scheme in section 7.4. It is anchored to textbook sections 7.1 to 7.4.
Subject: Computer Security · 88 slides · applied lesson
Open the interactive version of this deck · Homework for this lesson
Title
CS 161 · Lesson 22 of 45
what a hash is · preimage, second-preimage & collision resistance · integrity · SHA-2 vs SHA-3 & length extension · the lowest-hash trick
Objectives
Warm-up
Discussion prompt
Before we open L22 · Cryptographic Hash Functions: without looking back, what was the main idea of L21 · Padding (PKCS#7), Parallelization Trade-offs & IV Reuse, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
CS 161, Lesson 21, in 56 slides. It explains why CBC needs PKCS#7 padding while CTR does not, works the parallelization trade-offs between CBC and CTR, and shows why reusing an IV is catastrophic for CTR but merely contained for CBC. This is where IND-CPA is won or lost in practice. It is anchored to textbook sections 6.7 to 6.9.
Concept
We have ciphers that keep data secret. Hashes solve a different problem — proving data is unchanged — and they do it with no key at all.
Matching
Match the pairs
From Three questions this lesson answers — match each one to what it actually does. The descriptions have been shuffled.
Why: What is a hash?, What makes it secure?, What do we build with it? are easy to tell apart while they are sitting next to their descriptions and much harder afterwards, which is what this checks.
Section
Part 1 · §7.1 the fingerprint
Concept
Alice and Bob each have a copy of a 1 GB file and want to know if the copies are identical, without sending the whole gigabyte across the network.
Instead of comparing every word, each computes a tiny fixed-length fingerprint of their file and compares just that. A cryptographic hash function is what produces that fingerprint.
Counterexample
Discussion prompt
Alice and Bob each have a copy of a 1 GB file and want to know if the copies are identical, without sending the whole gigabyte across the network.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Instead of comparing every word, each computes a tiny fixed-length fingerprint of their file and compares just that. A cryptographic hash function is what produces that fingerprint.
Concept
Cryptographic hash function H — A function that maps a message M of ANY length to a fixed-length output H(M) — a 'fingerprint' or 'digest' of the input.
\[ H : \{0,1\}^* \longrightarrow \{0,1\}^n \]
Whatever you feed in — one byte or a terabyte — the output is the same fixed size. For SHA-256 that output is always 256 bits.
Analogy
Discussion prompt
Explain §7.1 A hash maps any input to fixed-length output by analogy to something with no Computer Security in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Whatever you feed in — one byte or a terabyte — the output is the same fixed size. For SHA-256 that output is always 256 bits.
Intuition
A checksum that grew with the file would be no easier to compare than the file itself. The power of a hash is that a terabyte and a tweet both collapse to the SAME small size.
That fixed width is what lets a hash stand in for data everywhere — a fingerprint you can store, send, sign, or compare in constant space, no matter how big the thing it names.
Ask yourself: why not just use the first 256 bits of the file as its 'fingerprint'? (Two files sharing a prefix would collide constantly, and tampering past byte 32 would go undetected — a real hash mixes the WHOLE input in.)
Explain it
Discussion prompt
Explain §7.1 Why fixed-length output is the whole point to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
A checksum that grew with the file would be no easier to compare than the file itself. The power of a hash is that a terabyte and a tweet both collapse to the SAME small size.
Concept
A hash is deterministic: the same input M always gives the same output H(M). That is exactly why Alice and Bob can compare fingerprints and trust a match.
A hash is also unkeyed: there is no secret. Anyone — Alice, Bob, or Eve — can compute H(M) for any M they hold. Keep that fact in mind; it limits what a bare hash can prove.
Intuition
Change one bit of the input — flip a single comma to a period — and the output looks completely different: about half the output bits flip, in no pattern you could predict.
This is the avalanche effect. There's no 'similar input, similar output' the way there is for an average or a checksum; a hash scatters its input across the whole digest.
Ask yourself: if two files differ in only one character, how related are their hashes? (Not at all — the digests look like two independent random strings.)
Ranking
Put in order
Put the moves of §7.1 Using a hash to compare two files into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Each hash is a fixed 256-bit fingerprint — tiny compared to a gigabyte — and deterministic, so identical files give identical fingerprints.
Worked example
Alice computes h_A = H(her 1 GB file); Bob computes h_B = H(his 1 GB file)
Why: Each hash is a fixed 256-bit fingerprint — tiny compared to a gigabyte — and deterministic, so identical files give identical fingerprints.
They exchange only the 256-bit hashes and compare
Why: 256 bits is 32 bytes; sending that is trivial next to sending or diffing a full gigabyte.
| case | h_A vs h_B | conclusion |
|---|---|---|
| files identical | h_A = h_B | same file (deterministic) |
| files differ | h_A ≠ h_B | different file (avalanche guarantees a mismatch shows) |
Verify: same hash ⇒ same file; different hash ⇒ different file
Why: §7.1: a hash mismatch always means the files differ. A match means they almost certainly match — finding two different files with the same hash is what collision resistance forbids (Part 2).
Comparison
Comparison matrix
From §7.1 Using a hash to compare two files: refill the h_A vs h_B column from what you know. The rest of the table is as it appeared.
| case | h_A vs h_B | conclusion |
|---|---|---|
| files identical | h_A = h_B | same file (deterministic) |
| files differ | h_A ≠ h_B | different file (avalanche guarantees a mismatch shows) |
Intuition
A hash is like a human fingerprint: tiny, identifies the thing it came from, but you cannot rebuild the person from the print. The gigabyte is not stored inside 256 bits.
By the pigeonhole principle, infinitely many inputs map to each output — so a hash cannot be a reversible compression. It's a one-way summary, not a zip file.
Ask yourself: can you reconstruct a 1 GB file from its 256-bit hash? (No — far more inputs than outputs; the information simply isn't there.)
Anomaly
Predict first
A student writes this, and it looks reasonable:
A student: 'A hash scrambles the data with a key, so with the key I can decrypt H(M) back into M.'
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Wrong on two counts. A hash is UNKEYED — there is no key to hold — and it is ONE-WAY: many inputs share each output, so there is nothing to invert back to.
A student: what kind of function is a hash, really?
Why: Wrong on two counts. A hash is UNKEYED — there is no key to hold — and it is ONE-WAY: many inputs share each output, so there is nothing to invert back to.
Trap
A student: 'A hash scrambles the data with a key, so with the key I can decrypt H(M) back into M.'
Treat H as a reversible cipher with a secret key
Why: Wrong on two counts. A hash is UNKEYED — there is no key to hold — and it is ONE-WAY: many inputs share each output, so there is nothing to invert back to.
A student: what kind of function is a hash, really?
Treat H as an unkeyed, one-way, deterministic fingerprint
Why: §7.1: a hash has no secret key (anyone can compute it) and cannot be reversed (fixed-length output, infinitely many inputs per output). Encryption hides and is reversible with a key; hashing summarizes and is not.
Section
Part 2 · §7.2 preimage, second-preimage, collision
Concept
Preimage resistance (one-way) — Given an output y = H(x), it is computationally infeasible to find ANY input x with H(x) = y.
You are handed only the fingerprint and asked to find something that hashes to it. A good hash makes that search hopeless — there is no shortcut to inverting H.
Intuition
If H were easy to invert, storing H(password) instead of the password would protect nothing — an attacker who steals the hash would just invert it. One-wayness is what lets a hash stand in for a secret.
Inverting means searching the input space, and the input space is astronomically large. With no pattern to exploit (avalanche), the only route is brute force.
Ask yourself: given just a 256-bit digest, how do you find an input that produces it? (Guess and hash — there is no smarter move against a good hash.)
Concept
Second-preimage resistance — Given a SPECIFIC input x, it is computationally infeasible to find a DIFFERENT input x′ ≠ x with H(x′) = H(x).
Here the first input is fixed and handed to you — say a legitimate document — and the attack is to find a second, different input that collides with it.
Concept
Collision resistance — It is computationally infeasible to find ANY pair x ≠ x′ with H(x) = H(x′). The attacker is free to choose BOTH inputs.
This is the strongest of the three: nothing is fixed in advance. The attacker wins by producing any two distinct messages that share a digest.
Definition probe
Sort into buckets
Every line below is part of the definition of Cryptographic hash function H or of Collision resistance — one or the other, never both. Put each where it belongs.
Concept
The properties differ in what the attacker is given versus what they must find. That single distinction is the whole table.
| property | given | hard to find |
|---|---|---|
| preimage | an output y | any x with H(x) = y |
| second-preimage | a specific input x | a different x′ ≠ x with H(x′) = H(x) |
| collision | nothing (free choice) | any x ≠ x′ with H(x) = H(x′) |
Trade off
Comparison matrix
From §7.2 The three properties side by side: every row here is a choice with a cost. Fill the given column, then say which row you would actually pick and what you give up for it.
| property | given | hard to find |
|---|---|---|
| preimage | an output y | any x with H(x) = y |
| second-preimage | a specific input x | a different x′ ≠ x with H(x′) = H(x) |
| collision | nothing (free choice) | any x ≠ x′ with H(x) = H(x′) |
Intuition
There are infinitely many possible inputs but only 2^256 possible outputs. By the pigeonhole principle, many different inputs MUST share each output — collisions are guaranteed to exist.
Collision resistance does not claim collisions don't exist. It claims they are infeasible to FIND — they're out there, but no efficient search lands on one.
Ask yourself: does 'collision resistant' mean no two inputs ever collide? (No — they certainly do; the point is you can't discover such a pair.)
Step zero
Discussion prompt
§7.2 Collision resistance implies second-preimage resistance — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Suppose H is NOT second-preimage resistant
Answer:
Worked example
Suppose H is NOT second-preimage resistant
Why: Assume an attacker, given some specific x, can efficiently find x′ ≠ x with H(x′) = H(x).
Use that ability to produce a collision
Why: Pick any x at all, run the second-preimage attack to get x′ ≠ x with H(x) = H(x′) — that pair (x, x′) is exactly a collision.
Conclude H is then NOT collision resistant either
Why: Breaking second-preimage hands you a collision, so a collision-resistant hash cannot have a second-preimage weakness.
Verify the implication: collision resistance ⇒ second-preimage resistance
Why: §7.2: the contrapositive we just proved. They are kept as SEPARATE properties because the difficulty levels differ — collisions are easier to find (birthday bound, Part 4), so a hash can be second-preimage resistant at a higher security level than it is collision resistant.
Blank canvas
Draw it
Draw what §7.2 Collision resistance implies second-preimage resistance just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Step zero
Discussion prompt
§7.2 Preimage vs collision: which is the attacker doing? — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Case A: an attacker has only a leaked digest y and wants any matching…
Answer:
Worked example
Case A: an attacker has only a leaked digest y and wants any matching input
Why: Nothing else is given — just the output. Finding any x with H(x) = y is a PREIMAGE attack.
Case B: an attacker is given a real document x and wants a fraudulent x′ ≠ x with the same hash
Why: The legitimate input is fixed; the attacker must match THAT target. This is a SECOND-PREIMAGE attack.
Case C: an attacker prepares two contracts — one benign, one malicious — engineered to share a hash, before anyone signs
Why: The attacker freely designs BOTH inputs to collide. This is a COLLISION attack, the easiest of the three.
| case | what's given | attack | work |
|---|---|---|---|
| A | an output y | preimage | ~2^n |
| B | a fixed input x | second-preimage | ~2^n |
| C | free choice of both | collision | ~2^(n/2) |
Verify: the property is named by what the attacker was HANDED, not just the goal
Why: §7.2: all three end in 'two inputs with the same hash', but how much the attacker controls in advance — and so the cost — differs. That control is the whole distinction.
Comparison
Comparison matrix
From §7.2 Preimage vs collision: which is the attacker doing?: refill the work column from what you know. The rest of the table is as it appeared.
| case | what's given | attack | work |
|---|---|---|---|
| A | an output y | preimage | ~2^n |
| B | a fixed input x | second-preimage | ~2^n |
| C | free choice of both | collision | ~2^(n/2) |
Anomaly
Predict first
A student writes this, and it looks reasonable:
A student: 'Second-preimage and collision are the same thing — both are two inputs with the same hash.'
It is wrong. Say what breaks — and say it before you turn the page.
Correct: They differ exactly there. Second-preimage GIVES you a specific x and asks for a partner.
A student: what is the precise difference?
Why: They differ exactly there. Second-preimage GIVES you a specific x and asks for a partner. Collision fixes nothing — the attacker picks BOTH inputs, which is far easier (birthday bound).
Trap
A student: 'Second-preimage and collision are the same thing — both are two inputs with the same hash.'
Ignore which input, if any, is fixed in advance
Why: They differ exactly there. Second-preimage GIVES you a specific x and asks for a partner. Collision fixes nothing — the attacker picks BOTH inputs, which is far easier (birthday bound).
A student: what is the precise difference?
Second-preimage: one input is fixed. Collision: attacker chooses both
Why: §7.2: in second-preimage you must match a TARGET you were handed; in a collision you only need ANY two distinct inputs that agree. More freedom ⇒ easier ⇒ a separate, weaker property.
Section
Part 3 · §7.2 and its limit
Concept
You download a 3 GB Ubuntu install image from a mirror you don't fully trust. How do you know the mirror — or a network attacker — didn't slip malware into it?
The Ubuntu developers publish the SHA-256 of the genuine image. You hash your download and compare. Match ⇒ your copy is the real image; mismatch ⇒ it was altered.
Hypothesis
Predict first
§7.2 Why collision resistance secures the ISO is about to be worked. State your hypothesis first: which rule or definition decides this one, and what is the first move it forces? Then watch whether the example agrees with you.
Correct: Developers compute h = H(genuine ISO) and publish h widely
Why: h is the 256-bit fingerprint of the exact bytes they intend you to run.
A hypothesis you wrote down is falsifiable; a vague sense of how it will go is not. If the example opens somewhere else, that gap is the thing worth chasing.
Worked example
Developers compute h = H(genuine ISO) and publish h widely
Why: h is the 256-bit fingerprint of the exact bytes they intend you to run.
You download some image, compute H(your download), and compare to h
Why: If your bytes differ from the genuine bytes by even one bit, avalanche makes your hash differ from h, so tampering shows up.
An attacker who wants a malicious ISO to pass must find an image with hash h
Why: To keep the published h while changing the bytes, the attacker needs a DIFFERENT input that hashes to the same value — a second-preimage / collision against H.
Verify: collision resistance blocks the attacker
Why: §7.2: because finding any colliding image is infeasible, the attacker cannot produce a tampered ISO that still matches h. The hash gives integrity.
Concept
Everything above assumes you hold the genuine h. But a hash is unkeyed — anyone can compute it. If the adversary can tamper with the published hash too, they simply hash their malicious ISO and post THAT value.
An unkeyed hash therefore gives integrity only over a trusted channel — a channel where you're sure the hash value itself wasn't replaced. Securing the hash against a tampering adversary is what MACs (L23) and signatures (L30) add.
Sorting
Sort into buckets
These are the pieces of L22 · Cryptographic Hash Functions, out of order. Put each one back under the part of the lesson it belongs to.
Intuition
Ubuntu publishes its hashes over many independent channels and signs them, so an attacker would have to corrupt all of them at once. The defense isn't the hash alone — it's the trust in the path the hash travelled.
Compare: if you download both the ISO and its hash from the SAME compromised mirror, you've checked nothing. The attacker controlled both halves.
Ask yourself: a bare hash proves a file came from WHOM? (From no one in particular — it only proves the bytes match a value you already trusted, not who produced them.)
Anomaly
Predict first
A student writes this, and it looks reasonable:
A student: 'If the file's hash matches the one posted next to it, I know the file is authentic and came from the developer.'
It is wrong. Say what breaks — and say it before you turn the page.
Correct: An unkeyed hash has no secret, so anyone can recompute it.
A student: when does matching the hash actually mean something?
Why: An unkeyed hash has no secret, so anyone can recompute it. An attacker who edits the file just recomputes and republishes the matching hash — a self-consistent forgery. The match proves nothing about the source.
Trap
A student: 'If the file's hash matches the one posted next to it, I know the file is authentic and came from the developer.'
Trust a hash that travelled with the file
Why: An unkeyed hash has no secret, so anyone can recompute it. An attacker who edits the file just recomputes and republishes the matching hash — a self-consistent forgery. The match proves nothing about the source.
A student: when does matching the hash actually mean something?
Only when the hash value itself comes from a TRUSTED channel
Why: §7.2: a hash gives integrity only against an attacker who can change the file but NOT the published hash. To authenticate against an attacker who can change both, you need a key — a MAC (L23) or a signature (L30).
Section
Part 4 · §7.3 SHA-2, SHA-3, the birthday bound
Concept
Not every hash has held up. MD5 was broken long ago (Wang & Yu, 2005). SHA-1 was broken in 2017, when researchers produced an actual SHA-1 collision (the SHAttered attack).
The secure families used today are SHA-2 (SHA-256/384/512) and SHA-3 (SHA3-256/384/512). For this class, the key difference between them is length extension.
Ranking
Put in order
Put the moves of §7.3 SHA-1 broke git's assumption — eventually into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. git names objects by their SHA-1 hash, assuming two different objects could never share a hash — i.e.
Worked example
Note that git uses SHA-1 to identify every file and commit by content
Why: git names objects by their SHA-1 hash, assuming two different objects could never share a hash — i.e. assuming collision resistance.
Observe this worked fine for years — until it didn't
Why: No one could FIND a SHA-1 collision, so the assumption held in practice even though collisions must exist by pigeonhole.
In 2017 SHAttered produced two distinct PDFs with the same SHA-1 hash
Why: Stevens et al. computed a real collision, turning 'infeasible to find' into 'found' and invalidating SHA-1 for security use.
Verify the lesson: 'not broken yet' is not 'secure forever'
Why: §7.3: a hash is only as good as the best known attack. MD5 and SHA-1 are now broken; migrate to SHA-2 or SHA-3.
Concept
Length-extension attack — Given H(M) and the LENGTH of M (but NOT M itself), an attacker can compute H(M ‖ M′) for a suffix M′ of their choice — without ever knowing M.
This works against SHA-2 because its output IS its full internal state: the digest hands the attacker everything needed to keep hashing more data onto the end.
Matching
Match the pairs
Match each term to the definition this lesson gave it — not the one you would guess from the word.
Why: These are the working definitions of Cryptographic hash function H, Preimage resistance (one-way), Second-preimage resistance, Collision resistance, Length-extension attack as L22 · Cryptographic Hash Functions uses them. Pairing them correctly is the test of whether you could state each one with the slide switched off.
Intuition
Think of SHA-2 as showing its entire working memory at the finish line — so anyone can pick up the pen and keep writing. The published digest is the machine's full state.
SHA-3 keeps part of its internal state hidden: the output is only a window onto a larger secret state. With the rest hidden, an attacker can't resume the computation, so length extension fails.
Ask yourself: why can the attacker extend a SHA-2 digest but not a SHA-3 one? (SHA-2 reveals its whole state in the output; SHA-3 holds part back.)
Concept
A tempting way to build a keyed integrity tag is the naive MAC Hash(K ‖ M): prepend a secret key, then hash. It looks like only someone with K could produce the tag.
But with SHA-2, length extension lets an attacker who sees Hash(K ‖ M) compute Hash(K ‖ M ‖ M′) and forge a tag for an extended message — without knowing K. This exact flaw is why we need HMAC (L23).
Concept
Inverting an n-bit hash (preimage) takes about 2^n work. But FINDING A COLLISION takes only about 2^(n/2) work — the birthday bound, from the birthday paradox.
\[ \text{collision work} \approx 2^{n/2}, \qquad \text{preimage work} \approx 2^{n} \]
So a 256-bit hash gives only about 2^128 collision resistance — not 2^256. That square-root loss is why we pick the output length we do.
Step zero
Discussion prompt
§7.3 Matching hash length to key length — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Pick a security target — say AES-128, which costs ~2^128 to…
Answer:
Worked example
Pick a security target — say AES-128, which costs ~2^128 to brute-force
Why: We want the hash's weakest property (collisions) to be no easier than breaking the cipher it's paired with.
Apply the birthday bound: an n-bit hash collides in ~2^(n/2)
Why: To make collisions cost 2^128, we need n/2 = 128, i.e. n = 256 — so SHA-256 pairs with AES-128.
Scale up for AES-256: the NSA pairs it with SHA-384
Why: SHA-384 collides at ~2^192, which is already far beyond feasible; you don't need the full 2^256 of SHA-512 to be safe.
| cipher | key brute-force | hash | collision work (birthday) |
|---|---|---|---|
| AES-128 | 2^128 | SHA-256 | 2^128 |
| AES-256 | 2^256 | SHA-384 | 2^192 (already impractical) |
Verify: hash output ≈ twice the key length keeps collisions as hard as a key search
Why: §7.3: because collisions cost only 2^(n/2), you double the bits to match a 2^k key. SHA-256 for AES-128, SHA-384 for AES-256.
Comparison
Comparison matrix
From §7.3 Matching hash length to key length: refill the key brute-force column from what you know. The rest of the table is as it appeared.
| cipher | key brute-force | hash | collision work (birthday) |
|---|---|---|---|
| AES-128 | 2^128 | SHA-256 | 2^128 |
| AES-256 | 2^256 | SHA-384 | 2^192 (already impractical) |
Concept
A quick reference for which hash to reach for and which to retire.
| algorithm | status | output sizes |
|---|---|---|
| MD5 | BROKEN (2005) — never use | 128 bits |
| SHA-1 | BROKEN (2017, SHAttered) — never use | 160 bits |
| SHA-2 | secure, but length-extendable | 256 / 384 / 512 bits |
| SHA-3 | secure, length-extension immune | 256 / 384 / 512 bits |
Discrimination
Sort into buckets
Sort these by output sizes, from memory, without looking back at §7.3 The algorithm scoreboard. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Anomaly
Predict first
A student writes this, and it looks reasonable:
A student: 'SHA-256 has 256-bit outputs, so finding a collision takes 2^256 tries — astronomically safe.'
It is wrong. Say what breaks — and say it before you turn the page.
Correct: 2^n is the cost of a PREIMAGE. Collisions follow the birthday bound at ~2^(n/2).
A student: how much work IS a SHA-256 collision?
Why: 2^n is the cost of a PREIMAGE. Collisions follow the birthday bound at ~2^(n/2). For SHA-256 that's ~2^128, not 2^256 — a square root smaller.
Trap
A student: 'SHA-256 has 256-bit outputs, so finding a collision takes 2^256 tries — astronomically safe.'
\[ \text{collision work} \stackrel{?}{=} 2^{256} \]
Use 2^n for collisions instead of the birthday bound
Why: Wrong. 2^n is the cost of a PREIMAGE. Collisions follow the birthday bound at ~2^(n/2). For SHA-256 that's ~2^128, not 2^256 — a square root smaller.
A student: how much work IS a SHA-256 collision?
\[ \text{collision work} \approx 2^{n/2} = 2^{128} \]
Use the birthday bound: ~2^128 for a 256-bit hash
Why: §7.3: by the birthday paradox, collisions appear after about 2^(n/2) hashes. That square-root loss is exactly why SHA-256 (collisions at 2^128) is matched with AES-128.
Notation
Annotate
From Trap: 'a 256-bit hash needs 2^256 work to collide' — read this one piece at a time. What is each part doing?
On: \( \text{collision work} \stackrel{?}{=} 2^{256} \)
Section
Part 5 · §7.4 counting without seeing
Concept
A hacker tells a journalist they stole 150 million records and demands payment for the leak. The journalist must judge whether the claim is plausible without ever seeing the records.
Hashes give a slick test. The trick rests on a counting intuition — start with balls in a box, then translate it to hashes.
Explain it
Discussion prompt
Explain §7.4 A scenario: verifying a hacker's claim to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
A hacker tells a journalist they stole 150 million records and demands payment for the leak. The journalist must judge whether the claim is plausible without ever seeing the records.
Intuition
A box holds balls numbered 1–100. You draw with replacement n times and remember the smallest number you ever saw. The more you draw, the lower that smallest number tends to be.
Draw twice and the minimum might be 40. Draw 150 million times and you will almost surely have seen ball 1. A claimed minimum of 50 after 150M draws is astronomically unlikely.
Ask yourself: after millions of draws from 1–100, what's the smallest you've almost certainly seen? (1 — with that many draws the minimum is essentially guaranteed to bottom out.)
Analogy
Discussion prompt
Explain §7.4 Balls in a box: the lowest you draw shrinks by analogy to something with no Computer Security in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
A box holds balls numbered 1–100. You draw with replacement n times and remember the smallest number you ever saw. The more you draw, the lower that smallest number tends to be.
Step zero
Discussion prompt
§7.4 Turn records into draws with a hash — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Have the hacker hash all their records; treat each digest as a random…
Answer:
Worked example
Have the hacker hash all their records; treat each digest as a random number
Why: Avalanche makes each hash look like an independent uniform draw from the huge space of n-bit strings — the 'balls in a box', just with an enormous box.
Ask the hacker to return the 10 LOWEST hashes among all their records
Why: With 150M genuine random-looking hashes, the lowest ones cluster extremely close to zero — and the journalist can compute how close they SHOULD be.
Check the 10 lowest values are consistent with 150M random bitstrings
Why: If the hacker really has 150M records, the minima sit in a predictable tiny range. If they only have a few thousand records, their lowest hashes will be far too large.
Verify: the lowest hashes are a count estimator
Why: §7.4: 'how low is the minimum hash' reveals 'how many items were hashed' — exactly the balls-in-a-box logic applied to digests. This is the idea behind MinHash-style counting.
Concept
A hacker could fake small hashes by searching for inputs that hash low. So the journalist also requires the actual records whose hashes are those 10 lowest values.
To produce 10 genuinely-lowest hashes AND valid-looking records for them, the hacker would need ~150M real, valid records — which is precisely the claim being tested. Faking it is as hard as actually having the data.
Counterexample
Discussion prompt
A hacker could fake small hashes by searching for inputs that hash low. So the journalist also requires the actual records whose hashes are those 10 lowest values.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Ranking
Put in order
Put the moves of §7.4 Block offline precomputation into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Given unlimited time before the challenge, an attacker could grind out inputs with tiny hashes and fake a large dataset.
Worked example
Worry: the hacker could precompute lots of low hashes in advance
Why: Given unlimited time before the challenge, an attacker could grind out inputs with tiny hashes and fake a large dataset.
The journalist sends a fresh random word — say 'fubar' — to PREPEND to every record before hashing
Why: Now the relevant hashes are H('fubar' ‖ record). The attacker couldn't have precomputed these because 'fubar' was unknown until challenge time.
Time-box the response
Why: A short deadline means the hacker cannot grind 'fubar'-prefixed low hashes on the spot — there isn't enough time to fake 150M of them.
Verify: the random prefix + deadline force an ONLINE, honest computation
Why: §7.4: salting with a fresh challenge value defeats precomputation — the same idea that protects password hashes against precomputed tables (L33).
Blank canvas
Draw it
Draw what §7.4 Block offline precomputation just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Anomaly
Predict first
A student writes this, and it looks reasonable:
A student: 'The journalist only needs the 10 lowest hash VALUES; asking for the underlying records is overkill.'
It is wrong. Say what breaks — and say it before you turn the page.
Correct: A cheater can search for inputs that hash low and submit those values without owning 150M records — and could even precompute them in advance.
A student: what makes the lowest-hash test actually sound?
Why: A cheater can search for inputs that hash low and submit those values without owning 150M records — and could even precompute them in advance. Bare hash values prove nothing about the dataset's size.
Trap
A student: 'The journalist only needs the 10 lowest hash VALUES; asking for the underlying records is overkill.'
Accept low hash values with no records and no fresh challenge
Why: A cheater can search for inputs that hash low and submit those values without owning 150M records — and could even precompute them in advance. Bare hash values prove nothing about the dataset's size.
A student: what makes the lowest-hash test actually sound?
Require the records for those hashes AND a fresh random prefix under a deadline
Why: §7.4: demanding the matching records forces real data; the random prefix ('fubar') plus a time limit forces an online computation, defeating precomputation. All three together make faking as hard as having the data.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
Constraint
Discussion prompt
Run The cryptographic-hash playbook with this step confiscated:
Algorithms: MD5 and SHA-1 are broken; use SHA-2 or SHA-3. SHA-2 is length-extendable (output = full state) ⇒ Hash(K‖M) MACs fail; SHA-3 hides state and is immune.
Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.
Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.
Answer:
Pattern
Edge cases
Discussion prompt
The cryptographic-hash playbook works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
Elimination
Eliminate the wrong options
Which property must hold to stop this attack, and roughly how much work does breaking it take for an n-bit hash?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: §7.2: the target input x is FIXED and given to the attacker, who must find a different x′ matching it — that is exactly second-preimage resistance, not collision resistance. Because the target is pinned, there is no birthday speedup: breaking it costs about 2^n work, the same as a preimage search, not the 2^(n/2) birthday cost of a free-choice collision.
Check
An attacker is handed one specific, legitimate contract x and wants a DIFFERENT contract x′ ≠ x that hashes to the same value, H(x′) = H(x), so the forged contract passes any integrity check tied to that hash. Think it through before clicking.
Check your understanding
Which property must hold to stop this attack, and roughly how much work does breaking it take for an n-bit hash?
Answer: A
Why: §7.2: the target input x is FIXED and given to the attacker, who must find a different x′ matching it — that is exactly second-preimage resistance, not collision resistance. Because the target is pinned, there is no birthday speedup: breaking it costs about 2^n work, the same as a preimage search, not the 2^(n/2) birthday cost of a free-choice collision.
Concept
Concept
Concept
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — What a Hash Is · The Three Security Properties · Hashes for Integrity · Real Algorithms & Length Extension · A Clever Application: The Lowest Hash. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
You can now define a cryptographic hash and its avalanche/deterministic/unkeyed nature, tell apart preimage, second-preimage and collision resistance, use a hash for integrity over a trusted channel, name the secure families and explain SHA-2's length-extension flaw and the birthday bound, and apply the lowest-hash scheme to verify a record count without seeing the data.
| Idea | § | The one-line version |
|---|---|---|
| What a hash is | 7.1 | Any-length M ⇒ fixed-length fingerprint; deterministic, unkeyed, avalanche |
| Preimage | 7.2 | Given y, infeasible to find any x with H(x)=y (~2^n) |
| Second-preimage | 7.2 | Given x, infeasible to find x′≠x with same hash (~2^n) |
| Collision | 7.2 | Find any x≠x′ colliding; birthday bound ~2^(n/2) |
| Integrity | 7.2 | Publish H(file) on a TRUSTED channel; bare hash authenticates no one |
| Algorithms | 7.3 | MD5/SHA-1 broken; SHA-2 length-extendable, SHA-3 immune |
| Match lengths | 7.3 | SHA-256 ↔ AES-128, SHA-384 ↔ AES-256 (birthday) |
| Lowest hash | 7.4 | Min of n random hashes counts n; demand records + fresh prefix |
Want this taught 1-on-1? Alexander tutors Computer Security — $55/session, free consultation.