This deck shows how Huffman coding builds an optimal variable-length, prefix-free code from symbol frequencies. It traces the merge-the-two-smallest algorithm by hand on a six-symbol alphabet, encodes and decodes a message, and sketches the exchange argument for why the greedy strategy is provably optimal. It targets codes that are not prefix-free, merging the wrong nodes or forgetting to reinsert the merged node, the reversal that would give frequent symbols longer codes, and the false belief that fixed-length coding is never worse than Huffman.
Subject: CS3000 Algorithms · 131 slides · symbolic lesson
Open the interactive version of this deck · Homework for this lesson
Title
CS3000 Algorithms · Greedy Algorithms
Building the shortest possible prefix-free code - and proving you can't do better.
Objectives
Huffman coding is the algorithm behind ZIP, JPEG, and countless other compressors. By the end of this lesson you can:
Warm-up
Discussion prompt
Before we open Huffman Coding: without looking back, what was the main idea of Greedy Algorithms & Exchange Arguments, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
That deck explains what a greedy algorithm is and traces two of them: interval scheduling by earliest finish time, and fractional knapsack by highest value-to-weight ratio. It then gives the two standard ways to prove a greedy algorithm optimal, greedy-stays-ahead and the exchange argument. It targets the beginner's biggest gap, which is getting a proof started, along with the misconceptions that greedy always works, that any scheduling criterion is as good as another, that a few working examples count as a proof, and that an exchange can be made carelessly.
Concept
Before any new material: cover the screen.
You have named 14 reusable moves so far. Say as many as you can out loud, by number, from memory.
Do not advance until you have actually tried. Getting four of eight is information; skipping the exercise is not.
Here they are. Score yourself.
Today adds no new moves. Every proof in this lesson is built out of the list above. That is the whole point of the list.
The question that starts every proof from here on is not how do I begin. It is which of these applies here?
Counterexample
Discussion prompt
You have named 14 reusable moves so far. Say as many as you can out loud, by number, from memory.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Do not advance until you have actually tried. Getting four of eight is information; skipping the exercise is not.
Section
Section 1
Concept
Suppose you must store or send symbols from some alphabet, and you have to pick one fixed number of bits for every symbol, no matter how often it shows up.
With a fixed-length code, every symbol costs the same, so the total number of bits is just the number of symbols times the code length.
\[ \text{total bits} = (\text{number of symbols}) \times (\text{bits per symbol}) \]
fixed-length code — A code where every symbol is represented using the exact same number of bits, regardless of how frequently that symbol occurs.
Analogy
Discussion prompt
Explain The cost of a fixed-length code by analogy to something with no CS3000 Algorithms in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Suppose you must store or send symbols from some alphabet, and you have to pick one fixed number of bits for every symbol, no matter how often it shows up.
Picture it
Animation
Shows: Against a fixed-length code — a rendered Manim animation.
Rendered with Manim.
Takeaway: The saving comes entirely from the frequency skew.
Ranking
Put in order
Put the moves of How many bits does a fixed-length code need for six symbols? into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Each bit doubles the number of available patterns, so k bits give two to the k possible patterns.
Worked example
Our running example alphabet has six symbols: a, b, c, d, e, f. A fixed-length code needs enough bits to give each one a distinct pattern.
Count how many distinct patterns k bits can make
Why: Each bit doubles the number of available patterns, so k bits give two to the k possible patterns.
\[ k \text{ bits} \Rightarrow 2^{k} \text{ distinct patterns} \]
Find the smallest k with two to the k at least six
Why: Two bits give only four patterns - not enough for six symbols. Three bits give eight patterns, which covers six with room to spare.
\[ 2^{2}=4 < 6 \le 8 = 2^{3} \]
Verify three bits is enough and two is not
Why: Since 4 is less than 6 and 8 is at least 6, three is the smallest number of bits that can label all six symbols distinctly.
\[ k = 3 \text{ bits per symbol} \]
Picture it
Animation
Shows: Each line of the worked example "How many bits does a fixed-length code need for six symbols?", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Since 4 is less than 6 and 8 is at least 6, three is the smallest number of bits that can label all six symbols distinctly.
Intuition
Picture a texting abbreviation system. If you text 'lol' fifty times a day but 'defenestrate' once a year, it makes no sense to give both the same length shortcut.
A fixed-length code treats the everyday word and the once-a-year word identically. Every use of the common symbol pays the same toll as the rare one - and the common symbol is used far more often, so that toll adds up fast.
Explain it
Discussion prompt
Explain Why waste bits on a symbol you rarely use? to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Picture a texting abbreviation system. If you text 'lol' fifty times a day but 'defenestrate' once a year, it makes no sense to give both the same length shortcut.
Concept
Huffman coding assigns short codes to symbols that occur often, and long codes to symbols that occur rarely. The total cost is dominated by the frequent symbols, so that is exactly where you want to save bits.
variable-length code — A code where different symbols may use different numbers of bits, chosen so that the total encoded length is as small as possible given the symbols' frequencies.
Picture it
Animation
Shows: Frequent symbols get short codes — a rendered Manim animation.
Rendered with Manim.
Takeaway: The rarest symbols merge earliest, so they end up deepest and longest.
Step zero
Discussion prompt
Total bits if every symbol used a fixed-length code — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Use the fixed-length size found earlier
Answer:
Worked example
Our alphabet's frequencies, out of 100 symbols total: a appears 45 times, b 13, c 12, d 16, e 9, f 5.
| Symbol | Frequency |
|---|---|
| a | 45 |
| b | 13 |
| c | 12 |
| d | 16 |
| e | 9 |
| f | 5 |
| Total | 100 |
Use the fixed-length size found earlier
Why: We showed three bits are needed to distinguish six symbols.
\[ 3 \text{ bits per symbol} \]
Multiply by the total symbol count
Why: Every one of the 100 symbols costs the same three bits under a fixed-length code.
\[ 100 \times 3 = 300 \text{ bits} \]
Verify by checking a single symbol's share
Why: Each of the 100 symbol occurrences contributes exactly 3 bits, and 100 occurrences times 3 bits really is 300 bits - no symbol gets a discount.
\[ 100 \text{ symbols} \times 3 \text{ bits} = 300 \text{ bits} \]
Picture it
Animation
Shows: Each line of the worked example "Total bits if every symbol used a fixed-length code", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Each of the 100 symbol occurrences contributes exactly 3 bits, and 100 occurrences times 3 bits really is 300 bits - no symbol gets a discount.
Section
Section 2
Concept
If codes can have different lengths, how does a reader know where one symbol's code ends and the next begins? The answer is a design rule on the codes themselves.
prefix-free code — A code in which no symbol's code is the beginning (prefix) of any other symbol's code. Every codeword is self-terminating: once you have matched one, you know it cannot be the start of a longer, different codeword.
Intuition
Imagine reading a sentence with no spaces between words. If one word were secretly the start of another word, you would not know when to stop reading one and start the next.
A prefix-free code guarantees this never happens: the moment the bits you have read so far match a complete code, that match is final. No other code shares that same beginning, so there is nothing left to wait for.
Estimation
Predict first
Three symbols, x, y, z, are assigned these codes: x gets 0, y gets 10, z gets 11.
Commit before you compute: what does Check whether a small code is prefix-free come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Verify the whole set is prefix-free
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. No codeword starts with another codeword's full pattern, so this code is prefix-free and any string of these codes can be decoded without ambiguity.
Worked example
Three symbols, x, y, z, are assigned these codes: x gets 0, y gets 10, z gets 11.
| Symbol | Code |
|---|---|
| x | 0 |
| y | 10 |
| z | 11 |
Compare every pair of codes
Why: A code is prefix-free only if no codeword is a leading chunk of a longer one. Check all three pairs.
Check x's code (0) against y's and z's
Why: y is 10 and z is 11 - neither begins with 0, so x's code is not a prefix of either.
Check y's code (10) against z's code (11)
Why: They share the first bit but differ in the second, and neither is a prefix of the other since both have the same length.
Verify the whole set is prefix-free
Why: No codeword starts with another codeword's full pattern, so this code is prefix-free and any string of these codes can be decoded without ambiguity.
Picture it
Animation
Shows: Why the code is prefix-free — a rendered Manim animation.
Rendered with Manim.
Takeaway: Which is what makes the stream decodable with no separators.
Anomaly
Predict first
A student writes this, and it looks reasonable:
A student assigns: x gets 0, y gets 1, z gets 01 - reusing 0 and 1 as a 'combo' for the third symbol.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Read left to right: the first bit, 0, already matches x's complete code.
Give every code a pattern that cannot be completed into a different codeword - here, letting z be 11 keeps every code the same length as its siblings at that branch.
Why: Read left to right: the first bit, 0, already matches x's complete code. But the whole string, 01, also matches z's complete code. The same two bits mean two different things.
Trap
A student assigns: x gets 0, y gets 1, z gets 01 - reusing 0 and 1 as a 'combo' for the third symbol.
| Symbol | Code |
|---|---|
| x | 0 |
| y | 1 |
| z | 01 |
Try to decode the bit string 01
Why: Read left to right: the first bit, 0, already matches x's complete code. But the whole string, 01, also matches z's complete code. The same two bits mean two different things.
Give every code a pattern that cannot be completed into a different codeword - here, letting z be 11 keeps every code the same length as its siblings at that branch.
| Symbol | Code |
|---|---|
| x | 0 |
| y | 10 |
| z | 11 |
Decode 011 with this fixed code
Why: Read 0 - matches x completely, output x, then continue with 11 - matches z completely. The result, x then z, is the only possible reading.
Pattern
Step through it
Step through Trap: a code that is not prefix-free one row at a time. What is driving the change, and what would the row after the last one be?
Concept
Every prefix-free code can be pictured as a binary tree: start at the root, and each bit of a code tells you whether to go to the left child (0) or the right child (1).
Symbols sit only at the leaves of the tree - the nodes with no children. Internal nodes are just junctions on the way down; they never carry a symbol themselves.
Intuition
If every symbol lives at a leaf, then reaching one symbol's leaf means the path stops there - there is no branch left to continue down.
That is exactly why a tree automatically gives you a prefix-free code: a leaf's path can never continue on into a deeper leaf's path, since a leaf has nowhere further to go.
Picture it
Figure (svg): A small binary tree with x as the left child of the root, and y and z as the two children of the root's right child
Discussion prompt
Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.
Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.
Answer:
Use the fixed x, y, z code from before: x is 0, y is 10, z is 11. Draw the tree that produces it.
Worked example
Use the fixed x, y, z code from before: x is 0, y is 10, z is 11. Draw the tree that produces it.
Place x at depth one on the left
Why: x's code is a single 0, meaning one left branch from the root lands directly on x's leaf.
Branch right from the root, then split again
Why: Both y and z start with 1, so they share the root's right branch; a second branch (0 for y, 1 for z) then separates them one level deeper.
Figure (svg): A small binary tree with x as the left child of the root, and y and z as the two children of the root's right child
Verify the leaf paths match the codes
Why: Root to x is just left (0); root to y is right-then-left (10); root to z is right-then-right (11). These match exactly the codes we started with, confirming the tree and the code agree.
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Verify the leaf paths match the codes
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Use the fixed x, y, z code from before: x is 0, y is 10, z is 11. Draw the tree that produces it.
Section
Section 3
Concept
From here on, every example uses the same six-symbol alphabet, with frequencies drawn from a classic textbook example, out of 100 total occurrences.
| Symbol | Frequency |
|---|---|
| a | 45 |
| b | 13 |
| c | 12 |
| d | 16 |
| e | 9 |
| f | 5 |
| Total | 100 |
Notice how spread out the frequencies are: a is nearly half of everything, while f shows up only five times. That spread is exactly what makes variable-length coding worth doing.
Pattern
Step through it
Step through Our alphabet and its frequencies one row at a time. What is driving the change, and what would the row after the last one be?
Picture it
Animation
Shows: The table has to travel with the data — a rendered Manim animation.
Rendered with Manim.
Takeaway: For short messages that overhead can exceed the saving.
Concept
Huffman's algorithm starts with one node per symbol and repeatedly fuses two nodes into a single new node, until only one node - the root - remains.
Each fusion reduces the node count by exactly one. Starting from six symbol-nodes and ending at a single root means exactly five fusions happen along the way.
\[ 6 \text{ nodes} \to 1 \text{ root} \Rightarrow 6 - 1 = 5 \text{ merges} \]
Intuition
What feels wrong about this?
The rule says: repeatedly combine the two least frequent symbols.
But merging pushes both of those symbols one level deeper in the tree, which makes their codes longer.
_Plain English only. No notation, no algebra. Just say what bothers you._
The feeling: it looks backwards. You are deliberately penalizing symbols, over and over, right at the start.
That feeling is the proof. It is not a substitute for the proof — it is the thing the proof writes down.
The resolution is that depth is a resource somebody has to spend. Every merge costs depth to exactly two symbols, so you spend it on the two that will be transmitted least often. The rarest symbols are the cheapest place to put the cost.
Sorting
Sort into buckets
These are the pieces of Huffman Coding, out of order. Put each one back under the part of the lesson it belongs to.
Concept
At every step, look at all the current nodes - original symbols and any nodes already merged - and combine the two with the smallest frequency.
The new node's frequency is just the sum of the two it replaced, and it goes back into the pool of nodes still waiting to be merged.
\[ \text{new node's frequency} = f_1 + f_2 \]
Concept
A priority queue and one loop. Every pass removes the two lightest nodes and puts back a single node weighing what they weighed together.
HUFFMAN(symbols, freq)
Q = priority queue of leaf nodes, keyed by freq
while Q.size > 1
a = Extract-Min(Q)
b = Extract-Min(Q)
z = new node
z.left = a
z.right = b
z.freq = a.freq + b.freq
Insert(Q, z)
return Extract-Min(Q)The queue shrinks by exactly one each pass: two out, one in. Starting from k symbols that means k minus one merges, and the last node standing is the root of the finished tree.
Notation
Every line of HUFFMAN says one thing. Read the line, then read what it does — not the other way round.
Annotate
Invariant
The total weight in the queue never changes; only the number of nodes falls. Depth in the finished tree is exactly how many merges a symbol survived.
Step through it
Before each merge, name the two lightest yourself.
Picture it
Animation
Shows: HUFFMAN executing: the current line of pseudocode is highlighted while the data it touches changes.
Rendered with Manim.
Takeaway: Repeatedly merge the two lightest nodes — rare symbols merge earliest, sit deepest, and so earn the longest codes.
Intuition
Picture weights on a table, each labeled with a symbol. You always pick up the two lightest weights, tape them together into one heavier weight, and set that combined weight back down among the others.
Because you always start with the lightest pair, the lightest original weights end up wrapped in the most layers of tape by the end - which is exactly why rare symbols end up deepest in the tree, with the longest codes.
Intuition
Watch me not know the answer. This is what the first two minutes actually look like.
Frequencies for five symbols: 45, 13, 12, 16, 9.
Try merging the two heaviest first: 45 and 16
Why: It is the symmetric rule and it takes the same amount of code to write, so it is worth ten seconds to see what it does.
It buries the most common symbol
Why: That merge sends the 45-weight symbol one level down immediately, and every later merge involving that combined node pushes it deeper still. The most frequent symbol ends up with one of the longest codes.
Dead end. Not a mistake — a move that was worth trying and did not pay off. This happens in most proofs.
Back up. Compute the cost of both and compare
Why: Do not argue about it — the weighted total is a number. Build both trees and add up frequency times depth.
Merging heaviest-first produces a strictly larger total. The symmetric rule is not merely unproven, it is wrong — and finding that out cost one arithmetic check rather than an argument.
Trying the mirror image of a rule is a cheap and underused move. It either produces a counterexample or it sharpens your sense of why the real rule works.
The expert does not see the whole path in advance. The expert tries something, reads the result, and adjusts. That is the skill.
Step zero
Discussion prompt
Trace the Huffman merges — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Merge 1: combine the two smallest, f (5) and e (9)
Answer:
Worked example
Starting pool: f is 5, e is 9, c is 12, b is 13, d is 16, a is 45.
Merge 1: combine the two smallest, f (5) and e (9)
Why: f and e are the lowest frequencies in the pool. Their combined node has frequency 5 plus 9.
\[ f(5) + e(9) = 14 \;\Rightarrow\; \text{new node } fe:14 \]
Merge 2: combine the two smallest now, c (12) and b (13)
Why: With fe:14 back in the pool, the current smallest two are c and b - fe (14) is now larger than both.
\[ c(12) + b(13) = 25 \;\Rightarrow\; \text{new node } cb:25 \]
Merge 3: combine fe (14) and d (16)
Why: The remaining pool is fe:14, d:16, cb:25, a:45. The smallest two are fe and d.
\[ fe(14) + d(16) = 30 \;\Rightarrow\; \text{new node } fed:30 \]
Merge 4: combine cb (25) and fed (30)
Why: The remaining pool is cb:25, fed:30, a:45. The smallest two are cb and fed.
\[ cb(25) + fed(30) = 55 \;\Rightarrow\; \text{new node } cbfed:55 \]
Merge 5: combine a (45) and cbfed (55)
Why: Only two nodes remain, a and cbfed, so they merge into the root.
\[ a(45) + cbfed(55) = 100 \;\Rightarrow\; \text{root}: 100 \]
Verify the merge count and the total
Why: Six symbols required exactly five merges, matching the count predicted earlier, and the root's frequency is 100 - the full original total, with nothing lost or double-counted.
\[ 5 \text{ merges}, \quad \text{root frequency} = 100 \]
Picture it
Animation
Shows: Huffman merges the two rarest symbols — a rendered Manim animation.
Rendered with Manim.
Takeaway: Built upward from the rarest pair, so the rarest symbols end up deepest.
Estimation
Predict first
Label every left branch 0 and every right branch 1. A symbol's code is the sequence of labels on the path from the root down to its leaf.
Commit before you compute: what does Read the codes off the finished tree come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Verify the codes are prefix-free
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Every code sits at a leaf of the same tree, and leaves have no children, so no code's path can continue into another leaf's path.
Worked example
Label every left branch 0 and every right branch 1. A symbol's code is the sequence of labels on the path from the root down to its leaf.
Trace a's path
Why: a is a direct child of the root - the left child, since a was the smaller of the last two nodes merged.
\[ a: 0 \]
Trace c and b's paths
Why: Both sit under the cb node, reached by going right from the root (into cbfed:55) then left (into cb:25). c is cb's left child and b is cb's right child.
\[ c: 100 \qquad b: 101 \]
Trace d, then f and e's paths
Why: The fed node is reached by going right then right again. d is fed's right child. fe is fed's left child; within fe, f is the left child and e is the right child.
\[ d: 111 \qquad f: 1100 \qquad e: 1101 \]
| Symbol | Frequency | Code | Length (bits) |
|---|---|---|---|
| a | 45 | 0 | 1 |
| b | 13 | 101 | 3 |
| c | 12 | 100 | 3 |
| d | 16 | 111 | 3 |
| e | 9 | 1101 | 4 |
| f | 5 | 1100 | 4 |
Verify the codes are prefix-free
Why: Every code sits at a leaf of the same tree, and leaves have no children, so no code's path can continue into another leaf's path. For example, a's code (0) never appears as the first digit of b, c, or d's codes, all of which start with 1.
Concept
Here is the whole tree at once, with each internal node's combined frequency and every 0/1 branch label shown.
Figure (svg): The full Huffman tree for symbols a, b, c, d, e, f with frequencies 45, 13, 12, 16, 9, 5, showing merge frequencies at each internal node and 0/1 edge labels down to each leaf
Notice the shape: a, the most frequent symbol, sits closest to the root with the shortest code. f and e, the two rarest symbols, sit at the bottom of the tree, sharing the longest codes.
Anomaly
Predict first
A student writes this, and it looks reasonable:
A student thinks 'Huffman merges nodes' and merges the two largest frequencies first instead of the two smallest - combining a (45) and d (16) right away.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: This buries the most frequent symbol, a, one level deeper than necessary right from the start, which can only raise its code length and its contribution to the total cost.
Always identify the two smallest frequencies currently in the pool, no matter whether they are original symbols or already-merged nodes.
Why: This buries the most frequent symbol, a, one level deeper than necessary right from the start, which can only raise its code length and its contribution to the total cost.
Trap
A student thinks 'Huffman merges nodes' and merges the two largest frequencies first instead of the two smallest - combining a (45) and d (16) right away.
\[ a(45) + d(16) = 61 \quad (\text{wrong first move}) \]
Push the two heaviest symbols together early
Why: This buries the most frequent symbol, a, one level deeper than necessary right from the start, which can only raise its code length and its contribution to the total cost.
A second version of the same mistake: correctly merging f (5) and e (9) into fe:14, but then forgetting to put fe:14 back into the pool - leaving it stranded and out of consideration.
Continue merging only the leftover originals c, b, d, a and never revisit fe
Why: With fe left out, the algorithm can no longer build one connected tree over all six symbols - some symbols end up unreachable from a single root.
Always identify the two smallest frequencies currently in the pool, no matter whether they are original symbols or already-merged nodes.
\[ f(5) + e(9) = 14 \quad (\text{correct first move}) \]
Merge the two smallest, then reinsert the result
Why: The new node's frequency (14) goes right back into the pool, competing on equal footing with every remaining node in the very next round.
Keep repeating until one node is left
Why: Each round always looks at the full current pool - including every node created so far - so nothing is ever left stranded outside the growing tree.
Notation
Annotate
From Trap: merging the wrong nodes — read this one piece at a time. What is each part doing?
On: \( a(45) + d(16) = 61 \quad (\text{wrong first move}) \)
Pattern
1. Put every symbol in a pool, each as its own node labeled with its frequency
Why: This is the starting point: one leaf-to-be for every symbol, nothing merged yet.
2. While more than one node remains, find the two nodes with the smallest frequency
Why: Always the two smallest in the current pool - which may include nodes created by earlier merges, not just original symbols.
3. Merge them into a new node whose frequency is their sum, and put that new node back into the pool
Why: Reinserting is essential - skip it and the algorithm cannot connect every symbol into one tree.
4. Repeat until a single node remains: that node is the root
Why: For n symbols this takes exactly n minus 1 merges, since each merge shrinks the pool by one.
5. Read off each symbol's code as the 0/1 path from the root down to its leaf
Why: Left branches contribute a 0, right branches a 1; the path length is the code's length in bits.
Real world
Discussion prompt
Outside this lesson: where does Huffman Coding actually turn up? Name one concrete situation — a job, a piece of software someone ships, a decision somebody has to make — and say which part of The Huffman algorithm, step by step is doing the work in it.
Hint: Vague is the failure mode here. "Engineering" is not a situation; "deciding whether this build is fast enough to ship" is.
Answer:
That deck shows how Huffman coding builds an optimal variable-length, prefix-free code from symbol frequencies. It traces the merge-the-two-smallest algorithm by hand on a six-symbol alphabet, encodes and decodes a message, and sketches the exchange argument for why the greedy strategy is provably optimal. It targets codes that are not prefix-free, merging the wrong nodes or forgetting to reinsert the merged node, the reversal that would give frequent symbols longer codes, and the false belief that fixed-length coding is never worse than Huffman.
Picture it
Animation
Shows: Ties give different trees, same cost — a rendered Manim animation.
Rendered with Manim.
Takeaway: Which is why two correct answers can look nothing alike.
Elimination
Eliminate the wrong options
Which two nodes does the Huffman algorithm merge first?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Huffman always merges the two nodes with the smallest frequency in the current pool. Here, x (3) and z (4) are the two lowest frequencies, so they merge first into a node of frequency 7.
Check
A new alphabet has four symbols with these frequencies: w is 7, x is 3, y is 9, z is 4.
| Symbol | Frequency |
|---|---|
| w | 7 |
| x | 3 |
| y | 9 |
| z | 4 |
Check your understanding
Which two nodes does the Huffman algorithm merge first?
Answer: A
Why: Huffman always merges the two nodes with the smallest frequency in the current pool. Here, x (3) and z (4) are the two lowest frequencies, so they merge first into a node of frequency 7.
Pattern
Step through it
Step through Check yourself: which two nodes merge first? one row at a time. What is driving the change, and what would the row after the last one be?
Section
Section 4
Concept
Once every symbol has a code, the total number of bits used to encode a source is found by weighting each code's length by how often that symbol occurs.
\[ L = \sum_{i} f_i \cdot \ell_i \]
In this formula, each symbol's frequency is multiplied by its own code's length in bits, and every one of those products is added together to get the grand total, L.
Ranking
Put in order
Put the moves of Compute the Huffman total for our alphabet into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. This gives how many bits that symbol contributes across all of its occurrences, shown in the table's last column.
Worked example
Use the code table built earlier: a is 1 bit, b is 3, c is 3, d is 3, e is 4, f is 4.
| Symbol | Frequency | Length (bits) | Frequency times length |
|---|---|---|---|
| a | 45 | 1 | 45 |
| b | 13 | 3 | 39 |
| c | 12 | 3 | 36 |
| d | 16 | 3 | 48 |
| e | 9 | 4 | 36 |
| f | 5 | 4 | 20 |
Multiply each frequency by its code length
Why: This gives how many bits that symbol contributes across all of its occurrences, shown in the table's last column.
Add every product together
Why: Summing the last column gives the total number of bits the whole 100-symbol source needs under this code.
\[ 45+39+36+48+36+20 = 224 \]
Verify the sum matches the weighted-sum formula
Why: 224 bits is exactly the value of L from the formula, computed directly from the table - no symbol was skipped or double-counted.
\[ L = 224 \text{ bits} \]
Picture it
Animation
Shows: Each line of the worked example "Compute the Huffman total for our alphabet", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: 224 bits is exactly the value of L from the formula, computed directly from the table - no symbol was skipped or double-counted.
Concept
Dividing the total bits by the number of symbols gives a single number you can compare across different codes: the average bits spent per symbol.
\[ \text{average} = \frac{L}{\text{number of symbols}} = \frac{224}{100} = 2.24 \text{ bits per symbol} \]
Compare that to a fixed-length code, which always spends exactly three bits per symbol, no matter which symbol it is. Huffman spends noticeably less on average - and that difference adds up across every symbol in a long message.
Intuition
Think of it like a grocery bill: a two-dollar item you buy fifty times costs you far more overall than a twenty-dollar item you buy once. The frequency, not just the price tag, decides your total spending.
The same logic drives Huffman's total: a symbol's contribution depends on both its code length and how often it shows up. That is why the algorithm fights hardest to shorten the codes of the busiest symbols.
Concept
A fixed-length code has to give every symbol in the alphabet its own distinct pattern, so its bit budget depends only on how many symbols exist - never on their frequencies.
\[ \text{bits per symbol} = \lceil \log_{2}(\text{number of symbols}) \rceil \]
This is why a fixed-length code cannot adapt: whether one symbol dominates the source or all symbols are equally common, the per-symbol cost never changes.
Hypothesis
Predict first
Compute the fixed-length total and compare is about to be worked. State your hypothesis first: which rule or definition decides this one, and what is the first move it forces? Then watch whether the example agrees with you.
Correct: Multiply the fixed bit cost by the total symbol count
Why: All 100 symbol occurrences pay the same fixed rate of three bits each.
A hypothesis you wrote down is falsifiable; a vague sense of how it will go is not. If the example opens somewhere else, that gap is the thing worth chasing.
Worked example
Our alphabet has six symbols, so, as shown earlier, a fixed-length code needs three bits per symbol.
Multiply the fixed bit cost by the total symbol count
Why: All 100 symbol occurrences pay the same fixed rate of three bits each.
\[ 100 \times 3 = 300 \text{ bits} \]
Compare the two totals
Why: Huffman's total was 224 bits; the fixed-length total is 300 bits. Subtracting gives the number of bits saved.
\[ 300 - 224 = 76 \text{ bits saved} \]
Verify the savings as a percentage
Why: 76 bits saved out of an original 300 is about one quarter of the total - Huffman uses roughly 25 percent fewer bits than the fixed-length code on this alphabet.
\[ \frac{76}{300} \approx 0.253 \;(\approx 25\%) \]
Worked example
A smaller alphabet, used only in the next two slides: p occurs 70 times, q occurs 20 times, r occurs 10 times, out of 100 total.
| Symbol | Frequency |
|---|---|
| p | 70 |
| q | 20 |
| r | 10 |
Merge the two smallest, q (20) and r (10)
Why: q and r are the lowest frequencies in this three-symbol pool.
\[ q(20)+r(10)=30 \;\Rightarrow\; qr:30 \]
Merge the remaining two nodes, p (70) and qr (30)
Why: Only p and the merged qr node remain, so they combine into the root.
\[ p(70)+qr(30)=100 \;\Rightarrow\; \text{root}: 100 \]
Read off the codes
Why: p is a direct child of the root, so its code is one bit. q and r sit one level deeper, under the qr node, so their codes are two bits each.
\[ p: 0 \; (1\text{ bit}) \qquad q: 10 \; (2\text{ bits}) \qquad r: 11 \; (2\text{ bits}) \]
Verify the total against the weighted-sum formula
Why: 70 times 1, plus 20 times 2, plus 10 times 2, gives the total bits Huffman needs for this alphabet.
\[ 70(1)+20(2)+10(2) = 70+40+20 = 130 \text{ bits} \]
Picture it
Animation
Shows: The greedy choice, and why it is safe — a rendered Manim animation.
Rendered with Manim.
Takeaway: An exchange argument again, in a different costume.
Anomaly
Predict first
A student writes this, and it looks reasonable:
A student reasons backwards: 'p shows up so much, it must need a more complicated, longer code to stand out. The rare symbols, q and r, can keep it simple.' They assign p a two-bit code and q a one-bit code.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Even though this reversed code happens to still be prefix-free, giving the busiest symbol the longer code is exactly backwards.
Huffman does the opposite on purpose: the symbol used most often, p, gets the shortest code, and the rarer symbols, q and r, get the longer ones.
Why: Even though this reversed code happens to still be prefix-free, giving the busiest symbol the longer code is exactly backwards.
Trap
A student reasons backwards: 'p shows up so much, it must need a more complicated, longer code to stand out. The rare symbols, q and r, can keep it simple.' They assign p a two-bit code and q a one-bit code.
| Symbol | Frequency | Reversed code |
|---|---|---|
| p | 70 | 00 |
| q | 20 | 1 |
| r | 10 | 01 |
Compute the total bits under this reversed assignment
Why: Even though this reversed code happens to still be prefix-free, giving the busiest symbol the longer code is exactly backwards.
\[ 70(2)+20(1)+10(2) = 140+20+20 = 180 \text{ bits} \]
Huffman does the opposite on purpose: the symbol used most often, p, gets the shortest code, and the rarer symbols, q and r, get the longer ones.
| Symbol | Frequency | Huffman code |
|---|---|---|
| p | 70 | 0 |
| q | 20 | 10 |
| r | 10 | 11 |
Compare the totals directly
Why: The correct assignment costs 130 bits, computed on the previous slide, while the reversed assignment costs 180 bits - fifty bits worse for encoding the exact same source.
\[ 130 \text{ bits (correct)} \; < \; 180 \text{ bits (reversed)} \]
Pattern
Step through it
Step through Trap: frequent symbols should get longer codes one row at a time. What is driving the change, and what would the row after the last one be?
Anomaly
Predict first
A student writes this, and it looks reasonable:
A student assumes a fixed-length code is always a safe, 'good enough' choice, reasoning that every symbol still gets represented either way, so the total cost should not really change.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: This assumption is never checked against real numbers, so the student never notices that fixed-length coding can cost meaningfully more.
Actually compute both totals for our six-symbol alphabet before assuming anything.
Why: This assumption is never checked against real numbers, so the student never notices that fixed-length coding can cost meaningfully more.
Trap
A student assumes a fixed-length code is always a safe, 'good enough' choice, reasoning that every symbol still gets represented either way, so the total cost should not really change.
Skip computing the fixed-length total, assuming it ties with Huffman
Why: This assumption is never checked against real numbers, so the student never notices that fixed-length coding can cost meaningfully more.
Actually compute both totals for our six-symbol alphabet before assuming anything.
| Code | Total bits |
|---|---|
| Fixed-length (3 bits each) | 300 |
| Huffman | 224 |
Compare the two totals honestly
Why: 300 is strictly more than 224. The fixed-length code is not 'just as good' here - it spends 76 more bits than Huffman does to encode the exact same 100 symbols.
\[ 300 > 224 \]
Comparison
Comparison matrix
From Trap: fixed-length is never worse than Huffman: refill the Total bits column from what you know. The rest of the table is as it appeared.
| Code | Total bits |
|---|---|
| Fixed-length (3 bits each) | 300 |
| Huffman | 224 |
Concept
When two nodes happen to tie in frequency, the algorithm can break the tie either way and still produce a valid Huffman tree. Left-right choices at any merge are also arbitrary - swapping which child is 0 and which is 1 changes the codes but not their lengths.
What stays fixed across every valid choice is the total encoded length. Different Huffman trees for the same frequencies can look different, but they all achieve the same minimum total cost.
Prediction
Predict first
What is the total number of bits Huffman uses to encode all 100 symbols?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: 170
Why: Multiply each frequency by its code length and add: 50 times 1 is 50, 30 times 2 is 60, 15 times 3 is 45, and 5 times 3 is 15. Adding 50, 60, 45, and 15 gives 170 total bits.
Check
A four-symbol alphabet has these frequencies and Huffman code lengths, out of 100 symbols total: A is 50 with a 1-bit code, B is 30 with a 2-bit code, C is 15 with a 3-bit code, D is 5 with a 3-bit code.
| Symbol | Frequency | Code length (bits) |
|---|---|---|
| A | 50 | 1 |
| B | 30 | 2 |
| C | 15 | 3 |
| D | 5 | 3 |
Check your understanding
What is the total number of bits Huffman uses to encode all 100 symbols?
Answer: A
Why: Multiply each frequency by its code length and add: 50 times 1 is 50, 30 times 2 is 60, 15 times 3 is 45, and 5 times 3 is 15. Adding 50, 60, 45, and 15 gives 170 total bits.
Pattern
Step through it
Step through Check yourself: which total is smaller? one row at a time. What is driving the change, and what would the row after the last one be?
Section
Section 5
Concept
To encode a message, look up each symbol's code in the table and write the codes down one after another, with no separators needed.
Because the code is prefix-free, the bits will never be mixed up later - you never need spaces, commas, or any other marker between codewords.
Step zero
Discussion prompt
Encode the message a-a-b-a — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Look up each symbol's code in order
Answer:
Worked example
Encode the four-symbol message a, a, b, a using the code table: a is 0, b is 101.
Look up each symbol's code in order
Why: a maps to 0, a maps to 0 again, b maps to 101, and the final a maps to 0.
\[ a{:}0, \quad a{:}0, \quad b{:}101, \quad a{:}0 \]
Concatenate the codes with no separators
Why: Write the four codes back to back in message order to get the full encoded bit string.
\[ 0 \; 0 \; 101 \; 0 \;\Rightarrow\; 001010 \]
Verify the bit count matches the code lengths
Why: The message has four symbols with code lengths 1, 1, 3, and 1 bit, which add up to six bits - exactly the length of 001010.
\[ 1+1+3+1 = 6 \text{ bits} \]
Concept
To decode, start at the root of the Huffman tree. Read one bit at a time: go left on a 0, right on a 1.
The moment you land on a leaf, output that leaf's symbol, then jump back to the root and keep reading the remaining bits the same way.
Picture it
Animation
Shows: Decoding walks the tree — a rendered Manim animation.
Rendered with Manim.
Takeaway: No lookahead, no separators — the tree does all the work.
Intuition
Because every symbol sits at a leaf, and leaves have no further branches, you always know the instant a symbol's code is complete - there is nowhere left to walk.
This is the payoff of prefix-free codes: no lookahead, no guessing, no separators. One pass through the bits, one symbol at a time, always unambiguous.
Estimation
Predict first
Decode the bit string 001010 using the same tree (a is 0, b is 101).
Commit before you compute: what does Decode the bits back to a-a-b-a come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Verify the decoded message matches the original
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Reading 001010 through the tree produced a, a, b, a in order - exactly the message that was encoded, confirming the round trip is lossless.
Worked example
Decode the bit string 001010 using the same tree (a is 0, b is 101).
Read bit 1: a 0
Why: From the root, a single 0 already lands on a's leaf. Output a, then return to the root.
\[ 0 \to a \]
Read bit 2: another 0
Why: Back at the root, the next bit is again 0, landing on a's leaf a second time. Output a.
\[ 0 \to a \]
Read bits 3 through 5: 101
Why: From the root: 1 goes right, 0 goes left, 1 goes right again - that three-bit path lands exactly on b's leaf. Output b.
\[ 101 \to b \]
Read the last bit: 0
Why: Back at the root, the final 0 lands on a's leaf once more. Output a.
\[ 0 \to a \]
Verify the decoded message matches the original
Why: Reading 001010 through the tree produced a, a, b, a in order - exactly the message that was encoded, confirming the round trip is lossless.
Picture it
Animation
Shows: Each line of the worked example "Decode the bits back to a-a-b-a", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Reading 001010 through the tree produced a, a, b, a in order - exactly the message that was encoded, confirming the round trip is lossless.
Anomaly
Predict first
A student writes this, and it looks reasonable:
While decoding 001010, a student reaches b's leaf after reading 101 - but then keeps trying to walk from b's leaf instead of jumping back to the root for the next symbol.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Leaves have no children to walk to, so continuing from a leaf is not just wrong - it is undefined.
Every time a leaf is reached and a symbol is output, decoding must restart from the root before reading the next bit.
Why: Leaves have no children to walk to, so continuing from a leaf is not just wrong - it is undefined. The remaining bit, the final 0, never gets matched to anything.
Trap
While decoding 001010, a student reaches b's leaf after reading 101 - but then keeps trying to walk from b's leaf instead of jumping back to the root for the next symbol.
Continue walking from wherever decoding stopped
Why: Leaves have no children to walk to, so continuing from a leaf is not just wrong - it is undefined. The remaining bit, the final 0, never gets matched to anything.
Every time a leaf is reached and a symbol is output, decoding must restart from the root before reading the next bit.
Return to the root after every symbol
Why: After outputting b, decoding restarts at the root, reads the final 0, and correctly lands on a's leaf, completing the message a, a, b, a.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
Step zero
Discussion prompt
Encode and decode a second message: f-a-c-e — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Concatenate the four codes in order
Answer:
Worked example
Encode f, a, c, e using the full six-symbol code: f is 1100, a is 0, c is 100, e is 1101.
Concatenate the four codes in order
Why: Write each symbol's code back to back: f's code, then a's, then c's, then e's.
\[ 1100 \; 0 \; 100 \; 1101 \;\Rightarrow\; 110001001101 \]
Decode the result by walking the tree from the root
Why: Reading left to right: 1100 lands on f's leaf, 0 lands on a's leaf, 100 lands on c's leaf, and 1101 lands on e's leaf, restarting at the root after each one.
\[ 1100{:}f, \quad 0{:}a, \quad 100{:}c, \quad 1101{:}e \]
Verify the round trip returns the original message
Why: Decoding 110001001101 reproduces f, a, c, e in order, matching the message that was encoded - a full check that encoding and decoding are true inverses of each other.
Picture it
Animation
Shows: Each line of the worked example "Encode and decode a second message: f-a-c-e", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Decoding 110001001101 reproduces f, a, c, e in order, matching the message that was encoded - a full check that encoding and decoding are true inverses of each other.
Prediction
Predict first
What message does the bit string 1011000 decode to?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: b, c, a
Why: Reading from the root: the first three bits, 101, match b's code exactly, restart at the root; the next three, 100, match c's code exactly, restart; the final bit, 0, matches a's code. The bits decode to b, then c, then a.
Check
Using the six-symbol code (a is 0, b is 101, c is 100, d is 111, e is 1101, f is 1100), decode the bit string 1011000.
Check your understanding
What message does the bit string 1011000 decode to?
Answer: A
Why: Reading from the root: the first three bits, 101, match b's code exactly, restart at the root; the next three, 100, match c's code exactly, restart; the final bit, 0, matches a's code. The bits decode to b, then c, then a.
Section
Section 6
Concept
A code is optimal for a given set of frequencies if no other prefix-free code achieves a smaller total encoded length for that same alphabet and those same frequencies.
\[ \text{optimal} \iff L(\text{this code}) \le L(\text{any other prefix-free code}) \]
Huffman's algorithm is not merely a reasonable greedy method - it always produces an optimal code, and that claim can be proven. The rest of this section sketches why.
Picture it
Animation
Shows: How close to optimal it gets — a rendered Manim animation.
Rendered with Manim.
Takeaway: Optimal among prefix codes; arithmetic coding beats it by dropping that restriction.
Intuition
Rather than building a tree and hoping it is optimal, the standard proof strategy assumes an optimal tree already exists, built by any method whatsoever, and studies its shape.
If we can show that some optimal tree always looks the way Huffman's first move would build it, that justifies making that move: it never rules out reaching an optimal answer.
Concept
Call the alphabet's two least-frequent symbols x and y. Assume, for the sake of argument, that some tree T is an optimal prefix-code tree for this alphabet - we do not assume T is the tree Huffman would build.
exchange argument — A proof technique that starts from an assumed optimal solution and shows it can be reshaped, without hurting its quality, to match the structure your algorithm is about to build. If the reshaping never makes things worse, your algorithm's choice is justified.
The goal of this setup: show that T can be rearranged into another optimal tree in which x and y sit as sibling leaves at the deepest level - exactly the pair Huffman would merge first.
Concept
Every proof of this kind has the same five or six moves in the same order. The order is not something you rediscover each time.
It is on the right. It will stay on the right through the worked examples that follow.
Why this matters: the structure is now handled. You are not spending working memory on what comes next — you are spending all of it on the one hard step.
Steps 3 and 5 carry the proof. Step 5 is where a botched exchange hides — showing the swap is legal is not the same as showing it does not lose anything, and you need both.
Concept
A Huffman tree is a full binary tree: every internal node has exactly two children. In any such tree, somewhere at the maximum depth, there is a pair of leaves that share the same parent.
Call those two symbols p and q. They may or may not be x and y - the point of the next two slides is to swap them until they are.
Intuition
What move should we make next?
We have assumed some optimal tree exists, and we know it has two deepest siblings.
\[ \text{optimal tree } \; T^{*}, \qquad \text{deepest sibling pair } \; (x, y) \]
Huffman's very first merge combines the two rarest symbols. The optimal tree might have done something else.
You have run this exact structure twice in the last lesson. Name the move, and say what you are about to swap with what.
_Look at your toolkit. Say a move number out loud before this slide advances._ A wrong guess is useful. A silent guess is not.
Picture it
Animation
Shows: Why no prefix code does better — a rendered Manim animation.
Rendered with Manim.
Takeaway: Two lemmas, then an exchange argument, then induction.
Fill the middle
Fill in the blanks
From Step 1: swap the rarest symbol into the deepest spot — finish the line. Write what belongs on the right of the equals sign before you look.
f_x \le f_p, \qquad d_p = \text{maximum depth in } T
Why: Producing the right-hand side unprompted is the difference between recognising this line and being able to use it. Move x down to depth d_p, and move p up to wherever x used to sit.
Concept
Since x is the single least-frequent symbol in the whole alphabet, its frequency is no greater than p's, and p sits at the tree's maximum depth.
\[ f_x \le f_p, \qquad d_p = \text{maximum depth in } T \]
Swap the leaves holding x and p
Why: Move x down to depth d_p, and move p up to wherever x used to sit. Every other leaf stays exactly where it was.
Compute how the total cost changes
Why: Only the two swapped leaves change depth, so only their two terms in the weighted sum change.
\[ \Delta = f_x d_p + f_p d_x - (f_x d_x + f_p d_p) = (d_p - d_x)(f_x - f_p) \]
Sign-check the change
Why: d_p is the maximum depth, so d_p minus d_x is at least zero; and f_x is at most f_p, so f_x minus f_p is at most zero. A non-negative number times a non-positive number is never positive.
\[ (d_p - d_x) \ge 0, \qquad (f_x - f_p) \le 0 \;\Rightarrow\; \Delta \le 0 \]
Translation
\( (d_p - d_x) \ge 0, \qquad (f_x - f_p) \le 0 \;\Rightarrow\; \Delta \le 0 \)
Draw it
Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.
Intuition
Picture moving a light box to a high, hard-to-reach shelf, and moving a heavier box down to the easy, close shelf it vacated. The extra climbing now falls on the box that weighs the least.
That is exactly the swap: the rarely used symbol takes on the extra depth, while the symbol that was needlessly deep and at least as heavy moves up to a shallower spot. Total effort cannot go up.
Concept
After step 1, x sits at the maximum depth, next to whichever leaf was already its sibling there. Call that sibling q - possibly the same q as before, possibly a new neighbor.
Repeat the same swap argument for y and q
Why: y is the second least-frequent symbol overall, so its frequency is no greater than q's. Swapping y into q's position, by the identical inequality as step 1, cannot increase the total cost.
\[ (d_q - d_y)(f_y - f_q) \le 0 \]
Now both x and y sit at the tree's maximum depth, as siblings sharing one parent - and every swap along the way only ever kept the cost the same or lower.
Concept
The move: #14 (Exchange argument), then #7 (Negate and assume).
Move #7 supplies the rival
Why: Assuming an optimal tree exists is what gives you an object to edit. Without that assumption there is nothing on the table to swap anything into.
Move #14 does the editing
Why: Swap the rarest symbol into the deepest position. Legal, because any leaf can hold any symbol. No worse, because you moved a low frequency to a deep spot and a higher frequency to a shallower one, which cannot increase the weighted total.
\[ (f_{\text{deep}} - f_{\text{rare}})(d_{\text{deep}} - d_{\text{rare}}) \ge 0 \]
Repeat for the second-rarest, then recurse
Why: After both swaps, some optimal tree has the two rarest symbols as deepest siblings — which is exactly what Huffman's first merge assumed. Then induct on the smaller alphabet.
That final induction is moves #5 and #6: merging two symbols into one peels the problem down to an alphabet one smaller, and the hypothesis is that Huffman is optimal there.
Notation
Annotate
From The same move, third subject — read this one piece at a time. What is each part doing?
On: \( (f_{\text{deep}} - f_{\text{rare}})(d_{\text{deep}} - d_{\text{rare}}) \ge 0 \)
Concept
Starting from any optimal tree T, we reshaped it, without raising its cost, into an optimal tree where the two least-frequent symbols, x and y, are sibling leaves at maximum depth.
That is exactly the pair Huffman's algorithm merges first. So the very first greedy choice is always safe: it never closes the door on reaching an optimal solution.
The same argument then applies to the smaller problem left after merging x and y into one combined symbol, and the one after that, all the way down - which is why induction on the number of symbols completes the full optimality proof.
Explain it
Discussion prompt
Explain Conclusion: Huffman's first merge always matches an optimal tree to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Starting from any optimal tree T, we reshaped it, without raising its cost, into an optimal tree where the two least-frequent symbols, x and y, are sibling leaves at maximum depth.
Ranking
Put in order
Put the moves of Verify the exchange lemma on our own tree into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Looking back at the finished tree, f and e sit four branches down from the root - the deepest level anywhere in the tree.
Worked example
Check the lemma against the tree we actually built for a, b, c, d, e, f. The two least-frequent symbols are f (5) and e (9).
Find the maximum depth in our tree
Why: Looking back at the finished tree, f and e sit four branches down from the root - the deepest level anywhere in the tree.
\[ \text{depth}(f) = \text{depth}(e) = 4 = \text{maximum depth} \]
Check that f and e are siblings
Why: Both hang directly off the same internal node, the one we labeled fe:14, so they share a parent, not just a depth.
Verify the lemma holds without needing any swaps
Why: f and e, the two rarest symbols, are already siblings at the tree's maximum depth in the tree Huffman actually built - precisely what the exchange argument guarantees is always achievable.
Picture it
Animation
Shows: Cost of building the tree — a rendered Manim animation.
Rendered with Manim.
Takeaway: The heap is what makes always-take-the-two-rarest cheap.
Intuition
What move should we make next?
The optimality proof is finished. Look back at it as a whole.
It used four moves you already owned, and introduced nothing new.
Name all four, in the order they appear in the proof, and say which line of the proof each one is doing.
_Look at your toolkit. Say a move number out loud before this slide advances._ A wrong guess is useful. A silent guess is not.
Analogy
Discussion prompt
Explain Decision point: account for the whole proof by analogy to something with no CS3000 Algorithms in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
The optimality proof is finished. Look back at it as a whole.
Elimination
Eliminate the wrong options
Why is it safe to swap x into p's position (and p into x's old position)?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: The cost change from this swap equals (depth of p minus depth of x) times (frequency of x minus frequency of p). Since p sits at the maximum depth, the first factor is at least zero; since x has the smallest frequency, the second factor is at most zero. A non-negative number times a non-positive number is never positive, so the swap can only keep the cost the same or lower it.
Check
Recall the setup: T is an assumed-optimal tree, x is the single rarest symbol, and p is a symbol currently sitting at T's maximum depth, with p's frequency at least as large as x's.
Check your understanding
Why is it safe to swap x into p's position (and p into x's old position)?
Answer: A
Why: The cost change from this swap equals (depth of p minus depth of x) times (frequency of x minus frequency of p). Since p sits at the maximum depth, the first factor is at least zero; since x has the smallest frequency, the second factor is at most zero. A non-negative number times a non-positive number is never positive, so the swap can only keep the cost the same or lower it.
Concept
Moves added today: none.
That is a result, not a gap. Everything in this lesson was proved with moves you already owned.
Moves you reused today:
Huffman's optimality proof is move #14 wrapped in move #7, then finished with moves #5 and #6. Every part of it is something you named in an earlier lesson.
Full toolkit so far: #1 through #14.
Next session opens with you naming every one of these from memory, before any new material.
Counterexample
Discussion prompt
Next session opens with you naming every one of these from memory, before any new material.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Why Variable-Length Codes · Prefix-Free Codes · Building the Huffman Tree · Costs & Comparisons · Encoding and Decoding · Why Huffman Is Optimal. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
Huffman coding builds the best possible prefix-free code for a given set of symbol frequencies.
| Idea | The one move |
|---|---|
| Prefix-free | No code is a prefix of another - decode without separators |
| Huffman merge | Always combine the two least-frequent nodes, then reinsert |
| Optimality | Swap the two rarest symbols into the deepest sibling spot without raising cost |
Want this taught 1-on-1? Alexander tutors CS3000 Algorithms — $55/session, free consultation.