Huffman Coding

This deck shows how Huffman coding builds an optimal variable-length, prefix-free code from symbol frequencies. It traces the merge-the-two-smallest algorithm by hand on a six-symbol alphabet, encodes and decodes a message, and sketches the exchange argument for why the greedy strategy is provably optimal. It targets codes that are not prefix-free, merging the wrong nodes or forgetting to reinsert the merged node, the reversal that would give frequent symbols longer codes, and the false belief that fixed-length coding is never worse than Huffman.

Subject: CS3000 Algorithms · 131 slides · symbolic lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Huffman Coding

Title

CS3000 Algorithms · Greedy Algorithms

Building the shortest possible prefix-free code - and proving you can't do better.

2. What you will be able to do

Objectives

Huffman coding is the algorithm behind ZIP, JPEG, and countless other compressors. By the end of this lesson you can:

  1. Explain why variable-length codes beat a fixed-length code when symbol frequencies differ.
  2. Define a prefix-free code and explain why that property makes decoding unambiguous.
  1. Trace the Huffman algorithm on a small alphabet: repeatedly merge the two least-frequent nodes into a tree.
  2. Read a set of codes off a finished Huffman tree, and use them to encode and decode a short message.
  1. Set up the exchange argument: assume an optimal tree exists, and show the two rarest symbols can be made deepest siblings without raising the cost.
  2. Compare a Huffman code's total length against a fixed-length code's, and explain why Huffman is never worse.

3. What survived from Greedy Algorithms & Exchange Arguments?

Warm-up

Discussion prompt

Before we open Huffman Coding: without looking back, what was the main idea of Greedy Algorithms & Exchange Arguments, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

That deck explains what a greedy algorithm is and traces two of them: interval scheduling by earliest finish time, and fractional knapsack by highest value-to-weight ratio. It then gives the two standard ways to prove a greedy algorithm optimal, greedy-stays-ahead and the exchange argument. It targets the beginner's biggest gap, which is getting a proof started, along with the misconceptions that greedy always works, that any scheduling criterion is as good as another, that a few working examples count as a proof, and that an exchange can be made carelessly.

4. Toolkit check-in: name them before you look

Concept

Before any new material: cover the screen.

You have named 14 reusable moves so far. Say as many as you can out loud, by number, from memory.

Do not advance until you have actually tried. Getting four of eight is information; skipping the exercise is not.

Here they are. Score yourself.

Today adds no new moves. Every proof in this lesson is built out of the list above. That is the whole point of the list.

The question that starts every proof from here on is not how do I begin. It is which of these applies here?

5. Break it if you can: Toolkit check-in: name them before you look

Counterexample

Discussion prompt

You have named 14 reusable moves so far. Say as many as you can out loud, by number, from memory.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Do not advance until you have actually tried. Getting four of eight is information; skipping the exercise is not.

6. Why Variable-Length Codes

Section

Section 1

7. The cost of a fixed-length code

Concept

Suppose you must store or send symbols from some alphabet, and you have to pick one fixed number of bits for every symbol, no matter how often it shows up.

With a fixed-length code, every symbol costs the same, so the total number of bits is just the number of symbols times the code length.

\[ \text{total bits} = (\text{number of symbols}) \times (\text{bits per symbol}) \]

fixed-length code — A code where every symbol is represented using the exact same number of bits, regardless of how frequently that symbol occurs.

8. By analogy: The cost of a fixed-length code

Analogy

Discussion prompt

Explain The cost of a fixed-length code by analogy to something with no CS3000 Algorithms in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Suppose you must store or send symbols from some alphabet, and you have to pick one fixed number of bits for every symbol, no matter how often it shows up.

9. Against a fixed-length code

Picture it

Animation

Shows: Against a fixed-length code — a rendered Manim animation.

Rendered with Manim.

Takeaway: The saving comes entirely from the frequency skew.

10. What has to happen first: How many bits does a fixed-length code need for six…

Ranking

Put in order

Put the moves of How many bits does a fixed-length code need for six symbols? into the order they have to happen.

  1. Count how many distinct patterns k bits can make
  2. Find the smallest k with two to the k at least six
  3. Verify three bits is enough and two is not

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Each bit doubles the number of available patterns, so k bits give two to the k possible patterns.

11. How many bits does a fixed-length code need for six symbols?

Worked example

Our running example alphabet has six symbols: a, b, c, d, e, f. A fixed-length code needs enough bits to give each one a distinct pattern.

Count how many distinct patterns k bits can make

Why: Each bit doubles the number of available patterns, so k bits give two to the k possible patterns.

\[ k \text{ bits} \Rightarrow 2^{k} \text{ distinct patterns} \]

Find the smallest k with two to the k at least six

Why: Two bits give only four patterns - not enough for six symbols. Three bits give eight patterns, which covers six with room to spare.

\[ 2^{2}=4 < 6 \le 8 = 2^{3} \]

Verify three bits is enough and two is not

Why: Since 4 is less than 6 and 8 is at least 6, three is the smallest number of bits that can label all six symbols distinctly.

\[ k = 3 \text{ bits per symbol} \]

12. How many bits does a fixed-length code need for six… — line by line

Picture it

Animation

Shows: Each line of the worked example "How many bits does a fixed-length code need for six symbols?", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: Since 4 is less than 6 and 8 is at least 6, three is the smallest number of bits that can label all six symbols distinctly.

13. Why waste bits on a symbol you rarely use?

Intuition

Picture a texting abbreviation system. If you text 'lol' fifty times a day but 'defenestrate' once a year, it makes no sense to give both the same length shortcut.

A fixed-length code treats the everyday word and the once-a-year word identically. Every use of the common symbol pays the same toll as the rare one - and the common symbol is used far more often, so that toll adds up fast.

14. Teach it back: Why waste bits on a symbol you rarely use?

Explain it

Discussion prompt

Explain Why waste bits on a symbol you rarely use? to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Picture a texting abbreviation system. If you text 'lol' fifty times a day but 'defenestrate' once a year, it makes no sense to give both the same length shortcut.

15. The big idea: frequent symbols get shorter codes

Concept

Huffman coding assigns short codes to symbols that occur often, and long codes to symbols that occur rarely. The total cost is dominated by the frequent symbols, so that is exactly where you want to save bits.

variable-length code — A code where different symbols may use different numbers of bits, chosen so that the total encoded length is as small as possible given the symbols' frequencies.

16. Frequent symbols get short codes

Picture it

Animation

Shows: Frequent symbols get short codes — a rendered Manim animation.

Rendered with Manim.

Takeaway: The rarest symbols merge earliest, so they end up deepest and longest.

17. Plan first: Total bits if every symbol used a fixed-length code

Step zero

Discussion prompt

Total bits if every symbol used a fixed-length code — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Use the fixed-length size found earlier

Answer:

  1. Use the fixed-length size found earlier
  2. Multiply by the total symbol count
  3. Verify by checking a single symbol's share

18. Total bits if every symbol used a fixed-length code

Worked example

Our alphabet's frequencies, out of 100 symbols total: a appears 45 times, b 13, c 12, d 16, e 9, f 5.

SymbolFrequency
a45
b13
c12
d16
e9
f5
Total100

Use the fixed-length size found earlier

Why: We showed three bits are needed to distinguish six symbols.

\[ 3 \text{ bits per symbol} \]

Multiply by the total symbol count

Why: Every one of the 100 symbols costs the same three bits under a fixed-length code.

\[ 100 \times 3 = 300 \text{ bits} \]

Verify by checking a single symbol's share

Why: Each of the 100 symbol occurrences contributes exactly 3 bits, and 100 occurrences times 3 bits really is 300 bits - no symbol gets a discount.

\[ 100 \text{ symbols} \times 3 \text{ bits} = 300 \text{ bits} \]

19. Total bits if every symbol used a fixed-length code — line by line

Picture it

Animation

Shows: Each line of the worked example "Total bits if every symbol used a fixed-length code", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: Each of the 100 symbol occurrences contributes exactly 3 bits, and 100 occurrences times 3 bits really is 300 bits - no symbol gets a discount.

20. Prefix-Free Codes

Section

Section 2

21. Prefix-free codes

Concept

If codes can have different lengths, how does a reader know where one symbol's code ends and the next begins? The answer is a design rule on the codes themselves.

prefix-free code — A code in which no symbol's code is the beginning (prefix) of any other symbol's code. Every codeword is self-terminating: once you have matched one, you know it cannot be the start of a longer, different codeword.

22. Reading without a pause button

Intuition

Imagine reading a sentence with no spaces between words. If one word were secretly the start of another word, you would not know when to stop reading one and start the next.

A prefix-free code guarantees this never happens: the moment the bits you have read so far match a complete code, that match is final. No other code shares that same beginning, so there is nothing left to wait for.

23. Guess the shape of the answer: Check whether a small code is prefix-free

Estimation

Predict first

Three symbols, x, y, z, are assigned these codes: x gets 0, y gets 10, z gets 11.

Commit before you compute: what does Check whether a small code is prefix-free come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Verify the whole set is prefix-free

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. No codeword starts with another codeword's full pattern, so this code is prefix-free and any string of these codes can be decoded without ambiguity.

24. Check whether a small code is prefix-free

Worked example

Three symbols, x, y, z, are assigned these codes: x gets 0, y gets 10, z gets 11.

SymbolCode
x0
y10
z11

Compare every pair of codes

Why: A code is prefix-free only if no codeword is a leading chunk of a longer one. Check all three pairs.

Check x's code (0) against y's and z's

Why: y is 10 and z is 11 - neither begins with 0, so x's code is not a prefix of either.

Check y's code (10) against z's code (11)

Why: They share the first bit but differ in the second, and neither is a prefix of the other since both have the same length.

Verify the whole set is prefix-free

Why: No codeword starts with another codeword's full pattern, so this code is prefix-free and any string of these codes can be decoded without ambiguity.

25. Why the code is prefix-free

Picture it

Animation

Shows: Why the code is prefix-free — a rendered Manim animation.

Rendered with Manim.

Takeaway: Which is what makes the stream decodable with no separators.

26. Something is wrong here: a code that is not prefix-free

Anomaly

Predict first

A student writes this, and it looks reasonable:

A student assigns: x gets 0, y gets 1, z gets 01 - reusing 0 and 1 as a 'combo' for the third symbol.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Read left to right: the first bit, 0, already matches x's complete code.

Give every code a pattern that cannot be completed into a different codeword - here, letting z be 11 keeps every code the same length as its siblings at that branch.

Why: Read left to right: the first bit, 0, already matches x's complete code. But the whole string, 01, also matches z's complete code. The same two bits mean two different things.

27. Trap: a code that is not prefix-free

Trap

The trap

A student assigns: x gets 0, y gets 1, z gets 01 - reusing 0 and 1 as a 'combo' for the third symbol.

SymbolCode
x0
y1
z01

Try to decode the bit string 01

Why: Read left to right: the first bit, 0, already matches x's complete code. But the whole string, 01, also matches z's complete code. The same two bits mean two different things.

The fix

Give every code a pattern that cannot be completed into a different codeword - here, letting z be 11 keeps every code the same length as its siblings at that branch.

SymbolCode
x0
y10
z11

Decode 011 with this fixed code

Why: Read 0 - matches x completely, output x, then continue with 11 - matches z completely. The result, x then z, is the only possible reading.

28. Watch it run: Trap: a code that is not prefix-free

Pattern

Step through it

Step through Trap: a code that is not prefix-free one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: Symbol is x
  2. Step 2: Symbol is y
  3. Step 3: Symbol is z

29. Codes live on a tree: symbols are leaves

Concept

Every prefix-free code can be pictured as a binary tree: start at the root, and each bit of a code tells you whether to go to the left child (0) or the right child (1).

Symbols sit only at the leaves of the tree - the nodes with no children. Internal nodes are just junctions on the way down; they never carry a symbol themselves.

30. Leaves have no children, so no code can swallow another

Intuition

If every symbol lives at a leaf, then reaching one symbol's leaf means the path stops there - there is no branch left to continue down.

That is exactly why a tree automatically gives you a prefix-free code: a leaf's path can never continue on into a deeper leaf's path, since a leaf has nowhere further to go.

31. Picture it first: Build the tree for the tiny prefix-free code

Picture it

Figure (svg): A small binary tree with x as the left child of the root, and y and z as the two children of the root's right child

Discussion prompt

Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.

Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.

Answer:

Use the fixed x, y, z code from before: x is 0, y is 10, z is 11. Draw the tree that produces it.

32. Build the tree for the tiny prefix-free code

Worked example

Use the fixed x, y, z code from before: x is 0, y is 10, z is 11. Draw the tree that produces it.

Place x at depth one on the left

Why: x's code is a single 0, meaning one left branch from the root lands directly on x's leaf.

Branch right from the root, then split again

Why: Both y and z start with 1, so they share the root's right branch; a second branch (0 for y, 1 for z) then separates them one level deeper.

Figure (svg): A small binary tree with x as the left child of the root, and y and z as the two children of the root's right child

Verify the leaf paths match the codes

Why: Root to x is just left (0); root to y is right-then-left (10); root to z is right-then-right (11). These match exactly the codes we started with, confirming the tree and the code agree.

33. Work backwards from the answer: Build the tree for the tiny prefix-free code

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Verify the leaf paths match the codes

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Use the fixed x, y, z code from before: x is 0, y is 10, z is 11. Draw the tree that produces it.

34. Building the Huffman Tree

Section

Section 3

35. Our alphabet and its frequencies

Concept

From here on, every example uses the same six-symbol alphabet, with frequencies drawn from a classic textbook example, out of 100 total occurrences.

SymbolFrequency
a45
b13
c12
d16
e9
f5
Total100

Notice how spread out the frequencies are: a is nearly half of everything, while f shows up only five times. That spread is exactly what makes variable-length coding worth doing.

36. Watch it run: Our alphabet and its frequencies

Pattern

Step through it

Step through Our alphabet and its frequencies one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: Symbol is a
  2. Step 2: Symbol is b
  3. Step 3: Symbol is c
  4. Step 4: Symbol is d
  5. Step 5: Symbol is e
  6. Step 6: Symbol is f
  7. Step 7: Symbol is Total

37. The table has to travel with the data

Picture it

Animation

Shows: The table has to travel with the data — a rendered Manim animation.

Rendered with Manim.

Takeaway: For short messages that overhead can exceed the saving.

38. Every merge combines two nodes into one

Concept

Huffman's algorithm starts with one node per symbol and repeatedly fuses two nodes into a single new node, until only one node - the root - remains.

Each fusion reduces the node count by exactly one. Starting from six symbol-nodes and ending at a single root means exactly five fusions happen along the way.

\[ 6 \text{ nodes} \to 1 \text{ root} \Rightarrow 6 - 1 = 5 \text{ merges} \]

39. What feels wrong about merging the two lightest?

Intuition

What feels wrong about this?

The rule says: repeatedly combine the two least frequent symbols.

But merging pushes both of those symbols one level deeper in the tree, which makes their codes longer.

_Plain English only. No notation, no algebra. Just say what bothers you._

The feeling: it looks backwards. You are deliberately penalizing symbols, over and over, right at the start.

That feeling is the proof. It is not a substitute for the proof — it is the thing the proof writes down.

The resolution is that depth is a resource somebody has to spend. Every merge costs depth to exactly two symbols, so you spend it on the two that will be transmitted least often. The rarest symbols are the cheapest place to put the cost.

40. Where does each piece belong: Huffman Coding

Sorting

Sort into buckets

These are the pieces of Huffman Coding, out of order. Put each one back under the part of the lesson it belongs to.

Why Variable-Length Codes
The cost of a fixed-length code; How many bits does a fixed-length code need for six symbols?; Why waste bits on a symbol you rarely use?
Prefix-Free Codes
Prefix-free codes; Reading without a pause button; Check whether a small code is prefix-free
Building the Huffman Tree
Our alphabet and its frequencies; Every merge combines two nodes into one; What feels wrong about merging the two lightest?
s1
Why Variable-Length Codes is where Huffman Coding puts The cost of a fixed-length code, How many bits does a fixed-length code need for six symbols?, Why waste bits on a symbol you rarely use?. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
Prefix-Free Codes is where Huffman Coding puts Prefix-free codes, Reading without a pause button, Check whether a small code is prefix-free. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
Building the Huffman Tree is where Huffman Coding puts Our alphabet and its frequencies, Every merge combines two nodes into one, What feels wrong about merging the two lightest?. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

41. The Huffman rule: always merge the two lightest

Concept

At every step, look at all the current nodes - original symbols and any nodes already merged - and combine the two with the smallest frequency.

The new node's frequency is just the sum of the two it replaced, and it goes back into the pool of nodes still waiting to be merged.

\[ \text{new node's frequency} = f_1 + f_2 \]

42. Huffman's algorithm in pseudocode

Concept

A priority queue and one loop. Every pass removes the two lightest nodes and puts back a single node weighing what they weighed together.

HUFFMAN(symbols, freq)
  Q = priority queue of leaf nodes, keyed by freq
  while Q.size > 1
    a = Extract-Min(Q)
    b = Extract-Min(Q)
    z = new node
    z.left = a
    z.right = b
    z.freq = a.freq + b.freq
    Insert(Q, z)
  return Extract-Min(Q)

The queue shrinks by exactly one each pass: two out, one in. Starting from k symbols that means k minus one merges, and the last node standing is the root of the finished tree.

43. Reading HUFFMAN line by line

Notation

Every line of HUFFMAN says one thing. Read the line, then read what it does — not the other way round.

Annotate

  • Every symbol starts as its own leaf. Nothing is a parent yet, so the tree is built strictly bottom-up.
  • Stop when one node remains. That survivor is the root, and it always weighs the total of every frequency.
  • The two lightest, always. This is the greedy choice, and the whole correctness proof is about justifying exactly this line.
  • The new node's weight is the sum. Rare symbols merge early and therefore sit deepest, which is why they get the longest codes.
  • Put the merged node back in as an ordinary competitor. It can be merged again immediately if it is still among the lightest.
  • Two out and one in each pass, so k symbols cost k minus one merges — and each queue operation is log k.

44. Step HUFFMAN yourself

Invariant

The total weight in the queue never changes; only the number of nodes falls. Depth in the finished tree is exactly how many merges a symbol survived.

Step through it

Before each merge, name the two lightest yourself.

  1. Line 2: four leaves, lightest first
  2. Line 4: extract the two lightest: 5 and 9
  3. Line 9: put back a node weighing 14
  4. Line 4: now 12 and 13 are lightest
  5. Line 9: put back 25
  6. Line 4: only two left
  7. Line 9: the root weighs the whole alphabet
  8. Line 11: 5 merged three times, so its code is three bits

45. The two lightest merge, over and over

Picture it

Animation

Shows: HUFFMAN executing: the current line of pseudocode is highlighted while the data it touches changes.

Rendered with Manim.

Takeaway: Repeatedly merge the two lightest nodes — rare symbols merge earliest, sit deepest, and so earn the longest codes.

46. Combine the two lightest weights first, like a scale

Intuition

Picture weights on a table, each labeled with a symbol. You always pick up the two lightest weights, tape them together into one heavier weight, and set that combined weight back down among the others.

Because you always start with the lightest pair, the lightest original weights end up wrapped in the most layers of tape by the end - which is exactly why rare symbols end up deepest in the tree, with the longest codes.

47. Process: what if we merged the two heaviest instead?

Intuition

Watch me not know the answer. This is what the first two minutes actually look like.

Frequencies for five symbols: 45, 13, 12, 16, 9.

Try merging the two heaviest first: 45 and 16

Why: It is the symmetric rule and it takes the same amount of code to write, so it is worth ten seconds to see what it does.

It buries the most common symbol

Why: That merge sends the 45-weight symbol one level down immediately, and every later merge involving that combined node pushes it deeper still. The most frequent symbol ends up with one of the longest codes.

Dead end. Not a mistake — a move that was worth trying and did not pay off. This happens in most proofs.

Back up. Compute the cost of both and compare

Why: Do not argue about it — the weighted total is a number. Build both trees and add up frequency times depth.

Merging heaviest-first produces a strictly larger total. The symmetric rule is not merely unproven, it is wrong — and finding that out cost one arithmetic check rather than an argument.

Trying the mirror image of a rule is a cheap and underused move. It either produces a counterexample or it sharpens your sense of why the real rule works.

The expert does not see the whole path in advance. The expert tries something, reads the result, and adjusts. That is the skill.

48. Plan first: Trace the Huffman merges

Step zero

Discussion prompt

Trace the Huffman merges — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Merge 1: combine the two smallest, f (5) and e (9)

Answer:

  1. Merge 1: combine the two smallest, f (5) and e (9)
  2. Merge 2: combine the two smallest now, c (12) and b (13)
  3. Merge 3: combine fe (14) and d (16)
  4. Merge 4: combine cb (25) and fed (30)
  5. Merge 5: combine a (45) and cbfed (55)
  6. Verify the merge count and the total

49. Trace the Huffman merges

Worked example

Starting pool: f is 5, e is 9, c is 12, b is 13, d is 16, a is 45.

Merge 1: combine the two smallest, f (5) and e (9)

Why: f and e are the lowest frequencies in the pool. Their combined node has frequency 5 plus 9.

\[ f(5) + e(9) = 14 \;\Rightarrow\; \text{new node } fe:14 \]

Merge 2: combine the two smallest now, c (12) and b (13)

Why: With fe:14 back in the pool, the current smallest two are c and b - fe (14) is now larger than both.

\[ c(12) + b(13) = 25 \;\Rightarrow\; \text{new node } cb:25 \]

Merge 3: combine fe (14) and d (16)

Why: The remaining pool is fe:14, d:16, cb:25, a:45. The smallest two are fe and d.

\[ fe(14) + d(16) = 30 \;\Rightarrow\; \text{new node } fed:30 \]

Merge 4: combine cb (25) and fed (30)

Why: The remaining pool is cb:25, fed:30, a:45. The smallest two are cb and fed.

\[ cb(25) + fed(30) = 55 \;\Rightarrow\; \text{new node } cbfed:55 \]

Merge 5: combine a (45) and cbfed (55)

Why: Only two nodes remain, a and cbfed, so they merge into the root.

\[ a(45) + cbfed(55) = 100 \;\Rightarrow\; \text{root}: 100 \]

Verify the merge count and the total

Why: Six symbols required exactly five merges, matching the count predicted earlier, and the root's frequency is 100 - the full original total, with nothing lost or double-counted.

\[ 5 \text{ merges}, \quad \text{root frequency} = 100 \]

50. Huffman merges the two rarest symbols

Picture it

Animation

Shows: Huffman merges the two rarest symbols — a rendered Manim animation.

Rendered with Manim.

Takeaway: Built upward from the rarest pair, so the rarest symbols end up deepest.

51. Guess the shape of the answer: Read the codes off the finished tree

Estimation

Predict first

Label every left branch 0 and every right branch 1. A symbol's code is the sequence of labels on the path from the root down to its leaf.

Commit before you compute: what does Read the codes off the finished tree come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Verify the codes are prefix-free

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Every code sits at a leaf of the same tree, and leaves have no children, so no code's path can continue into another leaf's path.

52. Read the codes off the finished tree

Worked example

Label every left branch 0 and every right branch 1. A symbol's code is the sequence of labels on the path from the root down to its leaf.

Trace a's path

Why: a is a direct child of the root - the left child, since a was the smaller of the last two nodes merged.

\[ a: 0 \]

Trace c and b's paths

Why: Both sit under the cb node, reached by going right from the root (into cbfed:55) then left (into cb:25). c is cb's left child and b is cb's right child.

\[ c: 100 \qquad b: 101 \]

Trace d, then f and e's paths

Why: The fed node is reached by going right then right again. d is fed's right child. fe is fed's left child; within fe, f is the left child and e is the right child.

\[ d: 111 \qquad f: 1100 \qquad e: 1101 \]

SymbolFrequencyCodeLength (bits)
a4501
b131013
c121003
d161113
e911014
f511004

Verify the codes are prefix-free

Why: Every code sits at a leaf of the same tree, and leaves have no children, so no code's path can continue into another leaf's path. For example, a's code (0) never appears as the first digit of b, c, or d's codes, all of which start with 1.

53. Picture the finished tree

Concept

Here is the whole tree at once, with each internal node's combined frequency and every 0/1 branch label shown.

Figure (svg): The full Huffman tree for symbols a, b, c, d, e, f with frequencies 45, 13, 12, 16, 9, 5, showing merge frequencies at each internal node and 0/1 edge labels down to each leaf

Notice the shape: a, the most frequent symbol, sits closest to the root with the shortest code. f and e, the two rarest symbols, sit at the bottom of the tree, sharing the longest codes.

54. Something is wrong here: merging the wrong nodes

Anomaly

Predict first

A student writes this, and it looks reasonable:

A student thinks 'Huffman merges nodes' and merges the two largest frequencies first instead of the two smallest - combining a (45) and d (16) right away.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: This buries the most frequent symbol, a, one level deeper than necessary right from the start, which can only raise its code length and its contribution to the total cost.

Always identify the two smallest frequencies currently in the pool, no matter whether they are original symbols or already-merged nodes.

Why: This buries the most frequent symbol, a, one level deeper than necessary right from the start, which can only raise its code length and its contribution to the total cost.

55. Trap: merging the wrong nodes

Trap

The trap

A student thinks 'Huffman merges nodes' and merges the two largest frequencies first instead of the two smallest - combining a (45) and d (16) right away.

\[ a(45) + d(16) = 61 \quad (\text{wrong first move}) \]

Push the two heaviest symbols together early

Why: This buries the most frequent symbol, a, one level deeper than necessary right from the start, which can only raise its code length and its contribution to the total cost.

A second version of the same mistake: correctly merging f (5) and e (9) into fe:14, but then forgetting to put fe:14 back into the pool - leaving it stranded and out of consideration.

Continue merging only the leftover originals c, b, d, a and never revisit fe

Why: With fe left out, the algorithm can no longer build one connected tree over all six symbols - some symbols end up unreachable from a single root.

The fix

Always identify the two smallest frequencies currently in the pool, no matter whether they are original symbols or already-merged nodes.

\[ f(5) + e(9) = 14 \quad (\text{correct first move}) \]

Merge the two smallest, then reinsert the result

Why: The new node's frequency (14) goes right back into the pool, competing on equal footing with every remaining node in the very next round.

Keep repeating until one node is left

Why: Each round always looks at the full current pool - including every node created so far - so nothing is ever left stranded outside the growing tree.

56. Decode the notation: Trap: merging the wrong nodes

Notation

Annotate

From Trap: merging the wrong nodes — read this one piece at a time. What is each part doing?

On: \( a(45) + d(16) = 61 \quad (\text{wrong first move}) \)

  • This buries the most frequent symbol, a, one level deeper than necessary right from the start, which can only raise its code length and its contribution to the total cost.
  • With fe left out, the algorithm can no longer build one connected tree over all six symbols - some symbols end up unreachable from a single root.
  • The new node's frequency (14) goes right back into the pool, competing on equal footing with every remaining node in the very next round.

57. The Huffman algorithm, step by step

Pattern

1. Put every symbol in a pool, each as its own node labeled with its frequency

Why: This is the starting point: one leaf-to-be for every symbol, nothing merged yet.

2. While more than one node remains, find the two nodes with the smallest frequency

Why: Always the two smallest in the current pool - which may include nodes created by earlier merges, not just original symbols.

3. Merge them into a new node whose frequency is their sum, and put that new node back into the pool

Why: Reinserting is essential - skip it and the algorithm cannot connect every symbol into one tree.

4. Repeat until a single node remains: that node is the root

Why: For n symbols this takes exactly n minus 1 merges, since each merge shrinks the pool by one.

5. Read off each symbol's code as the 0/1 path from the root down to its leaf

Why: Left branches contribute a 0, right branches a 1; the path length is the code's length in bits.

58. Where this shows up: Huffman Coding

Real world

Discussion prompt

Outside this lesson: where does Huffman Coding actually turn up? Name one concrete situation — a job, a piece of software someone ships, a decision somebody has to make — and say which part of The Huffman algorithm, step by step is doing the work in it.

Hint: Vague is the failure mode here. "Engineering" is not a situation; "deciding whether this build is fast enough to ship" is.

Answer:

That deck shows how Huffman coding builds an optimal variable-length, prefix-free code from symbol frequencies. It traces the merge-the-two-smallest algorithm by hand on a six-symbol alphabet, encodes and decodes a message, and sketches the exchange argument for why the greedy strategy is provably optimal. It targets codes that are not prefix-free, merging the wrong nodes or forgetting to reinsert the merged node, the reversal that would give frequent symbols longer codes, and the false belief that fixed-length coding is never worse than Huffman.

59. Ties give different trees, same cost

Picture it

Animation

Shows: Ties give different trees, same cost — a rendered Manim animation.

Rendered with Manim.

Takeaway: Which is why two correct answers can look nothing alike.

60. Rule out three: Check yourself: which two nodes merge first?

Elimination

Eliminate the wrong options

Which two nodes does the Huffman algorithm merge first?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. x and z
  • B. w and y
  • C. x and w
  • D. y and z

Survives elimination: A

Why: Huffman always merges the two nodes with the smallest frequency in the current pool. Here, x (3) and z (4) are the two lowest frequencies, so they merge first into a node of frequency 7.

61. Check yourself: which two nodes merge first?

Check

A new alphabet has four symbols with these frequencies: w is 7, x is 3, y is 9, z is 4.

SymbolFrequency
w7
x3
y9
z4

Check your understanding

Which two nodes does the Huffman algorithm merge first?

  • A. x and z (correct)
  • B. w and y
  • C. x and w
  • D. y and z

Answer: A

Why: Huffman always merges the two nodes with the smallest frequency in the current pool. Here, x (3) and z (4) are the two lowest frequencies, so they merge first into a node of frequency 7.

Why B tempts people
This picks w and y, the two largest frequencies (7 and 9) - the opposite of the rule. Huffman merges the smallest, not the largest.
Why C tempts people
This correctly picks the smallest node, x, but pairs it with the largest, w, instead of with the second-smallest node, z.
Why D tempts people
This pairs the largest frequency, y, with the second-smallest, z, ignoring that x (3) is smaller than both and must be part of the first merge.

62. Watch it run: Check yourself: which two nodes merge first?

Pattern

Step through it

Step through Check yourself: which two nodes merge first? one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: Symbol is w
  2. Step 2: Symbol is x
  3. Step 3: Symbol is y
  4. Step 4: Symbol is z

63. Costs & Comparisons

Section

Section 4

64. Total encoded length is a weighted sum

Concept

Once every symbol has a code, the total number of bits used to encode a source is found by weighting each code's length by how often that symbol occurs.

\[ L = \sum_{i} f_i \cdot \ell_i \]

In this formula, each symbol's frequency is multiplied by its own code's length in bits, and every one of those products is added together to get the grand total, L.

65. What has to happen first: Compute the Huffman total for our alphabet

Ranking

Put in order

Put the moves of Compute the Huffman total for our alphabet into the order they have to happen.

  1. Multiply each frequency by its code length
  2. Add every product together
  3. Verify the sum matches the weighted-sum formula

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. This gives how many bits that symbol contributes across all of its occurrences, shown in the table's last column.

66. Compute the Huffman total for our alphabet

Worked example

Use the code table built earlier: a is 1 bit, b is 3, c is 3, d is 3, e is 4, f is 4.

SymbolFrequencyLength (bits)Frequency times length
a45145
b13339
c12336
d16348
e9436
f5420

Multiply each frequency by its code length

Why: This gives how many bits that symbol contributes across all of its occurrences, shown in the table's last column.

Add every product together

Why: Summing the last column gives the total number of bits the whole 100-symbol source needs under this code.

\[ 45+39+36+48+36+20 = 224 \]

Verify the sum matches the weighted-sum formula

Why: 224 bits is exactly the value of L from the formula, computed directly from the table - no symbol was skipped or double-counted.

\[ L = 224 \text{ bits} \]

67. Compute the Huffman total for our alphabet — line by line

Picture it

Animation

Shows: Each line of the worked example "Compute the Huffman total for our alphabet", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: 224 bits is exactly the value of L from the formula, computed directly from the table - no symbol was skipped or double-counted.

68. Average bits per symbol

Concept

Dividing the total bits by the number of symbols gives a single number you can compare across different codes: the average bits spent per symbol.

\[ \text{average} = \frac{L}{\text{number of symbols}} = \frac{224}{100} = 2.24 \text{ bits per symbol} \]

Compare that to a fixed-length code, which always spends exactly three bits per symbol, no matter which symbol it is. Huffman spends noticeably less on average - and that difference adds up across every symbol in a long message.

69. Why weighted means busy symbols count more

Intuition

Think of it like a grocery bill: a two-dollar item you buy fifty times costs you far more overall than a twenty-dollar item you buy once. The frequency, not just the price tag, decides your total spending.

The same logic drives Huffman's total: a symbol's contribution depends on both its code length and how often it shows up. That is why the algorithm fights hardest to shorten the codes of the busiest symbols.

70. A fixed-length code needs room for every symbol

Concept

A fixed-length code has to give every symbol in the alphabet its own distinct pattern, so its bit budget depends only on how many symbols exist - never on their frequencies.

\[ \text{bits per symbol} = \lceil \log_{2}(\text{number of symbols}) \rceil \]

This is why a fixed-length code cannot adapt: whether one symbol dominates the source or all symbols are equally common, the per-symbol cost never changes.

71. State the rule before it runs: Compute the fixed-length total and compare

Hypothesis

Predict first

Compute the fixed-length total and compare is about to be worked. State your hypothesis first: which rule or definition decides this one, and what is the first move it forces? Then watch whether the example agrees with you.

Correct: Multiply the fixed bit cost by the total symbol count

Why: All 100 symbol occurrences pay the same fixed rate of three bits each.

A hypothesis you wrote down is falsifiable; a vague sense of how it will go is not. If the example opens somewhere else, that gap is the thing worth chasing.

72. Compute the fixed-length total and compare

Worked example

Our alphabet has six symbols, so, as shown earlier, a fixed-length code needs three bits per symbol.

Multiply the fixed bit cost by the total symbol count

Why: All 100 symbol occurrences pay the same fixed rate of three bits each.

\[ 100 \times 3 = 300 \text{ bits} \]

Compare the two totals

Why: Huffman's total was 224 bits; the fixed-length total is 300 bits. Subtracting gives the number of bits saved.

\[ 300 - 224 = 76 \text{ bits saved} \]

Verify the savings as a percentage

Why: 76 bits saved out of an original 300 is about one quarter of the total - Huffman uses roughly 25 percent fewer bits than the fixed-length code on this alphabet.

\[ \frac{76}{300} \approx 0.253 \;(\approx 25\%) \]

73. A second alphabet: Huffman on p, q, r

Worked example

A smaller alphabet, used only in the next two slides: p occurs 70 times, q occurs 20 times, r occurs 10 times, out of 100 total.

SymbolFrequency
p70
q20
r10

Merge the two smallest, q (20) and r (10)

Why: q and r are the lowest frequencies in this three-symbol pool.

\[ q(20)+r(10)=30 \;\Rightarrow\; qr:30 \]

Merge the remaining two nodes, p (70) and qr (30)

Why: Only p and the merged qr node remain, so they combine into the root.

\[ p(70)+qr(30)=100 \;\Rightarrow\; \text{root}: 100 \]

Read off the codes

Why: p is a direct child of the root, so its code is one bit. q and r sit one level deeper, under the qr node, so their codes are two bits each.

\[ p: 0 \; (1\text{ bit}) \qquad q: 10 \; (2\text{ bits}) \qquad r: 11 \; (2\text{ bits}) \]

Verify the total against the weighted-sum formula

Why: 70 times 1, plus 20 times 2, plus 10 times 2, gives the total bits Huffman needs for this alphabet.

\[ 70(1)+20(2)+10(2) = 70+40+20 = 130 \text{ bits} \]

74. The greedy choice, and why it is safe

Picture it

Animation

Shows: The greedy choice, and why it is safe — a rendered Manim animation.

Rendered with Manim.

Takeaway: An exchange argument again, in a different costume.

75. Something is wrong here: frequent symbols should get longer codes

Anomaly

Predict first

A student writes this, and it looks reasonable:

A student reasons backwards: 'p shows up so much, it must need a more complicated, longer code to stand out. The rare symbols, q and r, can keep it simple.' They assign p a two-bit code and q a one-bit code.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Even though this reversed code happens to still be prefix-free, giving the busiest symbol the longer code is exactly backwards.

Huffman does the opposite on purpose: the symbol used most often, p, gets the shortest code, and the rarer symbols, q and r, get the longer ones.

Why: Even though this reversed code happens to still be prefix-free, giving the busiest symbol the longer code is exactly backwards.

76. Trap: frequent symbols should get longer codes

Trap

The trap

A student reasons backwards: 'p shows up so much, it must need a more complicated, longer code to stand out. The rare symbols, q and r, can keep it simple.' They assign p a two-bit code and q a one-bit code.

SymbolFrequencyReversed code
p7000
q201
r1001

Compute the total bits under this reversed assignment

Why: Even though this reversed code happens to still be prefix-free, giving the busiest symbol the longer code is exactly backwards.

\[ 70(2)+20(1)+10(2) = 140+20+20 = 180 \text{ bits} \]

The fix

Huffman does the opposite on purpose: the symbol used most often, p, gets the shortest code, and the rarer symbols, q and r, get the longer ones.

SymbolFrequencyHuffman code
p700
q2010
r1011

Compare the totals directly

Why: The correct assignment costs 130 bits, computed on the previous slide, while the reversed assignment costs 180 bits - fifty bits worse for encoding the exact same source.

\[ 130 \text{ bits (correct)} \; < \; 180 \text{ bits (reversed)} \]

77. Watch it run: Trap: frequent symbols should get longer codes

Pattern

Step through it

Step through Trap: frequent symbols should get longer codes one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: Symbol is p
  2. Step 2: Symbol is q
  3. Step 3: Symbol is r

78. Something is wrong here: fixed-length is never worse than Huffman

Anomaly

Predict first

A student writes this, and it looks reasonable:

A student assumes a fixed-length code is always a safe, 'good enough' choice, reasoning that every symbol still gets represented either way, so the total cost should not really change.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: This assumption is never checked against real numbers, so the student never notices that fixed-length coding can cost meaningfully more.

Actually compute both totals for our six-symbol alphabet before assuming anything.

Why: This assumption is never checked against real numbers, so the student never notices that fixed-length coding can cost meaningfully more.

79. Trap: fixed-length is never worse than Huffman

Trap

The trap

A student assumes a fixed-length code is always a safe, 'good enough' choice, reasoning that every symbol still gets represented either way, so the total cost should not really change.

Skip computing the fixed-length total, assuming it ties with Huffman

Why: This assumption is never checked against real numbers, so the student never notices that fixed-length coding can cost meaningfully more.

The fix

Actually compute both totals for our six-symbol alphabet before assuming anything.

CodeTotal bits
Fixed-length (3 bits each)300
Huffman224

Compare the two totals honestly

Why: 300 is strictly more than 224. The fixed-length code is not 'just as good' here - it spends 76 more bits than Huffman does to encode the exact same 100 symbols.

\[ 300 > 224 \]

80. Fill in: Total bits for Trap: fixed-length is never worse than…

Comparison

Comparison matrix

From Trap: fixed-length is never worse than Huffman: refill the Total bits column from what you know. The rest of the table is as it appeared.

CodeTotal bits
Fixed-length (3 bits each)300
Huffman224

81. Huffman codes aren't unique, but their cost is

Concept

When two nodes happen to tie in frequency, the algorithm can break the tie either way and still produce a valid Huffman tree. Left-right choices at any merge are also arbitrary - swapping which child is 0 and which is 1 changes the codes but not their lengths.

What stays fixed across every valid choice is the total encoded length. Different Huffman trees for the same frequencies can look different, but they all achieve the same minimum total cost.

82. Answer it before you see the options: Check yourself: which total is smaller?

Prediction

Predict first

What is the total number of bits Huffman uses to encode all 100 symbols?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: 170

Why: Multiply each frequency by its code length and add: 50 times 1 is 50, 30 times 2 is 60, 15 times 3 is 45, and 5 times 3 is 15. Adding 50, 60, 45, and 15 gives 170 total bits.

83. Check yourself: which total is smaller?

Check

A four-symbol alphabet has these frequencies and Huffman code lengths, out of 100 symbols total: A is 50 with a 1-bit code, B is 30 with a 2-bit code, C is 15 with a 3-bit code, D is 5 with a 3-bit code.

SymbolFrequencyCode length (bits)
A501
B302
C153
D53

Check your understanding

What is the total number of bits Huffman uses to encode all 100 symbols?

  • A. 170 (correct)
  • B. 200
  • C. 150
  • D. 165

Answer: A

Why: Multiply each frequency by its code length and add: 50 times 1 is 50, 30 times 2 is 60, 15 times 3 is 45, and 5 times 3 is 15. Adding 50, 60, 45, and 15 gives 170 total bits.

Why B tempts people
200 is the fixed-length total (2 bits times 100 symbols, since 4 symbols need 2 bits each) - a different code entirely, not the Huffman total.
Why C tempts people
150 comes from mistakenly using a 2-bit length for C and D instead of their actual 3-bit codes (50 plus 60 plus 15 times 2 plus 5 times 2 equals 150).
Why D tempts people
165 comes from an arithmetic slip on the C term, computing 15 times 3 as 40 instead of 45, giving 50 plus 60 plus 40 plus 15.

84. Watch it run: Check yourself: which total is smaller?

Pattern

Step through it

Step through Check yourself: which total is smaller? one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: Symbol is A
  2. Step 2: Symbol is B
  3. Step 3: Symbol is C
  4. Step 4: Symbol is D

85. Encoding and Decoding

Section

Section 5

86. Encoding: concatenate the codes

Concept

To encode a message, look up each symbol's code in the table and write the codes down one after another, with no separators needed.

Because the code is prefix-free, the bits will never be mixed up later - you never need spaces, commas, or any other marker between codewords.

87. Plan first: Encode the message a-a-b-a

Step zero

Discussion prompt

Encode the message a-a-b-a — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Look up each symbol's code in order

Answer:

  1. Look up each symbol's code in order
  2. Concatenate the codes with no separators
  3. Verify the bit count matches the code lengths

88. Encode the message a-a-b-a

Worked example

Encode the four-symbol message a, a, b, a using the code table: a is 0, b is 101.

Look up each symbol's code in order

Why: a maps to 0, a maps to 0 again, b maps to 101, and the final a maps to 0.

\[ a{:}0, \quad a{:}0, \quad b{:}101, \quad a{:}0 \]

Concatenate the codes with no separators

Why: Write the four codes back to back in message order to get the full encoded bit string.

\[ 0 \; 0 \; 101 \; 0 \;\Rightarrow\; 001010 \]

Verify the bit count matches the code lengths

Why: The message has four symbols with code lengths 1, 1, 3, and 1 bit, which add up to six bits - exactly the length of 001010.

\[ 1+1+3+1 = 6 \text{ bits} \]

89. Decoding: walk the tree one bit at a time

Concept

To decode, start at the root of the Huffman tree. Read one bit at a time: go left on a 0, right on a 1.

The moment you land on a leaf, output that leaf's symbol, then jump back to the root and keep reading the remaining bits the same way.

90. Decoding walks the tree

Picture it

Animation

Shows: Decoding walks the tree — a rendered Manim animation.

Rendered with Manim.

Takeaway: No lookahead, no separators — the tree does all the work.

91. The tree tells you exactly when a symbol ends

Intuition

Because every symbol sits at a leaf, and leaves have no further branches, you always know the instant a symbol's code is complete - there is nowhere left to walk.

This is the payoff of prefix-free codes: no lookahead, no guessing, no separators. One pass through the bits, one symbol at a time, always unambiguous.

92. Guess the shape of the answer: Decode the bits back to a-a-b-a

Estimation

Predict first

Decode the bit string 001010 using the same tree (a is 0, b is 101).

Commit before you compute: what does Decode the bits back to a-a-b-a come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Verify the decoded message matches the original

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Reading 001010 through the tree produced a, a, b, a in order - exactly the message that was encoded, confirming the round trip is lossless.

93. Decode the bits back to a-a-b-a

Worked example

Decode the bit string 001010 using the same tree (a is 0, b is 101).

Read bit 1: a 0

Why: From the root, a single 0 already lands on a's leaf. Output a, then return to the root.

\[ 0 \to a \]

Read bit 2: another 0

Why: Back at the root, the next bit is again 0, landing on a's leaf a second time. Output a.

\[ 0 \to a \]

Read bits 3 through 5: 101

Why: From the root: 1 goes right, 0 goes left, 1 goes right again - that three-bit path lands exactly on b's leaf. Output b.

\[ 101 \to b \]

Read the last bit: 0

Why: Back at the root, the final 0 lands on a's leaf once more. Output a.

\[ 0 \to a \]

Verify the decoded message matches the original

Why: Reading 001010 through the tree produced a, a, b, a in order - exactly the message that was encoded, confirming the round trip is lossless.

94. Decode the bits back to a-a-b-a — line by line

Picture it

Animation

Shows: Each line of the worked example "Decode the bits back to a-a-b-a", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: Reading 001010 through the tree produced a, a, b, a in order - exactly the message that was encoded, confirming the round trip is lossless.

95. Something is wrong here: forgetting to restart at the root

Anomaly

Predict first

A student writes this, and it looks reasonable:

While decoding 001010, a student reaches b's leaf after reading 101 - but then keeps trying to walk from b's leaf instead of jumping back to the root for the next symbol.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Leaves have no children to walk to, so continuing from a leaf is not just wrong - it is undefined.

Every time a leaf is reached and a symbol is output, decoding must restart from the root before reading the next bit.

Why: Leaves have no children to walk to, so continuing from a leaf is not just wrong - it is undefined. The remaining bit, the final 0, never gets matched to anything.

96. Trap: forgetting to restart at the root

Trap

The trap

While decoding 001010, a student reaches b's leaf after reading 101 - but then keeps trying to walk from b's leaf instead of jumping back to the root for the next symbol.

Continue walking from wherever decoding stopped

Why: Leaves have no children to walk to, so continuing from a leaf is not just wrong - it is undefined. The remaining bit, the final 0, never gets matched to anything.

The fix

Every time a leaf is reached and a symbol is output, decoding must restart from the root before reading the next bit.

Return to the root after every symbol

Why: After outputting b, decoding restarts at the root, reads the final 0, and correctly lands on a's leaf, completing the message a, a, b, a.

97. Which of these survive contact with Huffman Coding?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
You have named 14 reusable moves so far. Say as many as you can out loud, by number, from memory.; Suppose you must store or send symbols from some alphabet, and you have to pick one fixed number of bits for every symbol, no matter how often it shows up.; Picture a texting abbreviation system. If you text 'lol' fifty times a day but 'defenestrate' once a year, it makes no sense to give both the same length shortcut.
Breaks
A student assigns: x gets 0, y gets 1, z gets 01 - reusing 0 and 1 as a 'combo' for the third symbol.; A student thinks 'Huffman merges nodes' and merges the two largest frequencies first instead of the two smallest - combining a (45) and d (16) right away.
sound
These are stated as this lesson states them — each one survives the edge cases Huffman Coding puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

98. Plan first: Encode and decode a second message: f-a-c-e

Step zero

Discussion prompt

Encode and decode a second message: f-a-c-e — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Concatenate the four codes in order

Answer:

  1. Concatenate the four codes in order
  2. Decode the result by walking the tree from the root
  3. Verify the round trip returns the original message

99. Encode and decode a second message: f-a-c-e

Worked example

Encode f, a, c, e using the full six-symbol code: f is 1100, a is 0, c is 100, e is 1101.

Concatenate the four codes in order

Why: Write each symbol's code back to back: f's code, then a's, then c's, then e's.

\[ 1100 \; 0 \; 100 \; 1101 \;\Rightarrow\; 110001001101 \]

Decode the result by walking the tree from the root

Why: Reading left to right: 1100 lands on f's leaf, 0 lands on a's leaf, 100 lands on c's leaf, and 1101 lands on e's leaf, restarting at the root after each one.

\[ 1100{:}f, \quad 0{:}a, \quad 100{:}c, \quad 1101{:}e \]

Verify the round trip returns the original message

Why: Decoding 110001001101 reproduces f, a, c, e in order, matching the message that was encoded - a full check that encoding and decoding are true inverses of each other.

100. Encode and decode a second message: f-a-c-e — line by line

Picture it

Animation

Shows: Each line of the worked example "Encode and decode a second message: f-a-c-e", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: Decoding 110001001101 reproduces f, a, c, e in order, matching the message that was encoded - a full check that encoding and decoding are true inverses of each other.

101. Answer it before you see the options: Check yourself: decode this bit string

Prediction

Predict first

What message does the bit string 1011000 decode to?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: b, c, a

Why: Reading from the root: the first three bits, 101, match b's code exactly, restart at the root; the next three, 100, match c's code exactly, restart; the final bit, 0, matches a's code. The bits decode to b, then c, then a.

102. Check yourself: decode this bit string

Check

Using the six-symbol code (a is 0, b is 101, c is 100, d is 111, e is 1101, f is 1100), decode the bit string 1011000.

Check your understanding

What message does the bit string 1011000 decode to?

  • A. b, c, a (correct)
  • B. c, b, a
  • C. b, d, a
  • D. a, b, c

Answer: A

Why: Reading from the root: the first three bits, 101, match b's code exactly, restart at the root; the next three, 100, match c's code exactly, restart; the final bit, 0, matches a's code. The bits decode to b, then c, then a.

Why B tempts people
This reverses the order of the decoded symbols - b and c are correct, but the message must be read in the order the bits were consumed, not backwards.
Why C tempts people
This mistakes c's code (100) for d's code (111), a mix-up between two similar-looking three-bit codes in the table.
Why D tempts people
This assumes decoding starts with a's short code first rather than following the actual bits from the root - the string does not even begin with a's code, 0.

103. Why Huffman Is Optimal

Section

Section 6

104. What optimal means here

Concept

A code is optimal for a given set of frequencies if no other prefix-free code achieves a smaller total encoded length for that same alphabet and those same frequencies.

\[ \text{optimal} \iff L(\text{this code}) \le L(\text{any other prefix-free code}) \]

Huffman's algorithm is not merely a reasonable greedy method - it always produces an optimal code, and that claim can be proven. The rest of this section sketches why.

105. How close to optimal it gets

Picture it

Animation

Shows: How close to optimal it gets — a rendered Manim animation.

Rendered with Manim.

Takeaway: Optimal among prefix codes; arithmetic coding beats it by dropping that restriction.

106. Prove it by assuming the best already exists

Intuition

Rather than building a tree and hoping it is optimal, the standard proof strategy assumes an optimal tree already exists, built by any method whatsoever, and studies its shape.

If we can show that some optimal tree always looks the way Huffman's first move would build it, that justifies making that move: it never rules out reaching an optimal answer.

107. Setup: assume an optimal tree exists

Concept

Call the alphabet's two least-frequent symbols x and y. Assume, for the sake of argument, that some tree T is an optimal prefix-code tree for this alphabet - we do not assume T is the tree Huffman would build.

exchange argument — A proof technique that starts from an assumed optimal solution and shows it can be reshaped, without hurting its quality, to match the structure your algorithm is about to build. If the reshaping never makes things worse, your algorithm's choice is justified.

The goal of this setup: show that T can be rearranged into another optimal tree in which x and y sit as sibling leaves at the deepest level - exactly the pair Huffman would merge first.

108. The Greedy Correctness skeleton

Concept

Every proof of this kind has the same five or six moves in the same order. The order is not something you rediscover each time.

It is on the right. It will stay on the right through the worked examples that follow.

Why this matters: the structure is now handled. You are not spending working memory on what comes next — you are spending all of it on the one hard step.

Steps 3 and 5 carry the proof. Step 5 is where a botched exchange hides — showing the swap is legal is not the same as showing it does not lose anything, and you need both.

109. Every full binary tree has two deepest siblings

Concept

A Huffman tree is a full binary tree: every internal node has exactly two children. In any such tree, somewhere at the maximum depth, there is a pair of leaves that share the same parent.

Call those two symbols p and q. They may or may not be x and y - the point of the next two slides is to swap them until they are.

110. Decision point: an optimal tree exists. Compare it to Huffman.

Intuition

What move should we make next?

We have assumed some optimal tree exists, and we know it has two deepest siblings.

\[ \text{optimal tree } \; T^{*}, \qquad \text{deepest sibling pair } \; (x, y) \]

Huffman's very first merge combines the two rarest symbols. The optimal tree might have done something else.

You have run this exact structure twice in the last lesson. Name the move, and say what you are about to swap with what.

_Look at your toolkit. Say a move number out loud before this slide advances._ A wrong guess is useful. A silent guess is not.

111. Why no prefix code does better

Picture it

Animation

Shows: Why no prefix code does better — a rendered Manim animation.

Rendered with Manim.

Takeaway: Two lemmas, then an exchange argument, then induction.

112. Complete the line: Step 1: swap the rarest symbol into the deepest spot

Fill the middle

Fill in the blanks

From Step 1: swap the rarest symbol into the deepest spot — finish the line. Write what belongs on the right of the equals sign before you look.

f_x \le f_p, \qquad d_p = \text{maximum depth in } T

Why: Producing the right-hand side unprompted is the difference between recognising this line and being able to use it. Move x down to depth d_p, and move p up to wherever x used to sit.

113. Step 1: swap the rarest symbol into the deepest spot

Concept

Since x is the single least-frequent symbol in the whole alphabet, its frequency is no greater than p's, and p sits at the tree's maximum depth.

\[ f_x \le f_p, \qquad d_p = \text{maximum depth in } T \]

Swap the leaves holding x and p

Why: Move x down to depth d_p, and move p up to wherever x used to sit. Every other leaf stays exactly where it was.

Compute how the total cost changes

Why: Only the two swapped leaves change depth, so only their two terms in the weighted sum change.

\[ \Delta = f_x d_p + f_p d_x - (f_x d_x + f_p d_p) = (d_p - d_x)(f_x - f_p) \]

Sign-check the change

Why: d_p is the maximum depth, so d_p minus d_x is at least zero; and f_x is at most f_p, so f_x minus f_p is at most zero. A non-negative number times a non-positive number is never positive.

\[ (d_p - d_x) \ge 0, \qquad (f_x - f_p) \le 0 \;\Rightarrow\; \Delta \le 0 \]

114. Say it in words: Step 1: swap the rarest symbol into the deepest…

Translation

\( (d_p - d_x) \ge 0, \qquad (f_x - f_p) \le 0 \;\Rightarrow\; \Delta \le 0 \)

Draw it

Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.

115. Why the swap cannot hurt

Intuition

Picture moving a light box to a high, hard-to-reach shelf, and moving a heavier box down to the easy, close shelf it vacated. The extra climbing now falls on the box that weighs the least.

That is exactly the swap: the rarely used symbol takes on the extra depth, while the symbol that was needlessly deep and at least as heavy moves up to a shallower spot. Total effort cannot go up.

116. Step 2: swap the second-rarest into the sibling spot

Concept

After step 1, x sits at the maximum depth, next to whichever leaf was already its sibling there. Call that sibling q - possibly the same q as before, possibly a new neighbor.

Repeat the same swap argument for y and q

Why: y is the second least-frequent symbol overall, so its frequency is no greater than q's. Swapping y into q's position, by the identical inequality as step 1, cannot increase the total cost.

\[ (d_q - d_y)(f_y - f_q) \le 0 \]

Now both x and y sit at the tree's maximum depth, as siblings sharing one parent - and every swap along the way only ever kept the cost the same or lower.

117. The same move, third subject

Concept

The move: #14 (Exchange argument), then #7 (Negate and assume).

Move #7 supplies the rival

Why: Assuming an optimal tree exists is what gives you an object to edit. Without that assumption there is nothing on the table to swap anything into.

Move #14 does the editing

Why: Swap the rarest symbol into the deepest position. Legal, because any leaf can hold any symbol. No worse, because you moved a low frequency to a deep spot and a higher frequency to a shallower one, which cannot increase the weighted total.

\[ (f_{\text{deep}} - f_{\text{rare}})(d_{\text{deep}} - d_{\text{rare}}) \ge 0 \]

Repeat for the second-rarest, then recurse

Why: After both swaps, some optimal tree has the two rarest symbols as deepest siblings — which is exactly what Huffman's first merge assumed. Then induct on the smaller alphabet.

That final induction is moves #5 and #6: merging two symbols into one peels the problem down to an alphabet one smaller, and the hypothesis is that Huffman is optimal there.

118. Decode the notation: The same move, third subject

Notation

Annotate

From The same move, third subject — read this one piece at a time. What is each part doing?

On: \( (f_{\text{deep}} - f_{\text{rare}})(d_{\text{deep}} - d_{\text{rare}}) \ge 0 \)

  • Assuming an optimal tree exists is what gives you an object to edit. Without that assumption there is nothing on the table to swap anything into.
  • Swap the rarest symbol into the deepest position. Legal, because any leaf can hold any symbol. No worse, because you moved a low frequency to a deep spot and a higher frequency to a shallower one, which cannot increase the weighted total.
  • After both swaps, some optimal tree has the two rarest symbols as deepest siblings — which is exactly what Huffman's first merge assumed. Then induct on the smaller alphabet.

119. Conclusion: Huffman's first merge always matches an optimal tree

Concept

Starting from any optimal tree T, we reshaped it, without raising its cost, into an optimal tree where the two least-frequent symbols, x and y, are sibling leaves at maximum depth.

That is exactly the pair Huffman's algorithm merges first. So the very first greedy choice is always safe: it never closes the door on reaching an optimal solution.

The same argument then applies to the smaller problem left after merging x and y into one combined symbol, and the one after that, all the way down - which is why induction on the number of symbols completes the full optimality proof.

120. Teach it back: Conclusion: Huffman's first merge always matches an optimal…

Explain it

Discussion prompt

Explain Conclusion: Huffman's first merge always matches an optimal tree to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Starting from any optimal tree T, we reshaped it, without raising its cost, into an optimal tree where the two least-frequent symbols, x and y, are sibling leaves at maximum depth.

121. What has to happen first: Verify the exchange lemma on our own tree

Ranking

Put in order

Put the moves of Verify the exchange lemma on our own tree into the order they have to happen.

  1. Find the maximum depth in our tree
  2. Check that f and e are siblings
  3. Verify the lemma holds without needing any swaps

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Looking back at the finished tree, f and e sit four branches down from the root - the deepest level anywhere in the tree.

122. Verify the exchange lemma on our own tree

Worked example

Check the lemma against the tree we actually built for a, b, c, d, e, f. The two least-frequent symbols are f (5) and e (9).

Find the maximum depth in our tree

Why: Looking back at the finished tree, f and e sit four branches down from the root - the deepest level anywhere in the tree.

\[ \text{depth}(f) = \text{depth}(e) = 4 = \text{maximum depth} \]

Check that f and e are siblings

Why: Both hang directly off the same internal node, the one we labeled fe:14, so they share a parent, not just a depth.

Verify the lemma holds without needing any swaps

Why: f and e, the two rarest symbols, are already siblings at the tree's maximum depth in the tree Huffman actually built - precisely what the exchange argument guarantees is always achievable.

123. Cost of building the tree

Picture it

Animation

Shows: Cost of building the tree — a rendered Manim animation.

Rendered with Manim.

Takeaway: The heap is what makes always-take-the-two-rarest cheap.

124. Decision point: account for the whole proof

Intuition

What move should we make next?

The optimality proof is finished. Look back at it as a whole.

It used four moves you already owned, and introduced nothing new.

Name all four, in the order they appear in the proof, and say which line of the proof each one is doing.

_Look at your toolkit. Say a move number out loud before this slide advances._ A wrong guess is useful. A silent guess is not.

125. By analogy: Decision point: account for the whole proof

Analogy

Discussion prompt

Explain Decision point: account for the whole proof by analogy to something with no CS3000 Algorithms in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

The optimality proof is finished. Look back at it as a whole.

126. Rule out three: Check yourself: the exchange argument

Elimination

Eliminate the wrong options

Why is it safe to swap x into p's position (and p into x's old position)?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Moving the lighter symbol deeper and the heavier one shallower cannot increase the total weighted cost, since the extra depth lands on the smaller frequency.
  • B. Because the least-frequent symbol must always sit at the root of an optimal tree.
  • C. Because a prefix-free code requires every symbol to sit at the same depth.
  • D. Because swapping any two leaves in any binary tree never changes its total cost.

Survives elimination: A

Why: The cost change from this swap equals (depth of p minus depth of x) times (frequency of x minus frequency of p). Since p sits at the maximum depth, the first factor is at least zero; since x has the smallest frequency, the second factor is at most zero. A non-negative number times a non-positive number is never positive, so the swap can only keep the cost the same or lower it.

127. Check yourself: the exchange argument

Check

Recall the setup: T is an assumed-optimal tree, x is the single rarest symbol, and p is a symbol currently sitting at T's maximum depth, with p's frequency at least as large as x's.

Check your understanding

Why is it safe to swap x into p's position (and p into x's old position)?

  • A. Moving the lighter symbol deeper and the heavier one shallower cannot increase the total weighted cost, since the extra depth lands on the smaller frequency. (correct)
  • B. Because the least-frequent symbol must always sit at the root of an optimal tree.
  • C. Because a prefix-free code requires every symbol to sit at the same depth.
  • D. Because swapping any two leaves in any binary tree never changes its total cost.

Answer: A

Why: The cost change from this swap equals (depth of p minus depth of x) times (frequency of x minus frequency of p). Since p sits at the maximum depth, the first factor is at least zero; since x has the smallest frequency, the second factor is at most zero. A non-negative number times a non-positive number is never positive, so the swap can only keep the cost the same or lower it.

Why B tempts people
This confuses the rarest symbol's depth with the root - Huffman actually puts the most frequent symbols closest to the root, and the rarest symbols deepest, the opposite of this claim.
Why C tempts people
This describes a fixed-length code, not a prefix-free code in general. Prefix-free codes can, and typically do, have leaves at different depths.
Why D tempts people
This overgeneralizes: swapping two leaves only preserves-or-lowers cost under the specific condition used here, that the deeper leaf's frequency is at least the shallower one's. Swapping arbitrary leaves can raise the cost.

128. Toolkit update

Concept

Moves added today: none.

That is a result, not a gap. Everything in this lesson was proved with moves you already owned.

Moves you reused today:

Huffman's optimality proof is move #14 wrapped in move #7, then finished with moves #5 and #6. Every part of it is something you named in an earlier lesson.

Full toolkit so far: #1 through #14.

Next session opens with you naming every one of these from memory, before any new material.

129. Break it if you can: Toolkit update

Counterexample

Discussion prompt

Next session opens with you naming every one of these from memory, before any new material.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

130. Connect it up: Huffman Coding

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Why Variable-Length Codes · Prefix-Free Codes · Building the Huffman Tree · Costs & Comparisons · Encoding and Decoding · Why Huffman Is Optimal. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

131. What you can do now

Recap

Huffman coding builds the best possible prefix-free code for a given set of symbol frequencies.

IdeaThe one move
Prefix-freeNo code is a prefix of another - decode without separators
Huffman mergeAlways combine the two least-frequent nodes, then reinsert
OptimalitySwap the two rarest symbols into the deepest sibling spot without raising cost

Sources

  1. D. A. Huffman, 'A Method for the Construction of Minimum-Redundancy Codes', Proceedings of the IRE, 40(9), 1952 — Original source of the Huffman coding algorithm.
  2. Cormen, Leiserson, Rivest, Stein, Introduction to Algorithms (CLRS), 3rd ed., Section 16.3 (Huffman Codes) - source of the six-symbol alphabet (a:45, b:13, c:12, d:16, e:9, f:5) used throughout this deck — MIT Press, 2009.
  3. Every merge, code length, encoded/decoded bit string, and the exchange-argument cost inequality was re-derived by hand and checked against the CLRS worked example before writing this deck. — Verified 2026-07-18.
  4. Northeastern University CS 3000, Algorithms and Data (Summer 2026) — course page and syllabus — course.ccs.neu.edu/cs3000su26. Sets Cormen, Leiserson, Rivest and Stein, Introduction to Algorithms (3rd ed.) as the textbook; listings follow its conventions.
  5. CS 3000 course notes and midterm references circulated by students — github.com/vigneshsaravanakumar404/CS-3000-Algorithms-Data. Notes are typeset with the algpseudocode package, which is the style the listings in this deck follow.

Want this taught 1-on-1? Alexander tutors CS3000 Algorithms — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108