This deck traces merge sort end to end, explains why the merge step is linear, sets up the merge sort recurrence, and works the Master Theorem's three cases by comparing the extra work f(n) against the watershed function n^log_b(a). It targets swapping a and b, mis-comparing f(n) to the watershed, assuming the merge step is free, and force-fitting recurrences - those with log-factor gaps or unequal splits - that do not satisfy the theorem's hypotheses.
Subject: CS3000 Algorithms · 135 slides · symbolic lesson
Open the interactive version of this deck · Homework for this lesson
Title
CS3000 Algorithms
How to prove a divide-and-conquer algorithm's running time -- and why merge sort lands at n log n.
Objectives
Divide and conquer is the workhorse pattern of this course. By the end of this lesson you can:
Warm-up
Discussion prompt
Before we open Merge Sort & the Master Theorem: without looking back, what was the main idea of Recurrences & Recursion Trees, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
How to turn a recursive algorithm into a recurrence, then solve it three ways: the recursion-tree method (work per level times number of levels), unrolling by repeated substitution, and the substitution method's proof by induction. Covers both equal and unequal subproblem sizes.
Concept
Before any new material: cover the screen.
You have named 11 reusable moves so far. Say as many as you can out loud, by number, from memory.
Do not advance until you have actually tried. Getting four of eight is information; skipping the exercise is not.
Here they are. Score yourself.
Today adds no new moves. Every proof in this lesson is built out of the list above. That is the whole point of the list.
The question that starts every proof from here on is not how do I begin. It is which of these applies here?
Counterexample
Discussion prompt
You have named 11 reusable moves so far. Say as many as you can out loud, by number, from memory.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Do not advance until you have actually tried. Getting four of eight is information; skipping the exercise is not.
Section
Section 1
Concept
A divide-and-conquer algorithm always follows the same three-part recipe: split the problem into smaller pieces, solve each piece the same way, and then combine the pieces back into a single answer.
divide and conquer — Divide: break the input into smaller subproblems. Conquer: solve each subproblem recursively (the same recipe, applied to a smaller input). Combine: stitch the solved subproblems into a solution for the whole input.
Concept
Merge sort fills in that recipe in one specific way: divide the array into two halves, conquer by recursively sorting each half, then combine by merging the two now-sorted halves into one sorted array.
Every other sorting choice merge sort could have made -- splitting into three pieces, splitting unevenly -- would give a different algorithm. This particular choice, an even two-way split, is what makes the later math so clean.
Picture it
Animation
Shows: Split until nothing is left to split — a rendered Manim animation.
Rendered with Manim.
Takeaway: Splitting costs nothing. All the work is in putting it back together.
Concept
Recursion has to bottom out somewhere. For merge sort, an array of size zero or one is already sorted by definition -- there is nothing to compare, so it is simply returned as-is.
Every recursive call keeps splitting smaller arrays into still-smaller ones until this base case is hit. That is what guarantees the recursion actually terminates instead of splitting forever.
Analogy
Discussion prompt
Explain The base case: when to stop dividing by analogy to something with no CS3000 Algorithms in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Recursion has to bottom out somewhere. For merge sort, an array of size zero or one is already sorted by definition -- there is nothing to compare, so it is simply returned as-is.
Intuition
Picture two piles of playing cards, each pile already arranged from lowest to highest. To combine them into one sorted pile, you do not need to re-sort anything from scratch.
You only ever look at the top card of each pile, take whichever one is smaller, and place it down. Repeat until both piles are empty. That single move -- compare the two fronts, take the smaller -- is the entire merge step.
Explain it
Discussion prompt
Explain Two piles of sorted cards to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Picture two piles of playing cards, each pile already arranged from lowest to highest. To combine them into one sorted pile, you do not need to re-sort anything from scratch.
Ranking
Put in order
Put the moves of Trace the divide phase into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. An 8-element array splits into two 4-element halves -- the left half keeps its original order, the right half keeps its original order.
Worked example
Start with the array below and repeatedly split it in half until every piece has just one element.
\[ [8,\,3,\,5,\,1,\,9,\,2,\,7,\,4] \]
Split once down the middle
Why: An 8-element array splits into two 4-element halves -- the left half keeps its original order, the right half keeps its original order.
\[ [8,3,5,1] \quad\big|\quad [9,2,7,4] \]
Split each half again
Why: Each 4-element piece splits into two 2-element pieces, the same rule applied one level deeper.
\[ [8,3]\ [5,1] \quad\big|\quad [9,2]\ [7,4] \]
Split once more to reach single elements
Why: Each 2-element piece splits into two 1-element pieces. A single element is the base case -- already sorted, nothing left to divide.
\[ [8]\,[3]\ [5]\,[1] \quad\big|\quad [9]\,[2]\ [7]\,[4] \]
| level | subarrays produced | size of each |
|---|---|---|
| 0 | [8,3,5,1,9,2,7,4] | 8 |
| 1 | [8,3,5,1] and [9,2,7,4] | 4 |
| 2 | [8,3] [5,1] [9,2] [7,4] | 2 |
| 3 | [8][3][5][1][9][2][7][4] | 1 |
Verify the divide phase is complete
Why: Every one of the original 8 numbers appears in exactly one singleton array, and no splitting rule was applied to a piece of size one -- the base case was respected at every branch.
Worked example
Now build the sorted array back up from the singletons produced above, merging pairs at every level.
Merge each pair of singletons
Why: Merging two single elements just means putting the smaller one first.
| step | left | right | merged |
|---|---|---|---|
| 1 | [8] | [3] | [3,8] |
| 2 | [5] | [1] | [1,5] |
| 3 | [9] | [2] | [2,9] |
| 4 | [7] | [4] | [4,7] |
Merge the resulting pairs
Why: Now each merge combines two already-sorted 2-element arrays into a sorted 4-element array, comparing fronts and taking the smaller each time.
| step | left | right | merged |
|---|---|---|---|
| 5 | [3,8] | [1,5] | [1,3,5,8] |
| 6 | [2,9] | [4,7] | [2,4,7,9] |
Merge the two halves into the final array
Why: The last merge combines the two sorted 4-element halves into the full sorted 8-element array.
\[ [1,3,5,8]\ \text{merged with}\ [2,4,7,9] \ \longrightarrow\ [1,2,3,4,5,7,8,9] \]
Verify the final array is correctly sorted
Why: It contains exactly the eight original numbers -- 8,3,5,1,9,2,7,4 -- now arranged from smallest to largest, with each value adjacent to its correct neighbors.
\[ 1<2<3<4<5<7<8<9\ \checkmark \]
Picture it
Animation
Shows: The space merge sort actually needs — a rendered Manim animation.
Rendered with Manim.
Takeaway: Which is the main reason quicksort is often preferred in practice.
Step zero
Discussion prompt
One full trace, start to finish — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Divide all the way down to singletons
Answer:
Worked example
Run the whole algorithm, divide then merge, on a fresh small array.
\[ [5,\,2,\,8,\,1] \]
Divide all the way down to singletons
Why: Split into halves, then split each half again, stopping once every piece has one element.
\[ [5,2,8,1] \to [5,2]\ [8,1] \to [5]\,[2]\ [8]\,[1] \]
Merge each pair of singletons
Why: Compare the two single elements and place the smaller one first.
\[ [5],[2] \to [2,5] \qquad [8],[1] \to [1,8] \]
Merge the two sorted pairs into the final array
Why: Compare fronts: 2 vs 1 (take 1), then 2 vs 8 (take 2), then 5 vs 8 (take 5), then only 8 remains.
\[ [2,5]\ \text{and}\ [1,8] \ \longrightarrow\ [1,2,5,8] \]
Verify by comparing to the original elements
Why: The result contains exactly the original four numbers -- 5, 2, 8, 1 -- now in non-decreasing order.
\[ 1\le 2 \le 5 \le 8\ \checkmark \]
Intuition
Every trace you just did fits inside a single picture: a tree where the root is the whole array, its two children are the two halves, and so on down to single elements at the bottom.
Figure (svg): A tree showing an array of size n splitting into two halves of size n over 2, then into four quarters of size n over 4
Splitting happens on the way down this tree; merging happens on the way back up, starting from the bottom. Every node in the tree eventually does one merge -- except the leaves, which are already sorted by having only one element.
Anomaly
Predict first
A student writes this, and it looks reasonable:
A student figures the hard work already happened during the recursive calls, so sticking two sorted halves together must be quick -- basically instant, no matter how big the halves are.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: This treats combining two sorted halves as one quick operation, as if the size of the halves did not matter at all.
Merging two sorted halves means walking through every element at least once -- comparing the two fronts and copying the smaller one into the output, over and over, until both halves are used up.
Why: This treats combining two sorted halves as one quick operation, as if the size of the halves did not matter at all.
Trap
A student figures the hard work already happened during the recursive calls, so sticking two sorted halves together must be quick -- basically instant, no matter how big the halves are.
\[ T(n) = 2T\!\left(\frac{n}{2}\right) + O(1) \]
Model the merge step as a fixed, tiny amount of work
Why: This treats combining two sorted halves as one quick operation, as if the size of the halves did not matter at all.
\[ \text{(wrong)}\quad T(n) = \Theta(n) \]
Merging two sorted halves means walking through every element at least once -- comparing the two fronts and copying the smaller one into the output, over and over, until both halves are used up.
\[ T(n) = 2T\!\left(\frac{n}{2}\right) + \Theta(n) \]
Model the merge step as work proportional to the total size
Why: The merge step touches every element being merged exactly once, so it costs an amount proportional to how many elements are being merged -- not a fixed amount.
\[ \text{(right)}\quad T(n) = \Theta(n\log n) \]
Notation
Annotate
From Trap: assuming the merge step is free — read this one piece at a time. What is each part doing?
On: \( T(n) = 2T\!\left(\frac{n}{2}\right) + O(1) \)
Concept
Merge sort: if the array has zero or one element, it is already sorted. Otherwise, split it into two halves, recursively sort each half, and merge the two sorted halves into the final sorted array.
Three moving parts -- a base case, two recursive calls, and one merge -- are everything the algorithm is. The rest of this lesson is about timing exactly how long it takes.
Concept
Six lines, and only one of them does any actual comparing. Merge sort itself just splits; the sorting all happens inside the merge on the last line.
MERGE-SORT(A, lo, hi)
if lo >= hi
return
mid = floor((lo + hi) / 2)
MERGE-SORT(A, lo, mid)
MERGE-SORT(A, mid + 1, hi)
MERGE(A, lo, mid, hi)Notice the order. Line 7 runs after both recursive calls have finished, so by the time anything is merged, both halves are already sorted. Merging two sorted halves is easy; that is the entire trick.
Notation
Every line of MERGE-SORT says one thing. Read the line, then read what it does — not the other way round.
Annotate
Invariant
Watch which parts of the array are settled. A range turns green only after its own merge has finished, and the ranges settle from the smallest outward.
Step through it
At each merge, say which two already-sorted ranges are being combined.
Picture it
Animation
Shows: MERGE-SORT executing: the current line of pseudocode is highlighted while the data it touches changes.
Rendered with Manim.
Takeaway: Split blindly at the middle, sort both halves, then merge — all the comparing happens in the merge, after the halves are already sorted.
Section
Section 2
Concept
The merge step keeps one pointer at the front of each sorted half. It repeatedly compares the two pointed-at values, copies the smaller one into the output, and advances only the pointer that lost the comparison.
two-pointer merge — When one half runs out entirely, every remaining element of the other half is already in sorted order, so it can simply be copied over directly with no more comparisons needed.
Concept
Two fingers, one on each sorted half, and a rule for which one moves. Every pass appends exactly one item and advances exactly one finger, which is why the whole merge is linear.
MERGE(L, R)
i = 1
j = 1
out = empty list
while i <= L.length and j <= R.length
if L[i] <= R[j]
append L[i] to out
i = i + 1
else
append R[j] to out
j = j + 1
append whatever remains of L or R
return outLine 12 is not an afterthought. When one side runs out, everything left on the other side is already sorted and already larger than everything appended so far, so it can be copied across without a single further comparison.
Notation
Every line of MERGE says one thing. Read the line, then read what it does — not the other way round.
Annotate
Invariant
The output list is sorted at every single step, and every item in it is smaller than everything still waiting in either half. That is the invariant the merge maintains.
Step through it
Before each step, predict which half the next item comes from.
Picture it
Animation
Shows: MERGE executing: the current line of pseudocode is highlighted while the data it touches changes.
Rendered with Manim.
Takeaway: One comparison per appended item, so merging two halves of total size n costs n — linear, not quadratic.
Picture it
Animation
Shows: Merge sort is stable, if you break ties left — a rendered Manim animation.
Rendered with Manim.
Takeaway: One comparison operator decides whether the sort is stable.
Intuition
Think of a zipper: each tooth from the left side interlocks with a tooth from the right side, one at a time, moving forward. No tooth is ever revisited, and the zipper never needs to back up.
The merge works the same way -- each pointer only ever moves forward, each element is looked at exactly once, and the process finishes the moment both sides are used up.
Estimation
Predict first
Return to the final merge from the divide-and-merge trace earlier: two sorted 4-element halves becoming one sorted 8-element array.
Commit before you compute: what does Count the work in one real merge come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Verify the total work is proportional to the combined size
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Seven comparisons plus one leftover copy touches all eight elements exactly once.
Worked example
Return to the final merge from the divide-and-merge trace earlier: two sorted 4-element halves becoming one sorted 8-element array.
\[ [1,3,5,8] \quad\text{and}\quad [2,4,7,9] \]
Walk the two pointers forward, one comparison at a time
Why: At each step, compare the two front values and copy the smaller one; that pointer then advances to its next value.
| step | compare | take | output so far |
|---|---|---|---|
| 1 | 1 vs 2 | 1 | [1] |
| 2 | 3 vs 2 | 2 | [1,2] |
| 3 | 3 vs 4 | 3 | [1,2,3] |
| 4 | 5 vs 4 | 4 | [1,2,3,4] |
| 5 | 5 vs 7 | 5 | [1,2,3,4,5] |
| 6 | 8 vs 7 | 7 | [1,2,3,4,5,7] |
| 7 | 8 vs 9 | 8 | [1,2,3,4,5,7,8] |
| 8 | (right side used up) | 9 | [1,2,3,4,5,7,8,9] |
Verify the total work is proportional to the combined size
Why: Seven comparisons plus one leftover copy touches all eight elements exactly once. That is proportional to n, not to n multiplied by itself -- confirming the merge step is linear.
\[ \text{comparisons} \le 4+4-1 = 7 \]
Picture it
Animation
Shows: Each line of the worked example "Count the work in one real merge", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Seven comparisons plus one leftover copy touches all eight elements exactly once. That is proportional to n, not to n multiplied by itself -- confirming the merge step is linear.
Concept
Every step of the merge consumes exactly one element from exactly one side, and a side's pointer never moves backward. So the number of steps can never exceed the combined size of the two halves.
\[ \text{total elements copied} = n_{\text{left}} + n_{\text{right}} = n, \qquad \text{comparisons} \le n-1 \]
Missing information
Discussion prompt
The two halves do not even need to be the same size for this to hold. Merge a sorted 5-element array with a sorted 3-element array.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
The rule never changed: compare the two front values, take the smaller, advance that pointer.
Worked example
The two halves do not even need to be the same size for this to hold. Merge a sorted 5-element array with a sorted 3-element array.
\[ [2,4,6,8,10] \quad\text{and}\quad [1,5,9] \]
Walk both pointers forward exactly as before
Why: The rule never changed: compare the two front values, take the smaller, advance that pointer.
| step | compare | take | output so far |
|---|---|---|---|
| 1 | 2 vs 1 | 1 | [1] |
| 2 | 2 vs 5 | 2 | [1,2] |
| 3 | 4 vs 5 | 4 | [1,2,4] |
| 4 | 6 vs 5 | 5 | [1,2,4,5] |
| 5 | 6 vs 9 | 6 | [1,2,4,5,6] |
| 6 | 8 vs 9 | 8 | [1,2,4,5,6,8] |
| 7 | 10 vs 9 | 9 | [1,2,4,5,6,8,9] |
| 8 | (right side used up) | 10 | [1,2,4,5,6,8,9,10] |
Verify the bound still holds
Why: Seven comparisons plus one leftover copy again equals exactly the combined total of eight elements -- the bound holds regardless of how uneven the two sides are.
\[ \text{comparisons} \le 5+3-1 = 7 \ \checkmark \]
Picture it
Animation
Shows: Why n log n is the floor — a rendered Manim animation.
Rendered with Manim.
Takeaway: No comparison sort can beat this. Counting sorts get around it by not comparing.
Section
Section 3
Concept
Sorting an array of size n breaks into: two recursive calls, each sorting an array half the size, plus the merge step at the end, which we just showed costs an amount proportional to n.
\[ T(n) = 2\,T\!\left(\frac{n}{2}\right) + \Theta(n), \qquad T(1) = \Theta(1) \]
Concept
It helps to name each piece before doing anything else with it.
T of n — The running time of the algorithm on an input of size n. This is the quantity the whole recurrence is trying to describe.
the 2 — How many recursive calls are made -- here, two, one for each half of the array.
n over 2 — The size of each subproblem -- here, each recursive call works on an array half the size of the original.
the extra Theta of n term — The work done outside the recursive calls -- for merge sort, this is exactly the merge step, which we already showed costs an amount proportional to n.
Definition probe
Sort into buckets
Every line below is part of the definition of divide and conquer or of T of n — one or the other, never both. Put each where it belongs.
Intuition
Draw one node for every recursive call. The root is the call on the full array; it has two children, one for each half; each of those has two children, one for each quarter; and so on, down to single elements.
At every node in this tree, some merge work happens once its children return. Adding up all of that work, node by node, gives the total running time.
Sorting
Sort into buckets
These are the pieces of Merge Sort & the Master Theorem, out of order. Put each one back under the part of the lesson it belongs to.
Picture it
Animation
Shows: The recursion tree of merge sort — a rendered Manim animation.
Rendered with Manim.
Takeaway: Each level does n work in total, and the tree is log n deep.
Estimation
Predict first
Take the concrete case of eight elements and count how many times the size can be cut in half before reaching one.
Commit before you compute: what does Count the levels of the tree come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Verify the count against the doubling relationship
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Three halvings means two raised to the third power gets back to eight, so the number of halvings is exactly the base-two logarithm of eight.
Worked example
Take the concrete case of eight elements and count how many times the size can be cut in half before reaching one.
\[ n = 8 \]
Halve repeatedly and count the halvings
Why: Eight becomes four, four becomes two, two becomes one -- three halvings in total.
\[ 8 \to 4 \to 2 \to 1 \qquad (3\ \text{halvings}) \]
Verify the count against the doubling relationship
Why: Three halvings means two raised to the third power gets back to eight, so the number of halvings is exactly the base-two logarithm of eight.
\[ 2^3 = 8 \ \Rightarrow\ \log_2 8 = 3\ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Count the levels of the tree", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Three halvings means two raised to the third power gets back to eight, so the number of halvings is exactly the base-two logarithm of eight.
Intuition
A recursion that only shrinks by removing one element at a time needs a number of levels equal to the size itself -- shrink a thousand-element problem down to nothing, one at a time, and that is a thousand levels.
A recursion that shrinks by cutting the size in half needs far fewer levels, because halving grows the remaining distance to the base case much faster: doubling the input size only adds a single extra level, not a whole new batch of them.
Intuition
What move should we make next?
For merge sort we now know both numbers:
\[ \text{work per level} = n, \qquad \text{number of levels} = \log_{2} n \]
You have done this exact step before, on a different lesson. Name the move.
_Look at your toolkit. Say a move number out loud before this slide advances._ A wrong guess is useful. A silent guess is not.
Missing information
Discussion prompt
Using the same eight-element case, add up the merge work done at each level of the recursion tree.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
At every level, the subproblems together still cover all n original elements, so the merge work at that level always adds up to n -- regardless of how many subproblems that level has been split into.
Worked example
Using the same eight-element case, add up the merge work done at each level of the recursion tree.
Total the merge work level by level
Why: At every level, the subproblems together still cover all n original elements, so the merge work at that level always adds up to n -- regardless of how many subproblems that level has been split into.
| level | subproblems | size each | merge work at level |
|---|---|---|---|
| 0 | 1 | 8 | 8 |
| 1 | 2 | 4 | 8 |
| 2 | 4 | 2 | 8 |
| 3 (base case) | 8 | 1 | 0 |
Verify the total matches n times the number of levels
Why: Three levels of merging, each totaling 8 units of work, gives 24 total -- exactly n multiplied by the level count computed in the previous slide.
\[ 8+8+8 = 24 = 8\cdot\log_2 8\ \checkmark \]
Picture it
Animation
Shows: Every level costs n — a rendered Manim animation.
Rendered with Manim.
Takeaway: n per level, log n levels. The product is the running time.
Concept
That trick -- count the levels, total the work per level, multiply -- worked because merge sort's recurrence has a very regular shape: the same number of subproblems at every branch, the same shrink factor every time.
The Master Theorem packages exactly that trick into a reusable rule, so it never has to be rebuilt by hand for every new recurrence of this shape.
Section
Section 4
Concept
The Master Theorem applies to any recurrence that fits one specific mold: some number of equally-sized subproblems, plus some extra work done outside of them.
\[ T(n) = a\,T\!\left(\frac{n}{b}\right) + f(n), \qquad a \ge 1,\ b > 1 \]
Before applying any of its three cases, the very first thing to do is confirm the recurrence actually has this shape -- one fixed number of calls, all on subproblems of the same size.
Concept
The letter a counts how many recursive calls are made at each step -- how many subproblems the input gets split into.
branching factor — The number of children each node in the recursion tree has. For merge sort, a is two -- every call splits into exactly two recursive calls.
Matching
Match the pairs
Match each term to the definition this lesson gave it — not the one you would guess from the word.
Why: These are the working definitions of T of n, the 2, n over 2, the extra Theta of n term, branching factor as Merge Sort & the Master Theorem uses them. Pairing them correctly is the test of whether you could state each one with the slide switched off.
Concept
The letter b measures something different: the factor by which the input size shrinks for each subproblem. It is not how many pieces you get -- it is how much smaller each piece is.
For merge sort, b is two, because each recursive call receives an array half the size of the original -- the same number, two, as a, but answering a completely different question. Mixing these two up is the single most common Master Theorem mistake.
Concept
The function f of n is everything the algorithm does outside the recursive calls at each step -- splitting the input apart and combining the results back together.
For merge sort, f of n is exactly the merge step, which was already shown to cost an amount proportional to n. Different divide-and-conquer algorithms plug in completely different functions here.
Picture it
Figure (svg): A recursion tree with the root highlighted in orange to show where the extra combining work happens, and the leaves highlighted in teal to show where the recursive subproblems bottom out
Discussion prompt
Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.
Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.
Answer:
Every recurrence of this shape has two competing sources of work: all the tiny bits of extra work happening at every node all the way down at the leaves, added together, versus the single extra-work cost happening at the very top, the root.
Intuition
Every recurrence of this shape has two competing sources of work: all the tiny bits of extra work happening at every node all the way down at the leaves, added together, versus the single extra-work cost happening at the very top, the root.
Figure (svg): A recursion tree with the root highlighted in orange to show where the extra combining work happens, and the leaves highlighted in teal to show where the recursive subproblems bottom out
Whichever side wins that comparison decides the shape of the final running time. The Master Theorem is really just a precise way of judging that comparison.
Concept
At depth k down the tree there are a-to-the-k subproblems, each of size n over b-to-the-k. The tree bottoms out at the depth where the size reaches one -- a depth of the base-b logarithm of n.
\[ \text{number of leaves} = a^{\log_b n} = n^{\log_b a} \]
That quantity, n raised to the base-b logarithm of a, is called the watershed function. It represents the total work the leaves alone would contribute if each did a small constant amount -- the natural yardstick to compare the extra work against.
\[ \text{watershed}(n) = n^{\log_b a} \]
Concept
Every proof of this kind has the same five or six moves in the same order. The order is not something you rediscover each time.
It is on the right. It will stay on the right through the worked examples that follow.
Why this matters: the structure is now handled. You are not spending working memory on what comes next — you are spending all of it on the one hard step.
Steps 1, 2 and 6 are lookup. Step 3 is the only judgment call, and step 5 is the one everybody skips — a case whose side condition fails does not apply, no matter how well f(n) seemed to match.
Picture it
Animation
Shows: The Master Theorem, three cases — a rendered Manim animation.
Rendered with Manim.
Takeaway: Compare the work at the leaves against the work at the root.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Given the recurrence below, a student reads the coefficient as the shrink factor and the denominator as the branching factor -- exactly backwards.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: This mixes up which number counts the subproblems and which number measures how much smaller each one is.
Read the recurrence in order: the coefficient in front of T is always a, the number dividing n inside T is always b.
Why: This mixes up which number counts the subproblems and which number measures how much smaller each one is.
Trap
Given the recurrence below, a student reads the coefficient as the shrink factor and the denominator as the branching factor -- exactly backwards.
\[ T(n) = 4\,T\!\left(\frac{n}{2}\right) + n \]
Swap the labels: treat 4 as b and 2 as a
Why: This mixes up which number counts the subproblems and which number measures how much smaller each one is.
\[ \text{(wrong)}\ a=2,\ b=4 \ \Rightarrow\ n^{\log_4 2} = n^{0.5} \]
Compare against the wrong watershed and pick a case
Why: Against the incorrect watershed n to the one-half, the extra work n looks polynomially bigger, wrongly pointing to the case where the root dominates.
\[ \text{(wrong conclusion)}\quad T(n) = \Theta(n) \]
Read the recurrence in order: the coefficient in front of T is always a, the number dividing n inside T is always b.
\[ T(n) = 4\,T\!\left(\frac{n}{2}\right) + n \]
Read a and b directly off the recurrence
Why: Four recursive calls are made (a equals four); each one works on an array half the size (b equals two).
\[ a=4,\ b=2 \ \Rightarrow\ n^{\log_2 4} = n^{2} \]
Compare against the correct watershed and pick a case
Why: Against the correct watershed n squared, the extra work n is polynomially smaller, correctly pointing to the case where the leaves dominate.
\[ \text{(right conclusion)}\quad T(n) = \Theta(n^2) \]
Section
Section 5
Concept
Once a, b, and f of n are correctly identified, the whole question comes down to one comparison: how does the extra work stack up against the watershed function?
There are exactly three possibilities: the extra work is much smaller (the leaves win), about the same (the two sides balance), or much bigger (the root wins). Each possibility has its own case and its own answer.
Concept
If the extra work is polynomially smaller than the watershed -- smaller by some fixed exponent, not just a little smaller -- then almost all the total work happens down at the leaves.
\[ f(n) = O\!\left(n^{\log_b a - \varepsilon}\right) \ \text{for some}\ \varepsilon>0 \]
In that case, the running time is simply the watershed function itself -- the extra work at every level barely matters next to how much the leaves contribute.
\[ \Rightarrow\ T(n) = \Theta\!\left(n^{\log_b a}\right) \]
Step zero
Discussion prompt
Apply Case 1 — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Read off a, b, and f(n)
Answer:
Worked example
Solve the recurrence below.
\[ T(n) = 9\,T\!\left(\frac{n}{3}\right) + n, \qquad T(1)=1 \]
Read off a, b, and f(n)
Why: Nine recursive calls are made, each on a subproblem a third the size, with n units of extra work outside the calls.
\[ a=9,\ b=3,\ f(n)=n \]
Compute the watershed
Why: Since 3 squared is 9, the base-three logarithm of nine is two.
\[ \log_3 9 = 2 \ \Rightarrow\ \text{watershed}(n) = n^2 \]
Compare f(n) to the watershed
Why: The extra work n is polynomially smaller than the watershed n squared -- it is at most n to the power one, which is n squared minus one -- so this is Case 1.
\[ n = O(n^{2-1}) \ \Rightarrow\ T(n) = \Theta(n^2) \]
Verify against the exact closed form
Why: Solving the recurrence directly gives a formula matching every value computed by hand: at n equals twenty-seven, the formula gives 1080, exactly matching direct recursive computation.
\[ T(n) = 1.5n^2 - 0.5n \ \Rightarrow\ T(27) = 1.5(729)-0.5(27) = 1080\ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Apply Case 1", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Solving the recurrence directly gives a formula matching every value computed by hand: at n equals twenty-seven, the formula gives 1080, exactly matching direct recursive computation.
Concept
If the extra work matches the watershed within a constant factor -- neither one polynomially bigger than the other -- then every level of the tree contributes about the same amount of total work.
\[ f(n) = \Theta\!\left(n^{\log_b a}\right) \]
Since there are a logarithmic number of levels, and each level contributes about the same amount, the total is the watershed function multiplied by the number of levels.
\[ \Rightarrow\ T(n) = \Theta\!\left(n^{\log_b a}\log n\right) \]
Intuition
What move should we make next?
Merge sort's recurrence, with its three ingredients read off:
\[ T(n) = 2T(n/2) + n \]
\[ a = 2, \qquad b = 2, \qquad f(n) = n \]
Do not name the case yet. First say what two things you are about to compare, and which is which.
_Look at your toolkit. Say a move number out loud before this slide advances._ A wrong guess is useful. A silent guess is not.
Estimation
Predict first
Solve merge sort's own recurrence using the Master Theorem, and see whether it matches the direct level-by-level count done earlier.
Commit before you compute: what does Apply Case 2 to merge sort itself come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Verify against the direct level-by-level count
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. For n equal to eight, this recurrence solves exactly to n times the base-two logarithm of n, which is 8 times 3, or 24 -- exactly the total computed earlier by summing the recursion tree level by level.
Worked example
Solve merge sort's own recurrence using the Master Theorem, and see whether it matches the direct level-by-level count done earlier.
\[ T(n) = 2\,T\!\left(\frac{n}{2}\right) + n, \qquad T(1)=0 \]
Read off a, b, and f(n)
Why: Two recursive calls, each on half the input, plus n units of merge work.
\[ a=2,\ b=2,\ f(n)=n \]
Compute the watershed
Why: The base-two logarithm of two is one, since two to the first power is two.
\[ \log_2 2 = 1 \ \Rightarrow\ \text{watershed}(n) = n^1 = n \]
Compare f(n) to the watershed
Why: The extra work n exactly matches the watershed n -- neither is polynomially bigger than the other -- so this is Case 2.
\[ f(n) = \Theta(n) = \Theta(n^{\log_2 2}) \ \Rightarrow\ T(n) = \Theta(n\log n) \]
Verify against the direct level-by-level count
Why: For n equal to eight, this recurrence solves exactly to n times the base-two logarithm of n, which is 8 times 3, or 24 -- exactly the total computed earlier by summing the recursion tree level by level.
\[ T(n) = n\log_2 n \ \Rightarrow\ T(8) = 8\cdot 3 = 24\ \checkmark \]
Picture it
Animation
Shows: Merging two sorted halves — a rendered Manim animation.
Rendered with Manim.
Takeaway: Linear work per level, and there are log n levels.
Concept
The move: #11 (Unroll and sum the levels), then #4 (Replace with the dominant term).
The watershed is the leaves' total work
Why: n to the log base b of a counts the work at the bottom of the tree. Comparing f(n) against it is asking whether the root or the leaves dominate.
\[ n^{\log_{2} 2} = n^{1} = n = f(n) \]
Case 2 means neither dominates, so every level costs the same
Why: That is precisely the merge sort tree you drew by hand: n per level, log n levels. The theorem did move #11 for you and cached the answer.
This is why no new toolkit move was added today. The theorem is a shortcut for a move you already own — and the moment a recurrence does not fit its shape, you are back to unrolling by hand.
Translation
\( n^{\log_{2} 2} = n^{1} = n = f(n) \)
Draw it
Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.
Picture it
Animation
Shows: When the Master Theorem does not apply — a rendered Manim animation.
Rendered with Manim.
Takeaway: The theorem needs subproblems of the same size.
Ranking
Put in order
Put the moves of Apply Case 2 to binary search into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Only one recursive call is made (the other half of the range is simply discarded), on half the remaining size, with one unit of extra comparison work.
Worked example
Binary search repeatedly cuts the remaining range in half, doing a constant amount of extra work -- one comparison -- at each step.
\[ T(n) = T\!\left(\frac{n}{2}\right) + 1, \qquad T(1)=1 \]
Read off a, b, and f(n)
Why: Only one recursive call is made (the other half of the range is simply discarded), on half the remaining size, with one unit of extra comparison work.
\[ a=1,\ b=2,\ f(n)=1 \]
Compute the watershed
Why: Any base-b logarithm of one is zero, since raising b to the power zero always gives one.
\[ \log_2 1 = 0 \ \Rightarrow\ \text{watershed}(n) = n^0 = 1 \]
Compare f(n) to the watershed
Why: The extra work, one comparison, exactly matches the watershed, one -- so this is Case 2, and the levels balance rather than the leaves or the root dominating.
\[ f(n) = \Theta(1) = \Theta(n^{\log_2 1}) \ \Rightarrow\ T(n) = \Theta(\log n) \]
Verify against the exact closed form
Why: Solving directly gives the base-two logarithm of n plus one, which matches every value computed by hand: at n equals sixteen, the formula gives 5, exactly matching direct recursive computation.
\[ T(n) = \log_2 n + 1 \ \Rightarrow\ T(16) = 4+1 = 5\ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Apply Case 2 to binary search", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Solving directly gives the base-two logarithm of n plus one, which matches every value computed by hand: at n equals sixteen, the formula gives 5, exactly matching direct recursive computation.
Concept
If the extra work is polynomially bigger than the watershed, then the single largest cost in the whole tree is the extra work done right at the root -- everything below it is comparatively negligible.
\[ f(n) = \Omega\!\left(n^{\log_b a + \varepsilon}\right) \ \text{for some}\ \varepsilon>0 \]
This case needs one more technical condition -- roughly, that the extra work does not blow up too fast when fed a smaller input -- but when it holds, the running time is just the extra work function itself.
\[ \text{(regularity)}\ a\,f\!\left(\frac{n}{b}\right) \le c\,f(n)\ \text{for some}\ c<1 \ \Rightarrow\ T(n)=\Theta(f(n)) \]
Step zero
Discussion prompt
Apply Case 3 — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Read off a, b, and f(n), and compute the watershed
Answer:
Worked example
Solve the recurrence below.
\[ T(n) = 2\,T\!\left(\frac{n}{2}\right) + n^2, \qquad T(1)=1 \]
Read off a, b, and f(n), and compute the watershed
Why: Two recursive calls on half-sized subproblems, watershed n to the base-two-log-of-two, which is n to the first power.
\[ a=2,\ b=2 \ \Rightarrow\ \text{watershed}(n) = n \]
Compare f(n) to the watershed
Why: The extra work n squared is polynomially bigger than the watershed n -- bigger by a full extra power -- so this is Case 3.
\[ n^2 = \Omega(n^{1+1}) \ \Rightarrow\ T(n) = \Theta(n^2) \]
Verify against the exact closed form
Why: Solving directly gives two n squared minus n, matching every value computed by hand: at n equals sixteen, the formula gives 496, exactly matching direct recursive computation.
\[ T(n) = 2n^2 - n \ \Rightarrow\ T(16) = 2(256)-16 = 496\ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Apply Case 3", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Solving directly gives two n squared minus n, matching every value computed by hand: at n equals sixteen, the formula gives 496, exactly matching direct recursive computation.
Step zero
Discussion prompt
Apply Case 3 again, with a=1 — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Read off a, b, and f(n), and compute the watershed
Answer:
Worked example
The leading coefficient does not need to be big for the root to win. Solve the recurrence below.
\[ T(n) = T\!\left(\frac{n}{2}\right) + n^2, \qquad T(1)=1 \]
Read off a, b, and f(n), and compute the watershed
Why: Only one recursive call is made, so the watershed is n to the base-two-log-of-one, which is n to the power zero -- just one.
\[ a=1,\ b=2 \ \Rightarrow\ \text{watershed}(n) = n^0 = 1 \]
Compare f(n) to the watershed
Why: The extra work n squared is polynomially bigger than the watershed 1 for any positive exponent, so this is Case 3.
\[ n^2 = \Omega(n^{0+1}) \ \Rightarrow\ T(n) = \Theta(n^2) \]
Verify against the exact closed form
Why: Solving directly gives four-thirds n squared minus one-third, matching every value computed by hand: at n equals sixteen, the formula gives 341, exactly matching direct recursive computation.
\[ T(n) = \tfrac{4}{3}n^2 - \tfrac{1}{3} \ \Rightarrow\ T(16) = \tfrac{4}{3}(256) - \tfrac{1}{3} = 341\ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Apply Case 3 again, with a=1", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Solving directly gives four-thirds n squared minus one-third, matching every value computed by hand: at n equals sixteen, the formula gives 341, exactly matching direct recursive computation.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Given the recurrence below, a student correctly computes the watershed but then misjudges the comparison, assuming a large-looking extra-work function must automatically dominate.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: This treats matching the watershed as the same thing as exceeding it, without actually checking whether the extra work is polynomially bigger.
Actually check the relationship: is the extra work polynomially bigger, polynomially smaller, or the same, within a constant factor, as the watershed?
Why: This treats matching the watershed as the same thing as exceeding it, without actually checking whether the extra work is polynomially bigger.
Trap
Given the recurrence below, a student correctly computes the watershed but then misjudges the comparison, assuming a large-looking extra-work function must automatically dominate.
\[ T(n) = 4\,T\!\left(\frac{n}{2}\right) + n^2, \qquad \text{watershed}(n) = n^2 \]
Reason "n squared is a big function, so it must dominate"
Why: This treats matching the watershed as the same thing as exceeding it, without actually checking whether the extra work is polynomially bigger.
\[ \text{(wrong)}\ \Rightarrow\ \text{Case 3} \ \Rightarrow\ T(n) = \Theta(n^2) \]
Actually check the relationship: is the extra work polynomially bigger, polynomially smaller, or the same, within a constant factor, as the watershed?
\[ T(n) = 4\,T\!\left(\frac{n}{2}\right) + n^2, \qquad \text{watershed}(n) = n^2 \]
Recognize f(n) exactly equals the watershed
Why: The extra work n squared is not bigger than the watershed n squared -- they are the same function, within a constant factor -- so the correct comparison is equality, which is Case 2.
\[ f(n) = \Theta(n^2) = \Theta(n^{\log_2 4}) \ \Rightarrow\ \text{Case 2} \]
Conclude with the log factor Case 2 requires
Why: Because the two sides balance rather than one dominating, an extra logarithmic factor appears that the wrong reasoning dropped entirely.
\[ \text{(right)}\quad T(n) = \Theta(n^2\log n) \]
Break the constraint
Discussion prompt
The rule this trap just fixed:
The extra work n squared is not bigger than the watershed n squared -- they are the same function, within a constant factor -- so the correct comparison is equality, which is Case 2.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
This treats matching the watershed as the same thing as exceeding it, without actually checking whether the extra work is polynomially bigger.
Concept
Every application of the Master Theorem comes down to filling in this table.
| Case | Relationship of f(n) to the watershed | Result |
|---|---|---|
| 1 | polynomially smaller (leaves win) | watershed function |
| 2 | matches within a constant factor (balanced) | watershed function times log n |
| 3 | polynomially bigger, plus regularity (root wins) | f(n) itself |
Comparison
Comparison matrix
From The three cases, side by side: refill the Relationship of f(n) to the watershed column from what you know. The rest of the table is as it appeared.
| Case | Relationship of f(n) to the watershed | Result |
|---|---|---|
| 1 | polynomially smaller (leaves win) | watershed function |
| 2 | matches within a constant factor (balanced) | watershed function times log n |
| 3 | polynomially bigger, plus regularity (root wins) | f(n) itself |
Section
Section 6
Intuition
What feels wrong about this?
Every ingredient looks like Case 3 — the extra work beats the watershed:
\[ T(n) = 2T(n/2) + n \log n \]
\[ \text{watershed: } \; n^{\log_{2} 2} = n, \qquad n \log n \; \text{grows faster than } \; n \]
_Plain English only. No notation, no algebra. Just say what bothers you._
The feeling: n log n is barely bigger than n. It beats n, but not by any actual power of n — it wins by a hair, not by a whole factor.
That feeling is the proof. It is not a substitute for the proof — it is the thing the proof writes down.
That hair is exactly what the Master Theorem's side condition is about. The cases need a polynomial gap, and log n is not a polynomial factor. So this recurrence falls between the cases and you solve it by hand.
Picture it
Animation
Shows: Why the depth is log n — a rendered Manim animation.
Rendered with Manim.
Takeaway: The same count that makes binary search fast makes merge sort log-deep.
Concept
Case 1 and Case 3 both require the extra work to differ from the watershed by a genuine polynomial gap -- a fixed positive exponent of difference, not merely "a little" smaller or bigger.
\[ \text{needs}\ n^{\log_b a - \varepsilon}\ \text{or}\ n^{\log_b a + \varepsilon}\ \text{for some fixed}\ \varepsilon>0 \]
A single logarithmic factor of difference is not a polynomial gap -- logarithms grow more slowly than any positive power of n, no matter how small that power is. When that is all that separates the two sides, none of the three cases fit.
Explain it
Discussion prompt
Explain Requirement 1: an actual polynomial gap to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Case 1 and Case 3 both require the extra work to differ from the watershed by a genuine polynomial gap -- a fixed positive exponent of difference, not merely "a little" smaller or bigger.
Fill the middle
Fill in the blanks
From Process: forcing a recurrence into a case — finish the line. Write what belongs on the right of the equals sign before you look.
T(n) = 2T(n/2) + n \log n
Why: Producing the right-hand side unprompted is the difference between recognising this line and being able to use it. n log n beats n. Case 3 says the answer is Theta of f(n), so write Theta of n log n and move on.
Intuition
Watch me not know the answer. This is what the first two minutes actually look like.
The gap recurrence again:
\[ T(n) = 2T(n/2) + n \log n \]
Try Case 3: the extra work is bigger than the watershed, so the root should win
Why: n log n beats n. Case 3 says the answer is Theta of f(n), so write Theta of n log n and move on.
The side condition fails, and so does the answer
Why: Case 3 needs f(n) to beat the watershed by a polynomial factor and needs the regularity condition. Neither holds. The answer it produced is also wrong — the real answer is Theta of n times log n squared.
Dead end. Not a mistake — a move that was worth trying and did not pay off. This happens in most proofs.
Back up. Unroll it by hand
Why: Level i has 2 to the i nodes, each doing n over 2 to the i times log of n over 2 to the i. That is about n log n per level near the top, over log n levels.
\[ T(n) = \Theta(n \log^{2} n) \]
The lesson is not memorize this recurrence. It is that step 5 of the skeleton is not optional — a case whose side condition fails gives a confidently wrong answer.
The expert does not see the whole path in advance. The expert tries something, reads the result, and adjusts. That is the skill.
Notation
Annotate
From Process: forcing a recurrence into a case — read this one piece at a time. What is each part doing?
On: \( T(n) = \Theta(n \log^{2} n) \)
Ranking
Put in order
Put the moves of A recurrence that fits no case cleanly into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. The ratio of the extra work to the watershed is the base-two logarithm of n, which grows without bound rather than staying near any fixed constant -- so f(n) is not Theta of the watershed.
Worked example
Attempt to classify the recurrence below.
\[ T(n) = 2\,T\!\left(\frac{n}{2}\right) + n\log n, \qquad \text{watershed}(n) = n \]
Test Case 2: does f(n) match the watershed within a constant factor?
Why: The ratio of the extra work to the watershed is the base-two logarithm of n, which grows without bound rather than staying near any fixed constant -- so f(n) is not Theta of the watershed. Case 2 fails.
\[ \frac{n\log n}{n} = \log n \to \infty \]
Test Case 3: is f(n) polynomially bigger?
Why: For Case 3, the extra work must beat n to the power one plus epsilon for some fixed positive epsilon. But n log n divided by n to that power shrinks to zero as n grows, for any fixed epsilon greater than zero -- a log factor alone never wins against a genuine extra power. Case 3 fails.
\[ \frac{n\log n}{n^{1+\varepsilon}} = \frac{\log n}{n^{\varepsilon}} \to 0 \]
Check Case 1: is f(n) polynomially smaller?
Why: Case 1 requires the extra work to be smaller than the watershed, but n log n is actually bigger than n, not smaller -- Case 1 fails immediately for the opposite reason.
\[ n\log n \ \text{is bigger than}\ n,\ \text{not smaller} \]
Check the conclusion: none of the three cases apply
Why: The extra work sits exactly in the gap just above Case 2 and just below Case 3 -- bigger than the watershed by only a logarithmic factor, never by a full polynomial one. The Master Theorem gives no answer here; a different technique is needed.
Picture it
Animation
Shows: Each line of the worked example "A recurrence that fits no case cleanly", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Case 1 requires the extra work to be smaller than the watershed, but n log n is actually bigger than n, not smaller -- Case 1 fails immediately for the opposite reason.
Concept
The Master Theorem's formula assumes a single fixed shrink factor b that applies to every one of the a recursive calls. If different calls shrink the input by different amounts, there is no single b to plug in.
\[ T(n) = a\,T\!\left(\frac{n}{b}\right) + f(n) \quad \text{requires ONE shared } b \text{ for every call} \]
Anomaly
Predict first
A student writes this, and it looks reasonable:
Given the recurrence below, with subproblems of two different sizes, a student forces it into the Master Theorem's form anyway, using just one of the two fractions as if it were the whole story.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: This treats the recurrence as if it were two equal calls of size n over three, silently discarding that the second call is actually on two-thirds of n, a much larger subproblem.
Recognize first that the two recursive calls shrink the input by different factors -- one third, and two thirds -- so this recurrence simply does not have the aT(n over b) shape at all.
Why: This treats the recurrence as if it were two equal calls of size n over three, silently discarding that the second call is actually on two-thirds of n, a much larger subproblem.
Trap
Given the recurrence below, with subproblems of two different sizes, a student forces it into the Master Theorem's form anyway, using just one of the two fractions as if it were the whole story.
\[ T(n) = T\!\left(\frac{n}{3}\right) + T\!\left(\frac{2n}{3}\right) + n \]
Force-fit using only the n over 3 term
Why: This treats the recurrence as if it were two equal calls of size n over three, silently discarding that the second call is actually on two-thirds of n, a much larger subproblem.
\[ \text{(wrong)}\ a=2,\ b=3 \ \Rightarrow\ \text{Case 1} \ \Rightarrow\ T(n)=\Theta(n^{\log_3 2}) \]
Recognize first that the two recursive calls shrink the input by different factors -- one third, and two thirds -- so this recurrence simply does not have the aT(n over b) shape at all.
\[ T(n) = T\!\left(\frac{n}{3}\right) + T\!\left(\frac{2n}{3}\right) + n \]
Refuse to force it into the Master Theorem's mold
Why: No single b describes both calls at once, no matter which values of a and b are chosen -- the Master Theorem's hypotheses are simply not met here.
Reach for a different tool instead
Why: A recursion-tree argument still works: the shallowest possible path bottoms out after a logarithmic number of levels, the deepest possible path also bottoms out after a logarithmic number of levels, and every fully-populated level still totals n units of work -- so the total across the whole tree still comes out on the order of n times the logarithm of n.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
Concept
When the Master Theorem's hypotheses are not met -- a gap between the cases, or subproblems of unequal size -- the recurrence usually is not unsolvable. It just needs a different technique.
Two of the most common fallbacks: build the recursion tree directly and bound the total work level by level, exactly as done for merge sort by hand; or guess a closed-form answer and confirm it by mathematical induction. Generalizations of the Master Theorem that handle unequal-sized subproblems exist too, for anyone who wants to go further.
Analogy
Discussion prompt
Explain What to reach for instead by analogy to something with no CS3000 Algorithms in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
When the Master Theorem's hypotheses are not met -- a gap between the cases, or subproblems of unequal size -- the recurrence usually is not unsolvable. It just needs a different technique.
Pattern
1. Confirm the recurrence has the right shape
Why: Every recursive call must be on a subproblem of the same fractional size. If not, stop -- the Master Theorem does not apply, and a different method is needed.
2. Read off a, b, and f(n) directly
Why: a is the coefficient in front of T (how many calls); b is the number dividing n inside T (how much smaller each subproblem is); f(n) is everything left over outside the recursive calls.
3. Compute the watershed function
Why: Take the base-b logarithm of a, and raise n to that power. This is the yardstick everything else gets compared against.
4. Compare f(n) to the watershed, and check for a genuine polynomial gap
Why: Polynomially smaller means Case 1; matching within a constant factor means Case 2; polynomially bigger (plus regularity) means Case 3. Anything less than a true polynomial gap means none of the cases apply.
5. Read off the closed form and sanity-check it
Why: State T(n) in the form the case gives, and, where possible, check it against a small direct computation the way every worked example in this deck did.
Real world
Discussion prompt
Outside this lesson: where does Merge Sort & the Master Theorem actually turn up? Name one concrete situation — a job, a piece of software someone ships, a decision somebody has to make — and say which part of The Master Theorem decision procedure is doing the work in it.
Hint: Vague is the failure mode here. "Engineering" is not a situation; "deciding whether this build is fast enough to ship" is.
Answer:
That deck traces merge sort end to end, explains why the merge step is linear, sets up the merge sort recurrence, and works the Master Theorem's three cases by comparing the extra work f(n) against the watershed function n^log_b(a). It targets swapping a and b, mis-comparing f(n) to the watershed, assuming the merge step is free, and force-fitting recurrences - those with log-factor gaps or unequal splits - that do not satisfy the theorem's hypotheses.
Picture it
Animation
Shows: Merge sort through the Master Theorem — a rendered Manim animation.
Rendered with Manim.
Takeaway: The balanced case, which is why the log appears at all.
Check
Consider the recurrence below.
\[ T(n) = 5\,T\!\left(\frac{n}{4}\right) + n^2 \]
Check your understanding
What are a and b for this recurrence?
Answer: A
Why: In T(n) = aT(n/b) + f(n), a is the number of recursive calls (5 here) and b is the number the input size is divided by (4 here). Reading them off directly from the recurrence gives a = 5 and b = 4.
Elimination
Eliminate the wrong options
What is the watershed function here, and which case applies?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Here a = 4 and b = 2, so the base-two logarithm of 4 is 2, giving watershed n squared. Since f(n) = n cubed is polynomially larger than n squared (it beats n to the power two plus one), and the regularity condition holds, this is Case 3: the root's own work dominates, so T(n) = Theta(n cubed).
Check
Consider the recurrence below.
\[ T(n) = 4\,T\!\left(\frac{n}{2}\right) + n^3 \]
Check your understanding
What is the watershed function here, and which case applies?
Answer: A
Why: Here a = 4 and b = 2, so the base-two logarithm of 4 is 2, giving watershed n squared. Since f(n) = n cubed is polynomially larger than n squared (it beats n to the power two plus one), and the regularity condition holds, this is Case 3: the root's own work dominates, so T(n) = Theta(n cubed).
Check
Consider this recurrence, which describes an algorithm making seven recursive calls on half-sized subproblems, plus a squared amount of extra combining work.
\[ T(n) = 7\,T\!\left(\frac{n}{2}\right) + n^2 \]
Check your understanding
Which case applies, and what is T(n)?
Answer: A
Why: Here a = 7 and b = 2, so the watershed is n to the base-two logarithm of 7, approximately n to the 2.81 power. f(n) = n squared is polynomially smaller than that (n squared is smaller by roughly n to the 0.8 power), so this is Case 1, where the leaves dominate: T(n) = Theta of the watershed itself, not Theta of f(n).
Check
Consider the recurrence below.
\[ T(n) = 4\,T\!\left(\frac{n}{2}\right) + n^2\log n \]
Check your understanding
Which case of the Master Theorem applies?
Answer: A
Why: The watershed is n squared (since the base-two logarithm of 4 is 2). Case 2 needs f(n) to be Theta of n squared exactly, but n squared log n grows unboundedly faster than n squared -- their ratio is log n, not a constant. Case 3 needs f(n) to be polynomially larger, by some fixed positive exponent, but n squared log n is only a log factor larger, not a polynomial factor -- so Case 3 fails too. None of the three cases apply.
Prediction
Predict first
Can the Master Theorem be applied directly to this recurrence?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: No -- the two recursive calls shrink the input by different factors, so there is no single b.
Why: The Master Theorem's form requires every recursive call to shrink the input by the SAME factor b. Here one call uses n over 4 and the other uses 3n over 4 -- two different shrink factors -- so the recurrence does not have the required form, no matter how a and b are chosen.
Check
Consider the recurrence below.
\[ T(n) = T\!\left(\frac{n}{4}\right) + T\!\left(\frac{3n}{4}\right) + n \]
Check your understanding
Can the Master Theorem be applied directly to this recurrence?
Answer: A
Why: The Master Theorem's form requires every recursive call to shrink the input by the SAME factor b. Here one call uses n over 4 and the other uses 3n over 4 -- two different shrink factors -- so the recurrence does not have the required form, no matter how a and b are chosen.
Elimination
Eliminate the wrong options
What is the watershed function n to the base-b-log-of-a for this recurrence?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: The watershed function is n raised to the base-b logarithm of a. With a = 8 and b = 2: the base-two logarithm of 8 asks '2 raised to what power gives 8?' -- the answer is 3, since 2 cubed is 8. So the watershed is n cubed.
Check
Consider a recurrence with a equal to 8 and b equal to 2.
Check your understanding
What is the watershed function n to the base-b-log-of-a for this recurrence?
Answer: A
Why: The watershed function is n raised to the base-b logarithm of a. With a = 8 and b = 2: the base-two logarithm of 8 asks '2 raised to what power gives 8?' -- the answer is 3, since 2 cubed is 8. So the watershed is n cubed.
Concept
Moves added today: none.
That is a result, not a gap. Everything in this lesson was proved with moves you already owned.
Moves you reused today:
The Master Theorem is not a new idea — it is move #11 pre-computed for every recurrence of one particular shape. When a recurrence falls outside that shape, you go back to unrolling by hand.
Full toolkit so far: #1 through #11.
Next session opens with you naming every one of these from memory, before any new material.
Counterexample
Discussion prompt
The Master Theorem is not a new idea — it is move #11 pre-computed for every recurrence of one particular shape. When a recurrence falls outside that shape, you go back to unrolling by hand.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Next session opens with you naming every one of these from memory, before any new material.
Picture it
Animation
Shows: On a linked list, merge sort wins outright — a rendered Manim animation.
Rendered with Manim.
Takeaway: Quicksort's partition wants random access; merge does not.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Merge Sort, Traced · Why Merging Is Linear · The Merge Sort Recurrence · The Master Theorem: General Form · The Three Cases · When It Doesn't Apply. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
Merge sort divides an array in half, recursively sorts each half, and merges the two sorted halves back together in an amount of time proportional to their combined size -- never in constant time.
| Situation | What to do |
|---|---|
| a and b | a is the call count; b is the shrink factor -- never swap them |
| f(n) vs watershed | smaller / equal / bigger decides Case 1, 2, or 3 |
| log-only gap or unequal sizes | Master Theorem does not apply -- use a recursion tree instead |
Want this taught 1-on-1? Alexander tutors CS3000 Algorithms — $55/session, free consultation.