This deck explains what a minimum spanning tree is, then gives the cut property and the cycle property that make greedy MST algorithms provably correct. It includes full hand-traced runs of Prim's and Kruskal's algorithms, the latter with union-find, along with their running times. It targets the misconceptions that an MST is always unique, that it is the same thing as a shortest-path tree, that Kruskal never needs a cycle check, and that union-find stores weights or paths rather than connectivity alone.
Subject: CS3000 Algorithms · 125 slides · symbolic lesson
Open the interactive version of this deck · Homework for this lesson
Objectives
Minimum spanning trees show up whenever you need to connect every point in a network as cheaply as possible - power lines, road networks, computer clusters. By the end of this lesson you can:
Warm-up
Discussion prompt
Before we open Minimum Spanning Trees: Prim & Kruskal: without looking back, what was the main idea of Shortest Paths: Dijkstra & Bellman-Ford, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
That deck builds single-source shortest paths from one shared primitive, edge relaxation. It traces Dijkstra's algorithm by hand with a priority queue and explains why its greedy choice is safe only for non-negative weights, then covers Bellman-Ford's V-1 rounds of relaxation and its negative-cycle detection. It targets the traps of trusting Dijkstra with a negative edge, relaxing in the wrong direction, misjudging why V-1 rounds are needed, and reviving a vertex that has already been settled.
Concept
Before any new material: cover the screen.
You have named 15 reusable moves so far. Say as many as you can out loud, by number, from memory.
Do not advance until you have actually tried. Getting four of eight is information; skipping the exercise is not.
Here they are. Score yourself.
Today adds no new moves. Every proof in this lesson is built out of the list above. That is the whole point of the list.
The question that starts every proof from here on is not how do I begin. It is which of these applies here?
Counterexample
Discussion prompt
You have named 15 reusable moves so far. Say as many as you can out loud, by number, from memory.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Do not advance until you have actually tried. Getting four of eight is information; skipping the exercise is not.
Section
Section 1
Concept
A graph is a collection of dots called vertices, connected by lines called edges. In a network of towns, each town is a vertex and each road between two towns is an edge.
vertex — A single point in a graph, such as one town, one computer, or one sensor.
edge — A connection between exactly two vertices, such as one road, one cable, or one wireless link.
Analogy
Discussion prompt
Explain Graph vocabulary: vertices and edges by analogy to something with no CS3000 Algorithms in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
A graph is a collection of dots called vertices, connected by lines called edges. In a network of towns, each town is a vertex and each road between two towns is an edge.
Concept
In a weighted graph, every edge carries a number called its weight - the cost of using that connection. It could be distance, price, or delay.
weight — The number attached to an edge, representing whatever quantity you are trying to minimize: the length of a cable, the price of a road, or the delay of a link.
Explain it
Discussion prompt
Explain Weighted edges represent cost to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
In a weighted graph, every edge carries a number called its weight - the cost of using that connection. It could be distance, price, or delay.
Intuition
Picture a map of small towns joined by roads, and each road has a toll you must pay to drive it. Some roads are cheap shortcuts, others are expensive detours.
Every algorithm in this lesson is really just answering one question: which roads should you build so every town is reachable, while paying as little total toll as possible?
Socratic
Discussion prompt
Picture a map of small towns joined by roads, and each road has a toll you must pay to drive it. Some roads are cheap shortcuts, others are expensive detours.
Suppose that were not true. What is the first thing in Minimum Spanning Trees: Prim & Kruskal that would stop working?
Hint: Follow it one step downstream. The answer is whatever was quietly relying on it.
Answer:
Every algorithm in this lesson is really just answering one question: which roads should you build so every town is reachable, while paying as little total toll as possible?
Concept
A graph is connected if you can get from any vertex to any other vertex by following some sequence of edges. If some vertex is stranded with no path to the rest, the graph is disconnected.
Every algorithm in this lesson assumes the input graph is connected - otherwise there is no way to reach every vertex at all, spanning or otherwise.
Concept
A spanning tree of a graph is a subset of its edges that touches every single vertex, contains no cycles, and stays connected.
spanning — Touching every vertex in the graph - none left out.
cycle — A path that starts and ends at the same vertex without reusing an edge. A tree, by definition, has none.
Matching
Match the pairs
Match each term to the definition this lesson gave it — not the one you would guess from the word.
Why: These are the working definitions of vertex, edge, weight, spanning, cycle as Minimum Spanning Trees: Prim & Kruskal uses them. Pairing them correctly is the test of whether you could state each one with the slide switched off.
Picture it
Animation
Shows: An MST is not a shortest-path tree — a rendered Manim animation.
Rendered with Manim.
Takeaway: Different objectives, and usually different trees on the same graph.
Intuition
Picture stripping a tangled web of roads down to the bare minimum needed to still reach every town - no loops, no redundant connections, just the skeleton holding everything together.
Any extra edge beyond that skeleton would create a cycle: a second way to get between two towns that you simply do not need.
Concept
A tree connecting some number of vertices always uses exactly one fewer edge than the number of vertices. One edge less would leave something unreachable; one edge more would create a cycle.
\[ |V| = n \ \Rightarrow\ |\text{tree edges}| = n - 1 \]
Every spanning tree of a graph with six vertices, for instance, has exactly five edges - never more, never fewer.
Concept
A minimum spanning tree, or MST, is a spanning tree whose edges add up to the smallest possible total weight, out of every spanning tree the graph has.
minimum spanning tree (MST) — A spanning tree of a weighted graph with the least possible sum of edge weights among all spanning trees of that graph.
Intuition
If every road has a toll, the MST is the specific set of roads you would build to connect every town while paying the smallest possible total toll - no detours, no redundant links, and no missed towns.
Two different sets of roads can connect the same towns; the MST is whichever connected, cycle-free set costs the least.
Ranking
Put in order
Put the moves of Verify a spanning tree is minimal into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. With 3 vertices, a spanning tree needs exactly 2 edges.
Worked example
Three sensors P, Q, and R can each be wired directly to the other two. The cable costs are given below.
\[ w(P,Q)=3,\quad w(Q,R)=4,\quad w(P,R)=5 \]
List every possible spanning tree
Why: With 3 vertices, a spanning tree needs exactly 2 edges. There are three ways to choose 2 of the 3 available cables.
| Tree | Edges used | Total weight |
|---|---|---|
| 1 | P-Q, Q-R | 7 |
| 2 | P-Q, P-R | 8 |
| 3 | Q-R, P-R | 9 |
Identify the minimum
Why: Comparing the three totals, tree 1 (P-Q and Q-R) has the smallest sum, so it is the MST; the excluded cable P-R is the most expensive one, consistent with it being left out.
\[ 7 \ <\ 8 \ <\ 9\ \Rightarrow\ \text{MST} = \{P\text{-}Q,\ Q\text{-}R\} \]
Verify by re-adding the excluded edge
Why: Adding P-R back to tree 1 would create the full triangle, a cycle - confirming that a genuine spanning tree cannot include it once the other two edges already connect all three sensors.
\[ P\text{-}Q \ +\ Q\text{-}R \ +\ P\text{-}R \ \Rightarrow\ \text{cycle, not a tree} \]
Intuition
What feels wrong about this?
Four vertices in a square, every edge of weight 1.
\[ \text{every spanning tree has total weight } 3 \]
_Plain English only. No notation, no algebra. Just say what bothers you._
The feeling: there are several different trees and they all cost the same, so there is no single right answer to point at.
That feeling is the proof. It is not a substitute for the proof — it is the thing the proof writes down.
So minimum describes the weight, which is unique, not the tree, which need not be. Watch where this bites: any proof step that says the MST contains this edge is wrong unless you show every optimal tree does.
Trap
A student assumes every graph has exactly one minimum spanning tree, the way a math problem usually has one right answer.
Figure (svg): A four-vertex square graph A, B, C, D with all four edges — A-B, B-C, C-D, and D-A — having equal weight 1, and no diagonal edges.
Assume there is a single correct MST to find
Why: But this graph has four equally-weighted edges forming a loop. Assuming uniqueness would make a student think one specific set of three edges is 'the' answer, when several tie.
\[ w(A,B)=w(B,C)=w(C,D)=w(D,A)=1 \]
Ties in edge weight mean ties in which spanning tree is minimum. An MST is unique only when every edge weight in the graph is distinct.
List the tied MSTs
Why: Any 3 of these 4 equal-weight edges connects all four vertices with no cycle, and every such tree weighs 3. Dropping any one of the four edges gives a valid, equally minimal spanning tree.
| Dropped edge | Remaining tree edges | Total weight |
|---|---|---|
| A-B | B-C, C-D, D-A | 3 |
| B-C | A-B, C-D, D-A | 3 |
| C-D | A-B, B-C, D-A | 3 |
| D-A | A-B, B-C, C-D | 3 |
State the correct rule
Why: The MST's total weight is always unique, but the specific edges achieving it are guaranteed unique only when no two edges share a weight. With ties, multiple different edge sets can all be minimum.
Invariant
Step through it
Step through Trap: assuming the MST is always unique one row at a time. One of these columns never changes — find it, and say why it cannot.
Picture it
Figure (svg): A four-vertex graph with A at top, B at left, C at bottom, D at right. Edges A-B, B-C, C-D each weight 1; edges A-C and A-D each weight 2.
Discussion prompt
Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.
Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.
Answer:
A student computes the shortest-path tree from vertex A (using Dijkstra's algorithm) and assumes it must be the same as the MST, since both seem to 'connect everything cheaply'.
Trap
A student computes the shortest-path tree from vertex A (using Dijkstra's algorithm) and assumes it must be the same as the MST, since both seem to 'connect everything cheaply'.
Figure (svg): A four-vertex graph with A at top, B at left, C at bottom, D at right. Edges A-B, B-C, C-D each weight 1; edges A-C and A-D each weight 2.
Build the shortest-path tree from A
Why: The shortest distance from A to C is 2 (direct, or via B - both give 2); the shortest distance from A to D is 2 using the direct edge, since going through C would cost 2 plus 1 more, which is longer.
\[ \text{dist}(A,D) = \min(2,\ 2+1) = 2 \ \Rightarrow\ \text{use edge } A\text{-}D \]
This tree costs more than necessary
Why: The shortest-path tree total is A-B plus B-C plus A-D, equal to 4. It optimizes distance FROM A to every vertex, not total cable cost - so it is not guaranteed to be an MST.
\[ 1+1+2 = 4 \]
The MST ignores distance from any single source and just minimizes total edge weight.
Figure (svg): The same four-vertex graph with A at top, B at left, C at bottom, D at right. Edges A-B, B-C, C-D each weight 1; edges A-C and A-D each weight 2.
Run Kruskal's algorithm instead
Why: Sorted edges: A-B(1), B-C(1), C-D(1), A-C(2), A-D(2). Adding the three cheapest edges that avoid a cycle connects all four vertices for a lower total.
\[ A\text{-}B(1) + B\text{-}C(1) + C\text{-}D(1) = 3 \]
Verify the two trees are genuinely different
Why: The MST uses C-D (weight 1) to reach D, while the shortest-path tree from A used A-D (weight 2) instead - a real difference in which edges are chosen, and the MST's total of 3 is cheaper than the shortest-path tree's total of 4.
\[ \text{MST total} = 3 \ <\ \text{shortest-path tree total} = 4 \]
Notation
Annotate
From Trap: an MST is not a shortest-path tree — read this one piece at a time. What is each part doing?
On: \( \text{dist}(A,D) = \min(2,\ 2+1) = 2 \ \Rightarrow\ \text{use edge } A\text{-}D \)
Section
Section 2
Concept
A cut splits every vertex in the graph into exactly two non-empty groups. It does not remove anything - it is just a way of dividing the vertices into one side and the other side.
cut — A partition of the graph's vertices into two non-empty sets, commonly called one side and the other side.
Intuition
Imagine drawing a line across your map of towns, sorting every town into camp one or camp two. The line itself does not touch any road - it only decides which camp each town falls into.
Different lines produce different cuts. There is nothing special about any one cut; the cut property will hold for every single one you could draw.
Concept
Once vertices are sorted into two sides of a cut, an edge is called a crossing edge if it has one endpoint on each side.
crossing edge — An edge with one endpoint in each of the cut's two groups. Edges with both endpoints on the same side do not cross the cut.
A cut always has at least one crossing edge, as long as the whole graph stays connected - otherwise the two sides would already be unreachable from each other.
Concept
This is the single most important fact in this whole lesson. For any cut of the graph, the cheapest crossing edge is guaranteed to belong to some minimum spanning tree.
cut property — For any cut of a connected weighted graph, the minimum-weight edge crossing that cut belongs to at least one MST of the graph.
This holds for every cut you could possibly draw - there is no special cut required. That freedom is exactly what makes greedy algorithms provably correct.
Picture it
Animation
Shows: The cut property is why greedy works here — a rendered Manim animation.
Rendered with Manim.
Takeaway: Every greedy choice is provably safe, which is rare and worth noticing.
Concept
Every proof of this kind has the same five or six moves in the same order. The order is not something you rediscover each time.
It is on the right. It will stay on the right through the worked examples that follow.
Why this matters: the structure is now handled. You are not spending working memory on what comes next — you are spending all of it on the one hard step.
Step 4 is the only creative step. Adding an edge to a tree always makes exactly one cycle, and that cycle is what hands you a legal edge to remove — you are never searching the whole graph.
Intuition
Here is the move you will make in every correctness argument this lesson needs: pick any cut such that the edges chosen so far all stay on one side or the other - never crossing it themselves.
Once you have such a cut, the cheapest edge that crosses it is always safe to add. That single move is the engine behind both Prim's and Kruskal's algorithms.
Intuition
What move should we make next?
Split the vertices into two camps. Look at the cheapest edge crossing between them.
\[ e = \text{minimum-weight edge crossing the cut} \]
\[ \text{claim: some minimum spanning tree contains } e \]
You proved a claim of exactly this shape six lessons ago, about intervals. Name the move and say what object you assume exists.
_Look at your toolkit. Say a move number out loud before this slide advances._ A wrong guess is useful. A silent guess is not.
Picture it
Figure (svg): A four-vertex graph split by a dashed cut line into side S containing vertices 1 and 2, and the other side containing vertices 3 and 4. Crossing edges are 1-3 weight 3, 2-3 weight 1, and 2-4 weight 4; non-crossing edges are 1-2 weight 2 and 3-4 weight 5.
Discussion prompt
Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.
Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.
Answer:
Take a graph on four vertices, and a cut splitting side S, holding vertices 1 and 2, from the other side, holding vertices 3 and 4.
Worked example
Take a graph on four vertices, and a cut splitting side S, holding vertices 1 and 2, from the other side, holding vertices 3 and 4.
Figure (svg): A four-vertex graph split by a dashed cut line into side S containing vertices 1 and 2, and the other side containing vertices 3 and 4. Crossing edges are 1-3 weight 3, 2-3 weight 1, and 2-4 weight 4; non-crossing edges are 1-2 weight 2 and 3-4 weight 5.
\[ \text{crossing edges: } (1,3)=3,\ (2,3)=1,\ (2,4)=4 \]
Identify the cheapest crossing edge
Why: Among the three crossing edges, edge (2,3) has weight 1, strictly less than the others. Call it e; the cut property claims e belongs to some MST.
\[ e = (2,3),\quad w(e)=1 \]
Assume for contradiction that some MST T excludes e
Why: Since T is a spanning tree, it must connect vertex 2's side to vertex 3's side somehow - so T contains at least one crossing edge of this cut. Say T uses edge (1,3) instead, since T excludes (2,3).
\[ f=(1,3) \in T,\quad w(f)=3 \]
Add e to T and watch a cycle appear
Why: T is already a spanning tree, so adding any new edge to it creates exactly one cycle. Adding e = (2,3) creates a cycle that runs through the path already in T connecting 2 and 3, and that path passes through f = (1,3).
\[ T \cup \{e\} \ \text{contains a cycle through } f \]
Swap e in for f
Why: Removing f from that cycle breaks the cycle while keeping every vertex connected - the result, T with f removed and e added, is still a spanning tree.
\[ T' = (T \setminus \{f\}) \cup \{e\} \]
Verify T' is cheaper, contradicting T's minimality
Why: T' weighs exactly the weight of f minus the weight of e less than T, which is 3 minus 1 equals 2 less. Since T was assumed minimum, no spanning tree can weigh less than T - a contradiction. So the assumption was false: e must belong to every MST after all.
\[ w(T') = w(T) - 3 + 1 = w(T) - 2\ <\ w(T)\ \checkmark \]
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Verify T' is cheaper, contradicting T's minimality
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Take a graph on four vertices, and a cut splitting side S, holding vertices 1 and 2, from the other side, holding vertices 3 and 4.
Explain it to yourself
Discussion prompt
In The same move, new subject this move is made:
Add your edge and get exactly one cycle
Why is that legal? Name the rule or definition it rests on before you read on.
Hint: If you can only say "because that is what you do", the rule is the thing to go and find.
Answer:
A spanning tree plus any edge has precisely one cycle. That cycle must cross the cut a second time, since your edge crossed it once and a cycle crosses any cut an even number of times.
Concept
The move: #14 (Exchange argument), then #7 (Negate and assume).
Assume a minimum spanning tree that does not contain your edge
Why: Move #7 supplies the rival. Without an assumed optimal tree there is nothing to swap into.
Add your edge and get exactly one cycle
Why: A spanning tree plus any edge has precisely one cycle. That cycle must cross the cut a second time, since your edge crossed it once and a cycle crosses any cut an even number of times.
Remove that other crossing edge
Why: It is at least as heavy as yours, because yours was the minimum crossing edge. Removing it restores a spanning tree.
\[ w(e) \le w(e') \;\Longrightarrow\; \text{new total} \le \text{old total} \]
That is the identical five-step shape as the interval-scheduling exchange. Prim and Kruskal are both just repeated applications of this one lemma — which is why proving them separately is unnecessary.
Concept
The cycle property is the cut property's mirror image. For any cycle in the graph, the heaviest edge on that cycle can be safely thrown away - it never belongs to any minimum spanning tree.
cycle property — For any cycle in a connected weighted graph, the maximum-weight edge on that cycle belongs to no MST of the graph, provided that edge's weight is strictly the largest on the cycle.
Picture it
Animation
Shows: And the cycle property is why edges get rejected — a rendered Manim animation.
Rendered with Manim.
Takeaway: Kruskal's rejection step is this property, applied edge by edge.
Intuition
A cycle always offers more than one way to get between two of its vertices. The most expensive edge on that loop is never necessary - you could always take the rest of the loop instead, for less.
So whenever you spot a cycle, you already know one useless edge: the priciest one on it. That is exactly the edge to leave out of your minimum spanning tree.
Intuition
What move should we make next?
Take any cycle in the graph and its uniquely heaviest edge.
\[ \text{claim: no minimum spanning tree contains that edge} \]
This is the mirror image of the cut property, and it needs the mirror-image argument.
Same move again — but the swap runs the other way. Which edge do you remove first this time, and which do you add?
_Look at your toolkit. Say a move number out loud before this slide advances._ A wrong guess is useful. A silent guess is not.
Picture it
Figure (svg): A triangle with vertices X, Y, Z. Edge X-Y weight 2, edge Y-Z weight 3, edge X-Z weight 7, highlighted red as the heaviest edge on the cycle.
Discussion prompt
Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.
Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.
Answer:
Three vertices X, Y, and Z are mutually connected, forming one cycle.
Worked example
Three vertices X, Y, and Z are mutually connected, forming one cycle.
Figure (svg): A triangle with vertices X, Y, Z. Edge X-Y weight 2, edge Y-Z weight 3, edge X-Z weight 7, highlighted red as the heaviest edge on the cycle.
\[ w(X,Y)=2,\quad w(Y,Z)=3,\quad w(X,Z)=7 \]
Identify the heaviest edge on the cycle
Why: Comparing the three weights, X-Z at 7 is strictly the largest. Call it edge e; the cycle property claims e belongs to no MST.
\[ e=(X,Z),\quad w(e)=7 \]
Assume for contradiction that some MST T contains e
Why: A spanning tree on 3 vertices has exactly 2 edges. If T contains e = X-Z, its one other edge must be either X-Y or Y-Z.
\[ T = \{(X,Z),\ (X,Y)\}\quad \text{or}\quad T=\{(X,Z),\ (Y,Z)\} \]
Remove e and look at what remains
Why: The other two cycle edges, X-Y and Y-Z together, already reconnect X to Z without using e at all - so swapping e out for whichever of X-Y or Y-Z is missing from T restores a spanning tree.
\[ T' = \{(X,Y),\ (Y,Z)\}\ \text{(a spanning tree without } e\text{)} \]
Verify T' is cheaper, contradicting T's minimality
Why: T' weighs 2 plus 3, equal to 5, strictly less than either candidate for T, which weighs 7 plus 2 equals 9, or 7 plus 3 equals 10. Since T was assumed minimum, this is a contradiction - so no MST can contain the heaviest cycle edge X-Z after all.
\[ w(T')=5\ <\ 9,\ 10\ \Rightarrow\ \text{contradiction}\ \checkmark \]
Translation
\( T = \{(X,Z),\ (X,Y)\}\quad \text{or}\quad T=\{(X,Z),\ (Y,Z)\} \)
Draw it
Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.
Explain it to yourself
Discussion prompt
In The greedy MST recipe this move is made:
1. Maintain a partial forest that respects some cut
Why is that legal? Name the rule or definition it rests on before you read on.
Hint: If you can only say "because that is what you do", the rule is the thing to go and find.
Answer:
At every point in either algorithm, the edges chosen so far never cross a certain cut you can identify - that is what makes the next move provably safe.
Pattern
1. Maintain a partial forest that respects some cut
Why: At every point in either algorithm, the edges chosen so far never cross a certain cut you can identify - that is what makes the next move provably safe.
2. Add the cheapest edge crossing that cut
Why: The cut property guarantees this edge belongs to some MST, so adding it can never be a mistake.
3. Never add the heaviest edge on a cycle
Why: The cycle property guarantees this edge belongs to no MST, so an algorithm that would otherwise create a cycle should skip that edge instead.
4. Repeat until the tree has one fewer edge than the vertex count
Why: At that point every vertex is connected with no cycles - a complete, minimum spanning tree.
Real world
Discussion prompt
Outside this lesson: where does Minimum Spanning Trees: Prim & Kruskal actually turn up? Name one concrete situation — a job, a piece of software someone ships, a decision somebody has to make — and say which part of The greedy MST recipe is doing the work in it.
Hint: Vague is the failure mode here. "Engineering" is not a situation; "deciding whether this build is fast enough to ship" is.
Answer:
That deck explains what a minimum spanning tree is, then gives the cut property and the cycle property that make greedy MST algorithms provably correct. It includes full hand-traced runs of Prim's and Kruskal's algorithms, the latter with union-find, along with their running times. It targets the misconceptions that an MST is always unique, that it is the same thing as a shortest-path tree, that Kruskal never needs a cycle check, and that union-find stores weights or paths rather than connectivity alone.
Picture it
Animation
Shows: Distinct weights give a unique MST — a rendered Manim animation.
Rendered with Manim.
Takeaway: With ties, several minimum trees can exist and all are correct.
Elimination
Eliminate the wrong options
By the cut property, which edge is guaranteed to belong to some MST of this graph?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: The cut property says the cheapest edge crossing any cut belongs to some MST. Among the three crossing edges, weights 6, 4, and 9, the minimum is 4, on edge (2,3). That edge is guaranteed safe to add.
Check
A cut splits vertices 1 and 2 away from vertices 3, 4, and 5.
\[ \text{crossing edges: } (1,3)=6,\ (2,3)=4,\ (2,4)=9 \]
Check your understanding
By the cut property, which edge is guaranteed to belong to some MST of this graph?
Answer: A
Why: The cut property says the cheapest edge crossing any cut belongs to some MST. Among the three crossing edges, weights 6, 4, and 9, the minimum is 4, on edge (2,3). That edge is guaranteed safe to add.
Prediction
Predict first
By the cycle property, which edge is guaranteed to belong to NO minimum spanning tree of this graph?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: (C,D), weight 9
Why: The cycle property excludes the single heaviest edge on a cycle from every MST. Among 5, 2, 9, and 3, the maximum is 9 on edge (C,D), so that edge cannot appear in any MST.
Check
A cycle has four edges.
\[ (A,B)=5,\ (B,C)=2,\ (C,D)=9,\ (D,A)=3 \]
Check your understanding
By the cycle property, which edge is guaranteed to belong to NO minimum spanning tree of this graph?
Answer: A
Why: The cycle property excludes the single heaviest edge on a cycle from every MST. Among 5, 2, 9, and 3, the maximum is 9 on edge (C,D), so that edge cannot appear in any MST.
Section
Section 3
Concept
Prim's algorithm builds the MST by growing a single tree outward, one vertex at a time, starting from any vertex you like.
At every step, look at every edge leaving the current tree to a vertex not yet in it, and add whichever one is cheapest.
Socratic
Discussion prompt
Prim's algorithm builds the MST by growing a single tree outward, one vertex at a time, starting from any vertex you like.
Suppose that were not true. What is the first thing in Minimum Spanning Trees: Prim & Kruskal that would stop working?
Hint: Follow it one step downstream. The answer is whatever was quietly relying on it.
Answer:
At every step, look at every edge leaving the current tree to a vertex not yet in it, and add whichever one is cheapest.
Picture it
Animation
Shows: Prim grows one tree outward — a rendered Manim animation.
Rendered with Manim.
Takeaway: Always take the cheapest edge leaving the tree built so far.
Intuition
Picture the tree-so-far as a blob that starts as a single dot and slowly expands. At each moment, the blob looks outward at every road leading to a town it has not yet absorbed.
It always absorbs the town at the end of the cheapest such road. The blob keeps growing this way until it has swallowed every town.
Concept
Maintain a set of vertices already in the tree, starting with just one vertex. Repeatedly find the minimum-weight edge that has exactly one endpoint inside that set, and add it - bringing its other endpoint into the set too.
\[ \text{while } |\text{tree edges}| < |V|-1: \ \text{add}\ \min\{(u,v): u\in \text{tree}, v\notin \text{tree}\} \]
Concept
Grow one tree from one vertex. At every step take the cheapest edge leaving the tree, and the tree is connected the entire time.
PRIM(G, r)
for each u in V
key[u] = INFINITY
key[r] = 0
Q = priority queue of all vertices, keyed by key
while Q is not empty
u = Extract-Min(Q)
for each v in Adj[u]
if v is in Q and w(u, v) < key[v]
key[v] = w(u, v)
parent[v] = uCompare line 9 with Dijkstra's. Dijkstra stores distance from the source; Prim stores the weight of a single edge. That one difference is the whole difference between a shortest-path tree and a minimum spanning tree.
Notation
Every line of PRIM says one thing. Read the line, then read what it does — not the other way round.
Annotate
Invariant
The vertices already removed from the queue form one connected tree, and it is part of some minimum spanning tree of the whole graph. That second half is the cut property.
Step through it
At each step, name every edge crossing out of the tree and pick the cheapest yourself.
Picture it
Animation
Shows: PRIM executing: the current line of pseudocode is highlighted while the data it touches changes.
Rendered with Manim.
Takeaway: Prim keys on the edge weight alone, Dijkstra on distance from the source — that single change turns one algorithm into the other.
Sorting
Sort into buckets
These are the pieces of Minimum Spanning Trees: Prim & Kruskal, out of order. Put each one back under the part of the lesson it belongs to.
Concept
Every step of Prim's algorithm is really an application of the cut property. The set of vertices already in the tree is one side of a cut; everything else is the other side.
The edge Prim adds is always the cheapest edge crossing that exact cut - which the cut property guarantees is safe. That is the entire correctness argument, repeated at every step.
Picture it
Figure (svg): A weighted graph with six stations A through F. Edges: A-B weight 4, A-C weight 2, B-C weight 1, B-D weight 5, C-D weight 8, C-E weight 10, D-E weight 2, D-F weight 6, E-F weight 3.
Discussion prompt
Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.
Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.
Answer:
Trace Prim's algorithm on this six-station sensor network, starting from station A.
Worked example
Trace Prim's algorithm on this six-station sensor network, starting from station A.
Figure (svg): A weighted graph with six stations A through F. Edges: A-B weight 4, A-C weight 2, B-C weight 1, B-D weight 5, C-D weight 8, C-E weight 10, D-E weight 2, D-F weight 6, E-F weight 3.
Add the cheapest edge leaving A
Why: A has two edges leaving the tree: A-B and A-C. Comparing their weights, A-C is cheaper, so it is safe to add by the cut property, where the cut here is A versus everyone else.
\[ \min\big(w(A,B){=}4,\ w(A,C){=}2\big) = 2\ \Rightarrow\ \text{add } A\text{-}C \]
Add the cheapest edge leaving {A,C}
Why: The edges now leaving the tree are A-B(4), B-C(1), C-D(8), and C-E(10). The cheapest is B-C, so add it and bring B in.
\[ \min(4,1,8,10) = 1\ \Rightarrow\ \text{add } B\text{-}C \]
Add the cheapest edge leaving {A,B,C}
Why: With B now inside the tree, B-D(5) becomes available; the other outgoing edges are C-D(8) and C-E(10). B-D is cheapest, so add it and bring D in.
\[ \min(5,8,10) = 5\ \Rightarrow\ \text{add } B\text{-}D \]
Add the cheapest edge leaving {A,B,C,D}
Why: D adds two new outgoing edges, D-E(2) and D-F(6); C-E(10) is still available too. D-E is cheapest, so add it and bring E in.
\[ \min(2,6,10) = 2\ \Rightarrow\ \text{add } D\text{-}E \]
Add the cheapest edge leaving {A,B,C,D,E}
Why: Only two outgoing edges remain, E-F(3) and D-F(6). E-F is cheaper, so add it and bring in the last vertex, F.
\[ \min(3,6) = 3\ \Rightarrow\ \text{add } E\text{-}F \]
Verify the total weight and structure
Why: Five edges were added, one fewer than the six vertices, and every vertex is now reachable with no cycle - a genuine spanning tree.
| Order | Edge | Weight | Vertex added |
|---|---|---|---|
| 1 | A-C | 2 | C |
| 2 | B-C | 1 | B |
| 3 | B-D | 5 | D |
| 4 | D-E | 2 | E |
| 5 | E-F | 3 | F |
\[ 2+1+5+2+3 = 13 \]
Pattern
Step through it
Step through Full Prim trace, starting from A one row at a time. What is driving the change, and what would the row after the last one be?
Picture it
Figure (svg): The same weighted six-station graph A through F, used to re-run Prim's algorithm from a different starting vertex.
Discussion prompt
Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.
Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.
Answer:
Now trace Prim's algorithm on the very same network, but starting from station D instead of A - to see whether the starting point changes the answer.
Worked example
Now trace Prim's algorithm on the very same network, but starting from station D instead of A - to see whether the starting point changes the answer.
Figure (svg): The same weighted six-station graph A through F, used to re-run Prim's algorithm from a different starting vertex.
Add the cheapest edge leaving D
Why: D's edges are D-B(5), D-C(8), D-E(2), D-F(6). The cheapest is D-E, weight 2.
\[ \min(5,8,2,6)=2 \ \Rightarrow\ \text{add } D\text{-}E \]
Add the cheapest edge leaving {D,E}
Why: New options from E are E-C(10) and E-F(3); D still offers D-B(5), D-C(8), D-F(6). The cheapest overall is E-F, weight 3.
\[ \min(5,8,10,3)=3 \ \Rightarrow\ \text{add } E\text{-}F \]
Add the cheapest edge leaving {D,E,F}
Why: F adds no new options beyond D-F and E-F, both already used or superseded. The cheapest remaining option is D-B, weight 5.
\[ \min(5,8)=5 \ \Rightarrow\ \text{add } D\text{-}B \]
Add the cheapest edge leaving {D,E,F,B}
Why: B adds two new options, A-B(4) and B-C(1); C is also still reachable via D-C(8). The cheapest is B-C, weight 1.
\[ \min(8,4,1)=1 \ \Rightarrow\ \text{add } B\text{-}C \]
Add the last edge, reaching A
Why: C adds a cheaper route to A: A-C(2), better than the earlier A-B(4). That is now the only vertex left outside the tree.
\[ \min(4,2)=2 \ \Rightarrow\ \text{add } A\text{-}C \]
Verify this matches the first trace
Why: The five edges added are D-E, E-F, D-B, B-C, and A-C - the exact same edge set as when starting from A, just discovered in a different order, confirming the MST does not depend on the starting vertex when all weights are distinct.
| Order | Edge | Weight |
|---|---|---|
| 1 | D-E | 2 |
| 2 | E-F | 3 |
| 3 | D-B | 5 |
| 4 | B-C | 1 |
| 5 | A-C | 2 |
\[ 2+3+5+1+2=13 \]
Pattern
Step through it
Step through A second full Prim trace, starting from D one row at a time. What is driving the change, and what would the row after the last one be?
Concept
Checking every fringe edge from scratch at each step would be slow. Real implementations keep a priority queue of candidate edges, so the cheapest option is always available instantly.
Whenever a vertex joins the tree, its new edges are added to the priority queue, and stale, more expensive edges to an already-reached vertex are simply ignored the next time they surface.
Check
Prim's algorithm has grown the tree to include vertices P, Q, and R. The edges leaving this tree are:
\[ (Q,S)=6,\ (R,S)=3,\ (R,T)=9,\ (P,T)=7 \]
Check your understanding
Which edge should Prim's algorithm add next?
Answer: A
Why: Prim always adds the cheapest edge leaving the current tree. Among the four candidates, (R,S) at weight 3 is the smallest, so it is added next, bringing S into the tree.
Section
Section 4
Concept
Kruskal's algorithm takes a completely different approach from Prim's. Instead of growing one tree, it looks at every edge in the whole graph, sorted from cheapest to most expensive.
It walks down that sorted list, adding each edge unless doing so would create a cycle - in which case it skips that edge and moves to the next.
Concept
Sort every edge by weight and walk the list once, taking any edge that does not close a cycle. The tree grows in several disconnected pieces that eventually join.
KRUSKAL(G)
T = empty set
for each u in V
Make-Set(u)
sort E by weight, ascending
for each edge (u, v) in sorted order
if Find(u) != Find(v)
add (u, v) to T
Union(u, v)
return TLine 7 is the cycle test, and it is why union-find exists. Both endpoints already in the same component means a path between them already exists, so this edge would close a cycle and must be skipped.
Notation
Every line of KRUSKAL says one thing. Read the line, then read what it does — not the other way round.
Annotate
Invariant
The chosen edges never contain a cycle, and at every moment they are the cheapest way to connect the vertices they have connected so far.
Step through it
At each edge, decide accept-or-skip before stepping, and name the two components involved.
Picture it
Animation
Shows: KRUSKAL executing: the current line of pseudocode is highlighted while the data it touches changes.
Rendered with Manim.
Takeaway: Take edges cheapest-first, skipping any whose endpoints are already connected — the answer is a forest until the last edge.
Intuition
Instead of one blob, picture many tiny islands - one per vertex - that slowly merge into bigger islands as cheap edges connect them.
Each edge Kruskal accepts joins two separate islands into one. An edge that would connect two vertices already on the same island is refused, since it would only create a cycle within that island, not connect anything new.
Concept
Kruskal's correctness also comes from the cut property, just applied a little differently than in Prim's algorithm.
When Kruskal considers an edge and finds its two endpoints are in different components, that edge is the cheapest one remaining anywhere in the graph - which makes it, in particular, the cheapest edge crossing the cut that separates those two components from each other.
Picture it
Animation
Shows: Kruskal sorts the edges instead — a rendered Manim animation.
Rendered with Manim.
Takeaway: Prim grows one tree; Kruskal merges a forest. Same answer.
Intuition
Watch me not know the answer. This is what the first two minutes actually look like.
Kruskal sorts all edges and adds each one whose endpoints are in different components.
Try proving the whole algorithm optimal in one argument
Why: Set up an induction over the number of edges added and try to show the partial forest is contained in some MST at every stage.
The induction step has nothing to lean on
Why: You reach the step and need to know that the next edge Kruskal picks is safe — which is a fact about cuts, and you have not stated it. The argument stalls, not because it is wrong but because the lemma is missing.
Dead end. Not a mistake — a move that was worth trying and did not pay off. This happens in most proofs.
Back up. Prove the lemma first, then the algorithm is three lines
Why: The cut property says the lightest edge crossing any cut is safe. Kruskal's next edge is the lightest crossing the cut between one endpoint's component and everything else.
Both Prim and Kruskal then take three lines each. Finding the right lemma is usually more of the work than the proof that uses it — and the signal that you need one is an induction step with nothing to lean on.
The expert does not see the whole path in advance. The expert tries something, reads the result, and adjusts. That is the skill.
Picture it
Figure (svg): The same weighted six-station graph A through F, used to examine why Kruskal's next edge choice is safe.
Discussion prompt
Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.
Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.
Answer:
Return to the six-station network. Kruskal has already accepted B-C(1), A-C(2), D-E(2), and E-F(3), forming two separate components: {A,B,C} and {D,E,F}. The next cheapest edge overall is B-D, weight 5.
Worked example
Return to the six-station network. Kruskal has already accepted B-C(1), A-C(2), D-E(2), and E-F(3), forming two separate components: {A,B,C} and {D,E,F}. The next cheapest edge overall is B-D, weight 5.
Figure (svg): The same weighted six-station graph A through F, used to examine why Kruskal's next edge choice is safe.
Find the cut this edge actually crosses
Why: Since B is in {A,B,C} and D is in {D,E,F}, edge B-D crosses exactly the cut separating those two components.
\[ S=\{A,B,C\},\quad V\setminus S=\{D,E,F\} \]
List every edge crossing that same cut
Why: Checking the full edge list, only three edges cross this cut: B-D(5), C-D(8), and C-E(10). No edge connects A to D, E, or F directly.
\[ (B,D)=5,\ (C,D)=8,\ (C,E)=10 \]
Verify B-D is the cheapest crossing edge
Why: Comparing the three, B-D at weight 5 is the minimum - so by the cut property, it is guaranteed to belong to some MST. This is exactly why Kruskal is always safe to add the next cheapest non-cycle-forming edge: it is always the cheapest edge crossing the cut between the two components it merges.
\[ \min(5,8,10)=5\ \checkmark \]
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Verify B-D is the cheapest crossing edge
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Return to the six-station network. Kruskal has already accepted B-C(1), A-C(2), D-E(2), and E-F(3), forming two separate components: {A,B,C} and {D,E,F}. The next cheapest edge overall is B-D, weight 5.
Concept
Kruskal's rule sounds simple - skip an edge if it would create a cycle - but checking that naively means tracing the whole tree built so far, every single time, which is slow.
What Kruskal actually needs is a fast yes-or-no answer to one question: are these two vertices already connected by edges chosen so far?
Concept
A union-find structure, also called a disjoint-set structure, exists to answer exactly one question, quickly: are two given vertices already in the same connected group?
union-find — A data structure that tracks a collection of groups of elements, supporting two fast operations: checking whether two elements are already in the same group, and merging two groups into one.
It does not store distances, weights, or paths. It is purely a same-group tester - and that is exactly the tool Kruskal needs to detect cycles.
Intuition
Picture every vertex as a person, initially on their own one-person team. Whenever Kruskal accepts an edge, the two teams containing its endpoints merge into one bigger team.
Before accepting an edge, Kruskal just asks: are these two people already on the same team? If yes, connecting them would only create a cycle within that team, so skip the edge.
Concept
Union-find offers exactly two operations. Find, given a vertex, returns which group, identified by a representative root, it currently belongs to.
Union, given two vertices, merges their two groups into one - typically by making one group's root point to the other's.
\[ \text{find}(u) = \text{find}(v) \ \iff\ u, v \text{ already connected} \]
Concept
Two small tricks keep union-find fast even after many merges. Union by rank always attaches the smaller group's root under the bigger group's root, keeping the structure shallow.
Path compression goes further: every time find walks up a chain to reach the root, it re-points every vertex along that chain directly to the root, so future find calls on those vertices are immediate.
Intuition
Without these tricks, repeated unions can build a long, spindly chain, making find slower and slower - like tracing a family tree back through many, many generations.
Path compression flattens that chain permanently the first time anyone climbs it, so every vertex on the chain gets a direct line straight to the root from then on.
Ranking
Put in order
Put the moves of Trace path compression on a union-find array into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Starting at 1, follow parent pointers: 1 to 2, 2 to 3, 3 to 4, 4 to 5, and 5 points to itself - so 5 is the root.
Worked example
Start with six vertices, each its own group. Four unions are performed in order: union(1,2), union(2,3), union(3,4), and union(4,5), each attaching the first vertex's root under the second's.
| Vertex | Parent (before find) |
|---|---|
| 1 | 2 |
| 2 | 3 |
| 3 | 4 |
| 4 | 5 |
| 5 | 5 (root) |
| 6 | 6 (root) |
Call find(1) and walk the chain
Why: Starting at 1, follow parent pointers: 1 to 2, 2 to 3, 3 to 4, 4 to 5, and 5 points to itself - so 5 is the root. That is four hops just to answer one find call.
\[ 1 \to 2 \to 3 \to 4 \to 5\ (\text{root}) \]
Apply path compression
Why: Path compression re-points every vertex visited on that walk - 1, 2, 3, and 4 - directly to the root, 5, so the next find on any of them is immediate.
| Vertex | Parent (after find(1)) |
|---|---|
| 1 | 5 |
| 2 | 5 |
| 3 | 5 |
| 4 | 5 |
| 5 | 5 (root) |
| 6 | 6 (root) |
Verify the group membership did not change
Why: Before and after compression, find(1) still returns root 5, and every vertex 1 through 5 is still in the same group - compression only shortens the pointers, it never changes which vertices are connected.
\[ \text{find}(1) = 5 \ \text{both before and after}\ \checkmark \]
Check
Using the union-find example from the previous slide, after find(1) triggers path compression, the parent array is updated.
Check your understanding
What is parent[2] immediately after find(1) completes with path compression?
Answer: A
Why: Path compression re-points every vertex on the path walked during find - which includes 1, 2, 3, and 4 - directly to the discovered root, 5. So parent[2] becomes 5, not just parent[1].
Worked example
Trace Kruskal's algorithm on the same six-station network, using union-find to test for cycles.
Figure (svg): The same weighted six-station graph A through F, used to trace Kruskal's algorithm.
Sort every edge from cheapest to most expensive
Why: Sorting once up front lets Kruskal always consider the next cheapest edge in order.
| Order | Edge | Weight |
|---|---|---|
| 1 | B-C | 1 |
| 2 | A-C | 2 |
| 3 | D-E | 2 |
| 4 | E-F | 3 |
| 5 | A-B | 4 |
| 6 | B-D | 5 |
| 7 | D-F | 6 |
| 8 | C-D | 8 |
| 9 | C-E | 10 |
Consider B-C: different groups, so add it
Why: find(B) and find(C) return different roots, each still its own group, so accepting this edge cannot create a cycle. Union B and C.
\[ \text{find}(B) \neq \text{find}(C) \ \Rightarrow\ \text{add }B\text{-}C \]
Consider A-C: different groups, so add it
Why: A is still alone; C is now grouped with B. Different groups, so add A-C and union A into that group.
\[ \text{find}(A) \neq \text{find}(C) \ \Rightarrow\ \text{add }A\text{-}C \]
Consider D-E: different groups, so add it
Why: D and E are each still their own group. Add D-E and union them.
\[ \text{find}(D) \neq \text{find}(E) \ \Rightarrow\ \text{add }D\text{-}E \]
Consider E-F: different groups, so add it
Why: F is still alone; E is grouped with D. Add E-F and union F into that group.
\[ \text{find}(E) \neq \text{find}(F) \ \Rightarrow\ \text{add }E\text{-}F \]
Consider A-B: same group, so skip it
Why: By now A, B, and C are all one group. find(A) and find(B) return the same root, so adding A-B would only close a cycle inside that group - skip it.
\[ \text{find}(A) = \text{find}(B) \ \Rightarrow\ \text{skip} \]
Consider B-D: different groups, so add it
Why: B belongs to {A,B,C}; D belongs to {D,E,F} - two different groups. Add B-D, merging both groups into one.
\[ \text{find}(B) \neq \text{find}(D) \ \Rightarrow\ \text{add }B\text{-}D \]
Verify the tree is complete
Why: Five edges have now been added - one fewer than the six vertices - and every vertex belongs to a single group. The remaining edges, D-F, C-D, and C-E, are never even considered, since the algorithm can stop once V minus 1 edges are chosen. Total weight matches the Prim traces exactly.
| Accepted edge | Weight |
|---|---|
| B-C | 1 |
| A-C | 2 |
| D-E | 2 |
| E-F | 3 |
| B-D | 5 |
\[ 1+2+2+3+5=13 \]
Pattern
Step through it
Step through Full Kruskal trace with union-find one row at a time. What is driving the change, and what would the row after the last one be?
Anomaly
Predict first
A student writes this, and it looks reasonable:
A student sorts the edges and adds every one in order, forgetting to check whether each new edge would create a cycle.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Adding 1-2, then 2-3, then 1-3 without any cycle check gives three edges for only three vertices - but a spanning tree on three vertices needs exactly two.
Check before every addition: would this edge connect two vertices already in the same group?
Why: Adding 1-2, then 2-3, then 1-3 without any cycle check gives three edges for only three vertices - but a spanning tree on three vertices needs exactly two. The result is not a tree at all; it is the full triangle, containing a cycle.
Trap
A student sorts the edges and adds every one in order, forgetting to check whether each new edge would create a cycle.
\[ (1,2)=1,\ (2,3)=2,\ (1,3)=3 \]
Add all three edges in sorted order
Why: Adding 1-2, then 2-3, then 1-3 without any cycle check gives three edges for only three vertices - but a spanning tree on three vertices needs exactly two. The result is not a tree at all; it is the full triangle, containing a cycle.
\[ 1+2+3 = 6\ \text{(has a cycle, not a tree)} \]
Check before every addition: would this edge connect two vertices already in the same group?
\[ (1,2)=1,\ (2,3)=2,\ (1,3)=3 \]
Add 1-2, then 2-3, then test 1-3
Why: After adding 1-2 and 2-3, vertices 1, 2, and 3 are already all one group. Testing 1-3 with union-find shows find(1) equals find(3) - adding it would only close a cycle, so it is correctly skipped.
\[ \text{find}(1) = \text{find}(3) \ \Rightarrow\ \text{skip } (1,3) \]
Confirm the correct, lower-weight tree
Why: Skipping the cycle-forming edge leaves exactly two edges, 1-2 and 2-3, for a valid spanning tree of weight 3 - cheaper and correct, unlike the cycle-containing result on the left.
\[ 1+2=3\ <\ 6 \]
Notation
Annotate
From Trap: adding an edge without checking for a cycle — read this one piece at a time. What is each part doing?
On: \( \text{find}(1) = \text{find}(3) \ \Rightarrow\ \text{skip } (1,3) \)
Anomaly
Predict first
A student writes this, and it looks reasonable:
A student thinks union-find keeps track of the cheapest edge between two groups, and tries to ask it directly for that information.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Union-find cannot answer this - it stores no weights and no edges at all, only which vertices currently belong to which group.
Union-find answers exactly one question: are these two vertices already in the same group? Nothing more.
Why: Union-find cannot answer this - it stores no weights and no edges at all, only which vertices currently belong to which group. This request is simply outside what the structure does.
Trap
A student thinks union-find keeps track of the cheapest edge between two groups, and tries to ask it directly for that information.
Ask union-find for 'the minimum weight edge between A's group and B's group'
Why: Union-find cannot answer this - it stores no weights and no edges at all, only which vertices currently belong to which group. This request is simply outside what the structure does.
Union-find answers exactly one question: are these two vertices already in the same group? Nothing more.
Use union-find only to test find(u) equals find(v)
Why: The weight comparison already happened when the edges were sorted at the very start. Union-find's only job during the scan is the yes-or-no cycle test - the sorted order handles picking the cheapest edge.
Keep the two jobs separate
Why: Sorting decides which edge to consider next, cheapest first; union-find decides whether that edge is safe to add, different groups, or must be skipped, same group, since it would form a cycle. Confusing the two roles is the core of this misconception.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
Prediction
Predict first
Kruskal considers these edges next, in this sorted order. Which one does it actually add?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: (P,S), weight 6 - because (Q,R) is skipped for forming a cycle, and (P,S) is the next cheapest edge joining two different groups
Why: (Q,R) is cheapest, but Q and R are already in the same group {P,Q,R}, so adding it would create a cycle - find(Q) equals find(R), so it is skipped. The next cheapest edge, (P,S), connects two different groups, {P,Q,R} and {S,T}, so it is safe and gets added.
Check
Kruskal is processing edges in sorted order and has already accepted enough edges to form two groups, {P,Q,R} and {S,T}. Vertex U is still alone. The next edges in sorted order, not yet processed, are:
\[ (Q,R)=2,\ (P,S)=6,\ (R,U)=7,\ (S,U)=9 \]
Check your understanding
Kruskal considers these edges next, in this sorted order. Which one does it actually add?
Answer: A
Why: (Q,R) is cheapest, but Q and R are already in the same group {P,Q,R}, so adding it would create a cycle - find(Q) equals find(R), so it is skipped. The next cheapest edge, (P,S), connects two different groups, {P,Q,R} and {S,T}, so it is safe and gets added.
Check
While running Kruskal's algorithm, you call find(X) and find(Y) for two vertices X and Y, and they return the same root.
Check your understanding
What does this tell you?
Answer: A
Why: Union-find's entire job is testing group membership. Equal roots mean X and Y are already in the same connected group, so any edge directly joining them would only close a cycle within that group - exactly the case Kruskal must skip.
Section
Section 5
Concept
Prim's running time depends on how the fringe edges are stored. With a binary heap holding candidate edges, each edge may be inserted or updated once, and each heap operation costs time proportional to the log of the number of vertices.
\[ O\big((|V|+|E|)\log|V|\big) \]
With a simple array instead of a heap, finding the minimum takes longer per step, but there is no log factor at all - giving a different tradeoff that tends to win on dense graphs.
\[ O(|V|^2) \]
Socratic
Discussion prompt
With a simple array instead of a heap, finding the minimum takes longer per step, but there is no log factor at all - giving a different tradeoff that tends to win on dense graphs.
Suppose that were not true. What is the first thing in Minimum Spanning Trees: Prim & Kruskal that would stop working?
Hint: Follow it one step downstream. The answer is whatever was quietly relying on it.
Concept
Kruskal's algorithm first sorts every edge, which costs time proportional to the number of edges times the log of the number of edges.
\[ O(|E|\log|E|) \]
Processing each edge afterward with union-find, using union by rank and path compression, costs barely more than a constant amount of time per operation - so the sort dominates the total running time.
\[ O(|E|\log|E|) = O(|E|\log|V|) \quad \text{since } |E| \le |V|^2 \]
Explain it
Discussion prompt
Explain Kruskal's running time to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Kruskal's algorithm first sorts every edge, which costs time proportional to the number of edges times the log of the number of edges.
Ranking
Put in order
Put the moves of Compute the running time for a concrete graph into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Since 1024 is exactly 2 to the 10th power, its base-2 log is exactly 10.
Worked example
A network has 1024 vertices and 2000 edges - noticeably sparse, since 2000 is not much bigger than 1024.
\[ |V|=1024=2^{10},\quad |E|=2000 \]
Compute the log factors
Why: Since 1024 is exactly 2 to the 10th power, its base-2 log is exactly 10. 2000 is a little under 2 to the 11th power, 2048, so its base-2 log is just under 11.
\[ \log_2 1024 = 10, \qquad \log_2 2000 \approx 10.97 \]
Estimate Kruskal's operation count
Why: Kruskal's dominant cost is sorting the edges: the number of edges times the log of the number of edges.
\[ |E|\log_2|E| \ \approx\ 2000 \times 10.97 \ \approx\ 21{,}940 \]
Estimate Prim's operation count with a binary heap
Why: Prim's cost is the vertex-plus-edge count times the log of the vertex count.
\[ (|V|+|E|)\log_2|V| \ \approx\ 3024 \times 10 \ =\ 30{,}240 \]
Verify which algorithm wins on this sparse graph
Why: Kruskal's estimate, about 21,940, is noticeably smaller than Prim's, about 30,240, because the extra vertex-count term in Prim's bound matters proportionally more when the graph is sparse. On a dense graph, where the edge count approaches the vertex count squared, the array-based Prim bound would instead pull ahead.
\[ 21{,}940 \ <\ 30{,}240\ \checkmark \]
Picture it
Animation
Shows: Each line of the worked example "Compute the running time for a concrete graph", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Kruskal's estimate, about 21,940, is noticeably smaller than Prim's, about 30,240, because the extra vertex-count term in Prim's bound matters proportionally more when the graph is sparse. On a dense graph, where the edge count approaches the vertex count squared, the array-based Prim bound would instead pull ahead.
Concept
Neither algorithm is universally faster - the right choice depends on how many edges the graph has, relative to its vertices.
Kruskal, dominated by sorting the edge list, tends to do better on sparse graphs, where the edge count stays close to the vertex count. Prim, especially with a simple array instead of a heap, tends to do better on dense graphs, where the vertex-squared bound avoids any log factor at all.
Analogy
Discussion prompt
Explain Choosing Prim vs Kruskal by analogy to something with no CS3000 Algorithms in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Neither algorithm is universally faster - the right choice depends on how many edges the graph has, relative to its vertices.
Elimination
Eliminate the wrong options
Which running-time bound is most likely to favor Prim's algorithm with a simple array, no heap, over Kruskal's algorithm here?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: This graph is dense - the edge count is close to the maximum possible. Prim's array-based bound, vertex count squared, is about 250,000 here and has no log factor, while Kruskal must sort around 120,000 edges, paying a log-of-edge-count factor on top of a similarly large edge count. On dense graphs like this, the array-based Prim bound often wins.
Check
A graph has 500 vertices and roughly 120,000 edges - close to the maximum possible for that many vertices, since 500 choose 2 is about 124,750.
Check your understanding
Which running-time bound is most likely to favor Prim's algorithm with a simple array, no heap, over Kruskal's algorithm here?
Answer: A
Why: This graph is dense - the edge count is close to the maximum possible. Prim's array-based bound, vertex count squared, is about 250,000 here and has no log factor, while Kruskal must sort around 120,000 edges, paying a log-of-edge-count factor on top of a similarly large edge count. On dense graphs like this, the array-based Prim bound often wins.
Concept
Moves added today: none.
That is a result, not a gap. Everything in this lesson was proved with moves you already owned.
Moves you reused today:
The cut property and the cycle property are the same move as lesson 13's interval scheduling proof, with an edge swapped instead of an interval. If they felt like new theorems, re-read the exchange-argument slide from that lesson.
Full toolkit so far: #1 through #15.
Next session opens with you naming every one of these from memory, before any new material.
Counterexample
Discussion prompt
Next session opens with you naming every one of these from memory, before any new material.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Foundations — Graphs, Trees, and Spanning Trees · The Cut Property and the Cycle Property · Prim's Algorithm · Kruskal's Algorithm and Union-Find · Running Times. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
You now have the two correctness lemmas and two algorithms that power every minimum-spanning-tree problem you will meet.
| Technique | The one move |
|---|---|
| Cut property | Cheapest crossing edge is always safe |
| Cycle property | Heaviest cycle edge is never needed |
| Prim | Grow one tree, add the cheapest leaving edge |
| Kruskal + union-find | Sort edges; add if find(u) does not equal find(v) |
Want this taught 1-on-1? Alexander tutors CS3000 Algorithms — $55/session, free consultation.