This deck builds single-source shortest paths from one shared primitive, edge relaxation. It traces Dijkstra's algorithm by hand with a priority queue and explains why its greedy choice is safe only for non-negative weights, then covers Bellman-Ford's V-1 rounds of relaxation and its negative-cycle detection. It targets the traps of trusting Dijkstra with a negative edge, relaxing in the wrong direction, misjudging why V-1 rounds are needed, and reviving a vertex that has already been settled.
Subject: CS3000 Algorithms · 129 slides · symbolic lesson
Open the interactive version of this deck · Homework for this lesson
Title
CS3000 Algorithms
Finding the cheapest way from one vertex to every other vertex.
Objectives
By the end of this lesson you can:
1. Explain what edge relaxation is and use it as the shared building block of both algorithms below.
2. Trace Dijkstra's algorithm on a small graph using a priority queue, and explain why its greedy choice is safe only when every edge weight is non-negative.
3. Trace Bellman-Ford's algorithm, explain why relaxing every edge V-1 times is exactly enough, and use one extra round to detect a negative-weight cycle.
4. State the running time of each algorithm and choose the right one for a given graph.
Warm-up
Discussion prompt
Before we open Shortest Paths: Dijkstra & Bellman-Ford: without looking back, what was the main idea of Topological Sort & Strongly Connected Components, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
Two ways to linearize a DAG (DFS finish-time order and Kahn's in-degree removal), how strongly connected components are found with Kosaraju's two-pass reverse-graph idea, and why the condensation of any graph's SCCs is always a DAG. Targets trying to sort a cyclic graph, the false belief that a topological order is unique, mixing up a directed SCC with an undirected connected component, and forgetting to reverse the graph in Kosaraju's algorithm.
Concept
Before any new material: cover the screen.
You have named 14 reusable moves so far. Say as many as you can out loud, by number, from memory.
Do not advance until you have actually tried. Getting four of eight is information; skipping the exercise is not.
Here they are. Score yourself.
Today adds one move to this list. Everything else you will need is already above.
The question that starts every proof from here on is not how do I begin. It is which of these applies here?
Counterexample
Discussion prompt
You have named 14 reusable moves so far. Say as many as you can out loud, by number, from memory.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Do not advance until you have actually tried. Getting four of eight is information; skipping the exercise is not.
Section
Section 1
Concept
A graph is a set of vertices (the dots, also called nodes) connected by edges (the links between them). In this lesson every edge points in one direction, from one vertex to another.
weight — A number attached to an edge, representing its cost: distance, time, price, or anything else you want to minimize along a route.
A path is a sequence of edges chained head to tail. The cost of a path is just the sum of the weights of the edges you used to walk it.
Analogy
Discussion prompt
Explain Graphs, vertices, and weighted edges by analogy to something with no CS3000 Algorithms in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
A graph is a set of vertices (the dots, also called nodes) connected by edges (the links between them). In this lesson every edge points in one direction, from one vertex to another.
Concept
The single-source shortest-path problem: pick one starting vertex, the source, and find the cheapest possible path from it to every other vertex in the graph, all at once.
The output is a distance estimate for every vertex, usually called dist. dist of a vertex is the total weight of the cheapest path found so far from the source to it.
Both algorithms in this lesson build and improve this same dist array until it is guaranteed correct.
Explain it
Discussion prompt
Explain What single-source shortest path means to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
The single-source shortest-path problem: pick one starting vertex, the source, and find the cheapest possible path from it to every other vertex in the graph, all at once.
Picture it
Animation
Shows: Why shortest paths decompose — a rendered Manim animation.
Rendered with Manim.
Takeaway: The property every shortest-path algorithm silently relies on.
Intuition
Picture a road map. Towns are vertices, roads are edges, and each road's toll or drive time is its weight. 'Shortest path' does not mean fewest roads — it means the cheapest total toll.
A route with three cheap roads can beat a route with one expensive one. That is exactly why we track running totals, dist, instead of counting hops.
Picture it
Figure (svg): A directed weighted graph with vertex S as the source, edges S to C weight 2, S to B weight 4, C to B weight 1, C to D weight 8, C to E weight 10, B to D weight 5, and D to E weight 2
Discussion prompt
Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.
Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.
Answer:
Both algorithms in this lesson will be traced on the same small graph, so you can compare their behavior directly. Source vertex S, plus four more vertices: B, C, D, E.
Concept
Both algorithms in this lesson will be traced on the same small graph, so you can compare their behavior directly. Source vertex S, plus four more vertices: B, C, D, E.
Figure (svg): A directed weighted graph with vertex S as the source, edges S to C weight 2, S to B weight 4, C to B weight 1, C to D weight 8, C to E weight 10, B to D weight 5, and D to E weight 2
Every weight here is non-negative, which matters a great deal, as you'll soon see. We'll call this graph G1 and reuse it for both algorithms.
Concept
Both Dijkstra and Bellman-Ford are built entirely out of one repeated move: relaxing an edge.
relax an edge — Given an edge from u to v with some weight, check whether going from the source to u, then across this edge, beats the best known way to reach v. If it does, update v's distance estimate and remember u as the way you got there.
The rule is a single comparison:
\[ \text{if } dist[u] + w(u,v) < dist[v]: \quad dist[v] \leftarrow dist[u] + w(u,v),\ \ prev[v] \leftarrow u \]
Notice the direction: you always add to the KNOWN side, dist[u], and compare against the side you might improve, dist[v]. Getting that backwards is a real and common bug — more on that shortly.
Definition probe
Sort into buckets
Every line below is part of the definition of weight or of relax an edge — one or the other, never both. Put each where it belongs.
Intuition
Think of dist[v] as the best rumor you've heard so far about how cheaply you can reach v. Relaxing an edge into v is hearing one more rumor: 'you can reach u for this much, and then it's one more hop to v.'
If the new rumor is cheaper than the best one you're holding, you update. If it's not, you ignore it and keep what you had. Nothing more mysterious than that.
Picture it
Figure (svg): Vertex u with distance 2 connected by an edge of weight 4 to vertex v with current distance 9
Discussion prompt
Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.
Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.
Answer:
Suppose we already know dist[u] = 2. There is an edge from u to v with weight 4, and the current best estimate is dist[v] = 9.
Worked example
Suppose we already know dist[u] = 2. There is an edge from u to v with weight 4, and the current best estimate is dist[v] = 9.
Figure (svg): Vertex u with distance 2 connected by an edge of weight 4 to vertex v with current distance 9
Compute the candidate distance through u
Why: Add u's known distance to the weight of the edge — this is the cost of the route that goes to u first, then hops to v.
\[ dist[u] + w(u,v) = 2 + 4 = 6 \]
Compare the candidate to the current estimate
Why: The candidate route costs 6, which is less than the 9 we currently believe. This route is a genuine improvement.
\[ 6 < 9 \]
Update dist[v] and prev[v]
Why: Since the comparison succeeded, replace v's estimate with the cheaper number, and record u as the vertex we'd arrive from.
\[ dist[v] \leftarrow 6, \quad prev[v] \leftarrow u \]
Verify the update really is an improvement
Why: 6 is less than the old value of 9, and it was computed honestly as dist[u] plus the real edge weight, so the update is legitimate. If a later relaxation finds something even cheaper, it will replace 6 the same way.
\[ 6 < 9\ \checkmark \]
Notation
Annotate
From Relax a single edge — read this one piece at a time. What is each part doing?
On: \( dist[u] + w(u,v) = 2 + 4 = 6 \)
Anomaly
Predict first
A student writes this, and it looks reasonable:
A student mixes up which side of the comparison is the 'known' side and which is the side being improved.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: This student writes the comparison as if v's distance plus the weight should beat u's distance — the roles are swapped.
Always add the weight to the vertex whose distance is already trusted, and compare against the vertex you might be improving.
Why: This student writes the comparison as if v's distance plus the weight should beat u's distance — the roles are swapped.
Trap
A student mixes up which side of the comparison is the 'known' side and which is the side being improved.
\[ dist[u]=2,\ \ dist[v]=9,\ \ w(u,v)=4 \]
Check the backwards condition
Why: This student writes the comparison as if v's distance plus the weight should beat u's distance — the roles are swapped.
\[ \text{if } dist[v] + w(u,v) < dist[u]:\ \ 9 + 4 = 13 < 2 \ ?\ \ \text{false} \]
The update never fires
Why: Because the backwards condition is false, dist[v] stays at 9 forever, even though a cheaper route through u genuinely exists. The bug doesn't crash — it just silently produces a worse answer.
\[ dist[v] \text{ stuck at } 9 \quad (\text{should be } 6) \]
Always add the weight to the vertex whose distance is already trusted, and compare against the vertex you might be improving.
\[ dist[u]=2,\ \ dist[v]=9,\ \ w(u,v)=4 \]
Check the correct condition
Why: u is the known side, so its distance plus the edge weight is the candidate for v.
\[ \text{if } dist[u] + w(u,v) < dist[v]:\ \ 2 + 4 = 6 < 9 \ ?\ \ \text{true} \]
The update fires correctly
Why: dist[v] becomes 6, matching the real cheapest known route. A quick way to remember it: the edge runs FROM u TO v, and the arithmetic follows the arrow the same way.
\[ dist[v] \leftarrow 6\ \checkmark \]
Anomaly
Predict first
A student writes this, and it looks reasonable:
A student initializes every vertex's distance to 0 'to be safe', instead of only the source.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: The candidate distance through S is 0 plus 2 equals 2, but C's current estimate is already 0, which looks smaller.
Only the source starts at 0. Every other vertex starts at infinity, meaning 'no route found yet'.
Why: The candidate distance through S is 0 plus 2 equals 2, but C's current estimate is already 0, which looks smaller. The comparison fails and NOTHING updates.
Trap
A student initializes every vertex's distance to 0 'to be safe', instead of only the source.
\[ dist[S]=0,\ dist[B]=0,\ dist[C]=0,\ dist[D]=0,\ dist[E]=0 \]
Try to relax the edge from S to C, weight 2
Why: The candidate distance through S is 0 plus 2 equals 2, but C's current estimate is already 0, which looks smaller. The comparison fails and NOTHING updates.
\[ dist[S] + w(S,C) = 0 + 2 = 2 \ \not< \ 0 = dist[C] \]
Every real relaxation gets blocked
Why: Because every non-source vertex already 'looks' reachable for free, no honest relaxation ever beats 0, and every dist value freezes at the wrong answer, 0.
Only the source starts at 0. Every other vertex starts at infinity, meaning 'no route found yet'.
\[ dist[S]=0,\ dist[B]=\infty,\ dist[C]=\infty,\ dist[D]=\infty,\ dist[E]=\infty \]
Relax the edge from S to C, weight 2
Why: Any finite number beats infinity, so the very first honest relaxation into an unreached vertex is guaranteed to succeed.
\[ dist[S] + w(S,C) = 0 + 2 = 2 \ < \ \infty\ \checkmark \]
Infinity means 'unknown', not 'free'
Why: Initializing to infinity guarantees that the first real path discovered always counts as an improvement, which is exactly what should happen.
Concept
Not every vertex has to be reachable from the source. If no chain of edges leads from the source to some vertex, no relaxation will ever fire for it.
That vertex's dist value simply stays at infinity forever, in both algorithms. Infinity there does not mean an error — it correctly reports 'no path exists'.
\[ dist[v] = \infty \ \text{at the end} \ \Rightarrow\ v \text{ is unreachable from the source} \]
Pattern
1. Initialize the source to 0, everyone else to infinity
Why: The source needs no path to reach itself; everyone else starts as 'unknown' so the first real route always wins.
2. For an edge from u to v, add dist[u] to the weight
Why: Always build the candidate from the KNOWN side, u, across the edge, toward the side you might improve, v.
3. Compare the candidate to dist[v]; update only if it's smaller
Why: This single comparison, repeated over and over on different edges, is the entire engine behind both algorithms in this lesson.
4. When you update dist[v], also update prev[v]
Why: Recording the vertex you arrived from lets you reconstruct the actual cheapest path later, not just its total cost.
Check
There is an edge from vertex p to vertex q with weight 3. Currently dist[p] = 5 and dist[q] = 6.
Check your understanding
What happens when this edge is relaxed?
Answer: D
Why: The candidate distance through p is dist[p] + weight = 5 + 3 = 8. Since 8 is not less than the current dist[q] = 6, the relaxation does not fire and dist[q] correctly stays at 6.
Section
Section 2
Concept
Dijkstra's algorithm repeatedly does one thing: find the unfinished vertex with the smallest current distance estimate, declare its distance final, and relax every edge leaving it.
settled — A vertex whose dist value has been declared final and will never be changed again for the rest of the algorithm.
This repeats — always picking the closest unsettled vertex next — until every vertex is settled.
Concept
BFS with a priority queue instead of a plain queue. Take the closest unsettled vertex, settle it forever, and relax its edges.
DIJKSTRA(G, s)
for each u in V
dist[u] = INFINITY
dist[s] = 0
Q = priority queue of all vertices, keyed by dist
while Q is not empty
u = Extract-Min(Q)
for each v in Adj[u]
if dist[u] + w(u, v) < dist[v]
dist[v] = dist[u] + w(u, v)
parent[v] = u
Decrease-Key(Q, v)Line 7 is the greedy choice and the load-bearing one. It claims the nearest unsettled vertex already has its final distance. That claim needs every weight to be non-negative — a negative edge could otherwise shorten a path after it was settled.
Notation
Every line of DIJKSTRA says one thing. Read the line, then read what it does — not the other way round.
Annotate
Invariant
Every settled vertex has its true shortest distance, permanently. Everything still in the queue holds the best distance found so far, which can still improve.
Step through it
Before each extraction, name which vertex you expect to be settled next.
Picture it
Animation
Shows: DIJKSTRA executing: the current line of pseudocode is highlighted while the data it touches changes.
Rendered with Manim.
Takeaway: Settle the nearest unsettled vertex and relax its edges — correct only because no edge weight is negative.
Intuition
Imagine dispatching couriers from a warehouse. You always send the next courier to whichever known town is currently cheapest to reach. Once a courier arrives, that town's delivery cost is locked in.
From that town, you discover new, possibly cheaper routes to its neighbors, and add those to your list of towns to consider next.
Concept
Dijkstra keeps two groups: settled vertices, whose distance is final, and unsettled vertices, whose distance is still just an estimate.
priority queue — A data structure that always hands you the item with the smallest key, efficiently. Here, the key is a vertex's current distance estimate.
Each step: pop the unsettled vertex with the smallest estimate from the priority queue, mark it settled, and relax its outgoing edges — which may push better estimates for its neighbors back into the queue.
Picture it
Animation
Shows: The heap is where the running time lives — a rendered Manim animation.
Rendered with Manim.
Takeaway: The algorithm is unchanged; only the data structure moves the bound.
Concept
Both Dijkstra and Bellman-Ford maintain the same two pieces of bookkeeping for every vertex.
prev — For a vertex v, prev[v] stores the vertex you arrive from on the cheapest known path to v. Following prev pointers backward from any vertex to the source reconstructs the actual path, not just its cost.
dist answers 'how much does it cost'; prev answers 'which way do I actually go'. You need both to fully solve the problem.
Intuition
What move should we make next?
Dijkstra settles the unsettled vertex with the smallest tentative distance and never revisits it.
\[ \text{claim: when } u \text{ is settled, } dist[u] = \delta(s, u) \]
A cheaper route to u would have to arrive through some vertex that is still unsettled.
This is a for-all claim about every settle. Which move, and what exactly would you assume goes wrong?
_Look at your toolkit. Say a move number out loud before this slide advances._ A wrong guess is useful. A silent guess is not.
Concept
Dijkstra's whole method rests on one guarantee: once you settle the closest unsettled vertex, its distance can never improve later. Why is that guaranteed?
Every unsettled vertex's current estimate is already at least as large as the one you're about to settle, since you always pick the smallest. Any future path to the settled vertex would have to pass through some unsettled vertex first.
\[ \text{any future path} = \text{through some unsettled } x, \quad dist[x] \ge dist[\text{chosen vertex}] \]
Because every edge weight is non-negative, continuing from x can only add more cost, never subtract it. So no future route can ever beat the one you already locked in.
\[ dist[x] + w(x, \ldots) \ge dist[x] \ge dist[\text{chosen vertex}] \]
Concept
Every proof of this kind has the same five or six moves in the same order. The order is not something you rediscover each time.
It is on the right. It will stay on the right through the worked examples that follow.
Why this matters: the structure is now handled. You are not spending working memory on what comes next — you are spending all of it on the one hard step.
Step 2 is what makes the proof finite. Reasoning about somewhere it fails gives you nothing; reasoning about the first failure hands you a fully-correct prefix to argue from.
Fill the middle
Fill in the blanks
From The new move, named — finish the line. Write what belongs on the right of the equals sign before you look.
dist[y] = \delta(s, y) \le \delta(s, u) \le dist[u]
Why: Producing the right-hand side unprompted is the difference between recognising this line and being able to use it. Suppose some vertex is settled with a wrong distance.
Concept
The move: #15 (Cut at the first failure).
Assume the claim fails and take the FIRST failure
Why: Suppose some vertex is settled with a wrong distance. Let u be the first such vertex, in settle order. Every vertex settled before u is therefore correct — and that prefix of correctness is the thing you get to use.
Look at the true shortest path to u
Why: Walk it from s toward u and find the first vertex y on it that is not yet settled. The vertex x just before y is settled, hence correct, hence the edge from x to y was already relaxed.
\[ dist[y] = \delta(s, y) \le \delta(s, u) \le dist[u] \]
Show the step could not have happened
Why: So y had a tentative distance no larger than u's, and Dijkstra would have settled y instead of u. The first failure could not have occurred, so there is no failure at all.
Non-negative weights are what make the middle inequality true. Allow one negative edge and that line is false, which is precisely why Dijkstra breaks there. Note that this is #8 (Take the extreme one) with the ordering supplied by the algorithm.
Translation
\( dist[y] = \delta(s, y) \le \delta(s, u) \le dist[u] \)
Draw it
Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.
Intuition
Picture every unsettled vertex as sitting at or beyond the distance of the one you're about to settle. Any path that detours through one of them can only add distance from there — like a toll booth that never gives refunds.
That 'never gives refunds' property is exactly what a negative weight would violate. Keep that thought — it's the whole reason for the trap coming up soon.
Ranking
Put in order
Put the moves of Dijkstra trace: initializing and the first two settles into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. S has the smallest distance, 0, so it settles first.
Worked example
Run Dijkstra on graph G1 from source S. Recall the edges: S to C weight 2, S to B weight 4, C to B weight 1, C to D weight 8, C to E weight 10, B to D weight 5, D to E weight 2.
Initialize
Why: The source starts at 0; everyone else starts at infinity, unsettled.
| step | settled | dist[S] | dist[B] | dist[C] | dist[D] | dist[E] |
|---|---|---|---|---|---|---|
| 0 | - | 0 | ∞ | ∞ | ∞ | ∞ |
Settle S and relax its edges
Why: S has the smallest distance, 0, so it settles first. Relaxing S to C (weight 2) and S to B (weight 4) gives both their first real estimates.
| step | settled | dist[S] | dist[B] | dist[C] | dist[D] | dist[E] |
|---|---|---|---|---|---|---|
| 1 | S | 0 | 4 | 2 | ∞ | ∞ |
Settle C and relax its edges
Why: Among unsettled vertices, C has the smallest estimate, 2, so it settles next. Relaxing C to B improves B from 4 to 3; relaxing C to D and C to E gives them their first estimates.
| step | settled | dist[S] | dist[B] | dist[C] | dist[D] | dist[E] |
|---|---|---|---|---|---|---|
| 2 | S, C | 0 | 3 | 2 | 10 | 12 |
Check the running tally after two settles
Why: So far: S is settled at 0, C is settled at 2. B, D, and E are only estimates (3, 10, 12) and could still improve, since they are not yet settled.
Worked example
Continuing from where we left off: S and C are settled. The priority queue currently holds B at 3, D at 10, and E at 12 (plus a stale leftover entry for B at 4, from before C improved it — more on that soon).
Settle B and relax its edges
Why: B has the smallest unsettled estimate, 3. Relaxing B to D (weight 5) gives candidate 3 + 5 = 8, which beats the current D estimate of 10.
| step | settled | dist[S] | dist[B] | dist[C] | dist[D] | dist[E] |
|---|---|---|---|---|---|---|
| 3 | S, C, B | 0 | 3 | 2 | 8 | 12 |
Settle D and relax its edges
Why: D now has the smallest unsettled estimate, 8. Relaxing D to E (weight 2) gives candidate 8 + 2 = 10, which beats the current E estimate of 12.
| step | settled | dist[S] | dist[B] | dist[C] | dist[D] | dist[E] |
|---|---|---|---|---|---|---|
| 4 | S, C, B, D | 0 | 3 | 2 | 8 | 10 |
Settle E
Why: E is the only unsettled vertex left, at 10. No outgoing edges remain to relax. Every vertex is now settled.
| step | settled | dist[S] | dist[B] | dist[C] | dist[D] | dist[E] |
|---|---|---|---|---|---|---|
| 5 | S, C, B, D, E | 0 | 3 | 2 | 8 | 10 |
Verify every vertex is settled and the distances are final
Why: All five vertices appear in the settled column and no priority-queue entries remain that could still improve them. Final distances: S=0, C=2, B=3, D=8, E=10.
\[ dist[S]=0,\ dist[C]=2,\ dist[B]=3,\ dist[D]=8,\ dist[E]=10 \]
Picture it
Animation
Shows: Prim and Dijkstra differ in one word — a rendered Manim animation.
Rendered with Manim.
Takeaway: Nearly the same code, and a completely different answer.
Step zero
Discussion prompt
Reconstructing the shortest path with prev pointers — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Start at the destination and follow prev backward
Answer:
Worked example
The trace above also recorded a prev pointer every time it updated a distance. Use them to reconstruct the actual cheapest path from S to E, not just its cost.
| vertex | prev |
|---|---|
| C | S |
| B | C |
| D | B |
| E | D |
Start at the destination and follow prev backward
Why: prev[E] is D, since E's final estimate of 10 came from relaxing the edge D to E.
\[ E \leftarrow D \]
Keep following prev pointers
Why: prev[D] is B, and prev[B] is C, chaining the path backward one hop at a time.
\[ E \leftarrow D \leftarrow B \leftarrow C \]
Reach the source and reverse the chain
Why: prev[C] is S, the source, so the chain stops there. Reversing it gives the path in forward order.
\[ S \rightarrow C \rightarrow B \rightarrow D \rightarrow E \]
Verify the reconstructed path costs exactly 10
Why: Add up the edge weights along this exact path: S to C is 2, C to B is 1, B to D is 5, D to E is 2. The total matches dist[E].
\[ 2 + 1 + 5 + 2 = 10\ \checkmark \]
Intuition
What feels wrong about this?
Dijkstra settles a vertex and promises never to look at it again.
\[ \text{settled} \;\Rightarrow\; \text{final} \]
Now add an edge of weight minus 5 somewhere further out in the graph.
_Plain English only. No notation, no algebra. Just say what bothers you._
The feeling: a promise made early can be broken by something discovered later. Nothing stops a bargain from turning up after you already stopped looking.
That feeling is the proof. It is not a substitute for the proof — it is the thing the proof writes down.
Look back at the proof: the step that fails is the inequality claiming a partial path is never longer than the whole path. With a negative edge, going further can cost less, and that single line collapses.
Trap
A student runs Dijkstra as usual on a graph with a negative edge, and trusts the result without question.
Figure (svg): A directed graph with source S, edge S to A weight 4, edge S to B weight 1, and edge A to B weight negative 10
Settle S, then settle B before A
Why: Relaxing from S gives dist[A]=4 and dist[B]=1. Since 1 is smaller than 4, Dijkstra settles B first and locks its distance in as final.
\[ dist[B] = 1 \ (\text{settled, locked in}) \]
Settle A and relax A to B anyway — too late
Why: Relaxing A to B gives candidate 4 + (-10) = -6, far cheaper than 1. But B is already settled, so a correct Dijkstra implementation never revisits it. The final answer stays wrong.
\[ dist[B] \text{ reported as } 1 \quad (\text{true shortest is } -6) \]
The real cheapest path is S to A to B, costing 4 + (-10) = -6, which beats the direct edge's cost of 1. Dijkstra can never discover this, because it never reopens a settled vertex.
\[ \min(1,\ 4 + (-10)) = \min(1, -6) = -6 \]
Recognize why the greedy proof breaks
Why: The safety argument relied on 'continuing past an unsettled vertex can only add cost'. A negative edge violates that directly: continuing past A actually subtracted 10 from the running total.
\[ dist[A] + w(A,B) = 4 + (-10) = -6 < dist[A] \]
Use Bellman-Ford instead
Why: Bellman-Ford never assumes a vertex is 'done' until all V-1 rounds finish, so it correctly finds -6. We'll trace this exact graph with Bellman-Ford later in this lesson.
Break the constraint
Discussion prompt
The rule this trap just fixed:
The safety argument relied on 'continuing past an unsettled vertex can only add cost'. A negative edge violates that directly: continuing past A actually subtracted 10 from the running total.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
Relaxing from S gives dist[A]=4 and dist[B]=1. Since 1 is smaller than 4, Dijkstra settles B first and locks its distance in as final.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Recall the trace on G1: B first got a tentative distance of 4 (from relaxing S to B), then improved to 3 (from relaxing C to B). Both entries, 4 and 3, ended up sitting in the priority queue at the same time.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: A buggy implementation forgets to check whether a popped vertex is already settled.
A correct implementation checks, immediately after popping a vertex from the priority queue, whether it is already settled — and if so, discards the entry and moves on.
Why: A buggy implementation forgets to check whether a popped vertex is already settled. After D is settled at 8, the leftover entry {B, 4} is popped next, since 4 is less than E's 10 — and the buggy code treats this as settling B all over again.
Trap
Recall the trace on G1: B first got a tentative distance of 4 (from relaxing S to B), then improved to 3 (from relaxing C to B). Both entries, 4 and 3, ended up sitting in the priority queue at the same time.
\[ \text{queue holds two entries for B: } 4 \text{ and } 3 \]
Pop the stale entry and 're-settle' B
Why: A buggy implementation forgets to check whether a popped vertex is already settled. After D is settled at 8, the leftover entry {B, 4} is popped next, since 4 is less than E's 10 — and the buggy code treats this as settling B all over again.
\[ dist[B] \text{ overwritten: } 3 \rightarrow 4 \quad (\text{wrong — B was already correctly settled at 3}) \]
The wrong distance and wrong prev pointer stick
Why: Because the code re-processed a stale, worse entry, it silently corrupts both dist[B] and prev[B] after they were already correct.
A correct implementation checks, immediately after popping a vertex from the priority queue, whether it is already settled — and if so, discards the entry and moves on.
\[ \text{queue holds two entries for B: } 4 \text{ and } 3 \]
Pop the stale entry and discard it
Why: When {B, 4} is popped later, the algorithm sees B is already settled (with the correct value 3) and simply skips this leftover entry, doing no work.
\[ \text{B already settled at } 3 \Rightarrow \text{discard the stale entry} \]
Never touch a settled vertex again
Why: This 'settled means permanently done' rule is exactly what the safety argument for non-negative weights depends on. Skipping stale entries costs nothing but a discarded check.
Concept
What if two unsettled vertices have the exact same distance estimate? Either one may be settled first — the final dist values come out identical either way.
This is safe because of the same non-negative-weight argument from before: whichever of the tied vertices you settle first, its distance still cannot be improved later, since every other unsettled vertex's estimate is at least as large.
Socratic
Discussion prompt
What if two unsettled vertices have the exact same distance estimate? Either one may be settled first — the final dist values come out identical either way.
Suppose that were not true. What is the first thing in Shortest Paths: Dijkstra & Bellman-Ford that would stop working?
Hint: Follow it one step downstream. The answer is whatever was quietly relying on it.
Concept
With a binary heap as the priority queue, Dijkstra runs in:
\[ O\big((V+E)\log V\big) \]
V is the number of vertices, E is the number of edges. This is fast enough to be practical even on large road networks with millions of intersections.
Concept
Each vertex is settled exactly once, which means exactly V extract-min operations over the whole run — each one costing logarithmic time in a binary heap.
\[ V \text{ extract-min calls}, \quad O(\log V) \text{ each} \]
Every edge is relaxed at most once (when its tail is settled), and each successful relaxation can trigger one decrease-key or insert, also logarithmic time.
\[ E \text{ relaxations}, \quad O(\log V) \text{ each} \]
Adding these two contributions together gives the total:
\[ O(V \log V) + O(E \log V) = O\big((V+E)\log V\big) \]
Pattern
1. Initialize dist[source]=0, everyone else infinity; all unsettled
Why: Standard relaxation setup, with a priority queue keyed on dist.
2. Repeatedly pop the smallest unsettled vertex from the queue
Why: If it's already settled (a stale entry), discard it and pop again.
3. Mark it settled and relax every outgoing edge
Why: This may improve neighbors' estimates and push new queue entries.
4. Stop when every vertex is settled
Why: Non-negative weights guarantee each settled distance is already final — the safety argument from earlier in this section.
Picture it
Animation
Shows: Dijkstra on the very same graph — a rendered Manim animation.
Rendered with Manim.
Takeaway: Different parents from Prim: Dijkstra minimises distance FROM A, not edge weight.
Elimination
Eliminate the wrong options
Why is it safe to treat this vertex's distance as permanently final right now?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Every unsettled vertex already has an estimate at least as large as the one being settled. With non-negative weights, any path detouring through an unsettled vertex can only get more expensive from there, so no cheaper route to the settled vertex can ever appear later.
Check
Dijkstra just popped the unsettled vertex with the smallest current distance estimate and is about to mark it settled.
Check your understanding
Why is it safe to treat this vertex's distance as permanently final right now?
Answer: A
Why: Every unsettled vertex already has an estimate at least as large as the one being settled. With non-negative weights, any path detouring through an unsettled vertex can only get more expensive from there, so no cheaper route to the settled vertex can ever appear later.
Prediction
Predict first
What is the shortest path from S to D, and what does dist[D] represent?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: The path is S to C to D, and 9 is the total weight of that exact path
Why: Following prev backward from D gives D, then C (since prev[D]=C), then S (since prev[C]=S). Reversing that chain gives the forward path S to C to D, and dist[D]=9 is the total summed weight of every edge along that exact path.
Check
In a Dijkstra run from source S, you're given: prev[D] = C, prev[C] = S, and dist[D] = 9.
Check your understanding
What is the shortest path from S to D, and what does dist[D] represent?
Answer: A
Why: Following prev backward from D gives D, then C (since prev[D]=C), then S (since prev[C]=S). Reversing that chain gives the forward path S to C to D, and dist[D]=9 is the total summed weight of every edge along that exact path.
Section
Section 3
Concept
Dijkstra is fast, but the negative-weight trap showed its core safety argument depends on every edge weight being non-negative. We need a different algorithm for graphs where that isn't true.
Bellman-Ford trades some speed for generality: it correctly handles negative edge weights, and can even detect when a graph has no correct shortest-path answer at all.
Intuition
Instead of greedily trusting the closest vertex right away, Bellman-Ford just relaxes every edge in the graph, over and over, for a fixed number of rounds — patiently letting good news propagate outward from the source.
No vertex is ever declared 'done' early. Everyone's estimate stays open to improvement until the fixed number of rounds is complete.
Concept
Bellman-Ford's algorithm, in full: initialize distances as usual, then relax every edge in the graph, one full pass. Repeat this full pass V-1 times in total, where V is the number of vertices.
\[ \text{repeat } (V-1) \text{ times: for every edge } (u,v),\ \text{relax it} \]
That's the entire algorithm. No priority queue, no notion of 'settled' — just brute, repeated relaxation of the whole edge list.
Concept
No priority queue, no cleverness about order. Relax every edge, then do it again, and repeat one time fewer than there are vertices.
BELLMAN-FORD(G, s)
for each u in V
dist[u] = INFINITY
dist[s] = 0
repeat V - 1 times
for each edge (u, v) in E
if dist[u] + w(u, v) < dist[v]
dist[v] = dist[u] + w(u, v)
parent[v] = u
for each edge (u, v) in E
if dist[u] + w(u, v) < dist[v]
report a negative-weight cycleLines 7 and 8 are the identical relaxation Dijkstra uses. The difference is the total absence of a settled set: because nothing is ever final until the end, a negative edge cannot invalidate an earlier decision.
Notation
Every line of BELLMAN-FORD says one thing. Read the line, then read what it does — not the other way round.
Annotate
Invariant
After round k, every vertex reachable by a shortest path of at most k edges holds its true distance. Round by round the correct answers spread one edge further out.
Step through it
After each round, say which vertices you can now be sure about.
Picture it
Animation
Shows: BELLMAN-FORD executing: the current line of pseudocode is highlighted while the data it touches changes.
Rendered with Manim.
Takeaway: Relax every edge V minus one times; one extra round that still improves anything proves a negative cycle.
Sorting
Sort into buckets
These are the pieces of Shortest Paths: Dijkstra & Bellman-Ford, out of order. Put each one back under the part of the lesson it belongs to.
Concept
simple path — A path that never repeats a vertex. Since it can visit at most V distinct vertices, a simple path has at most V-1 edges.
\[ \text{a simple path uses at most } V-1 \text{ edges} \]
As long as there is no negative-weight cycle, the true shortest path between any two vertices is always a simple path — repeating a vertex could only add cost by going around a loop, never save any.
Matching
Match the pairs
Match each term to the definition this lesson gave it — not the one you would guess from the word.
Why: These are the working definitions of weight, settled, priority queue, simple path as Shortest Paths: Dijkstra & Bellman-Ford uses them. Pairing them correctly is the test of whether you could state each one with the slide switched off.
Intuition
What move should we make next?
Bellman-Ford relaxes every edge, V minus 1 times.
\[ \text{a shortest simple path has at most } |V| - 1 \text{ edges} \]
That fact is on the table. Turning it into V minus 1 rounds are enough still takes an argument.
What is the claim you would prove by induction, and what does round k buy you? Say the claim precisely before naming the move.
_Look at your toolkit. Say a move number out loud before this slide advances._ A wrong guess is useful. A silent guess is not.
Concept
Here is the key fact: after k full rounds of relaxing every edge, dist is guaranteed correct for every vertex whose true shortest path uses at most k edges.
\[ \text{after round } k: \text{ correct for all shortest paths with} \le k \text{ edges} \]
Since the true shortest path to any vertex is simple (previous slide), it uses at most V-1 edges. So after V-1 rounds, every vertex's shortest distance is guaranteed correct.
\[ \text{longest possible shortest path} = V-1 \text{ edges} \ \Rightarrow\ V-1 \text{ rounds suffice} \]
Explain it to yourself
Discussion prompt
In One round buys one edge this move is made:
Peel the last edge off a shortest path
Why is that legal? Name the rule or definition it rests on before you read on.
Hint: If you can only say "because that is what you do", the rule is the thing to go and find.
Answer:
A shortest path to v using k+1 edges is a shortest path to some u using k edges, plus one edge. By the hypothesis u was correct after round k, and round k+1 relaxes the edge from u to v.
Concept
The move: #5 (Peel one off), then #6 (Substitute the hypothesis).
Index the claim by number of edges, not by round
Why: Claim: after k rounds, every vertex whose shortest path uses at most k edges has its correct distance. That is the statement with a number in it.
Base case
Why: After 0 rounds the source has distance 0, which is the correct distance for the 0-edge path.
Peel the last edge off a shortest path
Why: A shortest path to v using k+1 edges is a shortest path to some u using k edges, plus one edge. By the hypothesis u was correct after round k, and round k+1 relaxes the edge from u to v.
\[ dist[v] \le dist[u] + w(u, v) = \delta(s, u) + w(u, v) = \delta(s, v) \]
Since no simple path exceeds V minus 1 edges, V minus 1 rounds cover every vertex. The relaxation order inside a round does not matter — the induction never referred to it.
Intuition
Picture a relay race with V runners standing at V different vertices. The longest possible baton-passing chain, visiting every runner exactly once, has V-1 handoffs — one fewer than the number of runners.
Each round of Bellman-Ford is like giving the baton one more chance to be passed along. After V-1 rounds, even the longest possible relay chain has had enough chances to complete.
Ranking
Put in order
Put the moves of Bellman-Ford trace on the example graph: rounds 1 and 2 into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. D-E and C-E and C-D and B-D all fail, since D, C, or B are still infinity when reached in this order.
Worked example
Trace Bellman-Ford on graph G1 (S, B, C, D, E) from source S. Process the edges in this order each round: D-E, C-E, C-D, B-D, C-B, S-C, S-B.
| round | dist[S] | dist[C] | dist[B] | dist[D] | dist[E] |
|---|---|---|---|---|---|
| 0 (init) | 0 | ∞ | ∞ | ∞ | ∞ |
Round 1: relax every edge once, in order
Why: D-E and C-E and C-D and B-D all fail, since D, C, or B are still infinity when reached in this order. S-C succeeds (0+2=2 < ∞) and S-B succeeds (0+4=4 < ∞).
| round | dist[S] | dist[C] | dist[B] | dist[D] | dist[E] |
|---|---|---|---|---|---|
| 1 | 0 | 2 | 4 | ∞ | ∞ |
Round 2: relax every edge again, using this round's improving values
Why: C-E gives 2+10=12 (< ∞, update). C-D gives 2+8=10 (< ∞, update). B-D gives 4+5=9, which beats the 10 just set, so D improves again to 9. C-B gives 2+1=3, beating B's current 4.
| round | dist[S] | dist[C] | dist[B] | dist[D] | dist[E] |
|---|---|---|---|---|---|
| 2 | 0 | 2 | 3 | 9 | 12 |
Check the estimates after two rounds
Why: D and E are still improving (9 and 12 are not yet their final values), which makes sense — their true shortest paths use more than two edges, so they need more rounds to become correct.
Picture it
Animation
Shows: Bellman-Ford relaxes everything, repeatedly — a rendered Manim animation.
Rendered with Manim.
Takeaway: Slower, and it survives negative weights.
Step zero
Discussion prompt
Bellman-Ford trace on the example graph: rounds 3 and 4 — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Round 3: relax every edge again
Answer:
Worked example
Continuing from round 2: dist[S]=0, dist[C]=2, dist[B]=3, dist[D]=9, dist[E]=12. This graph has 5 vertices, so V-1 = 4 rounds total are required.
Round 3: relax every edge again
Why: D-E gives 9+2=11, beating the current 12. B-D gives 3+5=8, beating the current 9. The other edges no longer find any improvement this round.
| round | dist[S] | dist[C] | dist[B] | dist[D] | dist[E] |
|---|---|---|---|---|---|
| 3 | 0 | 2 | 3 | 8 | 11 |
Round 4: relax every edge one last time
Why: D-E gives 8+2=10, beating the current 11. No other edge finds any further improvement.
| round | dist[S] | dist[C] | dist[B] | dist[D] | dist[E] |
|---|---|---|---|---|---|
| 4 | 0 | 2 | 3 | 8 | 10 |
Verify Bellman-Ford's final distances match Dijkstra's
Why: S=0, C=2, B=3, D=8, E=10 — exactly the distances Dijkstra found earlier on the same graph. Two very different methods, same correct answer.
\[ dist[S]=0,\ dist[C]=2,\ dist[B]=3,\ dist[D]=8,\ dist[E]=10\ \checkmark \]
Worked example
To see exactly why V-1 rounds are needed, trace Bellman-Ford on a 5-vertex chain, each edge weight 1, source V1. Process edges in this order each round: V4-V5, V3-V4, V2-V3, V1-V2 — deliberately the reverse of the natural direction.
Figure (svg): A chain graph with five vertices V1 through V5 connected in a line, each edge weight 1
Rounds 1 and 2: progress crawls forward one hop at a time
Why: In round 1, only V1-V2 can fire (everything else still involves an infinity), so only V2 improves. In round 2, only V2-V3 can now fire, so V3 improves — but V4 and V5 are still infinity.
| round | dist[V1] | dist[V2] | dist[V3] | dist[V4] | dist[V5] |
|---|---|---|---|---|---|
| 1 | 0 | 1 | ∞ | ∞ | ∞ |
| 2 | 0 | 1 | 2 | ∞ | ∞ |
Rounds 3 and 4: the last two hops finally complete
Why: Round 3 lets V3-V4 fire, reaching V4. Only in round 4 does V4-V5 finally fire, reaching V5 for the first time.
| round | dist[V1] | dist[V2] | dist[V3] | dist[V4] | dist[V5] |
|---|---|---|---|---|---|
| 3 | 0 | 1 | 2 | 3 | ∞ |
| 4 | 0 | 1 | 2 | 3 | 4 |
Verify that V5's distance only finalizes in round four
Why: This 5-vertex chain has a longest simple path of exactly 4 edges (V1 to V5), and with this adversarial edge order, the algorithm genuinely needs all 4 = V-1 rounds — not one fewer — to reach V5 at all.
\[ V = 5 \ \Rightarrow\ V - 1 = 4 \text{ rounds needed}\ \checkmark \]
Notation
Annotate
From Watching Bellman-Ford propagate down a chain — read this one piece at a time. What is each part doing?
On: \( V = 5 \ \Rightarrow\ V - 1 = 4 \text{ rounds needed}\ \checkmark \)
Anomaly
Predict first
A student writes this, and it looks reasonable:
On the chain graph, a student runs only 2 rounds, notices dist[V4] and dist[V5] haven't changed in a while, and assumes the algorithm has converged.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: This student reports dist[V4] and dist[V5] as infinity — unreachable — which is completely wrong.
Always run the full V-1 rounds, regardless of how things look partway through. 'Nothing changed recently' does not mean 'nothing will change'.
Why: This student reports dist[V4] and dist[V5] as infinity — unreachable — which is completely wrong. Both are reachable, at distances 3 and 4 respectively.
Trap
On the chain graph, a student runs only 2 rounds, notices dist[V4] and dist[V5] haven't changed in a while, and assumes the algorithm has converged.
\[ \text{after round 2: } dist[V4] = \infty, \ dist[V5] = \infty \]
Report the distances after only 2 rounds
Why: This student reports dist[V4] and dist[V5] as infinity — unreachable — which is completely wrong. Both are reachable, at distances 3 and 4 respectively.
\[ \text{reported: unreachable} \quad (\text{actual: } dist[V4]=3,\ dist[V5]=4) \]
Always run the full V-1 rounds, regardless of how things look partway through. 'Nothing changed recently' does not mean 'nothing will change'.
\[ V = 5 \ \Rightarrow\ \text{run all } V - 1 = 4 \text{ rounds} \]
Run rounds 3 and 4 too
Why: As traced earlier, V4 only becomes reachable in round 3, and V5 only in round 4. Stopping at round 2 misses both, purely because of how the edges happened to be ordered.
\[ dist[V4]=3 \ (\text{round 3}), \quad dist[V5]=4 \ (\text{round 4}) \]
The V-1 bound is about the worst case, not the typical case
Why: Some edge orders converge faster, as seen with the earlier lucky ordering on G1. But the guarantee that ALL graphs are correctly solved only holds if you always run the full V-1 rounds.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
dist. dist of a vertex is the total weight of the cheapest path found so far from the source to it.Ranking
Put in order
Put the moves of Bellman-Ford succeeds where Dijkstra failed into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. S-A gives dist[A]=0+4=4. S-B gives dist[B]=0+1=1.
Worked example
Return to the negative-weight graph from the Dijkstra trap: S, A, B, with S to A weight 4, S to B weight 1, and A to B weight -10. This graph has 3 vertices, so V-1 = 2 rounds.
Round 1: relax S-A, S-B, A-B in order
Why: S-A gives dist[A]=0+4=4. S-B gives dist[B]=0+1=1. A-B then gives candidate 4+(-10)=-6, which beats the just-set 1 — dist[B] improves to -6, all within round 1.
| round | dist[S] | dist[A] | dist[B] |
|---|---|---|---|
| 1 | 0 | 4 | -6 |
Round 2: relax every edge again
Why: S-A gives 4, no change. S-B gives 1, which is not less than -6, no change. A-B gives 4+(-10)=-6, not less than -6, no change. Nothing improves — the values are stable.
| round | dist[S] | dist[A] | dist[B] |
|---|---|---|---|
| 2 | 0 | 4 | -6 |
Verify the true shortest distance to B is -6
Why: The two possible routes to B cost 1 (direct) and 4+(-10)=-6 (through A). The smaller of the two, -6, is what Bellman-Ford correctly reports — unlike Dijkstra, which incorrectly reported 1.
\[ \min(1,\ 4+(-10)) = \min(1,-6) = -6\ \checkmark \]
Picture it
Animation
Shows: The non-negativity requirement, precisely — a rendered Manim animation.
Rendered with Manim.
Takeaway: It is the finality that breaks, not the arithmetic.
Intuition
If Bellman-Ford handles more cases correctly, why not use it everywhere? Because it pays for that generality: it relaxes every edge V-1 times, regardless of how the graph is shaped, while Dijkstra homes in on the answer using a priority queue.
When you know every weight is non-negative — routes, flight prices, most real distance maps — Dijkstra is the faster tool for the job. Save Bellman-Ford for when negative weights are possible, or when you need to check for negative cycles.
Concept
Bellman-Ford relaxes every one of the E edges, once per round, for V-1 rounds:
\[ O(V \cdot E) \]
Compare this to Dijkstra's O((V+E) log V). On a dense graph with many edges, Bellman-Ford is noticeably slower — the price paid for correctly handling negative weights.
Explain it to yourself
Discussion prompt
In The Bellman-Ford recipe this move is made:
3. Never stop early, even if nothing seems to be changing
Why is that legal? Name the rule or definition it rests on before you read on.
Hint: If you can only say "because that is what you do", the rule is the thing to go and find.
Answer:
The V-1 bound is a worst-case guarantee; some edge orders converge sooner, but only running the full count guarantees correctness on every graph.
Pattern
1. Initialize dist[source]=0, everyone else infinity
Why: Same setup as Dijkstra — this part of relaxation never changes.
2. Repeat V-1 times: relax every edge in the graph, once each
Why: No priority queue, no 'settled' status — just brute-force repetition over the full edge list.
3. Never stop early, even if nothing seems to be changing
Why: The V-1 bound is a worst-case guarantee; some edge orders converge sooner, but only running the full count guarantees correctness on every graph.
4. Optionally, run one more round to check for a negative cycle
Why: Covered next: if any edge still improves after V-1 rounds, the graph has a negative-weight cycle reachable from the source.
Picture it
Animation
Shows: Stop as soon as nothing relaxes — a rendered Manim animation.
Rendered with Manim.
Takeaway: The bound is worst-case; the loop rarely needs all of it.
Check
You are running Bellman-Ford on a graph with 6 vertices.
Check your understanding
What is the minimum number of rounds that guarantees every dist value is correct, on any graph with 6 vertices and no negative cycle?
Answer: A
Why: The number of required rounds is V-1, since the longest possible simple path in a 6-vertex graph has at most 5 edges. With 6 vertices, V-1 = 5 rounds are guaranteed sufficient.
Section
Section 4
Concept
A negative-weight cycle is a closed loop of edges whose total weight is negative. If such a cycle is reachable from the source, 'shortest path' stops making sense for any vertex reachable through it.
You could always go around the loop one more time and lower your total cost further. There is no minimum — the true 'shortest distance' is unbounded below.
\[ \text{go around the cycle again} \Rightarrow \text{cost keeps dropping, forever} \]
Concept
A single negative edge, by itself, is not a problem for Bellman-Ford — the earlier worked example proved that directly, correctly finding a distance of -6.
It is only a cycle whose total weight is negative that breaks things. A graph can have many negative edges and still have a perfectly well-defined shortest path for every vertex, as long as none of those edges form a negative-weight loop.
Picture it
Animation
Shows: Why Dijkstra breaks on negative edges — a rendered Manim animation.
Rendered with Manim.
Takeaway: The greedy commitment is exactly what negative weights invalidate.
Explain it to yourself
Discussion prompt
In Process: detecting a negative cycle the direct way this move is made:
Back up. Use the theorem you just proved instead of the definition
Why is that legal? Name the rule or definition it rests on before you read on.
Hint: If you can only say "because that is what you do", the rule is the thing to go and find.
Answer:
V minus 1 rounds settle everything if the distances are well defined. So run one extra round: if anything still improves, no finite shortest path exists, which means a negative cycle is reachable.
Intuition
Watch me not know the answer. This is what the first two minutes actually look like.
We want to report whether the graph contains a negative-weight cycle reachable from the source.
Try enumerating cycles and adding up their weights
Why: It is the definition, so it is guaranteed correct. Find every cycle, total its weights, check for a negative one.
There can be exponentially many cycles
Why: A graph with many parallel routes has a number of cycles that grows exponentially in the vertex count. Correct and unusable.
Dead end. Not a mistake — a move that was worth trying and did not pay off. This happens in most proofs.
Back up. Use the theorem you just proved instead of the definition
Why: V minus 1 rounds settle everything if the distances are well defined. So run one extra round: if anything still improves, no finite shortest path exists, which means a negative cycle is reachable.
One extra pass over the edges, instead of an exponential enumeration. The detection test is a corollary of the correctness proof, which is why the proof was worth doing carefully.
The expert does not see the whole path in advance. The expert tries something, reads the result, and adjusts. That is the skill.
Concept
After the normal V-1 rounds finish, run one extra round of relaxing every edge.
If any edge still successfully relaxes — still finds an improvement — during that extra round, the graph contains a negative-weight cycle reachable from the source.
\[ \text{after } V-1 \text{ rounds: any edge still relaxes} \Rightarrow \text{negative cycle exists} \]
Why this works: V-1 rounds are enough for every SIMPLE path. If a distance keeps shrinking past that point, the only way that's possible is by looping through a cycle whose total weight is negative.
Explain it
Discussion prompt
Explain Detecting a negative cycle with one more round to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
After the normal V-1 rounds finish, run one extra round of relaxing every edge.
Intuition
One more full pass over every edge is a small, fixed amount of extra work compared to the V-1 rounds you already ran — it does not change the algorithm's overall running time.
For that tiny cost, you get a guarantee: either nothing improves, and your distances are trustworthy, or something improves, and you know for certain not to trust any distance touched by the cycle.
Analogy
Discussion prompt
Explain The extra round costs almost nothing by analogy to something with no CS3000 Algorithms in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
One more full pass over every edge is a small, fixed amount of extra work compared to the V-1 rounds you already ran — it does not change the algorithm's overall running time.
Picture it
Figure (svg): A directed triangle graph with X, Y, Z. Edge X to Y weight 1, edge Y to Z weight negative 3, edge Z to X weight 1, forming a negative-weight cycle
Discussion prompt
Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.
Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.
Answer:
Vertices X, Y, Z. Source X. Edges: X to Y weight 1, Y to Z weight -3, Z to X weight 1. The cycle X to Y to Z to X totals 1 + (-3) + 1 = -1, a negative-weight cycle.
Worked example
Vertices X, Y, Z. Source X. Edges: X to Y weight 1, Y to Z weight -3, Z to X weight 1. The cycle X to Y to Z to X totals 1 + (-3) + 1 = -1, a negative-weight cycle.
Figure (svg): A directed triangle graph with X, Y, Z. Edge X to Y weight 1, edge Y to Z weight negative 3, edge Z to X weight 1, forming a negative-weight cycle
Rounds 1 and 2 (V-1 = 2, since V = 3)
Why: Relaxing X-Y, Y-Z, Z-X in order, twice: round 1 gives X=-1, Y=1, Z=-2 (the Z-X relaxation even improves X itself, since -2+1=-1 beats 0). Round 2 gives X=-2, Y=0, Z=-3.
| round | dist[X] | dist[Y] | dist[Z] |
|---|---|---|---|
| 1 | -1 | 1 | -2 |
| 2 | -2 | 0 | -3 |
Round 3 (the extra detection round)
Why: Relax X-Y again: candidate is -2+1=-1, which is less than the current dist[Y]=0. An edge relaxed successfully after the normal V-1=2 rounds finished.
\[ dist[X] + w(X,Y) = -2 + 1 = -1\ <\ 0 = dist[Y] \]
Verify the extra round still finds an improvement
Why: Since round 3 (beyond the required 2 rounds) still relaxed an edge successfully, this graph contains a negative-weight cycle reachable from the source X — exactly matching the cycle we identified up front, with total weight -1.
\[ \text{round beyond } V{-}1 \text{ still improves} \Rightarrow \text{negative cycle confirmed}\ \checkmark \]
Comparison
Comparison matrix
From Bellman-Ford catches a negative cycle: refill the dist[Y] column from what you know. The rest of the table is as it appeared.
| round | dist[X] | dist[Y] | dist[Z] |
|---|---|---|---|
| 1 | -1 | 1 | -2 |
| 2 | -2 | 0 | -3 |
Picture it
Animation
Shows: One extra pass detects a negative cycle — a rendered Manim animation.
Rendered with Manim.
Takeaway: With a negative cycle there is no shortest path at all — you can always go round again.
Concept
The two algorithms solve overlapping but different problems, at different costs.
| Algorithm | Handles negative edges? | Detects negative cycles? | Running time |
|---|---|---|---|
| Dijkstra | No | No | O((V+E) log V) |
| Bellman-Ford | Yes | Yes | O(V · E) |
Dijkstra's speed comes precisely from the assumption it cannot safely give up: non-negative weights. Bellman-Ford gives up that speed to stay correct without it.
Comparison
Comparison matrix
From Running times side by side: refill the Running time column from what you know. The rest of the table is as it appeared.
| Algorithm | Handles negative edges? | Detects negative cycles? | Running time |
|---|---|---|---|
| Dijkstra | No | No | O((V+E) log V) |
| Bellman-Ford | Yes | Yes | O(V · E) |
Explain it to yourself
Discussion prompt
In Choosing between Dijkstra and Bellman-Ford this move is made:
4. If you also need to detect a negative cycle, run one extra Bellman-Ford round
Why is that legal? Name the rule or definition it rests on before you read on.
Hint: If you can only say "because that is what you do", the rule is the thing to go and find.
Answer:
Any edge that still relaxes after the standard V-1 rounds proves a negative cycle is reachable from the source.
Pattern
1. Check whether every edge weight is non-negative
Why: This single check decides everything that follows.
2. If yes, use Dijkstra with a priority queue
Why: Faster, at O((V+E) log V), and always correct when weights are non-negative.
3. If negative weights are possible, use Bellman-Ford
Why: Relax every edge V-1 times; slower at O(V · E), but correct even with negative edges.
4. If you also need to detect a negative cycle, run one extra Bellman-Ford round
Why: Any edge that still relaxes after the standard V-1 rounds proves a negative cycle is reachable from the source.
Real world
Discussion prompt
Outside this lesson: where does Shortest Paths: Dijkstra & Bellman-Ford actually turn up? Name one concrete situation — a job, a piece of software someone ships, a decision somebody has to make — and say which part of Choosing between Dijkstra and Bellman-Ford is doing the work in it.
Hint: Vague is the failure mode here. "Engineering" is not a situation; "deciding whether this build is fast enough to ship" is.
Answer:
That deck builds single-source shortest paths from one shared primitive, edge relaxation. It traces Dijkstra's algorithm by hand with a priority queue and explains why its greedy choice is safe only for non-negative weights, then covers Bellman-Ford's V-1 rounds of relaxation and its negative-cycle detection. It targets the traps of trusting Dijkstra with a negative edge, relaxing in the wrong direction, misjudging why V-1 rounds are needed, and reviving a vertex that has already been settled.
Commit first
Predict first
Can Bellman-Ford still compute correct shortest distances for this graph?
Commit to an answer, then rate it — certain, fairly sure, or guessing — and write the rating down before you turn the page.
Correct: Yes — a single negative edge is fine as long as it isn't part of a negative-weight cycle
Why: As shown in the worked example where Bellman-Ford correctly found a distance of -6, negative edges by themselves are not a problem. Only a cycle whose total weight is negative breaks the notion of a shortest path, since it lets you lower your cost forever by looping.
The rating matters as much as the answer: confident-and-wrong is the combination that survives revision, because nothing about it feels like it needs revisiting.
Check
A graph has exactly one negative-weight edge, and that edge is not part of any cycle at all.
Check your understanding
Can Bellman-Ford still compute correct shortest distances for this graph?
Answer: A
Why: As shown in the worked example where Bellman-Ford correctly found a distance of -6, negative edges by themselves are not a problem. Only a cycle whose total weight is negative breaks the notion of a shortest path, since it lets you lower your cost forever by looping.
Prediction
Predict first
During that extra pass, one edge still successfully relaxes, lowering some vertex's distance further. What does this tell you?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: The graph contains a negative-weight cycle reachable from the source
Why: V-1 rounds are proven sufficient for every simple path, the longest kind a true shortest path can be. If a distance can still shrink after that many rounds, the only explanation is a negative-weight cycle being looped through again and again.
Check
You run Bellman-Ford for the standard V-1 rounds, then perform one additional relaxation pass over every edge.
Check your understanding
During that extra pass, one edge still successfully relaxes, lowering some vertex's distance further. What does this tell you?
Answer: A
Why: V-1 rounds are proven sufficient for every simple path, the longest kind a true shortest path can be. If a distance can still shrink after that many rounds, the only explanation is a negative-weight cycle being looped through again and again.
Elimination
Eliminate the wrong options
Which algorithm should you use, and why?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Since every travel time is non-negative, Dijkstra's greedy safety argument holds, so it will produce correct results at its faster running time of O((V+E) log V) versus Bellman-Ford's O(V times E).
Check
You are given a large road network where every road's travel time is a non-negative number, and you need the single fastest algorithm to find shortest travel times from one city to every other city.
Check your understanding
Which algorithm should you use, and why?
Answer: A
Why: Since every travel time is non-negative, Dijkstra's greedy safety argument holds, so it will produce correct results at its faster running time of O((V+E) log V) versus Bellman-Ford's O(V times E).
Concept
Moves added today:
Moves you reused today:
Move #15 is move #8 specialized to things that come in sequences. Instead of the smallest counterexample, you take the earliest point along a path or a run where the claim breaks.
Full toolkit so far: #1 through #15.
Next session opens with you naming every one of these from memory, before any new material.
Counterexample
Discussion prompt
Move #15 is move #8 specialized to things that come in sequences. Instead of the smallest counterexample, you take the earliest point along a path or a run where the claim breaks.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Next session opens with you naming every one of these from memory, before any new material.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — The Shortest-Path Problem · Dijkstra's Algorithm · The Bellman-Ford Algorithm · Negative Cycles & Choosing an Algorithm. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
Both algorithms in this lesson are built on the exact same move, repeated in different ways.
| Situation | Use |
|---|---|
| All weights non-negative, need speed | Dijkstra: O((V+E) log V) |
| Negative weights possible | Bellman-Ford: O(V · E) |
| Need to detect a negative cycle | Bellman-Ford, plus one extra round |
Watch for the traps: never trust Dijkstra with a negative edge, never relax backwards, never stop Bellman-Ford before V-1 rounds, and never let a settled Dijkstra vertex be touched again.
Want this taught 1-on-1? Alexander tutors CS3000 Algorithms — $55/session, free consultation.